Jingjing Fu

dblp:95/4837 · DBLP profile ↗
← Back
41ranked-venue papers
10as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 21 · 5 first-author · 6 since 2021Systems, architecture and hardware · 11 · 5 first-authorArtificial intelligence and machine learning · 9 · 8 since 2021Computer networks · 5 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Drim-NeRF: Diffusion-Based Restoration for Improving Neural Radiance Fields
abstract
The rendering degradations produced by Neural Radiance Field (NeRF) is a long-standing but complex issue in the field of 3D implicit representation, which arises from a multitude of intricate causes and was not entirely solved by designing complicated scene parameterization methods before. In this paper, we present a diffusion-based restoration method for improving Neural Radiance Field (Drim-NeRF). We consider the NeRF enhancement issue from a low-level restoration perspective by viewing all types of rendering artifacts as a specific degradation model added to clean ground truths. By leveraging the powerful prior knowledge encapsulated in diffusion model, we could restore the high-realism improved renderings conditioned on the raw low-quality rendering counterparts. To further ensure the multi-view consistent rendering enhancement, we innovatively propose to adopt optical flow warping to reduce temporal inconsistency and employ feature-wrapping in VAE decoder to improve fidelity. Our proposed method is easy to implement and agnostic to various NeRF backbones. We conduct extensive experiments on challenging large-scale urban scenes and unbounded 360-degree scenes, as well as other baselines and datasets and achieve substantial qualitative and quantitative improvements, both in the restoration quality and the multi-view consistency perspective.
Ganlin Yang, Kaidong Zhang, Jingjing Fu, Dong Liu 0002
IEEE Trans. Circuits Syst. Video Technol.3
2025 OMGM: Orchestrate Multiple Granularities and Modalities for Efficient Multimodal Retrieval
abstract
Vision-language retrieval-augmented generation (RAG) has become an effective approach for tackling Knowledge-Based Visual Question Answering (KB-VQA), which requires external knowledge beyond the visual content presented in images.The effectiveness of Vision-language RAG systems hinges on multimodal retrieval, which is inherently challenging due to the diverse modalities and knowledge granularities in both queries and knowledge bases.Existing methods have not fully tapped into the potential interplay between these elements.We propose a multimodal RAG system featuring a coarse-to-fine, multi-step retrieval that harmonizes multiple granularities and modalities to enhance efficacy.Our system begins with a broad initial search aligning knowledge granularity for cross-modal retrieval, followed by a multimodal fusion reranking to capture the nuanced multimodal information for top entity selection.A text reranker then filters out the most relevant fine-grained section for augmented generation.Extensive experiments on the InfoSeek and Encyclopedic-VQA benchmarks show our method achieves state-of-the-art retrieval performance and highly competitive answering results, underscoring its effectiveness in advancing KB-VQA systems.
Jingjing Fu, Rui Wang 0028, Lei Song 0001, Jiang Bian 0002
ACL (1)2
2025 From Complex to Atomic: Enhancing Augmented Generation via Knowledge-Aware Dual Rewriting and Reasoning
abstract
Recent advancements in Retrieval-Augmented Generation (RAG) systems have significantly enhanced the capabilities of large language models (LLMs) by incorporating external knowledge retrieval. However, the sole reliance on retrieval is often inadequate for mining deep, domain-specific knowledge and for performing logical reasoning from specialized datasets. To tackle these challenges, we present an approach, which is designed to extract, comprehend, and utilize domain knowledge while constructing a coherent rationale. At the heart of our approach lie four pivotal components: a knowledge atomizer that extracts atomic questions from raw data, a query proposer that generates subsequent questions to facilitate the original inquiry, an atomic retriever that locates knowledge based on atomic knowledge alignments, and an atomic selector that determines which follow-up questions to pose guided by the retrieved information. Through this approach, we implement a knowledge-aware task decomposition strategy that adeptly extracts multifaceted knowledge from segmented data and iteratively builds the rationale in alignment with the initial query and the acquired knowledge. We conduct comprehensive experiments to demonstrate the efficacy of our approach across various benchmarks, particularly those requiring multihop reasoning steps. The results indicate a significant enhancement in performance, up to 12.6% over the second-best method, underscoring the potential of the approach in complex, knowledge-intensive applications.
Jingjing Fu, Rui Wang 0028, Lei Song 0001, Jiang Bian 0002
ICML2
2025 Similarity-Guided Rapid Deployment of Federated Intelligence Over Heterogeneous Edge Computing
Hansong Zhou, Jingjing Fu, Yukun Yuan 0001, Linke Guo, Xiaonan Zhang 0001
INFOCOM2
2025 Prototypical Distribution Divergence Loss for Image Restoration
abstract
Neural networks have achieved significant advances in the field of image restoration and much research has focused on designing new architectures for convolutional neural networks (CNNs) and Transformers. The choice of loss functions, despite being a critical factor when training image restoration networks, has attracted little attention. The existing losses are primarily based on semantic or hand-crafted representations. Recently, discrete representations have demonstrated strong capabilities in representing images. In this work, we explore the loss of discrete representations for image restoration. Specifically, we propose a Local Residual Quantized Variational AutoEncoder (Local RQ-VAE) to learn prototype vectors that represent the local details of high-quality images. Then we propose a Prototypical Distribution Divergence (PDD) loss that measures the Kullback-Leibler divergence between the prototypical distributions of the restored and target images. Experimental results demonstrate that our PDD loss improves the restored images in both PSNR and visual quality for state-of-the-art CNNs and Transformers on several image restoration tasks, including image super-resolution, image denoising, image motion deblurring, and defocus deblurring.
Jialun Peng, Jingjing Fu, Dong Liu 0002
IEEE Trans. Image Process.2
2024 Confidence-Based Iterative Generation for Real-World Image Super-Resolution
Jialun Peng, Jingjing Fu, Dong Liu 0002
ECCV (65)3
2024 FreeEM: Uncovering Parallel Memory EMR Covert Communication in Volatile Environments
abstract
Memory Electromagnetic Radiation (EMR) allows attackers to manipulate the DRAM of infiltrated systems to leak sensitive secret information. Although most of the existing works have demonstrated its feasibility, practical concerns, such as the ideal electromagnetic environment and stationary attacking layout, make the covert channel attack less convincing, especially in vulnerable sites such as offices and data centers. This work removes the above impractical assumptions to uncover the potential of memory EMR by proposing the first parallel EMR covert communication protocol. Our design reshapes the current "1-to-1" covert communication mode to "n-to-1" mode via a novel pattern-based 2-dimensional symbol encoding scheme, allowing multiple victim computers to simultaneously perform data exfiltration to one attacker (the receiver) without mutual interference. Meanwhile, this novel scheme design also enables the very first mobile attacker, i.e., a smartphone connected to a software-defined radio (SDR) dongle, to capture parallel memory EMR signals in a volatile environment. Extensive experiments are conducted to verify the performance in a volatile environment with different parameter configurations, distances, motion modes, shielding materials, orientations, hardware configurations, and SDR platforms. Our experimental results demonstrate that FreeEM can support up to 4 parallel memory EMR transmissions to achieve an overall throughput of 625Kbps and a decoding accuracy of 96.88%. The maximum communication distance can reach up to 20 meters.
Sihan Yu, Jingjing Fu, Chenxu Jiang, ChunChih Lin, Zhenkai Zhang 0002, Long Cheng 0005, Ming Li 0006, Xiaonan Zhang 0001, Linke Guo
MobiSys2
2024 A Single-Step, Sharpness-Aware Minimization is All You Need to Achieve Efficient and Accurate Sparse Training
abstract
Sparse training stands as a landmark approach in addressing the considerable training resource demands imposed by the continuously expanding size of Deep Neural Networks (DNNs). However, the training of a sparse DNN encounters great challenges in achieving optimal generalization ability despite the efforts from the state-of-the-art sparse training methodologies. To unravel the mysterious reason behind the difficulty of sparse training, we connect the network sparsity with neural loss functions structure, and identify the cause of such difficulty lies in chaotic loss surface. In light of such revelation, we propose $S^{2} - SAM$, characterized by a **S**ingle-step **S**harpness_**A**ware **M**inimization that is tailored for **S**parse training. For the first time, $S^{2} - SAM$ innovates the traditional SAM-style optimization by approximating sharpness perturbation through prior gradient information, incurring *zero extra cost*. Therefore, $S^{2} - SAM$ not only exhibits the capacity to improve generalization but also aligns with the efficiency goal of sparse training. Additionally, we study the generalization result of $S^{2} - SAM$ and provide theoretical proof for convergence. Through extensive experiments, $S^{2} - SAM$ demonstrates its universally applicable plug-and-play functionality, enhancing accuracy across various sparse training methods. Code available at https://github.com/jjsrf/SSAM-NEURIPS2024.
Gen Li 0012, Jingjing Fu, Fatemeh Afghah, Linke Guo, Xiaoyong Yuan
NeurIPS3
2024 Behaviors Speak More: Achieving User Authentication Leveraging Facial Activities via mmWave Sensing
abstract
Human faces have been widely adopted in many applications and systems requiring a high-security standard. Although face authentication is deemed to be mature nowadays, many existing works have demonstrated not only the privacy leakage of facial information but also the success of spoofing attacks on face biometrics. The critical reason behind this is the failure of liveness detection in biometrics. This work advances most biometric-based user authentication schemes by exploring dynamic biometrics (human facial activities) rather than traditional static biometrics (human faces). Inspired by observations from psychology, we propose the mmFaceID to leverage humans' dynamic facial activities when performing word reading for achieving robust, highly accurate, and effective user authentication via mmWave sensing. By addressing a series of technical challenges of capturing micro-level facial muscle movements using a mmWave sensor, we build a neural network to reconstruct facial activities via estimated expression parameters. Then, unique features can be extracted to enable robust user authentication regardless of relative distances and orientations. We conduct comprehensive experiments on 23 participants to evaluate mmFaceID in terms of distances/orientations, length of word lists, occlusion, and language backgrounds, demonstrating an authentication accuracy of 94.7%. We also extend our evaluation in a real IoT scenario. By speaking real IoT commends, the average authentication accuracy can reach up to 92.28%.
Chenxu Jiang, Sihan Yu, Jingjing Fu, ChunChih Lin, Huadi Zhu, Ming Li 0006, Linke Guo
SenSys3
2024 Exploiting Optical Flow Guidance for Transformer-Based Video Inpainting
abstract
Transformers have been widely used for video processing owing to the multi-head self attention (MHSA) mechanism. However, the MHSA mechanism encounters an intrinsic difficulty for video inpainting, since the features associated with the corrupted regions are degraded and incur inaccurate self attention. This problem, termed query degradation, may be mitigated by first completing optical flows and then using the flows to guide the self attention, which was verified in our previous work - flow-guided transformer (FGT). We further exploit the flow guidance and propose FGT++ to pursue more effective and efficient video inpainting. First, we design a lightweight flow completion network by using local aggregation and edge loss. Second, to address the query degradation, we propose a flow guidance feature integration module, which uses the motion discrepancy to enhance the features, together with a flow-guided feature propagation module that warps the features according to the flows. Third, we decouple the transformer along the temporal and spatial dimensions, where flows are used to select the tokens through a temporally deformable MHSA mechanism, and global tokens are combined with the inner-window local tokens through a dual-perspective MHSA mechanism. FGT++ is experimentally evaluated to be outperforming the existing video inpainting networks qualitatively and quantitatively.
Kaidong Zhang, Jialun Peng, Jingjing Fu, Dong Liu 0002
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 Template-guided Hierarchical Feature Restoration for Anomaly Detection
abstract
Targeting for detecting anomalies of various sizes for complicated normal patterns, we propose a Template-guided Hierarchical Feature Restoration method, which introduces two key techniques, bottleneck compression and template-guided compensation, for anomaly-free feature restoration. Specially, our framework compresses hierarchical features of an image by bottleneck structure to preserve the most crucial features shared among normal samples. We design template-guided compensation to restore the distorted features towards anomaly-free features. Particularly, we choose the most similar normal sample as the template, and leverage hierarchical features from the template to compensate the distorted features. The bottleneck could partially filter out anomaly features, while the compensation further converts the reminding anomaly features towards normal with template guidance. Finally, anomalies are detected in terms of the cosine distance between the pre-trained features of an inference image and the corresponding restored anomaly-free features. Experimental results demonstrate the effectiveness of our approach, which achieves the state-of-the-art performance on the MVTec LOCO AD dataset.
Hewei Guo, Liping Ren, Jingjing Fu, Yuwang Wang, Zhizheng Zhang 0004, Cuiling Lan, Haoqian Wang, Xinwen Hou
ICCV3
2022 Inertia-Guided Flow Completion and Style Fusion for Video Inpainting
abstract
Physical objects have inertia, which resists changes in the velocity and motion direction. Inspired by this, we introduce inertia prior that optical flow, which reflects object motion in a local temporal window, keeps unchanged in the adjacent preceding or subsequent frame. We propose a flow completion network to align and aggregate flow features from the consecutive flow sequences based on the inertia prior. The corrupted flows are completed under the supervision of customized losses on reconstruction, flow smoothness, and consistent ternary census transform. The completed flows with high fidelity give rise to significant improvement on the video inpainting quality. Nevertheless, the existing flow-guided cross-frame warping methods fail to consider the lightening and sharpness variation across video frames, which leads to spatial incoherence after warping from other frames. To alleviate such problem, we propose the Adaptive Style Fusion Network (ASFN), which utilizes the style information extracted from the valid regions to guide the gradient refinement in the warped regions. Moreover, we design a data simulation pipeline to reduce the training difficulty of ASFN. Extensive experiments show the superiority of our method against the state-of-the-art methods quantitatively and qualitatively. The project page is at https://github.com/hitachinsk/ISVI.
Kaidong Zhang, Jingjing Fu, Dong Liu 0002
CVPR2
2022 Flow-Guided Transformer for Video Inpainting
Kaidong Zhang, Jingjing Fu, Dong Liu 0002
ECCV (18)2
2018 Feature Selective Networks for Object Detection
abstract
Objects for detection usually have distinct characteristics in different sub-regions and different aspect ratios. However, in prevalent two-stage object detection methods, Region-of-Interest (RoI) features are extracted by RoI pooling with little emphasis on these translation-variant feature components. We present feature selective networks to reform the feature representations of RoIs by exploiting their disparities among sub-regions and aspect ratios. Our network produces the sub-region attention bank and aspect ratio attention bank for the whole image. The RoI-based sub-region attention map and aspect ratio attention map are selectively pooled from the banks, and then used to refine the original RoI features for RoI classification. Equipped with a lightweight detection subnetwork, our network gets a consistent boost in detection performance based on general ConvNet backbones (ResNet-101, GoogLeNet and VGG-16). Without bells and whistles, our detectors equipped with ResNet-101 achieve more than 3% mAP improvement compared to counterparts on PASCAL VOC 2007, PASCAL VOC 2012 and MS COCO datasets.
Yao Zhai, Jingjing Fu, Yan Lu 0001, Houqiang Li
CVPR2
2016 The Improved Algorithm Based on DFS and BFS for Indoor Trajectory Reconstruction
Jingjing Fu, Zhujun Zhang, Siye Wang
WASA2
2016 A High-Fidelity and Low-Interaction-Delay Screen Sharing System
abstract
The pervasive computing environment and wide network bandwidth provide users more opportunities to share screen content among multiple devices. In this article, we introduce a remote display system to enable screen sharing among multiple devices with high fidelity and responsive interaction. In the developed system, the frame-level screen content is compressed and transmitted to the client side for screen sharing, and the instant control inputs are simultaneously transmitted to the server side for interaction. Even if the screen responds immediately to the control messages and updates at a high frame rate on the server side, it is difficult to update the screen content with low delay and high frame rate in the client side due to non-negligible time consumption on the whole screen frame compression, transmission, and display buffer updating. To address this critical problem, we propose a layered structure for screen coding and rendering to deliver diverse screen content to the client side with an adaptive frame rate. More specifically, the interaction content with small region screen update is compressed by a blockwise screen codec and rendered at a high frame rate to achieve smooth interaction, while the natural video screen content is compressed by standard video codec and rendered at a regular frame rate for a smooth video display. Experimental results with real applications demonstrate that the proposed system can successfully reduce transmission bandwidth cost and interaction delay during screen sharing. Especially for user interaction in small regions, the proposed system can achieve a higher frame rate than most previous counterparts.
Dan Miao, Jingjing Fu, Yan Lu 0001, Shipeng Li 0001, Chang Wen Chen
ACM Trans. Multim. Comput. Commun. Appl.2
2015 Region-of-interest based coding scheme for synthesized video
abstract
In many multimedia applications, such as online speech, video chat and online conference, multiple source videos are synthesized in a single scene for explicit presentation and the synthesized video is compressed for transmission. The source video with important contents deserves more compression resources for quality preservation under the bandwidth constraint. To address this problem, a region-of-interest (ROI) based coding scheme for synthesized video is proposed in this paper aiming at achieve better and consistent quality for ROI source videos with the bitrate meeting the constraint bandwidth. In the proposed coding scheme, ROI based rate-distortion (R-D) models are established, in which different R-D models are built for different source video. Then an objective function is defined with respect to the video quality and the consistency of video quality. By minimizing the objective function, the optimal quantization parameters for the ROI and non-ROI source videos are obtained. The experimental results show that the proposed coding scheme achieves better and consistent quality for ROI source videos.
Wenbo Zhao 0004, Jingjing Fu, Yan Lu 0001, Shipeng Li 0001, Debin Zhao
VCIP2
2015 Layered Compression for High-Precision Depth Data
abstract
With the development of depth data acquisition technologies, access to high-precision depth with more than 8-b depths has become much easier and determining how to efficiently represent and compress high-precision depth is essential for practical depth storage and transmission systems. In this paper, we propose a layered high-precision depth compression framework based on an 8-b image/video encoder to achieve efficient compression with low complexity. Within this framework, considering the characteristics of the high-precision depth, a depth map is partitioned into two layers: 1) the most significant bits (MSBs) layer and 2) the least significant bits (LSBs) layer. The MSBs layer provides rough depth value distribution, while the LSBs layer records the details of the depth value variation. For the MSBs layer, an error-controllable pixel domain encoding scheme is proposed to exploit the data correlation of the general depth information with sharp edges and to guarantee the data format of LSBs layer is 8 b after taking the quantization error from MSBs layer. For the LSBs layer, standard 8-b image/video codec is leveraged to perform the compression. The experimental results demonstrate that the proposed coding scheme can achieve real-time depth compression with satisfactory reconstruction quality. Moreover, the compressed depth data generated from this scheme can achieve better performance in view synthesis and gesture recognition applications compared with the conventional coding schemes because of the error control algorithm.
Dan Miao, Jingjing Fu, Yan Lu 0001, Shipeng Li 0001, Chang Wen Chen
IEEE Trans. Image Process.2
2014 High frame rate screen video coding for screen sharing applications
abstract
In this paper, we propose a high frame rate screen video compression scheme aiming at improving the interactive user experience on screen sharing applications. The proposed screen video compression is performed as two-layer coding: a base layer coding using the conventional video codec and an enhancement layer coding using the proposed open-loop coding scheme. For efficient frame level layer selection and compression, the content update of each frame is evaluated through global motion detection. The screen frame with significant content update is fed to the conventional video encoder in base layer. In contrast, the frame with little update is compressed in enhancement layer in which the duplicate content is indicated by global motion vector and skip flag while the updated content is encoded by distinct intra modes in terms of inherent local features. The experimental results demonstrate that for the screen video containing interaction, the proposed coding scheme can achieve 3.09ms/frame encoding rate and 2.33ms/frame decoding rate with efficient rate distortion performance.
Dan Miao, Jingjing Fu, Yan Lu 0001, Shipeng Li 0001, Chang Wen Chen
ISCAS2
2014 An adaptive multi-layer low-latency transmission scheme for H.264 based screen sharing system
abstract
Virtual screen system is becoming an essential part in the mobile cloud computing platform. However, designing a low-latency interactive communication for the high-resolution screen content is still challenging due to the network dynamics and the unique characteristics of screen content. In this paper we propose a H.264 based low-latency screen sharing system. To achieve high play-out frame rate, we decouple the low-latency screen content communication problem into two parts, a scalable H.264 based encoding and an optimal scalable stream transmission scheduling. By leveraging the unique characteristics of screen content, a multi-layer scalable video encoding scheme is designed to achieve a certain error resilience while keeping good video coding efficiency. In the transmission scheduling module, an optimal frame skipping policy is proposed to schedule the frames in the buffer to maximize the play-out frame rate. In the performance evaluation, we simulate our system in both one-hop end-to-end topology and two-hop proxy-based topology. The simulation results show that the proposed scheme achieves much better performance on frame rate and average delay, especially in the low bandwidth condition.
Ming Yang 0018, Jingjing Fu, Yan Lu 0001, Jianfei Cai 0001, Chuan Heng Foh
ISCAS2
2013 Layered screen video coding leveraging hardware video codec
abstract
In this paper, we propose a layered screen video coding scheme based on existing video codecs to leverage hardware video codec for efficient screen video compression. In this scheme, the screen video compression is performed as two-layer coding: base layer coding and enhancement layer coding. The screen video is first analyzed in both frame and block levels for useful temporal and spatial information extraction to assist coding content selection in each layer. The non-skip screen frames are directly compressed by the conventional video codec in the base layer, while the screen contents sensitive to the video quality degradation are selected for improved coding in the enhancement layer. For contents to be enhanced, two intra coding modes are designed to improve the quality of the compressed text/graphics contents and suppress the artifacts introduced by chroma downsampling. The experimental results demonstrate that the screen video quality is improved objectively and subjectively by the proposed scheme with low cost on bitrate and computation complexity. Moreover, an average of 2.95dB coding gain is achieved in high bitrate.
Dan Miao, Jingjing Fu, Yan Lu 0001, Shipeng Li 0001, Chang Wen Chen
ICME2
2013 Rate-distortion optimized block classification and bit allocation in screen video compression
abstract
Due to the divergent characteristics of image contents and text contents in screen videos, how to make the joint optimization leveraging rate-distortion (R-D) optimized block classification and bit allocation is critical to the compression performance. In this paper, a general model-based solution is proposed as an attempt to solve this problem. The contributions of this paper are twofold: First, the rate and distortion characteristics of image blocks and text blocks in block-based content-adaptive screen video encoder (BASC) are carefully studied, and the rate and distortion models are proposed. Second, with the proposed rate and distortion models, the R-D optimized block classification and bit allocation are derived using bisection searched Lagrange multiplier method. Experimental results demonstrate that the proposed R-D optimized block classification and bit allocation algorithms are able to adapt to diverse screen contents, which results in a significant gain of up to 4.5dB in PSNR.
Oscar C. Au, Jingjing Fu, Yan Lu 0001, Shipeng Li 0001
ISCAS3
2013 Depth sensor assisted real-time gesture recognition for interactive presentation
Hanjie Wang, Jingjing Fu, Yan Lu 0001, Xilin Chen 0001, Shipeng Li 0001
J. Vis. Commun. Image Represent.2
2013 Kinect-Like Depth Data Compression
abstract
Unlike traditional RGB video, Kinect-like depth is characterized by its large variation range and instability. As a result, traditional video compression algorithms cannot be directly applied to Kinect-like depth compression with respect to coding efficiency. In this paper, we propose a lossy Kinect-like depth compression framework based on the existing codecs, aiming to enhance the coding efficiency while preserving the depth features for further applications. In the proposed framework, the Kinect-like depth is reformed first by divisive normalized bilateral filter (DNBL) to suppress the depth noises caused by disparity normalization, and then block-level depth padding is implemented for invalid depth region compensation in collaboration with mask coding to eliminate the sharp variation caused by depth measurement failures. Before the traditional video coding, the inter-frame correlation of reformed depth is explored by proposed 2D+T prediction, in which depth volume is developed to simulate 3D volume to generate pseudo 3D prediction reference for depth uniqueness detection. The unique depth region, called active region is fed into the video encoder for traditional intra and inter prediction with residual coding, while the inactive region is skipped during depth coding. The experimental results demonstrate that our compression scheme can save 55%-85% in terms of bit cost and reduce coding complexity by 20%-65% in comparison with the traditional video compression algorithms. The visual quality of the 3D reconstruction is also improved after employing our compression scheme.
Jingjing Fu, Dan Miao, Weiren Yu, Shiqi Wang 0001, Yan Lu 0001, Shipeng Li 0001
IEEE Trans. Multim.1
2012 Kinect-like depth denoising
abstract
Accuracy and stability of Kinect-like depth data is limited by its generating principle. In order to serve further applications with high quality depth, the preprocessing on depth data is essential. In this paper, we analyze the characteristics of the Kinect-like depth data by examing its generation principle and propose a spatial-temporal denoising algorithm taking into account its special properties. Both the intra-frame spatial correlation and the inter-frame temporal correlation are exploited to fill the depth hole and suppress the depth noise. Moreover, a divisive normalization approach is proposed to assist the noise filtering process. The 3D rendering results of the processed depth demonstrates that the lost depth is recovered in some hole regions and the noise is suppressed with depth features preserved.
Jingjing Fu, Shiqi Wang 0001, Yan Lu 0001, Shipeng Li 0001, Wenjun Zeng 0001
ISCAS1
2012 Texture-assisted Kinect depth inpainting
abstract
The emergence of Kinect facilitates the possibility of depth capture in real-time and with low cost by consumers. It also provides powerful tool and inspiration for researchers to engage in new array of technology development. However, the quality of the depth map captured from Kinect is still inadequate for many applications due to holes, noises and artifacts existing within the depth information. In this paper, we present a texture assisted Kinect depth inpainting framework, aiming at obtaining improved depth information. In this framework, the relationship between texture and depth is investigated, and the characteristics of depth are also exploited. More specifically, texture edge information is extracted to assist the depth inpainting. Furthermore, filtering and diffusion are designed for hole-filling and edge alignment. Experiment results demonstrate that the Kinect depth can be appropriately repaired in both smooth and edge region. Comparing with the original depth, the inpainted depth information enhances the quality of advanced processing such as 3D reconstruction.
Dan Miao, Jingjing Fu, Yan Lu 0001, Shipeng Li 0001, Chang Wen Chen
ISCAS2
2012 Content-aware layered compound video compression
abstract
Compound video compression is crucial for remote control and data assessment. In this paper, we propose a content-aware layered video coding scheme as an attempt to efficiently compress the compound video. In this scheme, the compound video is analyzed and processed progressively at three pyramid levels: block, object and layer. Firstly, the compound video is analyzed by a block type classification technique to access each block's spatial and temporal properties. Secondly, the natural video object is detected adaptively in each frame based on the block type. Finally, the compound video content is distributed into different layers and specifically designed video coding algorithms are employed to compress each layer. Experiments demonstrate that our proposed scheme can preserve the advantages of the employed compression algorithms for each layer and outperform each of them in the compound video compression.
Shiqi Wang 0001, Jingjing Fu, Yan Lu 0001, Shipeng Li 0001, Wen Gao 0001
ISCAS2
2012 Layered compression for high dynamic range depth
abstract
With the rapid development of depth data acquisition technology, the high precision depth becomes much easier to access in real-time by depth sensors, and the generated high dynamic range (HDR) depth is widely adopted to benefit the depth assistant applications. Accordingly, the HDR depth compression becomes essential for the efficient depth storage and transmission. In this paper, we introduce a layered compression framework for HDR depth to achieve efficient and low-complexity depth compression. To leverage the state-of-art 8-bit image/video encoders, the HDR depth is partitioned into two layers: most significant bit (MSB) layer and least significant bit (LSB) layer. For MSB layer, an error controllable pixel domain encoding scheme is proposed to guarantee the compatibility for existing 8-bit codec by controlling quantization errors added back to LSB layer. Meanwhile, the efficient major color extraction and adaptive quantization enhance the coding performance of MSB layer. For LSB layer, the layer data with limited dynamic range is compressed by normal 8-bit image/video based encoding scheme. The experimental results demonstrate that our coding scheme can achieve real-time depth compression with the satisfactory reconstruction quality. The encoding time is less than 31ms/frame and the decoding time is around 20ms/frame in average. Our compression scheme can be easily integrated into the real-time depth transmission system.
Dan Miao, Jingjing Fu, Yan Lu 0001, Shipeng Li 0001, Chang Wen Chen
VCIP2
2010 Tree Structure Based Analyses on Compressive Sensing for Binary Sparse Sources
abstract
This paper proposes a new approach to theoretically analyze compressive sensing directly from the randomly sampling matrix phi instead of a certain recovery algorithm. For simplifying our analyses, we assume both input source and random sampling matrix as binary. Taking anyone of source bits, we can constitute a tree by parsing the randomly sampling matrix, where the selected source bit as the root. In the rest of tree, measurement nodes and source nodes are connected alternatively according to phi. With the tree, we can formulate the probability if one source bit can be recovered from randomly sampling measurements. The further analyses upon the tree structure reveal the relation between the un-recovery probability with random measurements and the un-recovery probability with source sparsity. The conditions of successful recovery are proven on the parameter S-M plane. Then the results of the tree structure based analyses are compared with the actual recovery process.
Jingjing Fu, Zhouchen Lin
DCC1
2010 Decoding of directional DCT-coded images: A total variational approach with directionality
abstract
This paper presents a novel decoding approach for directional DCT-coded images. In contrast to the conventional approach of performing the inverse transform, our new approach takes the decoding as a general restoration problem in which the total variation (TV) based optimization has been utilized. In this TV-based approach, we propose to measure the gradient according to the directional coding mode determined for each image block at the encoder side. We also consider some neighboring blocks so that possible blocking artifacts can be reduced after the decoding. Experimental results are included to verify the effectiveness of this novel decoding approach.
Jingjing Fu
ISCAS1
2010 TV-based multi-scale super resolution using intra- and inter-scale correlations
abstract
A multi-scale approach is proposed in this paper for image super resolution (SR) from a single source image of a low resolution (LR). In this approach, a sequence of "coarse" SR images, called multi-scales, are first generated from the source image via a total variation (TV) based restoration scheme. All multi-scale images are then brought into the minimum mean-square-error (MMSE) estimation to compose the final SR image. We demonstrate that these multi-scale images contain some geometric information and are highly correlated with each other. Therefore, by making use of intra- and inter-scale correlations, the joint MMSE estimation is able to produce a better quality (both objective and subjective) in the composed SR image. Experimental results are provided to confirm the performance gain over several state-of-the-art methods.
Jiying Wu, Jingjing Fu
ISCAS2
2009 Analysis on Rate-Distortion Performance of Compressive Sensing for Binary Sparse Source
abstract
This paper proposes to use a bipartite graph to represent compressive sensing (CS). The evolution of nodes and edges in the bipartite graph, which is equivalent to the decoding process of compressive sensing, is characterized by a set of differential equations. One of main contributions in this paper is that we derive the close-form formulation of the evolution in statistics, which enable us to more accurately analyze the performance of compressive sensing. Based on the formulation, the distortion of random sampling and the rate needed to code measurements are analyzed briefly. Finally, numerical experiments verify our formulation of the evolution and the rate-distortion curves of compressive sensing are drawn to be compared with entropy coding.
Jingjing Fu, Zhouchen Lin
DCC2
2008 Spectral Information Recovery for Compressed Image Restoration
abstract
The restoration of compressed images typically focuses on the removing of spatial artifacts caused by quantization error rather than the recovery of spectral information in compressed images. In this paper, we attempt to solve the compressed image restoration by considering the recovery of spectral information in compressed images. To this end, we convert this recovery problem to a route searching process in the high dimensional vector space spanned by all DCT coefficients. Since there are huge number of routes in the search space, we present two crucial issues in route construction: how to decide the nodes along the searching route and how to restrict the route within a reasonable sub-space. Total variation (TV) based regularization is applied to determine all the nodes in the searching process, and two constraints on the DCT coefficients are proposed to prune the searching space. Experimental results of our algorithm show a remarkable improvement in both PSNR and visual quality.
Jingjing Fu
DCC1
2008 Cross-frequency spectral prediction for compressed image restoration
abstract
We propose a cross-domain correlation in compress images, and introduce a novel spectral prediction algorithm to restore lossy spectral information caused by compression. This cross-frequency spectral predication algorithm is inspired from the spatial correlation and the connection between discrete cosine transform and Hadamard transform. The relationship among cross-frequency coefficients is adopted to predict spectral coefficients. We apply the spectral prediction algorithm in compressed image restoration under the total variation (TV) based regularization. Experimental results of restoration with or without cross-frequency spectral predication are compared, remarkable improvement is observed from the results with cross-frequency spectral prediction.
Jingjing Fu
MMSP1
2008 Directional Discrete Cosine Transforms - A New Framework for Image Coding
abstract
Nearly all block-based transform schemes for image and video coding developed so far choose the 2-D discrete cosine transform (DCT) of a square block shape. With almost no exception, this conventional DCT is implemented separately through two 1-D transforms, one along the vertical direction and another along the horizontal direction. In this paper, we develop a new block-based DCT framework in which the first transform may choose to follow a direction other than the vertical or horizontal one. The coefficients produced by all directional transforms in the first step are arranged appropriately so that the second transform can be applied to the coefficients that are best aligned with each other. Compared with the conventional DCT, the resulting directional DCT framework is able to provide a better coding performance for image blocks that contain directional edges-a popular scenario in many image signals. By choosing the best from all directional DCTs (including the conventional DCT as a special case) for each image block, we will demonstrate that the rate-distortion coding performance can be improved remarkably. Finally, a brief theoretical analysis is presented to justify why certain coding gain (over the conventional DCT) results from this directional framework.
Jingjing Fu
IEEE Trans. Circuits Syst. Video Technol.2
2007 Directional Discrete Cosine Transforms: A Theoretical Analysis
abstract
Nearly all block-based transform techniques developed so far for image and video coding applications choose the 2-D discrete cosine transform (DCT) of a square block shape. With almost no exception, this conventional DCT is always implemented separately through two 1-D transforms, along the vertical and horizontal directions, respectively. In one of our recent works, we have developed a directional DCT framework in which the first transform may choose to follow a direction other than the vertical or horizontal one, while the second transform is arranged to be a horizontal one. Compared to the conventional DCT, our directional DCT framework has been demonstrated to provide a better coding performance for image blocks that contain directional edges - a popular scenario in many image and video signals. In this paper, we attempt to pursue an in-depth theoretical analysis to understand how the coding gain is produced in the directional DCT framework and how big it can be.
Jingjing Fu
ICASSP (1)1
2007 A Comparative Study of Compensation Techniques in Directional DCT's
abstract
A block-based directional DCT framework has been developed recently in which the first transform may choose to follow a direction other than the vertical or horizontal one – the default direction in the conventional 2D DCT, while the second transform is selected according to the first one. However, these directional DCT's would suffer from the so-called mean weighting defect because DCT's of different lengths have to be used at the first step along different directional lines within each image block, thus resulting in some unnecessary non-zero components during the second transform. This paper presents a comparative study of various techniques for compensating such defect, including different ways of modifying the scaling factors used in the DCT matrices and the so-called ΔDC correction method. These techniques are compared with each other in terms of coding performance and computational complexity.
Jingjing Fu
ISCAS1
2006 Directional Discrete Cosine Transforms for Image Coding
abstract
Nearly all block-based transform schemes for image and video coding developed so far choose the 2-D discrete cosine transform (DCT) of a square block shape. With almost no exception, this conventional DCT is implemented separately through two 1-D transforms, one along the vertical direction and another along the horizontal direction. In this paper, we develop a new block-based DCT framework in which the first transform may follow a direction other than the vertical or horizontal one, while the second transform is arranged to be a horizontal one. Compared to the conventional DCT, the resulting directional DCT framework is able to provide a better coding performance for image blocks that contain directional edges - a popular scenario in many image signals. By choosing the best from all directional DCT's (including the conventional DCT as a special case) for each image block, we will demonstrate that the rate-distortion coding performance can be improved remarkably.
Jingjing Fu
ICME2
2006 Time-domain analysis methodology for large-scale RLC circuits and its applications
Zuying Luo, Yici Cai, Sheldon X.-D. Tan, Xianlong Hong, Zhu Pan, Jingjing Fu
Sci. China Ser. F Inf. Sci.7
2005 VLSI on-chip power/ground network optimization considering decap leakage currents
abstract
In today's power/ground(P/G) network design, on-chip decoupling capacitors(decaps) are usually made of MOS transistors with source and drain connected together. The gate leakage current becomes worse as the gate oxide layer thickness continues to shrink below 20Å. As a result, decaps will become leaky due to the gate leakage from CMOS devices. In this paper, we take a first look at the leaky decaps in P/G network optimization. We propose a leakage current model for practical decaps and also present a new two-stage leakage-current-aware approach to efficiently optimize P/G networks in a more area efficient way.
Jingjing Fu, Zuying Luo, Xianlong Hong, Yici Cai, Sheldon X.-D. Tan, Zhu Pan
ASP-DAC1
2004 A fast decoupling capacitor budgeting algorithm for robust on-chip power delivery
Jingjing Fu, Zuying Luo, Xianlong Hong, Yici Cai, Sheldon X.-D. Tan, Zhu Pan
ASP-DAC1