Qionghai Dai

dblp:39/4543 · DBLP profile ↗
← Back
352ranked-venue papers
6as first author
42since 2021 · last 2026
0000-0001-7043-3061ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 259 · 5 first-author · 17 since 2021Artificial intelligence and machine learning · 107 · 32 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 12 · 1 since 2021Systems, architecture and hardware · 5Human-computer interaction and ubiquitous computing · 5
YearPublicationVenuePosition
2026 Creating Multimodal Interactive Digital Twin Characters From Videos: A Dataset and Baseline
abstract
In this paper, we introduce a novel framework for creating multimodal interactive digital twin characters, from dialogue videos of TV shows. Specifically, these digital twin characters are capable of responding to user inputs with harmonious textual, vocal, and visual content. They not only replicate the external characteristics, such as appearance and tone, but also capture internal attributes, including personality and habitual behaviors. To support this ambitious task, we collect the Multimodal Character-Centric Conversation Dataset, named MCCCD, which includes character-specific and high-quality multimodal dialogue data with detailed annotations, featuring 6.8 k utterances and 4.6 hours of audio/video per character. Notably, the MCCCD dataset is approximately ten times larger than existing datasets in terms of per-character data volume, facilitating the detailed modeling of complex character-centric traits. Further, we propose a baseline framework to create digital twin characters, consists of dialogue generation through large language models, voice generation via speech synthesis models, and visual representation with 3D talking head models. Experimental results demonstrate that our approach significantly outperforms existing methods in generating consistent and character-specific responses, setting a new benchmark for digital character creation. Our collected dataset and proposed baseline have paved the way for the creation of highly interactive and natural digital avatars, opening the door to extensive and practical applications of digital humans.
Meidai Xuanyuan, Yuwang Wang, Hanshi Qu, Zhongming Li, Danping Yan, Jianhua Tao 0001, Qionghai Dai
IEEE Trans. Pattern Anal. Mach. Intell.10
2026 DICFusion: Infrared and Visible Image Fusion via a Deep Integrated and Semantic-Coordinated Network
abstract
Infrared-visible image fusion (IVF) aims to integrate complementary information from infrared and visible sensors into a single, more informative representation. However, achieving both visual clarity and semantic consistency in the fused results remains a critical challenge, particularly for real-world applications like scene understanding. To address this, we propose DICFusion, a deep integrated and semantic-coordinated network, tailored for perceptually and semantically enriched infrared-visible fusion tasks. Firstly, the DICFusion employs a novel modality-aware fusion strategy to integrate infrared and visible modalities into a cohesive feature embedding. Secondly, the framework incorporates hybrid mamba-convolution blocks, which leverage the combined strengths of mamba and convolution neural networks to accurately capture both global context and localized details while maintaining computational efficiency. To mitigate the feature heterogeneity between fusion and downstream tasks, DICFusion adopts a comprehensive framework of deep integration and collaborative optimization. This design utilizes a unified multi-scale encoder to harmonize feature representations, followed by parallel fusion and segmentation branches to enhance both visual quality and task performance. Moreover, a semantic guidance module leveraging cross-attention mechanism is incorporated to refine the semantic consistency of the fused outcomes. Comprehensive experimental evaluations validate the performance and efficiency of DICFusion, demonstrating its superiority over contemporary state-of-the-art methods, both in terms of fusion visual quality and downstream task precision. The code is available at https://github.com/fd-qhwang/DICFusion.
Tianyun Wang, Yuhong Luo, Feng Bao 0002, Nan Chi, Qionghai Dai
IEEE Trans. Circuits Syst. Video Technol.8
2026 GBR: Generative Bundle Refinement for High-Fidelity Gaussian Splatting With Enhanced Mesh Reconstruction
abstract
Gaussian splatting has gained attention for its efficient representation and rendering of 3D scenes using continuous Gaussian primitives. However, it struggles with sparse-view inputs due to limited geometric and photometric information, causing ambiguities in depth, shape, and texture. We propose GBR: Generative Bundle Refinement, a method for high-fidelity Gaussian splatting and meshing using only 4–6 input views. GBR integrates a neural bundle adjustment module to enhance geometry accuracy and a generative depth refinement module to improve geometry fidelity. More specifically, the neural bundle adjustment module integrates a foundation network to produce initial 3D point maps and point matches from unposed images, followed by bundle adjustment optimization to improve multiview consistency and point cloud accuracy. The generative depth refinement module employs a diffusion-based strategy to enhance geometric details and fidelity while preserving the scale. Finally, for Gaussian splatting optimization, we propose a multimodal loss function incorporating depth and normal consistency, geometric regularization, and pseudo-view supervision, providing robust guidance under sparse-view conditions. Experiments on widely used datasets show that GBR significantly outperforms existing methods under sparse-view inputs. Additionally, GBR demonstrates the ability to reconstruct and render large-scale real-world scenes, such as the Pavilion of Prince Teng and the Great Wall, with remarkable details using only 6 views. More results can be found on our project page https://gbrnvs.github.io.
Qionghai Dai, Xiaoyun Yuan
IEEE Trans. Circuits Syst. Video Technol.4
2025 Neural Fluid Simulation on Geometric Surfaces
abstract
Incompressible fluid on the surface is an interesting research area in the fluid simulation, which is the fundamental building block in visual effects, design of liquid crystal films, scientific analyses of atmospheric and oceanic phenomena, etc. The task brings two key challenges: the extension of the physical laws on 3D surfaces and the preservation of the energy and volume. Traditional methods rely on grids or meshes for spatial discretization, which leads to high memory consumption and a lack of robustness and adaptivity for various mesh qualities and representations. Many implicit representations based simulators like INSR are proposed for the storage efficiency and continuity, but they face challenges in the surface simulation and the energy dissipation. We propose a neural physical simulation framework on the surface with the implicit neural representation. Our method constructs a parameterized vector field with the exterior calculus and Closest Point Method on the surfaces, which guarantees the divergence-free property and enables the simulation on different surface representations (e.g. implicit neural represented surfaces). We further adopt a corresponding covariant derivative based advection process for surface flow dynamics and energy preservation. Our method shows higher accuracy, flexibility and memory-efficiency in the simulations of various surfaces with low energy dissipation. Numerical studies also highlight the potential of our framework across different practical applications such as vorticity shape generation and vector field Helmholtz decomposition.
Haoxiang Wang 0006, Tao Yu 0007, Qionghai Dai
ICLR4
2025 Lightweight High-Speed Photography Built on Coded Exposure and Implicit Neural Representation of Videos
Zhihong Zhang 0004, Runzhao Yang, Jin-Li Suo, Yuxiao Cheng, Qionghai Dai
Int. J. Comput. Vis.5
2025 Event-Enhanced Snapshot Compressive Videography at 10K FPS
abstract
Video snapshot compressive imaging (SCI) encodes the target dynamic scene compactly into a snapshot and reconstructs its high-speed frame sequence afterward, greatly reducing the required data footprint and transmission bandwidth as well as enabling high-speed imaging with a low frame rate intensity camera. In implementation, high-speed dynamics are encoded via temporally varying patterns, and only frames at corresponding temporal intervals can be reconstructed, while the dynamics occurring between consecutive frames are lost. To unlock the potential of conventional snapshot compressive videography, we propose a novel hybrid "intensity event imaging scheme by incorporating an event camera into a video SCI setup. Our proposed system consists of a dual-path optical setup to record the coded intensity measurement and intermediate event signals simultaneously, which is compact and photon-efficient by collecting the half photons discarded in conventional video SCI. Correspondingly, we developed a dual-branch Transformer utilizing the reciprocal relationship between two data modes to decode dense video frames. Extensive experiments on both simulated and real-captured data demonstrate our superiority to state-of-the-art video SCI and video frame interpolation (VFI) methods. Benefiting from the new hybrid design leveraging both intrinsic redundancy in videos and the unique feature of event cameras, we achieve high-quality videography at 0.1ms time intervals with a low-cost CMOS image sensor working at 24 FPS.
Bo Zhang 0109, Jin-Li Suo, Qionghai Dai
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 WaveFusion: A Novel Wavelet Vision Transformer With Saliency-Guided Enhancement for Multimodal Image Fusion
abstract
Multi-modal image fusion aims to amalgamate pivotal information from various sensor sources to provide informative visual representation in imaging scenes. Rapid and precise fusion of images is crucial for practical applications in fields such as autonomous driving and medical diagnostics. However, the primary challenge lies in balancing computational costs with the effectiveness of feature extraction, while ensuring the robust integration of salient features across modalities. Here, this paper introduces WaveFusion, a wavelet vision transformer equipped with an advanced saliency-guided loss strategy to optimize multi-modal image fusion. Initially, to provide a comprehensive and efficient representation of multi-modal data, we introduce an adaptive wavelet transform module for feature decomposition and reconstruction. Following this, self-attention mechanisms and convolutional networks are naturally applied in parallel to process low-frequency and high-frequency components, resulting in the development of a wavelet-enhanced vision transformer. Secondly, WaveFusion utilizes a dual-aggregation attention approach that improves cross-modal feature complementarity and intra-modal feature coherence within a single fusion module. Furthermore, we propose a dynamic saliency-informed selective loss function to refine the optimization process, with the objective of enhancing critical feature retention and maintaining overall image consistency across fusion scenarios. The efficacy and versatility of our method are validated in both infrared-visible fusion and medical image fusion tasks. Experiment results demonstrate that WaveFusion provides a superior balanced approach that optimizes both fusion performance and cost-efficiency, and additionally improves performance in downstream tasks such as multi-modal semantic segmentation and object detection.
Nan Chi, Qionghai Dai
IEEE Trans. Circuits Syst. Video Technol.5
2025 Hard or False: Keep the Balance for Negative Sampling in Knowledge Graphs
abstract
Negative sampling is an essential part in knowledge graph embedding, which offers significant advantages to numerous downstream related tasks. There are two kinds of important negatives: hard and false negatives. Hard negatives are the negatives which are difficult to distinguish from positive samples, while false negatives are positive samples which are mistakenly identified as negatives. Harnessing hard negatives effectively can make the model more discriminative, and reducing false negatives can avoid misleading the model during training. Therefore, the two kinds of negatives are essential in high-quality negative sampling. However, the present negative sampling methods face two shortcomings: 1.judging one negative is hard or false mainly relies on score functions; 2. difficulty in balancing the impact of hard and false negatives. In this paper, we absorb bigram language model and propose a novel criterion to help verify the negatives are hard or false, and discuss how to keep the balance between hard and false negatives. Experiments on four representative score functions and two public datasets demonstrate the effects of the proposed negative sampling method.
Feihu Che, Jianhua Tao 0001, Qionghai Dai
IEEE Trans. Knowl. Data Eng.3
2025 Super-NeRF: View-Consistent Detail Generation for NeRF Super-Resolution
abstract
The neural radiance field (NeRF) achieved remarkable success in modeling 3D scenes and synthesizing high-fidelity novel views. However, existing NeRF-based methods focus more on making full use of high-resolution images to generate high-resolution novel views, but less considering the generation of high-resolution details given only low-resolution images. In analogy to the extensive usage of image super-resolution, NeRF super-resolution is an effective way to generate low-resolution-guided high-resolution 3D scenes and holds great potential applications. Up to now, such an important topic is still under-explored. In this article, we propose a NeRF super-resolution method, named Super-NeRF, to generate high-resolution NeRF from only low-resolution inputs. Given multi-view low-resolution images, Super-NeRF constructs a multi-view consistency-controlling super-resolution module to generate various view-consistent high-resolution details for NeRF. Specifically, an optimizable latent code is introduced for each input view to control the generated reasonable high-resolution 2D images satisfying view consistency. The latent codes of each low-resolution image are optimized synergistically with the target Super-NeRF representation to utilize the view consistency constraint inherent in NeRF construction. We verify the effectiveness of Super-NeRF on synthetic, real-world, and even AI-generated NeRFs. Super-NeRF achieves state-of-the-art NeRF super-resolution performance on high-resolution detail generation and cross-view consistency.
Yuqi Han, Tao Yu 0007, Xiaohang Yu, Di Xu 0012, Binge Zheng, Zonghong Dai, Changpeng Yang, Yuwang Wang, Qionghai Dai
IEEE Trans. Vis. Comput. Graph.9
2025 ImmersiveNeRF: Hybrid Radiance Fields for Unbounded Immersive Light Field Reconstruction
abstract
This article proposes a hybrid radiance field representation for unbounded immersive light field reconstruction which supports high-quality rendering and aggressive view extrapolation. The key idea is to first formally separate the foreground and the background and then adaptively balance learning of them during the training process. To fulfill this goal, we represent the foreground and background as two separate radiance fields with two different spatial mapping strategies. We further propose an adaptive sampling strategy and a segmentation regularizer for more clear segmentation and robust convergence. Finally, we contribute a novel immersive light field dataset, named THUImmersive, with the potential to achieve much larger space 6DoF immersive rendering effects compared with existing datasets, by capturing multiple neighboring viewpoints for the same scene, to stimulate the research and AR/VR applications in the immersive light field domain. Extensive experiments demonstrate the strong performance of our method for unbounded immersive light field reconstruction.
Xiaohang Yu, Haoxiang Wang 0006, Yuqi Han, Lei Yang 0045, Tao Yu 0007, Qionghai Dai
IEEE Trans. Vis. Comput. Graph.6
2024 CUTS+: High-Dimensional Causal Discovery from Irregular Time-Series
abstract
Causal discovery in time-series is a fundamental problem in the machine learning community, enabling causal reasoning and decision-making in complex scenarios. Recently, researchers successfully discover causality by combining neural networks with Granger causality, but their performances degrade largely when encountering high-dimensional data because of the highly redundant network design and huge causal graphs. Moreover, the missing entries in the observations further hamper the causal structural learning. To overcome these limitations, We propose CUTS+, which is built on the Granger-causality-based causal discovery method CUTS and raises the scalability by introducing a technique called Coarse-to-fine-discovery (C2FD) and leveraging a message-passing-based graph neural network (MPGNN). Compared to previous methods on simulated, quasi-real, and real datasets, we show that CUTS+ largely improves the causal discovery performance on high-dimensional data with different types of irregular sampling.
Yuxiao Cheng, Lianglong Li, Tingxiong Xiao, Zongren Li, Jin-Li Suo, Kunlun He, Qionghai Dai
AAAI7
2024 Neural Physical Simulation with Multi-Resolution Hash Grid Encoding
abstract
We explore the generalization of the implicit representation in the physical simulation task. Traditional time-dependent partial differential equations (PDEs) solvers for physical simulation often adopt the grid or mesh for spatial discretization, which is memory-consuming for high resolution and lack of adaptivity. Many implicit representations like local extreme machine or Siren are proposed but they are still too compact to suffer from limited accuracy in handling local details and a long time of convergence. We contribute a neural simulation framework based on multi-resolution hash grid representation to introduce hierarchical consideration of global and local information, simultaneously. Furthermore, we propose two key strategies: 1) a numerical gradient method for computing high-order derivatives with boundary conditions; 2) a range analysis sample method for fast neural geometry boundary sampling with dynamic topologies. Our method shows much higher accuracy and strong flexibility for various simulation problems: e.g., large elastic deformations, complex fluid dynamics, and multi-scale phenomena which remain challenging for existing neural physical solvers.
Haoxiang Wang 0006, Tao Yu 0007, Tianwei Yang, Qionghai Dai
AAAI5
2024 A Physics-Informed Low-Rank Deep Neural Network for Blind and Universal Lens Aberration Correction
abstract
High-end lenses, although offering high-quality images, suffer from both insufficient affordability and bulky design, which hamper their applications in low-budget scenarios or on low-payload platforms. A flexible scheme is to tackle the optical aberration of low-end lenses computationally. However, it is highly demanded but quite challenging to build a general model capable of handling non-stationary aberrations and covering diverse lenses, especially in a blind manner. To address this issue, we propose a universal solution by extensively utilizing the physical properties of camera lenses: (i) reducing the complexity of lens aberrations, i.e., lens-specific non-stationary blur, by warping annual-ring-shaped sub-images into rectangular stripes to transform non-uniform degenerations into a uniform one, (ii) building a low-dimensional nonnegative orthogonal representation of lens blur kernels to cover diverse lenses; (iii) designing a decoupling network to decompose the input low-quality image into several components degenerated by above kernel bases, and applying corresponding pretrained deconvolution networks to reverse the degeneration. Benefiting from the proper incorporation of lenses' physical properties and unique network design, the proposed method achieves superb imaging quality, wide applicability for various lenses, high running efficiency, and is totally free of kernel calibration. These advantages bring great potential for scenarios requiring lightweight high-quality photography.
Jin Gong, Runzhao Yang, Jin-Li Suo, Qionghai Dai
CVPR5
2024 A versatile Wavelet-Enhanced CNN-Transformer for improved fluorescence microscopy image restoration
Nan Chi, Qionghai Dai
Neural Networks5
2024 Hypergraph-Based Multi-Modal Representation for Open-Set 3D Object Retrieval
abstract
The traditional 3D object retrieval (3DOR) task is under the close-set setting, which assumes the categories of objects in the retrieval stage are all seen in the training stage. Existing methods under this setting may tend to only lazily discriminate their categories, while not learning a generalized 3D object embedding. Under such circumstances, it is still a challenging and open problem in real-world applications due to the existence of various unseen categories. In this paper, we first introduce the open-set 3DOR task to expand the applications of the traditional 3DOR task. Then, we propose the Hypergraph-Based Multi-Modal Representation (HGM$^{2}$R) framework to learn 3D object embeddings from multi-modal representations under the open-set setting. The proposed framework is composed of two modules, i.e., the Multi-Modal 3D Object Embedding (MM3DOE) module and the Structure-Aware and Invariant Knowledge Learning (SAIKL) module. By utilizing the collaborative information of modalities derived from the same 3D object, the MM3DOE module is able to overcome the distinction across different modality representations and generate unified 3D object embeddings. Then, the SAIKL module utilizes the constructed hypergraph structure to model the high-order correlation among 3D objects from both seen and unseen categories. The SAIKL module also includes a memory bank that stores typical representations of 3D objects. By aligning with those memory anchors in the memory bank, the aligned embeddings can integrate the invariant knowledge to exhibit a powerful generalized capacity toward unseen categories. We formally prove that hypergraph modeling has better representative capability on data correlation than graph modeling. We generate four multi-modal datasets for the open-set 3DOR task, i.e., OS-ESB-core, OS-NTU-core, OS-MN40-core, and OS-ABO-core, in which each 3D object contains three modality representations: multi-view, point clouds, and voxel. Experiments on these four datasets show that the proposed method can significantly outperform existing methods. In particular, the proposed method outperforms the state-of-the-art by 12.12%/12.88% in terms of mAP on the OS-MN40-core/OS-ABO-core dataset, respectively. Results and visualizations demonstrate that the proposed method can effectively extract the generalized 3D object embeddings on the open-set 3DOR task and achieve satisfactory performance.
Yifan Feng 0001, Shuyi Ji, Yu-Shen Liu, Shaoyi Du, Qionghai Dai, Yue Gao 0002
IEEE Trans. Pattern Anal. Mach. Intell.5
2024 Audio Matters Too! Enhancing Markerless Motion Capture with Audio Signals for String Performance Capture
abstract
In this paper, we touch on the problem of markerless multi-modal human motion capture especially for string performance capture which involves inherently subtle hand-string contacts and intricate movements. To fulfill this goal, we first collect a dataset, named String Performance Dataset (SPD), featuring cello and violin performances. The dataset includes videos captured from up to 23 different views, audio signals, and detailed 3D motion annotations of the body, hands, instrument, and bow. Moreover, to acquire the detailed motion annotations, we propose an audio-guided multi-modal motion capture framework that explicitly incorporates hand-string contacts detected from the audio signals for solving detailed hand poses. This framework serves as a baseline for string performance capture in a completely markerless manner without imposing any external devices on performers, eliminating the potential of introducing distortion in such delicate movements. We argue that the movements of performers, particularly the sound-producing gestures, contain subtle information often elusive to visual methods but can be inferred and retrieved from audio cues. Consequently, we refine the vision-based motion capture results through our innovative audio-guided approach, simultaneously clarifying the contact relationship between the performer and the instrument, as deduced from the audio. We validate the proposed framework and conduct ablation studies to demonstrate its efficacy. Our results outperform current state-of-the-art vision-based algorithms, underscoring the feasibility of augmenting visual motion capture with audio modality. To the best of our knowledge, SPD is the first dataset for musical instrument performance, covering fine-grained hand motion details in a multi-modal, large-scale collection. It holds significant implications and guidance for string instrument pedagogy, animation, and virtual concerts, as well as for both musical performance analysis and generation. Our code and SPD dataset are available at https://github.com/Yitongishere/string_performance.
Yitong Jin, Zhiping Qiu, Yi Shi 0009, Shuangpeng Sun, Chongwu Wang, Donghao Pan, Zhenghao Liang, Yuan Wang 0052, Feng Yu 0032, Tao Yu 0007, Qionghai Dai
ACM Trans. Graph.13
2023 SCI: A Spectrum Concentrated Implicit Neural Compression for Biomedical Data
abstract
Massive collection and explosive growth of biomedical data, demands effective compression for efficient storage, transmission and sharing. Readily available visual data compression techniques have been studied extensively but tailored for natural images/videos, and thus show limited performance on biomedical data which are of different features and larger diversity. Emerging implicit neural representation (INR) is gaining momentum and demonstrates high promise for fitting diverse visual data in target-data-specific manner, but a general compression scheme covering diverse biomedical data is so far absent. To address this issue, we firstly derive a mathematical explanation for INR's spectrum concentration property and an analytical insight on the design of INR based compressor. Further, we propose a Spectrum Concentrated Implicit neural compression (SCI) which adaptively partitions the complex biomedical data into blocks matching INR's concentrated spectrum envelop, and design a funnel shaped neural network capable of representing each block with a small number of parameters. Based on this design, we conduct compression via optimization under given budget and allocate the available parameters with high representation accuracy. The experiments show SCI's superior performance to state-of-the-art methods including commercial compressors, data-driven ones, and INR based counterparts on diverse biomedical data. The source code can be found at https://github.com/RichealYoung/ImplicitNeuralCompression.git.
Runzhao Yang, Tingxiong Xiao, Yuxiao Cheng, Qianni Cao, Jinyuan Qu, Jin-Li Suo, Qionghai Dai
AAAI7
2023 PARF: Primitive-Aware Radiance Fusion for Indoor Scene Novel View Synthesis
abstract
This paper proposes a method for fast scene radiance field reconstruction with strong novel view synthesis performance and convenient scene editing functionality. The key idea is to fully utilize semantic parsing and primitive extraction for constraining and accelerating the radiance field reconstruction process. To fulfill this goal, a primitive-aware hybrid rendering strategy was proposed to enjoy the best of both volumetric and primitive rendering. We further contribute a reconstruction pipeline conducts primitive parsing and radiance field learning iteratively for each input frame which successfully fuses semantic, primitive, and radiance information into a single framework. Extensive evaluations demonstrate the fast reconstruction ability, high rendering quality, and convenient editing functionality of our method.
Haiyang Ying, Baowei Jiang, Jinzhi Zhang, Di Xu 0012, Tao Yu 0007, Qionghai Dai, Lu Fang 0001
ICCV6
2023 CUTS: Neural Causal Discovery from Irregular Time-Series Data
Yuxiao Cheng, Runzhao Yang, Tingxiong Xiao, Zongren Li, Jin-Li Suo, Kunlun He, Qionghai Dai
ICLR7
2023 Triangulation Residual Loss for Data-efficient 3D Pose Estimation
abstract
This paper presents Triangulation Residual loss (TR loss) for multiview 3D pose estimation in a data-efficient manner. Existing 3D supervised models usually require large-scale 3D annotated datasets, but the amount of existing data is still insufficient to train supervised models to achieve ideal performance, especially for animal pose estimation. To employ unlabeled multiview data for training, previous epipolar-based consistency provides a self-supervised loss that considers only the local consistency in pairwise views, resulting in limited performance and heavy calculations. In contrast, TR loss enables self-supervision with global multiview geometric consistency. Starting from initial 2D keypoint estimates, the TR loss can fine-tune the corresponding 2D detector without 3D supervision by simply minimizing the smallest singular value of the triangulation matrix in an end-to-end fashion. Our method achieves the state-of-the-art 25.8mm MPJPE and competitive 28.7mm MPJPE with only 5\% 2D labeled training data on the Human3.6M dataset. Experiments on animals such as mice demonstrate our TR loss's data-efficient training ability.
Tao Yu 0007, Liang An 0002, Yipeng Huang 0005, Fang Deng, Qionghai Dai
NeurIPS6
2023 Generating Hypergraph-Based High-Order Representations of Whole-Slide Histopathological Images for Survival Prediction
abstract
Patient survival prediction based on gigapixel whole-slide histopathological images (WSIs) has become increasingly prevalent in recent years. A key challenge of this task is achieving an informative survival-specific global representation from those WSIs with highly complicated data correlation. This article proposes a multi-hypergraph based learning framework, called "HGSurvNet," to tackle this challenge. HGSurvNet achieves an effective high-order global representation of WSIs via multilateral correlation modeling in multiple spaces and a general hypergraph convolution network. It has the ability to alleviate over-fitting issues caused by the lack of training data by using a new convolution structure called hypergraph max-mask convolution. Extensive validation experiments were conducted on three widely-used carcinoma datasets: Lung Squamous Cell Carcinoma (LUSC), Glioblastoma Multiforme (GBM), and National Lung Screening Trial (NLST). Quantitative analysis demonstrated that the proposed method consistently outperforms state-of-the-art methods, coupled with the Bayesian Concordance Readjust loss. We also demonstrate the individual effectiveness of each module of the proposed framework and its application potential for pathology diagnosis and reporting empowered by its interpretability potential.
Donglin Di, Changqing Zou, Yifan Feng 0001, Rongrong Ji, Qionghai Dai, Yue Gao 0002
IEEE Trans. Pattern Anal. Mach. Intell.6
2023 SuperFast: 200× Video Frame Interpolation via Event Camera
abstract
Traditional frame-based video frame interpolation (VFI) methods rely on the linear motion assumption and brightness invariance assumption, which may lead to fatal errors confronting the scenarios with high-speed motions. To tackle the above challenge, inspired by the advantages of event cameras on asynchronously recording brightness changes at each pixel, we propose a Fast-Slow joint synthesis framework for event-enhanced high-speed video frame interpolation, named SuperFast, in this paper, which can generate high frame rate (5000 FPS, 200× faster) video from the input low frame rate (25 FPS) video and the corresponding event stream. In our framework, the task is divided into two sub-tasks, i.e., video frame interpolation for the contents with and without high-speed motions, which are tackled by two corresponding branches, i.e., the fast synthesis pathway and the slow synthesis pathway. The fast synthesis pathway leverages a spiking neural network to encode the input event stream, and combines boundary frames to generate intermediate results through synthesis and refinement, targeting on contents with high-speed motions. The slow synthesis pathway stacks the two input boundary frames and the event stream to synthesize intermediate results, focusing on relatively slow-motion contents. Finally, a fusion module with a comparison loss is utilized to generate the final video frame interpolation results. We also build a hybrid visual acquisition system containing an event camera and a high frame rate camera, and collect the first 5000 FPS High-Speed Event-enhanced Video frame Interpolation (THU[Formula: see text]) dataset. To evaluate the performance of our proposed framework, we have conducted experiments on our THU[Formula: see text] dataset and the existing HS-ERGB dataset. Experimental results demonstrate that our proposed framework can achieve state-of-the-art 200× video frame interpolation performance under high-speed motion scenarios.
Yue Gao 0002, Siqi Li 0001, Yandong Guo, Qionghai Dai
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 Action Recognition and Benchmark Using Event Cameras
abstract
Recent years have witnessed remarkable achievements in video-based action recognition. Apart from traditional frame-based cameras, event cameras are bio-inspired vision sensors that only record pixel-wise brightness changes rather than the brightness value. However, little effort has been made in event-based action recognition, and large-scale public datasets are also nearly unavailable. In this paper, we propose an event-based action recognition framework calledEV-ACT. The Learnable Multi-Fused Representation (LMFR) is first proposed to integrate multiple event information in a learnable manner. The LMFR with dual temporal granularity is fed into the event-based slow-fast network for the fusion of appearance and motion features. A spatial-temporal attention mechanism is introduced to further enhance the learning capability of action recognition. To prompt research in this direction, we have collected the largest event-based action recognition benchmark namedTHUE-ACT-50and the accompanyingTHUE-ACT-50-CHLdataset under challenging environments, including a total of over 12,830 recordings from 50 action categories, which is over 4 times the size of the previous largest dataset. Experimental results show that our proposed framework could achieve improvements of over 14.5%, 7.6%, 11.2%, and 7.4% compared to previous works on four benchmarks. We have also deployed our proposed EV-ACT framework on a mobile platform to validate its practicality and efficiency.
Yue Gao 0002, Jiaxuan Lu, Siqi Li 0001, Nan Ma 0012, Shaoyi Du, Qionghai Dai
IEEE Trans. Pattern Anal. Mach. Intell.7
2023 STORM: Structure-Based Overlap Matching for Partial Point Cloud Registration
abstract
Partial point cloud registration aims to transform partial scans into a common coordinate system. It is an important preprocessing step to generate complete 3D shapes. Although previous registration methods have made great progress in recent decades, traditional registration methods, such as Iterative Closest Point (ICP) and its variants, all these methods highly depend on the sufficient overlaps between two point clouds, because they cannot distinguish outlier correspondences. Note that the overlap between point clouds could always be small, which limits the application of these methods. To tackle this problem, we present a StrucTure-based OveRlap Matching (STORM) method for partial point cloud registration. In our method, an overlap prediction module with differentiable sampling is designed to detect points in overlap utilizing structure information, and facilitates exact partial correspondence generation, which is based on discriminative pointwise feature similarity. The pointwise features which contain effective structural information are extracted by graph-based methods. Experimental results and comparison with state-of-the-art methods demonstrate that STORM can achieve better performance. Moreover, most registration methods perform worse when the overlap ratio decreases, while STORM can still achieve satisfactory performance when the overlap ratio is small.
Chenggang Yan 0001, Yutong Feng, Shaoyi Du, Qionghai Dai, Yue Gao 0002
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 Computational Imaging and Artificial Intelligence: The Next Revolution of Mobile Vision
abstract
Signal capture is at the forefront of perceiving and understanding the environment; thus, imaging plays a pivotal role in mobile vision. Recent unprecedented progress in artificial intelligence (AI) has shown great potential in the development of advanced mobile platforms with new imaging devices. Traditional imaging systems based on the “capturing images first and processing afterward” mechanism cannot meet this explosive demand. On the other hand, computational imaging (CI) systems are designed to capture high-dimensional data in an encoded manner to provide more information for mobile vision systems. Thanks to AI, CI can now be used in real-life systems by integrating deep learning algorithms into the mobile vision platform to achieve a closed loop of intelligent acquisition, processing, and decision-making, thus leading to the next revolution of mobile vision. Starting from the history of mobile vision using digital cameras, this work first introduces the advancement of CI in diverse applications and then conducts a comprehensive review of current research topics combining CI and AI. Although new-generation mobile platforms, represented by smart mobile phones, have deeply integrated CI and AI for better image acquisition and processing, most mobile vision platforms, such as self-driving cars and drones only loosely connect CI and AI, and are calling for a closer integration. Motivated by this fact, at the end of this work, we propose some potential technologies and disciplines that aid the deep integration of CI and AI and shed light on new directions in the future generation of mobile vision platforms.
Jin-Li Suo, Jin Gong, Xin Yuan 0002, David J. Brady, Qionghai Dai
Proc. IEEE6
2023 Retrieving Object Motions From Coded Shutter Snapshot in Dark Environment
abstract
Video object detection is a widely studied topic and has made significant progress in the past decades. However, the feature extraction and calculations in existing video object detectors demand decent imaging quality and avoidance of severe motion blur. Under extremely dark scenarios, due to limited sensor sensitivity, we have to trade off signal-to-noise ratio for motion blur compensation or vice versa, and thus suffer from performance deterioration. To address this issue, we propose to temporally multiplex a frame sequence into one snapshot and extract the cues characterizing object motion for trajectory retrieval. For effective encoding, we build a prototype for encoded capture by mounting a highly compatible programmable shutter. Correspondingly, in terms of decoding, we design an end-to-end deep network called detection from coded snapshot (DECENT) to retrieve sequential bounding boxes from the coded blurry measurements of dynamic scenes. For effective network learning, we generate quasi-real data by incorporating physically-driven noise into the temporally coded imaging model, which circumvents the unavailability of training data and with high generalization ability on real dark videos. The approach offers multiple advantages, including low bandwidth, low cost, compact setup, and high accuracy. The effectiveness of the proposed approach is experimentally validated under low illumination vision and provide a feasible way for night surveillance.
Kaiming Dong, Runzhao Yang, Yuxiao Cheng, Jin-Li Suo, Qionghai Dai
IEEE Trans. Image Process.6
2023 INFWIDE: Image and Feature Space Wiener Deconvolution Network for Non-Blind Image Deblurring in Low-Light Conditions
abstract
Under low-light environment, handheld photography suffers from severe camera shake under long exposure settings. Although existing deblurring algorithms have shown promising performance on well-exposed blurry images, they still cannot cope with low-light snapshots. Sophisticated noise and saturation regions are two dominating challenges in practical low-light deblurring: the former violates the Gaussian or Poisson assumption widely used in most existing algorithms and thus degrades their performance badly, while the latter introduces non-linearity to the classical convolution-based blurring model and makes the deblurring task even challenging. In this work, we propose a novel non-blind deblurring method dubbed image and feature space Wiener deconvolution network (INFWIDE) to tackle these problems systematically. In terms of algorithm design, INFWIDE proposes a two-branch architecture, which explicitly removes noise and hallucinates saturated regions in the image space and suppresses ringing artifacts in the feature space, and integrates the two complementary outputs with a subtle multi-scale fusion network for high quality night photograph deblurring. For effective network training, we design a set of loss functions integrating a forward imaging model and backward reconstruction to form a close-loop regularization to secure good convergence of the deep neural network. Further, to optimize INFWIDE's applicability in real low-light conditions, a physical-process-based low-light noise model is employed to synthesize realistic noisy night photographs for model training. Taking advantage of the traditional Wiener deconvolution algorithm's physically driven characteristics and deep neural network's representation ability, INFWIDE can recover fine details while suppressing the unpleasant artifacts during deblurring. Extensive experiments on synthetic data and real data demonstrate the superior performance of the proposed approach.
Zhihong Zhang 0004, Yuxiao Cheng, Jin-Li Suo, Liheng Bian, Qionghai Dai
IEEE Trans. Image Process.5
2022 Predicting algorithm of attC site based on combination optimization strategy
abstract
Site-specific recombination systems are widely used as bioengineering tools. However, the traditional site-specific recombination system requires a consensus sequence for the specific site. Such sequence-level constraints limit effective recombination between sites. Therefore, in order to develop an efficient site-specific recombination system, we investigated the attC site of the bacterial integration subsystem and built a predictive model to infer the important features that contribute to recombination. Here, we design an attC site prediction algorithm based on a combination optimisation strategy. Based on the structural features of attC sites, the prediction algorithm realises the high-precision prediction of the recombination frequencies between sites and the screening of the top 20 important features that play a role in recombination, which are effective for improving the design method of attC sites. The algorithm has better portability and higher prediction accuracy compared with the existing advanced algorithms, among which the Pearson correlation coefficient is 0.87, explained variance score is 0.73, root mean square error is 0.006 and mean absolute error is 0.041. This can not only provide ideas for the research of efficient recombination systems but also provide a theoretical basis for developing genetic engineering further.
Dongyan Li, Xinrong Lv, Mengying Qin, Ke Bai 0006, Yurong Yang, Qionghai Dai
Connect. Sci.10
2022 Heterogeneous Hypergraph Variational Autoencoder for Link Prediction
abstract
Link prediction aims at inferring missing links or predicting future ones based on the currently observed network. This topic is important for many applications such as social media, bioinformatics and recommendation systems. Most existing methods focus on homogeneous settings and consider only low-order pairwise relations while ignoring either the heterogeneity or high-order complex relations among different types of nodes, which tends to lead to a sub-optimal embedding result. This paper presents a method named Heterogeneous Hypergraph Variational Autoencoder (HeteHG-VAE) for link prediction in heterogeneous information networks (HINs). It first maps a conventional HIN to a heterogeneous hypergraph with a certain kind of semantics to capture both the high-order semantics and complex relations among nodes, while preserving the low-order pairwise topology information of the original HIN. Then, deep latent representations of nodes and hyperedges are learned by a Bayesian deep generative framework from the heterogeneous hypergraph in an unsupervised manner. Moreover, a hyperedge attention module is designed to learn the importance of different types of nodes in each hyperedge. The major merit of HeteHG-VAE lies in its ability of modeling multi-level relations in heterogeneous settings. Extensive experiments on real-world datasets demonstrate the effectiveness and efficiency of the proposed method.
Haoyi Fan, Fengbin Zhang, Yuxuan Wei, Changqing Zou, Yue Gao 0002, Qionghai Dai
IEEE Trans. Pattern Anal. Mach. Intell.7
2022 Plug-and-Play Algorithms for Video Snapshot Compressive Imaging
abstract
We consider the reconstruction problem of video snapshot compressive imaging (SCI), which captures high-speed videos using a low-speed 2D sensor (detector). The underlying principle of SCI is to modulate sequential high-speed frames with different masks and then these encoded frames are integrated into a snapshot on the sensor and thus the sensor can be of low-speed. On one hand, video SCI enjoys the advantages of low-bandwidth, low-power and low-cost. On the other hand, applying SCI to large-scale problems (HD or UHD videos) in our daily life is still challenging and one of the bottlenecks lies in the reconstruction algorithm. Existing algorithms are either too slow (iterative optimization algorithms) or not flexible to the encoding process (deep learning based end-to-end networks). In this paper, we develop fast and flexible algorithms for SCI based on the plug-and-play (PnP) framework. In addition to the PnP-ADMM method, we further propose the PnP-GAP (generalized alternating projection) algorithm with a lower computational workload. We first employ the image deep denoising priors to show that PnP can recover a UHD color video with 30 frames from a snapshot measurement. Since videos have strong temporal correlation, by employing the video deep denoising priors, we achieve a significant improvement in the results. Furthermore, we extend the proposed PnP algorithms to the color SCI system using mosaic sensors, where each pixel only captures the red, green or blue channels. A joint reconstruction and demosaicing paradigm is developed for flexible and high quality reconstruction of color video SCI systems. Extensive results on both simulation and real datasets verify the superiority of our proposed algorithm.
Xin Yuan 0002, Yang Liu 0146, Jin-Li Suo, Frédo Durand, Qionghai Dai
IEEE Trans. Pattern Anal. Mach. Intell.5
2022 GigaMVS: A Benchmark for Ultra-Large-Scale Gigapixel-Level 3D Reconstruction
abstract
Multiview stereopsis (MVS) methods, which can reconstruct both the 3D geometry and texture from multiple images, have been rapidly developed and extensively investigated from the feature engineering methods to the data-driven ones. However, there is no dataset containing both the 3D geometry of large-scale scenes and high-resolution observations of small details to benchmark the algorithms. To this end, we present GigaMVS, the first gigapixel-image-based 3D reconstruction benchmark for ultra-large-scale scenes. The gigapixel images, with both wide field-of-view and high-resolution details, can clearly observe both thePalace-scale scene structure andRelievo-scale local details. The ground-truth geometry is captured by the laser scanner, which covers ultra-large-scale scenes with an average area of 8667 m$^2$and a maximum area of 32007 m$^2$. Owing to the extremely large scale, complex occlusion, and gigapixel-level images, GigaMVS exposes problems that emerge from the poor scalability and efficiency of the existing MVS algorithms. We thoroughly investigate the state-of-the-art methods in terms of geometric and textural measurements, which point to the weakness of the existing methods and promising opportunities for future works. We believe that GigaMVS can benefit the community of 3D reconstruction and support the development of novel algorithms balancing robustness, scalability and accuracy.
Jinzhi Zhang, Shi Mao, Mengqi Ji, Zequn Chen, Xiaoyun Yuan, Qionghai Dai, Lu Fang 0001
IEEE Trans. Pattern Anal. Mach. Intell.9
2022 PaMIR: Parametric Model-Conditioned Implicit Representation for Image-Based Human Reconstruction
abstract
Modeling 3D humans accurately and robustly from a single image is very challenging, and the key for such an ill-posed problem is the 3D representation of the human models. To overcome the limitations of regular 3D representations, we propose Parametric Model-Conditioned Implicit Representation (PaMIR), which combines the parametric body model with the free-form deep implicit function. In our PaMIR-based reconstruction framework, a novel deep neural network is proposed to regularize the free-form deep implicit function using the semantic features of the parametric model, which improves the generalization ability under the scenarios of challenging poses and various clothing topologies. Moreover, a novel depth-ambiguity-aware training loss is further integrated to resolve depth ambiguities and enable successful surface detail reconstruction with imperfect body reference. Finally, we propose a body reference optimization method to improve the parametric model estimation accuracy and to enhance the consistency between the parametric model and the implicit function. With the PaMIR representation, our framework can be easily extended to multi-image input scenarios without the need of multi-camera calibration and pose synchronization. Experimental results demonstrate that our method achieves state-of-the-art performance for image-based 3D human reconstruction in the cases of challenging poses and clothing types.
Zerong Zheng, Tao Yu 0007, Yebin Liu, Qionghai Dai
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 Memory Recall: A Simple Neural Network Training Framework Against Catastrophic Forgetting
abstract
It is widely acknowledged that biological intelligence is capable of learning continually without forgetting previously learned skills. Unfortunately, it has been widely observed that many artificial intelligence techniques, especially (deep) neural network (NN)-based ones, suffer from catastrophic forgetting problem, which severely forgets previous tasks when learning a new one. How to train NNs without catastrophic forgetting, which is termed continual learning, is emerging as a frontier topic and attracting considerable research interest. Inspired by memory replay and synaptic consolidation mechanism in brain, in this article, we propose a novel and simple framework termed memory recall (MeRec) for continual learning with deep NNs. In particular, we first analyze the feature stability across tasks in NN and show that NN can yield task stable features in certain layers. Then, based on this observation, we use a memory module to keep the feature statistics (mean and std) for each learned task. Based on the memory and statistics, we show that a simple replay strategy with Gaussian distribution-based feature regeneration can recall and recover the knowledge from previous tasks. Together with the weight regularization, MeRec preserves weights learned from previous tasks. Based on this simple framework, MeRec achieved leading performance with extremely small memory budget (only two feature vectors for each class) for continual learning on CIFAR-10 and CIFAR-100 datasets, with at least 50% accuracy drop reduction after several tasks compared to previous state-of-the-art approaches.
Baosheng Zhang, Haoqian Wang, Qionghai Dai
IEEE Trans. Neural Networks Learn. Syst.6
2021 Function4D: Real-Time Human Volumetric Capture From Very Sparse Consumer RGBD Sensors
abstract
Human volumetric capture is a long-standing topic in computer vision and computer graphics. Although high-quality results can be achieved using sophisticated off-line systems, real-time human volumetric capture of complex scenarios, especially using light-weight setups, remains challenging. In this paper, we propose a human volumetric capture method that combines temporal volumetric fusion and deep implicit functions. To achieve high-quality and temporal-continuous reconstruction, we propose dynamic sliding fusion to fuse neighboring depth observations together with topology consistency. Moreover, for detailed and complete surface generation, we propose detailpreserving deep implicit functions for RGBD input which can not only preserve the geometric details on the depth inputs but also generate more plausible texturing results. Results and experiments show that our method outperforms existing methods in terms of view sparsity, generalization capacity, reconstruction quality, and run-time efficiency.
Tao Yu 0007, Zerong Zheng, Qionghai Dai, Yebin Liu
CVPR5
2021 Deep Implicit Templates for 3D Shape Representation
abstract
Deep implicit functions (DIFs), as a kind of 3D shape representation, are becoming more and more popular in the 3D vision community due to their compactness and strong representation power. However, unlike polygon mesh-based templates, it remains a challenge to reason dense correspondences or other semantic relationships across shapes represented by DIFs, which limits its applications in texture transfer, shape analysis and so on. To overcome this limitation and also make DIFs more interpretable, we propose Deep Implicit Templates, a new 3D shape representation that supports explicit correspondence reasoning in deep implicit representations. Our key idea is to formulate DIFs as conditional deformations of a template implicit function. To this end, we propose Spatial Warping LSTM, which de-composes the conditional spatial transformation into multiple point-wise transformations and guarantees generalization capability. Moreover, the training loss is carefully designed in order to achieve high reconstruction accuracy while learning a plausible template with accurate correspondences in an unsupervised manner. Experiments show that our method can not only learn a common implicit tem-plate for a collection of shapes, but also establish dense correspondences across all the shapes simultaneously with-out any supervision.
Zerong Zheng, Tao Yu 0007, Qionghai Dai, Yebin Liu
CVPR3
2021 Universal and Flexible Optical Aberration Correction Using Deep-Prior Based Deconvolution
abstract
High quality imaging usually requires bulky and expensive lenses to compensate geometric and chromatic aberrations. This poses high constraints on the optical hash or low cost applications. Although one can utilize algorithmic reconstruction to remove the artifacts of low-end lenses, the degeneration from optical aberrations is spatially varying and the computation has to trade off efficiency for performance. For example, we need to conduct patch-wise optimization or train a large set of local deep neural networks to achieve high reconstruction performance across the whole image. In this paper, we propose a PSF aware deep network, which takes the aberrant image and PSF map as input and produces the latent high quality version via incorporating deep priors, thus leading to a universal and flexible optical aberration correction method. Specifically, we pre-train a base model from a set of diverse lenses and then adapt it to a given lens by quickly refining the parameters, which largely alleviates the time and memory consumption of model learning. The approach is of high efficiency in both training and testing stages. Extensive results verify the promising applications of our proposed approach for compact low-end cameras. The code is available at https://github.com/leehsiu/UABC
Xiu Li 0003, Jin-Li Suo, Xin Yuan 0002, Qionghai Dai
ICCV5
2021 DeepMultiCap: Performance Capture of Multiple Characters Using Sparse Multiview Cameras
abstract
We propose DeepMultiCap, a novel method for multi-person performance capture using sparse multi-view cameras. Our method can capture time varying surface details without the need of using pre-scanned template models. To tackle with the serious occlusion challenge for close interacting scenes, we combine a recently proposed pixel-aligned implicit function with parametric model for robust reconstruction of the invisible surface areas. An effective attention-aware module is designed to obtain the fine-grained geometry details from multi-view images, where high-fidelity results can be generated. In addition to the spatial attention method, for video inputs, we further propose a novel temporal fusion method to alleviate the noise and temporal inconsistencies for moving character reconstruction. For quantitative evaluation, we contribute a high quality multi-person dataset, MultiHuman, which consists of 150 static scenes with different levels of occlusions and ground truth 3D human models. Experimental results demonstrate the state-of-the-art performance of our method and the well generalization to real multiview video data, which outperforms the prior works by a large margin.
Ruizhi Shao, Yuxiang Zhang 0006, Tao Yu 0007, Zerong Zheng, Qionghai Dai, Yebin Liu
ICCV6
2021 Sinusoidal Sampling Enhanced Compressive Camera for High Speed Imaging
abstract
Compressive sensing technique allows capturing fast phenomena at a much higher frame rate than the camera sensor, by recovering a frame sequence from their encoded combination. However, most conventional compressive video sensing methods limit the achieved frame rate improvement to tenfold and only support low resolution recovery. Making use of the camera's redundant spatial resolution for further frame rate improve, here we report a novel compressive video acquisition technique termed Sinusoidal Sampling Enhanced Compressive Camera (S2EC2) to encode denser frames within a snapshot. Specifically, we decompose the dense frames into groups and apply combinational coding: random codes within each group for compressive acquisition; group specific sinusoidal codes to multiplex different groups onto the high resolution sensor. The sinusoidal codes designed for these groups would shift their frequency components by different offsets in the Fourier domain and staggered the dominant frequencies of the coded measurements of these groups. Correspondingly, the reconstruction successfully separate coded measurements of different groups and recovers frames within each group. Besides, we also solve the implementation problem of insufficient gray scale spatial light modulation speed, and build a prototype achieving 2000 fps reconstruction with a 15.6 fps camera (the actual compression ratio is 0.009). The extensive experiments validate the proposed approach.
Chao Deng 0005, Yuanlong Zhang, Yifeng Mao, Jingtao Fan, Jin-Li Suo, Qionghai Dai
IEEE Trans. Pattern Anal. Mach. Intell.7
2021 SurfaceNet+: An End-to-end 3D Neural Network for Very Sparse Multi-View Stereopsis
abstract
Multi-view stereopsis (MVS) tries to recover the 3D model from 2D images. As the observations become sparser, the significant 3D information loss makes the MVS problem more challenging. Instead of only focusing on densely sampled conditions, we investigate sparse-MVS with large baseline angles since the sparser sensation is more practical and more cost-efficient. By investigating various observation sparsities, we show that the classical depth-fusion pipeline becomes powerless for the case with a larger baseline angle that worsens the photo-consistency check. As another line of the solution, we present SurfaceNet+, a volumetric method to handle the 'incompleteness' and the 'inaccuracy' problems induced by a very sparse MVS setup. Specifically, the former problem is handled by a novel volume-wise view selection approach. It owns superiority in selecting valid views while discarding invalid occluded views by considering the geometric prior. Furthermore, the latter problem is handled via a multi-scale strategy that consequently refines the recovered geometry around the region with the repeating pattern. The experiments demonstrate the tremendous performance gap between SurfaceNet+ and state-of-the-art methods in terms of precision and recall. Under the extreme sparse-MVS settings in two datasets, where existing methods can only return very few points, SurfaceNet+ still works as well as in the dense MVS setting.
Mengqi Ji, Jinzhi Zhang, Qionghai Dai, Lu Fang 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2021 Model Study of Transient Imaging With Multi-Frequency Time-of-Flight Sensors
abstract
As an emerging imaging modality, transient imaging that records the transient information of light transport has significantly shaped our understanding of scenes. In spite of the great progress made in computer vision and optical imaging fields, commonly used multi-frequency time-of-flight (ToF) sensors are still afflicted with the band-limited modulation frequency and long acquisition process. To overcome such barriers, more effective image-formation schemes and reconstruction algorithms are highly desired. In this paper, we propose a compressive transient imaging model, without any priori knowledge, by constructing a near-tight-frame based representation of the ToF imaging principle. We prove that the compressibility of sensor measurements can be presented in the Fourier domain and held in the frame, and the ToF measurements possess multi-scale characteristics. Solving the inverse problems in transient imaging with our proposed model consists of two major steps, including a compressed-sensing-based approach for full measurement recovery, which essentially reduces the capture time, and a wavelet-based transient image reconstruction framework, which realizes adaptive transient image reconstruction and achieves highly accurate reconstruction results. The compressive transient imaging model is suitable for various existing multi-frequency ToF sensors and requires no hardware modifications. Experimental results using synthetic and real online datasets demonstrate its promising performance.
Hongman Wang, Rihui Wu, Yebin Liu, Qionghai Dai
IEEE Trans. Pattern Anal. Mach. Intell.6
2021 Gated Value Network for Multilabel Classification
abstract
We introduce a gated value network (GVN) for general multilabel classification (MLC) tasks. GVN was motivated by deep value network (DVN) that directly exploits the "compatibility" metric as the learning pursuit for MLC. Meanwhile, it further improves traditional DVN on twofold. First, GVN relaxes the complex variable optimization steps in DVN inference by incorporating a feedforward predictor for straightforward multilabel prediction. Second, GVN also introduces the gating mechanism to block confounding factors from the input data that allows more precise compatibility evaluations for data and their potential multilabels. The whole GVN framework is trained in an end-to-end manner with policy gradient approaches. We show the effectiveness and generalization of GVN on diverse learning tasks, including document classification, audio tagging, and image attribute prediction.
Yimin Hou 0001, Sen Wan, Feng Bao 0002, Zhiquan Ren, Yunfeng Dong, Qionghai Dai, Yue Deng 0001
IEEE Trans. Neural Networks Learn. Syst.6
2021 Human-in-the-Loop Low-Shot Learning
abstract
We consider a human-in-the-loop scenario in the context of low-shot learning. Our approach was inspired by the fact that the viability of samples in novel categories cannot be sufficiently reflected by those limited observations. Some heterogeneous samples that are quite different from existing labeled novel data can inevitably emerge in the testing phase. To this end, we consider augmenting an uncertainty assessment module into low-shot learning system to account into the disturbance of those out-of-distribution (OOD) samples. Once detected, these OOD samples are passed to human beings for active labeling. Due to the discrete nature of this uncertainty assessment process, the whole Human-In-the-Loop Low-shot (HILL) learning framework is not end-to-end trainable. We hence revisited the learning system from the aspect of reinforcement learning and introduced the REINFORCE algorithm to optimize model parameters via policy gradient. The whole system gains noticeable improvements over existing low-shot learning approaches.
Sen Wan, Yimin Hou 0001, Feng Bao 0002, Zhiquan Ren, Yunfeng Dong, Qionghai Dai, Yue Deng 0001
IEEE Trans. Neural Networks Learn. Syst.6
2020 PANDA: A Gigapixel-Level Human-Centric Video Dataset
abstract
We present PANDA, the first gigaPixel-level humAN-centric viDeo dAtaset, for large-scale, long-term, and multi-object visual analysis. The videos in PANDA were captured by a gigapixel camera and cover real-world scenes with both wide field-of-view (~1 square kilometer area) and high-resolution details (~gigapixel-level/frame). The scenes may contain 4k head counts with over 100× scale variation. PANDA provides enriched and hierarchical ground-truth annotations, including 15,974.6k bounding boxes, 111.8k fine-grained attribute labels, 12.7k trajectories, 2.2k groups and 2.9k interactions. We benchmark the human detection and tracking tasks. Due to the vast variance of pedestrian pose, scale, occlusion and trajectory, existing approaches are challenged by both accuracy and efficiency. Given the uniqueness of PANDA with both wide FoV and high resolution, a new task of interaction-aware group detection is introduced. We design a `global-to-local zoom-in' framework, where global trajectories and local interactions are simultaneously encoded, yielding promising results. We believe PANDA will contribute to the community of artificial intelligence and praxeology by understanding human behaviors and interactions in large-scale real-world scenes. PANDA Website: http://www.panda-dataset.com.
Xiya Zhang, Yinheng Zhu, Xiaoyun Yuan, Liuyu Xiang, Zerun Wang, Guiguang Ding, David J. Brady, Qionghai Dai, Lu Fang 0001
CVPR10
2020 Plug-and-Play Algorithms for Large-Scale Snapshot Compressive Imaging
abstract
Snapshot compressive imaging (SCI) aims to capture the high-dimensional (usually 3D) images using a 2D sensor (detector) in a single snapshot. Though enjoying the advantages of low-bandwidth, low-power and low-cost, applying SCI to large-scale problems (HD or UHD videos) in our daily life is still challenging. The bottleneck lies in the reconstruction algorithms; they are either too slow (iterative optimization algorithms) or not flexible to the encoding process (deep learning based end-to-end networks). In this paper, we develop fast and flexible algorithms for SCI based on the plug-and-play (PnP) framework. In addition to the widely used PnP-ADMM method, we further propose the PnP-GAP (generalized alternating projection) algorithm with a lower computational workload and prove the {global convergence} of PnP-GAP under the SCI hardware constraints. By employing deep denoising priors, we first time show that PnP can recover a UHD color video (3840×1644×48 with PNSR above 30dB) from a snapshot 2D measurement. Extensive results on both simulation and real datasets verify the superiority of our proposed algorithm.
Xin Yuan 0002, Yang Liu 0146, Jin-Li Suo, Qionghai Dai
CVPR4
2020 Multiscale-VR: Multiscale Gigapixel 3D Panoramic Videography for Virtual Reality
abstract
Creating virtual reality (VR) content with effective imaging systems has attracted significant attention worldwide following the broad applications of VR in various fields, including entertainment, surveillance, sports, etc. However, due to the inherent trade-off between field-of-view and resolution of the imaging system as well as the prohibitive computational cost, live capturing and generating multiscale 360° 3D video content at an eye-limited resolution to provide immersive VR experiences confront significant challenges. In this work, we propose Multiscale-VR, a multiscale unstructured camera array computational imaging system for high-quality gigapixel 3D panoramic videography that creates the six-degree-of-freedom multiscale interactive VR content. The Multiscale-VR imaging system comprises scalable cylindrical-distributed global and local cameras, where global stereo cameras are stitched to cover 360° field-of-view, and unstructured local monocular cameras are adapted to the global camera for flexible high-resolution video streaming arrangement. We demonstrate that a high-quality gigapixel depth video can be faithfully reconstructed by our deep neural network-based algorithm pipeline where the global depth via stereo matching and the local depth via high-resolution RGB-guided refinement are associated. To generate the immersive 3D VR content, we present a three-layer rendering framework that includes an original layer for scene rendering, a diffusion layer for handling occlusion regions, and a dynamic layer for efficient dynamic foreground rendering. Our multiscale reconstruction architecture enables the proposed prototype system for rendering highly effective 3D, 360° gigapixel live VR video at 30 fps from the captured high-throughput multiscale video sequences. The proposed multiscale interactive VR content generation approach by using a heterogeneous camera system design, in contrast to the existing single-scale VR imaging systems with structured homogeneous cameras, will open up new avenues of research in VR and provide an unprecedented immersive experience benefiting various novel applications.
Anke Zhang, Xiaoyun Yuan, Sebastian Beetschen, Lan Xu 0003, Qionghai Dai, Lu Fang 0001
ICCP9
2020 DoubleFusion: Real-Time Capture of Human Performances with Inner Body Shapes from a Single Depth Sensor
abstract
We propose DoubleFusion, a new real-time system that combines volumetric non-rigid reconstruction with data-driven template fitting to simultaneously reconstruct detailed surface geometry, large non-rigid motion and the optimized human body shape from a single depth camera. One of the key contributions of this method is a double-layer representation consisting of a complete parametric body model inside, and a gradually fused detailed surface outside. A pre-defined node graph on the body parameterizes the non-rigid deformations near the body, and a free-form dynamically changing graph parameterizes the outer surface layer far from the body, which allows more general reconstruction. We further propose a joint motion tracking method based on the double-layer representation to enable robust and fast motion tracking performance. Moreover, the inner parametric body is optimized online and forced to fit inside the outer surface layer as well as the live depth input. Overall, our method enables increasingly denoised, detailed and complete surface reconstructions, fast motion tracking performance and plausible inner body shape reconstruction in real-time. Experiments and comparisons show improved fast motion tracking and loop closure performance on more challenging scenarios. Two extended applications including body measurement and shape retargeting show the potential of our system in terms of practical use.
Tao Yu 0007, Jianhui Zhao 0002, Zerong Zheng, Qionghai Dai, Hao Li 0015, Gerard Pons-Moll, Yebin Liu
IEEE Trans. Pattern Anal. Mach. Intell.5
2020 Weighted Convolutional Motion-Compensated Frame Rate Up-Conversion Using Deep Residual Network
abstract
Frame rate up-conversion (FRUC) usually suffers from unreliable motion vectors due to the absence of the current frame to be interpolated. In addition, since the majority of video sequences are usually compressed by various coding standards to reduce the data volume, the quality of the generated frames in the FRUC will be further impaired. To address this problem, we proposed two FRUC algorithms based on deep residual network. We first present a deep residual network for the FRUC (DRNFRUC), which consists of feature extraction, feature recursive analysis, and image restoration parts with a skip connection between the input and the output of the network. The proposed DRNFRUC takes the result of an arbitrary existing FRUC method as the input and is able to significantly reduce the edge blurring and blocking artifacts when the motion of the block is violent. In addition, we proposed a deep residual network with weighted convolutional motion compensation (DRNWCMC) for the FRUC, where the convolution operations can be embedded into the motion compensation interpolation (MCI) in any existing MCI-based FRUC method. In DRNWCMC, we first devise two convolutional neural networks corresponding to the forward and backward motion compensated frames, respectively. And then, the adaptive interpolation coefficients for motion compensation are designed as two$1\times1$convolutional kernels. Finally, the interpolation result of WCMC is fed into another convolutional neural network to further improve the performance. All the parameters involved in the DRNWCMC are trained simultaneously under the same cost function. The experimental results show that the two proposed algorithms can remarkably improve both the objective and subjective quality of the interpolated frames.
Yongbing Zhang 0002, Lixin Chen, Chenggang Yan 0001, Peiwu Qin, Xiangyang Ji, Qionghai Dai
IEEE Trans. Circuits Syst. Video Technol.6
2020 Color-Guided Depth Image Recovery With Adaptive Data Fidelity and Transferred Graph Laplacian Regularization
abstract
Depth images play an important role and are prevalently used in many computer vision and computational imaging tasks. However, due to the limitation of active sensing technology, the captured depth images in practice usually suffer from low resolution and noise, which prevents its further applications. To remedy this problem, in this paper, we first propose an adaptive data fidelity formulation to optimally generate each depth pixel from a mixture probability distribution, characterizing the similarity both in the depth map and the corresponding high-resolution guided color image. The proposed method is able to fit the distribution of the input depth signal as an optimization problem by maximizing the mixture probability. Furthermore, to promote the piecewise property that depth images exhibit, we propose a transferred graph Laplacian model as a regularization term, which is general and able to handle various depth recovery tasks such as super-resolution and denoising well. Specifically, each pixel within the recovered depth image is represented as a vertex in a graph with weights in connected edges representing the similarity between vertices. By minimizing the squared variations of the image signal, the task of depth image recovery can be converted to the problem of graph-based image filtering. Since the proposed graph Laplacian regularization model is able to fully exploit a priori information about the depth image, a much more accurate and robust estimation of the underlying depth can be obtained. Extensive experiment evaluations verify that the proposed method obtains recovered depth with higher quality in terms of both objective and subjective criteria, compared with most of the state-of-the-art methods.
Yongbing Zhang 0002, Yihui Feng, Xianming Liu 0005, Deming Zhai, Xiangyang Ji, Haoqian Wang, Qionghai Dai
IEEE Trans. Circuits Syst. Video Technol.7
2020 Cooperative Deep Reinforcement Learning for Large-Scale Traffic Grid Signal Control
abstract
Exploiting reinforcement learning (RL) for traffic congestion reduction is a frontier topic in intelligent transportation research. The difficulty in this problem stems from the inability of the RL agent simultaneously monitoring multiple signal lights when taking into account complicated traffic dynamics in different regions of a traffic system. Such challenge is even more outstanding when forming control decisions on a large-scale traffic grid, where the RL action space grows exponentially with the number of intersections within the traffic grid. In this paper, we tackle such a problem by proposing a cooperative deep reinforcement learning (Coder) framework. The intuition behind Coder is to decompose the original difficult RL task as a number of subproblems with relatively easy RL goals. Accordingly, we implement Coder with multiple regional agents and a centralized global agent. Each regional agent learns its own RL policy and value functions over a small region with limited actions. Then, the centralized global agent hierarchically aggregates RL achievements from different regional agents and forms the final Q -function over the entire large-scale traffic grid. The experimental investigations demonstrate that the proposed Coder could reduce on average 30% congestions in terms of the number of waiting vehicles during high density traffic flows in simulations.
Tian Tan 0003, Feng Bao 0002, Yue Deng 0001, Alex Jin, Qionghai Dai, Jie Wang 0006
IEEE Trans. Cybern.5
2020 Parallax Tolerant Light Field Stitching for Hand-Held Plenoptic Cameras
abstract
Light field (LF) stitching is a potential solution to improve the field of view (FOV) for hand-held plenoptic cameras. Existing LF stitching methods cannot provide accurate registration for scenes with large depth variation. In this paper, a novel LF stitching method is proposed to handle parallax in the LFs more flexibly and accurately. First, a depth layer map (DLM) is proposed to guarantee adequate feature points on each depth layer. For the regions of nondeterministic depth, superpixel layer map (SLM) is proposed based on LF spatial correlation analysis to refine the depth layer assignments. Then, DLM-SLM-based LF registration is proposed to derive the location dependent homography transforms accurately and to warp LFs to its corresponding position without parallax interference. 4D graph-cut is further applied to fuse the registration results for higher LF spatial continuity and angular continuity. Horizontal, vertical and multi-LF stitching are tested for different scenes, which demonstrates the superior performance provided by the proposed method in terms of subjective quality of the stitched LFs, epipolar plane image consistency in the stitched LF, and perspective-averaged correlation between the stitched LF and the input LFs.
Xin Jin 0002, Qionghai Dai
IEEE Trans. Image Process.3
2020 PoNA: Pose-Guided Non-Local Attention for Human Pose Transfer
abstract
Human pose transfer, which aims at transferring the appearance of a given person to a target pose, is very challenging and important in many applications. Previous work ignores the guidance of pose features or only uses local attention mechanism, leading to implausible and blurry results. We propose a new human pose transfer method using a generative adversarial network (GAN) with simplified cascaded blocks. In each block, we propose a pose-guided non-local attention (PoNA) mechanism with a long-range dependency scheme to select more important regions of image features to transfer. We also design pre-posed image-guided pose feature update and post-posed pose-guided image feature update to better utilize the pose and image features. Our network is simple, stable, and easy to train. Quantitative and qualitative results on Market-1501 and DeepFashion datasets show the efficacy and efficiency of our model. Compared with state-of-the-art methods, our model generates sharper and more realistic images with rich details, while having fewer parameters and faster speed. Furthermore, our generated images can help to alleviate data insufficiency for person re-identification.
Kun Li 0001, Yebin Liu, Yukun Lai, Qionghai Dai
IEEE Trans. Image Process.5
2020 STAT: Spatial-Temporal Attention Mechanism for Video Captioning
abstract
Video captioning refers to automatic generate natural language sentences, which summarize the video contents. Inspired by the visual attention mechanism of human beings, temporal attention mechanism has been widely used in video description to selectively focus on important frames. However, most existing methods based on temporal attention mechanism suffer from the problems of recognition error and detail missing, because temporal attention mechanism cannot further catch significant regions in frames. In order to address above problems, we propose the use of a novel spatial-temporal attention mechanism (STAT) within an encoder-decoder neural network for video captioning. The proposed STAT successfully takes into account both the spatial and temporal structures in a video, so it makes the decoder to automatically select the significant regions in the most relevant temporal segments for word prediction. We evaluate our STAT on two well-known benchmarks: MSVD and MSR-VTT-10K. Experimental results show that our proposed STAT achieves the state-of-the-art performance with several popular evaluation metrics: BLEU-4, METEOR, and CIDEr.
Chenggang Yan 0001, Yunbin Tu, Xingzheng Wang, Yongbing Zhang 0002, Xinhong Hao, Yongdong Zhang 0001, Qionghai Dai
IEEE Trans. Multim.7
2020 Corrections to "STAT: Spatial-Temporal Attention Mechanism for Video Captioning"
abstract
Presents corrections to affiliations in the above named paper.
Chenggang Yan 0001, Yunbin Tu, Xingzheng Wang, Yongbing Zhang 0002, Xinhong Hao, Yongdong Zhang 0001, Qionghai Dai
IEEE Trans. Multim.7
2020 Learning Deep Landmarks for Imbalanced Classification
abstract
We introduce a deep imbalanced learning framework called learning DEep Landmarks in laTent spAce (DELTA). Our work is inspired by the shallow imbalanced learning approaches to rebalance imbalanced samples before feeding them to train a discriminative classifier. Our DELTA advances existing works by introducing the new concept of rebalancing samples in a deeply transformed latent space, where latent points exhibit several desired properties including compactness and separability. In general, DELTA simultaneously conducts feature learning, sample rebalancing, and discriminative learning in a joint, end-to-end framework. The framework is readily integrated with other sophisticated learning concepts including latent points oversampling and ensemble learning. More importantly, DELTA offers the possibility to conduct imbalanced learning with the assistancy of structured feature extractor. We verify the effectiveness of DELTA not only on several benchmark data sets but also on more challenging real-world tasks including click-through-rate (CTR) prediction, multi-class cell type classification, and sentiment analysis with sequential inputs.
Feng Bao 0002, Yue Deng 0001, Youyong Kong, Zhiquan Ren, Jin-Li Suo, Qionghai Dai
IEEE Trans. Neural Networks Learn. Syst.6
2019 Dual-View Ranking with Hardness Assessment for Zero-Shot Learning
abstract
Zero-shot learning (ZSL) is to build recognition models for previously unseen target classes which have no labeled data for training by transferring knowledge from some other related auxiliary source classes with abundant labeled samples to the target ones with class attributes as the bridge. The key is to learn a similarity based ranking function between samples and class labels using the labeled source classes so that the proper (unseen) class label for a test sample can be identified by the function. In order to learn the function, single-view ranking based loss is widely used which aims to rank the true label prior to the other labels for a training sample. However, we argue that the ranking can be performed from the other view, which aims to place the images belonging to a label before the images from the other classes. Motivated by it, we propose a novel DuAl-view RanKing (DARK) loss for zeroshot learning simultaneously ranking labels for an image by point-to-point metric and ranking images for a label by pointto-set metric, which is capable of better modeling the relationship between images and classes. In addition, we also notice that previous ZSL approaches mostly fail to well exploit the hardness of training samples, either using only very hard ones or using all samples indiscriminately. In this work, we also introduce a sample hardness assessment method to ZSL which assigns different weights to training samples based on their hardness, which leads to a more accurate and robust ZSL model. Experiments on benchmarks demonstrate that DARK outperforms the state-of-the-arts for (generalized) ZSL.
Guiguang Ding, Jungong Han, Xiaohan Ding, Sicheng Zhao, Zheng Wang 0001, Chenggang Yan 0001, Qionghai Dai
AAAI8
2019 SimulCap : Single-View Human Performance Capture With Cloth Simulation
abstract
This paper proposes a new method for live free-viewpoint human performance capture with dynamic details (e.g., cloth wrinkles) using a single RGBD camera. Our main contributions are: (i) a multi-layer representation of garments and body, and (ii) a physics-based performance capture procedure. We first digitize the performer using multi-layer surface representation, which includes the undressed body surface and separate clothing meshes. For performance capture, we perform skeleton tracking, cloth simulation, and iterative depth fitting sequentially for the incoming frame. By incorporating cloth simulation into the performance capture pipeline, we can simulate plausible cloth dynamics and cloth-body interactions even in the occluded regions, which was not possible in previous capture methods. Moreover, by formulating depth fitting as a physical process, our system produces cloth tracking results consistent with the depth observation while still maintaining physical constraints. Results and evaluations show the effectiveness of our method. Our method also enables new types of applications such as cloth retargeting, free-viewpoint video rendering and animations.
Tao Yu 0007, Zerong Zheng, Jianhui Zhao 0002, Qionghai Dai, Gerard Pons-Moll, Yebin Liu
CVPR5
2019 Light Field Image Compression Using Depth-based CNN in Intra Prediction
abstract
Recently, light field images have received extensive attention due to their potential applications. Since they take up a huge memory because of its super-high resolution, efficient compression methods are fundamentally required. In this paper, we propose a novel intra prediction mode by using depth-adaptive convolutional neuro network (DCNN). Light field projection finds the imaging response distribution for each object point using the depth estimated from each macropixel in the light field image. The highly correlated imaging responses are used to select the neural network structure. The network structure also adapts to the to-be-encoded block size. Adding the proposed DCNN-based prediction mode into the rate-distortion optimization loop with other 35 intra prediction modes of HEVC, the proposed encoding scheme achieves a significant bit-rate saving compared to representative compression approaches with limited computational complexity increment. Statistical data are also provided and analyzed to demonstrate the efficiency of the proposed method.
Tingting Zhong, Xin Jin 0002, Lingjun Li, Qionghai Dai
ICASSP4
2019 DeepHuman: 3D Human Reconstruction From a Single Image
abstract
We propose DeepHuman, an image-guided volume-to-volume translation CNN for 3D human reconstruction from a single RGB image. To reduce the ambiguities associated with the reconstruction of invisible areas, our method leverages a dense semantic representation generated from SMPL model as an additional input. One key feature of our network is that it fuses different scales of image features into the 3D space through volumetric feature transformation, which helps to recover accurate surface geometry. The surface details are further refined through a normal refinement network, which can be concatenated with the volume generation network using our proposed volumetric normal projection layer. We also contribute THuman, a 3D real-world human model dataset containing approximately 7000 models. The network is trained using training data generated from the dataset. Overall, due to the specific design of our network and the diversity in our dataset, our method enables 3D human model estimation given only a single image and outperforms state-of-the-art approaches.
Zerong Zheng, Tao Yu 0007, Yixuan Wei, Qionghai Dai, Yebin Liu
ICCV4
2019 Light Field Stitching Targeting Focal Length Inconsistency
abstract
Focal length inconsistency is a common problem in capturing multiple light fields (LFs), which reduces the accuracy in feature matching and LF registration, and limits the quality of LF stitching. In this paper, a novel method is proposed to stitch the LFs captured at different focal length accurately. First, refocus matching is proposed to find the best-matched-focused regions in the focal stack to improve the quality of feature matching. Then, depth distances and refocusing features are exploited to register central sub-aperture images (CSI) to handle focal length variation. Finally, CSI-based 4D warping is applied to preserve the LF angular consistency. Experimental results show that the proposed method outperforms the existing approaches obviously in terms of the stitched LFs and the coherence in epipolar plane image (EPI), especially for the LFs with the focal length inconsistency.
Xin Jin 0002, Qionghai Dai
ICIP3
2019 F-Number Adaptation for Maximizing the Sensor Usage of Light Field Cameras
abstract
Since the effective imaging area in micro lens array (MLA) of the light field camera cannot completely cover the MLA plane, the imaging response of MLA cannot fully cover the sensor plane, which decreases the sensor usage. Hence, an f-number adaptation model is proposed to maximize the sensor usage without changing the structure of light field cameras. By deriving the relationship between the main lens f-number, the macropixel diameter and the sensor usage, a constrained concave optimization model is proposed. A pixel extraction method is also proposed to make a full use of the increased effective pixels. The experimental results demonstrate the effectiveness of the proposed model in terms of the effective pixel number, the subaperture number and improving 3D reconstruction quality. It is also robust to different light field camera architectures and different arrangements of MLA.
Chuanpu Li, Xin Jin 0002, Junke Li, Qionghai Dai
ICME4
2019 Blind Calibration for Focused Plenoptic Cameras
abstract
Because of the subtle structure of focused plenoptic camera, the exact geometry parameters cannot be retrieved, which leads to inaccurate f-number matching or refocusing errors in light field processing. In this paper, a novel blind calibration method is proposed to calculate the geometry parameters for the focused plenoptic cameras. The blind calibration model is derived based on the geometry projection analysis to establish the relationship between the patch-size of each micro-image used in subaperture image rendering and geometry parameters. A triple-level calibration board is designed to realize calibration via single shot. A gradient-SSIM-based fractional-pixel matching is proposed to retrieve the precise rendering patch-size for the calibration model. Experimental results demonstrate that the proposed method is robust to different focused plenoptic cameras and can get the geometry parameters with high accuracy.
Xufu Sun, Xin Jin 0002, Yanqin Chen, Qionghai Dai
ICME5
2019 Real-time Indoor Scene Reconstruction with RGBD and Inertial Input
abstract
Camera motion estimation is a key technique for 3D scene reconstruction. Previous works usually assume slow camera motions, which limit the usage in many real cases. We propose an end-to-end 3D reconstruction system which combines color, depth and inertial measurements to achieve robust reconstruction with fast sensor motions. Our framework utilizes extended Kalman filter to fuse the three kinds of information and involve an iterative method to jointly optimize feature correspondences, camera poses and scene geometry. We also propose a novel geometry-aware patch deformation technique to adapt the feature appearance in image domain, leading to a more accurate feature matching under fast camera motions. Experiments show that our patch deformation method improves the accuracy of feature tracking, and our 3D reconstruction framework outperforms the state-of-the-art solutions under fast camera motions.
Zunjie Zhu, Feng Xu 0005, Chenggang Yan 0001, Xinhong Hao, Xiangyang Ji, Yongdong Zhang 0001, Qionghai Dai
ICME7
2019 Zero-shot Learning with Many Classes by High-rank Deep Embedding Networks
abstract
Zero-shot learning (ZSL) is a recently emerging research topic which aims to build classification models for unseen classes with knowledge from auxiliary seen classes. Though many ZSL works have shown promising results on small-scale datasets by utilizing a bilinear compatibility function, the ZSL performance on large-scale datasets with many classes (say, ImageNet) is still unsatisfactory. We argue that the bilinear compatibility function is a low-rank approximation of the true compatibility function such that it is not expressive enough especially when there are a large number of classes because of the rank limitation. To address this issue, we propose a novel approach, termed as High-rank Deep Embedding Networks (GREEN), for ZSL with many classes. In particular, we propose a feature-dependent mixture of softmaxes as the image-class compatibility function, which is a simple extension of the bilinear compatibility function, but yields much better results. It utilizes a mixture of non-linear transformations with feature-dependent latent variables to approximate the true function in a high-rank way, which makes GREEN more expressive. Experiments on several datasets including ImageNet demonstrate GREEN significantly outperforms the state-of-the-art approaches.
Guiguang Ding, Jungong Han, Qionghai Dai
IJCAI6
2019 Landmark Selection for Zero-shot Learning
abstract
Zero-shot learning (ZSL) is an emerging research topic whose goal is to build recognition models for previously unseen classes. The basic idea of ZSL is based on heterogeneous feature matching which learns a compatibility function between image and class features using seen classes. The function is constructed based on one-vs-all training in which each class has only one class feature and many image features. Existing ZSL works mostly treat all image features equivalently. However, in this paper we argue that it is more reasonable to use some representative cross-domain data instead of all. Motivated by this idea, we propose a novel approach, termed as Landmark Selection(LAST) for ZSL. LAST is able to identify representative cross-domain features which further lead to better image-class compatibility function. Experiments on several ZSL datasets including ImageNet demonstrate the superiority of LAST to the state-of-the-arts.
Guiguang Ding, Jungong Han, Chenggang Yan 0001, Jiyong Zhang 0001, Qionghai Dai
IJCAI6
2019 Live Demonstration: 4-DoF Parallax Tolerant Light Field Stitching
abstract
This demonstration shows a 4-Degree-of-Freedom (4-DoF) parallax tolerant light field (LF) stitching system to generate large field of view (FoV) LF with interactive LF experiences. It can capture the LFs in real-time, generate the large FoV LF with high visual quality and visualize the advantage of large FoV LF by interactive 3D applications like 4-DoF stereo panorama, digital refocusing and 3D point cloud reconstruction.
Xin Jin 0002, Qionghai Dai
ISCAS3
2019 Real-time indoor scene reconstruction with Manhattan assumption
Zunjie Zhu, Feng Xu 0005, Chenggang Yan 0001, Bingjian Gong, Yongdong Zhang 0001, Qionghai Dai
Multim. Tools Appl.7
2019 Rank Minimization for Snapshot Compressive Imaging
abstract
Snapshot compressive imaging (SCI) refers to compressive imaging systems where multiple frames are mapped into a single measurement, with video compressive imaging and hyperspectral compressive imaging as two representative applications. Though exciting results of high-speed videos and hyperspectral images have been demonstrated, the poor reconstruction quality precludes SCI from wide applications. This paper aims to boost the reconstruction quality of SCI via exploiting the high-dimensional structure in the desired signal. We build a joint model to integrate the nonlocal self-similarity of video/hyperspectral frames and the rank minimization approach with the SCI sensing process. Following this, an alternating minimization algorithm is developed to solve this non-convex problem. We further investigate the special structure of the sampling process in SCI to tackle the computational workload and memory issues in SCI reconstruction. Both simulation and real data (captured by four different SCI cameras) results demonstrate that our proposed algorithm leads to significant improvements compared with current state-of-the-art algorithms. We hope our results will encourage the researchers and engineers to pursue further in compressive imaging for real applications.
Yang Liu 0146, Xin Yuan 0002, Jin-Li Suo, David J. Brady, Qionghai Dai
IEEE Trans. Pattern Anal. Mach. Intell.5
2019 Light Field Reconstruction Using Convolutional Network on EPI and Extended Applications
abstract
In this paper, a novel convolutional neural network (CNN)-based framework is developed for light field reconstruction from a sparse set of views. We indicate that the reconstruction can be efficiently modeled as angular restoration on an epipolar plane image (EPI). The main problem in direct reconstruction on the EPI involves an information asymmetry between the spatial and angular dimensions, where the detailed portion in the angular dimensions is damaged by undersampling. Directly upsampling or super-resolving the light field in the angular dimensions causes ghosting effects. To suppress these ghosting effects, we contribute a novel "blur-restoration-deblur" framework. First, the "blur" step is applied to extract the low-frequency components of the light field in the spatial dimensions by convolving each EPI slice with a selected blur kernel. Then, the "restoration" step is implemented by a CNN, which is trained to restore the angular details of the EPI. Finally, we use a non-blind "deblur" operation to recover the spatial high frequencies suppressed by the EPI blur. We evaluate our approach on several datasets, including synthetic scenes, real-world scenes and challenging microscope light field data. We demonstrate the high performance and robustness of the proposed framework compared with state-of-the-art algorithms. We further show extended applications, including depth enhancement and interpolation for unstructured input. More importantly, a novel rendering approach is presented by combining the proposed framework and depth information to handle large disparities.
Gaochang Wu, Yebin Liu, Lu Fang 0001, Qionghai Dai, Tianyou Chai
IEEE Trans. Pattern Anal. Mach. Intell.4
2019 DECODE: Deep Confidence Network for Robust Image Classification
abstract
Recent years have witnessed the success of deep convolutional neural networks for image classification and many related tasks. It should be pointed out that the existing training strategies assume that there is a clean dataset for model learning. In elaborately constructed benchmark datasets, deep network has yielded promising performance under the assumption. However, in real-world applications, it is burdensome and expensive to collect sufficient clean training samples. On the other hand, collecting noisy labeled samples is very economical and practical, especially with the rapidly increasing amount of visual data in the web. Unfortunately, the accuracy of current deep models may drop dramatically even with 5%-10% label noise. Therefore, enabling label noise resistant classification has become a crucial issue in the data driven deep learning approaches. In this paper, we propose a DEep COnfiDEnce network (DECODE) to address this issue. In particular, based on the distribution of mislabeled data, we adopt a confidence evaluation module that is able to determine the confidence that a sample is mislabeled. With the confidence, we further use a weighting strategy to assign different weights to different samples so that the model pays less attention to low confidence data, which is more likely to be noise. In this way, the deep model is more robust to label noise. DECODE is designed to be general, such that it can be easily combined with existing studies. We conduct extensive experiments on several datasets, and the results validate that DECODE can improve the accuracy of deep models trained with noisy data.
Guiguang Ding, Kai Chen 0044, Chaoqun Chu, Jungong Han, Qionghai Dai
IEEE Trans. Image Process.6
2019 Learning Sheared EPI Structure for Light Field Reconstruction
abstract
Research in light field reconstruction focuses on synthesizing novel views with the assistance of depth information. In this paper, we present a learning-based light field reconstruction approach by fusing a set of sheared epipolar plane images (EPIs). We start by showing that a patch in a sheared EPI will exhibit a clear structure when the sheared value equals the depth of that patch. By taking advantage of this pattern, a convolutional neural network (CNN) is then trained to evaluate the sheared EPIs, and output a reference score for fusing the sheared EPIs. The proposed CNN is elaborately designed to learn the similarity degree between the input sheared EPI and the ground truth EPI. Therefore, no depth information is required for network training and reasoning. We demonstrate the high performance of the proposed method through evaluations on synthetic scenes, real-world scenes, and challenging microscope light fields. We also show a further application of our proposed network for depth inference.
Gaochang Wu, Yebin Liu, Qionghai Dai, Tianyou Chai
IEEE Trans. Image Process.3
2019 Cross-Modality Bridging and Knowledge Transferring for Image Understanding
abstract
The understanding of web images has been a hot research topic in both artificial intelligence and multimedia content analysis domains. The web images are composed of various complex foregrounds and backgrounds, which makes the design of an accurate and robust learning algorithm a challenging task. To solve the above significant problem, first, we learn a cross-modality bridging dictionary for the deep and complete understanding of a vast quantity of web images. The proposed algorithm leverages the visual features into the semantic concept probability distribution, which can construct a global semantic description for images while preserving the local geometric structure. To discover and model the occurrence patterns between intra- and inter-categories, multi-task learning is introduced for formulating the objective formulation with Capped-ℓ1penalty, which can obtain the optimal solution with a higher probability and outperform the traditional convex function-based methods. Second, we propose a knowledge-based concept transferring algorithm to discover the underlying relations of different categories. This distribution probability transferring among categories can bring the more robust global feature representation, and enable the image semantic representation to generalize better as the scenario becomes larger. Experimental comparisons and performance discussion with classical methods on the ImageNet, Caltech-256, SUN397, and Scene15 datasets show the effectiveness of our proposed method at three traditional image understanding tasks.
Chenggang Yan 0001, Liang Li 0003, Chunjie Zhang 0001, Bingtao Liu, Yongdong Zhang 0001, Qionghai Dai
IEEE Trans. Multim.6
2019 Collaborative Representation Cascade for Single-Image Super-Resolution
abstract
Most recent learning-based single-image superresolution methods first interpolate the low-resolution (LR) input, from which overlapped LR features are then extracted to reconstruct their high-resolution (HR) counterparts and the final HR image. However, most of them neglect to take advantage of the intermediate recovered HR image to enhance image quality further. We conduct principal component analysis (PCA) to reduce LR feature dimension. Then we find that the number of principal components after conducting PCA in the LR feature space from the reconstructed images is larger than that from the interpolated images by using bicubic interpolation. Based on this observation, we present an unsophisticated yet effective framework named collaborative representation cascade (CRC) that learns multilayer mapping models between LR and HR feature pairs. In particular, we extract the features from the intermediate recovered image to upscale and enhance LR input progressively. In the learning phase, for each cascade layer, we use the intermediate recovered results and their original HR counterparts to learn single-layer mapping model. Then, we use this single-layer mapping model to super-resolve the original LR inputs. And the intermediate HR outputs are regarded as training inputs for the next cascade layer, until we obtain multilayer mapping models. In the reconstruction phase, we extract multiple sets of LR features from the LR image and intermediate recovered. Then, in each cascade layer, mapping model is utilized to pursue HR image. Our experiments on several commonly used image SR testing datasets show that our proposed CRC method achieves state-of-the-art image SR results.
Yongbing Zhang 0002, Yulun Zhang 0001, Jian Zhang 0018, Dong Xu 0001, Yun Fu 0001, Yisen Wang 0001, Xiangyang Ji, Qionghai Dai
IEEE Trans. Syst. Man Cybern. Syst.8
2018 A PID Controller Approach for Stochastic Optimization of Deep Networks
abstract
Deep neural networks have demonstrated their power in many computer vision applications. State-of-the-art deep architectures such as VGG, ResNet, and DenseNet are mostly optimized by the SGD-Momentum algorithm, which updates the weights by considering their past and current gradients. Nonetheless, SGD-Momentum suffers from the overshoot problem, which hinders the convergence of network training. Inspired by the prominent success of proportional-integral-derivative (PID) controller in automatic control, we propose a PID approach for accelerating deep network optimization. We first reveal the intrinsic connections between SGD-Momentum and PID based controller, then present the optimization algorithm which exploits the past, current, and change of gradients to update the network parameters. The proposed PID method reduces much the overshoot phenomena of SGD-Momentum, and it achieves up to 50% acceleration on popular deep network architectures with competitive accuracy, as verified by our experiments on the benchmark datasets including CIFAR10, CIFAR100, and Tiny-ImageNet.
Wangpeng An, Haoqian Wang, Qingyun Sun, Jun Xu 0019, Qionghai Dai, Lei Zhang 0006
CVPR5
2018 DoubleFusion: Real-Time Capture of Human Performances With Inner Body Shapes From a Single Depth Sensor
abstract
We propose DoubleFusion, a new real-time system that combines volumetric dynamic reconstruction with data-driven template fitting to simultaneously reconstruct detailed geometry, non-rigid motion and the inner human body shape from a single depth camera. One of the key contributions of this method is a double layer representation consisting of a complete parametric body shape inside, and a gradually fused outer surface layer. A pre-defined node graph on the body surface parameterizes the non-rigid deformations near the body, and a free-form dynamically changing graph parameterizes the outer surface layer far from the body, which allows more general reconstruction. We further propose a joint motion tracking method based on the double layer representation to enable robust and fast motion tracking performance. Moreover, the inner body shape is optimized online and forced to fit inside the outer surface layer. Overall, our method enables increasingly denoised, detailed and complete surface reconstructions, fast motion tracking performance and plausible inner body shape reconstruction in real-time. In particular, experiments show improved fast motion tracking and loop closure performance on more challenging scenarios.
Tao Yu 0007, Zerong Zheng, Jianhui Zhao 0002, Qionghai Dai, Hao Li 0015, Gerard Pons-Moll, Yebin Liu
CVPR5
2018 HybridFusion: Real-Time Performance Capture Using a Single Depth Sensor and Sparse IMUs
Zerong Zheng, Tao Yu 0007, Hao Li 0015, Qionghai Dai, Lu Fang 0001, Yebin Liu
ECCV (9)5
2018 High-Speed Light Field Image Formation Analysis Using Wavefield Modeling with Flexible Sampling
abstract
Understanding the image formation inside plenoptic cameras is significant for the investigations of improving the low spatial resolution. Most researches explore the image formation from the perspective of geometric optics. However, as the hardware components in combination with low-aperture optical systems become smaller and smaller, geometric analysis will no longer be valid due to diffraction effects. In this paper, a wave-optic-based model is proposed that uses the Fresnel diffraction equation to propagate the whole object field into the plenoptic systems. The proposed model employs averaging of intensities on the sensor from uncorrelated coherent wave to avoid interference during propagations among the optical component planes. Besides, by utilizing the method of multiple partial propagations, the proposed model is much flexible at sampling on propagation planes. In order to verify the effectiveness of the proposed model, numerical simulations are conducted by comparing with existing wave optic model under different optical configurations of plenoptic cameras. Results demonstrate that the proposed model can describe the light field image formation properly. In addition, the time for image formation has been reduced by a factor of 19.22 using the proposed model.
Yanqin Chen, Xin Jin 0002, Qionghai Dai
ICASSP4
2018 Light Field Stitching for Parallax Tolerance
abstract
In this paper, a novel light field (LF) stitching method is proposed to handle parallax more flexibly and accurately. The depth guided feature point filtering and depth-based motion model are proposed to warp the 4D meshes in the LFs adapting to the relative depth layer distance between the mesh center and the feature points. Also, 4D graph-cut is applied in LF fusion to reduce ghosting effects and produce better refocusing effects. Experimental results show that the proposed method works well without producing visual artifacts like ghosting or misalignments. It outperforms the existing approaches obviously in terms of the subjective quality of the stitched LFs, RMSE and the coherence in epipolar plane image (EPI).
Xin Jin 0002, Chuanpu Li, Yanqin Chen, Qionghai Dai
ICIP5
2018 Fast, Robust, and Accurate Image Denoising via Very Deeply Cascaded Residual Networks
abstract
Patch based image modelings have shown great potential in image denoising. They mainly exploit the nonlocal self-similarity (NSS) of either input degraded images or clean natural ones when training models, while failing to learn the mappings between them. More seriously, these algorithms have very high time complexity and poor robustness when handling images with different noise variances and resolutions. To address these problems, in this paper, we propose very deeply cascaded residual networks (VDCRN) to build the precise relationships between the noisy images and their corresponding noise-free ones. It adopts a new residual unit with an identity skip connection (shortcut) to make training easy and improve generalization. The introduction of shortcut is helpful to avoid the problem of gradient vanishing and preserve more image details. By cascading three such residual units, we build the VDCRN to deploy deeper and larger convolutional networks. Based on such a residual network, our VDCRN achieves very fast speed and good robustness. Experimental results demonstrate that our model outperforms a lot of state-of-the-art denoising algorithms quantitively and qualitively.
Yongbing Zhang 0002, Xingzheng Wang, Haoqian Wang, Qionghai Dai
MMSP5
2018 Generating VR Live Videos with Tripod Panoramic Rig
abstract
Recent breakthrough in consumer-level virtual reality (VR) devices brings an increasing demand of VR live content. As converting real life content into VR need complex computations, current techniques can not synthesize 360° 3D VR content with high performance, not to mention real time. We propose an end-to-end system that records a scene using a tripod panoramic rig and broadcasts 360° stereo panorama videos in real time. The system performs a panorama stitching technique which pre-compute 3 stitching seam candidates for dynamic seam switching in the live broadcasting. This technique achieves high frame rates (>30fps) with minimum foreground cutoff and temporal jittering artifacts. Stereo vision quality is also better preserved by a proposed weighting-based image alignment scheme. We demonstrate the effectiveness of our approach on a variety of videos delivering live events. And our system has been successfully used in broadcasting live shows to mobile phone users on a professional live broadcasting platform with about 390 million user visits per month.
Feng Xu 0005, Bicheng Luo, Qionghai Dai
VR4
2018 Probabilistic natural mapping of gene-level tests for genome-wide association studies
abstract
Genome-wide association studies (GWASs) generally focus on a single marker, which limits the elucidation of the genetic architecture of complex traits. Herein, we present a new computational framework, termed probabilistic natural mapping (PALM), for performing gene-level association tests. PALM robustly reveals the inherent genomic structures of genes and generates feature representations that can be seamlessly incorporated into conventional statistic tests. Our approach substantially improves the effectiveness of uncovering associations derived from a subgroup of variants with weak effects, which represents a known challenge associated with existing methods. We applied PALM in a gastric cancer GWAS and identified two additional gastric cancer-associated susceptibility genes, NOC3L and RUNDC2A. The robust susceptibility discoveries of PALM are widely supported by existing studies from other biological perspectives. PALM will be useful for further GWAS analytical strategies that use gene-level analyses.
Feng Bao 0002, Yue Deng 0001, Mulong Du, Zhiquan Ren, Yanyu Zhao, Jin-Li Suo, Meilin Wang, Qionghai Dai
Briefings Bioinform.10
2018 Predicting Model and Algorithm in RNA Folding Structure Including Pseudoknots
abstract
The prediction of RNA structure with pseudoknots is a nondeterministic polynomial-time hard (NP-hard) problem; according to minimum free energy models and computational methods, we investigate the RNA-pseudoknotted structure. Our paper presents an efficient algorithm for predicting RNA structure with pseudoknots, and the algorithm takes O([Formula: see text]) time and O([Formula: see text]) space, the experimental tests in Rfam10.1 and PseudoBase indicate that the algorithm is more effective and precise. The predicting accuracy, the time complexity and space complexity outperform existing algorithms, such as Maximum Weight Matching (MWM) algorithm, PKNOTS algorithm and Inner Limiting Layer (ILM) algorithm, and the algorithm can predict arbitrary pseudoknots. And there exists a [Formula: see text] ([Formula: see text]) polynomial time approximation scheme in searching maximum number of stackings, and we give the proof of the approximation scheme in RNA-pseudoknotted structure. We have improved several types of pseudoknots considered in RNA folding structure, and analyze their possible transitions between types of pseudoknots.
Daming Zhu, Qionghai Dai
Int. J. Pattern Recognit. Artif. Intell.3
2018 ACID: Association Correction for Imbalanced Data in GWAS
abstract
Genome-wide association study (GWAS) has been widely witnessed as a powerful tool for revealing suspicious loci from various diseases. However, real world GWAS tasks always suffer from the data imbalance problem of sufficient control samples and limited case samples. This imbalance issue can cause serious biases to the result and thus leads to losses of significance for true causal markers. To tackle this problem, we proposed a computational framework to perform association correction for imbalanced data (ACID) that could potentially improve the performance of GWAS under the imbalance condition. ACID is inspired by the imbalance learning theory but is particularly modified to address the task of association discovery from sequential genomic data. Simulation studies demonstrate ACID can dramatically improve the power of traditional GWAS method on the dataset with severe imbalances. We further applied ACID to two imbalanced datasets (gastric cancer and bladder cancer) to conduct genome wide association analysis. Experimental results indicate that our method has better abilities in identifying suspicious loci than the regression approach and shows consistencies with existing discoveries.
Feng Bao 0002, Yue Deng 0001, Qionghai Dai
IEEE ACM Trans. Comput. Biol. Bioinform.3
2018 Convolutional Sparse Coding for RGB+NIR Imaging
abstract
Emerging sensor designs increasingly rely on novel color filter arrays (CFAs) to sample the incident spectrum in unconventional ways. In particular, capturing a near-infrared (NIR) channel along with conventional RGB color is an exciting new imaging modality. RGB+NIR sensing has broad applications in computational photography, such as low-light denoising, it has applications in computer vision, such as facial recognition and tracking, and it paves the way toward low-cost single-sensor RGB and depth imaging using structured illumination. However, cost-effective commercial CFAs suffer from severe spectral cross talk. This cross talk represents a major challenge in high-quality RGB+NIR imaging, rendering existing spatially multiplexed sensor designs impractical. In this work, we introduce a new approach to RGB+NIR image reconstruction using learned convolutional sparse priors. We demonstrate high-quality color and NIR imaging for challenging scenes, even including high-frequency structured NIR illumination. The effectiveness of the proposed method is validated on a large data set of experimental captures, and simulated benchmark results which demonstrate that this work achieves unprecedented reconstruction quality.
Felix Heide, Qionghai Dai, Gordon Wetzstein
IEEE Trans. Image Process.3
2018 Plenoptic Image Coding Using Macropixel-Based Intra Prediction
abstract
The plenoptic image in a super high resolution is composed of a number of macropixels recording both spatial and angular light radiance. Based on the analysis of spatial correlations of macropixel structure, this paper proposes a macropixel-based intra prediction method for plenoptic image coding. After applying an invertible image reshaping method to the plenoptic image, the macropixel structures are aligned with the coding unit grids of a block-based video coding standard. The reshaped and regularized image is compressed by the video encoder comprising the proposed macropixel-based intra prediction, which includes three modes: multi-block weighted prediction mode (MWP), co-located single-block prediction mode (CSP), and boundary matching based prediction mode (BMP). In the MWP mode and BMP mode, the predictions are generated by minimizing spatial Euclidean distance and boundary error among the reference samples, respectively, which can fully exploit spatial correlations among the pixels beneath the neighboring microlens. The proposed approach outperforms HEVC by an average of 47.0% bitrate reduction. Compared with other state-of-the-art methods, like pseudo-video based on tiling and arrangement method (PVTA), intra block copy (IBC) mode, and locally linear embedding (LLE) based prediction, it can also achieve 45.0%, 27.7% and 22.7% bitrate savings on average, respectively.
Xin Jin 0002, Haixu Han, Qionghai Dai
IEEE Trans. Image Process.3
2018 Residual Highway Convolutional Neural Networks for in-loop Filtering in HEVC
abstract
High efficiency video coding (HEVC) standard achieves half bit-rate reduction while keeping the same quality compared with AVC. However, it still cannot satisfy the demand of higher quality in real applications, especially at low bit rates. To further improve the quality of reconstructed frame while reducing the bitrates, a residual highway convolutional neural network (RHCNN) is proposed in this paper for in-loop filtering in HEVC. The RHCNN is composed of several residual highway units and convolutional layers. In the highway units, there are some paths that could allow unimpeded information across several layers. Moreover, there also exists one identity skip connection (shortcut) from the beginning to the end, which is followed by one small convolutional layer. Without conflicting with deblocking filter (DF) and sample adaptive offset (SAO) filter in HEVC, RHCNN is employed as a high-dimension filter following DF and SAO to enhance the quality of reconstructed frames. To facilitate the real application, we apply the proposed method to I frame, P frame, and B frame, respectively. For obtaining better performance, the entire quantization parameter (QP) range is divided into several QP bands, where a dedicated RHCNN is trained for each QP band. Furthermore, we adopt a progressive training scheme for the RHCNN where the QP band with lower value is used for early training and their weights are used as initial weights for QP band of higher values in a progressive manner. Experimental results demonstrate that the proposed method is able to not only raise the PSNR of reconstructed frame but also prominently reduce the bit-rate compared with HEVC reference software.
Yongbing Zhang 0002, Xiangyang Ji, Yun Zhang 0002, Ruiqin Xiong, Qionghai Dai
IEEE Trans. Image Process.6
2018 Adaptive Residual Networks for High-Quality Image Restoration
abstract
Image restoration methods based on convolutional neural networks have shown great success in the literature. However, since most of networks are not deep enough, there is still some room for the performance improvement. On the other hand, though some models are deep and introduce shortcuts for easy training, they ignore the importance of location and scaling of different inputs within the shortcuts. As a result, existing networks can only handle one specific image restoration application. To address such problems, we propose a novel adaptive residual network (ARN) for high-quality image restoration in this paper. Our ARN is a deep residual network, which is composed of convolutional layers, parametric rectified linear unit layers, and some adaptive shortcuts. We assign different scaling parameters to different inputs of the shortcuts, where the scaling is considered as part parameters of the ARN and trained adaptively according to different applications. Due to the special construction of ARN, it can solve many image restoration problems and have superior performance. We demonstrate its capabilities with three representative applications, including Gaussian image denoising, single image super resolution, and JPEG image deblocking. Experimental results prove that our model greatly outperforms numerous state-of-the-art restoration methods in terms of both peak signal-to-noise ratio and structure similarity index metrics, e.g., it achieves 0.2-0.3 dB gain in average compared with the second best method at a wide range of situations.
Yongbing Zhang 0002, Chenggang Yan 0001, Xiangyang Ji, Qionghai Dai
IEEE Trans. Image Process.5
2018 Effective Uyghur Language Text Detection in Complex Background Images for Traffic Prompt Identification
abstract
Text detection in complex background images is a challenging task for intelligent vehicles. Actually, almost all the widely-used systems focus on commonly used languages while for some minority languages, such as the Uyghur language, text detection is paid less attention. In this paper, we propose an effective Uyghur language text detection system in complex background images. First, a new channel-enhanced maximally stable extremal regions (MSERs) algorithm is put forward to detect component candidates. Second, a two-layer filtering mechanism is designed to remove most non-character regions. Third, the remaining component regions are connected into short chains, and the short chains are extended by a novel extension algorithm to connect the missed MSERs. Finally, a two-layer chain elimination filter is proposed to prune the non-text chains. To evaluate the system, we build a new data set by various Uyghur texts with complex backgrounds. Extensive experimental comparisons show that our system is obviously effective for Uyghur language text detection in complex background images. The F-measure is 85%, which is much better than the state-of-the-art performance of 75.5%.
Chenggang Yan 0001, Hongtao Xie 0001, Jian Yin 0003, Yongdong Zhang 0001, Qionghai Dai
IEEE Trans. Intell. Transp. Syst.6
2018 Supervised Hash Coding With Deep Neural Network for Environment Perception of Intelligent Vehicles
abstract
Image content analysis is an important surround perception modality of intelligent vehicles. In order to efficiently recognize the on-road environment based on image content analysis from the large-scale scene database, relevant images retrieval becomes one of the fundamental problems. To improve the efficiency of calculating similarities between images, hashing techniques have received increasing attentions. For most existing hash methods, the suboptimal binary codes are generated, as the hand-crafted feature representation is not optimally compatible with the binary codes. In this paper, a one-stage supervised deep hashing framework (SDHP) is proposed to learn high-quality binary codes. A deep convolutional neural network is implemented, and we enforce the learned codes to meet the following criterions: 1) similar images should be encoded into similar binary codes, and vice versa; 2) the quantization loss from Euclidean space to Hamming space should be minimized; and 3) the learned codes should be evenly distributed. The method is further extended into SDHP+ to improve the discriminative power of binary codes. Extensive experimental comparisons with state-of-the-art hashing algorithms are conducted on CIFAR-10 and NUS-WIDE, the MAP of SDHP reaches to 87.67% and 77.48% with 48 b, respectively, and the MAP of SDHP+ reaches to 91.16%, 81.08% with 12 b, 48 b on CIFAR-10 and NUS-WIDE, respectively. It illustrates that the proposed method can obviously improve the search accuracy.
Chenggang Yan 0001, Hongtao Xie 0001, Dongbao Yang, Jian Yin 0003, Yongdong Zhang 0001, Qionghai Dai
IEEE Trans. Intell. Transp. Syst.6
2018 Depth Assisted Adaptive Workload Balancing for Parallel View Synthesis
abstract
Depth image-based rendering has been adopted by MPEG as the recommended view synthesis technique for free viewpoint TV applications. In this paper, a workload balancing algorithm is proposed for parallel view synthesis on multicore platforms. First, view synthesis workload is defined as the function of the number of hole-pixels in the warped images. Then, a novel depth assisted prediction method is proposed to predict the number of hole-pixels in the current frame by exploiting the depth differences between the neighboring frames, which reflects the movement of objects in video content. Feeding the predicted workload to the proposed cost function, each input frame is partitioned adaptively to balance the synthesis workload among the cores. The proposed workload prediction method outperforms the existing approaches both in terms of frame average prediction error and standard deviation in prediction error. Applying the proposed workload balancing method, the parallel view synthesis system provides higher acceleration ratio and better synchronization performance among the cores compared with other parallel processing systems without sacrificing the subjective and objective quality. It is also robust to different platforms, which shows high potential in being applied to mobile oriented applications.
Xin Jin 0002, Zhanqi Liu, Qionghai Dai
IEEE Trans. Multim.4
2018 A Fast Uyghur Text Detector for Complex Background Images
abstract
Uyghur text localization in images with complex backgrounds is a challenging yet important task for many applications. Generally, Uyghur characters in images consist of strokes with uniform features, and they are distinct from backgrounds in color, intensity, and texture. Based on these differences, we propose a FASTroke keypoint extractor, which is fast and stroke-specific. Compared with the commonly used MSER detector, FASTroke produces less than twice the amount of components and recognizes at least 10% more characters. While the characters in a line usually have uniform features such as size, color, and stroke width, a component similarity based clustering is presented without component-level classification. It incurs no extra errors by incorporating a component-level classifier while the computing cost is drastically reduced. The experiments show that the proposed method can achieve the best performance on the UICBI-500 benchmark dataset.
Chenggang Yan 0001, Hongtao Xie 0001, Zhengjun Zha, Xinhong Hao, Yongdong Zhang 0001, Qionghai Dai
IEEE Trans. Multim.7
2018 Robust Non-Rigid Motion Tracking and Surface Reconstruction Using L0 Regularization
abstract
We present a new motion tracking technique to robustly reconstruct non-rigid geometries and motions from a single view depth input recorded by a consumer depth sensor. The idea is based on the observation that most non-rigid motions (especially human-related motions) are intrinsically involved in articulate motion subspace. To take this advantage, we propose a novel based motion regularizer with an iterative solver that implicitly constrains local deformations with articulate structures, leading to reduced solution space and physical plausible deformations. The strategy is integrated into the available non-rigid motion tracking pipeline, and gradually extracts articulate joints information online with the tracking, which corrects the tracking errors in the results. The information of the articulate joints is used in the following tracking procedure to further improve the tracking accuracy and prevent tracking failures. Extensive experiments over complex human body motions with occlusions, facial and hand motions demonstrate that our approach substantially improves the robustness and accuracy in motion tracking.
Feng Xu 0005, Yangang Wang 0001, Yebin Liu, Qionghai Dai
IEEE Trans. Vis. Comput. Graph.5
2018 Errata to "Robust Non-Rigid Motion Tracking and Surface Reconstruction Using L0 Regularization"
abstract
Presents corrections to grant number information from the paper, “Robust non-rigid motion tracking and surface reconstruction using L0 regularization,” (Guo, K., et al), IEEE Trans. Vis. Comput. Graph., vol. 24, no. 5, pp. 1770–1783, May 2018.
Feng Xu 0005, Yangang Wang 0001, Yebin Liu, Qionghai Dai
IEEE Trans. Vis. Comput. Graph.5
2018 Outdoor Markerless Motion Capture with Sparse Handheld Video Cameras
abstract
We present a method for outdoor markerless motion capture with sparse handheld video cameras. In the simplest setting, it only involves two mobile phone cameras following the character. This setup can maximize the flexibilities of data capture and broaden the applications of motion capture. To solve the character pose under such challenge settings, we exploit the generative motion capture methods and propose a novel model-view consistency that considers both foreground and background in the tracking stage. The background is modeled as a deformable 2D grid, which allows us to compute the background-view consistency for sparse moving cameras. The 3D character pose is tracked with a global-local optimization through minimizing our consistency cost. A novel motion regularizer is also proposed in the optimization to constrain the solution pose space. The whole process of the proposed method is simple as frame by frame video segmentation is not required. Our method outperforms several alternative methods on various examples demonstrated in the paper.
Yangang Wang 0001, Yebin Liu, Xin Tong 0001, Qionghai Dai, Ping Tan 0002
IEEE Trans. Vis. Comput. Graph.4
2018 FlyCap: Markerless Motion Capture Using Multiple Autonomous Flying Cameras
abstract
Aiming at automatic, convenient and non-instrusive motion capture, this paper presents a new generation markerless motion capture technique, the FlyCap system, to capture surface motions of moving characters using multiple autonomous flying cameras (autonomous unmanned aerial vehicles(UAVs) each integrated with an RGBD video camera). During data capture, three cooperative flying cameras automatically track and follow the moving target who performs large-scale motions in a wide space. We propose a novel non-rigid surface registration method to track and fuse the depth of the three flying cameras for surface motion tracking of the moving target, and simultaneously calculate the pose of each flying camera. We leverage the using of visual-odometry information provided by the UAV platform, and formulate the surface tracking problem in a non-linear objective function that can be linearized and effectively minimized through a Gaussian-Newton method. Quantitative and qualitative experimental results demonstrate the plausible surface and motion reconstruction results.
Lan Xu 0003, Yebin Liu, Guyue Zhou, Qionghai Dai, Lu Fang 0001
IEEE Trans. Vis. Comput. Graph.6
2017 Light Field Reconstruction Using Deep Convolutional Network on EPI
abstract
In this paper, we take advantage of the clear texture structure of the epipolar plane image (EPI) in the light field data and model the problem of light field reconstruction from a sparse set of views as a CNN-based angular detail restoration on EPI. We indicate that one of the main challenges in sparsely sampled light field reconstruction is the information asymmetry between the spatial and angular domain, where the detail portion in the angular domain is damaged by undersampling. To balance the spatial and angular information, the spatial high frequency components of an EPI is removed using EPI blur, before feeding to the network. Finally, a non-blind deblur operation is used to recover the spatial detail suppressed by the EPI blur. We evaluate our approach on several datasets including synthetic scenes, real-world scenes and challenging microscope light field data. We demonstrate the high performance and robustness of the proposed framework compared with the state-of-the-arts algorithms. We also show a further application for depth enhancement by using the reconstructed light field.
Gaochang Wu, Mandan Zhao, Liangyong Wang, Qionghai Dai, Tianyou Chai, Yebin Liu
CVPR4
2017 Dynamic cloud Offloading for View Synthesis
abstract
In this paper, a dynamic offloading model is proposed to minimize the energy consumption of mobile devices by exploiting cloud computational resources for view synthesis. The computational complexity of view synthesis, the processing capability of the cloud, the processing capability and the power consumption of the mobile are considered jointly into the model to provide an optimized solution. Several simulations based on parameters of real mobile devices demonstrate that the proposed method can save an average of 42.59%(4-partition case) and 46.40%(8-partition case) of total energy on different mobile devices and an average of 67.58%(4-partition case) and 69.76%(8-partition case) of total energy under different transmitting rates than the existing algorithms for view synthesis, respectively.
Xin Jin 0002, Zhanqi Liu, Qionghai Dai
ICASSP4
2017 Enhanced depth estimation for hand-held light field cameras
abstract
Conventional depth estimation methods are confined by the occluder and homogenous regions in the scene. In this paper, we propose a new depth estimation and enhancement method. The raw depth is calculated from analyzing the Consistency Metric Range (CMR) in the angular patch. Confident depth map is obtained by analyzing the variation of CMR within a neighborhood around the lowest CMR curves. Confident depth points are propagated to the whole image by global optimization with weighted neighborhood smoothness, gradient and second derivative constraints. Finally, depth is enhanced by using weighted median filter. The experimental results demonstrated the effectiveness of the proposed approach in providing much clearer transitions of texture regions and much smoother homogenous regions after depth propagation.
Yanwen Qin, Xin Jin 0002, Yanqin Chen, Qionghai Dai
ICASSP4
2017 Multiscale gigapixel video: A cross resolution image matching and warping approach
abstract
We present a multi-scale camera array to capture and synthesize gigapixel videos in an efficient way. Our acquisition setup contains a reference camera with a short-focus lens to get a large field-of-view video and a number of unstructured long-focus cameras to capture local-view details. Based on this new design, we propose an iterative feature matching and image warping method to independently warp each local-view video to the reference video. The key feature of the proposed algorithm is its robustness to and high accuracy for the huge resolution gap (more than 8x resolution gap between the reference and the local-view videos), camera parallaxes, complex scene appearances and color inconsistency among cameras. Experimental results show that the proposed multi-scale camera array and cross resolution video warping scheme is capable of generating seamless gigapixel video without the need of camera calibration and large overlapping area constraints between the local-view cameras.
Xiaoyun Yuan, Lu Fang 0001, Qionghai Dai, David J. Brady, Yebin Liu
ICCP3
2017 BodyFusion: Real-Time Capture of Human Motion and Surface Geometry Using a Single Depth Camera
abstract
We propose BodyFusion, a novel real-time geometry fusion method that can track and reconstruct non-rigid surface motion of a human performance using a single consumer-grade depth camera. To reduce the ambiguities of the non-rigid deformation parameterization on the surface graph nodes, we take advantage of the internal articulated motion prior for human performance and contribute a skeleton-embedded surface fusion (SSF) method. The key feature of our method is that it jointly solves for both the skeleton and graph-node deformations based on information of the attachments between the skeleton and the graph nodes. The attachments are also updated frame by frame based on the fused surface geometry and the computed deformations. Overall, our method enables increasingly denoised, detailed, and complete surface reconstruction as well as the updating of the skeleton and attachments as the temporal depth frames are fused. Experimental results show that our method exhibits substantially improved nonrigid motion fusion performance and tracking robustness compared with previous state-of-the-art fusion methods. We also contribute a dataset for the quantitative evaluation of fusion-based dynamic scene reconstruction algorithms using a single depth camera.
Tao Yu 0007, Feng Xu 0005, Zhaoqi Su, Jianhui Zhao 0002, Qionghai Dai, Yebin Liu
ICCV8
2017 Single depth image super-resolution and denoising based on sparse graphs via structure tensor
abstract
The existing single depth image super-resolution (SR) methods suppose that the image to be interpolated is noise free. However, the supposition is invalid in practice because noise will be inevitably introduced in the depth image acquisition process. In this paper, we address the problem of image denoising and SR jointly based on designing sparse graphs that are useful for describing the geometric structures of data domains. In our method, we first cluster similar patches in a noisy depth image and compute an average patch. Different from the majority of the graph Fourier transform (GFT) that assumed an underlying 4-connected graph structure with vertical and horizontal edges only, we select more general sparse graph structures and edges weights based on the difference of the blocks' structure tensors. For the average patch, a graph template with edges orthogonal to the principal gradient is designed. Finally, the graph based transform (GBT) dictionary is learned from the derived correlation graph for signal representation. As shown in our experimental results, the proposed method obtains a lot of improvement in performance.
Yihui Feng, Xianming Liu 0005, Yongbing Zhang 0002, Qionghai Dai
ICIP4
2017 Lenslet image compression using adaptive macropixel prediction
abstract
In this paper, an efficient compression method is proposed for lenslet images captured by plenoptic cameras for recording the spatial and angular light information at a super-high-resolution. After applying a reversible image reshaping method to the lenslet image, a reshaped and regularized image will be generated and compressed by the video codec comprising the proposed adaptive macropixel prediction mode. Based on the analysis of spatial correlations among adjacent macropixels, two spatial prediction modes are proposed as: multi-block weighted prediction mode and co-located single-block prediction mode, to predict the coding unit by minimizing the coding cost. The multi-block weighted prediction is formulated by minimizing the Euclidean distance between the coding unit and co-located blocks in the macropixel structure. Performance evaluations have shown that the proposed method achieves 50.9% of bit-savings on average compared to HEVC. It also outperforms state-of-the-art coding methods drastically.
Haixu Han, Xin Jin 0002, Qionghai Dai
ICIP3
2017 Lenslet image compression based on image reshaping and macro-pixel Intra prediction
abstract
Lenslet images that record both spatial and angular light radiance in a super high definition with distinct macropixel structures desire efficient compression methods for promoting the applications of handheld plenoptic cameras urgently. In this paper, a lenslet image compression method is proposed. First, a reversible image reshaping and adaptive interpolation is proposed to align the macropixel structures with coding unit grids in the block based video coding standards. Then, based on the reshaped and regularized lenslet images, a macro-pixel Intra prediction mode, in which the coding unit is predicted by minimizing spatial boundary error among the adjacent macropixels, is proposed to fully exploit spatial correlations among the pixels beneath the neighboring microlens. The proposed approach outperforms HEVC by an average of 35.56% bitrate reduction. Compared with the existing coding approaches, like Intra block coding (IBC) and locally linear embedding-based (LLE) prediction, it achieves an average of 17.72%/23.06% bitrate reduction, which demonstrates its efficiency explicitly.
Haixu Han, Xin Jin 0002, Qionghai Dai
ICME3
2017 Exponential decay sine wave learning rate for fast deep neural network training
abstract
Most state-of-the-art results on image classification tasks were obtained by residual neural networks, which use stochastic gradient descent (SGD) with momentum for training. In most cases, the learning rate drops by a constant factor every pre-defined number of epochs. However, it is difficult and time-consuming to estimate how many epochs to drop the learning rate. To tackle this problem, cyclical learning rate is gaining popularity in gradient-based optimization to improve the convergence speed in accelerated gradient schemes. But cyclical learning rate scheme scans a broad range of learning rate, some of which are not suitable for deep neural network training. In this paper, we propose a simple yet effective exponential decay sine wave like learning rate technique for SGD to improve its convergence speed. In the training process, the learning rate would vary in sine wave way. While the maximum value of sine wave would decay exponentially along with training epochs. An ensemble of wide residual nets with our proposed learning scheme achieves 3.01% and 16.03% errors on CIFAR-10 and CIFAR-100 respectively. Furthermore, our proposed method uses far less number of epochs than most recent learning rate strategies, accelerating neural network training tremendously.
Wangpeng An, Haoqian Wang, Yulun Zhang 0001, Qionghai Dai
VCIP4
2017 Non-invasive imaging based on speckle pattern estimation and deconvolution
abstract
Non-invasive imaging through scattering media is still a challenging task, especially when the imaging target is complex. This paper presents a novel imaging method through scattering media based on speckle pattern estimation and deconvolution. Different from previous frameworks based on PSF-adjusting or PSF-capturing that need to invasively put a point source at the imaging area, we utilize the non-invasive imaging system based on speckle scanning and memory effect. Phase retrieval results of simple targets behind the scattering media are used as the input of the proposed speckle pattern estimation model, in which speckle modeling and constrained least square optimization are applied to estimate the distribution of speckle pattern. The estimated speckle pattern is exploited in deconvoluting the integrated intensity matrices (IIMs) of the scattered images to recover the complex targets. Experimental results show that the proposed method can recover the imaging targets behind the scattering media much more accurate and stable than the existing phase retrieval methods.
Zhouping Wang, Xin Jin 0002, Yifu Hu, Qionghai Dai
VCIP4
2017 Special feature on computational photography
Qionghai Dai
Frontiers Inf. Technol. Electron. Eng.1
2017 Emerging theories and technologies on computational imaging
abstract
Computational imaging describes the whole imaging process from the perspective of light transport and information transmission, features traditional optical computing capabilities, and assists in breaking through the limitations of visual information recording. Progress in computational imaging promotes the development of diverse basic and applied disciplines. In this review, we provide an overview of the fundamental principles and methods in computational imaging, the history of this field, and the important roles that it plays in the development of science. We review the most recent and promising advances in computational imaging, from the perspective of different dimensions of visual signals, including spatial dimension, temporal dimension, angular dimension, spectral dimension, and phase. We also discuss some topics worth studying for future developments in computational imaging.
Jin-Li Suo, Qionghai Dai
Frontiers Inf. Technol. Electron. Eng.4
2017 Frequency-Domain Transient Imaging
abstract
A transient image is the optical impulse response of a scene, which also visualizes the propagation of light during an ultra-short time interval. In contrast to the previous transient imaging which samples in the time domain using an ultra-fast imaging system, this paper proposes transient imaging in the frequency domain using a multi-frequency time-of-flight (ToF) camera. Our analysis reveals the Fourier relationship between transient images and the measurements of a multi-frequency ToF camera, and identifies the causes of the systematic error-non-sinusoidal and frequency-varying waveforms and limited frequency range of the modulation signal. Based on the analysis we propose a novel framework of frequency-domain transient imaging. By removing the systematic error and exploiting the harmonic components inside the measurements, we achieves high quality reconstruction results. Moreover, our technique significantly reduces the computational cost of ToF camera based transient image reconstruction, especially reduces the memory usage, such that it is feasible for the reconstruction of transient images at extremely small time steps. The effectiveness of frequency-domain transient imaging is tested on synthetic data, real data from the web, and real data acquired by our prototype camera.
Yebin Liu, Jin-Li Suo, Qionghai Dai
IEEE Trans. Pattern Anal. Mach. Intell.4
2017 Depth Estimation by Parameter Transfer With a Lightweight Model for Single Still Images
abstract
In this paper, we propose a novel method for automatic depth estimation from color images using parameter transfer. By modeling the correlation between color images and their depth maps with a set of parameters, we get a database of parameter sets. Given an input image, we extract the high-level features to find the best matched image sets from the database. Then the set of parameters corresponding to the best match are used to estimate the depth of the input image. Compared with the past learning-based methods, our trained model consists only of trained features and parameter sets, which occupy little space. We evaluate our depth estimation method on several benchmark RGB-D (RGB + depth) data sets. The experimental results are comparable to the state-of-the-art results, while the model size is very small and very suitable for mobile devices, demonstrating the promising performance of our proposed method.
Hongwei Qin, Xiu Li 0001, Yangang Wang 0001, Yongbing Zhang 0002, Qionghai Dai
IEEE Trans. Circuits Syst. Video Technol.5
2017 Efficient Method for High-Quality Removal of Nonuniform Blur in the Wavelet Domain
abstract
This paper presents a novel nonuniform deblurring approach, which defines the blur model and calculates regularized nonuniform deconvolution in the wavelet domain to achieve high efficiency and high accuracy simultaneously. Targeting high computation efficiency, we derive a wavelet-domain hierarchical blur model, which can be calculated efficiently by exploiting the sparsity property of natural images in the wavelet domain. Correspondingly, the blur model is incorporated into a multilayer framework and at each layer spatially varying step sizes are introduced to further accelerate the convergence of the algorithm. In addition to the efficiency advantages, the proposed approach deals with intensely nonuniform blur with high accuracy due to the intrinsic tight supportness of wavelet basis. We conduct a series of experiments and comparisons to validate the efficiency and effectiveness of our algorithm.
Tao Yue 0003, Jin-Li Suo, Xun Cao, Qionghai Dai
IEEE Trans. Circuits Syst. Video Technol.4
2017 Light-Field Depth Estimation via Epipolar Plane Image Analysis and Locally Linear Embedding
abstract
In this paper, we propose a novel method for 4D light-field (LF) depth estimation exploiting the special linear structure of an epipolar plane image (EPI) and locally linear embedding (LLE). Without high computational complexity, depth maps are locally estimated by locating the optimal slope of each line segmentation on the EPIs, which are projected by the corresponding scene points. For each pixel to be processed, we build and then minimize the matching cost that aggregates the intensity pixel value, gradient pixel value, spatial consistency, as well as reliability measure to select the optimal slope from a predefined set of directions. Next, a subangle estimation method is proposed to further refine the obtained optimal slope of each pixel. Furthermore, based on a local reliability measure, all the pixels are classified into reliable and unreliable pixels. For the unreliable pixels, LLE is employed to propagate the missing pixels by the reliable pixels based on the assumption of manifold preserving property maintained by natural images. We demonstrate the effectiveness of our approach on a number of synthetic LF examples and real-world LF data sets, and show that our experimental results can achieve higher performance than the typical and recent state-of-the-art LF stereo matching methods.
Yongbing Zhang 0002, Huijin Lv, Yebin Liu, Haoqian Wang, Xingzheng Wang, Qian Huang 0008, Xinguang Xiang, Qionghai Dai
IEEE Trans. Circuits Syst. Video Technol.8
2017 Video-Based Outdoor Human Reconstruction
abstract
A human body scanning system of great practical convenience, which can be used in an outdoor environment, is proposed. The system uses only a single conventional video camera without the aid of special sensors or controlled illuminations. We leverage the structure from motion calibration results directly and improve the available video-based dense 3D reconstruction by integrating the surface smoothness constraints. The point cloud reinforcement is proposed to detect and adjust the conflict point data for the slender and shaky body parts. Combined with the silhouette adaptation, the proposed point cloud reinforcement achieves reasonable and plausible mesh reconstruction on these challenging parts. We further introduce the close-shot frames to refine the prereconstructed mesh model, leading to a colored watertight model. The overall system is approximate to automatic since only one or two times of painting brush interaction are required for robust and high-quality multiview image segmentation. The experiment results on various test sequences demonstrate the effectiveness and the robustness of the proposed method, even under very challenging scenarios when shaking body, varying illumination, and textureless regions occur.
Hao Zhu 0004, Yebin Liu, Jingtao Fan, Qionghai Dai, Xun Cao
IEEE Trans. Circuits Syst. Video Technol.4
2017 Discriminant Kernel Assignment for Image Coding
abstract
This paper proposes discriminant kernel assignment (DKA) in the bag-of-features framework for image representation. DKA slightly modifies existing kernel assignment to learn width-variant Gaussian kernel functions to perform discriminant local feature assignment. When directly applying gradient-descent method to solve DKA, the optimization may contain multiple time-consuming reassignment implementations in iterations. Accordingly, we introduce a more practical way to locally linearize the DKA objective and the difficult task is cast as a sequence of easier ones. Since DKA only focuses on the feature assignment part, it seamlessly collaborates with other discriminative learning approaches, e.g., discriminant dictionary learning or multiple kernel learning, for even better performances. Experimental evaluations on multiple benchmark datasets verify that DKA outperforms other image assignment approaches and exhibits significant efficiency in feature coding.
Yue Deng 0001, Yanyu Zhao, Zhiquan Ren, Youyong Kong, Feng Bao 0002, Qionghai Dai
IEEE Trans. Cybern.6
2017 Toward Simultaneous Visual Comfort and Depth Sensation Optimization for Stereoscopic 3-D Experience
abstract
Visual comfort and depth sensation are two important incongruent counterparts in determining the overall stereoscopic 3-D experience. In this paper, we proposed a novel simultaneous visual comfort and depth sensation optimization approach for stereoscopic images. The main motivation of the proposed optimization approach is to enhance the overall stereoscopic 3-D experience. Toward this end, we propose a two-stage solution to address the optimization problem. In the first layer-independent disparity adjustment process, we iteratively adjust the disparity range of each depth layer to satisfy with visual comfort and depth sensation constraints simultaneously. In the following layer-dependent disparity process, disparity adjustment is implemented based on a defined total energy function built with intra-layer data, inter-layer data and just noticeable depth difference terms. Experimental results on perceptually uncomfortable and comfortable stereoscopic images demonstrate that in comparison with the existing methods, the proposed method can achieve a reasonable performance balance between visual comfort and depth sensation, leading to promising overall stereoscopic 3-D experience.
Feng Shao 0001, Weisi Lin, Zhutuan Li, Gangyi Jiang, Qionghai Dai
IEEE Trans. Cybern.5
2017 A Hierarchical Fused Fuzzy Deep Neural Network for Data Classification
abstract
Deep learning (DL) is an emerging and powerful paradigm that allows large-scale task-driven feature learning from big data. However, typical DL is a fully deterministic model that sheds no light on data uncertainty reductions. In this paper, we show how to introduce the concepts of fuzzy learning into DL to overcome the shortcomings of fixed representation. The bulk of the proposed fuzzy system is a hierarchical deep neural network that derives information from both fuzzy and neural representations. Then, the knowledge learnt from these two respective views are fused altogether forming the final data representation to be classified. The effectiveness of the model is verified on three practical tasks of image categorization, high-frequency financial data prediction and brain MRI segmentation that all contain high level of uncertainties in the raw data. The fuzzy dDL paradigm greatly outperforms other nonfuzzy and shallow learning approaches on these tasks.
Yue Deng 0001, Zhiquan Ren, Youyong Kong, Feng Bao 0002, Qionghai Dai
IEEE Trans. Fuzzy Syst.5
2017 Learning Sparse Representation for No-Reference Quality Assessment of Multiply Distorted Stereoscopic Images
abstract
Binocular combination under different distortion types poses a great challenge to three-dimensional image quality assessment (3D-IQA). However, the research works on 3D-IQA with multiple distortion types are very limited. In this paper, we first construct a new multiply distorted stereoscopic image database (NBU-MDSID), which is composed of 270 multiply distorted stereoscopic images and 90 singly distorted stereoscopic images that are corrupted simultaneously and independently by blurring, JPEG compression, and noise injection. We then propose a new multimodal blind metric for quality assessment of multiply distorted stereoscopic images. Inspired by multimodal sparse representation framework, modality-specific dictionaries and the corresponding projection matrices are learned from the singly distorted training database at the training stage, and the testing stage only needs to estimate the quality score based on the reconstruction errors. Experimental results demonstrate the effectiveness of our blind metric.
Feng Shao 0001, Weijun Tian, Weisi Lin, Gangyi Jiang, Qionghai Dai
IEEE Trans. Multim.5
2017 Deep Direct Reinforcement Learning for Financial Signal Representation and Trading
abstract
Can we train the computer to beat experienced traders for financial assert trading? In this paper, we try to address this challenge by introducing a recurrent deep neural network (NN) for real-time financial signal representation and trading. Our model is inspired by two biological-related learning concepts of deep learning (DL) and reinforcement learning (RL). In the framework, the DL part automatically senses the dynamic market condition for informative feature learning. Then, the RL module interacts with deep representations and makes trading decisions to accumulate the ultimate rewards in an unknown environment. The learning system is implemented in a complex NN that exhibits both the deep and recurrent structures. Hence, we propose a task-aware backpropagation through time method to cope with the gradient vanishing issue in deep training. The robustness of the neural system is verified on both the stock and the commodity future markets under broad testing conditions.
Yue Deng 0001, Feng Bao 0002, Youyong Kong, Zhiquan Ren, Qionghai Dai
IEEE Trans. Neural Networks Learn. Syst.5
2017 Real-Time Geometry, Albedo, and Motion Reconstruction Using a Single RGB-D Camera
abstract
This article proposes a real-time method that uses a single-view RGB-D input (a depth sensor integrated with a color camera) to simultaneously reconstruct a casual scene with a detailed geometry model, surface albedo, per-frame non-rigid motion, and per-frame low-frequency lighting, without requiring any template or motion priors. The key observation is that accurate scene motion can be used to integrate temporal information to recover the precise appearance, whereas the intrinsic appearance can help to establish true correspondence in the temporal domain to recover motion. Based on this observation, we first propose a shading-based scheme to leverage appearance information for motion estimation. Then, using the reconstructed motion, a volumetric albedo fusing scheme is proposed to complete and refine the intrinsic appearance of the scene by incorporating information from multiple frames. Since the two schemes are iteratively applied during recording, the reconstructed appearance and motion become increasingly more accurate. In addition to the reconstruction results, our experiments also show that additional applications can be achieved, such as relighting, albedo editing, and free-viewpoint rendering of a dynamic scene, since geometry, appearance, and motion are all reconstructed by our technique.
Feng Xu 0005, Tao Yu 0007, Qionghai Dai, Yebin Liu
ACM Trans. Graph.5
2017 The Light Field Attachment: Turning a DSLR into a Light Field Camera Using a Low Budget Camera Ring
abstract
We propose a concept for a lens attachment that turns a standard DSLR camera and lens into a light field camera. The attachment consists of eight low-resolution, low-quality side cameras arranged around the central high-quality SLR lens. Unlike most existing light field camera architectures, this design provides a high-quality 2D image mode, while simultaneously enabling a new high-quality light field mode with a large camera baseline but little added weight, cost, or bulk compared with the base DSLR camera. From an algorithmic point of view, the high-quality light field mode is made possible by a new light field super-resolution method that first improves the spatial resolution and image quality of the side cameras and then interpolates additional views as needed. At the heart of this process is a super-resolution method that we call iterative Patch- And Depth-based Synthesis (iPADS), which combines patch-based and depth-based synthesis in a novel fashion. Experimental results obtained for both real captured data and synthetic data confirm that our method achieves substantial improvements in super-resolution for side-view images as well as the high-quality and view-coherent rendering of dense and high-resolution light fields.
Yuwang Wang, Yebin Liu, Wolfgang Heidrich, Qionghai Dai
IEEE Trans. Vis. Comput. Graph.4
2016 Deep Convolutional Neural Network for Decompressed Video Enhancement
abstract
Block-wise intra/inter prediction, transformation and quantization used in block-based hybrid video coding will inevitably result in blocking artifacts, especially at the low bit rate. To address this problem, this paper employs a deep convolutional neural network (CNN) to approximate the reverse function of video compression, motived by the great success of deep learning in computer vision fields recently. The proposed method establishes an end-to-end mapping, represented as the CNN, which takes the decompressed frame as input and outputs the enhanced one. Employing numerous sequences compressed by H.264 and HEVC reference software, the proposed CNN learns the connections between the lossy frame and the original one in an implicit way under different quantization parameters (QP). Figure 1 shows the architecture of our CNN and the pipeline of the network training. We build our network with convolution layers and ReLU layer and the weights and biases of all the convolution layers in our model are updated by minimizing the loss using stochastic gradient descent with the standard backpropagation. We implement the CNN as a post-loop deblocking filter and explore varying CNN parameters for different QPs. Various experimental results demonstrate that the proposed method is able to significantly improve the quality of enhanced frames in terms of both objective and subjective criterions.
Rongqun Lin, Yongbing Zhang 0002, Haoqian Wang, Xingzheng Wang, Qionghai Dai
DCC5
2016 Depth fused from intensity range and blur estimation for light-field cameras
abstract
Light-field cameras attract great attention because of its refocusing and perspective-shifting functions after capturing. The special 4D-structured data contains depth information. In this paper, a novel depth estimation algorithm is proposed for light-field cameras by fully exploiting the characteristics of 4D light-field data. A novel tensor, intensity range of pixels within a microlens, is proposed, which presents strong correlation with the transition on focus, especially for texture-complex regions. Meanwhile, the other tensor, defocus blur amount is utilized to estimate the focus level, which generates more accurate depth estimation especially for homogeneous regions. Then, the depths calculated from the two tensors are fused according to the variation scale of intensity range and the minimal defocus blur amount under spatial smoothness constraints. Compared with the representative approaches, the depth generated by the proposed approach presents richer details for texture regions and higher consistency for unified regions.
Yatong Xu, Xin Jin 0002, Qionghai Dai
ICASSP3
2016 Imaging through scattering media with intensity modulated incoherent sources
abstract
Imaging through scattering media is a significant challenge in computational imaging. Recently, a breakthrough technique based on speckle scanning was proposed with outstanding imaging performance. However, the dense angular scanning of the incident laser beam leads to a lengthy scanning process. In this paper, we propose a method based on compressive sensing (CS) to accelerate the data acquisition process. A new imaging system is proposed which adopts mutually incoherent laser beams as light sources and the intensity of each beam is modulated to perform CS measurement. High-resolution fluctuation of the total fluorescence can be well reconstructed from CS measurements with the learned dictionary as basis function. Experimental results show that the proposed method can reduce nearly 70 percent of the data acquisition complexity without impairing the imaging quality.
Yifu Hu, Xin Jin 0002, Kaiyun Wei, Qionghai Dai
ICIP4
2016 Fourier ptychographic reconstruction using weighted replacement in the fourier domain
abstract
Fourier ptychographic microscopy (FPM) is an attractive method to extend the resolution beyond the conventional limit defined by a microscope optics, sharing properties with ptychographic, synthetic aperture imaging and phase retrieval. The algorithm uses a sequence of low-resolution (LR) images acquired under angularly varying illumination to reconstruct a high-resolution (HR) image. However, traditional FPM may trap to the sub-optimal solution, since brute-force replacement in the Fourier domain is applied. To address this problem, we propose here a weighted replacement for Fourier ptychographic microscopy (WFPM). We employ the weighted average of spectrums corresponding to different illumination angle to replace the overlapped regions in the Fourier domain. A series of experimental results demonstrate that the reconstructed image using the proposed WFPM shows a better quality and a faster convergence compared with the results obtained by FPM.
Pengming Song, Weixin Jiang, Yongbing Zhang 0002, Qionghai Dai
ICIP4
2016 Parameterized reconstruction based Fourier Ptychography
abstract
Fourier ptychography (FP) is recently proposed as a computational imaging technique, which aims at enhancing the space-bandwidth product (SBP) of the optical imaging system. Specifically, the FP recovery routine iteratively stitches together a number of low-resolution images, which are captured under angularly varying illumination, to produce a wide-field, high-resolution image. However, the reconstruction procedure of the FP recovery routine is based on a low efficient phase retrieval algorithm and may severely degrade quality of the reconstructions. To address this problem, in this paper, we develop and test a Parameterized Reconstruction based Fourier Ptychography (PR-FP), which parameterizes the reconstruction procedure by introducing an updating-coefficient. Meanwhile, a convergence-related metric, which measures how good the reconstruction matches the input dataset, is proposed to help determine a proper updating-coefficient. Extensive experimental results demonstrate that the proposed PR-FP algorithm achieves superior reconstructions both on simulated dataset and real captured dataset.
Weixin Jiang, Yongbing Zhang 0002, Qionghai Dai
ICME3
2016 A SVR based quality metric for depth quality assessment
abstract
A depth map generally can be divided into the region with sharp edges and the region consisting of nearly constant or slowly varying samples. In this paper, in order to investigate how the distortion in the two regions affect the perceived 3D quality of synthesized stereopairs, a dataset is first built based on distinctive coding for two regions, and then a subjective test is conducted. Based on the subjective evaluation results, a support vector regression (SVR) based model is built to estimate the perceived 3D quality of the synthesized stereopairs with the features extracted from the depth maps. As a metric for depth quality assessment, the proposed model outperforms the conventional 2D QA metrics applied to the depth maps as well as that applied to the stereopairs.
Xin Jin 0002, Qionghai Dai
ISCAS3
2016 Efficient imaging through scattering media by random sampling
abstract
Imaging through scattering media is a tough task in computational imaging. A recent breakthrough technique based on speckle scanning was proposed with outstanding imaging performance. However, to achieve high imaging quality, dense sampling of the integrated intensity matrix is needed, which leads to a time-consuming scanning process. In this paper, we propose a method that exploits spatial redundancy of the integrated intensity matrix and reconstructs the complete matrix from few random samples. A reconstruction model that jointly penalizes total variation and weighted sum of nuclear norm of local patches is built with improved reconstruction quality. Experiments are performed to verify the effectiveness of the proposed method and results demonstrate that the proposed method can achieve a same imaging quality with 80% reduction of the data acquisition complexity.
Yifu Hu, Xin Jin 0002, Qionghai Dai
MMSP3
2016 Single image super-resolution via projective dictionary learning with anchored neighborhood regression
abstract
We propose a novel single image super-resolution (SR) algorithm based on the projective dictionary pair learning with anchored neighborhood regression. Different from previous dictionary learning methods that aim to learn only a synthesis or an analysis dictionary, our method would learn both types of dictionaries jointly for regression to achieve image SR. We first cluster the training features into K clusters in order to learn synthesis and analysis dictionaries. Moreover, we learn the regressions with the training samples at training phase and use them on reconstruction stage. As shown in our experimental results, the proposed method obtains high-quality SR results quantitatively and visually against state-of-the-art methods.
Yihui Feng, Yongbing Zhang 0002, Yulun Zhang 0001, Qionghai Dai
VCIP5
2016 Decompressed video enhancement via accurate regression prior
abstract
There is an increasing need for high-quality multimedia applications based on block-based hybrid video coding. Inevitably, the frame will degrade during the process of block-wise intra/inter prediction, transformation, and quantization, especially when the bit rate is low. In this paper, we propose an efficient decompressed video enhancement algorithm based on the adjusted anchored neighborhood regression (A+) method. In our work, first, we learn offline linear regressors, i.e. projection matrices from the decompressed to original video frames in the training phase. For grouping anchored neighborhoods more accurately, we adopt MI-KSVD rather than KSVD to learn the dictionary. Moreover, we exploit the mutual coherence between dictionary atoms and training samples to find the nearest neighbors. Second, in the enhancement phase, we boost the quality of input decompressed videos offline by learned regression priors. To verify the robustness of our enhancement method, extensive experiments are conducted. As shown in our experimental results, the proposed enhancement method yields superior performance both objectively and subjectively.
Yulun Zhang 0001, Yongbing Zhang 0002, Xingzheng Wang, Haoqian Wang, Qionghai Dai
VCIP6
2016 A 3D subjective quality prediction model based on depth distortion
abstract
Depth map quality plays an important role in 3D subjective quality. In this paper, a 3D subjective quality prediction model is proposed to estimate the 3D quality of synthesized stereopairs based on depth map distortion and neural mechanism, instead of performing view synthesis directly, which benefits 3D processing. In order to build the model, a dataset is first built to include distinctive distortion features for depth map coding, and then a subjective test is conducted. Based on the subjective evaluation results, a prediction model is built to estimate the perceived 3D quality of the synthesized stereopairs with the features extracted from the texture characteristics of decoded depth maps and neural population coding model. In terms of the correlation coefficient with actual 3D subjective quality, the proposed model outperforms the conventional 2D QA metrics applied to the depth maps as well as that applied to the stereopairs.
Xin Jin 0002, Qionghai Dai
VCIP3
2016 Re-Compositable Panoramic Selfie with Robust Multi-Frame Segmentation and Stitching
abstract
Abstract It is a challenging task for ordinary users to capture selfies with a good scene composition, given the limited freedom to position the camera. Creative hardware (e.g., selfie sticks) and software (e.g., panoramic selfie apps) solutions have been proposed to extend the background coverage of a selife, but to achieve a perfect composition on the spot when the selfie is captured remains to be difficult. In this paper, we propose a system that allows the user to shoot a selfie video by rotating the body first, then produce a final panoramic selfie image with user‐guided scene composition as postprocessing. Our key technical contribution is a fully Automatic, robust multi‐frame segmentation and stitching framework that is tailored towards the special characteristics of selfie images. We analyze the sparse feature points and employ a spatial‐temporal optimization for bilayer feature segmentation, which leads to more reliable background alignment than previous image stitching techniques. The sparse classification is then propagated to all pixels to create dense foreground masks for person‐background composition. Finally, based on a user‐selected foreground position, our system uses content‐preserving warping to produce a panoramic seflie with minimal distortion to the face region. Experimental results show that our approach can reliably generate high quality panoramic selfies, while a simple combination of previous image stitching and segmentation approaches often fails.
Kai Li 0016, Jue Wang 0001, Yebin Liu, Qionghai Dai
Comput. Graph. Forum5
2016 Local visual feature fusion via maximum margin multimodal deep neural network
Zhiquan Ren, Yue Deng 0001, Qionghai Dai
Neurocomputing3
2016 Directed Adaptive Graphical Lasso for causality inference
Zhiquan Ren, Feng Bao 0002, Yue Deng 0001, Qionghai Dai
Neurocomputing5
2016 Sampling-based causal inference in cue combination and its neural implementation
Zhaofei Yu, Feng Chen 0007, Jianwu Dong, Qionghai Dai
Neurocomputing4
2016 Robust subspace segmentation via nonconvex low rank representation
Wei Jiang 0007, Heng Qi, Qionghai Dai
Inf. Sci.4
2016 Signal-dependent noise removal for color videos using temporal and cross-channel priors
Jin-Li Suo, Liheng Bian, Feng Chen 0007, Qionghai Dai
J. Vis. Commun. Image Represent.4
2016 Special issue: When social media meets physical world
Rongrong Ji, Yue Gao 0002, Qi Tian 0001, Qionghai Dai, Ralf Steinmetz
Multim. Syst.4
2016 Recent advances in social multimedia big data mining and applications
Yue Gao 0002, Bing-Kun Bao, Cees Snoek, Qionghai Dai
Multim. Syst.5
2016 Online distribution and interaction of video data in social multimedia network
Xiangyang Ji, Qifei Wang, Bo-Wei Chen, Seungmin Rho, C.-C. Jay Kuo, Qionghai Dai
Multim. Tools Appl.6
2016 Normalized filter pool for prior modeling of nature images
Yangang Wang 0001, Jin-Li Suo, Qionghai Dai
Mach. Vis. Appl.3
2016 Depth dithering based on texture edge-assisted classification
Xin Jin 0002, Yatong Xu, Qionghai Dai
Signal Process. Image Commun.3
2016 A Polynomial Approximation Motion Estimation Model for Motion-Compensated Frame Interpolation
abstract
Motion-compensated frame interpolation (MCFI) usually finds the most matched blocks by minimizing pixel intensity discrepancies between neighboring frames along the motion trajectory. However, the quality of interpolated frames is susceptible to inaccurate motion vectors possibly for regions with complex texture patterns, irregularly shaped objects, repeated patterns, motion blurring or aliasing, and so on. It is believed that pixel intensity across adjacent frames varies gradually and smoothly, and therefore can be modeled mathematically by a continuous and differentiable function. Thus, the pixel intensity within one frame can be expressed as either a forward polynomial approximation (FWPA) or a backward polynomial approximation (BWPA) by the Taylor expansion in this paper. The discrepancy between the FWPA and BWPA is employed to find the best motion vector. In addition, a motion-aligned partial derivative is proposed to calculate the Taylor expansion along the motion trajectory. The proposed method is applicable to any existing MCFI schemes and achieve superior performance by consuming relatively more buffer memories and computational resources. Extensive experimentation with comparison with previous techniques validates our method in terms of both objective and subjective criteria.
Yongbing Zhang 0002, Long Xu 0001, Xiangyang Ji, Qionghai Dai
IEEE Trans. Circuits Syst. Video Technol.4
2016 Learning Receptive Fields and Quality Lookups for Blind Quality Assessment of Stereoscopic Images
abstract
Blind quality assessment of 3D images encounters more new challenges than its 2D counterparts. In this paper, we propose a blind quality assessment for stereoscopic images by learning the characteristics of receptive fields (RFs) from perspective of dictionary learning, and constructing quality lookups to replace human opinion scores without performance loss. The important feature of the proposed method is that we do not need a large set of samples of distorted stereoscopic images and the corresponding human opinion scores to learn a regression model. To be more specific, in the training phase, we learn local RFs (LRFs) and global RFs (GRFs) from the reference and distorted stereoscopic images, respectively, and construct their corresponding local quality lookups (LQLs) and global quality lookups (GQLs). In the testing phase, blind quality pooling can be easily achieved by searching optimal GRF and LRF indexes from the learnt LQLs and GQLs, and the quality score is obtained by combining the LRF and GRF indexes together. Experimental results on three publicly 3D image quality assessment databases demonstrate that in comparison with the existing methods, the devised algorithm achieves high consistent alignment with subjective assessment.
Feng Shao 0001, Weisi Lin, Gangyi Jiang, Mei Yu 0001, Qionghai Dai
IEEE Trans. Cybern.6
2016 Deep and Structured Robust Information Theoretic Learning for Image Analysis
abstract
This paper presents a robust information theoretic (RIT) model to reduce the uncertainties, i.e., missing and noisy labels, in general discriminative data representation tasks. The fundamental pursuit of our model is to simultaneously learn a transformation function and a discriminative classifier that maximize the mutual information of data and their labels in the latent space. In this general paradigm, we, respectively, discuss three types of the RIT implementations with linear subspace embedding, deep transformation, and structured sparse learning. In practice, the RIT and deep RIT are exploited to solve the image categorization task whose performances will be verified on various benchmark data sets. The structured sparse RIT is further applied to a medical image analysis task for brain magnetic resonance image segmentation that allows group-level feature selections on the brain tissues.
Yue Deng 0001, Feng Bao 0002, XueSong Deng, Ruiping Wang 0001, Youyong Kong, Qionghai Dai
IEEE Trans. Image Process.6
2016 Toward a Blind Deep Quality Evaluator for Stereoscopic Images Based on Monocular and Binocular Interactions
abstract
During recent years, blind image quality assessment (BIQA) has been intensively studied with different machine learning tools. Existing BIQA metrics, however, do not design for stereoscopic images. We believe this problem can be resolved by separating 3D images and capturing the essential attributes of images via deep neural network. In this paper, we propose a blind deep quality evaluator (DQE) for stereoscopic images (denoted by 3D-DQE) based on monocular and binocular interactions. The key technical steps in the proposed 3D-DQE are to train two separate 2D deep neural networks (2D-DNNs) from 2D monocular images and cyclopean images to model the process of monocular and binocular quality predictions, and combine the measured 2D monocular and cyclopean quality scores using different weighting schemes. Experimental results on four public 3D image quality assessment databases demonstrate that in comparison with the existing methods, the devised algorithm achieves high consistent alignment with subjective assessment.
Feng Shao 0001, Weijun Tian, Weisi Lin, Gangyi Jiang, Qionghai Dai
IEEE Trans. Image Process.5
2016 Fast and High Quality Highlight Removal From a Single Image
abstract
Specular reflection exists widely in photography and causes the recorded color deviating from its true value, thus, fast and high quality highlight removal from a single nature image is of great importance. In spite of the progress in the past decades in highlight removal, achieving wide applicability to the large diversity of nature scenes is quite challenging. To handle this problem, we propose an analytic solution to highlight removal based on an L2chromaticity definition and corresponding dichromatic model. Specifically, this paper derives a normalized dichromatic model for the pixels with identical diffuse color: a unit circle equation of projection coefficients in two subspaces that are orthogonal to and parallel with the illumination, respectively. In the former illumination orthogonal subspace, which is specular-free, we can conduct robust clustering with an explicit criterion to determine the cluster number adaptively. In the latter, illumination parallel subspace, a property called pure diffuse pixels distribution rule helps map each specular-influenced pixel to its diffuse component. In terms of efficiency, the proposed approach involves few complex calculation, and thus can remove highlight from high resolution images fast. Experiments show that this method is of superior performance in various challenging cases.
Jin-Li Suo, Dongsheng An, Xiangyang Ji, Haoqian Wang, Qionghai Dai
IEEE Trans. Image Process.5
2016 Clustering-Based Content Adaptive Tiles Under On-chip Memory Constraints
abstract
Tiles have been introduced to the next generation video coding standard, high-efficiency video coding (HEVC) standard, as a fundamental tool to reduce on-chip memory requirement during encoding and decoding high-definition video. In this paper, a content adaptive tile partitioning approach is proposed to improve the compression efficiency for HEVC under the on-chip memory constraint. Local competition optimization-based rectangular clustering is proposed to partition the frames into a required number of tiles adapting to content variations. Under the same memory constraint, the adaptive scheme improves compression efficiency by up to 1.8% bitrate saving relative to uniformly spaced tiles with negligible complexity increment. It especially benefits the videos with regional high spatial correlations and no penalty in compression efficiency is observed for other types of videos.
Xin Jin 0002, Qionghai Dai
IEEE Trans. Multim.2
2016 Learning Blind Quality Evaluator for Stereoscopic Images Using Joint Sparse Representation
abstract
Perceptual quality prediction for stereoscopic images is of fundamental importance in determining the level of quality perceived by humans in terms of the 3D viewing experience. However, the existing no-reference quality assessment (NR-IQA) framework has its limitation in addressing binocular combination for stereoscopic images. In this paper, we propose a new NR-IQA for stereoscopic images using joint sparse representation. We analyze the relationship between left and right quality predictors, and formulate stereoscopic quality prediction as a combination of feature-prior and feature-distribution. Based on this finding, we extract feature vector that handles different features to be interacted by joint sparse representation, and use support vector regression to characterize feature-prior. Meanwhile, we implement feature-distribution using sparsity regularization as the basis of weights for binocular combination to derive the overall quality score. Experimental results on five public 3D IQA databases demonstrate that in comparison with the existing methods, the devised algorithm achieves high consistent alignment with subjective assessment.
Feng Shao 0001, Kemeng Li, Weisi Lin, Gangyi Jiang, Qionghai Dai
IEEE Trans. Multim.5
2016 CCR: Clustering and Collaborative Representation for Fast Single Image Super-Resolution
abstract
Clustering and collaborative representation (CCR) have recently been used in fast single image super-resolution (SR). In this paper, we propose an effective and fast single image super-resolution (SR) algorithm by combining clustering and collaborative representation. In particular, we first cluster the feature space of low-resolution (LR) images into multiple LR feature subspaces and group the corresponding high-resolution (HR) feature subspaces. The local geometry property learned from the clustering process is used to collect numerous neighbor LR and HR feature subsets from the whole feature spaces for each cluster center. Multiple projection matrices are then computed via collaborative representation to map LR feature subspaces to HR subspaces. For an arbitrary input LR feature, the desired HR output can be estimated according to the projection matrix, whose corresponding LR cluster center is nearest to the input. Moreover, by learning statistical priors from the clustering process, our clustering-based SR algorithm would further decrease the computational time in the reconstruction phase. Extensive experimental results on commonly used datasets indicate that our proposed SR algorithm obtains compelling SR images quantitatively and qualitatively against many state-of-the-art methods.
Yongbing Zhang 0002, Yulun Zhang 0001, Jian Zhang 0018, Qionghai Dai
IEEE Trans. Multim.4
2016 Image Categorization by Learning a Propagated Graphlet Path
abstract
Spatial pyramid matching is a standard architecture for categorical image retrieval. However, its performance is largely limited by the prespecified rectangular spatial regions when pooling local descriptors. In this paper, we propose to learn object-shaped and directional receptive fields for image categorization. In particular, different objects in an image are seamlessly constructed by superpixels, while the direction captures human gaze shifting path. By generating a number of superpixels in each image, we construct graphlets to describe different objects. They function as the object-shaped receptive fields for image comparison. Due to the huge number of graphlets in an image, a saliency-guided graphlet selection algorithm is proposed. A manifold embedding algorithm encodes graphlets with the semantics of training image tags. Then, we derive a manifold propagation to calculate the postembedding graphlets by leveraging visual saliency maps. The sequentially propagated graphlets constitute a path that mimics human gaze shifting. Finally, we use the learned graphlet path as receptive fields for local image descriptor pooling. The local descriptors from similar receptive fields of pairwise images more significantly contribute to the final image kernel. Thorough experiments demonstrate the advantage of our approach.
Richang Hong, Yue Gao 0002, Rongrong Ji, Qionghai Dai, Xuelong Li 0001
IEEE Trans. Neural Networks Learn. Syst.5
2015 Blind optical aberration correction by exploring geometric and visual priors
abstract
Optical aberration widely exists in optical imaging systems, especially in consumer-level cameras. In contrast to previous solutions using hardware compensation or pre-calibration, we propose a computational approach for blind aberration removal from a single image, by exploring various geometric and visual priors. The global rotational symmetry allows us to transform the non-uniform degeneration into several uniform ones by the proposed radial splitting and warping technique. Locally, two types of symmetry constraints, i.e. central symmetry and reflection symmetry are defined as geometric priors in central and surrounding regions, respectively. Furthermore, by investigating the visual artifacts of aberration degenerated images captured by consumer-level cameras, the non-uniform distribution of sharpness across color channels and the image lattice is exploited as visual priors, resulting in a novel strategy to utilize the guidance from the sharpest channel and local image regions to improve the overall performance and robustness. Extensive evaluation on both real and synthetic data suggests that the proposed method outperforms the state-of-the-art techniques.
Tao Yue 0003, Jin-Li Suo, Jue Wang 0001, Xun Cao, Qionghai Dai
CVPR5
2015 Light field from micro-baseline image pair
abstract
We present a novel phase-based approach for reconstructing 4D light field from a micro-baseline stereo pair. Our approach takes advantage of the unique property of complex steerable pyramid filters in micro-baseline stereo. We first introduce a Disparity Assisted Phase based Synthesis (DAPS) strategy that can integrate disparity information into the phase term of a reference image to warp it to its close neighbor views. Based on the DAPS, an “analysis by synthesis” approach is proposed to warp from one of the input binocular images to the other, and iteratively optimize the disparity map to minimize the phase differences between the warped one and the ground truth input. Finally, the densely and regularly spaced, high quality light field images can be reconstructed using the proposed DAPS according to the refined disparity map. Our approach also solves the problems of disparity inconsistency and ringing artifact in available phase-based view synthesis methods. Experimental results demonstrate that our approach substantially improves both the quality of disparity map and light field, compared with the state-of-the-art stereo matching and image based rendering approaches.
Zhoutong Zhang, Yebin Liu, Qionghai Dai
CVPR3
2015 Image colorization using hybrid domain transform
abstract
Image colorization is the process of spreading the user specified colors to all the expected regions in image. It is a great challenge to distinguish between texture edge and object boundary. To improve the performance of colorization, we incorporate depth and texture to accurately extract boundary information in an implicit way. Inspired by the low complexity and high efficiency properties of recently proposed domain transform, we proposed a hybrid domain transform, taking corresponding depth image of the processed texture image into account, to perform colorization and recoloring. Various experimental results demonstrate that hybrid domain transform is able to achieve better edge-aware property while maintaining the property of low complexity.
Hongbo Ao, Yongbing Zhang 0002, Qionghai Dai
ICASSP3
2015 A workload balanced parallel view synthesis for FTV
abstract
In this paper, a parallel system together with an adaptive workload balancing algorithm is proposed for view synthesis on multi-core platforms. Based on system level data parallelism, an adaptive workload balancing method is proposed for depth image based rendering by evaluating the number of non-hole pixels after warping. Experimental results demonstrated that with the proposed workload balancing algorithm, the workload difference among the cores is reduced by 90.65% on average for 2-core systems and by 79.57% on average for 4-core systems, respectively. Compared with the parallel system without the proposed balancing algorithm, synthesis speed is further improved by 7.5% for 2-core systems and 8.9% for 4-core systems at maximum, respectively, without degradation in the subjective and objective quality.
Zhanqi Liu, Xin Jin 0002, Qionghai Dai
ICASSP4
2015 Robust Non-rigid Motion Tracking and Surface Reconstruction Using L0 Regularization
abstract
We present a new motion tracking method to robustly reconstruct non-rigid geometries and motions from single view depth inputs captured by a consumer depth sensor. The idea comes from the observation of the existence of intrinsic articulated subspace in most of non-rigid motions. To take advantage of this characteristic, we propose a novel L0based motion regularizer with an iterative optimization solver that can implicitly constrain local deformation only on joints with articulated motions, leading to reduced solution space and physical plausible deformations. The L0strategy is integrated into the available non-rigid motion tracking pipeline, forming the proposed L0-L2non-rigid motion tracking method that can adaptively stop the tracking error propagation. Extensive experiments over complex human body motions with occlusions, face and hand motions demonstrate that our approach substantially improves tracking robustness and surface reconstruction accuracy.
Feng Xu 0005, Yangang Wang 0001, Yebin Liu, Qionghai Dai
ICCV5
2015 Depth estimation by analyzing intensity distribution for light-field cameras
abstract
In this paper, a novel depth estimation algorithm is proposed for light-field cameras by fully exploiting the characteristics of 4D light-field data. Based on a standard light-field acquisition system, a novel tensor, intensity range of pixels within a microlens, is proposed, which presents much stronger correlation with the transition on focus. Then, the tendency of intensity range is combined with the texture gradient to further improve the accuracy and consistency of the estimated depth and preserve edges through global optimization. Compared with the representative approaches, the depth generated by the proposed approach presents clearer object boundaries with much richer details both for indoor and outdoor contents. Moreover, the execution time is on average 12 times lower than the existing approaches.
Yatong Xu, Xin Jin 0002, Qionghai Dai
ICIP3
2015 Image super-resolution based on dictionary learning and anchored neighborhood regression with mutual incoherence
abstract
In this paper, we employ unified mutual coherence between the dictionary atoms and atoms/samples when learning the dictionary and sampling anchored neighborhoods respectively for image super-resolution (SR) application algorithm. On one hand, an incoherence promoting term in dictionary learning for SR is introduced to encourage dictionary atoms, associated to different anchored regressors, to be as independent as possible, while still allowing for different regressors to share same samples. On the other hand, a unified form with mutual coherence between dictionary atoms and training samples is proposed when we group neighborhoods of samples centered on each atom and find the nearest neighbors for input samples in image super-resolution. Extensive experimental results on commonly used datasets demonstrate that our method outperforms state-of-the-art methods by obtaining compelling results with improved quality, such as sharper edges, finer textures and higher structural similarity.
Yulun Zhang 0001, Kaiyu Gu, Yongbing Zhang 0002, Jian Zhang 0018, Qionghai Dai
ICIP5
2015 Generalized iterative phase retrieval algorithms and their applications
abstract
It is well known that the phase contains more important information about the field in comparison with the amplitude. Therefore the imaging of phase is encountered in many branches of modern science and engineering. Direct measurement of the phase is easy in the long wavelength regime of the electromagnetic spectrum, but is difficult in the short regime such as the visible light due to the limited bandwidth of imaging sensors. One must employ computational techniques to extract the phase from the captured intensity. So far many methods have been proposed for this task. These algorithms can be basically classified into three categories: Holography, deterministic algorithms such as the transport of intensity equation, and iterative algorithms such as the Gerchberg-Saxton-Fienup-type algorithm. Each of these algorithms has its own advantages and disadvantages. This paper mainly focuses on the our previous works on iterative phase retrieval techniques, and their applications in the calculation of computer-generated holograms, microscopic imaging, and optical signal processing.
Guohai Situ, Jin-Li Suo, Qionghai Dai
INDIN3
2015 Non-invasive imaging based on sparse representation
abstract
In this paper, we propose a method based on sparse representation to accelerate the scanning technique to perform non-invasive imaging through scattering layers. The scanning time is proportional to the size of the collected integrated intensity pattern and it usually takes tens of hours to finish the collecting process. To speed up the scanning technique, only a much smaller integrated intensity pattern is collected. A training set of integrated intensity pattern pairs with sizes corresponding to the necessary integrated intensity pattern that makes the scanning technique work well and the much smaller one is constructed. And a pair of dictionaries is trained from the constructed set exploiting the K-SVD algorithm. Based on sparse representation, the necessary integrated intensity pattern can be recovered from the much smaller one using the trained dictionaries, thus realizing non-invasive imaging successfully. Experimental results show that our method can successfully make the scanning time reduced by 8/9 without deteriorating the imaging quality of the scanning technique.
Kaiyun Wei, Xin Jin 0002, Yifu Hu, Qionghai Dai
MMSP4
2015 Region adaptive workload prediction for parallel view synthesis
abstract
In this paper, a parallel system together with a real-time workload balancing algorithm is proposed for view synthesis on multi-core platforms. First, a numerical relationship between the number of holes after warping and the workload of view synthesis is derived based on correlation analysis for the texture regions. Then, according to the location of the holes (whether they lie around an edge or at a homogeneous region), different models are proposed to predict the synthesis workload accurately. Experimental results show that the workload difference among cores is reduced largely and higher speedup ratio is achieved with negligible quality degradation by the proposed workload balancing system.
Zhanqi Liu, Xin Jin 0002, Qionghai Dai
VCIP3
2015 Adaptive local nonparametric regression for fast single image super-resolution
abstract
We propose a fast single image super-resolution algorithm based on adaptive local nonparametric regression. Making use of dictionary learning and regression, we learn multiple projection matrices mapping low-resolution features to their corresponding high-resolution ones directly. Different from previous linear regression that needs some constant parameters, our method would not use extra parameters for regression. We use the mutual coherence between dictionary atom and low-resolution feature as a label to reconstruct more sophisticated high-resolution feature. As we use the same form of mutual coherence as labels in both training and testing phases, our method would lead to an adaptive local linear regression model. Moreover, we investigate the statistical property of the dictionary atoms from the training features. Utilizing the learned statistical priors, our method would not only obtain more useful dictionary atoms, but also further decrease the computational time. As shown in our experimental results, the proposed method yields high-quality super-resolution images quantitatively and visually against state-of-the-art methods.
Yulun Zhang 0001, Yongbing Zhang 0002, Jian Zhang 0018, Haoqian Wang, Xingzheng Wang, Qionghai Dai
VCIP6
2015 Hybrid fusion and interpolation algorithm with near-infrared image
Xiaoyan Luo, Jun Zhang 0007, Qionghai Dai
Frontiers Comput. Sci.3
2015 Learning for 3D understanding
Yue Gao 0002, Rongrong Ji, Wei Liu 0005, Qionghai Dai
Neurocomputing4
2015 Structuring Lecture Videos by Automatic Projection Screen Localization and Analysis
abstract
We present a fully automatic system for extracting the semantic structure of a typical academic presentation video, which captures the whole presentation stage with abundant camera motions such as panning, tilting, and zooming. Our system automatically detects and tracks both the projection screen and the presenter whenever they are visible in the video. By analyzing the image content of the tracked screen region, our system is able to detect slide progressions and extract a high-quality, non-occluded, geometrically-compensated image for each slide, resulting in a list of representative images that reconstruct the main presentation structure. Afterwards, our system recognizes text content and extracts keywords from the slides, which can be used for keyword-based video retrieval and browsing. Experimental results show that our system is able to generate more stable and accurate screen localization results than commonly-used object tracking methods. Our system also extracts more accurate presentation structures than general video summarization methods, for this specific type of video.
Kai Li 0016, Jue Wang 0001, Haoqian Wang, Qionghai Dai
IEEE Trans. Pattern Anal. Mach. Intell.4
2015 A fast encoder of frame-compatible format based on content similarity for 3D distribution
Xin Jin 0002, Zhuoying Zeng, Satoshi Goto, Qionghai Dai
Signal Process. Image Commun.4
2015 Discriminative Clustering and Feature Selection for Brain MRI Segmentation
abstract
Automatic segmentation of brain tissues from MRI is of great importance for clinical application and scientific research. Recent advancements in supervoxel-level analysis enable robust segmentation of brain tissues by exploring the inherent information among multiple features extracted on the supervoxels. Within this prevalent framework, the difficulties still remain in clustering uncertainties imposed by the heterogeneity of tissues and the redundancy of the MRI features. To cope with the aforementioned two challenges, we propose a robust discriminative segmentation method from the view of information theoretic learning. The prominent goal of the method is to simultaneously select the informative feature and to reduce the uncertainties of supervoxel assignment for discriminative brain tissue segmentation. Experiments on two brain MRI datasets verified the effectiveness and efficiency of the proposed approach.
Youyong Kong, Yue Deng 0001, Qionghai Dai
IEEE Signal Process. Lett.3
2015 Extracting Depth and Radiance From a Defocused Video Pair
abstract
We present a novel iterative feedback approach for the simultaneous estimation of depth and all-in-focus (AIF) videos from a defocused video pair by joint spatiotemporal optimization. Depth and AIF videos benefit each other in the iterative optimization. First, for the recovery of AIF video, the sparse prior of natural video is incorporated to ensure a high-quality defocus blur removal even under inaccurate depth estimation. Second, in depth estimation step, we feed back the spatial and temporal constraints from the high-quality AIF video and adopt a numerical solution, which is robust to the inaccuracy of AIF recovery to further boost the performance of depth from the defocus algorithm. Benefitting from the incorporation of AIF video priors and the temporal consistency constraint, the proposed framework can effectively reconstruct the depth of the textureless region and is insensitive to camera parameter changes. Our approach provides better temporal consistency and higher depth accuracy than the conventional method that applies postsmoothing to the sequential frame estimation. We not only demonstrate the feasibility of our approach via real experimentation but also provide visual and quantitative evaluation on synthetic data.
Jin-Li Suo, Qionghai Dai
IEEE Trans. Circuits Syst. Video Technol.3
2015 Sparse Coding-Inspired Optimal Trading System for HFT Industry
abstract
The financial industry has witnessed an exceptionally fast progress of incorporating information processing techniques in designing knowledge-based automated systems for high-frequency trading (HFT). This paper proposes a sparse coding-inspired optimal trading (SCOT) system for real-time high-frequency financial signal representation and trading. Mathematically, SCOT simultaneously learns the dictionary, sparse features, and the trading strategy in a joint optimization, yielding optimal feature representations for the specific trading objective. The learning process is modeled as a bilevel optimization and solved by the online gradient descend method with fast convergence. In this dynamic context, the system is tested on the real financial market to trade the index futures in the Shanghai exchange center.
Yue Deng 0001, Youyong Kong, Feng Bao 0002, Qionghai Dai
IEEE Trans. Ind. Informatics4
2015 Toward Naturalistic 2D-to-3D Conversion
abstract
Natural scene statistics (NSSs) models have been developed that make it possible to impose useful perceptually relevant priors on the luminance, colors, and depth maps of natural scenes. We show that these models can be used to develop 3D content creation algorithms that can convert monocular 2D videos into statistically natural 3D-viewable videos. First, accurate depth information on key frames is obtained via human annotation. Then, both forward and backward motion vectors are estimated and compared to decide the initial depth values, and a compensation process is applied to further improve the depth initialization. Then, the luminance/chrominance and initial depth map are decomposed by a Gabor filter bank. Each subband of depth is modeled to produce a NSS prior term. The statistical color-depth priors are combined with the spatial smoothness constraint in the depth propagation target function as a prior regularizing term. The final depth map associated with each frame of the input 2D video is optimized by minimizing the target function over all subbands. In the end, stereoscopic frames are rendered from the color frames and their associated depth maps. We evaluated the quality of the generated 3D videos using both subjective and objective quality assessment methods. The experimental results obtained on various sequences show that the presented method outperforms several state-of-the-art 2D-to-3D conversion methods.
Xun Cao, Ke Lu 0002, Qionghai Dai, Alan C. Bovik
IEEE Trans. Image Process.4
2015 Full-Reference Quality Assessment of Stereoscopic Images by Learning Binocular Receptive Field Properties
abstract
Quality assessment of 3D images encounters more challenges than its 2D counterparts. Directly applying 2D image quality metrics is not the solution. In this paper, we propose a new full-reference quality assessment for stereoscopic images by learning binocular receptive field properties to be more in line with human visual perception. To be more specific, in the training phase, we learn a multiscale dictionary from the training database, so that the latent structure of images can be represented as a set of basis vectors. In the quality estimation phase, we compute sparse feature similarity index based on the estimated sparse coefficient vectors by considering their phase difference and amplitude difference, and compute global luminance similarity index by considering luminance changes. The final quality score is obtained by incorporating binocular combination based on sparse energy and sparse complexity. Experimental results on five public 3D image quality assessment databases demonstrate that in comparison with the most related existing methods, the devised algorithm achieves high consistency with subjective assessment.
Feng Shao 0001, Kemeng Li, Weisi Lin, Gangyi Jiang, Mei Yu 0001, Qionghai Dai
IEEE Trans. Image Process.6
2015 Depth Error Elimination for RGB-D Cameras
abstract
The rapid spreading of RGB-D cameras has led to wide applications of 3D videos in both academia and industry, such as 3D entertainment and 3D visual understanding. Under these circumstances, extensive research efforts have been dedicated to RGB-D camera--oriented topics. In these topics, quality promotion of depth videos with the temporal characteristic is emerging and important. Due to the limited exposure time of RGB-D cameras, object movement can easily lead to motion blurs in intensive images, which can further result in obvious artifacts (holes or fake boundaries) in the corresponding depth frames. With regard to this problem, we propose a depth error elimination method based on time series analysis to remove the artifacts in depth images. In this method, we first locate the regions with erroneous depths in intensive images by using motion blur detection based on a time series analysis model. This is based on the fact that the depth image is calculated by intensive color images that are captured synchronously by RGB-D cameras. Then, the artifacts, such as holes or fake boundaries, are fixed by a depth error elimination method. To evaluate the performance of the proposed method, we conducted experiments on 250 images. Experimental results demonstrate that the proposed method can locate the error regions correctly and eliminate these artifacts effectively. The quality of depth video can be improved significantly by using the proposed method.
Yue Gao 0002, You Yang 0002, Yi Zhen, Qionghai Dai
ACM Trans. Intell. Syst. Technol.4
2015 Probabilistic Skimlets Fusion for Summarizing Multiple Consumer Landmark Videos
abstract
It is difficult to develop a computational model that can accurately predict the quality of the video summary. This paper proposes a novel algorithm to summarize one-shot landmark videos. The algorithm can optimally combine multiple unedited consumer video skims into an aesthetically pleasing summary. In particular, to effectively select the representative key frames from multiple videos, an active learning algorithm is derived by taking advantage of the locality of the frames within each video. Toward a smooth video summary, we define skimlet, a video clip with adjustable length, starting frame, and positioned by each skim. Thereby, a probabilistic framework is developed to transfer the visual cues from a collection of aesthetically pleasing photos into the video summary. The length and the starting frame of each skimlet are calculated to maximally smoothen the video summary. At the same time, the unstable frames are removed from each skimlet. Experiments on multiple videos taken from different sceneries demonstrated the aesthetics, the smoothness, and the stability of the generated summary.
Yue Gao 0002, Richang Hong, Yuxing Hu, Rongrong Ji, Qionghai Dai
IEEE Trans. Multim.6
2014 DEPT: Depth Estimation by Parameter Transfer for Single Still Images
Xiu Li 0001, Hongwei Qin, Yangang Wang 0001, Yongbing Zhang 0002, Qionghai Dai
ACCV (2)5
2014 Action-Gons: Action Recognition with a Discriminative Dictionary of Structured Elements with Varying Granularity
Yuwang Wang, Baoyuan Wang, Yizhou Yu, Qionghai Dai, Zhuowen Tu
ACCV (5)4
2014 Fourier Analysis on Transient Imaging with a Multifrequency Time-of-Flight Camera
abstract
A transient image is the optical impulse response of a scene which visualizes light propagation during an ultra-short time interval. In this paper we discover that the data captured by a multifrequency time-of-flight (ToF) camera is the Fourier transform of a transient image, and identify the sources of systematic error. Based on the discovery we propose a novel framework of frequency-domain transient imaging, as well as algorithms to remove systematic error. The whole process of our approach is of much lower computational cost, especially lower memory usage, than Heide et al.'s approach using the same device. We evaluate our approach on both synthetic and real-datasets.
Yebin Liu, Matthias B. Hullin, Qionghai Dai
CVPR4
2014 Transparent Object Reconstruction via Coded Transport of Intensity
abstract
Capturing and understanding visual signals is one of the core interests of computer vision. Much progress has been made w.r.t. many aspects of imaging, but the reconstruc-tion of refractive phenomena, such as turbulence, gas and heat flows, liquids, or transparent solids, has remained a challenging problem. In this paper, we derive an intuitive formulation of light transport in refractive media using light fields and the transport of intensity equation. We show how coded illumination in combination with pairs of recorded images allow for robust computational reconstruction of dy-namic two and three-dimensional refractive phenomena. 1.
Chenguang Ma, Jin-Li Suo, Qionghai Dai, Gordon Wetzstein
CVPR4
2014 Hybrid Image Deblurring by Fusing Edge and Power Spectrum Information
Tao Yue 0003, Sunghyun Cho, Jue Wang 0001, Qionghai Dai
ECCV (7)4
2014 Recovering Scene Geometry under Wavy Fluid via Distortion and Defocus Analysis
Mohit Gupta 0001, Jin-Li Suo, Qionghai Dai
ECCV (5)5
2014 A novel distortion model for depth coding in 3D-HEVC
abstract
In 3D-HEVC, the latest 3D video coding project of MPEG, the coding mode of depth maps is determined by view synthesis optimization (VSO) which integrates view synthesis distortion into rate distortion optimization (RDO) to improve the compression efficiency and synthesis quality simultaneously. However, it introduces a big burden in coding complexity of depth maps. In this paper, a novel distortion model for depth coding in 3D-HEVC is proposed which estimates the distortion of synthesized views using an adaptive model for different pixel intervals. It outperforms existing depth distortion estimation algorithms by 16.6% BD-BR saving in max and 10.2% BD-BR saving on average. Compared with VSO, it reduces encoding complexity of depth maps by 46.2% with the highest efficiency-complexity performance among existing depth distortion estimation algorithms.
Xin Jin 0002, Qionghai Dai
ICIP3
2014 Automatic inpainting of linearly related video frames
abstract
This paper addresses automatic inpainting of a specific but common kind of videos captured by imaging a far or planar scene with a moving camera. The projective model tells that the frames of such videos can be approximately aligned by linear mappings except for some to-be-inpainted small regions. Mathematically, we treat inpainting as a global optimization with a linear system incorporating both the temporal consistency and the priors of the inpainting regions: (i) temporally registered frames form a low rank matrix; (ii) the pixels in the given inpainting regions destroy the low rank-ness with gross sparse errors. Besides, we also use a soft mask to ensure consistent global brightness before and after inpainting. Further, we propose a numerical solution to above optimization based on Augmented Lagrangian Method. The experiment results demonstrated our advantageous in both preserving thin scene structures and the details prone to be smoothed out by previous methods.
Yudong Xiao, Jin-Li Suo, Liheng Bian, Qionghai Dai
ICIP5
2014 Deblur a blurred RGB image with a sharp NIR image through local linear mapping
abstract
Image acquisition in a low light environment requires long exposure to achieve acceptable signal-to-noise ratio, which however causes blurry effect. This paper addresses this problem by using a sharp near-infrared (NIR) image when the environment has sufficient NIR light. We assume that an RGB and NIR image pair has a linear mapping in a local area and that the mapping function is valid for both the blur and sharp image pairs. Using this property, we solve the sharp RGB images from a blurred RGB image and the corres ponding s harp NIR image. The effectiveness of the proposed algorithm is verified with both synthetic and real captured datasets.
Tao Yue 0003, Ming-Ting Sun, Zhengyou Zhang, Jin-Li Suo, Qionghai Dai
ICME5
2014 Dynamic visual localization and tracking method based on RGB-D information
abstract
This paper proposes a new dynamic visual localization and tracking method using RGB-D camera. The proposed method makes use of band-width matrix, combines particle filter and mean shift with a new strategy, avoids falling into local optimum, and maintains particle diversity with only a few sampled points. A fast object searching strategy is successfully used to find the missing object. Experiments show that the proposed method runs robustly in complex scenes. The object 3-D parameters relative to the camera center can be estimated with a RGB-D camera, and that makes significant sense in spatial localization.
Chunxia Yin, Cai Luo, Qionghai Dai
ICRA4
2014 Robust Image Restoration via Reweighted Low-Rank Matrix Recovery
YiGang Peng, Jin-Li Suo, Qionghai Dai, Wenli Xu
MMM (1)3
2014 A fast coding algorithm based on inter-view correlations for 3D-HEVC
abstract
The newly published 3D-HEVC has received a remarkable response due to its high compression efficiency which is based on High Efficiency Video Coding (HEVC). However, the complexity of its encoding process is also large as a result of introducing the coding units (CU) size decision process together with the rate distortion optimization (RDO) process. In this paper, a fast coding algorithm making good use of the interview correlations is proposed. With the inter-view correlation statistical analysis, the CU depth candidates of the dependent views can be predicted from the independent view instead of the brute force RDO process in determining CU depth. The experimental results show that the proposed method saves 51% time in texture coding and the loss is negligible.
Guangsheng Chi, Xin Jin 0002, Qionghai Dai
VCIP3
2014 Depth map super-resolution via iterative joint-trilateral-upsampling
abstract
In this paper, we propose a new approach to solve the depth map super-resolution (SR) and denoising problems simultaneously. Inspired by joint-bilateral-upsampling (JBU), we devised the joint-trilateral-upsampling (JTU), which takes edge of the initial depth map, texture of the corresponding high-resolution color image and the values of the surrounding depth pixels, into consideration during the process of SR. To preserve the sharp edge of the up-sampled depth map and remove the noise, we introduce an iterative implementation, where current up-sampled depth map is fed into the next iteration, to refine the filter coefficients of JTU. The iterative JTU presents a high performance at many aspects such as sharping edge, denoising and none texture copying, etc. To demonstrate the superiority of the proposed method, we carry out various experiments and show an across-the-board quality improvement by both of subjective and objective evaluations compared with previous state-of-art methods.
Lei Zhang 0006, Yongbing Zhang 0002, Huiming Xuan, Qionghai Dai
VCIP5
2014 Synthesis-guided depth super resolution
abstract
Depth map, as important auxiliary information in 3D procession, is used to synthesize virtual view rather than exhibition. Inspired by this, a synthesis-guided depth super resolution (SGDSR) algorithm is proposed. Employing the synthesis error between virtual view and corresponding original one as the criteria, the best super-resolved result is selected among numerous candidate super resolution (SR) results. To fully exploit varying property within different regions of an image, a patch-based SGDSR is further devised in this paper. Experimental results demonstrate the effectiveness of our method subjectively and objectively on both single view and two views platform based on depth-image-based rendering (DIBR).
Huijin Lv, Yongbing Zhang 0002, Kai Li 0016, Xingzheng Wang, Huiming Xuan, Qionghai Dai
VCIP6
2014 Real-time air quality estimation based on color image processing
abstract
This paper address the problem of efficient, realtime estimation of the particulate mass concentration, exactly PM2.5 (particles with aerodynamic diameters less than 2.5 μm) from a superb view image. And the proposed method is to achieve high degree of accuracy at the cost of only modest user's effort by analyzing the relationship between the PM2.5 and the degradation of the observed image. With the fitting algorithm with experimental data, the PM2.5 could be real-time estimated by a general camera with little artificial participation, and the correlation coefficient produced by our data set and the standard observation will be as high as 0.8219, as the MSE (Mean Squared Error) value 51.2324 μg/m3.
Haoqian Wang, Xin Yuan 0002, Xingzheng Wang, Yongbing Zhang 0002, Qionghai Dai
VCIP5
2014 Acquisition of High Spatial and Spectral Resolution Video with a Hybrid Camera System
Chenguang Ma, Xun Cao, Xin Tong 0001, Qionghai Dai, Stephen Lin 0001
Int. J. Comput. Vis.4
2014 Decomposing Global Light Transport Using Time of Flight Imaging
Di Wu 0006, Andreas Velten, Matthew O'Toole, Belén Masiá, Amit K. Agrawal, Qionghai Dai, Ramesh Raskar
Int. J. Comput. Vis.6
2014 Ultra-fast Lensless Computational Imaging through 5D Frequency Analysis of Time-resolved Light Transport
Di Wu 0006, Gordon Wetzstein, Christopher Barsi, Thomas Willwacher, Qionghai Dai, Ramesh Raskar
Int. J. Comput. Vis.5
2014 Free-viewpoint video relighting from multi-view sequence under general illumination
Yebin Liu, Qionghai Dai
Mach. Vis. Appl.3
2014 Texture aided depth frame interpolation
Yongbing Zhang 0002, Jian Zhang 0018, Qionghai Dai
Signal Process. Image Commun.3
2014 Separable Coded Aperture for Depth from a Single Image
abstract
We propose the use of a separable coded aperture to estimate depth from a single defocused image, and derive a criterion for evaluating aperture patterns with respect to depth discrimination. With a separable coded aperture, two-dimensional (2D) point spread functions (PSFs) are separated into the product of horizontal and vertical one-dimensional (1D) PSFs. Depth is recovered by finding the scale of the 1D PSFs at each pixels based on the space-frequency analysis. Benefiting from 1D PSFs and the optimal code, our approach obtains more accurate depth especially on depth boundaries, and is faster since 2D computation is reduced to 1D. Experimental results on synthetic and real-captured data demonstrates the effectiveness of our approach.
Xiangyang Ji, Qionghai Dai
IEEE Signal Process. Lett.4
2014 A Parametric Model for Describing the Correlation Between Single Color Images and Depth Maps
abstract
This letter introduces a new approach for modeling the correlation between a single color image and its depth map with a set of parameters. The proposed model treats the color image as a set of patches and describes the correlation with a kernel function in a non-linear mapping space. We also present how to estimate the model parameters from sampled color image patches as well as the corresponding depth values. The proposed approach is tested on different color images and experimental results are comparable to the state-of-the-art, which demonstrates the power of the proposed method. Furthermore, we validate the efficiency of the proposed parametric model by evaluating each of its component, including the filters optimization, the choice of the patches and the kernel function.
Yangang Wang 0001, Ruiping Wang 0001, Qionghai Dai
IEEE Signal Process. Lett.3
2014 A Highly Parallel Framework for HEVC Coding Unit Partitioning Tree Decision on Many-core Processors
abstract
High Efficiency Video Coding (HEVC) uses a very flexible tree structure to organize coding units, which leads to a superior coding efficiency compared with previous video coding standards. However, such a flexible coding unit tree structure also places a great challenge for encoders. In order to fully exploit the coding efficiency brought by this structure, huge amount of computational complexity is needed for an encoder to decide the optimal coding unit tree for each image block. One way to achieve this is to use parallel computing enabled by many-core processors. In this paper, we analyze the challenge to use many-core processors to make coding unit tree decision. Through in-depth understanding of the dependency among different coding units, we propose a parallel framework to decide coding unit trees. Experimental results show that, on the Tile64 platform, our proposed method achieves averagely more than 11 and 16 times speedup for 1920x1080 and 2560x1600 video sequences, respectively, without any coding efficiency degradation.
Chenggang Yan 0001, Yongdong Zhang 0001, Jizheng Xu, Liang Li 0003, Qionghai Dai, Feng Wu 0001
IEEE Signal Process. Lett.6
2014 Efficient Parallel Framework for HEVC Motion Estimation on Many-Core Processors
abstract
High Efficiency Video Coding (HEVC) provides superior coding efficiency than previous video coding standards at the cost of increasing encoding complexity. The complexity increase of motion estimation (ME) procedure is rather significant, especially when considering the complicated partitioning structure of HEVC. To fully exploit the coding efficiency brought by HEVC requires a huge amount of computations. In this paper, we analyze the ME structure in HEVC and propose a parallel framework to decouple ME for different partitions on many-core processors. Based on local parallel method (LPM), we first use the directed acyclic graph (DAG)-based order to parallelize coding tree units (CTUs) and adopt improved LPM (ILPM) within each CTU (DAGILPM), which exploits the CTU-level and prediction unit (PU)-level parallelism. Then, we find that there exist completely independent PUs (CIPUs) and partially independent PUs (PIPUs). When the degree of parallelism (DP) is smaller than the maximum DP of DAGILPM, we process the CIPUs and PIPUs, which further increases the DP. The data dependencies and coding efficiency stay the same as LPM. Experiments show that on a 64-core system, compared with serial execution, our proposed scheme achieves more than 30 and 40 times speedup for 1920 × 1080 and 2560 × 1600 video sequences, respectively.
Chenggang Yan 0001, Yongdong Zhang 0001, Jizheng Xu, Jun Zhang 0007, Qionghai Dai, Feng Wu 0001
IEEE Trans. Circuits Syst. Video Technol.6
2014 Visual Words Assignment Via Information-Theoretic Manifold Embedding
abstract
Codebook-based learning provides a flexible way to extract the contents of an image in a data-driven manner for visual recognition. One central task in such frameworks is codeword assignment, which allocates local image descriptors to the most similar codewords in the dictionary to generate histogram for categorization. Nevertheless, existing assignment approaches, e.g., nearest neighbors strategy (hard assignment) and Gaussian similarity (soft assignment), suffer from two problems: 1) too strong Euclidean assumption and 2) neglecting the label information of the local descriptors. To address the aforementioned two challenges, we propose a graph assignment method with maximal mutual information (GAMI) regularization. GAMI takes the power of manifold structure to better reveal the relationship of massive number of local features by nonlinear graph metric. Meanwhile, the mutual information of descriptor-label pairs is ultimately optimized in the embedding space for the sake of enhancing the discriminant property of the selected codewords. According to such objective, two optimization models, i.e., inexact-GAMI and exact-GAMI, are respectively proposed in this paper. The inexact model can be efficiently solved with a closed-from solution. The stricter exact-GAMI nonparametrically estimates the entropy of descriptor-label pairs in the embedding space and thus leads to a relatively complicated but still trackable optimization. The effectiveness of GAMI models are verified on both the public and our own datasets.
Yue Deng 0001, Yanjun Qian, Xiangyang Ji, Qionghai Dai
IEEE Trans. Cybern.5
2014 Reweighted Low-Rank Matrix Recovery and its Application in Image Restoration
abstract
In this paper, we propose a reweighted low-rank matrix recovery method and demonstrate its application for robust image restoration. In the literature, principal component pursuit solves low-rank matrix recovery problem via a convex program of mixed nuclear norm and l1 norm. Inspired by reweighted l1 minimization for sparsity enhancement, we propose reweighting singular values to enhance low rank of a matrix. An efficient iterative reweighting scheme is proposed for enhancing low rank and sparsity simultaneously and the performance of low-rank matrix recovery is prompted greatly. We demonstrate the utility of the proposed method both on numerical simulations and real images/videos restoration, including single image restoration, hyperspectral image restoration, and background modeling from corrupted observations. All of these experiments give empirical evidence on significant improvements of the proposed algorithm over previous work on low-rank matrix recovery.
YiGang Peng, Jin-Li Suo, Qionghai Dai, Wenli Xu
IEEE Trans. Cybern.3
2014 Hyperspectral Image Classification Through Bilayer Graph-Based Learning
abstract
Hyperspectral image classification with limited number of labeled pixels is a challenging task. In this paper, we propose a bilayer graph-based learning framework to address this problem. For graph-based classification, how to establish the neighboring relationship among the pixels from the high dimensional features is the key toward a successful classification. Our graph learning algorithm contains two layers. The first-layer constructs a simple graph, where each vertex denotes one pixel and the edge weight encodes the similarity between two pixels. Unsupervised learning is then conducted to estimate the grouping relations among different pixels. These relations are subsequently fed into the second layer to form a hypergraph structure, on top of which, semisupervised transductive learning is conducted to obtain the final classification results. Our experiments on three data sets demonstrate the merits of our proposed approach, which compares favorably with state of the art.
Yue Gao 0002, Rongrong Ji, Peng Cui 0001, Qionghai Dai, Gang Hua 0001
IEEE Trans. Image Process.4
2014 Weakly Supervised Visual Dictionary Learning by Harnessing Image Attributes
abstract
Bag-of-features (BoFs) representation has been extensively applied to deal with various computer vision applications. To extract discriminative and descriptive BoF, one important step is to learn a good dictionary to minimize the quantization loss between local features and codewords. While most existing visual dictionary learning approaches are engaged with unsupervised feature quantization, the latest trend has turned to supervised learning by harnessing the semantic labels of images or regions. However, such labels are typically too expensive to acquire, which restricts the scalability of supervised dictionary learning approaches. In this paper, we propose to leverage image attributes to weakly supervise the dictionary learning procedure without requiring any actual labels. As a key contribution, our approach establishes a generative hidden Markov random field (HMRF), which models the quantized codewords as the observed states and the image attributes as the hidden states, respectively. Dictionary learning is then performed by supervised grouping the observed states, where the supervised information is stemmed from the hidden states of the HMRF. In such a way, the proposed dictionary learning approach incorporates the image attributes to learn a semantic-preserving BoF representation without any genuine supervision. Experiments in large-scale image retrieval and classification tasks corroborate that our approach significantly outperforms the state-of-the-art unsupervised dictionary learning approaches.
Yue Gao 0002, Rongrong Ji, Wei Liu 0005, Qionghai Dai, Gang Hua 0001
IEEE Trans. Image Process.4
2014 Joint Non-Gaussian Denoising and Superresolving of Raw High Frame Rate Videos
abstract
High frame rate cameras capture sharp videos of highly dynamic scenes by trading off signal-noise-ratio and image resolution, so combinational super-resolving and denoising is crucial for enhancing high speed videos and extending their applications. The solution is nontrivial due to the fact that two deteriorations co-occur during capturing and noise is nonlinearly dependent on signal strength. To handle this problem, we propose conducting noise separation and super resolution under a unified optimization framework, which models both spatiotemporal priors of high quality videos and signal-dependent noise. Mathematically, we align the frames along temporal axis and pursue the solution under the following three criterion: 1) the sharp noise-free image stack is low rank with some missing pixels denoting occlusions; 2) the noise follows a given nonlinear noise model; and 3) the recovered sharp image can be reconstructed well with sparse coefficients and an over complete dictionary learned from high quality natural images. In computation aspects, we propose to obtain the final result by solving a convex optimization using the modern local linearization techniques. In the experiments, we validate the proposed approach in both synthetic and real captured data.
Jin-Li Suo, Yue Deng 0001, Liheng Bian, Qionghai Dai
IEEE Trans. Image Process.4
2014 High-Dimensional Camera Shake Removal With Given Depth Map
abstract
Camera motion blur is drastically nonuniform for large depth-range scenes, and the nonuniformity caused by camera translation is depth dependent but not the case for camera rotations. To restore the blurry images of large-depth-range scenes deteriorated by arbitrary camera motion, we build an image blur model considering 6-degrees of freedom (DoF) of camera motion with a given scene depth map. To make this 6D depth-aware model tractable, we propose a novel parametrization strategy to reduce the number of variables and an effective method to estimate high-dimensional camera motion as well. The number of variables is reduced by temporal sampling motion function, which describes the 6-DoF camera motion by sampling the camera trajectory uniformly in time domain. To effectively estimate the high-dimensional camera motion parameters, we construct the probabilistic motion density function (PMDF) to describe the probability distribution of camera poses during exposure, and apply it as a unified constraint to guide the convergence of the iterative deblurring algorithm. Specifically, PMDF is computed through a back projection from 2D local blur kernels to 6D camera motion parameter space and robust voting. We conduct a series of experiments on both synthetic and real captured data, and validate that our method achieves better performance than existing uniform methods and nonuniform methods on large-depth-range scenes.
Tao Yue 0003, Jin-Li Suo, Qionghai Dai
IEEE Trans. Image Process.3
2014 Actively Learning Human Gaze Shifting Paths for Semantics-Aware Photo Cropping
abstract
Photo cropping is a widely used tool in printing industry, photography, and cinematography. Conventional cropping models suffer from the following three challenges. First, the deemphasized role of semantic contents that are many times more important than low-level features in photo aesthetics. Second, the absence of a sequential ordering in the existing models. In contrast, humans look at semantically important regions sequentially when viewing a photo. Third, the difficulty of leveraging inputs from multiple users. Experience from multiple users is particularly critical in cropping as photo assessment is quite a subjective task. To address these challenges, this paper proposes semantics-aware photo cropping, which crops a photo by simulating the process of humans sequentially perceiving semantically important regions of a photo. We first project the local features (graphlets in this paper) onto the semantic space, which is constructed based on the category information of the training photos. An efficient learning algorithm is then derived to sequentially select semantically representative graphlets of a photo, and the selecting process can be interpreted by a path, which simulates humans actively perceiving semantics in a photo. Furthermore, we learn a prior distribution of such active graphlet paths from training photos that are marked as aesthetically pleasing by multiple users. The learned priors enforce the corresponding active graphlet path of a test photo to be maximally similar to those from the training photos. Experimental results show that: 1) the active graphlet path accurately predicts human gaze shifting, and thus is more indicative for photo aesthetics than conventional saliency maps and 2) the cropped photos produced by our approach outperform its competitors in both qualitative and quantitative comparisons.
Yue Gao 0002, Rongrong Ji, Yingjie Xia, Qionghai Dai, Xuelong Li 0001
IEEE Trans. Image Process.5
2014 A Data-Driven Approach for Facial Expression Retargeting in Video
abstract
This paper presents a data-driven approach for facial expression retargeting in video, i.e., synthesizing a face video of a target subject that mimics the expressions of a source subject in the input video. Our approach takes advantage of a pre-existing facial expression database of the target subject to achieve realistic synthesis. First, for each frame of the input video, a new facial expression similarity metric is proposed for querying the expression database of the target person to select multiple candidate images that are most similar to the input. The similarity metric is developed using a metric learning approach to reliably handle appearance difference between different subjects. Secondly, we employ an optimization approach to choose the best candidate image for each frame, resulting in a retrieved sequence that is temporally coherent. Finally, a spatio-temporal expression mapping method is employed to further improve the synthesized sequence. Experimental results show that our system is capable of generating high quality facial expression videos that match well with the input sequences, even when the source and target subjects have big identity difference. In addition, extensive evaluations demonstrate the high accuracy of the learned expression similarity metric and the effectiveness of our retrieval strategy.
Kai Li 0016, Qionghai Dai, Ruiping Wang 0001, Yebin Liu, Feng Xu 0005, Jue Wang 0001
IEEE Trans. Multim.2
2014 Spatial-spectral encoded compressive hyperspectral imaging
abstract
This paper proposes a novel compressive hyperspectral (HS) imaging approach that allows for high-resolution HS images to be captured in a single image. The proposed architecture comprises three key components: spatial-spectral encoded optical camera design, over-complete HS dictionary learning and sparse-constraint computational reconstruction. Our spatial-spectral encoded sampling scheme provides a higher degree of randomness in the measured projections than previous compressive HS imaging approaches; and a robust nonlinear sparse reconstruction method is employed to recover the HS images from the coded projection with higher performance. To exploit the sparsity constraint on the nature HS images for computational reconstruction, an over-complete HS dictionary is learned to represent the HS images in a sparser way than previous representations. We validate the proposed approach on both synthetic and real captured data, and show successful recovery of HS images for both indoor and outdoor scenes. In addition, we demonstrate other applications for the over-complete HS dictionary and sparse coding techniques, including 3D HS images compression and denoising.
Yebin Liu, Qionghai Dai
ACM Trans. Graph.4
2014 Intrinsic video and applications
abstract
We present a method to decompose a video into its intrinsic components of reflectance and shading, plus a number of related example applications in video editing such as segmentation, stylization, material editing, recolorization and color transfer. Intrinsic decomposition is an ill-posed problem, which becomes even more challenging in the case of video due to the need for temporal coherence and the potentially large memory requirements of a global approach. Additionally, user interaction should be kept to a minimum in order to ensure efficiency. We propose a probabilistic approach, formulating a Bayesian Maximum a Posteriori problem to drive the propagation of clustered reflectance values from the first frame, and defining additional constraints as priors on the reflectance and shading. We explicitly leverage temporal information in the video by building a causal-anticausal, coarse-to-fine iterative scheme, and by relying on optical flow information. We impose no restrictions on the input video, and show examples representing a varied range of difficult cases. Our method is the first one designed explicitly for video; moreover, it naturally ensures temporal consistency, and compares favorably against the state of the art in this regard.
Genzhi Ye, Elena Garces 0001, Yebin Liu, Qionghai Dai, Diego Gutierrez
ACM Trans. Graph.4
2013 Coded focal stack photography
abstract
We present coded focal stack photography as a computational photography paradigm that combines a focal sweep and a coded sensor readout with novel computational algorithms. We demonstrate various applications of coded focal stacks, including photography with programmable non-planar focal surfaces and multiplexed focal stack acquisition. By leveraging sparse coding techniques, coded focal stacks can also be used to recover a full-resolution depth and all-in-focus (AIF) image from a single photograph. Coded focal stack photography is a significant step towards a computational camera architecture that facilitates high-resolution post-capture refocusing, flexible depth of field, and 3D imaging.
Jin-Li Suo, Gordon Wetzstein, Qionghai Dai, Ramesh Raskar
ICCP4
2013 High-rank coded aperture projection for extended depth of field
abstract
Projectors require large apertures to maximize light throughput. Unfortunately, this leads to shallow depths of field (DOF), hence blurry images, when projecting on non-planar surfaces, such as cultural heritage sites, curved screens, or when sharing visual information in everyday environments. We introduce high-rank coded aperture projectors - a new computational display technology that combines optical designs with computational processing to overcome depth of field limitations of conventional devices. In particular, we employ high-speed spatial light modulators (SLMs) on the image plane and in the aperture of modified projectors. The patterns displayed on these SLMs are computed with a new mathematical framework that uses high-rank light field factorizations and directly exploits the limited temporal resolution and contrast sensitivity of the human visual system. With an experimental prototype projector, we demonstrate significantly increased DOF as compared to conventional technology.
Chenguang Ma, Jin-Li Suo, Qionghai Dai, Ramesh Raskar, Gordon Wetzstein
ICCP3
2013 Non-uniform image deblurring using an optical computing system
Tao Yue 0003, Jin-Li Suo, Qionghai Dai
Comput. Graph.3
2013 A Progressive Tri-level Segmentation Approach for Topology-Change-Aware Video Matting
abstract
Abstract Previous video matting approaches mostly adopt the “binary segmentation + matting” strategy, i.e., first segment each frame into foreground and background regions, then extract the fine details of the foreground boundary using matting techniques. This framework has several limitations due to the fact that binary segmentation is employed. In this paper, we propose a new supervised video matting approach. Instead of applying binary segmentation, we explicitly model segmentation uncertainty in a novel tri‐level segmentation procedure. The segmentation is done progressively, enabling us to handle difficult cases such as large topology changes, which are challenging to previous approaches. The tri‐level segmentation results can be naturally fed into matting techniques to generate the final alpha mattes. Experimental results show that our system can generate high quality results with less user inputs than the state‐of‐theart methods.
Jinlong Ju, Jue Wang 0001, Yebin Liu, Haoqian Wang, Qionghai Dai
Comput. Graph. Forum5
2013 Capturing Relightable Human Performances under General Uncontrolled Illumination
abstract
Abstract We present a novel approach to create relightable free‐viewpoint human performances from multi‐view video recorded under general uncontrolled and uncalibated illumination. We first capture a multi‐view sequence of an actor wearing arbitrary apparel and reconstruct a spatio‐temporal coherent coarse 3D model of the performance using a marker‐less tracking approach. Using these coarse reconstructions, we estimate the low‐frequency component of the illumination in a spherical harmonics (SH) basis as well as the diffuse reflectance, and then utilize them to estimate the dynamic geometry detail of human actors based on shading cues. Given the high‐quality time‐varying geometry, the estimated illumination is extended to the all‐frequency domain by re‐estimating it in the wavelet basis. Finally, the high‐quality all‐frequency illumination is utilized to reconstruct the spatially‐varying BRDF of the surface. The recovered time‐varying surface geometry and spatially‐varying non‐Lambertian reflectance allow us to generate high‐quality model‐based free view‐point videos of the actor under novel illumination conditions. Our method enables plausible reconstruction of relightable dynamic scene models without a complex controlled lighting apparatus, and opens up a path towards relightable performance capture in less constrained environments and using less complex acquisition setups.
Chenglei Wu, Carsten Stoll, Yebin Liu, Kiran Varanasi, Qionghai Dai, Christian Theobalt
Comput. Graph. Forum6
2013 Robust blind motion deblurring using near-infrared flash image
Jun Zhang 0007, Qionghai Dai
J. Vis. Commun. Image Represent.3
2013 Markerless Motion Capture of Multiple Characters Using Multiview Image Segmentation
abstract
Capturing the skeleton motion and detailed time-varying surface geometry of multiple, closely interacting peoples is a very challenging task, even in a multicamera setup, due to frequent occlusions and ambiguities in feature-to-person assignments. To address this task, we propose a framework that exploits multiview image segmentation. To this end, a probabilistic shape and appearance model is employed to segment the input images and to assign each pixel uniquely to one person. Given the articulated template models of each person and the labeled pixels, a combined optimization scheme, which splits the skeleton pose optimization problem into a local one and a lower dimensional global one, is applied one by one to each individual, followed with surface estimation to capture detailed nonrigid deformations. We show on various sequences that our approach can capture the 3D motion of humans accurately even if they move rapidly, if they wear wide apparel, and if they are engaged in challenging multiperson motions, including dancing, wrestling, and hugging.
Yebin Liu, Juergen Gall, Carsten Stoll, Qionghai Dai, Hans-Peter Seidel, Christian Theobalt
IEEE Trans. Pattern Anal. Mach. Intell.4
2013 Retinex based visual identicalness detection for videos corrupted by imaging noise
Xin Jin 0002, Satoshi Goto, Qionghai Dai
Signal Process. Image Commun.3
2013 Complexity Reduction and Performance Improvement for Geometry Partitioning in Video Coding
abstract
Geometry partitioning for video coding involves establishing a partition line boundary within each block-shaped region and applying motion-compensated prediction to the two sub-regions created by the partition line. This paper presents techniques for enhancing the effectiveness and reducing the complexity of geometry partitioning schemes. A texture-difference-based approach is described to simplify the process of selecting the partition lines. Applying this approach together with a described skipping strategy for blocks with uniform texture can achieve a 94% reduction of encoding time while retaining a similar rate-distortion (R-D) performance to the full-search partitioning approach, when implemented for wedge-based geometry partitioning (WGP) in the context of H.264/MPEG-4 AVC JM 16.2. A bit rate improvement of approximately 6% is shown relative to not using geometry partitioning. For further R-D improvement, we describe a background-compensated prediction scheme to reduce the number of overhead bits used for motion vectors. Additionally, for systems in which high-quality depth maps are available, we incorporate depth map usage into the described approaches to generate a more accurate partitioning. Using these approaches with object-boundary-based geometry partitioning can achieve about 9% bit rate savings relative to using WGP, while keeping a similar computational complexity to the described complexity-reduced WGP.
Qifei Wang, Xiangyang Ji, Ming-Ting Sun, Gary J. Sullivan, Qionghai Dai
IEEE Trans. Circuits Syst. Video Technol.6
2013 Stereo Interleaving Video Coding With Content Adaptive Image Subsampling
abstract
Stereo interleaving video coding, in which both left and right view frames are subsampled into half size and multiplexed into one single frame before being encoded by a traditional 2-D video encoder, is an efficient encoding scenario for stereoscopic video. Many existing stereo interleaving video coding methods subsample each frame by utilizing fixed subsampling filter coefficients. Such methods are easy to implement; however, the varying property of the frame signal is ignored. By jointly considering the influences of subsampling and compression, a rate and distortion analysis about stereo interleaving video coding is proposed. The final distortion in stereo interleaving video coding is the summation of errors caused by subsampling (causing distortion between subsampling-interpolated image and the original full resolution one) and by quantization during compression. Based on the provided rate distortion analysis, a content adaptive image subsampling (CAIS) is also proposed. In CAIS, the half-size frames are generated by the optimal subsampling filters, which are calculated based on frame contents and the targeted interpolation coefficients. Experimental results demonstrate that the proposed CAIS is able to greatly improve compression efficiency of stereo interleaving video coding.
Yongbing Zhang 0002, Xiangyang Ji, Haoqian Wang, Qionghai Dai
IEEE Trans. Circuits Syst. Video Technol.4
2013 Free-Viewpoint Video of Human Actors Using Multiple Handheld Kinects
abstract
We present an algorithm for creating free-viewpoint video of interacting humans using three handheld Kinect cameras. Our method reconstructs deforming surface geometry and temporal varying texture of humans through estimation of human poses and camera poses for every time step of the RGBZ video. Skeletal configurations and camera poses are found by solving a joint energy minimization problem, which optimizes the alignment of RGBZ data from all cameras, as well as the alignment of human shape templates to the Kinect data. The energy function is based on a combination of geometric correspondence finding, implicit scene segmentation, and correspondence finding using image features. Finally, texture recovery is achieved through jointly optimization on spatio-temporal RGB data using matrix completion. As opposed to previous methods, our algorithm succeeds on free-viewpoint video of human actors under general uncontrolled indoor scenes with potentially dynamic background, and it succeeds even if the cameras are moving.
Genzhi Ye, Yebin Liu, Yue Deng 0001, Nils Hasler, Xiangyang Ji, Qionghai Dai, Christian Theobalt
IEEE Trans. Cybern.6
2013 Absolute Depth Estimation From a Single Defocused Image
abstract
Shape from defocus (SFD) is one of the most popular techniques in monocular 3D vision. While most SFD approaches require two or more images of the same scene captured at a fixed view point, this paper presents an efficient approach to estimate absolute depth from a single defocused image. Instead of directly measuring defocus level of each pixel, we propose to design a sequence of aperture-shape filters to segment a defocused image by defocus level. A boundary-weighted belief propagation algorithm is employed to obtain a smooth depth map. We also give an estimation of depth error. Extensive experiments show that our approach outperforms the state-of-the-art single-image SFD approaches both in precision of the estimated absolute depth and running time.
Xiangyang Ji, Wenli Xu, Qionghai Dai
IEEE Trans. Image Process.4
2013 Joint Bit Allocation and Rate Control for Coding Multi-View Video Plus Depth Based 3D Video
abstract
In three-dimensional (3D) video coding, distortion in texture video and depth maps can all affect the quality of the synthesized virtual views. Therefore, under the total bitrate constraint, effective bit allocation between texture and depth information is very important for 3D video coding. In this paper, the major technical contribution is to formulate view synthesis quality for optimal resource allocation in 3D video coding, since such quality is what that matters most to the ultimate user (i.e., the viewer) of the system; to be more specific, a new joint bit allocation and rate control method for multi-view video plus depth (MVD) based 3D video coding is proposed accordingly. We firstly derive a view synthesis distortion model to characterize the effect of coding distortion of texture video and depth maps on the synthesized virtual views. Based on this model, we derive a rate-distortion model to characterize the relationship between the bitrate and the view synthesis distortion, and the optimal bitrate ratio between texture and depth is established adaptively by solving the associated optimization problem. Finally, the rate control algorithm is performed on view level, texture/depth level and frame level. Experimental results show that compared with other methods, the proposed bit allocation method obtains higher performance of view synthesis. Moreover, the proposed rate control method can accurately control the bitrate to satisfy the total bitrate constraint.
Feng Shao 0001, Gangyi Jiang, Weisi Lin, Mei Yu 0001, Qionghai Dai
IEEE Trans. Multim.5
2013 Low-Rank Structure Learning via Nonconvex Heuristic Recovery
abstract
In this paper, we propose a nonconvex framework to learn the essential low-rank structure from corrupted data. Different from traditional approaches, which directly utilizes convex norms to measure the sparseness, our method introduces more reasonable nonconvex measurements to enhance the sparsity in both the intrinsic low-rank structure and the sparse corruptions. We will, respectively, introduce how to combine the widely used ℓp norm (0 < p < 1) and log-sum term into the framework of low-rank structure learning. Although the proposed optimization is no longer convex, it still can be effectively solved by a majorization-minimization (MM)-type algorithm, with which the nonconvex objective function is iteratively replaced by its convex surrogate and the nonconvex problem finally falls into the general framework of reweighed approaches. We prove that the MM-type algorithm can converge to a stationary point after successive iterations. The proposed model is applied to solve two typical problems: robust principal component analysis and low-rank representation. Experimental results on low-rank structure learning demonstrate that our nonconvex heuristic methods, especially the log-sum heuristic recovery algorithm, generally perform much better than the convex-norm-based method (0 < p < 1) for both data with higher rank and with denser corruptions.
Yue Deng 0001, Qionghai Dai, Risheng Liu, Zengke Zhang, Sanqing Hu
IEEE Trans. Neural Networks Learn. Syst.2
2013 Video-based hand manipulation capture through composite motion control
abstract
This paper describes a new method for acquiring physically realistic hand manipulation data from multiple video streams. The key idea of our approach is to introduce a composite motion control to simultaneously model hand articulation, object movement, and subtle interaction between the hand and object. We formulate video-based hand manipulation capture in an optimization framework by maximizing the consistency between the simulated motion and the observed image data. We search an optimal motion control that drives the simulation to best match the observed image data. We demonstrate the effectiveness of our approach by capturing a wide range of high-fidelity dexterous manipulation data. We show the power of our recovered motion controllers by adapting the captured motion data to new objects with different properties. The system achieves superior performance against alternative methods such as marker-based motion capture and kinematic hand motion tracking.
Yangang Wang 0001, Jianyuan Min, Jianjie Zhang, Yebin Liu, Feng Xu 0005, Qionghai Dai, Jinxiang Chai
ACM Trans. Graph.6
2012 Iterative Feedback Estimation of Depth and Radiance from Defocused Images
Jin-Li Suo, Xun Cao, Qionghai Dai
ACCV (4)4
2012 Visual words assignment on a graph via minimal mutual information loss
abstract
Visual codewords assignment plays an important role in many Bag of Features (BoF) models for image understanding and visual recognition. It allocates image descriptors to the most similar codewords in the pre-configured visual dictionary to generate descriptive histogram for the consequent categorization. Nevertheless, existing assignment approaches, e.g. nearest neighbors strategy and Gaussian similarity, suffer from two problems:1) too strong Euclidean assumption and 2) neglecting the label information of the local features. Accordingly, in this paper, we propose an assignment method to simultaneously consider the above two issues in a unified model via graph learning and information theoretic criterions. For learning, the proposed model can be efficiently solved in a closed-form with the reasonable graph topology invariant approximation. Moreover, the learned projections enable us to extend the assignment ability to the out-of-sample visual features beyond the initial training graph. Experiments on our own manifold dataset and two benchmarks verify the effectiveness of the proposed graph assignment method.
Yanjun Qian, Yue Deng 0001, Qionghai Dai, Guihua Er
BMVC3
2012 A data-driven approach for facial expression synthesis in video
abstract
This paper presents a method to synthesize a realistic facial animation of a target person, driven by a facial performance video of another person. Different from traditional facial animation approaches, our system takes advantage of an existing facial performance database of the target person, and generates the final video by retrieving frames from the database that have similar expressions to the input ones. To achieve this we develop an expression similarity metric for accurately measuring the expression difference between two video frames. To enforce temporal coherence, our system employs a shortest path algorithm to choose the optimal image for each frame from a set of candidate frames determined by the similarity metric. Finally, our system adopts an expression mapping method to further minimize the expression difference between the input and retrieved frames. Experimental results show that our system can generate high quality facial animation using the proposed data-driven approach.
Kai Li 0016, Feng Xu 0005, Jue Wang 0001, Qionghai Dai, Yebin Liu
CVPR4
2012 Covariance discriminative learning: A natural and efficient approach to image set classification
abstract
We propose a novel discriminative learning approach to image set classification by modeling the image set with its natural second-order statistic, i.e. covariance matrix. Since nonsingular covariance matrices, a.k.a. symmetric positive definite (SPD) matrices, lie on a Riemannian manifold, classical learning algorithms cannot be directly utilized to classify points on the manifold. By exploring an efficient metric for the SPD matrices, i.e., Log-Euclidean Distance (LED), we derive a kernel function that explicitly maps the covariance matrix from the Riemannian manifold to a Euclidean space. With this explicit mapping, any learning method devoted to vector space can be exploited in either its linear or kernel formulation. Linear Discriminant Analysis (LDA) and Partial Least Squares (PLS) are considered in this paper for their feasibility for our specific problem. We further investigate the conventional linear subspace based set modeling technique and cast it in a unified framework with our covariance matrix based modeling. The proposed method is evaluated on two tasks: face recognition and object categorization. Extensive experimental results show not only the superiority of our method over state-of-the-art ones in both accuracy and efficiency, but also its stability to two real challenges: noisy set data and varying set size.
Ruiping Wang 0001, Huimin Guo, Larry Davis 0001, Qionghai Dai
CVPR4
2012 A Single Frame Super-Resolution Method Based on Matrix Completion
abstract
Efficiently exploring the linear relationship among neighboring pixels is a pervasive way to reconstruct high-resolution image from low-resolution one. However, it is a challenge to determine the order of linear model. According to the theory of matrix completion, we propose a single frame super-resolution algorithm by minimizing the sum of all the augmented matrices' rank, which can reflect the order of the region aware linear model. Various experiments demonstrate the images reconstructed by the proposed method have superior PSNR and visual quality, benefitting from its desirable ability of depressing the ringing noise and other artifacts.
Changjun Fu, Xiangyang Ji, Yongbing Zhang 0002, Qionghai Dai
DCC4
2012 Packet Video Error Concealment Based on Compressed Sensing and Regularized Least Squares
abstract
Error concealment (EC) is an important post processing technique to deal with the packet loss during the transmission of compressed video stream. This paper aims to address the problem of recovering the missing block in the decoded video stream from the perspective of compressed sensing. The missing block is assumed to be sparsely represented by a dictionary of prototype signal atoms. The atoms are generated by the motion-compensated blocks with a range of motion displacements from the temporally previously reconstructed frame. To avoid inefficient exploration for the prior of sparsity due to the potential coherency among atoms, the regularized least square is incorporated into the compressed sensing reconstruction for the recovery of the missing block. Experimental results demonstrate the superiority of the proposed EC method in terms of objective (PSNR) and subjective quality compared to the existing methods.
Changjun Fu, Xiangyang Ji, Yongbing Zhang 0002, Qionghai Dai
DCC4
2012 Content Adaptive Subsampling for Stereo Interleaving Video Coding
abstract
Stereo interleaving video coding receives considerable attention due to its desirable property of being compatible with 2D video coding standards. The errors caused by sub sampling (causing distortion between subsampling interpolated image and the original full resolution one) and by quantization during compression lead to the final distortion in stereo interleaving video coding. In this paper, the rate and distortion analysis in stereo interleaving video coding is provided. It proves that appropriate sub sampling in stereo interleaving video coding is able to obtain good compression performance. Subsequently, a content adaptive sub sampling (CAS) is proposed. In CAS, the half resolution frames are generated by decimation, where the down sampling filter coefficients are calculated based on frame contents and the targeted interpolation coefficients. Experiment results demonstrate that the CAS is able to achieve high compression efficiency of stereo interleaving encoding scheme for stereoscopic videos.
Yongbing Zhang 0002, Xiangyang Ji, Haoqian Wang, Lei Zhang 0006, Qionghai Dai
DCC5
2012 Frequency Analysis of Transient Light Transport with Applications in Bare Sensor Imaging
Di Wu 0006, Gordon Wetzstein, Christopher Barsi, Thomas Willwacher, Matthew O'Toole, Nikhil Naik 0003, Qionghai Dai, Kiriakos N. Kutulakos, Ramesh Raskar
ECCV (1)7
2012 Performance Capture of Interacting Characters with Handheld Kinects
Genzhi Ye, Yebin Liu, Nils Hasler, Xiangyang Ji, Qionghai Dai, Christian Theobalt
ECCV (2)5
2012 3D spatial reconstruction and communication from vision field
abstract
Vision field describes the real world visual information by summarizing the seven-dimensional plenoptic function into three domains: view, light, time, from which a better understanding of previous 3D capture and reconstruction systems can be provided. In this paper, we first show how to reconstruct 3D spatial information from all the three attributes of the vision field, namely full-space vision field reconstruction. Then, based on Laplacian iterative geometry prediction, a 3D mesh coding algorithm with cascaded quantization is presented to facilitate the communication of the reconstructed 3D models from vision field. At last, experimental results of both the 3D spatial reconstruction and the 3D mesh coding are demonstrated.
Xun Cao, Qifei Wang, Xiangyang Ji, Qionghai Dai
ICASSP4
2012 A novel method for 2D-to-3D video conversion using bi-directional motion estimation
abstract
In this paper, we proposed a novel semi-automatic 2D-to-3D video conversion method. Our method requires just a few user-scribbles to generate depth maps for key frames and propagates these depth maps to non-key frames automatically. For key frames, foreground objects and the corresponding depth maps can be obtained by an interactive method. Then, both forward and backward motion vectors are estimated and compared to decide the depth propagation strategy. For pixels that failed the motion vectors comparison, a compensation process is adopted to refine their depth propagation results. Finally, stereoscopic pairs are generated by the warping method based on the original frames and associated depth maps. Our method is validated by both subjective and objective quality assessments. The experimental results show that our method outperforms several state-of-the-art 2D-to-3D video conversion methods.
Zhenyao Li, Xun Cao, Qionghai Dai
ICASSP3
2012 Geometric mapping assisted multi-view depth video coding
abstract
Multi-view plus depth (MVD), as a video representation supporting view synthesis based on depth video, has attracted more and more attention for the free view video (FVV) application. It is a challenge to efficiently compress the multi-view depth data in MVD format. In this paper, we explore the geometric relationships in 3D space and propose a geometric mapping assisted (GMA) multi-view depth video coding algorithm. The proposed GMA utilizes the mapped depth image as a reference candidate during prediction. Furthermore, the inpainting method is employed to fill in the holes in mapped depth images. Experimental results demonstrate the gains of up to 2.45 dB for the depth coding, as well as better quality of synthesized views.
Qiong Liu 0001, Yongbing Zhang 0002, Xiangyang Ji, Qionghai Dai
ICASSP4
2012 Super-resolution from unregistered aliased images with unknown scalings and shifts
abstract
We consider the problem of super-resolution from unregistered aliased images with unknown spatial scaling factors and shifts. Due to the limitation of pixel size in the image sensor, the sampling rate for each image is lower than the Nyquist rate of the scene. Thus, we have aliasing in captured images, which makes it hard to register the low-resolution images and then generate a high-resolution image. To work out this problem, we formulate it as a multichannel sampling and reconstruction problem with unknown parameters, spatial scaling factors and shifts. We can estimate the unknown parameters and then reconstruct the high-resolution image by solving a nonlinear least square problem using the variable projection method. Experiments with synthesized 1-D signals and 2-D images show the effectiveness of the proposed algorithm.
YiGang Peng, Qionghai Dai, Wenli Xu, Martin Vetterli
ICASSP3
2012 Robust joint reconstruction in compressed multi-view imaging
abstract
The newly emerging sampling methodology of compressed sensing opens a door to obtain compressed data directly. How to efficiently reconstruct the original signal from the compressed data becomes a new challenge. Many reconstruction works have been proposed on mono-view images by exploring the sparsity of the original image. However, it is a challenge to efficiently explore the correlations among different views in compressed multi-view imaging systems. With the aid of inter-view disparity information at receiver end, a joint reconstruction approach is presented for independently captured view-point images via compressed imaging. In the proposed approach, a robust reconstruction is obtained by formulating the occurrences of outliers, usually caused by illumination change, mismatch and discontinuity in disparity estimation, as a sparse model, which can be efficiently solved by a proximal sub-gradient algorithm bas ed on l1-norm minimization. Experimental results show that the joint reconstruction of compressed multi-view images can achieve significantly better recovery quality than the independently reconstructed ones.
Qionghai Dai, Changjun Fu, Xiangyang Ji, Yongbing Zhang 0002
PCS1
2012 Performance Capture of High-Speed Motion Using Staggered Multi-View Recording
abstract
Abstract We present a markerless performance capture system that can acquire the motion and the texture of human actors performing fast movements using only commodity hardware. To this end we introduce two novel concepts: First, a staggered surround multi‐view recording setup that enables us to perform model‐based motion capture on motion‐blurred images, and second, a model‐based deblurring algorithm which is able to handle disocclusion, self‐occlusion and complex object motions. We show that the model‐based approach is not only a powerful strategy for tracking but also for deblurring highly complex blur patterns.
Di Wu 0006, Yebin Liu, Ivo Ihrke, Qionghai Dai, Christian Theobalt
Comput. Graph. Forum4
2012 An overview of computational photography
Jin-Li Suo, Xiangyang Ji, Qionghai Dai
Sci. China Inf. Sci.3
2012 Relay-assisted hierarchical adaptation scheme for multi-user scalable video delivery to heterogeneous mobile devices
Hongjiang Xiao, Xiangyang Ji, Qionghai Dai
Sci. China Inf. Sci.3
2012 Commute time guided transformation for feature extraction
Yue Deng 0001, Qionghai Dai, Ruiping Wang 0001, Zengke Zhang
Comput. Vis. Image Underst.2
2012 A Concatenational Graph Evolution Aging Model
abstract
Modeling the long-term face aging process is of great importance for face recognition and animation, but there is a lack of sufficient long-term face aging sequences for model learning. To address this problem, we propose a CONcatenational GRaph Evolution (CONGRE) aging model, which adopts decomposition strategy in both spatial and temporal aspects to learn long-term aging patterns from partially dense aging databases. In spatial aspect, we build a graphical face representation, in which a human face is decomposed into mutually interrelated subregions under anatomical guidance. In temporal aspect, the long-term evolution of the above graphical representation is then modeled by connecting sequential short-term patterns following the Markov property of aging process under smoothness constraints between neighboring short-term patterns and consistency constraints among subregions. The proposed model also considers the diversity of face aging by proposing probabilistic concatenation strategy between short-term patterns and applying scholastic sampling in aging prediction. In experiments, the aging prediction results generated by the learned aging models are evaluated both subjectively and objectively to validate the proposed model.
Jin-Li Suo, Xilin Chen 0001, Shiguang Shan, Wen Gao 0001, Qionghai Dai
IEEE Trans. Pattern Anal. Mach. Intell.5
2012 A regional image fusion based on similarity characteristics
Xiaoyan Luo, Jun Zhang 0007, Qionghai Dai
Signal Process.3
2012 Free Viewpoint Video Coding With Rate-Distortion Analysis
abstract
To improve free viewpoint video (FVV) coding efficiency and optimize the quality of the synthesized virtual view video, this paper proposes a depth-assisted FVV coding framework and analyzes the rate-distortion (R-D) property of the synthesized virtual view video in FVV coding. In the depth-assisted FVV coding framework, the depth assigned disparity compensated prediction is introduced to exploit the correlation between multiview video (MVV) and depth. To model the R-D property of the synthesized virtual view video, a region-based view synthesis distortion estimation approach is investigated with respect to the distortion of MVV and depth. Subsequently, the general R-D property estimation models of MVV and depth are analyzed. Finally, a rate-allocation scheme is designed to optimize the quantization parameter pair of MVV and depth in FVV coding. The simulation results demonstrate that the proposed depth-assisted FVV coding framework can improve the FVV coding efficiency. The region-based view synthesis distortion estimation approach and the general R-D model are able to precisely approximate the R-D property of synthesized virtual view video in the multiview video plus depth based FVV coding frameworks. The proposed rate-allocation scheme can optimize the overall FVV coding efficiency to achieve a high-quality reconstructed video at the desired viewpoint with a given rate constraint.
Qifei Wang, Xiangyang Ji, Qionghai Dai, Naiyao Zhang
IEEE Trans. Circuits Syst. Video Technol.3
2012 Adaptive Compressed Sensing Recovery Utilizing the Property of Signal's Autocorrelations
abstract
Perfect compressed sensing (CS) recovery can be achieved when a certain basis space is found to sparsely represent the original signal. However, due to the diversity of the signals, there does not exist a universal predetermined basis space that can sparsely represent all kinds of signals, which results in an unsatisfying performance. To improve the accuracy of recovered signal, this paper proposes an adaptive basis CS reconstruction algorithm by minimizing the rank of an accumulated matrix (MRAM), whose eigenvectors approximate the optimal basis sparsely representing the original signal. The accumulated matrix is constructed to efficiently exploit the second-order statistical property of the signal's autocorrelations. Based on the theory of matrix completion, MRAM reconstructs the original signal from its random projections under the observation that the constructed accumulated matrix is of low rank for most natural signals such as periodic signals and those coming from an autoregressive stationary process. Experimental results show that the proposed MRAM efficiently improves the reconstruction quality compared with the existing algorithms.
Changjun Fu, Xiangyang Ji, Qionghai Dai
IEEE Trans. Image Process.3
2012 Camera Constraint-Free View-Based 3-D Object Retrieval
abstract
Recently, extensive research efforts have been dedicated to view-based methods for 3-D object retrieval due to the highly discriminative property of multiviews for 3-D object representation. However, most of state-of-the-art approaches highly depend on their own camera array settings for capturing views of 3-D objects. In order to move toward a general framework for 3-D object retrieval without the limitation of camera array restriction, a camera constraint-free view-based (CCFV) 3-D object retrieval algorithm is proposed in this paper. In this framework, each object is represented by a free set of views, which means that these views can be captured from any direction without camera constraint. For each query object, we first cluster all query views to generate the view clusters, which are then used to build the query models. For a more accurate 3-D object comparison, a positive matching model and a negative matching model are individually trained using positive and negative matched samples, respectively. The CCFV model is generated on the basis of the query Gaussian models by combining the positive matching model and the negative matching model. The CCFV removes the constraint of static camera array settings for view capturing and can be applied to any view-based 3-D object database. We conduct experiments on the National Taiwan University 3-D model database and the ETH 3-D object database. Experimental results show that the proposed scheme can achieve better performance than state-of-the-art methods.
Yue Gao 0002, Jinhui Tang 0001, Richang Hong, Shuicheng Yan, Qionghai Dai, Naiyao Zhang, Tat-Seng Chua
IEEE Trans. Image Process.5
2012 3-D Object Retrieval and Recognition With Hypergraph Analysis
abstract
View-based 3-D object retrieval and recognition has become popular in practice, e.g., in computer aided design. It is difficult to precisely estimate the distance between two objects represented by multiple views. Thus, current view-based 3-D object retrieval and recognition methods may not perform well. In this paper, we propose a hypergraph analysis approach to address this problem by avoiding the estimation of the distance between objects. In particular, we construct multiple hypergraphs for a set of 3-D objects based on their 2-D views. In these hypergraphs, each vertex is an object, and each edge is a cluster of views. Therefore, an edge connects multiple vertices. We define the weight of each edge based on the similarities between any two views within the cluster. Retrieval and recognition are performed based on the hypergraphs. Therefore, our method can explore the higher order relationship among objects and does not use the distance between objects. We conduct experiments on the National Taiwan University 3-D model dataset and the ETH 3-D object collection. Experimental results demonstrate the effectiveness of the proposed method by comparing with the state-of-the-art methods.
Yue Gao 0002, Meng Wang 0001, Dacheng Tao, Rongrong Ji, Qionghai Dai
IEEE Trans. Image Process.5
2012 Manifold-Manifold Distance and its Application to Face Recognition With Image Sets
abstract
In this paper, we address the problem of classifying image sets for face recognition, where each set contains images belonging to the same subject and typically covering large variations. By modeling each image set as a manifold, we formulate the problem as the computation of the distance between two manifolds, called manifold-manifold distance (MMD). Since an image set can come in three pattern levels, point, subspace, and manifold, we systematically study the distance among the three levels and formulate them in a general multilevel MMD framework. Specifically, we express a manifold by a collection of local linear models, each depicted by a subspace. MMD is then converted to integrate the distances between pairs of subspaces from one of the involved manifolds. We theoretically and experimentally study several configurations of the ingredients of MMD. The proposed method is applied to the task of face recognition with image sets, where identification is achieved by seeking the minimum MMD from the probe to the gallery of image sets. Our experiments demonstrate that, as a general set similarity measure, MMD consistently outperforms other competing nondiscriminative methods and is also promisingly comparable to the state-of-the-art discriminative methods.
Ruiping Wang 0001, Shiguang Shan, Xilin Chen 0001, Qionghai Dai, Wen Gao 0001
IEEE Trans. Image Process.4
2012 Three-Dimensional Motion Estimation via Matrix Completion
abstract
Three-dimensional motion estimation from multiview video sequences is of vital importance to achieve high-quality dynamic scene reconstruction. In this paper, we propose a new 3-D motion estimation method based on matrix completion. Taking a reconstructed 3-D mesh as the underlying scene representation, this method automatically estimates motions of 3-D objects. A "separating + merging" framework is introduced to multiview 3-D motion estimation. In the separating step, initial motions are first estimated for each view with a neighboring view. Then, in the merging step, the motions obtained by each view are merged together and optimized by low-rank matrix completion method. The most accurate motion estimation for each vertex in the recovered matrix is further selected by three spatiotemporal criteria. Experimental results on data sets with synthetic motions and real motions show that our method can reliably estimate 3-D motions.
Kun Li 0001, Qionghai Dai, Wenli Xu, Jing-Yu Yang 0002, Jianmin Jiang
IEEE Trans. Syst. Man Cybern. Part B2
2011 High resolution multispectral video capture with a hybrid camera system
abstract
We present a new approach to capture video at high spatial and spectral resolutions using a hybrid camera system. Composed of an RGB video camera, a grayscale video camera and several optical elements, the hybrid camera system simultaneously records two video streams: an RGB video with high spatial resolution, and a multispectral video with low spatial resolution. After registration of the two video streams, our system propagates the multispectral information into the RGB video to produce a video with both high spectral and spatial resolution. This propagation between videos is guided by color similarity of pixels in the spectral domain, proximity in the spatial domain, and the consistent color of each scene point in the temporal domain. The propagation algorithm is designed for rapid computation to allow real-time video generation at the original frame rate, and can thus facilitate real-time video analysis tasks such as tracking and surveillance. Hardware implementation details and design tradeoffs are discussed. We evaluate the proposed system using both simulations with ground truth data and on real-world scenes. The utility of this high resolution multispectral video data is demonstrated in dynamic white balance adjustment and tracking.
Xun Cao, Xin Tong 0001, Qionghai Dai, Stephen Lin 0001
CVPR3
2011 Exploring aligned complementary image pair for blind motion deblurring
abstract
Camera shake during long exposure is ineluctable in light-limited situations, and results in a blurry observation. Recovering the blur kernel and the latent image from the blurred image is an inherently ill-posed problem. In this paper, we analyze the image acquisition model to capture two blurred images simultaneously with different blur kernels. The image pair is well-aligned and the kernels have a certain relationship. Such strategy overcomes the challenge of blurry image alignment and reduces the ambiguity of blind deblurring. Thanks to the aided hardware, the algorithm based on such image pair can give high-quality kernel estimation and image restoration. The experiments on both synthetic and real images demonstrate the effectiveness of our image capture strategy, and show that the kernel estimation is accurate enough to restore superior latent image, which contains more details and fewer ringing artifacts.
Jun Zhang 0007, Qionghai Dai
CVPR3
2011 Compressed Multi-view Imaging with Joint Reconstruction
abstract
The newly emerging sampling methodology of compressed sensing opens a door to obtain compressed data directly. How to efficiently reconstruct the original signal from the compressed data becomes a new challenge problem. Many reconstruction works have been proposed on mono-view images by exploring the sparsity of the original image, how ever it is a challenge to efficiently explore the correlations between different views in compressed multiview imaging systems. With the aid of inter-view disparity information at receiver end, a joint reconstruction approach is presented for independently captured view point images via compressed imaging.
Changjun Fu, Xiangyang Ji, Qionghai Dai
DCC3
2011 Cooperative cross-layer transmission for scalable video to resource-constrained receiver
abstract
For resource-constrained wireless scenarios, this paper proposes a cooperative cross-layer scheme to improve the performance of scalable video transmission. At the application layer, the extraction parameters of scalable video coding are first optimized to maximize the visual quality and fulfill the resource constraints of decoding capacity and bandwidth. At the link layer, a priority-based cooperative scheduling strategy is further proposed to transmit the video packets of extracted substream. In this strategy, the relay link may substitute for the unreliable direct link if it has a relatively smaller packet error rate (PER). The packet priorities are determined by both the link PERs and the video layers properties. Simulations demonstrate the effectiveness of our proposed scheme.
Hongjiang Xiao, Qiong Liu 0001, Qionghai Dai
MUM3
2011 Vision field capture for advanced 3DTV applications
abstract
The seven-dimensional plenoptic function provides a full description of the visual information for the real world. In this paper, we present a novel concept called vision field, which simplifies the seven-dimensional plenoptic function into its three subspaces, namely, view, light, time. Based on this concept, we found that most previous 3D capture systems can be related to the vision field capture. This paper first gives a brief survey on the previous 3D capture systems, categorizes them from the vision field perspective. Then, we introduce a system which is able to capture the vision field. A Multi-View-Multi-Lighting (MVML) capture system is built to obtain the multiview images of the 3D scenes or objects under different steerable light conditions. Finally, we show how the vision field capture can be used for advanced 3DTV applications.
Xun Cao, Yebin Liu, Xiangyang Ji, Qionghai Dai
VCIP4
2011 Video-object segmentation and 3D-trajectory estimation for monocular video sequences
Feng Xu 0005, Kin-Man Lam 0001, Qionghai Dai
Image Vis. Comput.3
2011 A Prism-Mask System for Multispectral Video Acquisition
abstract
This paper presents a prism-mask system for capturing multispectral videos. The system is composed of a triangular prism, a monochrome camera, and an occlusion mask. Incoming light beams from the scene are sampled by the occlusion mask, dispersed into their constituent spectra by the triangular prism, and then captured by the monochrome camera. Our system is capable of capturing frames with high spectral resolution at video rates. It also allows for different trade-offs between spectral and spatial resolution by adjusting the focal length of the camera. We demonstrate multispectral video acquisition with various spectral resolutions and spatial resolutions, as well as different frame rates. The effectiveness of our system is further evaluated with several applications, including human skin detection, physical material recognition, video segmentation, RGB video generation, and illumination identification.
Xun Cao, Hao Du 0004, Xin Tong 0001, Qionghai Dai, Stephen Lin 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2011 3D model retrieval using weighted bipartite graph matching
Yue Gao 0002, Qionghai Dai, Meng Wang 0001, Naiyao Zhang
Signal Process. Image Commun.2
2011 Collaborative color calibration for multi-camera systems
Kun Li 0001, Qionghai Dai, Wenli Xu
Signal Process. Image Commun.2
2011 Video denoising using shape-adaptive sparse representation over similar spatio-temporal patches
Jun Zhang 0007, Qionghai Dai
Signal Process. Image Commun.3
2011 Markerless Shape and Motion Capture From Multiview Video Sequences
abstract
We propose a new markerless shape and motion capture approach from multiview video sequences. The shape recovery method consists of two steps: separating and merging. In the separating step, the depth map represented with a point cloud for each view is generated by solving a proposed variational model, which is regularized by four constraints to ensure the accuracy and completeness of the reconstruction. Then, in the merging step, the point clouds of all the views are merged together and reconstructed into a 3-D mesh using a marching cubes method with silhouette constraints. Experiments show that the geometric details are faithfully preserved in each estimated depth map. The 3-D meshes reconstructed from the estimated depth maps are watertight and present rich geometric details, even for non-convex objects. Taking the reconstructed 3-D mesh as the underlying scene representation, a volumetric deformation method with a new positional-constraint computation scheme is proposed to automatically capture motions of the 3-D object. Our method can capture non-rigid motions even for loosely dressed humans without the aid of markers.
Kun Li 0001, Qionghai Dai, Wenli Xu
IEEE Trans. Circuits Syst. Video Technol.2
2011 Graph Laplace for Occluded Face Completion and Recognition
abstract
This paper proposes a spectral-graph-based algorithm for face image repairing, which can improve the recognition performance on occluded faces. The face completion algorithm proposed in this paper includes three main procedures: 1) sparse representation for partially occluded face classification; 2) image-based data mining; and 3) graph Laplace (GL) for face image completion. The novel part of the proposed framework is GL, as named from graphical models and the Laplace equation, and can achieve a high-quality repairing of damaged or occluded faces. The relationship between the GL and the traditional Poisson equation is proven. We apply our face repairing algorithm to produce completed faces, and use face recognition to evaluate the performance of the algorithm. Experimental results verify the effectiveness of the GL method for occluded face completion.
Yue Deng 0001, Qionghai Dai, Zengke Zhang
IEEE Trans. Image Process.2
2011 Occlusion-Aware Motion Layer Extraction Under Large Interframe Motions
abstract
Extracting motion layers from videos is an important task for video representation, analysis, and compression. For videos with large interframe motions, motion layer extraction is challenging in two respects: the estimation of large disparity motions and the awareness of large occluded regions. In this paper, we propose an effective method for motion layer extraction under large disparity motions. To robustly estimate large displacement motions, we have developed an efficient voting-based method that estimates planar homographies from sparse feature matches. To handle occlusions, we first integrate color and motion consistency into a Markov random field framework to achieve per-pixel assignment with occlusion detection. Then, we perform motion-color segmentation and an earth mover's distance-based comparison to determine motion labels for occluded pixels. Experimental results show that our proposed method achieves good performance in automatically extracting multiple moving objects under large disparity motions while maintaining a low computational cost.
Feng Xu 0005, Qionghai Dai
IEEE Trans. Image Process.2
2011 Less is More: Efficient 3-D Object Retrieval With Query View Selection
abstract
The explosively increasing 3-D objects make their efficient retrieval technology highly desired. Extensive research efforts have been dedicated to view-based 3-D object retrieval for its advantage of using 2-D views to represent 3-D objects. In this paradigm, typically the retrieval is accomplished by matching the views of the query object with the objects in database. However, using all the query views may not only introduce difficulty in rapid retrieval but also degrade retrieval accuracy when there is a mismatch between the query views and the object views in the database. In this work, we propose an interactive 3-D object retrieval scheme. Given a set of query views, we first perform clustering to obtain several candidates. We then incrementally select query views for object matching: in each round of relevance feedback, we only add the query view that is judged to be the most informative one based on the labeling information. In addition, we also propose an efficient approach to learn a distance metric for the newly selected query view and the weights for combining all of the selected query views. We conduct experiments on the National Taiwan University 3D Model database, ETH 3D object collection, and Shape Retrieval Content of Non-Rigid 3D Model, and results demonstrated that our approach not only significantly speeds up the retrieval process but also achieves encouraging retrieval performance.
Yue Gao 0002, Meng Wang 0001, Zhengjun Zha, Qi Tian 0001, Qionghai Dai, Naiyao Zhang
IEEE Trans. Multim.5
2011 Causality Analysis of Neural Connectivity: Critical Examination of Existing Methods and Advances of New Methods
abstract
Granger causality (GC) is one of the most popular measures to reveal causality influence of time series and has been widely applied in economics and neuroscience. Especially, its counterpart in frequency domain, spectral GC, as well as other Granger-like causality measures have recently been applied to study causal interactions between brain areas in different frequency ranges during cognitive and perceptual tasks. In this paper, we show that: 1) GC in time domain cannot correctly determine how strongly one time series influences the other when there is directional causality between two time series, and 2) spectral GC and other Granger-like causality measures have inherent shortcomings and/or limitations because of the use of the transfer function (or its inverse matrix) and partial information of the linear regression model. On the other hand, we propose two novel causality measures (in time and frequency domains) for the linear regression model, called new causality and new spectral causality, respectively, which are more reasonable and understandable than GC or Granger-like measures. Especially, from one simple example, we point out that, in time domain, both new causality and GC adopt the concept of proportion, but they are defined on two different equations where one equation (for GC) is only part of the other (for new causality), thus the new causality is a natural extension of GC and has a sound conceptual/theoretical basis, and GC is not the desired causal influence at all. By several examples, we confirm that new causality measures have distinct advantages over GC or Granger-like measures. Finally, we conduct event-related potential causality analysis for a subject with intracranial depth electrodes undergoing evaluation for epilepsy surgery, and show that, in the frequency domain, all measures reveal significant directional event-related causality, but the result from new spectral causality is consistent with event-related time-frequency power spectrum activity. The spectral GC as well as other Granger-like measures are shown to generate misleading results. The proposed new causality measures may have wide potential applications in economics and neuroscience.
Sanqing Hu, Guojun Dai, Gregory A. Worrell, Qionghai Dai, Hualou Liang
IEEE Trans. Neural Networks4
2011 Video-based characters: creating new human performances from a multi-view video database
abstract
We present a method to synthesize plausible video sequences of humans according to user-defined body motions and viewpoints. We first capture a small database of multi-view video sequences of an actor performing various basic motions. This database needs to be captured only once and serves as the input to our synthesis algorithm. We then apply a marker-less model-based performance capture approach to the entire database to obtain pose and geometry of the actor in each database frame. To create novel video sequences of the actor from the database, a user animates a 3D human skeleton with novel motion and viewpoints. Our technique then synthesizes a realistic video sequence of the actor performing the specified motion based only on the initial database. The first key component of our approach is a new efficient retrieval strategy to find appropriate spatio-temporally coherent database frames from which to synthesize target video frames. The second key component is a warping-based texture synthesis approach that uses the retrieved most-similar database frames to synthesize spatio-temporally coherent target video frames. For instance, this enables us to easily create video sequences of actors performing dangerous stunts without them being placed in harm's way. We show through a variety of result videos and a user study that we can synthesize realistic videos of people, even if the target motions and camera views are different from the database content.
Feng Xu 0005, Yebin Liu, Carsten Stoll, James Tompkin 0001, Gaurav Bharaj, Qionghai Dai, Hans-Peter Seidel, Jan Kautz, Christian Theobalt
ACM Trans. Graph.6
2011 Fusing Multiview and Photometric Stereo for 3D Reconstruction under Uncalibrated Illumination
abstract
We propose a method to obtain a complete and accurate 3D model from multiview images captured under a variety of unknown illuminations. Based on recent results showing that for Lambertian objects, general illumination can be approximated well using low-order spherical harmonics, we develop a robust alternating approach to recover surface normals. Surface normals are initialized using a multi-illumination multiview stereo algorithm, then refined using a robust alternating optimization method based on the l(1) metric. Erroneous normal estimates are detected using a shape prior. Finally, the computed normals are used to improve the preliminary 3D model. The reconstruction system achieves watertight and robust 3D reconstruction while neither requiring manual interactions nor imposing any constraints on the illumination. Experimental results on both real world and synthetic data show that the technique can acquire accurate 3D models for Lambertian surfaces, and even tolerates small violations of the Lambertian assumption.
Chenglei Wu, Yebin Liu, Qionghai Dai, Bennett Wilburn
IEEE Trans. Vis. Comput. Graph.3
2010 Multi-View Stereo Reconstruction with High Dynamic Range Texture
Feng Lu 0005, Xiangyang Ji, Qionghai Dai, Guihua Er
ACCV (2)3
2010 Multiview video depth estimation with spatial-temporal consistency
Mingjin Yang, Xun Cao, Qionghai Dai
BMVC3
2010 Region Based Rate-Distortion Analysis for 3D Video Coding
abstract
Summary form only given. In 3D video (3DV), the virtual view images are commonly synthesized by the color and depth images of the reference views with image based rendering (IBR). Thus, in 3DV coding, to provide the high-quality interactive viewpoint video to audience, it is necessary to jointly optimize coding efficiency of color and depth images at a given bit-rate by rate-distortion (R-D) property analysis of 3DV coding.to calculate Edw, a region based distortion model is proposed firstly. In IBR, depth quantization error will cause the disparities between the pixels of the virtual view and the correspondent pixels of the reference views changed.
Qifei Wang, Xiangyang Ji, Qionghai Dai, Naiyao Zhang
DCC3
2010 High quality color calibration for multi-camera systems with an omnidirectional color checker
abstract
This paper proposes a new color calibration method for multi-camera systems with a novel omnidirectional color checker. The designed cylindrical color checker contains a periodic array of color patches, and is visible for all the cameras without manual adjustment. For color calibration, accurate global correspondences are first generated by local descriptors and area-based correlation methods. Then, the multi-camera color calibration problem is formulated as an overdetermined linear system, in which the dynamic range shaping is incorporated to ensure the high contrasts of captured images. The cameras are calibrated with the parameters obtained by solving the linear system. According to experimental results on a real multi-camera system, the proposed method shows high performance in achieving inter-camera color consistency and high dynamic range.
Kun Li 0001, Qionghai Dai, Wenli Xu
ICASSP2
2010 W2Go: a travel guidance system by automatic landmark ranking
abstract
In this paper, we present a travel guidance system W2Go (Where to Go), which can automatically recognize and rank the landmarks for travellers. In this system, a novel Automatic Landmark Ranking (ALR) method is proposed by utilizing the tag and geo-tag information of photos in Flickr and user knowledge from Yahoo Travel Guide. ALR selects the popular tourist attractions (landmarks) based on not only the subjective opinion of the travel editors as is currently done on sites like WikiTravel and Yahoo Travel Guide, but also the ranking derived from popularity among tourists. Our approach utilizes geo-tag information to locate the positions of the tag-indicated places, and computes the probability of a tag being a landmark/site name. For potential landmarks, impact factors are calculated from the frequency of tags, user numbers in Flickr, and user knowledge in Yahoo Travel Guide. These tags are then ranked based on the impact factors. Several representative views for popular landmarks are generated from the crawled images with geo-tags to describe and present them in context of information derived from several relevant reference sources. The experimental comparisons to the other systems are conducted on eight famous cities over the world. User-based evaluation demonstrates the effectiveness of the proposed ALR method and the W2Go system.
Yue Gao 0002, Jinhui Tang 0001, Richang Hong, Qionghai Dai, Tat-Seng Chua, Ramesh Jain 0001
ACM Multimedia4
2010 Intelligent query: open another door to 3d object retrieval
abstract
The increasing number of available 3D objects makes their efficient retrieval technology highly desired. Extensive research has been dedicated to view-based 3D object retrieval because of its advantage of 2D views for 3D object content representation. In this paradigm, typically the retrieval is accomplished based a set of different views of the query object, and focuses on the 3D object representation, matching and indexing. In this work, we present another aspect towards 3D object retrieval: intelligent query. Intelligent query includes query selection, query description and combination, and assistive query. We will show how this scheme is ideally suit for the 3D object retrieval problem. We conduct experiments on the National Taiwan University 3D Model database and results demonstrated that our approach can improve retrieval performance. Finally, we give insight into the future of the intelligent query for 3D object retrieval.
Yue Gao 0002, Meng Wang 0001, Jialie Shen 0001, Qionghai Dai, Naiyao Zhang
ACM Multimedia4
2010 Representative views re-ranking for 3D model retrieval with multi-bipartite graph reinforcement model
abstract
In this paper, we propose a multi-bipartite graph reinforcement model for representative views re-ranking in 3D model retrieval. Given the views of one query 3D model, all query views are grouped into clusters to generate representative views and corresponding original weights. In the retrieval procedure, labeled positive retrieval results are employed to refine the query information. Each group of views from positive retrieval results and the group of representative query views are employed to construct a bipartite graph, and a multi-bipartite graph reinforcement algorithm is performed on these bipartite graphs to re-rank all views. Then the weights of all representative query views are updated. Experimental results on two 3D model databases are provided to justify the effectiveness of the proposed method.
Yue Gao 0002, You Yang 0002, Qionghai Dai, Naiyao Zhang
ACM Multimedia3
2010 3D object retrieval with bag-of-region-words
abstract
View-based method becomes an essential approach to 3D object retrieval in recent years. In the view-based 3D object retrieval framework, each object is described by a set of views and representative features are extracted from these views to match the objects in database. In this paper, we propose a novel 3D multi-view representation method, Bag-of-Region-Words (BoRW). It first gridly selects points in each view and extracts local SIFT features. Each local feature is encoded into a visual word with a trained visual vocabulary. Then each view is split into several regions, and each region is represented by a bag-of-visual-words feature vector. All the obtained regions are further grouped into clusters based on the bag-of-visual-words feature, and one feature is selected from each cluster with corresponding weight. In this way, each object is described by a set of BoRW. The Earth Movers Distance is employed to estimate the distance between two BoRW feature vectors. Experimental results show that the proposed method can achieve better retrieval performance than existing methods.
Yue Gao 0002, You Yang 0002, Qionghai Dai, Naiyao Zhang
ACM Multimedia3
2010 Vision field capturing and its applications in 3DTV
abstract
3D video capturing acquires the visual information in 3D manner, which possesses the first step of the entire 3DTV system chain before 3D coding, transmission and visualization. The 3D capturing plays an important role because precise 3D visual capturing will benefit the whole 3DTV system. During the past decades, various kinds of capturing system have been built for different applications such as FTV[1], 3DTV, 3D movie, etc. As the cost of sensors reduces in recent years, a lot of systems utilize multiple cameras to acquire visual information, which is called multiview capturing. 3D information can be further extracted through multiview geometry. We will first give a brief review of these multiview systems and analyze their relationship from the perspective of plenoptic function [2]. Along with the multiple cameras, a lot of systems also make use of multiple lights to control the illumination condition. A new concept of vision field is presented in this talk according to the view-light-time subspace, which can be derived from the plenoptic function. The features and applications for each capturing system will be emphasized as well as the important issues in capturing like synchronization and calibration. Besides the multiple camera systems, some new techniques using TOF (time-off-light) camera [3] and 3D scanner will also be included in this talk.
Qionghai Dai, Xiangyang Ji, Xun Cao
PCS1
2010 Key technologies of light field capture for 3D reconstruction in microscopic scene
Xiangyang Ji, Qionghai Dai
Sci. China Inf. Sci.3
2010 View-based 3D model retrieval with probabilistic graph model
Yue Gao 0002, Jinhui Tang 0001, Qionghai Dai, Naiyao Zhang
Neurocomputing4
2010 3D model comparison using spatial structure circular descriptor
Yue Gao 0002, Qionghai Dai, Naiyao Zhang
Pattern Recognit.2
2010 Fast adaptive wavelet packets using interscale embedding of decomposition structures
Jing-Yu Yang 0002, Wenli Xu, Qionghai Dai
Pattern Recognit. Lett.3
2010 Statistical modeling and many-to-many matching for view-based 3D object retrieval
Qionghai Dai, Wenli Xu, Guihua Er
Signal Process. Image Commun.2
2010 A Novel JSCC Framework With Diversity-Multiplexing-Coding Gain Tradeoff for Scalable Video Transmission Over Cooperative MIMO
abstract
The multiple-input-multiple-output (MIMO) and cooperative communication are two state-of-the-art techniques to provide high-rate high-quality video communication services. By taking advantage of both techniques, this paper presents a novel joint source-channel coding (JSCC) framework for scalable video transmission over cooperative MIMO. In this framework, we first propose a cooperative MIMO architecture, which employs macro-micro power control strategy as a relaying protocol to determine the on/off mode of relays and the specific power allocation among them either by equal power amplification or by cooperative beamforming. Then, an unequal error protection structure is proposed to protect the video layers with different importance levels by concatenating the rate-variable low-density parity-check codes and diversity-embedded space-time block codes. Moreover, for the purpose of channel adaptation, the switch of space-time codes is designed to achieve different diversity and multiplexing tradeoff points. Finally, the JSCC algorithm integrated with diversity-multiplexing-coding gain tradeoff is proposed to optimize the resources of cooperative system to improve the transmission quality of scalable video. Experimental results demonstrate the effectiveness of our proposed schemes.
Hongjiang Xiao, Qionghai Dai, Xiangyang Ji, Wenwu Zhu 0001
IEEE Trans. Circuits Syst. Video Technol.2
2010 On the Recording Reference Contribution to EEG Correlation, Phase Synchorony, and Coherence
abstract
The degree of synchronization in electroencephalography (EEG) signals is commonly characterized by the time-series measures, namely, correlation, phase synchrony, and magnitude squared coherence (MSC). However, it is now well established that the interpretation of the results from these measures are confounded by the recording reference signal and that this problem is not mitigated by the use of other EEG montages, such as bipolar and average reference. In this paper, we analyze the impact of reference signal amplitude and power on EEG signal correlation, phase synchrony, and MSC. We show that, first, when two nonreferential signals have negative correlation, the phase synchrony and the absolute value of the correlation of the two referential signals may have two regions of behavior characterized by a monotonic decrease to zero and then a monotonic increase to one as the amplitude of the reference signal varies in [0, +∞). It is notable that even a small change of the amplitude may lead to significant impact on these two measures. Second, when two nonreferential signals have positive correlation, the correlation and phase-synchrony values of the two referential signals can monotonically increase to one (or monotonically decrease to some positive value and then monotonically increase to one) as the amplitude of the reference signal varies in [0, +∞). Third, when two nonreferential signals have negative cross-power, the MSC of the two referential signals can monotonically decrease to zero and then monotonically increase to one as reference signal power varies in [0, +∞). Fourth, when two nonreferential signals have positive cross-power, the MSC of the two referential signals can monotonically increase to one as the reference signal power varies in [0, +∞). In general, the reference signal with small amplitude or power relative to the signals of interest may decrease or increase the values of correlation, phase synchrony, and MSC. However, the reference signal with high relative amplitude or power will always increase each of the three measures. In our previous paper, we developed a method to identify and extract the reference signal contribution to intracranial EEG (iEEG) recordings. In this paper, we apply this approach to referential iEEG recorded from human subjects and directly investigate the contribution of recording reference on correlation, phase synchrony, and MSC. The experimental results demonstrate the significant impact that the recording reference may have on these bivariate measures.
Sanqing Hu, Matt Stead, Qionghai Dai, Gregory A. Worrell
IEEE Trans. Syst. Man Cybern. Part B3
2010 A Point-Cloud-Based Multiview Stereo Algorithm for Free-Viewpoint Video
abstract
This paper presents a robust multiview stereo (MVS) algorithm for free-viewpoint video. Our MVS scheme is totally point-cloud-based and consists of three stages: point cloud extraction, merging, and meshing. To guarantee reconstruction accuracy, point clouds are first extracted according to a stereo matching metric which is robust to noise, occlusion, and lack of texture. Visual hull information, frontier points, and implicit points are then detected and fused with point fidelity information in the merging and meshing steps. All aspects of our method are designed to counteract potential challenges in MVS data sets for accurate and complete model reconstruction. Experimental results demonstrate that our technique produces the most competitive performance among current algorithms under sparse viewpoint setups according to both static and motion MVS data sets.
Yebin Liu, Qionghai Dai, Wenli Xu
IEEE Trans. Vis. Comput. Graph.2
2009 Sketch realizing: lifelike portrait synthesis from sketch
abstract
People usually visualize their imaginations or memories through sketching. However, it might be difficult to record color and texture details powered by a large database of photographs gathered from the Web. This paper deals with the imagination visualization problem for human faces. A framework for synthesizing lifelike portraits from user specifications and input sketches is proposed which is an inverse process of sketch generation. The proposed framework synthesizes the realistic appearance of a face by taking parts from an annotated face library of photographs and stitching them together followed by further deformation.
Di Wu 0006, Qionghai Dai
CGI2
2009 Continuous depth estimation for multi-view stereo
abstract
Depth-map merging approaches have become more and more popular in multi-view stereo (MVS) because of their flexibility and superior performance. The quality of depth map used for merging is vital for accurate 3D reconstruction. While traditional depth map estimation has been performed in a discrete manner, we suggest the use of a continuous counterpart. In this paper, we first integrate silhouette information and epipolar constraint into the variational method for continuous depth map estimation. Then, several depth candidates are generated based on a multiple starting scales (MSS) framework. From these candidates, refined depth maps for each view are synthesized according to path-based NCC (normalized cross correlation) metric. Finally, the multiview depth maps are merged to produce 3D models. Our algorithm excels at detail capture and produces one of the most accurate results among the current algorithms for sparse MVS datasets according to the Middlebury benchmark. Additionally, our approach shows its outstanding robustness and accuracy in free-viewpoint video scenario.
Yebin Liu, Xun Cao, Qionghai Dai, Wenli Xu
CVPR3
2009 Slepian-Wolf Coding of Binary Finite Memory Source Using Burrows-Wheeler Transform
abstract
Summary form only given: In all existing codec designs for asymmetric Slepian-Wolf coding (SWC), it is assumed that the source sequence is i.i.d and equiprobable. When it comes to more complex source statistics, the encoder should firstly remove the redundancy within the source. However, this increases the complexity of the encoder. In this paper, we propose an asymmetric SWC scheme which explores the redundancy of the binary finite memory source (FMS) at the decoder. Specifically, inspired by the Burrows-Wheeler transform (BWT)-based source-controlled channel decoding algorithm proposed, we iteratively apply the LDPC decoding and BWT to the side information. In our codec implementation, the encoder is identical to the conventional LDPC-based SWC encoder. At the decoder, conventional LDPC-based SWC decoding algorithm and Burrows-Wheeler transform (BWT) are iteratively applied to the side information for decoding. BWT can asymptotically permute a FMS into a piece-wise i.i.d binary sequence. In other words, by applying BWT to the decoder side information, the redundancy in the memory is transformed into the redundancy in the marginal distributions of the output i.i.d segments. The marginal distributions of every i.i.d segment can be used as the a priori information for SWC decoding. To explore the marginal distribution, a segmentation algorithm is employed to adaptively partition the output sequence of BWT into i.i.d segments. The bias parameter of each segment is then empirically computed. Using these parameters, the a priori information of the FMS source can be derived and incorporated in the next iteration of SWC decoding. Experimental results show that our scheme performs significantly better than the scheme which does not utilize the a priori information for decoding.
Xiangyang Ji, Qionghai Dai, Xiaodong Liu 0005
DCC3
2009 Partially occluded face completion and recognition
abstract
This paper proposes a spectral graph based algorithm for face image repairing, which can improve the recognition performance on occluded faces. Our algorithm is called `guided label-learning', so named from graphical models, and can achieve a high-quality repairing of damaged or occluded faces. We apply our face repairing algorithm in order to produce completed faces, and then use face recognition to evaluate the performance of our algorithm. Experiment results show that, at most, a nearly 30%-increase in the recognition rate can be achieved for occluded faces with the use of our algorithm.
Yue Deng 0001, Dong Li 0028, Xudong Xie, Kin-Man Lam 0001, Qionghai Dai
ICIP5
2009 Image fusion in compressed sensing
abstract
This paper proposes an efficient image fusion scheme for compressed sensing (CS) imaging, in which fusion is performed on the random projections before reconstruction. Specifically, the measurements of multiple input images are fused into composite measurements via weighted average, in which the weights are calculated based on entropy metrics of the original measurements. Then the fused image with transformation coefficients in a selected basis is reconstructed from the composite measurements by the gradient projection for sparse reconstruction (GPSR) algorithm. The proposed scheme is implemented in a block-based CS framework. Simulation results show that our scheme provides promising fusion performance with a low computational complexity.
Xiaoyan Luo, Jun Zhang 0007, Jing-Yu Yang 0002, Qionghai Dai
ICIP4
2009 Joint resources allocation for cooperative video transmission
abstract
This paper proposes a cooperative multiple-input multiple-output (MIMO) architecture to transmit the video reliably. By exploring the configurable resources of the scalable video and the coded MIMO system, we further propose a novel joint source-channel coding (JSCC) framework with optimal tradeoff among diversity, multiplexing and coding gains. In this framework, the concatenated low-density parity-check codes (LDPC) and diversity-embedded space-time block codes (DE-STBC) provide double unequal error protection for the video layers, and the STBCs switching enables the adaptability to the varying channel. Experiments demonstrate the superiority of cooperative architecture and the effectiveness of our JSCC algorithm.
Hongjiang Xiao, Qionghai Dai, Xiangyang Ji
ICIP2
2009 Point-cloud refinement via exact matching
abstract
In many multi-view stereo (MVS) algorithms, a point-cloud evolution is performed, based on the matching process. For most of them, an assumption is usually employed for the matching, which indicates that the matching windows have the same shape. This assumption lays a great limit to the quality of the reconstructed result. To improve the point-cloud obtained from other algorithms, and break the limit laid by the regular matching, we propose our refinement method using exact matching. The exact matching enables more accurate matching windows for the points, by taking the normal vector into consideration. By maximizing the exact matching result, the point's coordinate and normal vectors are optimized, and we can thus make the original point-cloud much better.
Xiaoduan Feng, Yebin Liu, Qionghai Dai
ICME3
2009 Gabor Boost Linear Discriminant Analysis for face recognition
abstract
This paper proposes an innovative algorithm named Gabor Boost Linear Discriminant Analysis (GBLDA) for face recognition. In our method, we want to estimate the distribution of high dimensional Gabor wavelet (GW) features in a low dimensional LDA subspace without computing the GW feature of an input image. The computational complexity can be reduced significantly. Hence, GBLDA is suitable for real-time applications. Experimental results show that our proposed method not only possesses the advantages of linear subspace analysis approaches such as low computational complexity, but also has the advantage of a high recognition performance in the Gabor based methods.
Xudong Xie, Qionghai Dai, Zhigang Jin
ICME3
2009 Multi-view reconstruction under varying illumination conditions
abstract
This paper addresses the problem of complete and detailed 3D model reconstruction of objects filmed by multiple cameras under varying illumination. Firstly, initial normal maps are obtained to enhance the correspondence mapping. Then, the depth for every pixel is estimated by combining photometric constraint with occlusion robust photo-consistency. Finally, after filtering the point cloud, a Poisson surface reconstruction is applied to obtain a watertight mesh. In contrast with traditional photometric stereo techniques, the proposed algorithm does not directly calculate the photometric normal but integrates the photometric constraint into the depth estimation. Furthermore, different from classic multi-view stereo(MVS), we consider the counterpart under changing light conditions. The algorithm has been implemented based on our multi-camera and multi-light acquisition system. We validate the method by complete reconstruction of challenging real objects and show experimentally that this technique can greatly improve on correspondence-based MVS results.
Chenglei Wu, Yebin Liu, Xiangyang Ji, Qionghai Dai
ICME4
2009 Ways to sparse representation: An overview
Jing-Yu Yang 0002, YiGang Peng, Wenli Xu, Qionghai Dai
Sci. China Ser. F Inf. Sci.4
2009 Learning nonlinear manifolds based on mixtures of localized linear manifolds under a self-organizing framework
Huicheng Zheng, Qionghai Dai, Sanqing Hu, Zheming Lu 0001
Neurocomputing3
2009 Weighted Subspace Distance and Its Applications to Object Recognition and Retrieval With Image Sets
abstract
We address the problem of measuring the distance between two subspaces, each of which is spanned by an image set. In the existing methods, only the orthonormal basis is used to represent the subspace. However, the images are usually distributed in a limited area, rather than the whole subspace. Therefore, the characteristics of the distribution should also be considered. In this letter, a weighted subspace distance (WSD) is proposed, in which the principal component values of the data set are adopted to calculate the weights. Experimental results on object recognition and retrieval with image sets demonstrate the effectiveness of our proposal.
Qionghai Dai, Wenli Xu, Guihua Er
IEEE Signal Process. Lett.2
2009 Early Determination of Zero-Quantized 8 , ×, 8 DCT Coefficients
abstract
This paper proposes a novel approach to early determination of zero-quantized 8 × 8 discrete cosine transform (DCT) coefficients for fast video encoding. First, with the dynamic range analysis of DCT coefficients at different frequency positions, several sufficient conditions are derived to early determine whether a prediction error block (8 × 8) is an all-zero or a partial-zero block, i.e., the DCT coefficients within the block are all or partially zero-quantized. Being different from traditional methods that utilize the sum of absolute difference (SAD) of the entire prediction error block, the sufficient conditions are derived based on the SAD of each row of the prediction error block. For partial-zero blocks, fast DCT/IDCT algorithms are further developed by pruning conventional 8-point butterfly-based DCT/IDCT algorithms. Experimental results exhibit that the proposed early determination algorithm greatly reduces computational complexity in terms of DCT/IDCT, quantization, and inverse quantization, as compared with existing algorithms.
Xiangyang Ji, Sam Kwong, Debin Zhao, Hanli Wang, C.-C. Jay Kuo, Qionghai Dai
IEEE Trans. Circuits Syst. Video Technol.6
2009 Image and Video Denoising Using Adaptive Dual-Tree Discrete Wavelet Packets
abstract
We investigate image and video denoising using adaptive dual-tree discrete wavelet packets (ADDWP), which is extended from the dual-tree discrete wavelet transform (DDWT). With ADDWP, DDWT subbands are further decomposed into wavelet packets with anisotropic decomposition, so that the resulting wavelets have elongated support regions and more orientations than DDWT wavelets. To determine the decomposition structure, we develop a greedy basis selection algorithm for ADDWP, which has significantly lower computational complexity than a previously developed optimal basis selection algorithm, with only slight performance loss. For denoising the ADDWP coefficients, a statistical model is used to exploit the dependency between the real and imaginary parts of the coefficients. The proposed denoising scheme gives better performance than several state-of-the-art DDWT-based schemes for images with rich directional features. Moreover, our scheme shows promising results without using motion estimation in video denoising. The visual quality of images and videos denoised by the proposed scheme is also superior.
Jing-Yu Yang 0002, Yao Wang 0001, Wenli Xu, Qionghai Dai
IEEE Trans. Circuits Syst. Video Technol.4
2008 Shot-based similarity measure for content-based video summarization
abstract
The rapid development of multimedia applications over the past decade requires efficient methods for video browsing. In this paper, we present an algorithm for video summarization with shot comparison. We analyze video content in the shot level, and we calculate the shot distance using the advanced Hausdorff distance. The advanced Hausdorff distance combines the Hausdorff distance and Boolean model, and it could compare two shots from the global view. When the shot similarity matrix is obtained, we group these video shots into several clusters using the affinity propagation cluster method to remove redundant video content. Performance evaluation on ten video sequences are given to illustrate the proposed algorithm.
Yue Gao 0002, Qionghai Dai
ICIP2
2008 Inequivalent manifold ranking for content-based image retrieval
abstract
We propose to improve the effectiveness and scalability of graph based manifold ranking methods in image retrieval applications by emphasizing reliable images while damping the effect of noisy or irregular ones. Label information is firstly passed between most reliable data points, then propagated to less reliable ones on manifold structure. By treating these images inequivalently, undesirable effect of noisy samples is greatly reduced, thus effectiveness of manifold ranking algorithms is enhanced. Also, graph size in terms of number of nodes and edges is dramatically reduced, resulting in a great speed-up of the algorithm. Our experiment on real world image data set demonstrates the effectiveness of the proposed approach.
Guihua Er, Qionghai Dai
ICIP3
2008 2-D anisotropic dual-tree complex wavelet packets and its application to image denoising
abstract
In this paper, we extend the 2-D dual-tree complex wavelet transform (DTCWT) to an adaptive anisotropic dual-tree complex wavelet packets (ADTCWP). The DTCWT subbands are iteratively decomposed into anisotropic complex wavelet packets, generating anisotropic wavelets and increasing the number of wavelets orientations without introducing extra redundancy. Then a basis selection procedure is applied so that the selected complex wavelet packets are well adapted to image characteristics. The effectiveness of ADTCWP is examined in image denoising with a bivariate statistical model. The ADTCWP-based denoising scheme shows better denoising results than several DTCWT-based methods for image with rich directional features. The denoised images recovered by the proposed scheme are visually more appealing.
Jing-Yu Yang 0002, Wenli Xu, Yao Wang 0001, Qionghai Dai
ICIP4
2008 Face recognition using anisotropic dual-tree complex wavelet packets
abstract
In this paper, we propose a novel face recognition method based on anisotropic dual-tree complex wavelet packets(ADT-CWP). 2-D dual-tree complex wavelet transform(DT-CWT) provides a geometrically oriented decomposition for image representation as well as shift invariance. By applying anisotropic wavelet packet decomposition on DT-CWT further, ADT-CWP can be used to extract facial features better, which turns out to benefit for face recognition. With adaptively assigning different weights to different wavelet subbands, consistent best performances can be obtained based on different face databases which are under different conditions, such as varying illuminations and expressions, compared to PCA and other face recognition methods, especially Gabor-based method. Furthermore, in addition to the consistent and promising classification performances, our proposed ADT-CWP-based method has a really low computational complexity.
YiGang Peng, Xudong Xie, Wenli Xu, Qionghai Dai
ICPR4
2008 Locally Linear Online Mapping for Mining Low-Dimensional Data Manifolds
Huicheng Zheng, Qionghai Dai, Sanqing Hu
PAKDD3
2008 Image-based Material Weathering
abstract
Abstract The appearance manifold [WTL*06] is an efficient approach for modeling and editing time‐variant appearance of materials from the BRDF data captured at single time instance. However, this method is difficult to apply in images in which weathering and shading variations are combined. In this paper, we present a technique for modeling and editing the weathering effects of an object in a single image with appearance manifolds. In our approach, we formulate the input image as the product of reflectance and illuminance. An iterative method is then developed to construct the appearance manifold in color space (i.e., Lab space) for modeling the reflectance variations caused by weathering. Based on the appearance manifold, we propose a statistical method to robustly decompose reflectance and illuminance for each pixel. For editing, we introduce a “pixel‐walking” scheme to modify the pixel reflectance according to its position on the manifold, by which the detailed reflectance variations are well preserved. We illustrate our technique in various applications, including weathering transfer between two images that is first enabled by our technique. Results show that our technique can produce much better results than existing methods, especially for objects with complex geometry and shading effects.
Su Xue, Jiaping Wang, Xin Tong 0001, Qionghai Dai, Baining Guo
Comput. Graph. Forum4
2008 Image Coding Using Dual-Tree Discrete Wavelet Transform
abstract
In this paper, we explore the application of 2-D dual-tree discrete wavelet transform (DDWT), which is a directional and redundant transform, for image coding. Three methods for sparsifying DDWT coefficients, i.e., matching pursuit, basis pursuit, and noise shaping, are compared. We found that noise shaping achieves the best nonlinear approximation efficiency with the lowest computational complexity. The interscale, intersubband, and intrasubband dependency among the DDWT coefficients are analyzed. Three subband coding methods, i.e., SPIHT, EBCOT, and TCE, are evaluated for coding DDWT coefficients. Experimental results show that TCE has the best performance. In spite of the redundancy of the transform, our DDWT _ TCE scheme outperforms JPEG2000 up to 0.70 dB at low bit rates and is comparable to JPEG2000 at high bit rates. The DDWT _TCE scheme also outperforms two other image coders that are based on directional filter banks. To further improve coding efficiency, we extend the DDWT to an anisotropic dual-tree discrete wavelet packets (ADDWP), which incorporates adaptive and anisotropic decomposition into DDWT. The ADDWP subbands are coded with TCE coder. Experimental results show that ADDWP _ TCE provides up to 1.47 dB improvement over the DDWT _TCE scheme, outperforming JPEG2000 up to 2.00 dB. Reconstructed images of our coding schemes are visually more appealing compared with DWT-based coding schemes thanks to the directionality of wavelets.
Jing-Yu Yang 0002, Yao Wang 0001, Wenli Xu, Qionghai Dai
IEEE Trans. Image Process.4
2008 Multilabel Neighborhood Propagation for Region-Based Image Retrieval
abstract
Content-based image retrieval (CBIR) has been an active research topic in the last decade. As one of the promising approaches, graph-based semi-supervised learning has attracted many researchers. However, while the related work mainly focused on global visual features, little attention has been paid to region-based image retrieval (RBIR). In this paper, a framework based on multilabel neighborhood propagation is proposed for RBIR, which can be characterized by three key properties: (1) For graph construction, in order to determine the edge weights robustly and automatically, mixture distribution is introduced into the Earth mover's distance (EMD) and a linear programming framework is involved. (2) Multiple low-level labels for each image can be obtained based on a generative model, and the correlations among different labels are explored when the labels are propagated simultaneously on the weighted graph. (3) By introducing multilayer semantic representation (MSR) and support vector machine (SVM) into the long-term learning, more exact weighted graph for label propagation and more meaningful high-level labels to describe the images can be calculated. Experimental results, including comparisons with the state-of-the-art retrieval systems, demonstrate the effectiveness of our proposal.
Qionghai Dai, Wenli Xu, Guihua Er
IEEE Trans. Multim.2
2007 Motion Information Exploitation in H.264 Frame Skipping Transcoding
Xiaodong Liu 0005, Qionghai Dai
ACIVS3
2007 All-Clear Image Based Synthesis using Clarity Degree
abstract
Image based rendering (IBR) usually produces severe artifacts, typically blur and ghost, for objects not located on the focal plane. To obtain an all clear result, previous works either endeavor to recover scene depth or rest on iterations. These methods are difficult and time-consuming, not suitable for real-time IBR applications. In this paper, we propose a novel all-clear synthesis scheme bypassing tedious depth estimation. Our algorithm directly measures the clarity degree of local regions in pre-rendered images and then optimal blocks are combined. This measurement is motivated by the observation that mean change energies (MCE) are well consistent with the human vision system's feeling of "clarity". Furthermore, an efficient pseudo-depth filtering is proposed to alleviate the block effect during combination. Our algorithm gives outstanding performances on both synthetic and real world data. It is also proved to be fast enough for real-time implementation.
Xun Cao, Su Xue, Qionghai Dai
ICASSP (1)3
2007 A New Scalable Free Viewpoint Video Streaming System Over IP Network
abstract
Free viewpoint video (FVV), as a new type of multimedia, is a very challenging research area. In order to achieve free viewpoint navigation in realistic scenes, many image based rendering (IBR) technologies are introduced into FVV systems. Tow most important IBR technologies are light field rendering (LFR) and depth image based rendering (DIBR). This paper will show a new LFR based FVV streaming system over broadband IP networks. Our system uses jump frame method to support scalable quality of service (QoS) to users. An I frame retransmission method in application layer collaborates with RTP/RTCP technology in video streaming ensures different level of QoS.
Zhun Han, Qionghai Dai
ICASSP (2)2
2007 Correlated Probabilistic Label Propagation for Region-Based Image Retrieval
abstract
Label propagation and manifold ranking have been successfully adopted in content-based image retrieval (CBIR) in recent years. However, while the global low-level features are widely utilized in current systems, region-based features have received little attention. In this paper, a novel transductive framework based on correlated probabilistic label propagation is proposed for region-based images retrieval (RBIR), which can be characterized by three key properties: (1) unified feature matching (UFM) is chosen to measure the similarity between two segmented images. (2) To represent the segmented images in a uniform feature space, a generative model is adopted and the probabilistic labels of each image can be obtained. (3) In the retrieval process, multiple probabilistic labels of training samples are propagated simultaneously on the weighted graph, and the correlation among different labels are explored. Experimental results on 10000 images show that our algorithm can greatly improve the retrieval performance of the RBIR system.
Qionghai Dai, Wenli Xu, Guihua Er
ICASSP (1)2
2007 Multi-View Images Coding Based on Multiterminal Source Coding
abstract
In this paper, we proposed a multi-view images coding method based on multiterminal source coding (MSC). Due to separate encoding in MSC, our coding scheme can achieve good random access performance and the spatial redundancy can be exploited even if the encoders can not communicate with each other. Because of joint decoding, the compression performance of our scheme is promising, far better than that of separate encoding and decoding scheme. Compared to multi-view coding based on Wyner Ziv coding, our coding scheme is more flexible and can be easily extended to N views coding. There is no need to classify the view images into key images and Wyner Ziv images, images can be compressed in the same way and we can easily change the compression rate of each view to adapt to the resource conditions, like network bandwidth or storage. Experiment results show that the compression performance of our scheme is better than that of JPEG encoder and decoder scheme.
Qionghai Dai, Guiguang Ding
ICASSP (1)2
2007 Nonlinear Poisson Image Completion using Color Manifold
abstract
While most image completion methods focus on filling regions with structures or stationary textures, few are suitable for completing large-scale missing parts on complex background with nonlinearly progressive color changes. In this paper, we propose a novel approach, termed as nonlinear Poisson completion, to solve this problem. The visible parts of the background serve as a training set, from which we learn the embedding nonlinear subspace of progressive colors, namely color manifold. A Poisson image completion procedure, which works efficiently for smoothly linear interpolation, is extended to nonlinearly recover the missing regions with iteration solution confined to the manifold. In some especially challenging cases, a simple post-processing serves to generate more natural-looking results. Experiments on both synthetic and real images verify the effectiveness of the proposed algorithm.
Su Xue, Qionghai Dai
ICIP (4)2
2007 Image Coding using 2-D Anisotropic Dual-Tree Discrete Wavelet Transform
abstract
We propose an image coding scheme using 2-D anisotropic dual-tree discrete wavelet transform (DDWT). First, we extend 2-D DDWT to anisotropic decomposition, and obtain more directional subbands. Second, an iterative projection-based noise shaping algorithm is employed to further sparsify anisotropic DDWT coefficients. At last, the resulting coefficients are rearranged to preserve zero-tree relationship so that they can be efficiently coded with SPIHT. Experimental results show that our proposed scheme outperforms JPEG2000 and SPIHT at low bit rates despite the redundancy of DDWT.
Jing-Yu Yang 0002, Jizheng Xu, Feng Wu 0001, Qionghai Dai, Yao Wang 0001
ICIP (3)4
2007 Performance Modeling and Evaluation of Prediction Structures in Multi-View Video Coding
abstract
So far, lots of prediction structures have been proposed to exploit the temporal and inter-view redundancy of multi-view video. However, none of them can adapt themselves to different video contents. In this sense, it is desirable to adaptively select different prediction structure for different multi-view video contents. In this paper, a solution is proposed to evaluate the performance of prediction structures. With this solution, the performance of any kind of prediction structure can be evaluated before actually applying them. The solution can also be used for the tradeoff optimization between compression efficiency and other functionalities such as random-access ability and view scalability.
Yebin Liu, Qionghai Dai, Xiaodong Liu 0005
ICME3
2007 Image Compression using 2D Dual-tree Discrete Wavelet Transform (DDWT)
abstract
In this paper, we investigate image compression using 2D dual-tree discrete wavelet transform (DDWT), which is an overcomplete transform with direction-selective basis functions. To further sparsify DDWT coefficients, an iterative projection-based noise shaping method is employed. We analyze the statistics of DDWT coefficients as well as the inter-scale, inter-subband, and intra-subband dependency among the DDWT coefficients. We further evaluate the application of SPIHT and EBCOT for coding DDWT coefficients. Experimental results show that SPIHT is more effective than EBCOT for DDWT, and the DDWT-SPIHT coder outperforms JPEG2000 at low bit rates and is comparable to JPEG2000 at high bit rates.
Jing-Yu Yang 0002, Wenli Xu, Qionghai Dai, Yao Wang 0001
ISCAS3
2007 A New Multi-view Learning Algorithm Based on ICA Feature for Image Retrieval
Qionghai Dai
MMM (1)2
2007 A sender-driven time-stamp controlling based dynamic light field streaming service
abstract
Light Field Rendering (LFR) now plays a very important role in Free View-point Video (FVV) service, which is a new type of multi-media. Supporting Light Field Video (LFV) streaming over IP network is a very challenging research area. This paper shows a sender-driven streaming service that can support dynamic light field video service to multiple users over broad-band IP networks using a time-stamp controlling algorithm. Results show that system built based on our algorithm can support more than 50 users in a 100Mb band-width on server side.
Zhun Han, Qionghai Dai, Yebin Liu
VCIP2
2007 Region-based hidden Markov models for image categorization and retrieval
abstract
Hidden Markov models (HMMs) have been widely used in various fields, including image categorization and retrieval. Most of the existing methods train HMMs by low-level features of image blocks; however, the blockbased features can not reflect high-level semantic concepts well. This paper proposes a new method to train HMMs by region-based features, which can be obtained after image segmentation. Our work can be characterized by two key properties: (1) Region-based HMM is adopted to achieve better categorization performance, for the region-based features accord with the human perception better. (2) Multi-layer semantic representation (MSR) is introduced to couple with region-based HMM in a long-term relevance feedback framework for image retrieval. The experimental results demonstrate the effectiveness of our proposal in both aspects of categorization and retrieval.
Qionghai Dai, Wenli Xu
VCIP2
2007 A comparative study of image compression based on directional wavelets
abstract
Discrete wavelet transform is an effective tool to generate scalable stream, but it cannot efficiently represent edges which are not aligned in horizontal or vertical directions, while natural images often contain rich edges and textures of this kind. Hence, recently, intensive research has been focused particularly on the directional wavelets which can effectively represent directional attributes of images. Specifically, there are two categories of directional wavelets: redundant wavelets (RW) and adaptive directional wavelets (ADW). One representative redundant wavelet is the dual-tree discrete wavelet transform (DDWT), while adaptive directional wavelets can be further categorized into two types: with or without side information. In this paper, we briefly introduce directional wavelets and compare their directional bases and image compression performances.
Kun Li 0001, Wenli Xu, Qionghai Dai, Yao Wang 0001
VCIP3
2007 Rate-prediction structure complexity analysis for multi-view video coding using hybrid genetic algorithms
abstract
Efficient exploitation of the temporal and inter-view correlation is critical to multi-view video coding (MVC), and the key to it relies on the design of prediction chain structure according to the various pattern of correlations. In this paper, we propose a novel prediction structure model to design optimal MVC coding schemes along with tradeoff analysis in depth between compression efficiency and prediction structure complexity for certain standard functionalities. Focusing on the representation of the entire set of possible chain structures rather than certain typical ones, the proposed model can given efficient MVC schemes that adaptively vary with the requirements of structure complexity and video source characteristics (the number of views, the degrees of temporal and interview correlations). To handle large scale problem in model optimization, we deploy a hybrid genetic algorithm which yields satisfactory results shown in the simulations.
Yebin Liu, Qionghai Dai, Zhixiang You, Wenli Xu
VCIP2
2007 A novel approach to fuzzy rough sets based on a fuzzy covering
Tingquan Deng, Yanmei Chen, Wenli Xu, Qionghai Dai
Inf. Sci.4
2007 Histogram mining based on Markov chain and its application to image categorization
Qionghai Dai, Wenli Xu, Guihua Er
Signal Process. Image Commun.2
2006 Dynamic Light Field Compression Using Shared Fields and Region Blocks for Streaming Service
Yebin Liu, Qionghai Dai, Wenli Xu, Zhihong Liao
ACIVS2
2006 Composite-Block Model and Joint-Prediction Algorithm for Inter-Frame Video Coding
abstract
Block-matching algorithm (BMA) is widely adopted in video coding. However, many blocks contain both background pixels and moving-object pixels, which often differ significantly in true physical motions. Using BMA for these blocks may lead to large prediction error and thus decreased coding efficiency. More accurate prediction can be achieved by object-based coding, which segments a video frame into different objects and codes these objects separately. However, additional bits are required for encoding the shape of the objects, and the computation complexity of object segmentation is often very high. In this paper, a composite-block model and a joint-prediction algorithm are proposed. This algorithm can get accurate prediction without shape coding, and its computation complexity is very close to that of the BMA. The experimental result shows that the new H.264 coder enhanced by our algorithm outperforms original H.264 coder by approximately 0.8~1.5 dB
Qionghai Dai, Wenli Xu, Dongdong Zhu
ICASSP (2)3
2006 Color Light Field Block Truncation Compression using Hierarchical Bit-Plane Prediction
abstract
Compression and calculation efficiency are two key problems for light field coding. While nearly all works focus on one quality and neglect the other, this paper introduces a novel color block truncation algorithm using hierarchical bit-plane prediction (CBTC-HBPP) to chase for them both. This approach first performs two-color-truncation in block classes extracted from the whole light field. The yielding bit-plane and a const index table make decoding and rendering easy and fast. Hierarchical bit-plane prediction is employed to compressed bit-plane further. Different from motion and disparity estimation, this bit prediction has only a two-value result and it can be carried out fast. Both moderate compression and calculation efficiency are achieved.
Su Xue, Zhixiang You, Yebin Liu, Qionghai Dai
ICIP4
2006 A Decentralized Key Management Scheme in Overlay Multicast Network
abstract
The recent growth of the World Wide Web has sparked new research into using the Internet for novel types of group communication, like multiparty videoconferencing and real-time streaming. Multicast has the potential to be very useful, but it suffers from many problems like security. To achieve secure multicast communications, key management is one of the most critical problems. So far, a lot of multicast key management schemes have been proposed and most of them are centralized, which have the problem of "one point failure" and that the group controller is the bottleneck of the group. In order to solve these two problems, we propose a decentralized key management scheme (DKMS), using RSA key system as auxiliary keys. We analyze this scheme and find it has appropriate performance in security and scalability
Xingfeng Guo, Xiaodong Liu 0005, Qionghai Dai
ICME3
2006 Improved Similarity-Based Online Feature Selection in Region-Based Image Retrieval
abstract
To bridge the gap between high level semantic concepts and low level visual features in content-based image retrieval (CBIR), online feature selection is really required. An effective similarity-based online feature selection algorithm in region-based image retrieval (RBIR) systems was proposed by W. Jiang etc., but some parts of the algorithm need to be improved. In this paper, the above algorithm is modified in two aspects: (1) Adaptive mixture models based on mutual information theory are adopted to determine the codebook size. (2) A new method is proposed, which can select not only feature axes parallel to the original ones, but also combined feature axes. Experimental results on 10000 images show that the proposed method can improve the retrieval performance, and save the computational time
Qionghai Dai, Wenli Xu
ICME2
2006 An Improved Resource Reservation Algorithm for IEEE 802.15.3
abstract
In this paper, we propose an improved resource reservation algorithm ESRPT (enhanced shortest remaining processing time) for bursty traffic based on IEEE 802.15.3 standard. In this algorithm, each transmitter reports current fragments number of the first MSDU (MAC service data unit) and all of the fragments number of the remainder MSDUs in the pending transmission queue to PNC (piconet coordinator). In the next superframe, PNC firstly allocates part CTAs (channel time allocation) for each stream based on the remainder fragments number of the first MSDU by SRPT rule, then allocates remainder CTAs for each stream based on all fragments number of remainder MSDUs by the same SRPT rule. This algorithm decreases the job failure rates of delay-sensitive video streaming efficiently at the cost of slightly increasing communication overheads. Simulation results show that our proposed ESRPT method achieves better performance in QoS for multimedia streams compared to the existing SRPT schemes
Qionghai Dai, Qiufeng Wu
ICME2
2006 A Real Time Interactive Dynamic Light Field Transmission System
abstract
The ability to interactively and seamlessly roam in the scenario while watching a video through IP network is an exciting visual experience. In this work, we implemented a 3D TV system with real-time data acquisition, compression, Internet transmission, light field rendering, and free-viewpoint control of dynamic scenes. Our system consists of an 8times8 light field camera array, 16 producer PCs, a streaming server system and several clients. Multiple video streams are coded in a real time manner that each client can freely selects the streams for novel view rendering. Also, our system minimize the per-user transmission bit rate while maintaining multi-view simul-switching ability for each user. We believe that this is the first real-time Internet streaming system that can simultaneously guarantee real time free-view point control, data storage and support arbitrary number of users. The average transmission bit rate for end user is lower than 2 Mbps which is suitable for the broadband IP network
Yebin Liu, Qionghai Dai, Wenli Xu
ICME2
2006 A Neural Network Based Application Layer Multicast Routing Protocol
Qionghai Dai, Qiufeng Wu
ISNN (2)2
2006 A Neural Network Decision-Making Mechanism for Robust Video Transmission over 3G Wireless Network
Jianwei Wen, Qionghai Dai, Yihui Jin
ISNN (2)2
2006 Similarity-based online feature selection in content-based image retrieval
abstract
Content-based image retrieval (CBIR) has been more and more important in the last decade, and the gap between high-level semantic concepts and low-level visual features hinders further performance improvement. The problem of online feature selection is critical to really bridge this gap. In this paper, we investigate online feature selection in the relevance feedback learning process to improve the retrieval performance of the region-based image retrieval system. Our contributions are mainly in three areas. 1) A novel feature selection criterion is proposed, which is based on the psychological similarity between the positive and negative training sets. 2) An effective online feature selection algorithm is implemented in a boosting manner to select the most representative features for the current query concept and combine classifiers constructed over the selected features to retrieve images. 3) To apply the proposed feature selection method in region-based image retrieval systems, we propose a novel region-based representation to describe images in a uniform feature space with real-valued fuzzy features. Our system is suitable for online relevance feedback learning in CBIR by meeting the three requirements: learning with small size training set, the intrinsic asymmetry property of training samples, and the fast response requirement. Extensive experiments, including comparisons with many state-of-the-arts, show the effectiveness of our algorithm in improving the retrieval performance and saving the processing time.
Wei Jiang 0007, Guihua Er, Qionghai Dai, Jinwei Gu
IEEE Trans. Image Process.3
2005 Relevance Feedback Learning With Feature Selection In Region-Based Image Retrieval
abstract
Region-based image retrieval and relevance feedback are two important methods to bridge the gap between the low-level visual features and the high-level semantic concepts in content-based image retrieval. In this paper, we address the issue of introducing the relevance feedback mechanism into the region-based image retrieval with online feature selection during each feedback round. Our contribution is two-fold. (1) A novel region-based image representation is proposed. Based on a generative model, a fuzzy codebook is extracted from the original region-based features, which represents the images in a uniform real-value feature space. (2) A feature selection criterion is developed and an effective relevance feedback algorithm is implemented in a boosting manner to simultaneously select the optimal features from the fuzzy codebook, and generate a strong ensemble classifier over the selected features. Experimental results show that the proposed scheme can substantially improve the retrieval performance of the region-based image retrieval system.
Wei Jiang 0007, Guihua Er, Qionghai Dai, Lian Zhong, Yao Hou
ICASSP (2)3
2005 An Adaptive Hierarchical Clustering Protocol for Multimedia Overlay Multicast Applications
abstract
Adaptive Hierarchical Clustering Algorithm (AHCA) maps a flat topology to a hierarchical tree though output trees are incontrollable and are not suitable for multimedia. Prune-Relocate operation and Top Topologies operation are proposed in this paper to improve AHCA protocol and generate OM-AHCA trees. Numerical simulations show that OM-AHCA trees are compromise between AHCA trees and single-level topology flat protocol trees, which optimize the overall performance of single-level flat protocol and improve the degree metric of AHCA trees.
Qionghai Dai, Qiufeng Wu
ICME2
2005 Channel-adaptive hybrid ARQ/FEC for robust video transmission over 3G
abstract
This paper addresses the important issues of error control for video transmission over 3G. Based on the time-varying wireless channel conditions and the essential defects of the traditional hybrid ARQ for real-time service, the architecture of the channel-adaptive hybrid ARQ/FEC is presented, moreover an algorithm for encoder is given to automatically adjust the parity data length and the maximum number of retransmissions. The experimental studies show that the transmission efficiency of the proposed algorithm increased 13% than the traditional hybrid ARQ.
Jianwei Wen, Qionghai Dai, Yihui Jin
ICME2
2005 Fuzzy Neural Network for VBR MPEG Video Traffic Prediction
Xiaodong Liu 0005, Xiaokang Lin, Qionghai Dai
ISNN (3)4
2005 Affine-Invariant Image Retrieval Based on Wavelet Interest Points
abstract
This paper presents an affine-in variant image retrieval approach based on wavelet-based detector, which uses the space-tree property of the transform coefficients to estimate the interest points. Meanwhile, in order to retrieve images compressed by wavelet algorithm such as JPEG2000, the detector only uses the partial bit-planes of the wavelet coefficients to detect the interest points. To provide affine-invariant image matching, annular color histogram, annular texture histogram and spatial cohesion based on interest points are presented to describe image features. A series of experiments based on an image database consisting of 1000 images are performed to confirm the effectiveness of our method
Guiguang Ding, Qionghai Dai, Wenli Xu
MMSP2
2005 Adaptive Key Frame Selection Wyner-Ziv Video Coding
abstract
In Wyner-Ziv video coding, efficient compression is achieved by exploiting source statistics at the decoder only, which is radically different from conventional video coding. The performance of a Wyner-Ziv video codec is greatly dependent on the quality of reconstructed side information, which is an estimation version of current frame. Therefore the correlation between Wyner-Ziv (WZ) frame and key frame implicitly affects the performance of a Wyner-Ziv video codec. In this paper, an adaptive key frame selection method is proposed. Firstly, a simply interest points detector is utilized to detect interest points of a video frame. Then, a kind of interest points based measure, which could represent the correlation between video frames, is performed. If the number of different interest points between two video frames is below a threshold, then it should be transmitted as a key frame, otherwise as a WZ frame. Our experimental results are promising. About 1 dB gain in the quality of reconstructed frames has been achieved
Guiguang Ding, Qionghai Dai, Yaguang Yin
MMSP3
2005 Hidden annotation for image retrieval with long-term relevance feedback learning
Wei Jiang 0007, Guihua Er, Qionghai Dai, Jinwei Gu
Pattern Recognit.3
2004 Multiple boosting SVM active learning for image retrieval
abstract
Content-based image retrieval can be viewed as a classification problem, and the small sample size leaning difficulty makes it difficult for most CBIR classifiers to get satisfactory performance. In this paper, using the SVM classifier as the component classifier, the method of ensemble of classifiers is incorporated into the relevance feedback process to alleviate this problem from two aspects: (1) within each feedback round, multiple parallel component classifiers are constructed, one over one feature subspace individually, and then are merged together to get an ensemble classifier; (2) during feedback rounds, a boosting method is incorporated to sequentially combine the component classifiers over each feature subspace respectively, which further improves the classification result. Experiments over 5000 images show that the proposed method can improve the retrieval performance consistently, without loss of efficiency.
Wei Jiang 0007, Guihua Er, Qionghai Dai
ICASSP (3)3
2004 Improved rate allocation method based on sliding window for FGS video bit-stream
abstract
This paper proposes an improved rate allocation method, based on a sliding window, for FGS (fine granular scalability) video bit-streams. It can allocate the rate for every frame in the FGS video bit-stream rapidly, to smooth video quality under dynamical channel conditions. The paper formulates the FGS rate allocation problem under time-varying channel conditions using a sliding window to record the available network bandwidth. Under the constrained bandwidth, the proposed method first allocates rates in advance and creates a reference window by solving the problem of constant quality constrained FGS rate allocation. Then by refreshing the reference and sliding windows, the rate allocation of other frames can be solved rapidly. Theoretical analysis and simulation results for the method demonstrate its low complexity and high performance.
Guihua Er, Qionghai Dai
ICASSP (5)3
2004 A line-diamond parallel search algorithm for block motion estimation
abstract
The widespread use of block matching motion estimation (BMME) in video coding is due to its effectiveness and simplicity of implementation. This paper presents a novel fast BMME algorithm called the line-diamond parallel search (LDPS). The algorithm is based on the following two properties: the special directionality of the SAD distribution and the characteristics of the center-biased motion vector distribution. In addition, in order to increase the speed of search, the parallel processing idea is used in LDPS. That is to say LDPS realizes the coarse orientation and the accurate search in the same step. Our experimental results show that not only the processing speed of the LDPS algorithm is much higher than that of other fast algorithms, but also its accuracy of motion compensation is as nearly good as that of full search (FS).
Guiguang Ding, Qionghai Dai
ICIG2
2004 Fast mode decision for inter prediction in H.264
abstract
This paper proposes a fast inter prediction mode decision method for H.264. By pre-encoding a downsampled small image, the candidate inter block modes can be reduced to a small subset. The simulation result shows that our algorithm can achieve up to 50% complexity reduction with less than 0.2dB PSNR decrease.
Qionghai Dai, Dongdong Zhu
ICIP1
2004 Light field compression based on prediction propagating and wavelet packet
abstract
A novel data compression scheme of light field is presented. Different from the prior codecs, we perform image predicting and data coding in a set of subbands not directly in original images. In our approach, the original images are decomposed into subbands using wavelet packet transform, and the corresponding wavelet packet bases are divided into two parts: the predictable bases and the unpredictable bases by some criterions. In coding, a propagating algorithm and a rhombuses structure is used, and the subbands corresponding to the basis in predictable and unpredictable bases are added in sequence by relative energy until the reconstructing images meet the pre-established reconstruction quality. Experiments for two standard light fields verify the efficiency of our approach.
Qionghai Dai, Wenli Xu
ICIP2
2004 Multi-layer semantic representation learning for image retrieval
abstract
Long-term relevance feedback learning is an important learning mechanism in content-based image retrieval. In this paper, our work has two contributions: (1) A multilayer semantic representation (MSR) is proposed and an algorithm is implemented to automatically build the MSR for image database through long-term relevance feedback learning. (2) The accumulated MSR is incorporated with the short-term feedback learning to help subsequent users' retrieval. The MSR memorizes the multicorrelation among images and integrates these memories to build hidden semantic concepts for images, which are distributed in multiple semantic layers. In experiment, an MSR is built based on the real retrieval from 10 different users, which can precisely describe the hidden concepts underlying images and help to bridge the gap between high-level concepts and low-level features and thus improve the retrieval performance significantly.
Wei Jiang 0007, Guihua Er, Qionghai Dai
ICIP3
2004 Background-frame based motion compensation for video compression
abstract
This work presents a new kind of reference frame, i.e., the "background frame", to improve the accuracy of motion compensation in video compression. When an object moves in the scene of a video sequence, a certain part of the background is overlapped at first. After several frames, it begins to reappear. In the frames the background reappears, the blocks contained in the background part may be poorly predicted by neighboring blocks which have been overlapped in the previous frames. To solve this problem, the "background frame" is constructed, based on blocks that have kept unchanged in a certain number of continuous frames. The background frame is free from the influence of moving objects. So, when it is used as a reference jointly with the traditional previous frame, better prediction would be obtained. The experimental results show that the background frame based motion compensation (BFMC) algorithm can obtain increased coding efficiency compared to the JVT/AVC/H.264 standard with one reference frame or two reference frames.
Qionghai Dai, Wenli Xu, Dongdong Zhu
ICME2
2004 Data compression of light field using wavelet packet
abstract
A disparity-compensated compression scheme for light fields is presented. In contrast to prior disparity-compensated based codecs we perform disparity estimating and image predicting in subbands not directly in the original images. In our approach, firstly the original images are decomposed into subbands using the wavelet packet transform, and then the corresponding wavelet packet bases are divided into predictable bases and unpredictable bases by some criteria. In coding, the hierarchical structure proposed by Magnor (1999) is used, and the subbands corresponding to the predictable and unpredictable basis are added sequentially by relative energy until the reconstructed images meet the pre-established reconstruction quality. Except for high compression ratios our approach also provides scalability to some extent. Experiments for two standard light fields verify the efficiency of our approach
Qionghai Dai, Wenli Xu
ICME2
2004 Temporal scalable video transmission using multi-reference prediction chain coding
abstract
The work presents a novel multiple reference picture selection scheme, named multi-reference prediction chain, to encode video. Temporal scalability is obtained from the prediction scheme in which prediction chains are constructed for the pattern of prediction dependency between P-frames. The inter-chain independence also endows the video with error recovery ability. We dynamically manage the prediction chains or a group of frames using a rate-distortion optimization framework that adapts to the channel bandwidth so that the expected end-to-end video quality may be achieved under a given rate constraint. The expected error resilience capability may be achieved by the entire recovery of the error chains. In our algorithm, the video is pre-encoded to be transported partly with temporal scalability. Simulation results show that our algorithm is feasible in temporal scalability and has high performance in error resilience.
Xiaosen Lin, Qionghai Dai
ICME2
2004 Fast inter prediction mode decision for H.264
abstract
This work proposes a fast inter prediction mode decision method for H.264. By pre-encoding a down-sampled small image, the candidate inter block modes can be reduced to a small subset. The experimental result shows that our algorithm can reduce nearly 50% of encoding time with PSNR decrease less than 0.2dB.
Dongdong Zhu, Qionghai Dai
ICME2
2004 New algorithm for modulated complex lapped transform with symmetrical window function
abstract
An algorithm for the fast computation of modulated complex lapped transform (MCLT) with symmetrical window function is proposed. The method is based on two discrete cosine transforms (DCTs), two stages of butterfly operations and additional multiplications. The real or imaginary part of the MCLT coefficients can be independently obtained from one of the two blocks of DCT coefficients. The proposed algorithm not only reduces the computational complexity compared with previous algorithms, but also has regular architecture which is suitable for hardware implementations.
Qionghai Dai, Xinjian Chen 0002
IEEE Signal Process. Lett.1
2004 A novel VLSI architecture for multidimensional discrete wavelet transform
abstract
A novel VLSI architecture for multidimensional discrete wavelet transform (mD DWT) based on a systolic array is proposed. We divide the input mD image data into 2/sup m/ independent data streams, and then simultaneously pipeline them into a multi-filter chip, and finally obtain 2/sup m/ samples which are from different DWT subbands per clock cycles (ccs). The proposed architecture performs a decomposition of an N/sub 1//spl times/N/sub 2//spl times/.../spl times/N/sub m/ image in about N/sub 1/N/sub 2/...N/sub m//(2/sup m/-1) ccs and requires relatively lower hardware cost than previous architectures. Besides, the advantages of the proposed architecture include very simple hardware complexity, regular data flow and low control complexity.
Qionghai Dai, Xinjian Chen 0002, Chuang Lin 0002
IEEE Trans. Circuits Syst. Video Technol.1
2003 A novel VLSI architecture for multidimensional discrete wavelet transform
abstract
In this short paper we propose a novel VLSI architecture for multidimensional discrete wavelet transform (m-D DWT) based on systolic array and non-separable approach. The proposed architecture performs a decomposition of an N/sub 1/ /spl times/ N/sub 2/ /spl times/ ... /spl times/ N/sub m/ image in about N/sub 1/ N/sub 2/...N/sub m//(2/sup m/ $1) clock cycles (ccs). This result considerably speeds up other known architectures. Besides, the advantages of the proposed architecture include very simple hardware complexity, regular data flow and low control complexity.
Xinjian Chen 0002, Qionghai Dai
ICME2
2003 An improved RM algorithm for preventing streaming media tasks from starvation
abstract
A streaming media system is a soft real-time system according to streaming media service features. Some soft real-time tasks and best-effort tasks will be starved when using traditional RM algorithm to schedule streaming media services. Aiming at this problem, an improved RM algorithm with elastic priority is proposed in this paper. The algorithm can prevent periodic streaming media tasks from starvation, improving QoS of streaming media system. Some simulation experiments are given in the paper, and these experiments illustrate that the algorithm can reach our anticipative target.
Shuhua Peng, Xiaodong Liu 0005, Qionghai Dai
ICME3
2002 Fast tracking of semantic video object based on motion prediction and subregion extraction
abstract
Extraction quality and speed are two fundamental problems for semantic video object extraction. The paper introduces a novel fast-speed tracking algorithm for extracting semantic video objects from image sequences. First, the subregion that covers the contour of a semantic video object is extracted by using motion prediction and mathematical morphology operators to generate the inner and outer contour of the subregion. Then, an optimized tracking scheme is employed to track this particular subregion instead of the whole image. Results on real sequences show that this method greatly improves the processing speed compared to earlier tracking algorithms and that it keeps the quality of the extracted object. Therefore, it is a promising approach for real time processing systems based on video objects.
Qionghai Dai, Guihua Er
ICIP (3)2
2002 Subspaces of FMmlet transform
abstract
The subspaces of FM m let transform are investigated. It is shown that some of the existing transforms like the Fourier transform, short-time Fourier transform, Gabor transform, wavelet transform, chirplet transform, the mean of signal, and the FM −1 let transform, and the butterfly subspace are all special cases of FM m let transform. Therefore the FM m let transform is more flexible for delineating both the linear and nonlinear time-varying structures of a signal.
Hongxing Zou, Qionghai Dai, Guiming Chen, Yanda Li
Sci. China Ser. F Inf. Sci.2
2002 Nonexistence of cross-term free time-frequency distribution with concentration of Wigner-Ville distribution
abstract
Wigner-Ville distribution (WVD) is recognized as being a powerful tool and a nucleus in time-frequency representation (TFR) which gives an excellent time-frequency concentration, and more importantly, has many desirable properties. A major shortcoming of WVD is the inherent cross-term (CT) interference. Although solutions to this problem from the bulk of contributions to the literature concerning TFR are currently available, none has been able to completely eliminate the CT’s in WVD. It is therefore a common belief that if there exists an auxiliary time-frequency distribution (TFD) which has the same auto-terms (AT’s) as that in WVD, but has CT’s with the opposite sign, then, by adding the auxiliary TFD to WVD, an ideal TFD, which preserves the concentration of WVD while annihilating the CT’s, is readily obtained. However, we prove that the auxiliary TFD does not exist. Moreover, it is found that in general, CT free joint distributions with their concentrations close to that of WVD do not exist either.
Hongxing Zou, Xuguang Lu, Qionghai Dai, Yanda Li
Sci. China Ser. F Inf. Sci.3
2001 Parametric TFR via windowed exponential frequency modulated atoms
abstract
We propose a new atom, namely, the dilated and translated windowed exponential frequency modulated functions (FM/sup m/let) for compactly characterizing both the signal's time-invariant and time-varying spectral contents. The superiority of the proposed method to some existing time-frequency distributions (TFDs) is demonstrated using a bat sonar signal.
Hongxing Zou, Qionghai Dai, Renming Wang, Yanda Li
IEEE Signal Process. Lett.2