EDBT 2026 Demo / reviewers in the wild / expert
Lu Fang 0001
dblp:33/8116-1
· DBLP profile ↗
112ranked-venue papers
16as first author
36since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 89 · 14 first-author · 23 since 2021Artificial intelligence and machine learning · 39 · 1 first-author · 27 since 2021Systems, architecture and hardware · 7 · 2 first-authorDatabases, data management, data science and information retrieval · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RealLiFe: Real-Time Light Field Reconstruction via Hierarchical Sparse Gradient DescentabstractWith the rise of Extended Reality (XR) technology, there is a growing need for real-time light field reconstruction from sparse view inputs. Existing methods can be classified into offline techniques, which can generate high-quality novel views but at the cost of long inference/training time, and online methods, which either lack generalizability or produce unsatisfactory results. However, we have observed that the intrinsic sparse manifold of Multi-plane Images (MPI) enables a significant acceleration of light field reconstruction while maintaining rendering quality. Based on this insight, we introduce RealLiFe, a novel light field optimization method, which leverages the proposed Hierarchical Sparse Gradient Descent (HSGD) to produce high-quality light fields from sparse input images in real time. Technically, the coarse MPI of a scene is first generated using a 3D CNN, and it is further optimized leveraging only the scene content aligned sparse MPI gradients in a few iterations. Extensive experiments demonstrate that our method achieves comparable visual quality while being 100x faster on average than state-of-the-art offline methods and delivers better performance (about 2 dB higher in PSNR) compared to other online approaches. Yijie Deng, Tianpeng Lin, Jinzhi Zhang, Lu Fang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | Geo-NI: Geometry-Aware Neural Interpolation for Light Field RenderingabstractWe present a novel Geometry-aware Neural Interpolation (Geo-NI) framework for light field rendering. Previous learning-based approaches either perform direct interpolation via neural networks, which we dubbed Neural Interpolation (NI), or explore scene geometry for novel view synthesis, also known as Depth Image-Based Rendering (DIBR). Both kinds of approaches have their own strengths and weaknesses in addressing non-Lambert effect and large disparity problems. In this paper, we incorporate the ideas behind these two kinds of approaches by launching the NI within a specific DIBR pipeline. Specifically, a DIBR network in the proposed Geo-NI serves to construct a novel reconstruction cost volume for neural interpolated light fields sheared by different depth hypotheses. The reconstruction cost can be interpreted as an indicator reflecting the reconstruction quality under a certain depth hypothesis, and is further applied to guide the rendering of the final high angular resolution light field. To implement the Geo-NI framework more practically, we further propose an efficient modeling strategy to encode high-dimensional cost volumes using a lower-dimension network. By combining the superiorities of NI and DIBR, the proposed Geo-NI is able to render views with large disparities with the help of scene geometry while also reconstructing the non-Lambertian effect when depth is prone to be ambiguous. Extensive experiments on various datasets demonstrate the superior performance of the proposed geometry-aware light field rendering framework. Gaochang Wu, Yuemei Zhou, Lu Fang 0001, Yebin Liu, Tianyou Chai |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | TransGI: Real-Time Dynamic Global Illumination With Object-Centric Neural Transfer ModelabstractNeural rendering algorithms have revolutionized computer graphics, yet their impact on real-time rendering under arbitrary lighting conditions remains limited due to strict latency constraints in practical applications. The key challenge lies in formulating a compact yet expressive material representation. To address this, we propose TransGI, a novel neural rendering method for real-time, high-fidelity global illumination. It comprises an object-centric neural transfer model for material representation and a radiance-sharing lighting system for efficient illumination. Traditional BSDF representations and spatial neural material representations lack expressiveness, requiring thousands of ray evaluations to converge to noise-free colors. Conversely, real-time methods trade quality for efficiency by supporting only diffuse materials. In contrast, our object-centric neural transfer model achieves compactness and expressiveness through an MLP-based decoder and vertex-attached latent features, supporting glossy effects with low memory overhead. For dynamic, varying lighting conditions, we introduce local light probes capturing scene radiance, coupled with an across-probe radiance-sharing strategy for efficient probe generation. We implemented our method in a real-time rendering engine, combining compute shaders and CUDA-based neural networks. Experimental results demonstrate that our method achieves real-time performance of less than 10 ms to render a frame and significantly improved rendering quality compared to baseline methods. Yijie Deng, Lu Fang 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2024 | GigaHumanDet: Exploring Full-Body Detection on Gigapixel-Level ImagesabstractPerforming person detection in super-high-resolution images has been a challenging task. For such a task, modern detectors, which usually encode a box using center and width/height, struggle with accuracy due to two factors: 1) Human characteristic: people come in various postures and the center with high freedom is difficult to capture robust visual pattern; 2) Image characteristic: due to vast scale diversity of input (gigapixel-level), distance regression (for width and height) is hard to pinpoint, especially for a person, with substantial scale, who is near the camera. To address these challenges, we propose GigaHumanDet, an innovative solution aimed at further enhancing detection accuracy for gigapixel-level images. GigaHumanDet employs the corner modeling method to avoid the potential issues of a high degree of freedom in center pinpointing. To better distinguish similar-looking persons and enforce instance consistency of corner pairs, an instance-guided learning approach is designed to capture discriminative individual semantics. Further, we devise reliable shape-aware bodyness equipped with a multi-precision strategy as the human corner matching guidance to be appropriately adapted to the single-view large scene. Experimental results on PANDA and STCrowd datasets show the superiority and strong applicability of our design. Notably, our model achieves 82.4% in term of AP, outperforming current state-of-the-arts by more than 10%. Jinze Yang, Wenxi Li, Lu Fang 0001 |
AAAI | 7 |
| 2024 | GigaTraj: Predicting Long-term Trajectories of Hundreds of Pedestrians in Gigapixel Complex ScenesabstractPedestrian trajectory prediction is a well-established task with significant recent advancements. However, existing datasets are unable to fulfill the demand for studying minute-level long-term trajectory prediction, mainly due to the lack of high-resolution trajectory observation in the wide field of view (FoV). To bridge this gap, we introduce a novel dataset named GigaTraj, featuring videos capturing a wide FoV with ~4 ×104m2and high-resolution imagery at the gigapixel level. Furthermore, GigaTraj in-cludes comprehensive annotations such as bounding boxes, identity associations, world coordinates, group/interaction relationships, and scene semantics. Leveraging these multimodal annotations, we evaluate and validate the state-of-the-art approaches for minute-level long-term trajectory prediction in large-scale scenes. Extensive experiments and analyses have revealed that long-term prediction for pedestrian trajectories presents numerous challenges, indicating a vital new direction for trajectory research. The dataset is available at WWW.gigavision ai. Haozhe Lin, Chunyu Wei, Yunqi Zhao, Shanglong Li, Lu Fang 0001 |
CVPR | 7 |
| 2024 | When Visual Grounding Meets Gigapixel-Level Large-Scale Scenes: Benchmark and ApproachabstractVisual grounding refers to the process of associating natural language expressions with corresponding regions within an image. Existing benchmarks for visual grounding primarily operate within small-scale scenes with a few objects. Nevertheless, recent advances in imaging technology have enabled the acquisition of gigapixel-level images, providing high-resolution details in large-scale scenes containing numerous objects. To bridge this gap between imaging and computer vision benchmarks and make grounding more practically valuable, we introduce a novel dataset, named GigaGrounding, designed to challenge visual grounding models in gigapixel-level large-scale scenes. We extensively analyze and compare the dataset with existing benchmarks, demonstrating that GigaGrounding presents unique challenges such as large-scale scene understanding, gigapixel-level resolution, significant variations in object scales, and the “multi-hop expressions”. Furthermore, we introduced a simple yet effective grounding approach, which employs a “glance-to-zoom-in” paradigm and exhibits enhanced capabilities for addressing the GigaGrounding task. The dataset is available at www.gigavision.ai. M. Tao, Haozhe Lin, Heyuan Wang 0001, Lu Fang 0001 |
CVPR | 7 |
| 2024 | SPECAT: SPatial-spEctral Cumulative-Attention Transformer for High-Resolution Hyperspectral Image ReconstructionabstractCompressive spectral image reconstruction is a critical method for acquiring images with high spatial and spectral resolution. Current advanced methods, which involve de-signing deeper networks or adding more self-attention mod-ules, are limited by the scope of attention modules and the irrelevance of attentions across different dimensions. This leads to difficulties in capturing non-local mutation features in the spatial-spectral domain and results in a signif-icant parameter increase but only limited performance im-provement. To address these issues, we propose SPECAT, a SPatial-spEctral Cumulative-Attention Transformer de-signed for high-resolution hyperspectral image reconstruction. SPECAT utilizes Cumulative-Attention Blocks (CABs) within an efficient hierarchical framework to extract features from non-local spatial-spectral details. Furthermore, it employs a projection-object Dual-domain Loss Function (DLF) to integrate the optical path constraint, a physical aspect often overlooked in current methodologies. Ulti-mately, SPECAT not only significantly enhances the reconstruction quality of spectral details but also breaks through the bottleneck of mutual restriction between the cost and accuracy in existing algorithms. Our experimental re-sults demonstrate the superiority of SPECAT, achieving 40.3 dB in hyperspectral reconstruction benchmarks, out-performing the state-of-the-art (SOTA) algorithms by 1.2 dB while using only 5% of the network parameters and 10% of the computational cost. The code is available at https://github.com/THU-luvisionISPECAT. Zhiyang Yao, Xiaoyun Yuan, Lu Fang 0001 |
CVPR | 4 |
| 2024 | OmniSeg3D: Omniversal 3D Segmentation via Hierarchical Contrastive LearningabstractTowards holistic understanding of 3D scenes, a general 3D segmentation method is needed that can segment diverse objects without restrictions on object quantity or categories, while also reflecting the inherent hierarchical structure. To achieve this, we propose OmniSeg3D, an omniversal segmentation method aims for segmenting anything in 3D all at once. The key insight is to lift multi-view inconsistent 2D segmentations into a consistent 3D feature field through a hierarchical contrastive learning framework, which is accomplished by two steps. Firstly, we design a novel hierarchical representation based on category-agnostic 2D segmentations to model the multi-level relationship among pixels. Secondly, image features rendered from the 3D feature field are clustered at different levels, which can be further drawn closer or pushed apart according to the hierarchical relationship between different levels. In tackling the challenges posed by inconsistent 2D segmentations, this framework yields a global consistent 3D feature field, which further enables hierarchical segmentation, multi-object selection, and global discretization. Extensive experiments demonstrate the effectiveness of our method on high-quality 3D segmentation and accurate hierarchical structure understanding. A graphical user interface further facilitates flexible interaction for omniversal 3D segmentation. Haiyang Ying, Yixuan Yin 0001, Jinzhi Zhang, Tao Yu 0007, Ruqi Huang, Lu Fang 0001 |
CVPR | 7 |
| 2024 | Detecting Objects as Cascade CornersabstractThe corner-based detection paradigm enjoys the potential to produce high-quality boxes. But the development is constrained by three factors: 1) Hard to match corners. Heuristic corner matching algorithms can lead to incorrect boxes, especially when similar-looking objects co-occur. 2) Poor instance context. Two separate corners preserve few instance semantics, so it is difficult to guarantee getting both two class-specific corners on the same heatmap channel. 3) Unfriendly backbone. The training cost of the hourglass network is high. Accordingly, we build a novel corner-based framework, named Corner2Net. To achieve the corner-matching-free manner, we devise the cascade corner pipeline which progressively predicts the associated corner pair in two steps instead of synchronously searching two independent corners via parallel heads. Corner2Net decouples corner localization and object classification. Both two corners are class-agnostic and the instance-specific bottom-right corner further simplifies its search space. Meanwhile, RoI features with rich semantics are extracted for classification. Popular backbones (e.g., ResNeXt) can be easily connected to Corner2Net. Experimental results on COCO show Corner2Net surpasses all existing corner-based detectors by a large margin in accuracy and speed. Haorao Wei, Jinze Yang, Liangyu Xu, Lu Fang 0001 |
ECAI | 7 |
| 2024 | DynamicTrack: Advancing Gigapixel Tracking in Crowded ScenesabstractTracking in gigapixel scenarios holds numerous potential applications in video surveillance and pedestrian analysis. Existing algorithms attempt to perform tracking in crowded scenes by utilizing multiple cameras or group relationships. However, their performance significantly degrades when confronted with complex interaction and occlusion inherent in gigapixel images. In this paper, we introduce DynamicTrack, a dynamic tracking framework designed to address gigapixel tracking challenges in crowded scenes. In particular, we propose a dynamic detector that utilizes contrastive learning to jointly detect the head and body of pedestrians. Building upon this, we design a dynamic association algorithm that effectively utilizes head and body information for matching purposes. Extensive experiments show that our tracker achieves state-of-the-art performance on widely used tracking benchmarks specifically designed for gigapixel crowded scenes. Yunqi Zhao, Zheng Cao 0005, Ruqi Huang, Lu Fang 0001 |
ICME | 6 |
| 2024 | SparseFormer: Detecting Objects in HRW Shots via Sparse Vision TransformerabstractRecent years have seen an increase in the use of gigapixel-level image and video capture systems and benchmarks with high-resolution wide (HRW) shots. However, unlike close-up shots in the MS COCO dataset, the higher resolution and wider field of view raise unique challenges, such as extreme sparsity and huge scale changes, causing existing close-up detectors inaccuracy and inefficiency. In this paper, we present a novel model-agnostic sparse vision transformer, dubbed SparseFormer, to bridge the gap of object detection between close-up and HRW shots. The proposed SparseFormer selectively uses attentive tokens to scrutinize the sparsely distributed windows that may contain objects. In this way, it can jointly explore global and local attention by fusing coarse- and fine-grained features to handle huge scale changes. SparseFormer also benefits from a novel Cross-slice non-maximum suppression (C-NMS) algorithm to precisely localize objects from noisy windows and a simple yet effective multi-scale strategy to improve accuracy. Extensive experiments on two HRW benchmarks, PANDA and DOTA-v1.0, demonstrate that the proposed SparseFormer significantly improves detection accuracy (up to 5.8%) and speed (up to 3x) over the state-of-the-art approaches. Wenxi Li, Jilai Zheng, Haozhe Lin, Chao Ma 0004, Lu Fang 0001, Xiaokang Yang 0001 |
ACM Multimedia | 6 |
| 2024 | Bridging the gap between object detection in close-up and high-resolution wide shots
Wenxi Li, Jilai Zheng, Haozhe Lin, Chao Ma 0004, Lu Fang 0001, Xiaokang Yang 0001 |
Comput. Vis. Image Underst. | 6 |
| 2024 | GiganticNVS: Gigapixel Large-Scale Neural Rendering With Implicit Meta-Deformed ManifoldabstractThe rapid advances of high-performance sensation empowered gigapixel-level imaging/videography for large-scale scenes, yet the abundant details in gigapixel images were rarely valued in 3d reconstruction solutions. Bridging the gap between the sensation capacity and that of reconstruction requires to attack the large-baseline challenge imposed by the large-scale scenes, while utilizing the high-resolution details provided by the gigapixel images. This paper introduces GiganticNVS for gigapixel large-scale novel view synthesis (NVS). Existing NVS methods suffer from excessively blurred artifacts and fail on the full exploitation of image resolution, due to their inefficacy of recovering a faithful underlying geometry and the dependence on dense observations to accurately interpolate radiance. Our key insight is that, a highly-expressive implicit field with view-consistency is critical for synthesizing high-fidelity details from large-baseline observations. In light of this, we propose meta-deformed manifold, where meta refers to the locally defined surface manifold whose geometry and appearance are embedded into high-dimensional latent space. Technically, meta can be decoded as neural fields using an MLP (i.e., implicit representation). Upon this novel representation, multi-view geometric correspondence can be effectively enforced with featuremetric deformation and the reflectance field can be learned purely on the surface. Experimental results verify that the proposed method outperforms state-of-the-art methods both quantitatively and qualitatively, not only on the standard datasets containing complex real-world scenes with large baseline angles, but also on the challenging gigapixel-level ultra-large-scale benchmarks. Jinzhi Zhang, Ruqi Huang, Lu Fang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | DartBlur: Privacy Preservation with Detection Artifact SuppressionabstractNowadays, privacy issue has become a top priority when training AI algorithms. Machine learning algorithms are expected to benefit our daily life, while personal information must also be carefully protected from exposure. Facial information is particularly sensitive in this regard. Multiple datasets containing facial information have been taken offline, and the community is actively seeking solutions to remedy the privacy issues. Existing methods for privacy preservation can be divided into blur-based and face replacement-based methods. Owing to the advantages of review convenience and good accessibility, blur-based based methods have become a dominant choice in practice. However, blur-based methods would inevitably introduce training artifacts harmful to the performance of downstream tasks. In this paper, we propose a novel De-artifact Blurring (DartBlur) privacy-preserving method, which capitalizes on a DNN architecture to generate blurred faces. DartBlur can effectively hide facial privacy information while detection artifacts are simultaneously suppressed. We have designed four training objectives that particularly aim to improve review convenience and maximize detection artifact suppression. We associate the algorithm with an adversarial training strategy with a second-order optimization pipeline. Experimental results demonstrate that DartBlur outperforms the existing face-replacement method from both perspectives of review convenience and accessibility, and also shows an exclusive advantage in suppressing the training artifact compared to traditional blur-based methods. Our implementation is available at https://github.com/JaNg2333/DartBlur. Baowei Jiang, Haozhe Lin, Lu Fang 0001 |
CVPR | 6 |
| 2023 | Crowd3D: Towards Hundreds of People Reconstruction from a Single ImageabstractImage-based multi-person reconstruction in wide-field large scenes is critical for crowd analysis and security alert. However, existing methods cannot deal with large scenes containing hundreds of people, which encounter the challenges of large number of people, large variations in human scale, and complex spatial distribution. In this paper, we propose Crowd3D, the first framework to reconstruct the 3D poses, shapes and locations of hundreds of people with global consistency from a single large-scene image. The core of our approach is to convert the problem of complex crowd localization into pixel localization with the help of our newly defined concept, Human-scene Virtual Interaction Point (HVIP). To reconstruct the crowd with global consistency, we propose a progressive reconstruction network based on HVIP by pre-estimating a scene-level camera and a ground plane. To deal with a large number of persons and various human sizes, we also design an adaptive human-centric cropping scheme. Besides, we contribute a benchmark dataset, LargeCrowd, for crowd reconstruction in a large scene. Experimental results demonstrate the effectiveness of the proposed method. The code and the dataset are available at http://cic.tju.edu.cn/faculty/likun/projects/Crowd3D. Huili Cui, Haozhe Lin, Yukun Lai, Lu Fang 0001, Kun Li 0001 |
CVPR | 6 |
| 2023 | RealGraph: A Multiview Dataset for 4D Real-world Context Graph GenerationabstractUnderstanding 4D scene context in real world has become urgently critical for deploying sophisticated AI systems. In this paper, we propose a brand new scene understanding paradigm called "Context Graph Generation (CGG)", aiming at abstracting holistic semantic information in the complicated 4D world. The CGG task capitalizes on the calibrated multiview videos of a dynamic scene, and targets at recovering semantic information (coordination, trajectories and relationships) of the presented objects in the form of spatio-temporal context graph in 4D space. We also present a benchmark 4D video dataset "RealGraph", the first dataset tailored for the proposed CGG task. The raw data of RealGraph is composed of calibrated and synchronized multiview videos. We exclusively provide manual annotations including object 2D&3D bounding boxes, category labels and semantic relationships. We also make sure the annotated ID for every single object is temporally and spatially consistent. We propose the first CGG baseline algorithm, Multiview-based Context Graph Generation Network (MCGNet), to empirically investigate the legitimacy of CGG task on RealGraph dataset. We nevertheless reveal the great challenges behind this task and encourage the community to explore beyond our solution. Our project page is at https://github.com/THU-luvision/RealGraph. Haozhe Lin, Zequn Chen, Jinzhi Zhang, Ruqi Huang, Lu Fang 0001 |
ICCV | 7 |
| 2023 | PARF: Primitive-Aware Radiance Fusion for Indoor Scene Novel View SynthesisabstractThis paper proposes a method for fast scene radiance field reconstruction with strong novel view synthesis performance and convenient scene editing functionality. The key idea is to fully utilize semantic parsing and primitive extraction for constraining and accelerating the radiance field reconstruction process. To fulfill this goal, a primitive-aware hybrid rendering strategy was proposed to enjoy the best of both volumetric and primitive rendering. We further contribute a reconstruction pipeline conducts primitive parsing and radiance field learning iteratively for each input frame which successfully fuses semantic, primitive, and radiance information into a single framework. Extensive evaluations demonstrate the fast reconstruction ability, high rendering quality, and convenient editing functionality of our method. Haiyang Ying, Baowei Jiang, Jinzhi Zhang, Di Xu 0012, Tao Yu 0007, Qionghai Dai, Lu Fang 0001 |
ICCV | 7 |
| 2023 | RobustFusion: Robust Volumetric Performance Reconstruction Under Human-Object Interactions From Monocular RGBD StreamabstractHigh-quality 4D reconstruction of human performance with complex interactions to various objects is essential in real-world scenarios, which enables numerous immersive VR/AR applications. However, recent advances still fail to provide reliable performance reconstruction, suffering from challenging interaction patterns and severe occlusions, especially for the monocular setting. To fill this gap, in this paper, we propose RobustFusion, a robust volumetric performance reconstruction system for human-object interaction scenarios using only a single RGBD sensor, which combines various data-driven visual and interaction cues to handle the complex interaction patterns and severe occlusions. We propose a semantic-aware scene decoupling scheme to model the occlusions explicitly, with a segmentation refinement and robust object tracking to prevent disentanglement uncertainty and maintain temporal consistency. We further introduce a robust performance capture scheme with the aid of various data-driven cues, which not only enables re-initialization ability, but also models the complex human-object interaction patterns in a data-driven manner. To this end, we introduce a spatial relation prior to prevent implausible intersections, as well as data-driven interaction cues to maintain natural motions, especially for those regions under severe human-object occlusions. We also adopt an adaptive fusion scheme for temporally coherent human-object reconstruction with occlusion analysis and human parsing cue. Extensive experiments demonstrate the effectiveness of our approach to achieve high-quality 4D human performance reconstruction under complex human-object interactions whilst still maintaining the lightweight monocular setting. Zhuo Su 0006, Lan Xu 0003, Dawei Zhong, Zhong Li 0007, Fan Deng 0005, Shuxue Quan, Lu Fang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2022 | INS-Conv: Incremental Sparse Convolution for Online 3D SegmentationabstractWe propose INS-Conv, an INcremental Sparse Convolutional network which enables online accurate 3D semantic and instance segmentation. Benefiting from the incremental nature of RGB-D reconstruction, we only need to update the residuals between the reconstructed scenes of consecutive frames, which are usually sparse. For layer design, we define novel residual propagation rules for sparse convolution operations, achieving close approximation to standard sparse convolution. For network architecture, an uncertainty term is proposed to adaptively select which residual to update, further improving the inference accuracy and efficiency. Based on INS-Conv, an online joint 3D semantic and instance segmentation pipeline is proposed, reaching an inference speed of 15 FPS on GPU and 10 FPS on CPU. Experiments on ScanNetv2 and SceneNN datasets show that the accuracy of our method surpasses previous online methods by a large margin, and is on par with state-of-the-art offline methods. A live demo on portable devices further shows the superior performance of INS-Conv. Leyao Liu, Yun-Jou Lin, Lu Fang 0001 |
CVPR | 5 |
| 2022 | ParseMVS: Learning Primitive-aware Surface Representations for Sparse Multi-view StereopsisabstractMulti-view stereopsis (MVS) recovers 3D surfaces by finding dense photo-consistent correspondences from densely sampled images. In this paper, we tackle the challenging MVS task from sparsely sampled views (up to an order of magnitude fewer images), which is more practical and cost-efficient in applications. The major challenge comes from the significant correspondence ambiguity introduced by the severe occlusions and the highly skewed patches. On the other hand, such ambiguity can be resolved by incorporating geometric cues from the global structure. In light of this, we propose ParseMVS, boosting sparse MVS by learning the P rimitive-A waR e S urface rE presentation. In particular, on top of being aware of global structure, our novel representation further allows for the preservation of fine details including geometry, texture, and visibility. More specifically, the whole scene is parsed into multiple geometric primitives. On each of them, the geometry is defined as the displacement along the primitives' normal directions, together with the texture and visibility along each view direction. An unsupervised neural network is trained to learn these factors by progressively increasing the photo-consistency and render-consistency among all input images. Since the surface properties are changed locally in the 2D space of each primitive, ParseMVS can preserve global primitive structures while optimizing local details, handling the 'incompleteness' and the 'inaccuracy' problems. We experimentally demonstrate that ParseMVS constantly outperforms the state-of-the-art surface reconstruction method in both completeness and the overall score under varying sampling sparsity, especially under the extreme sparse-MVS settings. Beyond that, ParseMVS also shows great potential in compression, robustness, and efficiency. Haiyang Ying, Jinzhi Zhang, Zheng Cao 0005, Jing Xiao 0006, Ruqi Huang, Lu Fang 0001 |
ACM Multimedia | 7 |
| 2022 | ElasticMVS: Learning elastic part representation for self-supervised multi-view stereopsisabstractSelf-supervised multi-view stereopsis (MVS) attracts increasing attention for learning dense surface predictions from only a set of images without onerous ground-truth 3D training data for supervision. However, existing methods highly rely on the local photometric consistency, which fails to identify accurately dense correspondence in broad textureless and reflectance areas.In this paper, we show that geometric proximity such as surface connectedness and occlusion boundaries implicitly inferred from images could serve as reliable guidance for pixel-wise multi-view correspondences. With this insight, we present a novel elastic part representation which encodes physically-connected part segmentations with elastically-varying scales, shapes and boundaries. Meanwhile, a self-supervised MVS framework namely ElasticMVS is proposed to learn the representation and estimate per-view depth following a part-aware propagation and evaluation scheme. Specifically, the pixel-wise part representation is trained by a contrastive learning-based strategy, which increases the representation compactness in geometrically concentrated areas and contrasts otherwise. ElasticMVS iteratively optimizes a part-level consistency loss and a surface smoothness loss, based on a set of depth hypotheses propagated from the geometrically concentrated parts. Extensive evaluations convey the superiority of ElasticMVS in the reconstruction completeness and accuracy, as well as the efficiency and scalability. Particularly, for the challenging large-scale reconstruction benchmark, ElasticMVS demonstrates significant performance gain over both the supervised and self-supervised approaches. Jinzhi Zhang, Ruofan Tang, Zheng Cao 0005, Jing Xiao 0006, Ruqi Huang, Lu Fang 0001 |
NeurIPS | 6 |
| 2022 | Real-Time Globally Consistent Dense 3D Reconstruction With Online TexturingabstractHigh-quality reconstruction of 3D geometry and texture plays a vital role in providing immersive perception of the real world. Additionally, online computation enables the practical usage of 3D reconstruction for interaction. We present an RGBD-based globally-consistent dense 3D reconstruction approach, where high-quality (i.e., the spatial resolution of the RGB image) texture patches are mapped on high-resolution ([Formula: see text]) geometric models online. The whole pipeline uses merely the CPU computing of a portable device. For real-time geometric reconstruction with online texturing, we propose to solve the texture optimization problem with a simplified incremental MRF solver in the context of geometric reconstruction pipeline using sparse voxel sampling strategy. An efficient reference-based color adjustment scheme is also proposed to achieve consistent texture patch colors under inconsistent luminance situations. Quantitative and qualitative experiments demonstrate that our online scheme achieves a realistic visualization of the environment with more abundant details, while taking fairly compact memory consumption and much lower computational complexity than existing solutions. Siyuan Gu, Dawei Zhong, Shuxue Quan, Lu Fang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Revisiting Light Field Rendering With Deep Anti-Aliasing Neural NetworkabstractThe light field (LF) reconstruction is mainly confronted with two challenges, large disparity and the non-Lambertian effect. Typical approaches either address the large disparity challenge using depth estimation followed by view synthesis or eschew explicit depth information to enable non-Lambertian rendering, but rarely solve both challenges in a unified framework. In this paper, we revisit the classic LF rendering framework to address both challenges by incorporating it with advanced deep learning techniques. First, we analytically show that the essential issue behind the large disparity and non-Lambertian challenges is the aliasing problem. Classic LF rendering approaches typically mitigate the aliasing with a reconstruction filter in the Fourier domain, which is, however, intractable to implement within a deep learning pipeline. Instead, we introduce an alternative framework to perform anti-aliasing reconstruction in the image domain and analytically show comparable efficacy on the aliasing issue. To explore the full potential, we then embed the anti-aliasing framework into a deep neural network through the design of an integrated architecture and trainable parameters. The network is trained through end-to-end optimization using a peculiar training set, including regular LFs and unstructured LFs. The proposed deep learning pipeline shows a substantial superiority in solving both the large disparity and the non-Lambertian challenges compared with other state-of-the-art approaches. In addition to the view interpolation for an LF, we also show that the proposed pipeline also benefits light field view extrapolation. Gaochang Wu, Yebin Liu, Lu Fang 0001, Tianyou Chai |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | SurRF: Unsupervised Multi-View Stereopsis by Learning Surface Radiance FieldabstractThe recent success in supervised multi-view stereopsis (MVS) relies on the onerously collected real-world 3D data. While the latest differentiable rendering techniques enable unsupervised MVS, they are restricted to discretized (e.g., point cloud) or implicit geometric representation, suffering from either low integrity for a textureless region or less geometric details for complex scenes. In this paper, we propose SurRF, an unsupervised MVS pipeline by learning Surface Radiance Field, i.e., a radiance field defined on a continuous and explicit 2D surface. Our key insight is that, in a local region, the explicit surface can be gradually deformed from a continuous initialization along view-dependent camera rays by differentiable rendering. That enables us to define the radiance field only on a 2D deformable surface rather than in a dense volume of 3D space, leading to compact representation while maintaining complete shape and realistic texture for large-scale complex scenes. We experimentally demonstrate that the proposed SurRF produces competitive results over the-state-of-the-art on various real-world challenging scenes, without any 3D supervision. Moreover, SurRF shows great potential in owning the joint advantages of mesh (scene manipulation), continuous surface (high geometric resolution), and radiance field (realistic rendering). Jinzhi Zhang, Mengqi Ji, Zhiwei Xu 0002, Shengjin Wang, Lu Fang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2022 | GigaMVS: A Benchmark for Ultra-Large-Scale Gigapixel-Level 3D ReconstructionabstractMultiview stereopsis (MVS) methods, which can reconstruct both the 3D geometry and texture from multiple images, have been rapidly developed and extensively investigated from the feature engineering methods to the data-driven ones. However, there is no dataset containing both the 3D geometry of large-scale scenes and high-resolution observations of small details to benchmark the algorithms. To this end, we present GigaMVS, the first gigapixel-image-based 3D reconstruction benchmark for ultra-large-scale scenes. The gigapixel images, with both wide field-of-view and high-resolution details, can clearly observe both thePalace-scale scene structure andRelievo-scale local details. The ground-truth geometry is captured by the laser scanner, which covers ultra-large-scale scenes with an average area of 8667 m$^2$and a maximum area of 32007 m$^2$. Owing to the extremely large scale, complex occlusion, and gigapixel-level images, GigaMVS exposes problems that emerge from the poor scalability and efficiency of the existing MVS algorithms. We thoroughly investigate the state-of-the-art methods in terms of geometric and textural measurements, which point to the weakness of the existing methods and promising opportunities for future works. We believe that GigaMVS can benefit the community of 3D reconstruction and support the development of novel algorithms balancing robustness, scalability and accuracy. Jinzhi Zhang, Shi Mao, Mengqi Ji, Zequn Chen, Xiaoyun Yuan, Qionghai Dai, Lu Fang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 10 |
| 2022 | BuildingFusion: Semantic-Aware Structural Building-Scale 3D ReconstructionabstractScalable geometry reconstruction and understanding is an important yet unsolved task. Current methods often suffer from false loop closures when there are similar-looking rooms in the scene, and often lack online scene understanding. We propose BuildingFusion, a semantic-aware structural building-scale reconstruction system, which not only allows building-scale dense reconstruction collaboratively, but also provides semantic and structural information on-the-fly. Technically, the robustness to similar places is enabled by a novel semantic-aware room-level loop closure detection(LCD) method. The insight lies in that even though local views may look similar in different rooms, the objects inside and their locations are usually different, implying that the semantic information forms a unique and compact representation for place recognition. To achieve that, a 3D convolutional network is used to learn instance-level embeddings for similarity measurement and candidate selection, followed by a graph matching module for geometry verification. On the system side, we adopt a centralized architecture to enable collaborative scanning. Each agent reconstructs a part of the scene, and the combination is activated when the overlaps are found using room-level LCD, which is performed on the server. Extensive comparisons demonstrate the superiority of the semantic-aware room-level LCD over traditional image-based LCD. Live demo on the real-world building-scale scenes shows the feasibility of our method with robust, collaborative, and real-time performance. Lan Xu 0003, Lu Fang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Learning Residual Color for Novel View SynthesisabstractScene Representation Networks (SRN) have been proven as a powerful tool for novel view synthesis in recent works. They learn a mapping function from the world coordinates of spatial points to radiance color and the scene's density using a fully connected network. However, scene texture contains complex high-frequency details in practice that is hard to be memorized by a network with limited parameters, leading to disturbing blurry effects when rendering novel views. In this paper, we propose to learn 'residual color' instead of 'radiance color' for novel view synthesis, i.e., the residuals between surface color and reference color. Here the reference color is calculated based on spatial color priors, which are extracted from input view observations. The beauty of such a strategy lies in that the residuals between radiance color and reference are close to zero for most spatial points thus are easier to learn. A novel view synthesis system that learns the residual color using SRN is presented in this paper. Experiments on public datasets demonstrate that the proposed method achieves competitive performance in preserving high-resolution details, leading to visually more pleasant results than the state of the arts. Dawei Zhong, Lu Fang 0001 |
IEEE Trans. Image Process. | 5 |
| 2022 | MulayCap: Multi-Layer Human Performance Capture Using a Monocular Video CameraabstractWe introduce MulayCap, a novel human performance capture method using a monocular video camera without the need for pre-scanning. The method uses "multi-layer" representations for geometry reconstruction and texture rendering, respectively. For geometry reconstruction, we decompose the clothed human into multiple geometry layers, namely a body mesh layer and a garment piece layer. The key technique behind is a Garment-from-Video (GfV) method for optimizing the garment shape and reconstructing the dynamic cloth to fit the input video sequence, based on a cloth simulation model which is effectively solved with gradient descent. For texture rendering, we decompose each input image frame into a shading layer and an albedo layer, and propose a method for fusing a fixed albedo map and solving for detailed garment geometry using the shading layer. Compared with existing single view human performance capture systems, our "multi-layer" approach bypasses the tedious and time consuming scanning step for obtaining a human specific mesh template. Experimental results demonstrate that MulayCap produces realistic rendering of dynamically changing details that has not been achieved in any previous monocular video camera systems. Benefiting from its fully semantic modeling, MulayCap can be applied to various important editing applications, such as cloth editing, re-targeting, relighting, and AR applications. Zhaoqi Su, Weilin Wan 0001, Tao Yu 0007, Lingjie Liu, Lu Fang 0001, Wenping Wang 0001, Yebin Liu |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2021 | A2-FPN: Attention Aggregation Based Feature Pyramid Network for Instance SegmentationabstractLearning pyramidal feature representations is crucial for recognizing object instances at different scales. Feature Pyramid Network (FPN) is the classic architecture to build a feature pyramid with high-level semantics throughout. However, intrinsic defects in feature extraction and fusion inhibit FPN from further aggregating more discriminative features. In this work, we propose Attention Aggregation based Feature Pyramid Network (A2-FPN), to improve multi-scale feature learning through attention-guided feature aggregation. In feature extraction, it extracts discriminative features by collecting-distributing multi-level global context features, and mitigates the semantic information loss due to drastically reduced channels. In feature fusion, it aggregates complementary information from adjacent features to generate location-wise reassembly kernels for content-aware sampling, and employs channel-wise reweighting to enhance the semantic consistency before element-wise addition. A2-FPN shows consistent gains on different instance segmentation frameworks. By replacing FPN with A2-FPN in Mask R-CNN, our model boosts the performance by 2.1% and 1.6% mask AP when using ResNet-50 and ResNet-101 as backbone, respectively. Moreover, A2-FPN achieves an improvement of 2.0% and 1.4% mask AP when integrated into the strong baselines such as Cascade Mask R-CNN and Hybrid Task Cascade. Yali Li 0001, Lu Fang 0001, Shengjin Wang |
CVPR | 3 |
| 2021 | Data-Uncertainty Guided Multi-Phase Learning for Semi-Supervised Object DetectionabstractIn this paper, we delve into semi-supervised object detection where unlabeled images are leveraged to break through the upper bound of fully-supervised object detection. Previous semi-supervised methods based on pseudo labels are severely degenerated by noise and prone to overfit to noisy labels, thus are deficient in learning different unlabeled knowledge well. To address this issue, we propose a data-uncertainty guided multi-phase learning method for semisupervised object detection. We comprehensively consider divergent types of unlabeled images according to their difficulty levels, utilize them in different phases, and ensemble models from different phases together to generate ultimate results. Image uncertainty guided easy data selection and region uncertainty guided RoI Re-weighting are involved in multi-phase learning and enable the detector to concentrate on more certain knowledge. Through extensive experiments on PASCAL VOC and MS COCO, we demonstrate that our method behaves extraordinarily compared to baseline approaches and outperforms them by a large margin, more than 3% on VOC and 2% on COCO. Zhenyu Wang 0005, Yali Li 0001, Lu Fang 0001, Shengjin Wang |
CVPR | 4 |
| 2021 | LocalTrans: A Multiscale Local Transformer Network for Cross-Resolution Homography EstimationabstractCross-resolution image alignment is a key problem in multiscale gigapixel photography, which requires to estimate homography matrix using images with large resolution gap. Existing deep homography methods concatenate the input images or features, neglecting the explicit formulation of correspondences between them, which leads to degraded accuracy in cross-resolution challenges. In this paper, we consider the cross-resolution homography estimation as a multimodal problem, and propose a local transformer network embedded within a multiscale structure to explicitly learn correspondences between the multimodal inputs, namely, input images with different resolutions. The proposed local transformer adopts a local attention map specifically for each position in the feature. By combining the local transformer with the multiscale structure, the network is able to capture long-short range correspondences efficiently and accurately. Experiments on both the MS-COCO dataset and the real-captured cross-resolution dataset show that the proposed network outperforms existing state-of-the-art feature-based and deep-learning-based homography estimation methods, and is able to accurately align images under 10× resolution gap. Ruizhi Shao, Gaochang Wu, Yuemei Zhou, Ying Fu 0001, Lu Fang 0001, Yebin Liu |
ICCV | 5 |
| 2021 | SurfaceNet+: An End-to-end 3D Neural Network for Very Sparse Multi-View StereopsisabstractMulti-view stereopsis (MVS) tries to recover the 3D model from 2D images. As the observations become sparser, the significant 3D information loss makes the MVS problem more challenging. Instead of only focusing on densely sampled conditions, we investigate sparse-MVS with large baseline angles since the sparser sensation is more practical and more cost-efficient. By investigating various observation sparsities, we show that the classical depth-fusion pipeline becomes powerless for the case with a larger baseline angle that worsens the photo-consistency check. As another line of the solution, we present SurfaceNet+, a volumetric method to handle the 'incompleteness' and the 'inaccuracy' problems induced by a very sparse MVS setup. Specifically, the former problem is handled by a novel volume-wise view selection approach. It owns superiority in selecting valid views while discarding invalid occluded views by considering the geometric prior. Furthermore, the latter problem is handled via a multi-scale strategy that consequently refines the recovered geometry around the region with the repeating pattern. The experiments demonstrate the tremendous performance gap between SurfaceNet+ and state-of-the-art methods in terms of precision and recall. Under the extreme sparse-MVS settings in two datasets, where existing methods can only return very few points, SurfaceNet+ still works as well as in the dense MVS setting. Mengqi Ji, Jinzhi Zhang, Qionghai Dai, Lu Fang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2021 | CrossNet++: Cross-Scale Large-Parallax Warping for Reference-Based Super-ResolutionabstractThe ability of camera arrays to efficiently capture higher space-bandwidth product than single cameras has led to various multiscale and hybrid systems. These systems play vital roles in computational photography, including light field imaging, 360 VR camera, gigapixel videography, etc. One of the critical tasks in multiscale hybrid imaging is matching and fusing cross-resolution images from different cameras under perspective parallax. In this paper, we investigate the reference-based super-resolution (RefSR) problem associated with dual-camera or multi-camera systems. RefSR consists of super-resolving a low-resolution (LR) image given an external high-resolution (HR) reference image, where they suffer both a significant resolution gap ( 8×) and large parallax ( ∼ 10% pixel displacement). We present CrossNet++, an end-to-end network containing novel two-stage cross-scale warping modules, image encoder and fusion decoder. The stage I learns to narrow down the parallax distinctively with the strong guidance of landmarks and intensity distribution consensus. Then the stage II operates more fine-grained alignment and aggregation in feature domain to synthesize the final super-resolved image. To further address the large parallax, new hybrid loss functions comprising warping loss, landmark loss and super-resolution loss are proposed to regularize training and enable better convergence. CrossNet++ significantly outperforms the state-of-art on light field datasets as well as real dual-camera data. We further demonstrate the generalization of our framework by transferring it to video super-resolution and video denoising. Haitian Zheng, Yinheng Zhu, Xiaoyun Yuan, David J. Brady, Lu Fang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2021 | Boosting Single Image Super-Resolution Learnt From Implicit Multi-Image PriorabstractLearning-based single image super-resolution (SISR) aims to learn a versatile mapping from low resolution (LR) image to its high resolution (HR) version. The critical challenge is to bias the network training towards continuous and sharp edges. For the first time in this work, we propose an implicit boundary prior learnt from multi-view observations to significantly mitigate the challenge in SISR we outline. Specifically, the multi-image prior that encodes both disparity information and boundary structure of the scene supervise a SISR network for edge-preserving. For simplicity, in the training procedure of our framework, light field (LF) serves as an effective multi-image prior, and a hybrid loss function jointly considers the content, structure, variance as well as disparity information from 4D LF data. Consequently, for inference, such a general training scheme boosts the performance of various SISR networks, especially for the regions along edges. Extensive experiments on representative backbone SISR architectures constantly show the effectiveness of the proposed method, leading to around 0.6 dB gain without modifying the network architecture. Dingjian Jin, Mengqi Ji, Lan Xu 0003, Gaochang Wu, Lu Fang 0001 |
IEEE Trans. Image Process. | 6 |
| 2021 | Spatial-Angular Attention Network for Light Field ReconstructionabstractTypical learning-based light field reconstruction methods demand in constructing a large receptive field by deepening their networks to capture correspondences between input views. In this paper, we propose a spatial-angular attention network to perceive non-local correspondences in the light field, and reconstruct high angular resolution light field in an end-to-end manner. Motivated by the non-local attention mechanism (Wang et al., 2018; Zhang et al., 2019), a spatial-angular attention module specifically for the high-dimensional light field data is introduced to compute the response of each query pixel from all the positions on the epipolar plane, and generate an attention map that captures correspondences along the angular dimension. Then a multi-scale reconstruction structure is proposed to efficiently implement the non-local attention in the low resolution feature space, while also preserving the high frequency components in the high-resolution feature space. Extensive experiments demonstrate the superior performance of the proposed spatial-angular attention network for reconstructing sparsely-sampled light fields with Non-Lambertian effects. Gaochang Wu, Yingqian Wang 0002, Yebin Liu, Lu Fang 0001, Tianyou Chai |
IEEE Trans. Image Process. | 4 |
| 2021 | FlyFusion: Realtime Dynamic Scene Reconstruction Using a Flying Depth CameraabstractWhile dynamic scene reconstruction has made revolutionary progress from the earliest setup using a mass of static cameras in studio environment to the latest egocentric or hand-held moving camera based schemes, it is still restricted by the recording volume, user comfortability, human labor and expertise. In this paper, a novel solution is proposed through a real-time and robust dynamic fusion scheme using a single flying depth camera, denoted as FlyFusion. By proposing a novel topology compactness strategy for effectively regularizing the complex topology changes, and the Geometry And Motion Energy (GAME) metric for guiding the viewpoint optimization in the volumetric space, FlyFusion succeeds to enable intelligent viewpoint selection based on the immediate dynamic reconstruction result. The merit of FlyFusion lies in its concurrent robustness, efficiency, and adaptation in producing fused and denoised 3D geometry and motions of a moving target interacting with different non-rigid objects in a large space. Lan Xu 0003, Yebin Liu, Lu Fang 0001 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2020 | OccuSeg: Occupancy-Aware 3D Instance Segmentationabstract3D instance segmentation, with a variety of applications in robotics and augmented reality, is in large demands these days. Unlike 2D images that are projective observations of the environment, 3D models provide metric reconstruction of the scenes without occlusion or scale ambiguity. In this paper, we define “3D occupancy size”, as the number of voxels occupied by each instance. It owns advantages of robustness in prediction, on which basis, OccuSeg, an occupancy-aware 3D instance segmentation scheme is proposed. Our multi-task learning produces both occupancy signal and embedding representations, where the training of spatial and feature embeddings varies with their difference in scale-aware. Our clustering scheme benefits from the reliable comparison between the predicted occupancy size and the clustered occupancy size, which encourages hard samples being correctly clustered and avoids over segmentation. The proposed approach achieves state-of-theart performance on 3 real-world datasets, i.e. ScanNetV2, S3DIS and SceneNN, while maintaining high efficiency. Lan Xu 0003, Lu Fang 0001 |
CVPR | 4 |
| 2020 | PANDA: A Gigapixel-Level Human-Centric Video DatasetabstractWe present PANDA, the first gigaPixel-level humAN-centric viDeo dAtaset, for large-scale, long-term, and multi-object visual analysis. The videos in PANDA were captured by a gigapixel camera and cover real-world scenes with both wide field-of-view (~1 square kilometer area) and high-resolution details (~gigapixel-level/frame). The scenes may contain 4k head counts with over 100× scale variation. PANDA provides enriched and hierarchical ground-truth annotations, including 15,974.6k bounding boxes, 111.8k fine-grained attribute labels, 12.7k trajectories, 2.2k groups and 2.9k interactions. We benchmark the human detection and tracking tasks. Due to the vast variance of pedestrian pose, scale, occlusion and trajectory, existing approaches are challenged by both accuracy and efficiency. Given the uniqueness of PANDA with both wide FoV and high resolution, a new task of interaction-aware group detection is introduced. We design a `global-to-local zoom-in' framework, where global trajectories and local interactions are simultaneously encoded, yielding promising results. We believe PANDA will contribute to the community of artificial intelligence and praxeology by understanding human behaviors and interactions in large-scale real-world scenes. PANDA Website: http://www.panda-dataset.com. Xiya Zhang, Yinheng Zhu, Xiaoyun Yuan, Liuyu Xiang, Zerun Wang, Guiguang Ding, David J. Brady, Qionghai Dai, Lu Fang 0001 |
CVPR | 11 |
| 2020 | EventCap: Monocular 3D Capture of High-Speed Human Motions Using an Event CameraabstractThe high frame rate is a critical requirement for capturing fast human motions. In this setting, existing markerless image-based methods are constrained by the lighting requirement, the high data bandwidth and the consequent high computation overhead. In this paper, we propose EventCap - the first approach for 3D capturing of high-speed human motions using a single event camera. Our method combines model-based optimization and CNN-based human pose detection to capture high frequency motion details and to reduce the drifting in the tracking. As a result, we can capture fast motions at millisecond resolution with significantly higher data efficiency than using high frame rate videos. Experiments on our new event-based fast human motion dataset demonstrate the effectiveness and accuracy of our method, as well as its robustness to challenging lighting conditions. Lan Xu 0003, Weipeng Xu, Vladislav Golyanik, Marc Habermann, Lu Fang 0001, Christian Theobalt |
CVPR | 5 |
| 2020 | RobustFusion: Human Volumetric Capture with Data-Driven Visual Cues Using a RGBD Camera
Zhuo Su 0006, Lan Xu 0003, Zerong Zheng, Tao Yu 0007, Yebin Liu, Lu Fang 0001 |
ECCV (4) | 6 |
| 2020 | Multiscale-VR: Multiscale Gigapixel 3D Panoramic Videography for Virtual RealityabstractCreating virtual reality (VR) content with effective imaging systems has attracted significant attention worldwide following the broad applications of VR in various fields, including entertainment, surveillance, sports, etc. However, due to the inherent trade-off between field-of-view and resolution of the imaging system as well as the prohibitive computational cost, live capturing and generating multiscale 360° 3D video content at an eye-limited resolution to provide immersive VR experiences confront significant challenges. In this work, we propose Multiscale-VR, a multiscale unstructured camera array computational imaging system for high-quality gigapixel 3D panoramic videography that creates the six-degree-of-freedom multiscale interactive VR content. The Multiscale-VR imaging system comprises scalable cylindrical-distributed global and local cameras, where global stereo cameras are stitched to cover 360° field-of-view, and unstructured local monocular cameras are adapted to the global camera for flexible high-resolution video streaming arrangement. We demonstrate that a high-quality gigapixel depth video can be faithfully reconstructed by our deep neural network-based algorithm pipeline where the global depth via stereo matching and the local depth via high-resolution RGB-guided refinement are associated. To generate the immersive 3D VR content, we present a three-layer rendering framework that includes an original layer for scene rendering, a diffusion layer for handling occlusion regions, and a dynamic layer for efficient dynamic foreground rendering. Our multiscale reconstruction architecture enables the proposed prototype system for rendering highly effective 3D, 360° gigapixel live VR video at 30 fps from the captured high-throughput multiscale video sequences. The proposed multiscale interactive VR content generation approach by using a heterogeneous camera system design, in contrast to the existing single-scale VR imaging systems with structured homogeneous cameras, will open up new avenues of research in VR and provide an unprecedented immersive experience benefiting various novel applications. Anke Zhang, Xiaoyun Yuan, Sebastian Beetschen, Lan Xu 0003, Qionghai Dai, Lu Fang 0001 |
ICCP | 10 |
| 2020 | Zoom in to the Details of Human-Centric VideosabstractPresenting high-resolution (HR) human appearance is always critical for the human-centric videos. However, current imagery equipment can hardly capture HR details all the time. Existing super-resolution algorithms barely mitigate the problem by only considering universal and low-level priors of image patches. In contrast, our algorithm is under bias towards the human body super-resolution by taking advantage of high-level prior defined by HR human appearance. Firstly, a motion analysis module extracts inherent motion pattern from the HR reference video to refine the pose estimation of the low-resolution (LR) sequence. Furthermore, a human body reconstruction module maps the HR texture in the reference frames onto a 3D mesh model. Consequently, the input LR videos get super-resolved HR human sequences are generated conditioned on the original LR videos as well as few HR reference frames. Experiments on an existing dataset and real-world data captured by hybrid cameras show that our approach generates superior visual quality of human body compared with the traditional method. Guanghan Li, Mengqi Ji, Xiaoyun Yuan, Lu Fang 0001 |
ICIP | 5 |
| 2020 | All-in-depth via Cross-baseline Light Field CameraabstractLight-field (LF) camera holds great promise for passive/general depth estimation benefited from high angular resolution, yet suffering small baseline for distanced region. While stereo solution with large baseline is superior to handle distant scenarios, the problem of limited angular resolution becomes bothering for near objects. Aiming for all-in-depth solution, we propose a cross-baseline LF camera using a commercial LF camera and a monocular camera, which naturally form a 'stereo camera' enabling compensated baseline for LF camera. The idea is simple yet non-trivial, due to the significant angular resolution gap and baseline gap between LF and stereo cameras. Dingjian Jin, Anke Zhang, Gaochang Wu, Haoqian Wang, Lu Fang 0001 |
ACM Multimedia | 6 |
| 2020 | UnstructuredFusion: Realtime 4D Geometry and Texture Reconstruction Using Commercial RGBD CamerasabstractA high-quality 4D geometry and texture reconstruction for human activities usually requires multiview perceptions via highly structured multi-camera setup, where both the specifically designed cameras and the tedious pre-calibration restrict the popularity of professional multi-camera systems for daily applications. In this paper, we propose UnstructuredFusion, a practicable realtime markerless human performance capture method using unstructured commercial RGBD cameras. Along with the flexible hardware setup using simply three unstructured RGBD cameras without any careful pre-calibration, the challenge 4D reconstruction through multiple asynchronous videos is solved by proposing three novel technique contributions, i.e., online multi-camera calibration, skeleton warping based non-rigid tracking, and temporal blending based atlas texturing. The overall insights behind lie in the solid global constraints of human body and human motion which are modeled by the skeleton and the skeleton warping, respectively. Extensive experiments such as allocating three cameras flexibly in a handheld way demonstrate that the proposed UnstructuredFusion achieves high-quality 4D geometry and texture reconstruction without tiresome pre-calibration, liberating the cumbersome hardware and software restrictions in conventional structured multi-camera system, while eliminating the inherent occlusion issues of the single camera setup. Lan Xu 0003, Zhuo Su 0006, Tao Yu 0007, Yebin Liu, Lu Fang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2020 | Guest Editorial Introduction to the Special Section on Intelligent Visual Content Analysis and UnderstandingabstractVisual content analysis and understanding attract tremendous attention because of its potentially wide range of applications including human activity analysis, automated photo face tagging, multicamera tracking, crowded counting, and biometric security. With recent progress in end-to-end differentiable learning, the accuracy of algorithms has been significantly improved and even outperforms humans in some tasks. In addition, multimodality methods, targeting on making full use of various visual data sources, are further investigated. These developments contribute to the innovations of two core modules for a typical intelligent vision system, i.e., image and video description and recognition, which are critical for the success of the visual content analysis and understanding in more complex and challenging open world. Hongliang Li 0001, Lu Fang 0001, Tianzhu Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | Live Semantic 3D Perception for Immersive Augmented RealityabstractSemantic understanding of 3D environments is critical for both the unmanned system and the human involved virtual/augmented reality (VR/AR) immersive experience. Spatially-sparse convolution, taking advantage of the intrinsic sparsity of 3D point cloud data, makes high resolution 3D convolutional neural networks tractable with state-of-the-art results on 3D semantic segmentation problems. However, the exhaustive computations limits the practical usage of semantic 3D perception for VR/AR applications in portable devices. In this paper, we identify that the efficiency bottleneck lies in the unorganized memory access of the sparse convolution steps, i.e., the points are stored independently based on a predefined dictionary, which is inefficient due to the limited memory bandwidth of parallel computing devices (GPU). With the insight that points are continuous as 2D surfaces in 3D space, a chunk-based sparse convolution scheme is proposed to reuse the neighboring points within each spatially organized chunk. An efficient multi-layer adaptive fusion module is further proposed for employing the spatial consistency cue of 3D data to further reduce the computational burden. Quantitative experiments on public datasets demonstrate that our approach works 11× faster than previous approaches with competitive accuracy. By implementing both semantic and geometric 3D reconstruction simultaneously on a portable tablet device, we demo a foundation platform for immersive AR applications. Yinheng Zhu, Lan Xu 0003, Lu Fang 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2019 | SPI-Optimizer: An Integral-Separated PI Controller for Stochastic OptimizationabstractTo overcome the oscillation problem in the classical momentum-based optimizer, recent work associates it with the proportional-integral (PI) controller, and artificially adds D term producing a PID controller. It suppresses oscillation with the sacrifice of introducing extra hyper-parameter. In this paper, we analyze that the fluctuation problem relates to the lag effect of the integral (I) term, and propose SPI-Optimizer, an integral-Separated PI controller based optimizer WITHOUT introducing extra hyper-parameter. It separates momentum term adaptively when the inconsistency of current and historical gradient direction occurs. Extensive experiments demonstrate that SPI-Optimizer generalizes well on popular network architectures to eliminate the oscillation, and owns competitive performance with faster convergence speed (up to 40% epochs reduction ratio) and more accurate classification result on MNIST, CIFAR10, and CIFAR100 (up to 27.5% error reduction ratio) than state-of-the-art methods. Mengqi Ji, Haoqian Wang, Lu Fang 0001 |
ICIP | 5 |
| 2019 | iDFusion: Globally Consistent Dense 3D Reconstruction from RGB-D and Inertial MeasurementsabstractWe present a practical fast, globally consistent and robust dense 3D reconstruction system, iDFusion, by exploring the joint benefit of both the visual (RGB-D) solution and inertial measurement unit (IMU). A global optimization considering all the previous states is adopted to maintain high localization accuracy and global consistency, yet its complexity of being linear to the number of all previous camera/IMU observations seriously impedes real-time implementation. We show that the global optimization can be solved efficiently at the complexity linear to the number of keyframes, and further realize a real-time dense 3D reconstruction system given the estimated camera states. Meanwhile, for the sake of robustness, we propose a novel loop-validity detector based on the estimated bias of the IMU state. By checking the consistency of camera movements, a false loop closure constraint introduces manifest inconsistency between the camera movements and IMU measurements. Experiments reveal that iDFusion owns superior reconstruction performance running in 25 fps on CPU computing of portable devices, under challenging yet practical scenarios including texture-less, motion blur, and repetitive contents. Dawei Zhong, Lu Fang 0001 |
ACM Multimedia | 3 |
| 2019 | Light Field Reconstruction Using Convolutional Network on EPI and Extended ApplicationsabstractIn this paper, a novel convolutional neural network (CNN)-based framework is developed for light field reconstruction from a sparse set of views. We indicate that the reconstruction can be efficiently modeled as angular restoration on an epipolar plane image (EPI). The main problem in direct reconstruction on the EPI involves an information asymmetry between the spatial and angular dimensions, where the detailed portion in the angular dimensions is damaged by undersampling. Directly upsampling or super-resolving the light field in the angular dimensions causes ghosting effects. To suppress these ghosting effects, we contribute a novel "blur-restoration-deblur" framework. First, the "blur" step is applied to extract the low-frequency components of the light field in the spatial dimensions by convolving each EPI slice with a selected blur kernel. Then, the "restoration" step is implemented by a CNN, which is trained to restore the angular details of the EPI. Finally, we use a non-blind "deblur" operation to recover the spatial high frequencies suppressed by the EPI blur. We evaluate our approach on several datasets, including synthetic scenes, real-world scenes and challenging microscope light field data. We demonstrate the high performance and robustness of the proposed framework compared with state-of-the-art algorithms. We further show extended applications, including depth enhancement and interpolation for unstructured input. More importantly, a novel rendering approach is presented by combining the proposed framework and depth information to handle large disparities. Gaochang Wu, Yebin Liu, Lu Fang 0001, Qionghai Dai, Tianyou Chai |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2019 | Real-Time Global Registration for Globally Consistent RGB-D SLAMabstractReal-time globally consistent camera localization is critical for visual simultaneous localization and mapping (SLAM) applications. Regardless the popularity of high efficient pose graph optimization as a backend in SLAM, its deficiency in accuracy can hardly benefit the reconstruction application. An alternative solution for the sake of high accuracy would be global registration, which minimizes the alignment error of all the corresponding observations, yet suffers from high complexity due to the tremendous observations that need to be considered. In this paper, we start by analyzing the complexity bottleneck of global point cloud registration problem, i.e., each observation (three-dimensional point feature) has to be linearized based on its local coordinate (camera poses), which however is nonlinear and dynamically changing, resulting in extensive computation during optimization. We further prove that such nonlinearity can be decoupled into linear component (feature position) and nonlinear components (camera poses), where the former linear one can be effectively represented by its compact second-order statistics, while the latter nonlinear one merely requires six degrees of freedom for each camera pose. Benefiting from the decoupled representation, the complexity can be significantly reduced without sacrifice in accuracy. Experiments show that the proposed algorithm achieves globally consistent pose estimation in real-time via CPU computing, and owns comparable accuracy as state-of-the-art that use GPU computing, enabling the practical usage of globally consistent RGB-D SLAM on highly computationally constrained devices. Lan Xu 0003, Dmytro Bobkov, Eckehard G. Steinbach, Lu Fang 0001 |
IEEE Trans. Robotics | 5 |
| 2018 | CrossNet: An End-to-End Reference-Based Super Resolution Network Using Cross-Scale Warping
Haitian Zheng, Mengqi Ji, Haoqian Wang, Yebin Liu, Lu Fang 0001 |
ECCV (6) | 5 |
| 2018 | HybridFusion: Real-Time Performance Capture Using a Single Depth Sensor and Sparse IMUs
Zerong Zheng, Tao Yu 0007, Hao Li 0015, Qionghai Dai, Lu Fang 0001, Yebin Liu |
ECCV (9) | 6 |
| 2018 | A Natural Shape-Preserving Stereoscopic Image StitchingabstractThis paper presents a method for stereoscopic image stitching, which can make stereoscopic images look as natural as possible. Our method combines a constrained projective warp and a shape-preserving warp to reduce the projective distortion and the vertical disparity of the stitched image. In addition to provide a good alignment accuracy and maintain the consistency of input stereoscopic images, we add a specific restriction into the projective warp, which establishes the connection between target left and right images. To optimize the whole warp, a energy term is designed. It can constrain the shape of straight line and vertical disparity. Experimental results on a variety of stereoscopic images can ensure the efficiency of the proposed method. Haoqian Wang, YaZing Zhou, Xingzheng Wang, Lu Fang 0001 |
ICASSP | 4 |
| 2018 | Magnify-Net for Multi-Person 2D Pose EstimationabstractWe propose a novel method for multi-person 2D pose estimation. Our model zooms in the image gradually, which we refer to as the Magnify-Net, to solve the bottleneck problem of mean average precision (mAP) versus pixel error. Moreover, we squeeze the network efficiently by an inspired design that increases the mAP while saving the processing time. It is a simple, yet robust, bottom-up approach consisting of one stage. The architecture is designed to detect the part position and their association jointly via two branches of the same sequential prediction process, resulting in a remarkable performance and efficiency rise. Our method outcompetes the previous state-of-the-art results on the challenging COCO key-points task and MPII Multi-Person Dataset. Haoqian Wang, W. P. An, Xingzheng Wang, Lu Fang 0001, Jiahui Yuan |
ICME | 4 |
| 2018 | iHuman3D: Intelligent Human Body 3D Reconstruction using a Single Flying CameraabstractAiming at autonomous, adaptive and real-time human body reconstruction technique, this paper presents iHuman3D: an intelligent human body 3D reconstruction system using a single aerial robot integrated with an RGB-D camera. Specifically, we propose a real-time and active view planning strategy based on a highly efficient ray casting algorithm in GPU and a novel information gain formulation directly in TSDF. We also propose the human body reconstruction module by revising the traditional volumetric fusion pipeline with a compactly-designed non-rigid deformation for slight motion of the human target. We unify both the active view planning and human body reconstruction in the same TSDF volume-based representation. Quantitative and qualitative experiments are conducted to validate that the proposed iHuman3D system effectively removes the constraint of extra manual labor, enabling real-time and autonomous reconstruction of human body. Lan Xu 0003, Yuanfang Guo, Lu Fang 0001 |
ACM Multimedia | 5 |
| 2018 | Computation and Memory Efficient Image SegmentationabstractIn this paper, we address the segmentation problem under limited computation and memory resources. Given a segmentation algorithm, we propose a framework that can reduce its computation time and memory requirementsimultaneously, while preserving its accuracy. The proposed framework uses standard pixel-domain downsampling and includes two main steps.Coarse segmentationis first performed on the downsampled image.Refinementis then applied to the coarse segmentation results. We make two novel contributions to enable competitive accuracy using this simple framework. First, we rigorously examine the effect of downsampling on segmentation using a signal processing analysis. The analysis helps to determine theuncertain regions, which are small image regions where pixel labels are uncertain after the coarse segmentation. Second, we propose an efficient minimum spanning tree-based algorithm to propagate the labels into the uncertain regions. We perform extensive experiments using several standard data sets. The experimental results show that our segmentation accuracy is comparable to state-of-the-art methods, while requiring much less computation time and memory than those methods. Yiren Zhou, Thanh-Toan Do, Haitian Zheng, Ngai-Man Cheung, Lu Fang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2018 | Halftone Image Watermarking by Content Aware Double-Sided Embedding Error DiffusionabstractIn this paper, we carry out a performance analysis from a probabilistic perspective to introduce the error diffusion-based halftone visual watermarking (EDHVW) methods' expected performances and limitations. Then, we propose a new general EDHVW method, content aware double-sided embedding error diffusion (CaDEED), via considering the expected watermark decoding performance with specific content of the cover images and watermark, different noise tolerance abilities of various cover image content, and the different importance levels of every pixel (when being perceived) in the secret pattern (watermark). To demonstrate the effectiveness of CaDEED, we propose CaDEED with expectation constraint (CaDEED-EC) and CaDEED-noise visibility function (NVF) and importance factor (IF) (CaDEED-N&I). Specifically, we build CaDEED-EC by only considering the expected performances of specific cover images and watermark. By adopting the NVF and proposing the IF to assign weights to every embedding location and watermark pixel, respectively, we build the specific method CaDEED-N&I. In the experiments, we select the optimal parameters for NVF and IF via extensive experiments. In both the numerical and visual comparisons, the experimental results demonstrate the superiority of our proposed work. Yuanfang Guo, Oscar C. Au, Rui Wang 0032, Lu Fang 0001, Xiaochun Cao |
IEEE Trans. Image Process. | 4 |
| 2018 | FlyCap: Markerless Motion Capture Using Multiple Autonomous Flying CamerasabstractAiming at automatic, convenient and non-instrusive motion capture, this paper presents a new generation markerless motion capture technique, the FlyCap system, to capture surface motions of moving characters using multiple autonomous flying cameras (autonomous unmanned aerial vehicles(UAVs) each integrated with an RGBD video camera). During data capture, three cooperative flying cameras automatically track and follow the moving target who performs large-scale motions in a wide space. We propose a novel non-rigid surface registration method to track and fuse the depth of the three flying cameras for surface motion tracking of the moving target, and simultaneously calculate the pose of each flying camera. We leverage the using of visual-odometry information provided by the UAV platform, and formulate the surface tracking problem in a non-linear objective function that can be linearized and effectively minimized through a Gaussian-Newton method. Quantitative and qualitative experimental results demonstrate the plausible surface and motion reconstruction results. Lan Xu 0003, Yebin Liu, Guyue Zhou, Qionghai Dai, Lu Fang 0001 |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2017 | Learning Cross-scale Correspondence and Patch-based Synthesis for Reference-based Super-Resolution
Haitian Zheng, Mengqi Ji, Ziwei Xu 0001, Haoqian Wang, Yebin Liu, Lu Fang 0001 |
BMVC | 7 |
| 2017 | Multiscale gigapixel video: A cross resolution image matching and warping approachabstractWe present a multi-scale camera array to capture and synthesize gigapixel videos in an efficient way. Our acquisition setup contains a reference camera with a short-focus lens to get a large field-of-view video and a number of unstructured long-focus cameras to capture local-view details. Based on this new design, we propose an iterative feature matching and image warping method to independently warp each local-view video to the reference video. The key feature of the proposed algorithm is its robustness to and high accuracy for the huge resolution gap (more than 8x resolution gap between the reference and the local-view videos), camera parallaxes, complex scene appearances and color inconsistency among cameras. Experimental results show that the proposed multi-scale camera array and cross resolution video warping scheme is capable of generating seamless gigapixel video without the need of camera calibration and large overlapping area constraints between the local-view cameras. Xiaoyun Yuan, Lu Fang 0001, Qionghai Dai, David J. Brady, Yebin Liu |
ICCP | 2 |
| 2017 | SurfaceNet: An End-to-End 3D Neural Network for Multiview StereopsisabstractThis paper proposes an end-to-end learning framework for multiview stereopsis. We term the network SurfaceNet. It takes a set of images and their corresponding camera parameters as input and directly infers the 3D model. The key advantage of the framework is that both photo-consistency as well geometric relations of the surface structure can be directly learned for the purpose of multiview stereopsis in an end-to-end fashion. SurfaceNet is a fully 3D convolutional network which is achieved by encoding the camera parameters together with the images in a 3D voxel representation. We evaluate SurfaceNet on the large-scale DTU benchmark. Mengqi Ji, Juergen Gall, Haitian Zheng, Yebin Liu, Lu Fang 0001 |
ICCV | 5 |
| 2017 | MILD: Multi-index hashing for appearance based loop closure detectionabstractLoop Closure Detection (LCD) has been proved to be extremely useful in global consistent visual Simultaneously Localization and Mapping (SLAM) and appearance-based robot relocalization. Methods exploiting binary features in bag of words representation have recently gained a lot of popularity for their efficiency, but suffer from low recall due to the inherent drawback that high dimensional binary feature descriptors lack well-defined centroids. In this paper, we propose a realtime LCD approach called MILD (Multi-Index Hashing for Loop closure Detection), in which image similarity is measured by feature matching directly to achieve high recall without introducing extra computational complexity with the aid of Multi-Index Hashing (MIH). A theoretical analysis of the approximate image similarity measurement using MIH is presented, which reveals the trade-off between efficiency and accuracy from a probabilistic perspective. Extensive comparisons with state-of-the-art LCD methods demonstrate the superiority of MILD in both efficiency and accuracy. Lu Fang 0001 |
ICME | 2 |
| 2017 | Beyond SIFT using binary features in Loop Closure DetectionabstractIn this paper a binary feature based Loop Closure Detection (LCD) method is proposed, which for the first time achieves higher precision-recall (PR) performance compared with state-of-the-art SIFT feature based approaches. The proposed system originates from our previous work Multi-Index hashing for Loop closure Detection (MILD), which employs Multi-Index Hashing (MIH) [1] for Approximate Nearest Neighbor (ANN) search of binary features. As the accuracy of MILD is limited by repeating textures and inaccurate image similarity measurement, burstiness handling is introduced to solve this problem and achieves considerable accuracy improvement. Additionally, a comprehensive theoretical analysis on MIH used in MILD is conducted to further explore the potentials of hashing methods for ANN search of binary features from probabilistic perspective. This analysis provides more freedom on best parameter choosing in MIH for different application scenarios. Experiments on popular public datasets show that the proposed approach achieved the highest accuracy compared with state-of-the-art while running at 30Hz for databases containing thousands of images. Guyue Zhou, Lan Xu 0003, Lu Fang 0001 |
IROS | 4 |
| 2017 | Magic Glasses: From 2D to 3DabstractThis paper proposes a virtual 3D eyeglasses try-on system driven by a 2D Internet image of a human face wearing with a pair of eyeglasses. The main technical challenge of this system is the automatic 3D eyeglasses model reconstruction from the 2D glasses on a frontal human face. Against this challenge, this paper first proposes an eyeglasses segmentation method using a convolutional neural network-based parsing algorithm to label the glasses pixels, followed with a proposed symmetry-based level-set optimization algorithm to refine the contour of the eyeglasses. With the precisely extracted silhouette image, we take advantages of the smoothness and the symmetry priors of the eyeglasses, and propose a silhouette-based 3D deformation method to deform a 3D eyeglasses model selected from a predefined 3D model database. The obtained model is plausibly approximated to the input 2D eyeglasses after a texture mapping step. Finally, we develop a virtual try-on system to interactively synthesize the reconstructed eyeglasses on a moving target face in real time. The experimental results demonstrate the efficient and convincing virtual try-on performance of our approach and the commercial potential of our proposed system. Xiaoyun Yuan, Difei Tang, Yebin Liu, Lu Fang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2017 | Adaptive Multispectral Demosaicking Based on Frequency-Domain Analysis of Spectral CorrelationabstractColor filter array (CFA) interpolation, or three-band demosaicking, is a process of interpolating the missing color samples in each band to reconstruct a full color image. In this paper, we are concerned with the challenging problem of multispectral demosaicking, where each band is significantly undersampled due to the increment in the number of bands. Specifically, we demonstrate a frequency-domain analysis of the subsampled color-difference signal and observe that the conventional assumption of highly correlated spectral bands for estimating undersampled components is not precise. Instead, such a spectral correlation assumption is image dependent and rests on the aliasing interferences among the various color-difference spectra. To address this problem, we propose an adaptive spectral-correlation-based demosaicking (ASCD) algorithm that uses a novel anti-aliasing filter to suppress these interferences, and we then integrate it with an intra-prediction scheme to generate a more accurate prediction for the reconstructed image. Our ASCD is computationally very simple, and exploits the spectral correlation property much more effectively than the existing algorithms. Experimental results conducted on two data sets for multispectral demosaicking and one data set for CFA demosaicking demonstrate that the proposed ASCD outperforms the state-of-the-art algorithms. Sunil Prasad Jaiswal, Lu Fang 0001, Vinit Jakhetiya, Jiahao Pang, Klaus Mueller 0001, Oscar C. Au |
IEEE Trans. Image Process. | 2 |
| 2016 | Optimized high-frequency based interpolation for multispectral demosaickingabstractMultispectral demosaicking, which is an extension of color demosaicking, is a challenging problem because each band is significantly undersampled and thus precise reconstruction is needed for the restoration of high-frequency components, such as edges, textures etc. In general, existing algorithms borrow high-frequency information either from different bands via inter-color correlation or from within the bands, and produces artifact in the reconstructed image. To meet this inherent shortcoming, we propose to incorporate two different high-frequency components and integrate them optimally in the linear minimum mean square sense (LMMSE) for the precise reconstruction of undersampled components. Experimental results demonstrate that the proposed algorithm based on the optimized high-frequency achieves superior performance compared to existing algorithms both in terms of objective and subjective quality. Sunil Prasad Jaiswal, Lu Fang 0001, Vinit Jakhetiya, Manohar Kuse, Oscar C. Au |
ICIP | 2 |
| 2016 | Estimation of Virtual View Synthesis Distortion Toward Virtual View PositionabstractWe propose an analytical model to estimate the depth-error-induced virtual view synthesis distortion (VVSD) in 3D video, taking the distance between reference and virtual views (virtual view position) into account. In particular, we start with a comprehensive preanalysis and discussion over several possible VVSD scenarios. Taking intrinsic characteristic of each scenario into consideration, we specifically classify them into four clusters: 1) overlapping region; 2) disocclusion and boundary region; 3) edge region; and 4) infrequent region. We propose to model VVSD as the linear combination of the distortion under different scenarios (DDSs) weighted by the probability under different scenarios (PDSs). We show analytically that DDS and PDS can be related to the virtual view position using quadratic/biquadratic models and linear models, respectively. Experimental results verify that the proposed model is capable of estimating the relationship between VVSD and the distance between reference and virtual views. Therefore, our model can be used to inform a reference view setup for capturing, or distortion at certain virtual view positions, when depth information is compressed. Lu Fang 0001, Yijian Xiang, Ngai-Man Cheung, Feng Wu 0001 |
IEEE Trans. Image Process. | 1 |
| 2016 | Robust Blur Kernel Estimation for License Plate Images From Fast Moving VehiclesabstractAs the unique identification of a vehicle, license plate is a key clue to uncover over-speed vehicles or the ones involved in hit-and-run accidents. However, the snapshot of over-speed vehicle captured by surveillance camera is frequently blurred due to fast motion, which is even unrecognizable by human. Those observed plate images are usually in low resolution and suffer severe loss of edge information, which cast great challenge to existing blind deblurring methods. For license plate image blurring caused by fast motion, the blur kernel can be viewed as linear uniform convolution and parametrically modeled with angle and length. In this paper, we propose a novel scheme based on sparse representation to identify the blur kernel. By analyzing the sparse representation coefficients of the recovered image, we determine the angle of the kernel based on the observation that the recovered image has the most sparse representation when the kernel angle corresponds to the genuine motion angle. Then, we estimate the length of the motion kernel with Radon transform in Fourier domain. Our scheme can well handle large motion blur even when the license plate is unrecognizable by human. We evaluate our approach on real-world images and compare with several popular state-of-the-art blind image deblurring algorithms. Experimental results demonstrate the superiority of our proposed approach in terms of effectiveness and robustness. Qingbo Lu, Wengang Zhou 0001, Lu Fang 0001, Houqiang Li |
IEEE Trans. Image Process. | 3 |
| 2016 | Subpixel-Based Image Scaling for Grid-like Subpixel Arrangements: A Generalized Continuous-Domain Analysis ModelabstractSubpixel-based image scaling can improve the apparent resolution of displayed images by controlling individual subpixels rather than whole pixels. However, improved luminance resolution brings chrominance distortion, making it crucial to suppress color error while maintaining sharpness. Moreover, it is challenging to develop a scheme that is applicable for various subpixel arrangements and for arbitrary scaling factors. In this paper, we address the aforementioned issues by proposing a generalized continuous-domain analysis model, which considers the low-pass nature of the human visual system (HVS). Specifically, given a discrete image and a grid-like subpixel arrangement, the signal perceived by the HVS is modeled as a 2D continuous image. Minimizing the difference between the perceived image and the continuous target image leads to the proposed scheme, which we call continuous-domain analysis for subpixel-based scaling (CASS). To eliminate the ringing artifacts caused by the ideal low-pass filtering in CASS, we propose an improved scheme, which we call CASS with Laplacian-of-Gaussian filtering. Experiments show that the proposed methods provide sharp images with negligible color fringing artifacts. Our methods are comparable with the state-of-the-art methods when applied on the RGB stripe arrangement, and outperform existing methods when applied on other subpixel arrangements. Jiahao Pang, Lu Fang 0001, Jin Zeng 0004, Yuanfang Guo, Ketan Tang |
IEEE Trans. Image Process. | 2 |
| 2016 | Subpixel Image Quality Assessment Syncretizing Local Subpixel and Global Pixel FeaturesabstractThe subpixel rendering technology increases the apparent resolution of an LCD/OLED screen by exploiting the physical property that a pixel is composed of RGB individually addressable subpixels. Due to the intrinsic intercoordination between apparent luminance resolution and color fringing artifact, a common method of subpixel image assessment is subjective evaluation. In this paper, we propose a unified subpixel image quality assessment metric called subpixel image assessment (SPA), which syncretizes local subpixel and global pixel features. Specifically, comprehensive subjective studies are conducted to acquire data of user preferences. Accordingly, a collection of low-level features is designed under extensive perceptual validation, capturing subpixel and pixel features, which reflect local details and global distance from the original image. With the features and their measurements as the basis, the SPA is obtained, which leads to a good representation of the subpixel image characteristics. The experimental results justify the effectiveness and the superiority of the SPA. The SPA is also successfully adopted in a variety of applications, including content adaptive sampling and metric-guided image compression. Jin Zeng 0004, Lu Fang 0001, Jiahao Pang, Houqiang Li, Feng Wu 0001 |
IEEE Trans. Image Process. | 2 |
| 2016 | Deep Learning for Surface Material Classification Using Haptic and Visual InformationabstractWhen a user scratches a hand-held rigid tool across an object surface, an acceleration signal can be captured, which carries relevant information about the surface material properties. More importantly, such haptic acceleration signals can be used together with surface images to jointly recognize the surface material. In this paper, we present a novel deep learning method dealing with the surface material classification problem based on a fully convolutional network, which takes the aforementioned acceleration signal and a corresponding image of the surface texture as inputs. Compared to the existing surface material classification solutions which rely on a careful design of hand-crafted features, our method automatically extracts discriminative features utilizing advanced deep learning methodologies. Experiments performed on the TUM surface material database demonstrate that our method achieves state-of-the-art classification accuracy robustly and efficiently. Haitian Zheng, Lu Fang 0001, Mengqi Ji, Matti Strese, Yigitcan Özer, Eckehard G. Steinbach |
IEEE Trans. Multim. | 2 |
| 2015 | Depth Error Induced Virtual View Synthesis Distortion Estimation for 3D Video CodingabstractWe propose an analytical model to estimate the depth-error-induced virtual view synthesis distortion (VVSD) in 3D video, taking into account the configuration of the cameras. Focusing on view synthesis under depth error, we carefully analyze the merging operations under different situations that affect pixel availability: overlapping region, disocclusion and boundary region, disparity error region, and infrequent region. The analysis leads to quadratic/biquadratic models and linear models that explicitly relate the distance between camera positions (reference/virtual view) to Distortion under Different Situations (DDS) and Probability under Different Situations (PDS), respectively. We also show that VVSD is the linear combination of DDS weighted by PDS. Our careful analysis results in state-of-the-art estimation accuracy. Experimental results verify that the proposed model is capable to produce accurate estimates of VVSD based on the distance between reference/virtual views. Therefore, our model can effectively inform camera setup for capturing, in particular, the set-up of the cameras in situation where depth information will be compressed subsequently. Yijian Xiang, Lu Fang 0001, Ngai-Man Cheung |
DCC | 2 |
| 2015 | Stereo Matching with Optimal Local Adaptive Radiometric CompensationabstractA common assumption in stereo matching is that the corresponding pixels in stereo images have similar pixel values. Unfortunately, such an assumption may not be true due to radiometric variations in different views, leading to severely degraded matching results. In this letter, we propose a radiometrically invariant stereo matching algorithm called Optimal Local Adaptive Radiometric Compensation (LARAC). In LARAC, we approximate the spatially varying Pixel Value Correspondence Function (PVCF) between a corresponding pixel pair as a locally consistent polynomial within an optimal local adaptive window. The optimal polynomial coefficients are obtained for each candidate disparity value and are used to compute the matching cost. Meanwhile, a self-correction property is achieved by the proposed LARAC, leading to reduced matching errors for the outlier pixels. Experimental results suggest that the proposed LARAC outperforms other state-of-the-art stereo matching algorithms. Lingfeng Xu, Oscar C. Au, Wenxiu Sun, Lu Fang 0001, Feng Zou 0006 |
IEEE Signal Process. Lett. | 4 |
| 2015 | Deblurring Saturated Night Image With Function-Form KernelabstractDeblurring saturated night images are a challenging problem because such images have low contrast combined with heavy noise and saturated regions. Unlike the deblurring schemes that discard saturated regions when estimating blur kernels, this paper proposes a novel scheme to deduce blur kernels from saturated regions via a novel kernel representation and advanced algorithms. Our key technical contribution is the proposed function-form representation of blur kernels, which regularizes existing matrix-form kernels using three functional components: 1) trajectory; 2) intensity; and 3) expansion. From automatically detected saturated regions, their skeleton, brightness, and width are fitted into the corresponding three functional components of blur kernels. Such regularization significantly improves the quality of kernels deduced from saturated regions. Second, we propose an energy minimizing algorithm to select and assign the deduced function-form kernels to partitioned image regions as the initialization for non-uniform deblurring. Finally, we convert the assigned function-form kernels into matrix form for more detailed estimation in a multi-scale deconvolution. Experimental results show that our scheme outperforms existing schemes on challenging real examples. Xiaoyan Sun 0001, Lu Fang 0001, Feng Wu 0001 |
IEEE Trans. Image Process. | 3 |
| 2014 | Separable Kernel for Image DeblurringabstractIn this paper, we deal with the image deblurring problem in a completely new perspective by proposing separable kernel to represent the inherent properties of the camera and scene system. Specifically, we decompose a blur kernel into three individual descriptors (trajectory, intensity and point spread function) so that they can be optimized separately. To demonstrate the advantages, we extract one-pixel-width trajectories of blur kernels and propose a random perturbation algorithm to optimize them but still keeping their continuity. For many cases, where current deblurring approaches fall into local minimum, excellent deblurred results and correct blur kernels can be obtained by individually optimizing the kernel trajectories. Our work strongly suggests that more constraints and priors should be introduced to blur kernels in solving the deblurring problem because blur kernels have lower dimensions than images. Lu Fang 0001, Feng Wu 0001, Xiaoyan Sun 0001, Houqiang Li |
CVPR | 1 |
| 2014 | Joint Denoising and demosaicking of noisy CFA images based on inter-color correlationabstractMost digital cameras use a single sensor coupled with a Color Filter Array (CFA) to capture images, and apply demosaicking to interpolate the full color images. In reality, the CFA image is noisy, which causes problems in the demosaicking process. This paper proposes a Joint Denoising and Demo-saicking based on inter-Color correlation (JDDC) scheme. We propose a new framework that linearly combines an extracted luminance image and a low-passed RGB images to get a full color image. Given the noise in the extracted luminance image and the low-passed RGB images are non-stationary and partially correlated, we modify the classical Non-Local Means (NLM) filter to denoise the extracted luminance image and the low-passed RGB images before the combination. Experimental results verify the effectiveness of the proposed scheme both objectively and subjectively. Ming-Ting Sun, Lu Fang 0001, Oscar C. Au |
ICASSP | 3 |
| 2014 | Analytical model for camera distance related 3D virtual view distortion estimationabstractWe propose an analytical model to estimate the depth-error-induced synthesis distortion in 3D video, taking into account the configuration of the cameras. In particular, the model mathematically relates the Distance between camera positions (reference view and virtual view) to the Virtual View Distortion (VVD), thus it is denoted as DVVD model. Specifically, the DVVD model accounts for two modules: distribution of disparity errors and shift-induced distortion. The former one is derived under a Laplacian distribution assumption of depth errors, and the latter one is estimated under a Quadratic model. We further propose a linear Steady-State model by performing Taylor series approximation of the DVVD model over a region of practical interest. Experiment results demonstrate that both the DVVD and Steady-State models are capable of estimating the relationship between VVD and the distance between virtual/reference view. Therefore, our model can effectively inform camera setup for capturing, in particular, the setup of the cameras in situation where depth information will be compressed subsequently. Yijian Xiang, Ngai-Man Cheung, Juyong Zhang, Lu Fang 0001 |
ICIP | 4 |
| 2014 | Fast algorithm of arbitrary factor subpixel downsampling based on frequency analysisabstractSubpixel-based downsampling has shown its advantages over pixel-based downsampling in terms of preserving more spatial details along edges and generating sharper images, at the cost of certain amount of color-fringing artifacts in the downsampled image. To balance the sharpness and color-fringing artifacts, some algorithms are proposed to design optimal anti-aliasing (AA) filters, which are either image independent, or computationally too expensive. And all of the existing AA filters are designed for fixed downsampling factor, which makes them impractical for real applications. In this paper we propose two fast algorithms to design AA filter for arbitrary factor subpixel downsampling based on frequency analysis of the input image. The proposed algorithms generate image dependent AA filter which is as good as the state-of-the-art algorithm, but much faster. Ketan Tang, Oscar C. Au, Lu Fang 0001, Jiahao Pang, Yuanfang Guo |
ICME | 3 |
| 2014 | A low latency cloud gaming system using edge preserved image homographyabstractThe emerging cloud gaming technology has been growing fast, driving up huge mobile consumer demands. The video streaming based cloud gaming scenario renders the game scenes in the cloud servers, and streams the encoded sequences to the thin clints where the game scenes are decoded and displayed to the players. However, current existing clouding gaming services have some problems, such as the latency and bandwidth limitation. The size of the video stream is usually quite large which requires heavy transmission. Worse still, the frame data rate will burst when the game scenes contain fast translation or rotation, resulting in strong latency problem. In this paper, we propose a novel video streaming based cloud gaming algorithm which reduces the burst of the frame rate significantly. There are mainly two innovations in this paper. Firstly, based on the analysis of the motion estimation strategy in the video codec, we introduce image homography technique for better motion prediction. Meanwhile, according to the rasterization rules of the game engine, we present a special designed interpolation algorithm named Edge Preserved Interpolation (EPI), for more accurate edge interpolation and further reduce the residues in the edge regions. The proposed algorithm is implemented on the x264 platform. Experimental results show that our algorithm has 18.0% BD-rate reduction compared with x264. Lingfeng Xu, Xun Guo 0002, Yan Lu 0001, Shipeng Li 0001, Oscar C. Au, Lu Fang 0001 |
ICME | 6 |
| 2014 | An Analytical Model for Synthesis Distortion Estimation in 3D VideoabstractWe propose an analytical model to estimate the synthesized view quality in 3D video. The model relates errors in the depth images to the synthesis quality, taking into account texture image characteristics, texture image quality, and the rendering process. Especially, we decompose the synthesis distortion into texture-error induced distortion and depth-error induced distortion. We analyze the depth-error induced distortion using an approach combining frequency and spatial domain techniques. Experiment results with video sequences and coding/rendering tools used in MPEG 3DV activities show that our analytical model can accurately estimate the synthesis noise power. Thus, the model can be used to estimate the rendering quality for different system designs. Lu Fang 0001, Ngai-Man Cheung, Dong Tian, Anthony Vetro, Huifang Sun, Oscar C. Au |
IEEE Trans. Image Process. | 1 |
| 2013 | Low Bit-Rate Subpixel-Based Color Image CompressionabstractWe propose a novel low bit-rate compression scheme with sub pixel-based down-sampling and reconstruction (SPDR) for full color images. In the encoder stage, a decoder-dependent multi-channel sub pixel-based down-sampling is proposed, which is more effective in retaining high frequency detail than conventional pixel-based process. The decoder first decompresses the low-resolution image and then up-converts it to the original resolution using encoder dependent sub pixel-based reconstruction scheme by jointly considering the sub pixel-based down-sampling effect and the compression degradation. Compared to existing algorithms with comparable encoder and decoder complexity, the proposed SPDR offers complete standard compliance, competitive rate-distortion performance, and superior subjective quality. Lu Fang 0001, Ngai-Man Cheung, Oscar C. Au, Houqiang Li, Ketan Tang |
DCC | 1 |
| 2013 | Synthesis distortion estimation in 3D video using frequency and spatial analysisabstractWe propose an analytical model to estimate the synthesized view quality in 3D video. Specifically, we estimate the depth-error induced distortion using an approach that combines frequency and spatial domain analysis. We also propose to decompose the spatial-variant video signals into gradient-based representations to capture the interaction between image gradients, depth errors and synthesis distortion. Experiment results with video sequences and coding/rendering tools used in MPEG 3DV activities show that our analytical model can accurately estimate the synthesis noise power. Lu Fang 0001, Ngai-Man Cheung, Dong Tian, Anthony Vetro, Huifang Sun, Lu Yu 0003 |
ICIP | 1 |
| 2013 | Arbitrary factor image interpolation by convolution kernel constrained 2-D autoregressive modelingabstractAmong existing interpolation methods, convolution-based methods are able to perform arbitrary factor interpolation but the results are usually blurry or jaggy, adaptive interpolation methods usually can reduce the blurry and jaggy artifacts but cannot handle arbitrary factor interpolation. In this paper we propose an arbitrary factor adaptive interpolation algorithm by combining 2-D piecewise autoregressive (PAR) modeling and convolution kernel constraint. PAR model ensures local geometries are well preserved thus the resultant image is not blurry or jaggy. Convolution kernel constraint ensures the recovered high resolution image consistent with the low resolution image, and also provides the flexibility to handle arbitrary interpolation factor. Experiment results show that our algorithm achieves state-of-the-art performance for any interpolation factor. Ketan Tang, Oscar C. Au, Yuanfang Guo, Jiahao Pang, Lu Fang 0001 |
ICIP | 6 |
| 2013 | A parallel deblocking filter based on H.264/AVC video coding standardabstractThe deblocking filter in H.264/AVC is one of the most time consuming part of video decoder as its high content adaptation and data dependency lead to lots of computation. In this paper, we propose a novel parallel deblocking filter design based on the H.264/AVC video coding standard, taking the advantage that the data dependency of the deblocking filter are “periodic” in one dimension. Our proposed architecture successfully reduces the dependency between horizontal and vertical filters and utilizes the “periodic” property to achieve pixel-level parallelism. Algorithm analysis and experiment results on JM and GPU show that the proposed deblocking filter keeps as good a coding efficiency as that in H.264/AVC, and its high parallelism is suitable and promising in multi-core/multi-thread computing. Oscar C. Au, Lu Fang 0001, Lin Sun 0004, Wenxiu Sun, Dinuka Soysa |
ISCAS | 3 |
| 2013 | Stereo matching by adaptive weighting selection based cost aggregationabstractCost aggregation is the most essential step for dense stereo correspondence searching, which measures the similarity between pixels in the stereo images. In this paper, based on the analysis of the optimal adaptive weight, we propose a novel support aggregation strategy by adaptive weighting selection. The proposed method calculates the aggregation cost by the joint optimization of both left and right matching cost. By assigning more reasonable weighting coefficients, we exclude the occlusion pixels while preserving sufficient support region for accurate matching. The proposed optimal strategy can be integrated by any other adaptive weighting based cost aggregation method to generate more reasonable similarity measurement. Experimental results show that, compare with traditional methods, our algorithm can reduce the foreground fatten phenomenon while increasing the accuracy in the high texture regions. Lingfeng Xu, Oscar C. Au, Wenxiu Sun, Lu Fang 0001, Ketan Tang, Yuanfang Guo |
ISCAS | 4 |
| 2013 | Chroma Replacing and adaptive Chroma Blending for subpixel-based downsamplingabstractSubpixel-based downsampling generates images with higher apparent resolution with the expense of annoying color-fringing artifacts near strong edges. In this paper we propose two methods that find a balance in the tradeoff of apparent resolution and color-fringing artifacts. The first method is called Chroma Replacing in which the color-fringing artifacts are completely removed but the subpixel rendering effect is also removed. The second one is called Chroma Blending in which only the color-fringing artifacts that are strong enough to be noticed are removed, and also the subpixel rendering effect is retained. We also propose two objective measures for measuring the similarity of downsampled image to the original image. Experiment results show that the proposed methods are effective in removing color-fringing artifacts, without harming the high apparent resolution. Ketan Tang, Oscar C. Au, Lu Fang 0001, Yuanfang Guo, Jiahao Pang |
MMSP | 3 |
| 2013 | An analytical study of subpixel-based image down-sampling patterns in frequency domainabstractSubpixel-based image down-sampling is a class of methods that can provide improved apparent resolution of the down-scaled image compared to the pixel-based methods. The frequency characteristics of all possible subpixel-based down-sampling patterns for RGB vertical stripes are analytically studied in this paper. Our proposed algorithm reveals that there are merely seven equivalent energy distributions in the luminance frequency spectrum. To achieve higher luminance resolution, we then calculate and choose the optimal down-sampling pattern with anti-aliasing low-pass filter designed for it so as to maximize the energy of the luminance component within the cut-off shape. Experimental results show that the proposed method provides sharper images compared to the state-of-art subpixel-based methods, with little color distortion. Yonggen Ling, Oscar C. Au, Ketan Tang, Jiahao Pang, Jin Zeng 0004, Lu Fang 0001 |
VCIP | 6 |
| 2013 | Multichannel Nonlocal Means Fusion for Color Image DenoisingabstractIn this paper, we propose an advanced color image denoising scheme called multichannel nonlocal means fusion (MNLF), where noise reduction is formulated as the minimization of a penalty function. An inherent feature of color images is the strong interchannel correlation, which is introduced into the penalty function as additional prior constraints to expect a better performance. The optimal solution of the minimization problem is derived, consisting of constructing and fusing multiple nonlocal means (NLM) spanning all three channels. The weights in the fusion are optimized to minimize the overall mean squared denoising error, with the help of the extended and adapted Stein's unbiased risk estimator (SURE). Simulations on representative test images under various noise levels verify the improvement brought by the multichannel NLM, compared to the traditional single-channel NLM. In the meantime, MNLF provides competitive performance both in terms of the color peak signal-to-noise ratio and in perceptual quality when compared with other state-of-the-art benchmarks. Jingjing Dai, Oscar C. Au, Lu Fang 0001, Feng Zou 0006 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2013 | Luma-Chroma Space Filter Design for Subpixel-Based Monochrome Image DownsamplingabstractIn general, subpixel-based downsampling can achieve higher apparent resolution of the down-sampled images on LCD or OLED displays than pixel-based downsampling. With the frequency domain analysis of subpixel-based downsampling, we discover special characteristics of the luma-chroma color transform choice for monochrome images. With these, we model the anti-aliasing filter design for subpixel-based monochrome image downsampling as a human visual system-based optimization problem with a two-term cost function and obtain a closed-form solution. One cost term measures the luminance distortion and the other term measures the chrominance aliasing in our chosen luma-chroma space. Simulation results suggest that the proposed method can achieve sharper down-sampled gray/font images compared with conventional pixel and subpixel-based methods, without noticeable color fringing artifacts. Lu Fang 0001, Oscar C. Au, Ngai-Man Cheung, Aggelos K. Katsaggelos, Houqiang Li, Feng Zou 0006 |
IEEE Trans. Image Process. | 1 |
| 2012 | Color image denoising based on multichannel non-local means fusionabstractIn this paper, we investigate the problem of color image denoising, and propose a novel algorithm called multichannel non-local means fusion (MNLMF), building on the grayscale denoiser non-local means filter. By analyzing and modeling the inter-channel correlation in color images, we formulate the color noise reduction as a minimization problem with a specifically-designed penalty function which fully takes advantages of the inter-channel prior information. The optimal solution is derived consisting of constructing multiple non-local means spanning all three channels and fusing them together. The weights in the fusion are optimized to minimize the overall denoising error. Simulation results under various noise levels demonstrate that when compared to other state-of-the-art algorithms, the proposed MNLMF achieves competitive performance both in terms of the color peak signal-to-noise ratio (cPSNR) and in perceptual quality. Jingjing Dai, Oscar C. Au, Feng Zou 0006, Lu Fang 0001 |
ICIP | 5 |
| 2012 | Analytical study of RGB vertical stripe and RGBX square-shaped subpixel arrangementsabstractThe frequency characteristics of subpixel-based decimation with RGB vertical stripe and RGBX square-shaped subpixel arrangements are studied. To achieve higher apparent resolution than pixel-based decimation, the sampling locations are specially chosen for each of two subpixel arrangements, resulting in relatively small magnitudes of horizontal and vertical aliasing spectra in frequency domain. Thanks to 2-D RGBX square-shaped subpixel arrangement, all the horizontal, vertical, diagonal and anti-diagonal aliasing spectra merely contain low-frequency information, indicating that subpixel-based decimation with RGBX square-shaped panel is more effective in retaining original high frequency details than RGB vertical stripe subpixel arrangement. Lu Fang 0001, Oscar C. Au, Jingjing Dai, Hanli Wang, Ngai-Man Cheung |
ICIP | 1 |
| 2012 | From 2D Extrapolation to 1D Interpolation: Content Adaptive Image Bit-Depth ExpansionabstractIn this paper, we address the problem of image bit-depth expansion and present a novel method to generate high bit-depth (HBD) images from a single low bit-depth (LBD) image. We expand image bit-depth by reconstructing the least significant bits (LSBs) for the LBD image after it is rescaled to high bit-depth. For image regions whose intensities are neither locally maximum nor minimum, neighborhood flooding is applied to convert 2D interpolation problem into 1D interpolation, for local maxima/minima (LMM) regions where interpolation is not applicable, a virtual skeleton marking algorithm is proposed to convert problematic 2D extrapolation problem into 1D interpolation. At last, a content-adaptive reconstruction model is proposed to obtain the output HBD image. The experimental results show that proposed method significantly outperforms existing methods in PSNR and SSIM without contouring artifacts. Pengfei Wan 0001, Oscar C. Au, Ketan Tang, Yuanfang Guo, Lu Fang 0001 |
ICME | 5 |
| 2012 | Novel 2-D MMSE Subpixel-Based Image Down-SamplingabstractSubpixel-based down-sampling is a method that can potentially improve apparent resolution of a down-scaled image on LCD by controlling individual subpixels rather than pixels. However, the increased luminance resolution comes at price of chrominance distortion. A major challenge is to suppress color fringing artifacts while maintaining sharpness. We propose a new subpixel-based down-sampling pattern called diagonal direct subpixel-based down-sampling (DDSD) for which we design a 2-D image reconstruction model. Then, we formulate subpixel-based down-sampling as a MMSE problem and derive the optimal solution called minimum mean square error for subpixel-based down-sampling (MMSE-SD). Unfortunately, straightforward implementation of MMSE-SD is computational intensive. We thus prove that the solution is equivalent to a 2-D linear filter followed by DDSD, which is much simpler. We further reduce computational complexity using a smallk×kfilter to approximate the much larger MMSE-SD filter. To compare the performances of pixel and subpixel-based down-sampling methods, we propose two novel objective measures: normalizedl1high frequency energy for apparent luminance sharpness and PSNRU(V)for chrominance distortion. Simulation results show that both MMSE-SD and MMSE-SD(k) can give sharper images compared with conventional down-sampling methods, with little color fringing artifacts. Lu Fang 0001, Oscar C. Au, Ketan Tang, Hanli Wang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2012 | Antialiasing Filter Design for Subpixel Downsampling via Frequency-Domain AnalysisabstractIn this paper, we are concerned with image downsampling using subpixel techniques to achieve superior sharpness for small liquid crystal displays (LCDs). Such a problem exists when a high-resolution image or video is to be displayed on low-resolution display terminals. Limited by the low-resolution display, we have to shrink the image. Signal-processing theory tells us that optimal decimation requires low-pass filtering with a suitable cutoff frequency, followed by downsampling. In doing so, we need to remove many useful image details causing blurring. Subpixel-based downsampling, taking advantage of the fact that each pixel on a color LCD is actually composed of individual red, green, and blue subpixel stripes, can provide apparent higher resolution. In this paper, we use frequency-domain analysis to explain what happens in subpixel-based downsampling and why it is possible to achieve a higher apparent resolution. According to our frequency-domain analysis and observation, the cutoff frequency of the low-pass filter for subpixel-based decimation can be effectively extended beyond the Nyquist frequency using a novel antialiasing filter. Applying the proposed filters to two existing subpixel downsampling schemes called direct subpixel-based downsampling (DSD) and diagonal DSD (DDSD), we obtain two improved schemes, i.e., DSD based on frequency-domain analysis (DSD-FA) and DDSD based on frequency-domain analysis (DDSD-FA). Experimental results verify that the proposed DSD-FA and DDSD-FA can provide superior results, compared with existing subpixel or pixel-based downsampling methods. Lu Fang 0001, Oscar C. Au, Ketan Tang, Aggelos K. Katsaggelos |
IEEE Trans. Image Process. | 1 |
| 2012 | Joint Demosaicing and Subpixel-Based Down-Sampling for Bayer Images: A Fast Frequency-Domain Analysis ApproachabstractA portable device such as a digital camera with a single sensor and Bayer color filter array (CFA) requires demosaicing to reconstruct a full color image. To display a high resolution image on a low resolution LCD screen of the portable device, it must be down-sampled. The two steps, demosaicing and down-sampling, influence each other. On one hand, the color artifacts introduced in demosaicing may be magnified when followed by down-sampling; on the other hand, the detail removed in the down-sampling cannot be recovered in the demosaicing. Therefore, it is very important to consider simultaneous demosaicing and down-sampling. Lu Fang 0001, Oscar C. Au, Yan Chen 0007, Aggelos K. Katsaggelos, Hanli Wang |
IEEE Trans. Multim. | 1 |
| 2011 | Anti-aliasing filter for subpixel down-sampling based on frequency analysisabstractNowadays, digital pictures are usually captured at very high resolution ranged up to 12 mega-pixels. Limited by low-resolution display, we have to shrink the image. Signal processing theory tells us that optimal decimation requires low-pass filtering with a suit able cut-off frequency followed by down-sampling. In doing so, we need to remove lots of details. Subpixel-based down-sampling, taking advantage of the fact that each pixel on a color LCD is actually composed of individual red, green, and blue subpixel stripes, can provide apparent higher resolution. In this paper, we use frequency domain analysis to explain what happens in subpixel-based down sampling and why it is possible to achieve a higher apparent resolution. According to our frequency domain analysis and observation, the cut-off frequency of the low-pass filter for subpixel-based decimation can be effectively extended beyond the Nyquist frequency using a novel anti-aliasing filter. Experimental results verify that the proposed subpixel down-sampling scheme based on frequency analysis (SDSFA) can give superior results compared with existing pixel-based down-sampling methods. Lu Fang 0001, Ketan Tang, Oscar C. Au, Aggelos K. Katsaggelos |
ICASSP | 1 |
| 2011 | Image Interpolation Using Autoregressive Model and Gauss-Seidel OptimizationabstractIn this paper we propose a simple yet effective image interpolation algorithm based on autoregressive model. Unlike existing algorithms which rely on low resolution pixels to estimate interpolation coefficients, we optimize the interpolation coefficients and high resolution pixel values jointly from one optimization problem. Although the two sets of variables are coupled in the cost function, the problem can be effectively solved using Gauss-Seidel method. We prove the iterations are guaranteed to converge. Experiments show that on average we have over 3dB gain compared to bicubic interpolation and over 0.1dB gain compared to SAI. Ketan Tang, Oscar C. Au, Lu Fang 0001, Zhiding Yu, Yuanfang Guo |
ICIG | 3 |
| 2011 | Multi-scale analysis of color and texture for salient object detectionabstractIn this paper we propose a multi-scale segment-based framework for salient object detection. In this framework texture and color features are used together to provide diverse information of salient object. Segmentation is performed on three different scales so that the object boundary can be accurately captured with high probability. Besides, we propose a novel adaptive feature combination mechanism to combine the saliency maps produced with different features, in which the combining weight of each saliency map is learned using online learning. Experiment results demonstrate that the proposed method significantly outperforms the state-of-the-art methods. Ketan Tang, Oscar C. Au, Lu Fang 0001, Zhiding Yu, Yuanfang Guo |
ICIP | 3 |
| 2011 | Adaptive joint demosaicing and Subpixel-based Down-sampling for Bayer imageabstractA digital camera provided with a Bayer pattern single sensor needs color interpolation to reconstruct a full color image. To show high resolution image on a lower resolution display, it must then be down-sampled. These two steps influence each other, i.e., the color artifacts introduced in demosaicing may be magnified in subsequent down-sampling process and vice versa. Thanks to the fact that LCD displays are actually composed of separable subpixels, which can be individually addressed to achieve a higher effective apparent resolution. This paper presents an Adaptive Joint Demosaicing and Subpixel-based Down-sampling scheme (AJDSD) for single-sensor camera image, where the subpixel-based down-sampling is adaptively and directly applied in Bayer domain, without the process of demosaicing. Simulation results demonstrate that when compared with conventional “demosaicing-first and down-sampling-later” methods, AJDSD achieves superior performance improvement in terms of computational complexity. As for visual quality, AJDSD is more effective in preserving high frequency details, leading to much sharper and clearer results. Lu Fang 0001, Oscar C. Au, Aggelos K. Katsaggelos |
ICME | 1 |
| 2011 | Data hiding in dot diffused halftone imagesabstractIn this paper, we propose two halftone image watermarking methods. Data Hiding by Conjugate Dot Diffusion (DHCDD) and Data Hiding by Dual Conjugate Dot Diffusion (DHD-CDD). DHDCDD is an improved method of DHCDD. Both of these two methods can embed a secret pattern into two halftone images. When the two halftone images are overlaid, the secret pattern will be revealed. Compared to the recent method Noise Balanced Dot Diffusion, the experimental results show that the proposed methods are better in both Correct Decoding Rate and the visual quality of the revealed hidden pattern. Yuanfang Guo, Oscar C. Au, Ketan Tang, Lu Fang 0001, Zhiding Yu |
ICME | 4 |
| 2011 | How anti-aliasing filter affects image contrast: An analysis from majorization theory perspectiveabstractWhen we design an anti-aliasing low pass filter, it is usually an IIR filter. We need to truncate the filter to an FIR filter. One may think that the more taps there are, the better the image quality is. However, we find that there exists an optimal value of tap number that will give the best visual quality. Filters with larger or smaller number of taps will degrade the image quality, due to the fact that the image contrast is reduced. In this paper we analyze this phenomenon using majorization theory and find that the image contrast can be formulated as a Schur convex function on filter coefficients. We also propose an effective method to choose the best filter so that the image contrast is maximized, so as to give best visual quality. Ketan Tang, Oscar C. Au, Lu Fang 0001, Zhiding Yu, Yuanfang Guo |
ICME | 3 |
| 2011 | Sub-pixel downsampling of video with matching highly data re-use hardware architectureabstractSubpixel-based down-sampling is a method that can potentially improve the apparent resolution of a down-scaled image by controlling individual subpixels rather than pixels. However, the increased luminance resolution often comes at the price of chrominance distortion. A major challenge is to suppress color fringing artifacts while maintaining sharpness. In [1], we proposed a novel human visual quality based (HVS) subpixel downsampling method. In this paper, we propose a hardware-friendly subpixel based downsampling scheme based on our previous work which can achieve similar performance as [1] but eliminate all floating point operations with limited bandwidth and much better performance than Direct Pixel based Downsampling (DPD) and Pixel-based downsampling with Anti-aliasing Filter (PDAF). We further propose a hardware architecture for our subpixel downsampling method which is highly data re-useable. We are the first few, if not the first, to implement subpixel downsampling method to video by hardware. The design is implemented with TSMC 0.18um CMOS technology and costs 244k gates. At a clock frequency of 63 MHz, the architecture achieves real-time 1920×1080 subpixel downsampling at 30fps. Oscar C. Au, Jiang Xu 0001, Lu Fang 0001, Run Cha |
ISCAS | 4 |
| 2011 | Novel RD-Optimized VBSME With Matching Highly Data Re-Usable Hardware ArchitectureabstractTo achieve superior performance, rate-distortion optimized motion estimation (ME) for variable block size (RDO VBSME) is often used in state-of-the-art video coding systems such as the H.264 JM software. However, the complexity of RDO-VBSME is very high both for software and hardware implementations. In this paper, we propose a hardware-friendly ME algorithm called RDOMFS with a novel hardware-friendly rate-distortion (RD)-like cost function, and a hardware-friendly modified motion vector predictor. Simulation results suggest that the proposed RDOMFS can achieve essentially the same RD performance as RDO-VBSME in JM. We also propose a matching hardware architecture with a novel Smart Snake Scanning order which can achieve very high data re-use ratio and data throughout. It is also reconfigurable because it can achieve variable data re-use ratio and can process variable frame size. The design is implemented with TSMC 0.18 μm CMOS technology and costs 103 k gates. At a clock frequency of 63 MHz, the architecture achieves real-time 1920 × 1080 RDO-VBSME at 30 frames/s. At a maximum clock frequency of 250 MHz, it can process 4096 × 2160 at 30 frames/s. Oscar C. Au, Jiang Xu 0001, Lu Fang 0001, Run Cha |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2010 | Novel 2-D MMSE subpixel-based image down-sampling for matrix displaysabstractSubpixel-based down-sampling is a method that can potentially improve the apparent resolution of a down-scaled image by controlling individual subpixels rather than pixels. However, the increased luminance resolution often comes at the price of chrominance distortion. In this paper, we propose a new subpixel-based down-sampling scheme for which we design a two-dimensional image reconstruction model. Then, we formulate subpixel-based down-sampling as a MMSE problem and derive the optimal solution called MMSESD. To compare the performance of subpixel-based down-sampling methods, we propose novel objective measures for the apparent luminance resolution and chrominance distortion. Simulation results show that MMSE-SD can give sharper images compared with the conventional down-sampling methods, with little color fringing artifacts. Lu Fang 0001, Oscar C. Au |
ICASSP | 1 |
| 2010 | A highly data reusable and standard-compliant motion estimation hardware architectureabstractMotion Estimation (ME) is the most computationally intensive part in the whole video compression process. The ME algorithms can be divided into full search ME (FS) and fast ME (FME). The FS is not suitable for high definition (HD) frame size videos because its relevant high computation load and hard to deal with complex motions in limited search range. A lot of FME algorithms have been proposed which can significantly reduce the computation load compared to FS. Though many kinds of hardware implementations of ME have been proposed, almost all of them fail to consider about the motion vector field (MVF) coherence and rate-distortion (RD) cost which have significant impact to the coding efficiency. In this paper, we propose a hardware friendly ME algorithm and corresponding highly data reusable hardware architecture. Simulation results show that the proposed ME algorithm performs better RD performance than conventional FME algorithm. The proposed reconfigurable ME hardware is implemented in VHDL and mapped to a low cost Xilinx XC3S1500 FPGA. It works at 100MHz and is capable to process 1920 × 1080 of 30fps video format in real time and have very high data reuse ratio. Oscar C. Au, Jiang Xu 0001, Lu Fang 0001, Run Cha |
ICME | 4 |
| 2010 | Subpixel-based down-sampling via Min-Max Directional ErrorabstractSubpixel-based down-sampling is a method that can potentially improve the apparent resolution of a down-scaled image by controlling individual subpixels rather than pixels. However, the increased luminance resolution often comes at the expense of chrominance distortion. In this paper, we formulate the subpixel-based down-sampling as a Min-Max problem (Min-Max Directional Error) which we call MMDE. Unfortunately, the solution of MMDE is computational intensive, especially for large images. We thus relax the MMDE by determining the maximum error based on HVS, which largely reduces the number of constraints. We call such relaxation as MMDE-VR (Visual Relaxation). Simulation results illustrate that MMDE-VR can effectively reduce visible color fringing artifacts while still maintaining sharpness. Lu Fang 0001, Oscar C. Au |
ISCAS | 1 |
| 2009 | Subpixel-based image downsampling-some analysis and observationabstractOften we need to shrink a high resolution image (e.g. 10-mega pixel) in order to display it on a low resolution display (e.g mobile phone). Signal processing theory tells us that optimal decimation requires low-pass filtering with a suitable cutoff frequency followed by downsampling. In doing so, we need to remove lots of details in the original high resolution image. In this paper, we review some little known results on an interesting topic called subpixel rendering, which can provide apparent higher resolution at the expense of color fringing artifacts. We attempt to explain what happens and why this is even possible. Lu Fang 0001, Oscar C. Au, Yi Yang 0041, Weiran Tang |
ICME | 1 |
| 2009 | LMMSE frequency merging for demosaickingabstractFor raw images captured by most digital cameras, every pixel has only on color in R, G and B. Kinds demosaicking algorithms are proposed for interpolating the missing two colors. In this article, the relationships inter and intra color channels are analyzed, and basing on the features, we propose a method to divide raw images into sub images and merge them in frequency domain with linear combination. Optimal weights are calculated with estimation values and raw values based on minimum mean square error criteria. Experiments results with different estimations are presented and discussed. Weiran Tang, Oscar C. Au, Yi Yang 0041, Lu Fang 0001 |
ICME | 5 |
| 2009 | Perceptual compressive sensing for image signalsabstractHuman eyes have different sensitivity to different frequency components of image signals, typically, low frequency components are relatively more crucial to the perceptual quality of images than high frequency components. Based on this observation, we propose a novel sampling scheme for compressive sensing framework by designing a weighting scheme for the sampling matrix. By adjusting the weighting coefficients, we can tune the structure of the sampling matrix to favor the frequency components that are important to human perception, so that those components could be more precisely recovered in the reconstruction procedure. Experimental results reveal that our proposed scheme can greatly enhance the performance of compressive sensing framework in both PSNR and visual quality without increasing the complexity of the framework structure or computational procedure. Yi Yang 0041, Oscar C. Au, Lu Fang 0001, Weiran Tang |
ICME | 3 |
| 2009 | A New Adaptive Subpixel-based Downsampling Scheme using Edge DetectionabstractIn this paper, a new adaptive subpixel-based downsampling scheme is proposed. Inside this scheme, we take full advantage of subpixels by adaptively choosing the sample directions based on edge information which has not been addressed before. Then, an adaptive filter is designed to suppress color fringing artifacts. Moreover, a good cut-off frequency is derived and deployed in our filter to obtain extra information. Simulation results illustrate that the proposed adaptive subpixel-based downsampling scheme successfully improves the resolution while efficiently removes visible color fringing artifacts. Lu Fang 0001, Oscar C. Au, Yi Yang 0041, Weiran Tang |
ISCAS | 1 |
| 2009 | Frequency Selection and Merging with Universal Matrices for Color Filter Array DemosaickingabstractIn most digital cameras, color filter array (CFA) is used for sampling only one color value for each pixel on CCD. So the other two color values need to be interpolated. The interpolation process is commonly known as demosaicking. In this paper, we discuss the relationship among the down sampled CFA images in frequency domain. According to the correlation, we propose an interpolation method that select specified parts from every down sampled frequency CFA image and merge them together. For more accurate selection, we calculate the minimum error of points in frequency images, and produce some universal matrices to record the source of every point in merging. And then, interpolation is based on the universal matrices. This algorithm is a non-adaptive method and regular in hardware implementation. Weiran Tang, Oscar C. Au, Yi Yang 0041, Lu Fang 0001 |
ISCAS | 5 |
| 2009 | Reweighted Compressive Sampling for image compressionabstractCompressive Sampling (CS), is an emerging theory which points us a promising direction of designing novel efficient data compression techniques. However, the conventional CS adopts a non-discriminated sampling scheme which usually gives poor performance on realistic complex signals. In this paper we propose a reweighted Compressive Sampling for image compression. It introduces a weighting scheme into the conventional CS framework whose coefficients are determined in encoding side according to the statistics of image signals. Experimental results demonstrate that our proposed method notably outperforms the conventional Compressive Sampling framework in coding performance in the sense that the reconstruction quality is greatly enhanced with same number of measurements and computational complexity. Yi Yang 0041, Oscar C. Au, Lu Fang 0001, Weiran Tang |
PCS | 3 |