VLDB 2026 Research / reviewers in the wild / expert
Nianjuan Jiang
dblp:01/5601
· DBLP profile ↗
23ranked-venue papers
5as first author
7since 2021 · last 2024
0000-0001-8338-0245ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 18 · 5 first-author · 4 since 2021Artificial intelligence and machine learning · 16 · 4 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
13 papers |
3D vision · 84% Face, body and person analysis · 10% Efficient and distributed learning · 6% | |
| Computer graphics and multimedia
9 papers |
Image and video processing · 57% Rendering · 18% Geometric modeling and processing · 14% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Hardware accelerators and domain-specific architectures · 50% GPUs and heterogeneous computing · 50% |
Topics — the 30 heaviest of 45, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Image and video processing
image restoration |
1.9 | 3 | 2024 | From NeRFLiX to NeRFLiX++: A General NeRF-Agnostic Restorer Paradigm · IEEE Trans. Pattern Anal. Mach. Intell. 2024 NeRFLiX: High-Quality Neural View Synthesis by Learning a Degradation-Driven Inter-viewpoint MiXer · CVPR 2023 LAPAR: Linearly-Assembled Pixel-Adaptive Regression Network for Single Image Super-resolution and Beyond · NeurIPS 2020 |
Computer vision › 3D vision
point cloud analysis |
1.2 | 2 | 2023 | CRIN: Rotation-Invariant Point Cloud Analysis and Rotation Estimation via Centrifugal Reference Frame · AAAI 2023 ScatterNet: Point Cloud Learning via Scatters · ACM Multimedia 2022 |
Computer vision › Face, body and person analysis
human pose estimation |
1.0 | 2 | 2022 | HEMlets PoSh: Learning Part-Centric Heatmap Triplets for 3D Human Pose and Shape Estimation · IEEE Trans. Pattern Anal. Mach. Intell. 2022 HEMlets Pose: Learning Part-Centric Heatmap Triplets for Accurate 3D Human Pose Estimation · ICCV 2019 |
Image and video processing › image restoration
degradation modeling |
0.8 | 1 | 2024 | From NeRFLiX to NeRFLiX++: A General NeRF-Agnostic Restorer Paradigm · IEEE Trans. Pattern Anal. Mach. Intell. 2024 |
Rendering
neural radiance fields |
0.8 | 1 | 2024 | From NeRFLiX to NeRFLiX++: A General NeRF-Agnostic Restorer Paradigm · IEEE Trans. Pattern Anal. Mach. Intell. 2024 |
Rendering
novel view synthesis |
0.8 | 1 | 2024 | From NeRFLiX to NeRFLiX++: A General NeRF-Agnostic Restorer Paradigm · IEEE Trans. Pattern Anal. Mach. Intell. 2024 |
Computer vision › 3D vision
neural radiance field |
0.7 | 1 | 2023 | NeRFLiX: High-Quality Neural View Synthesis by Learning a Degradation-Driven Inter-viewpoint MiXer · CVPR 2023 |
Computer vision › 3D vision
novel view synthesis |
0.7 | 1 | 2023 | NeRFLiX: High-Quality Neural View Synthesis by Learning a Degradation-Driven Inter-viewpoint MiXer · CVPR 2023 |
Computer vision › 3D vision › pose estimation
rotation estimation |
0.7 | 1 | 2023 | CRIN: Rotation-Invariant Point Cloud Analysis and Rotation Estimation via Centrifugal Reference Frame · AAAI 2023 |
Computer vision › 3D vision › local feature descriptor
rotation-invariant descriptor |
0.7 | 1 | 2023 | CRIN: Rotation-Invariant Point Cloud Analysis and Rotation Estimation via Centrifugal Reference Frame · AAAI 2023 |
Image and video processing › super-resolution › video super-resolution
real-time video super-resolution |
0.7 | 1 | 2023 | A High-Performance Accelerator for Super-Resolution Processing on Embedded GPU · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2023 |
Image and video processing
super-resolution |
0.7 | 1 | 2023 | A High-Performance Accelerator for Super-Resolution Processing on Embedded GPU · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2023 |
GPUs and heterogeneous computing
embedded GPU |
0.7 | 1 | 2023 | A High-Performance Accelerator for Super-Resolution Processing on Embedded GPU · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2023 |
Hardware accelerators and domain-specific architectures › vision accelerator
super-resolution accelerator |
0.7 | 1 | 2023 | A High-Performance Accelerator for Super-Resolution Processing on Embedded GPU · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2023 |
Computer vision › 3D vision
structure from motion |
0.7 | 4 | 2015 | Direct structure estimation for 3D reconstruction · CVPR 2015 A Global Linear Method for Camera Pose Registration · ICCV 2013 Seeing double without confusion: Structure-from-motion in highly ambiguous scenes · CVPR 2012 |
Computer vision › 3D vision
human mesh recovery |
0.6 | 1 | 2022 | HEMlets PoSh: Learning Part-Centric Heatmap Triplets for 3D Human Pose and Shape Estimation · IEEE Trans. Pattern Anal. Mach. Intell. 2022 |
Computer vision › 3D vision › local feature descriptor
local descriptor learning |
0.6 | 1 | 2022 | ScatterNet: Point Cloud Learning via Scatters · ACM Multimedia 2022 |
Image and video processing › super-resolution › image super-resolution
single image super-resolution |
0.4 | 1 | 2020 | LAPAR: Linearly-Assembled Pixel-Adaptive Regression Network for Single Image Super-resolution and Beyond · NeurIPS 2020 |
Computer vision › 3D vision
3d reconstruction |
0.4 | 2 | 2015 | Direct structure estimation for 3D reconstruction · CVPR 2015 A Global Linear Method for Camera Pose Registration · ICCV 2013 |
Computer vision › 3D vision › 3d human pose estimation
single-image 3d pose estimation |
0.4 | 1 | 2019 | HEMlets Pose: Learning Part-Centric Heatmap Triplets for Accurate 3D Human Pose Estimation · ICCV 2019 |
Image and video processing
image registration |
0.4 | 2 | 2017 | Direct Photometric Alignment by Mesh Deformation · CVPR 2017 SEAGULL: Seam-Guided Local Alignment for Parallax-Tolerant Image Stitching · ECCV (3) 2016 |
Computational photography and imaging
image stitching |
0.3 | 2 | 2017 | SEAGULL: Seam-Guided Local Alignment for Parallax-Tolerant Image Stitching · ECCV (3) 2016 Direct Photometric Alignment by Mesh Deformation · CVPR 2017 |
Geometric modeling and processing › shape registration
mesh registration |
0.3 | 1 | 2017 | Direct Photometric Alignment by Mesh Deformation · CVPR 2017 |
Geometric modeling and processing › mesh deformation
mesh warping |
0.3 | 1 | 2017 | Direct Photometric Alignment by Mesh Deformation · CVPR 2017 |
Virtual and augmented reality › tracking and registration
photometric registration |
0.3 | 1 | 2017 | Direct Photometric Alignment by Mesh Deformation · CVPR 2017 |
Image and video processing
video stabilization |
0.3 | 1 | 2017 | Direct Photometric Alignment by Mesh Deformation · CVPR 2017 |
Geometric modeling and processing
3d reconstruction |
0.2 | 1 | 2016 | RepMatch: Robust Feature Matching and Pose for Reconstructing Modern Cities · ECCV (1) 2016 |
Computational photography and imaging › image stitching
parallax-tolerant image stitching |
0.2 | 1 | 2016 | SEAGULL: Seam-Guided Local Alignment for Parallax-Tolerant Image Stitching · ECCV (3) 2016 |
Geometric modeling and processing › 3d reconstruction
urban reconstruction |
0.2 | 1 | 2016 | RepMatch: Robust Feature Matching and Pose for Reconstructing Modern Cities · ECCV (1) 2016 |
Computer vision › 3D vision
camera pose estimation |
0.2 | 1 | 2015 | Direct structure estimation for 3D reconstruction · CVPR 2015 |
Methods — techniques the papers use, named apart from their topics
inter-viewpoint aggregation · 2.1hardware-aware kernel optimization · 2.0dictionary slimming · 2.08-bit integer quantization · 2.0neural radiance field · 1.3integral operation · 1.0convolutional network · 1.0degradation simulation · 0.8deep neural network · 0.8centrifugal reference frame · 0.7attention-based down-sampling · 0.7SMPL regression · 0.6pixel-adaptive regression · 0.4linear coefficient regression · 0.4filter basis dictionary · 0.4euclidean rigidity constraint · 0.2direct structure estimation · 0.2missing correspondence optimization · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | From NeRFLiX to NeRFLiX++: A General NeRF-Agnostic Restorer ParadigmabstractNeural radiance fields (NeRF) have shown great success in novel view synthesis. However, recovering high-quality details from real-world scenes is still challenging for the existing NeRF-based approaches, due to the potential imperfect calibration information and scene representation inaccuracy. Even with high-quality training frames, the synthetic novel views produced by NeRF models still suffer from notable rendering artifacts, such as noise and blur. To address this, we propose NeRFLiX, a general NeRF-agnostic restorer paradigm that learns a degradation-driven inter-viewpoint mixer. Specially, we design a NeRF-style degradation modeling approach and construct large-scale training data, enabling the possibility of effectively removing NeRF-native rendering artifacts for deep neural networks. Moreover, beyond the degradation removal, we propose an inter-viewpoint aggregation framework that fuses highly related high-quality training images, pushing the performance of cutting-edge NeRF models to entirely new levels and producing highly photo-realistic synthetic views. Based on this paradigm, we further present NeRFLiX++ with a stronger two-stage NeRF degradation simulator and a faster inter-viewpoint mixer, achieving superior performance with significantly improved computational efficiency. Notably, NeRFLiX++ is capable of restoring photo-realistic ultra-high-resolution outputs from noisy low-resolution NeRF-rendered views. Extensive experiments demonstrate the excellent restoration ability of NeRFLiX++ on various novel view synthesis benchmarks. Kun Zhou 0001, Wenbo Li 0002, Nianjuan Jiang, Xiaoguang Han 0001, Jiangbo Lu |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | CRIN: Rotation-Invariant Point Cloud Analysis and Rotation Estimation via Centrifugal Reference FrameabstractVarious recent methods attempt to implement rotation-invariant 3D deep learning by replacing the input coordinates of points with relative distances and angles. Due to the incompleteness of these low-level features, they have to undertake the expense of losing global information. In this paper, we propose the CRIN, namely Centrifugal Rotation-Invariant Network. CRIN directly takes the coordinates of points as input and transforms local points into rotation-invariant representations via centrifugal reference frames. Aided by centrifugal reference frames, each point corresponds to a discrete rotation so that the information of rotations can be implicitly stored in point features. Unfortunately, discrete points are far from describing the whole rotation space. We further introduce a continuous distribution for 3D rotations based on points. Furthermore, we propose an attention-based down-sampling strategy to sample points invariant to rotations. A relation module is adopted at last for reinforcing the long-range dependencies between sampled points and predicts the anchor point for unsupervised rotation estimation. Extensive experiments show that our method achieves rotation invariance, accurately estimates the object rotation, and obtains state-of-the-art results on rotation-augmented classification and part segmentation. Ablation studies validate the effectiveness of the network design. Yujing Lou, Zelin Ye, Yang You 0004, Nianjuan Jiang, Jiangbo Lu, Lizhuang Ma, Cewu Lu |
AAAI | 4 |
| 2023 | NeRFLiX: High-Quality Neural View Synthesis by Learning a Degradation-Driven Inter-viewpoint MiXerabstractNeural radiance fields (NeRF) show great success in novel view synthesis. However, in real-world scenes, recovering high-quality details from the source images is still challenging for the existing NeRF-based approaches, due to the potential imperfect calibration information and scene representation inaccuracy. Even with high-quality training frames, the synthetic novel views produced by NeRF models still suffer from notable rendering artifacts, such as noise, blur, etc. Towards to improve the synthesis quality of NeRF-based approaches, we propose NeRFLiX, a general NeRF-agnostic restorer paradigm by learning a degradation-driven inter-viewpoint mixer. Specially, we design a NeRF-style degradation modeling approach and construct large-scale training data, enabling the possibility of effectively removing NeRF-native rendering artifacts for existing deep neural networks. Moreover, beyond the degradation removal, we propose an inter-viewpoint aggregation framework that is able to fuse highly related high-quality training images, pushing the performance of cutting-edge NeRF models to entirely new levels and producing highly photo-realistic synthetic views. Kun Zhou 0001, Wenbo Li 0001, Yi Wang 0074, Tao Hu 0011, Nianjuan Jiang, Xiaoguang Han 0001, Jiangbo Lu |
CVPR | 5 |
| 2023 | A High-Performance Accelerator for Super-Resolution Processing on Embedded GPUabstractOver the past few years, super-resolution (SR) processing has achieved astonishing progress along with the development of deep learning. Nevertheless, the rigorous requirement for real-time inference, especially for video tasks, leaves a harsh challenge for both the model architecture design and the hardware-level implementation. In this article, we propose a hardware-aware acceleration on embedded GPU devices as a full-stack SR deployment framework. The most critical stage with dictionary learning applied in SR flow was analyzed in details and optimized with a tailored dictionary slimming strategy. Moreover, we also delve into the programming architecture of hardware while analyzing the model structure to optimize the computation kernels to reduce inference latency and maximize the throughput given restricted computing power. In addition, we further accelerate the model with 8-bit integer inference by quantizing the weights in the compressed model. An adaptive 8-bit quantization flow for SR task enables the quantized model to achieve a comparable result with the full-precision baselines. With the help of our approaches, the computation and communication bottlenecks in the deep dictionary learning-based SR models can be overcome effectively. The experiments on both edge embedded device NVIDIA NX and 2080Ti prove that our framework exceeds the performance of state-of-the-art NVIDIA TensorRT significantly and can achieve real-time performance. Wenqian Zhao 0002, Qi Sun 0002, Wenbo Li 0002, Haisheng Zheng, Nianjuan Jiang, Jiangbo Lu, Bei Yu 0001, Martin D. F. Wong |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2022 | ScatterNet: Point Cloud Learning via ScattersabstractDesign of point cloud shape descriptors is a challenging problem in practical applications due to the sparsity and the inscrutable distribution of the point clouds. In this paper, we propose ScatterNet, a novel 3D local feature learning approach for exploring and aggregating hypothetical scatters of the point clouds. Scatters of relational points are first organized in point cloud via guided explorations, and then propagated back to extend the capacity in representing the point-wise characteristics. We provide an practical implementation of the ScatterNet, which involves an unique scatter exploration operator and a scatter convolution operator. Our method achieves the state-of-the-art performance on several point cloud analysis tasks like classification, part segmentation and normal estimation. The source code of ScatterNet is available in supplementary materials. Nianjuan Jiang, Jiangbo Lu, Mingang Chen, Ran Yi 0002, Lizhuang Ma |
ACM Multimedia | 2 |
| 2022 | HEMlets PoSh: Learning Part-Centric Heatmap Triplets for 3D Human Pose and Shape EstimationabstractEstimating 3D human pose from a single image is a challenging task. This work attempts to address the uncertainty of lifting the detected 2D joints to the 3D space by introducing an intermediate state - Part-Centric Heatmap Triplets (HEMlets), which shortens the gap between the 2D observation and the 3D interpretation. The HEMlets utilize three joint-heatmaps to represent the relative depth information of the end-joints for each skeletal body part. In our approach, a Convolutional Network (ConvNet) is first trained to predict HEMlets from the input image, followed by a volumetric joint-heatmap regression. We leverage on the integral operation to extract the joint locations from the volumetric heatmaps, guaranteeing end-to-end learning. Despite the simplicity of the network design, the quantitative comparisons show a significant performance improvement over the best-of-grade methods (e.g., 20 percent on Human3.6M). The proposed method naturally supports training with "in-the-wild" images, where only weakly-annotated relative depth information of skeletal joints is available. This further improves the generalization ability of our model, as validated by qualitative comparisons on outdoor images. Leveraging the strength of the HEMlets pose estimation, we further design and append a shallow yet effective network module to regress the SMPL parameters of the body pose and shape. We term the entire HEMlets-based human pose and shape recovery pipeline HEMlets PoSh. Extensive quantitative and qualitative experiments on the existing human body recovery benchmarks justify the state-of-the-art results obtained with our HEMlets PoSh approach. Kun Zhou 0001, Xiaoguang Han 0001, Nianjuan Jiang, Kui Jia, Jiangbo Lu |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | JOLO-GCN: Mining Joint-Centered Light-Weight Information for Skeleton-Based Action RecognitionabstractSkeleton-based action recognition has attracted research attentions in recent years. One common drawback in currently popular skeleton-based human action recognition methods is that the sparse skeleton information alone is not sufficient to fully characterize human motion. This limitation makes several existing methods incapable of correctly classifying action categories which exhibit only subtle motion differences. In this paper, we propose a novel framework for employing human pose skeleton and joint-centered light-weight information jointly in a two-stream graph convolutional network, namely, JOLO-GCN. Specifically, we use Joint-aligned optical Flow Patches (JFP) to capture the local subtle motion around each joint as the pivotal joint-centered visual information. Compared to the pure skeleton-based baseline, this hybrid scheme effectively boosts performance, while keeping the computational and memory overheads low. Experiments on the NTU RGB+D, NTU RGB+D 120, and the Kinetics-Skeleton dataset demonstrate clear accuracy improvements attained by the proposed method over the state-of-the-art skeleton-based methods. Jinmiao Cai, Nianjuan Jiang, Xiaoguang Han 0001, Kui Jia, Jiangbo Lu |
WACV | 2 |
| 2020 | LAPAR: Linearly-Assembled Pixel-Adaptive Regression Network for Single Image Super-resolution and BeyondabstractSingle image super-resolution (SISR) deals with a fundamental problem of upsampling a low-resolution (LR) image to its high-resolution (HR) version. Last few years have witnessed impressive progress propelled by deep learning methods. However, one critical challenge faced by existing methods is to strike a sweet spot of deep model complexity and resulting SISR quality. This paper addresses this pain point by proposing a linearly-assembled pixel-adaptive regression network (LAPAR), which casts the direct LR to HR mapping learning into a linear coefficient regression task over a dictionary of multiple predefined filter bases. Such a parametric representation renders our model highly lightweight and easy to optimize while achieving state-of-the-art results on SISR benchmarks. Moreover, based on the same idea, LAPAR is extended to tackle other restoration tasks, e.g., image denoising and JPEG image deblocking, and again, yields strong performance. Wenbo Li 0002, Kun Zhou 0001, Lu Qi 0001, Nianjuan Jiang, Jiangbo Lu, Jiaya Jia |
NeurIPS | 4 |
| 2019 | HEMlets Pose: Learning Part-Centric Heatmap Triplets for Accurate 3D Human Pose EstimationabstractEstimating 3D human pose from a single image is a challenging task. This work attempts to address the uncertainty of lifting the detected 2D joints to the 3D space by introducing an intermediate state - Part-Centric Heatmap Triplets (HEMlets), which shortens the gap between the 2D observation and the 3D interpretation. The HEMlets utilize three joint-heatmaps to represent the relative depth information of the end-joints for each skeletal body part. In our approach, a Convolutional Network(ConvNet) is first trained to predict HEMlests from the input image, followed by a volumetric joint-heatmap regression. We leverage on the integral operation to extract the joint locations from the volumetric heatmaps, guaranteeing end-to-end learning. Despite the simplicity of the network design, the quantitative comparisons show a significant performance improvement over the best-of-grade method (by 20% on Human3.6M). The proposed method naturally supports training with "in-the-wild'' images, where only weakly-annotated relative depth information of skeletal joints is available. This further improves the generalization ability of our model, as validated by qualitative comparisons on outdoor images. Kun Zhou 0001, Xiaoguang Han 0001, Nianjuan Jiang, Kui Jia, Jiangbo Lu |
ICCV | 3 |
| 2018 | Robust Video Background Identification by Dominant Rigid Motion Estimation
Kaimo Lin, Nianjuan Jiang, Loong Fah Cheong, Jiangbo Lu, Xun Xu 0002 |
ACCV (2) | 2 |
| 2018 | Locating 3D Object Proposals: A Depth-Based Online Approachabstract2D object proposals, quickly detected regions in an image that likely contain an object of interest, are an effective approach for improving the computational efficiency and accuracy of object detection in color images. In this paper, we propose a novel online method that generates 3D object proposals in an RGB-D video sequence. Our main observation is that depth images provide important information about the geometry of the scene. Diverging from the traditional goal of 2D object proposals to provide a high recall, we aim for precise 3D proposals. We leverage on depth information per frame and multiview scene information to obtain accurate 3D object proposals. Using efficient but robust registration enables us to combine multiple frames of a scene in near real time and generate 3D bounding boxes for potential 3D regions of interest. Using standard metrics, such as precision-recall (P-R) curves and F-measure, we show that the proposed approach is significantly more accurate than the current state-of-the-art techniques. Our online approach can be integrated into simultaneous localization and mapping-based video processing for quick 3D object localization. Our method takes less than a second in MATLAB on the UW-RGBD scene data set on a single thread CPU and, thus, has potential to be used in low-power chips in unmanned aerial vehicles, quadcopters, and drones. Ramanpreet Singh Pahwa, Jiangbo Lu, Nianjuan Jiang, Tian-Tsong Ng, Minh N. Do |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2017 | Direct Photometric Alignment by Mesh DeformationabstractThe choice of motion models is vital in applications like image/video stitching and video stabilization. Conventional methods explored different approaches ranging from simple global parametric models to complex per-pixel optical flow. Mesh-based warping methods achieve a good balance between computational complexity and model flexibility. However, they typically require high quality feature correspondences and suffer from mismatches and low-textured image content. In this paper, we propose a mesh-based photometric alignment method that minimizes pixel intensity difference instead of Euclidean distance of known feature correspondences. The proposed method combines the superior performance of dense photometric alignment with the efficiency of mesh-based image warping. It achieves better global alignment quality than the feature-based counterpart in textured images, and more importantly, it is also robust to low-textured image content. Abundant experiments show that our method can handle a variety of images and videos, and outperforms representative state-of-the-art methods in both image stitching and video stabilization tasks. Kaimo Lin, Nianjuan Jiang, Shuaicheng Liu, Loong Fah Cheong, Minh N. Do, Jiangbo Lu |
CVPR | 2 |
| 2016 | SEAGULL: Seam-Guided Local Alignment for Parallax-Tolerant Image Stitching
Kaimo Lin, Nianjuan Jiang, Loong Fah Cheong, Minh N. Do, Jiangbo Lu |
ECCV (3) | 2 |
| 2016 | RepMatch: Robust Feature Matching and Pose for Reconstructing Modern Cities
Wen-Yan Lin, Nianjuan Jiang, Minh N. Do, Jiangbo Lu |
ECCV (1) | 3 |
| 2015 | Linear Global Translation Estimation with Feature TracksabstractGlobal structure-from-motion (SfM) algorithms register all cameras simultaneously, which are potentially more efficient and less prone to drifting than incremental SfM methods. Global SfM methods often solve the camera orientations and positions separately. This paper focuses on the problem of global position (i.e. translation) estimation. Essential matrix based global translation estimation methods (e.g. [1]) usually degenerate at collinear camera motion because the translation scale is not determined by an essential matrix. Trifocal tensor based methods (e.g. [3]) usually rely on a strongly connected camera-triplet graph, where two triplets are connected by their common edge. The 3D reconstruction will distort or break into disconnected components when such strong association among images does not exist. The recent 1DSfM method [4] designs a smart filter to discard outlier essential matrices and solves scene points and cameras together by enforcing orientation consistency. However, this method requires abundant association between input images, e.g.∼O(n2) essential matrices for n cameras, which is more suitable for Internet images and often fails on sequentially captured data. The data association problem of [4] and [3] is exemplified in Figure 1. The Street example on the top is a sequential data where each image is only matched upto 4 neighbors. 1DSfM fails on this example due to insufficient image association. In the Seville example on the bottom, those Internet images are mostly captured from two viewpoints (see the two representative sample images) with weak affinity between images at different viewpoints. This weak data association causes seriously distorted reconstruction for the triplet-based method in [3]. This paper introduces a direct linear algorithm to address the presented challenges. It avoids degeneracy at collinear motion and deals with weakly associated data. Our method capitalizes on constraints from essential matrices and feature tracks. As shown in Figure 2 (a), the location of a scene point p can be computed as the middle point of the mutual perpendicular line segment AB of the two rays passing through p’s image projections: Zhaopeng Cui, Nianjuan Jiang, Chengzhou Tang, Ping Tan 0002 |
BMVC | 2 |
| 2015 | Direct structure estimation for 3D reconstructionabstractMost conventional structure-from-motion (SFM) techniques require camera pose estimation before computing any scene structure. In this work we show that when combined with single/multiple homography estimation, the general Euclidean rigidity constraint provides a simple formulation for scene structure recovery without explicit camera pose computation. This direct structure estimation (DSE) opens a new way to design a SFM system that reverses the order of structure and motion estimation. We show that this alternative approach works well for recovering scene structure and camera poses from sideway motion given planar or general man-made scenes. Nianjuan Jiang, Wen-Yan Lin, Minh N. Do, Jiangbo Lu |
CVPR | 1 |
| 2013 | A Global Linear Method for Camera Pose RegistrationabstractWe present a linear method for global camera pose registration from pair wise relative poses encoded in essential matrices. Our method minimizes an approximate geometric error to enforce the triangular relationship in camera triplets. This formulation does not suffer from the typical `unbalanced scale' problem in linear methods relying on pair wise translation direction constraints, i.e. an algebraic error, nor the system degeneracy from collinear motion. In the case of three cameras, our method provides a good linear approximation of the trifocal tensor. It can be directly scaled up to register multiple cameras. The results obtained are accurate for point triangulation and can serve as a good initialization for final bundle adjustment. We evaluate the algorithm performance with different types of data and demonstrate its effectiveness. Our system produces good accuracy, robustness, and outperforms some well-known systems on efficiency. Nianjuan Jiang, Zhaopeng Cui, Ping Tan 0002 |
ICCV | 1 |
| 2012 | Seeing double without confusion: Structure-from-motion in highly ambiguous scenesabstract3D reconstruction from an unordered set of images may fail due to incorrect epipolar geometries (EG) between image pairs arising from ambiguous feature correspondences. Previous methods often analyze the consistency between different EGs, and regard the largest subset of self-consistent EGs as correct. However, as demonstrated in [14], such a largest self-consistent set often corresponds to incorrect result, especially when there are duplicate structures in the scene. We propose a novel optimization criteria based on the idea of `missing correspondences'. The global minimum of our optimization objective function is associated with the correct solution. We then design an efficient algorithm for minimization, whose convergence to a local minimum is guaranteed. Experimental results show our method outperforms the state-of-the-art. Nianjuan Jiang, Loong Fah Cheong |
CVPR | 1 |
| 2011 | Multi-view repetitive structure detectionabstractSymmetry, especially repetitive structures in architecture are universally demonstrated across countries and cultures. Existing detection methods mainly focus on the detection of planar patterns from a single image. It is difficult to apply them to detect repetitive structures in architecture, which abounds with non-planar 3D repetitive elements (such as balconies and windows) and curved surfaces. We study the repetitive structure detection problem from multiple images of such architecture. Our method jointly analyzes these images and a set of 3D points reconstructed from them by structure-from-motion algorithms. 3D points help to rectify geometric deformations and hypothesize possible lattice structures, while images provide denser color and texture information to evaluate and confirm these hypotheses. In the experiments, we compare our method with existing algorithm. We also show how our results might be used to assist image-based modeling. Nianjuan Jiang, Loong Fah Cheong |
ICCV | 1 |
| 2009 | Automatic camera calibration of broadcast tennis video with applications to 3D virtual content insertion and ball detection and tracking
Xinguo Yu, Nianjuan Jiang, Loong Fah Cheong, Hon Wai Leong, Xin Yan 0001 |
Comput. Vis. Image Underst. | 2 |
| 2009 | Symmetric architecture modeling with a single imageabstractWe present a method to recover a 3D texture-mapped architecture model from a single image. Both single image based modeling and architecture modeling are challenging problems. We handle these difficulties by employing constraints derived from shape symmetries, which are prevalent in architecture. We first present a novel algorithm to calibrate the camera from a single image by exploiting symmetry. Then a set of 3D points is recovered according to the calibration and the underlying symmetry. With these reconstructed points, the user interactively marks out components of the architecture structure, whose shapes and positions are automatically determined according to the 3D points. Lastly, we texture the 3D model according to the input image, and we enhance the texture quality at those foreshortened and occluded regions according to their symmetric counterparts. The modeling process requires only a few minutes interaction. Multiple examples are provided to demonstrate the presented method. Nianjuan Jiang, Loong Fah Cheong |
ACM Trans. Graph. | 1 |
| 2007 | Accurate and Stable Camera Calibration of Broadcast Tennis VideoabstractThis paper presents an original algorithm for accurate and stable camera calibration of broadcast tennis video (BTV). That frame-data of BTV is often erroneous results in wildly fluctuating camera parameters. To meet this challenge, we propose aframegroupingtechnique, which groups frames together according to camera viewpoint. We then use a group-wise data analysis to obtain more stable parameters. Recognizing the fact that some of these parameters do vary somewhat even if they have a similar camera viewpoint, we further employ aHough-likesearch to tune them, maximizing the reprojection similarity. This two-tiered process gains stability of the camera parameters, and yet ensures large reprojection similarity via the tuning step. The experimental results show that our algorithm is able to acquire accurate camera matrix. Xinguo Yu, Nianjuan Jiang, Loong Fah Cheong |
ICIP (3) | 2 |
| 2007 | Trajectory-based ball detection and tracking with aid of homography in broadcast tennis videoabstractBall-detection-and-tracking in broadcast tennis video (BTV) is a crucial but challenging task in tennis video semantics analysis. Informally, the challenges are due to camera motion and the other causes such as the presence of many ball-like objects and the small size of the tennis ball. The trajectory-based approach proposed by us in our previous papers mainly counteracted the challenges imposed by causes other than camera motion and achieves a good performance. This paper proposes an improved trajectory-based ball detection and tracking algorithm in BTV with the aid of homography, which counteracts the challenges caused by camera motion and bring us multiple new merits. Firstly, it acquires an accurate homography, which transforms each frame into the "standard" frame. Secondly, it achieved higher accuracy of ball identification. Thirdly, it obtains the ball projection position in the real world, instead of ball location in the image. Lastly, it also identifies landing frames and positions of the ball. The experimental results show that the improved algorithm can obtain not only higher accuracy in ball identification and in ball position alike, but also ball landing frames and positions. With the intent of using homography to improve the video-based event detection for smart home we also do some experiments on acquiring the homography for home surveillance video. Xinguo Yu, Nianjuan Jiang, Ee-Luang Ang |
VCIP | 2 |