EDBT 2026 Demo / reviewers in the wild / expert
Chengzhou Tang
dblp:133/4073
· DBLP profile ↗
13ranked-venue papers
8as first author
4since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 6 first-author · 4 since 2021Artificial intelligence and machine learning · 8 · 3 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
3D vision · 78% Generative modeling · 14% Optimization for machine learning · 4% | |
| Computer graphics and multimedia
3 papers |
Image and video processing · 56% Computational photography and imaging · 21% Virtual and augmented reality · 14% |
Topics — the 24 heaviest of 25, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision
3d reconstruction |
1.1 | 2 | 2024 | MVDiffusion++: A Dense High-Resolution Multi-view Diffusion Model for Single or Sparse-View 3D Object Reconstruction · ECCV (16) 2024 BA-Net: Dense Bundle Adjustment Networks · ICLR 2019 |
Computer vision › 3D vision
3d scene understanding |
1.0 | 2 | 2022 | RCP: Recurrent Closest Point for Point Cloud · CVPR 2022 SANet: Scene Agnostic Network for Camera Localization · ICCV 2019 |
Computer vision › 3D vision
visual localization |
0.9 | 2 | 2021 | Learning Camera Localization via Dense Scene Matching · CVPR 2021 SANet: Scene Agnostic Network for Camera Localization · ICCV 2019 |
Machine learning › Generative modeling
diffusion model |
0.8 | 1 | 2024 | MVDiffusion++: A Dense High-Resolution Multi-view Diffusion Model for Single or Sparse-View 3D Object Reconstruction · ECCV (16) 2024 |
Machine learning › Generative modeling › diffusion model › 3d-aware diffusion
multi-view diffusion |
0.8 | 1 | 2024 | MVDiffusion++: A Dense High-Resolution Multi-view Diffusion Model for Single or Sparse-View 3D Object Reconstruction · ECCV (16) 2024 |
Computer vision › 3D vision › 3d reconstruction
multi-view reconstruction |
0.8 | 1 | 2024 | MVDiffusion++: A Dense High-Resolution Multi-view Diffusion Model for Single or Sparse-View 3D Object Reconstruction · ECCV (16) 2024 |
Computer vision › 3D vision
point cloud processing |
0.6 | 1 | 2022 | RCP: Recurrent Closest Point for Point Cloud · CVPR 2022 |
Computer vision › 3D vision
point cloud registration |
0.6 | 1 | 2022 | RCP: Recurrent Closest Point for Point Cloud · CVPR 2022 |
Computer vision › 3D vision
scene flow estimation |
0.6 | 1 | 2022 | RCP: Recurrent Closest Point for Point Cloud · CVPR 2022 |
Image and video processing
image restoration |
0.6 | 1 | 2022 | Learning to Zoom Inside Camera Imaging Pipeline · CVPR 2022 |
Computational photography and imaging
image signal processing |
0.6 | 1 | 2022 | Learning to Zoom Inside Camera Imaging Pipeline · CVPR 2022 |
Image and video processing › super-resolution › image super-resolution
single image super-resolution |
0.6 | 1 | 2022 | Learning to Zoom Inside Camera Imaging Pipeline · CVPR 2022 |
Machine learning › Optimization for machine learning
energy minimization |
0.4 | 1 | 2020 | LSM: Learning Subspace Minimization for Low-Level Vision · CVPR 2020 |
Computer vision › Segmentation and scene understanding
interactive segmentation |
0.4 | 1 | 2020 | LSM: Learning Subspace Minimization for Low-Level Vision · CVPR 2020 |
Computer vision › 3D vision
low-level vision |
0.4 | 1 | 2020 | LSM: Learning Subspace Minimization for Low-Level Vision · CVPR 2020 |
Computer vision › 3D vision › motion estimation
optical flow |
0.4 | 1 | 2020 | LSM: Learning Subspace Minimization for Low-Level Vision · CVPR 2020 |
Computer vision › 3D vision › stereo vision
stereo matching |
0.4 | 1 | 2020 | LSM: Learning Subspace Minimization for Low-Level Vision · CVPR 2020 |
Computer vision › 3D vision › structure from motion
bundle adjustment |
0.4 | 1 | 2019 | BA-Net: Dense Bundle Adjustment Networks · ICLR 2019 |
Computer vision › 3D vision › visual localization
scene coordinate regression |
0.4 | 1 | 2019 | SANet: Scene Agnostic Network for Camera Localization · ICCV 2019 |
Computer vision › 3D vision
structure from motion |
0.4 | 1 | 2019 | BA-Net: Dense Bundle Adjustment Networks · ICLR 2019 |
Virtual and augmented reality
immersive video |
0.4 | 1 | 2019 | Joint Stabilization and Direction of 360° Videos · ACM Trans. Graph. 2019 |
Image and video processing
video stabilization |
0.4 | 1 | 2019 | Joint Stabilization and Direction of 360° Videos · ACM Trans. Graph. 2019 |
Visual content generation and editing
3d content creation |
0.2 | 1 | 2024 | MVDiffusion++: A Dense High-Resolution Multi-view Diffusion Model for Single or Sparse-View 3D Object Reconstruction · ECCV (16) 2024 |
Computer vision › 3D vision
depth estimation |
0.1 | 1 | 2021 | Learning Camera Localization via Dense Scene Matching · CVPR 2021 |
Methods — techniques the papers use, named apart from their topics
cost volume · 1.1convolutional neural network · 1.1subspace constraint · 0.6recurrent network · 0.6point-wise optimization · 0.6degradation model estimation · 0.6pnp algorithm · 0.5learnable subspace constraint · 0.4data-driven regularization · 0.4spherical warping · 0.4optimization · 0.4neural network · 0.4motion estimation · 0.4hierarchical scene representation · 0.4dense prediction · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | MVDiffusion++: A Dense High-Resolution Multi-view Diffusion Model for Single or Sparse-View 3D Object Reconstruction
Shitao Tang, Dilin Wang, Chengzhou Tang, Fuyang Zhang, Yuchen Fan 0001, Vikas Chandra, Yasutaka Furukawa |
ECCV (16) | 4 |
| 2022 | RCP: Recurrent Closest Point for Point Cloudabstract3D motion estimation including scene flow and point cloud registration has drawn increasing interest. Inspired by 2D flow estimation, recent methods employ deep neural networks to construct the cost volume for estimating accurate 3D flow. However, these methods are limited by the fact that it is difficult to define a search window on point clouds because of the irregular data structure. In this paper, we avoid this irregularity by a simple yet effective method. We decompose the problem into two interlaced stages, where the 3D flows are optimized point-wisely at the first stage and then globally regularized in a recurrent network at the second stage. Therefore, the recurrent network only receives the regular point-wise information as the input. In the experiments, we evaluate the proposed method on both the 3D scene flow estimation and the point cloud registration task. For 3D scene flow estimation, we make comparisons on the widely used FlyingThings3D [32] and KITTI [33] datasets. For point cloud registration, we follow previous works and evaluate the data pairs with large pose and partially overlapping from ModelNet40 [65]. The results show that our method outperforms the previous method and achieves a new state-of-the-art performance on both 3D scene flow estimation and point cloud registration, which demonstrates the superiority of the proposed zero-order method on irregular point cloud data. Our source code is available at https://github.com/gxd1994/RCP. Xiaodong Gu 0004, Chengzhou Tang, Weihao Yuan 0001, Zuozhuo Dai, Siyu Zhu 0001, Ping Tan 0002 |
CVPR | 2 |
| 2022 | Learning to Zoom Inside Camera Imaging PipelineabstractExisting single image super-resolution methods are either designed for synthetic data, or for real data but in the RGB-to-RGB or the RAW-to-RGB domain. This paper proposes to zoom an image from RAW to RAW inside the camera imaging pipeline. The RAW-to-RAW domain closes the gap between the ideal and the real degradation models. It also excludes the image signal processing pipeline, which refocuses the model learning onto the super-resolution. To these ends, we design a method that receives a low-resolution RAW as the input and estimates the desired higher-resolution RAW jointly with the degradation model. In our method, two convolutional neural networks are learned to constrain the high-resolution image and the degradation model in lower-dimensional subspaces. This subspace constraint converts the ill-posed SISR problem to a well-posed one. To demonstrate the superiority of the proposed method and the RAW-to-RAW domain, we conduct evaluations on the RealSR and the SR-RAW datasets. The results show that our method performs superiorly over the state-of-the-arts both qualitatively and quantitatively, and it also generalizes well and enables zero-shot transfer across different sensors. Chengzhou Tang, Yuqiang Yang, Bing Zeng 0001, Ping Tan 0002, Shuaicheng Liu |
CVPR | 1 |
| 2021 | Learning Camera Localization via Dense Scene MatchingabstractCamera localization aims to estimate 6 DoF camera poses from RGB images. Traditional methods detect and match interest points between a query image and a prebuilt 3D model. Recent learning-based approaches encode scene structures into a specific convolutional neural network (CNN) and thus are able to predict dense coordinates from RGB images. However, most of them require re-training or re-adaption for a new scene and have difficulties in handling large-scale scenes due to limited network capacity. We present a new method for scene agnostic camera localization using dense scene matching (DSM), where a cost volume is constructed between a query image and a scene. The cost volume and the corresponding coordinates are processed by a CNN to predict dense coordinates. Camera poses can then be solved by PnP algorithms. In addition, our method can be extended to temporal domain, which leads to extra performance boost during testing time. Our scene-agnostic approach achieves comparable accuracy as the existing scene-specific approaches, such as KFNet, on the 7scenes and Cambridge benchmark. This approach also remarkably outperforms state-of-the-art scene-agnostic dense coordinate regression network SANet. The Code is available at https://github.com/Tangshitao/DenseScene-Matching. Shitao Tang, Chengzhou Tang, Rui Huang 0001, Siyu Zhu 0001, Ping Tan 0002 |
CVPR | 2 |
| 2020 | LSM: Learning Subspace Minimization for Low-Level VisionabstractWe study the energy minimization problem in low-level vision tasks from a novel perspective. We replace the heuristic regularization term with a data-driven learnable subspace constraint, and preserve the data term to exploit domain knowledge derived from the first principles of a task. This learning subspace minimization (LSM) framework unifies the network structures and the parameters for many different low-level vision tasks, which allows us to train a single network for multiple tasks simultaneously with shared parameters, and even generalizes the trained network to an unseen task as long as the data term can be formulated. We validate our LSM frame on four low-level tasks including edge detection, interactive segmentation, stereo matching, and optical flow, and validate the network on various datasets. The experiments demonstrate that the proposed LSM generates state-of-the-art results with smaller model size, faster training convergence, and real-time inference. Chengzhou Tang, Lu Yuan 0001, Ping Tan 0002 |
CVPR | 1 |
| 2019 | SANet: Scene Agnostic Network for Camera LocalizationabstractThis paper presents a scene agnostic neural architecture for camera localization, where model parameters and scenes are independent from each other.Despite recent advancement in learning based methods, most approaches require training for each scene one by one, not applicable for online applications such as SLAM and robotic navigation, where a model must be built on-the-fly.Our approach learns to build a hierarchical scene representation and predicts a dense scene coordinate map of a query RGB image on-the-fly given an arbitrary scene. The 6D camera pose of the query image can be estimated with the predicted scene coordinate map. Additionally, the dense prediction can be used for other online robotic and AR applications such as obstacle avoidance. We demonstrate the effectiveness and efficiency of our method on both indoor and outdoor benchmarks, achieving state-of-the-art performance. Luwei Yang, Ziqian Bai, Chengzhou Tang, Honghua Li, Yasutaka Furukawa |
ICCV | 3 |
| 2019 | BA-Net: Dense Bundle Adjustment Networks
Chengzhou Tang |
ICLR | 1 |
| 2019 | Joint Stabilization and Direction of 360° VideosabstractThree-hundred-sixty-degree (360°) video provides an immersive experience for viewers, allowing them to freely explore the world by turning their head. However, creating high-quality 360° video content can be challenging, as viewers may miss important events by looking in the wrong direction, or they may see things that ruin the immersion, such as stitching artifacts and the film crew. We take advantage of the fact that not all directions are equally likely to be observed; most viewers are more likely to see content located at “true north,” i.e., in front of them, due to ergonomic constraints. We therefore propose 360° video direction, where the video is jointly optimized to orient important events to the front of the viewer and visual clutter behind them, while producing smooth camera motion. Unlike traditional video, viewers can still explore the space as desired, but with the knowledge that the most important content is likely to be in front of them. Constraints can be user guided, either added directly on the equirectangular projection or by recording “guidance” viewing directions while watching the video in a VR headset or automatically computed, such as via visual saliency or forward-motion direction. To accomplish this, we propose a new motion estimation technique specifically designed for 360° video that outperforms the commonly used five-point algorithm on wide-angle video. We additionally formulate the direction problem as an optimization where a novel parametrization of spherical warping allows us to correct for some degree of parallax effects. We compare our approach to recent methods that address stabilization-only and converting 360° video to narrow field-of-view video. Our pipeline can also enable the viewing of wide-angle non-360° footage in a spherical 360° space, giving an immersive “virtual cinema” experience for a wide range of existing content filmed with first-person cameras. Chengzhou Tang, Oliver Wang, Feng Liu 0015, Ping Tan 0002 |
ACM Trans. Graph. | 1 |
| 2017 | GSLAM: Initialization-Robust Monocular Visual SLAM via Global Structure-from-MotionabstractMany monocular visual SLAM algorithms are derived from incremental structure-from-motion (SfM) methods. This work proposes a novel monocular SLAM method which integrates recent advances made in global SfM. In particular, we present two main contributions to visual SLAM. First, we solve the visual odometry problem by a novel rank-1 matrix factorization technique which is more robust to the errors in map initialization. Second, we adopt a recent global SfM method for the pose-graph optimization, which leads to a multi-stage linear formulation and enables L1 optimization for better robustness to false loops. The combination of these two approaches generates more robust reconstruction and is significantly faster (4X) than recent state-of-the-art SLAM systems. We also present a new dataset recorded with ground truth camera motion in a Vicon motion capture room, and compare our method to prior systems on it and established benchmark datasets. Chengzhou Tang, Oliver Wang |
3DV | 1 |
| 2015 | Linear Global Translation Estimation with Feature TracksabstractGlobal structure-from-motion (SfM) algorithms register all cameras simultaneously, which are potentially more efficient and less prone to drifting than incremental SfM methods. Global SfM methods often solve the camera orientations and positions separately. This paper focuses on the problem of global position (i.e. translation) estimation. Essential matrix based global translation estimation methods (e.g. [1]) usually degenerate at collinear camera motion because the translation scale is not determined by an essential matrix. Trifocal tensor based methods (e.g. [3]) usually rely on a strongly connected camera-triplet graph, where two triplets are connected by their common edge. The 3D reconstruction will distort or break into disconnected components when such strong association among images does not exist. The recent 1DSfM method [4] designs a smart filter to discard outlier essential matrices and solves scene points and cameras together by enforcing orientation consistency. However, this method requires abundant association between input images, e.g.∼O(n2) essential matrices for n cameras, which is more suitable for Internet images and often fails on sequentially captured data. The data association problem of [4] and [3] is exemplified in Figure 1. The Street example on the top is a sequential data where each image is only matched upto 4 neighbors. 1DSfM fails on this example due to insufficient image association. In the Seville example on the bottom, those Internet images are mostly captured from two viewpoints (see the two representative sample images) with weak affinity between images at different viewpoints. This weak data association causes seriously distorted reconstruction for the triplet-based method in [3]. This paper introduces a direct linear algorithm to address the presented challenges. It avoids degeneracy at collinear motion and deals with weakly associated data. Our method capitalizes on constraints from essential matrices and feature tracks. As shown in Figure 2 (a), the location of a scene point p can be computed as the middle point of the mutual perpendicular line segment AB of the two rays passing through p’s image projections: Zhaopeng Cui, Nianjuan Jiang, Chengzhou Tang, Ping Tan 0002 |
BMVC | 3 |
| 2014 | Sparse moving factorization for subspace video stabilizationabstractThis paper presents a new method for calculating the low-rank approximation of a highly incomplete trajectory matrix for subspace video stabilization. We extend moving factorization proposed in [1], which is a streamable method based on least squares. By utilizing sparse representation of trajectories, the proposed factorization method is more accurate while still streamable. We test our sparse moving factorization on synthetic data as well as real videos. Experiments on synthetic sequence demonstrate the numerical properties of our method, and stabilized videos show that our method outperforms moving factorization for subspace video stabilization. In addition, our results are also better than the ones from some other state-of-the-art video stabilization methods. Chengzhou Tang, Ronggang Wang |
ICASSP | 1 |
| 2014 | Local subspace video stabilizationabstractVideo stabilization enhances video quality by stabilizing unstable motion. This paper proposes a new video stabilization method that simultaneously factors and smooths motion trajectories. We model the trajectories with a time-variant local subspace constraint. Every column of the trajectory matrix is factored and smoothed in separate local subspace. This model makes our method more flexible and accurate than subspace video stabilization. In addition, we design a novel outlier detection technique that utilizes the relationship between consecutive local subspaces. Experiments on synthetic data validate the numerical performance of our factorization. Quantitative comparisons on real videos show that our local method is better than subspace video stabilization. Moreover, our stabilized videos are comparable with the public results from some other state-of-the-art methods. Chengzhou Tang, Ronggang Wang |
ICME | 1 |
| 2013 | Adaptive motion estimation order for frame rate up-conversionabstractThis paper proposes an adaptive motion estimation (ME) order for frame rate up-conversion (FRUC). Almost all existing FRUC methods adopt a raster scan order for ME. The ME is performed from top-left blocks to bottom-right blocks in raster scan order. Such an order can propagate some wrongly estimated motion vectors (MV) through a frame. The proposed method first detects the blocks rich in features (feature blocks) and estimates their MVs. Then ME is performed on the other blocks according to their distance to feature blocks. The closer a block to feature blocks is; the earlier the ME is performed on it. In this adaptive order, MVs of feature blocks are propagated to its neighbors. It makes the estimated motions of a frame close to the true motions. In order to demonstrate the efficiency of the proposed method, we estimate the MVs with diamond search in the proposed adaptive ME order. In the experiments, the quality of frame rate up converted videos have been significantly improved compared with the ones using traditional raster scan order. Moreover, the adaptive ME order can be easily combined with various ME methods applied in previous FRUC. Chengzhou Tang, Ronggang Wang, Wenmin Wang 0001 |
ISCAS | 1 |