VLDB 2026 Research / reviewers in the wild / expert
Simon Chen
dblp:155/6097
· DBLP profile ↗
6ranked-venue papers
0as first author
5since 2021 · last 2023
0000-0003-0409-5978ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
3D vision · 87% Segmentation and scene understanding · 10% Transfer learning and domain adaptation · 3% | |
| Computer graphics and multimedia
1 paper |
Visual content generation and editing · 100% |
Topics — the 9 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision › depth estimation
monocular depth estimation |
1.3 | 3 | 2023 | Towards Accurate Reconstruction of 3D Scene Shape From A Single Monocular Image · IEEE Trans. Pattern Anal. Mach. Intell. 2023 Learning To Recover 3D Scene Shape From a Single Image · CVPR 2021 Layered Depth Refinement with Mask Guidance · CVPR 2022 |
Computer vision › 3D vision
3d scene reconstruction |
1.2 | 2 | 2023 | Towards Accurate Reconstruction of 3D Scene Shape From A Single Monocular Image · IEEE Trans. Pattern Anal. Mach. Intell. 2023 Learning To Recover 3D Scene Shape From a Single Image · CVPR 2021 |
Computer vision › 3D vision › depth estimation › monocular depth estimation
metric depth estimation |
1.2 | 2 | 2023 | Towards Accurate Reconstruction of 3D Scene Shape From A Single Monocular Image · IEEE Trans. Pattern Anal. Mach. Intell. 2023 Learning To Recover 3D Scene Shape From a Single Image · CVPR 2021 |
Computer vision › 3D vision
depth estimation |
0.6 | 1 | 2022 | Layered Depth Refinement with Mask Guidance · CVPR 2022 |
Computer vision › 3D vision › depth estimation
depth map refinement |
0.6 | 1 | 2022 | Layered Depth Refinement with Mask Guidance · CVPR 2022 |
Visual content generation and editing › image editing
image compositing |
0.5 | 1 | 2021 | SSH: A Self-Supervised Framework for Image Harmonization · ICCV 2021 |
Visual content generation and editing › image editing › image compositing
image harmonization |
0.5 | 1 | 2021 | SSH: A Self-Supervised Framework for Image Harmonization · ICCV 2021 |
Computer vision › 3D vision › camera calibration
focal length estimation |
0.2 | 1 | 2023 | Towards Accurate Reconstruction of 3D Scene Shape From A Single Monocular Image · IEEE Trans. Pattern Anal. Mach. Intell. 2023 |
Machine learning › Transfer learning and domain adaptation
zero-shot transfer |
0.1 | 1 | 2021 | Learning To Recover 3D Scene Shape From a Single Image · CVPR 2021 |
Methods — techniques the papers use, named apart from their topics
3d color lookup table · 1.0relative depth training · 0.7normal-based geometry loss · 0.7image-level normalized regression loss · 0.7self-supervised learning · 0.6layered refinement · 0.6inpainting · 0.6representation fusion · 0.5point cloud encoder · 0.5normalized regression loss · 0.5geometry loss · 0.5data augmentation · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | PRN: Panoptic Refinement NetworkabstractPanoptic segmentation is the task of uniquely assigning every pixel in an image to either a semantic label or an individual object instance, generating a coherent and complete scene description. Many current panoptic segmentation methods, however, predict masks of semantic classes and object instances in separate branches, yielding inconsistent predictions. Moreover, because state-of-the-art panoptic segmentation models rely on box proposals, the instance masks predicted are often of low-resolution. To overcome these limitations, we propose the Panoptic Refinement Network (PRN), which takes masks from base panoptic segmentation models and refines them jointly to produce coherent results. PRN extends the offset map-based architecture of Panoptic-Deeplab with several novel ideas including a foreground mask and instance bounding box offsets, as well as coordinate convolutions for improved spatial prediction. Experimental results on COCO and Cityscapes show that PRN can significantly improve already accurate results from a variety of panoptic segmentation networks. Jason Kuen, Zhe Lin 0001, Philippos Mordohai, Simon Chen |
WACV | 5 |
| 2023 | Towards Accurate Reconstruction of 3D Scene Shape From A Single Monocular ImageabstractDespite significant progress made in the past few years, challenges remain for depth estimation using a single monocular image. First, it is nontrivial to train a metric-depth prediction model that can generalize well to diverse scenes mainly due to limited training data. Thus, researchers have built large-scale relative depth datasets that are much easier to collect. However, existing relative depth estimation models often fail to recover accurate 3D scene shapes due to the unknown depth shift caused by training with the relative depth data. We tackle this problem here and attempt to estimate accurate scene shapes by training on large-scale relative depth data, and estimating the depth shift. To do so, we propose a two-stage framework that first predicts depth up to an unknown scale and shift from a single monocular image, and then exploits 3D point cloud data to predict the depth shift and the camera's focal length that allow us to recover 3D scene shapes. As the two modules are trained separately, we do not need strictly paired training data. In addition, we propose an image-level normalized regression loss and a normal-based geometry loss to improve training with relative depth annotation. We test our depth model on nine unseen datasets and achieve state-of-the-art performance on zero-shot evaluation. Code is available at: https://github.com/aim-uofa/depth/. Wei Yin 0006, Jianming Zhang 0001, Oliver Wang, Simon Niklaus, Simon Chen, Yifan Liu 0001, Chunhua Shen |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Layered Depth Refinement with Mask GuidanceabstractDepth maps are used in a wide range of applications from 3D rendering to 2D image effects such as Bokeh. However, those predicted by single image depth estimation (SIDE) models often fail to capture isolated holes in objects and/or have inaccurate boundary regions. Meanwhile, high-quality masks are much easier to obtain, using commercial auto-masking tools or off-the-shelf methods of segmentation and matting or even by manual editing. Hence, in this paper, we formulate a novel problem of mask-guided depth refinement that utilizes a generic mask to refine the depth prediction of SIDE models. Our framework performs layered refinement and inpainting/outpainting, decomposing the depth map into two separate layers signified by the mask and the inverse mask. As datasets with both depth and mask annotations are scarce, we propose a self-supervised learning scheme that uses arbitrary masks and RGB-D datasets. We empirically show that our method is robust to different types of masks and initial depth predictions, accurately refining depth values in inner and outer mask boundary regions. We further analyze our model with an ablation study and demonstrate results on real applications. More information can be found on our project page.11https://sooyekim.github.io/MaskDepth/ Soo Ye Kim, Jianming Zhang 0001, Simon Niklaus, Simon Chen, Zhe Lin 0001, Munchurl Kim |
CVPR | 5 |
| 2021 | Learning To Recover 3D Scene Shape From a Single ImageabstractDespite significant progress in monocular depth estimation in the wild, recent state-of-the-art methods cannot be used to recover accurate 3D scene shape due to an unknown depth shift induced by shift-invariant reconstruction losses used in mixed-data depth prediction training, and possible unknown camera focal length. We investigate this problem in detail, and propose a two-stage framework that first predicts depth up to an unknown scale and shift from a single monocular image, and then use 3D point cloud encoders to predict the missing depth shift and focal length that allow us to recover a realistic 3D scene shape. In addition, we propose an image-level normalized regression loss and a normal-based geometry loss to enhance depth prediction models trained on mixed datasets. We test our depth model on nine unseen datasets and achieve state-of-the-art performance on zero-shot dataset generalization. Code is available at: https://git.io/Depth Wei Yin 0006, Jianming Zhang 0001, Oliver Wang, Simon Niklaus, Long Mai, Simon Chen, Chunhua Shen |
CVPR | 6 |
| 2021 | SSH: A Self-Supervised Framework for Image HarmonizationabstractImage harmonization aims to improve the quality of image compositing by matching the "appearance" (e.g., color tone, brightness and contrast) between foreground and background images. However, collecting large-scale annotated datasets for this task requires complex professional retouching. Instead, we propose a novel Self-Supervised Harmonization framework (SSH) that can be trained using just "free" natural images without being edited. We reformulate the image harmonization problem from a representation fusion perspective, which separately processes the foreground and background examples, to address the background occlusion issue. This framework design allows for a dual data augmentation method, where diverse [foreground, background, pseudo GT] triplets can be generated by cropping an image with perturbations using 3D color lookup tables (LUTs). In addition, we build a real-world harmonization dataset as carefully created by expert users, for evaluation and benchmarking purposes. Our results show that the proposed self-supervised method outperforms previous state-of-the-art methods in terms of reference metrics, visual quality, and subject user study. Code and dataset are available at https://github.com/VITA-Group/SSHarmonization. Yifan Jiang 0001, He Zhang 0004, Jianming Zhang 0001, Yilin Wang 0002, Zhe Lin 0001, Kalyan Sunkavalli, Simon Chen, Sohrab Amirghodsi, Sarah Kong, Zhangyang Wang |
ICCV | 7 |
| 2011 | Modeling and removing spatially-varying optical blurabstractPhoto deblurring has been a major research topic in the past few years. So far, existing methods have focused on removing the blur due to camera shake and object motion. In this paper, we show that the optical system of the camera also generates significant blur, even with professional lenses. We introduce a method to estimate the blur kernel densely over the image and across multiple aperture and zoom settings. Our measures show that the blur kernel can have a non-negligible spread, even with top-of-the-line equipment, and that it varies nontrivially over this domain. In particular, the spatial variations are not radially symmetric and not even left-right symmetric. We develop and compare two models of the optical blur, each of them having its own advantages. We show that our models predict accurate blur kernels that can be used to restore photos. We demonstrate that we can produce images that are more uniformly sharp unlike those produced with spatially-invariant deblurring techniques. Eric Kee, Sylvain Paris, Simon Chen, Jue Wang 0001 |
ICCP | 3 |