VLDB 2026 Research / reviewers in the wild / expert
Nianyi Li
dblp:126/0742
· DBLP profile ↗
17ranked-venue papers
6as first author
7since 2021 · last 2024
0000-0002-4172-4940ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 13 · 5 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Turb-Seg-Res: A Segment-then-Restore Pipeline for Dynamic Videos with Atmospheric TurbulenceabstractTackling image degradation due to atmospheric turbu-lence, particularly in dynamic environments, remains a challenge for long-range imaging systems. Existing techniques have been primarily designed for static scenes or scenes with small motion. This paper presents the first segment-then-restore pipeline for restoring the videos of dy-namic scenes in turbulent environments. We leverage mean optical flow with an unsupervised motion segmentation method to separate dynamic and static scene components prior to restoration. After camera shake compensation and segmentation, we introduce foreground/background en-hancement leveraging the statistics of turbulence strength and a transformer model trained on a novel noise-based procedural turbulence generator for fast dataset augmen-tation. Benchmarked against existing restoration meth-ods, our approach restores most of the geometric distortion and enhances the sharpness of videos. We make our code, simulator, and data publicly available to ad-vance the field of video restoration from turbulence: riponcs.github.io/TurbSegRes Ripon K. Saha, Dehao Qin, Nianyi Li, Jinwei Ye, Suren Jayasuriya |
CVPR | 3 |
| 2024 | Unsupervised Moving Object Segmentation with Atmospheric Turbulence
Dehao Qin, Ripon K. Saha, Woojeh Chung, Suren Jayasuriya, Jinwei Ye, Nianyi Li |
ECCV (6) | 6 |
| 2024 | Unsupervised Coordinate-Based Video DenoisingabstractIn this paper, we introduce a novel unsupervised video denoising deep learning approach that can help to mitigate data scarcity issues and show robustness against different noise patterns, enhancing its broad applicability. Our method comprises three modules: a Feature generator creating feature maps, a Denoise-Net generating denoised but slightly blurry reference frames, and a Refine-Net re-introducing high-frequency details. By leveraging the coordinate-based network, we can greatly simplify the network structure while preserving high-frequency details in the denoised video frames. Extensive experiments on both simulated and real-captured videos demonstrate that our method can effectively denoise real-world calcium imaging video sequences without prior knowledge of noise models and data augmentation during training. Mary Damilola Aiyetigbo, Dineshchandar Ravichandran, Reda Chalhoub, Peter Kalivas, Nianyi Li |
ICIP | 6 |
| 2023 | Free-view Face Relighting Using a Hybrid Parametric Neural Model on a SMALL-OLAT DatasetabstractAbstract The development of neural relighting techniques has by far outpaced the rate of their corresponding training data (e.g., OLAT) generation. For example, high-quality relighting from a single portrait image still requires supervision from comprehensive datasets covering broad diversities in gender, race, complexion, and facial geometry. We present a hybrid parametric neural relighting (PN-Relighting) framework for single portrait relighting, using a much smaller OLAT dataset or SMOLAT. At the core of PN-Relighting, we employ parametric 3D faces coupled with appearance inference and implicit material modelling to enrich SMOLAT for handling in-the-wild images. Specifically, we tailor an appearance inference module to generate detailed geometry and albedo on top of the parametric face and develop a neural rendering module to first construct an implicit material representation from SMOLAT and then conduct self-supervised training on in-the-wild image datasets. Comprehensive experiments show that PN-Relighting produces comparable high-quality relighting to TotalRelighting (Pandey et al., 2021), but with a smaller dataset. It further improves shape estimation and naturally supports free-viewpoint rendering and partial skin material editing. PN-Relighting also serves as a data augmenter to produce rich OLAT datasets beyond the original capture. Youjia Wang, Taotao Zhou 0006, Kaixin Yao, Nianyi Li, Lan Xu 0003, Jingyi Yu 0001 |
Int. J. Comput. Vis. | 5 |
| 2022 | NIMBLE: a non-rigid hand model with bones and musclesabstractEmerging Metaverse applications demand reliable, accurate, and photorealistic reproductions of human hands to perform sophisticated operations as if in the physical world. While real human hand represents one of the most intricate coordination between bones, muscle, tendon, and skin, state-of-the-art techniques unanimously focus on modeling only the skeleton of the hand. In this paper, we present NIMBLE, a novel parametric hand model that includes the missing key components, bringing 3D hand model to a new level of realism. We first annotate muscles, bones and skins on the recent Magnetic Resonance Imaging hand (MRI-Hand) dataset [Li et al. 2021] and then register a volumetric template hand onto individual poses and subjects within the dataset. NIMBLE consists of 20 bones as triangular meshes, 7 muscle groups as tetrahedral meshes, and a skin mesh. Via iterative shape registration and parameter learning, it further produces shape blend shapes, pose blend shapes, and a joint regressor. We demonstrate applying NIMBLE to modeling, rendering, and visual inference tasks. By enforcing the inner bones and muscles to match anatomic and kinematic rules, NIMBLE can animate 3D hands to new poses at unprecedented realism. To model the appearance of skin, we further construct a photometric HandStage to acquire high-quality textures and normal maps to model wrinkles and palm print. Finally, NIMBLE also benefits learning-based hand pose and shape estimation by either synthesizing rich data or acting directly as a differentiable layer in the inference network. Longwen Zhang, Zesong Qiu, Yingwenqi Jiang, Nianyi Li, Yuexin Ma, Yuyao Zhang 0005, Lan Xu 0003, Jingyi Yu 0001 |
ACM Trans. Graph. | 5 |
| 2021 | Unsupervised Non-Rigid Image Distortion Removal via Grid DeformationabstractMany computer vision problems face difficulties when imaging through turbulent refractive media (e.g., air and water) due to the refraction and scattering of light. These effects cause geometric distortion that requires either handcrafted physical priors or supervised learning methods to remove. In this paper, we present a novel unsupervised network to recover the latent distortion-free image. The key idea is to model non-rigid distortions as deformable grids. Our network consists of a grid deformer that estimates the distortion field and an image generator that outputs the distortion-free image. By leveraging the positional encoding operator, we can simplify the network structure while maintaining fine spatial details in the recovered images. Our method doesn't need to be trained on labeled data and has good transferability across various turbulent image datasets with different types of distortions. Extensive experiments on both simulated and real-captured turbulent images demonstrate that our method can remove both air and water distortions without much customization. Nianyi Li, Simron Thapa, Cameron Whyte, Albert W. Reed, Suren Jayasuriya, Jinwei Ye |
ICCV | 1 |
| 2021 | Learning to Remove Refractive Distortions from Underwater ImagesabstractThe fluctuation of the water surface causes refractive distortions that severely downgrade the image of an underwater scene. Here, we present the distortion-guided network (DG-Net) for restoring distortion-free underwater images. The key idea is to use a distortion map to guide network training. The distortion map models the pixel displacement caused by water refraction. We first use a physically constrained convolutional network to estimate the distortion map from the refracted image. We then use a generative adversarial network guided by the distortion map to restore the sharp distortion-free image. Since the distortion map indicates correspondences between the distorted image and the distortion-free one, it guides the network to make better predictions. We evaluate our network on several real and synthetic underwater image datasets and show that it out-performs the state-of-the-art algorithms, especially in presence of large distortions. We also show results of complex scenarios, including outdoor swimming pool images captured by drone and indoor aquarium images taken by cellphone camera. Simron Thapa, Nianyi Li, Jinwei Ye |
ICCV | 2 |
| 2020 | Dynamic Fluid Surface Reconstruction Using Deep Neural NetworkabstractRecovering the dynamic fluid surface is a long-standing challenging problem in computer vision. Most existing image-based methods require multiple views or a dedicated imaging system. Here we present a learning-based single-image approach for 3D fluid surface reconstruction. Specifically, we design a deep neural network that estimates the depth and normal maps of a fluid surface by analyzing the refractive distortion of a reference background image. Due to the dynamic nature of fluid surfaces, our network uses recurrent layers that carry temporal information from previous frames to achieve spatio-temporally consistent reconstruction given a video input. Due to the lack of fluid data, we synthesize a large fluid dataset using physics-based fluid modeling and rendering techniques for network training and validation. Through experiments on simulated and real captured fluid images, we demonstrate that our proposed deep neural network trained on our fluid dataset can recover dynamic 3D fluid surfaces with high accuracy. Simron Thapa, Nianyi Li, Jinwei Ye |
CVPR | 2 |
| 2019 | Jittered Exposures for Light Field Super-Resolution
Nianyi Li, Scott McCloskey, Jingyi Yu 0001 |
ICIP | 1 |
| 2019 | Personalized Saliency and Its PredictionabstractNearly all existing visual saliency models by far have focused on predicting a universal saliency map across all observers. Yet psychology studies suggest that visual attention of different observers can vary significantly under specific circumstances, especially a scene is composed of multiple salient objects. To study such heterogenous visual attention pattern across observers, we first construct a personalized saliency dataset and explore correlations between visual attention, personal preferences, and image contents. Specifically, we propose to decompose a personalized saliency map (referred to as PSM) into a universal saliency map (referred to as USM) predictable by existing saliency detection models and a new discrepancy map across users that characterizes personalized saliency. We then present two solutions towards predicting such discrepancy maps, i.e., a multi-task convolutional neural network (CNN) framework and an extended CNN with Person-specific Information Encoded Filters (CNN-PIEF). Extensive experimental results demonstrate the effectiveness of our models for PSM prediction as well their generalization capability for unseen observers. Yanyu Xu 0001, Shenghua Gao, Nianyi Li, Jingyi Yu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2017 | Beyond Universal Saliency: Personalized Saliency Prediction with Multi-task CNNabstractSaliency detection is a long standing problem in computer vision. Tremendous efforts have been focused on exploring a universal saliency model across users despite their differences in gender, race, age, etc. Yet recent psychology studies suggest that saliency is highly specific than universal: individuals exhibit heterogeneous gaze patterns when viewing an identical scene containing multiple salient objects. In this paper, we first show that such heterogeneity is common and critical for reliable saliency prediction. Our study also produces the first database of personalized saliency maps (PSMs). We model PSM based on universal saliency map (USM) shared by different participants and adopt a multi-task CNN framework to estimate the discrepancy between PSM and USM. Comprehensive experiments demonstrate that our new PSM model and prediction scheme are effective and reliable. Yanyu Xu 0001, Nianyi Li, Jingyi Yu 0001, Shenghua Gao |
IJCAI | 2 |
| 2017 | Saliency Detection on Light FieldabstractExisting saliency detection approaches use images as inputs and are sensitive to foreground/background similarities, complex background textures, and occlusions. We explore the problem of using light fields as input for saliency detection. Our technique is enabled by the availability of commercial plenoptic cameras that capture the light field of a scene in a single shot. We show that the unique refocusing capability of light fields provides useful focusness, depths, and objectness cues. We further develop a new saliency detection algorithm tailored for light fields. To validate our approach, we acquire a light field database of a range of indoor and outdoor scenes and generate the ground truth saliency map. Experiments show that our saliency detection scheme can robustly handle challenging scenarios such as similar foreground and background, cluttered background, complex occlusions, etc., and achieve high accuracy and robustness. Nianyi Li, Jinwei Ye, Yu Ji 0001, Haibin Ling, Jingyi Yu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2016 | Rotational Crossed-Slit Light FieldsabstractLight fields (LFs) are image-based representation that records the radiance along all rays along every direction through every point in space. Traditionally LFs are acquired by using a 2D grid of evenly spaced pinhole cameras or by translating a pinhole camera along the 2D grid using a robot arm. In this paper, we present a novel LF sampling scheme by exploiting a special non-centric camera called the crossed-slit or XSlit camera. An XSlit camera acquires rays that simultaneously pass through two oblique slits. We show that, instead of translating the camera as in the pinhole case, we can effectively sample the LF by rotating individual or both slits while keeping the camera fixed. This leads a "fixed-location" LF acquisition scheme. We further show through theoretical analysis and experiments that the resulting XSlit LFs provide several advantages: they provide more dense spatial-angular sampling, are amenable multi-view stereo matching and volumetric reconstruction, and can synthesize unique refocusing effects. Nianyi Li, Haiting Lin, Bilin Sun, Mingyuan Zhou, Jingyi Yu 0001 |
CVPR | 1 |
| 2015 | A weighted sparse coding framework for saliency detectionabstractThere is an emerging interest on using high-dimensional datasets beyond 2D images in saliency detection. Examples include 3D data based on stereo matching and Kinect sensors and more recently 4D light field data. However, these techniques adopt very different solution frameworks, in both type of features and procedures on using them. In this paper, we present a unified saliency detection framework for handling heterogenous types of input data. Our approach builds dictionaries using data-specific features. Specifically, we first select a group of potential foreground superpixels to build a primitive saliency dictionary. We then prune the outliers in the dictionary and test on the remaining superpixels to iteratively refine the dictionary. Comprehensive experiments show that our approach universally outperforms the state-of-the-art solution on all 2D, 3D and 4D data. Nianyi Li, Bilin Sun, Jingyi Yu 0001 |
CVPR | 1 |
| 2015 | Salient object detection via background contrastabstractThis paper addresses the problem of salient object detection. We introduce a novel framework which aims to automatically identify salient regions in natural images based on two key ideas. The first one is to consider the statistical spatial distribution of saliency and non-saliency regions as two complementary processes. The second one is based on the assumption that contrast saliency with respect to background regions outperforms those with respect to entire image. Experimental results demonstrate the effectiveness of our approach over 12 state-of-the-art models. Quan Zhou 0004, Nianyi Li, Shu Cai, Longin Jan Latecki |
ICASSP | 2 |
| 2014 | Saliency Detection on Light FieldabstractExisting saliency detection approaches use images as inputs and are sensitive to foreground/background similarities, complex background textures, and occlusions. We explore the problem of using light fields as input for saliency detection. Our technique is enabled by the availability of commercial plenoptic cameras that capture the light field of a scene in a single shot. We show that the unique refocusing capability of light fields provides useful focusness, depths, and objectness cues. We further develop a new saliency detection algorithm tailored for light fields. To validate our approach, we acquire a light field database of a range of indoor and outdoor scenes and generate the ground truth saliency map. Experiments show that our saliency detection scheme can robustly handle challenging scenarios such as similar foreground and background, cluttered background, complex occlusions, etc, and achieve high accuracy and robustness. Nianyi Li, Jinwei Ye, Yu Ji 0001, Haibin Ling, Jingyi Yu 0001 |
CVPR | 1 |
| 2012 | Corner-surround Contrast for saliency detection
Quan Zhou 0004, Nianyi Li, Pan Chen 0004, Wenyu Liu 0001 |
ICPR | 2 |