EDBT 2026 Demo / reviewers in the wild / expert
Jinwei Ye
dblp:117/4793
· DBLP profile ↗
36ranked-venue papers
7as first author
18since 2021 · last 2026
0000-0001-7780-7943ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 27 · 5 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 5 first-author · 9 since 2021Systems, architecture and hardware · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Snapshot 3D Gaussian Splatting for Miniature ScenesabstractWe present a snapshot imaging technique for recovering 3D surrounding views of miniature scenes. Due to their intricacy, miniature scenes with objects sized in millimeters are difficult to reconstruct. Yet miniatures are common in life and their 3D digitalization is desirable. We design a catadioptric imaging system with a single camera and multiple pairs of planar mirrors for snapshot 3D reconstruction of miniature scenes from a dollhouse perspective. We first present an in-depth analysis on how to configure catadioptric imaging systems that use a pair of planar mirrors. We derive mirror parameters (e.g., orientation and position) by solving a viewpoint mapping problem, which aims to determine viable mirror configuration that is able to provide reflection image from a desired viewpoint. By applying the design principle, We show a full-surround snapshot catadioptric imaging system built with eight pairs of planar mirrors. Specifically, we place the mirror pairs on nested pyramid surfaces with different angles for capturing surrounding multi-view images in a single shot. This would allow 3D reconstruction of dynamic scenes. Our mirror design is customizable based on the size of the scene for optimized view coverage. We use the 3D Gaussian Splatting (3DGS) representation for scene reconstruction and novel view synthesis. We overcome the challenge posed by our sparse view input by integrating visual hull-derived depth constraint. We perform experiments on rendered synthetic images and real images of a variety of miniature scenes that are captured by our custom-built imaging system. Experimental results demonstrate that our method outperforms state-of-the-art sparse-view and calibration-free 3DGS methods. We also show novel view synthesis results on dynamic miniature scenes (e.g., a live, moving insect). Yufan Zhang 0001, Yu Ji 0001, Yu Guo 0007, Jinwei Ye |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | De-Decay: Defusing Computer Vision Model Degradation through Scalable and Actionable Human-Data AlignmentabstractComputer Vision (CV) models can become outdated after deployment as real-world data evolves, requiring intensive attention from AI engineers to address degraded performance through tasks like data relabeling to update models with new human perceptions. Interactive human-in-the-loop systems have considerable potential to enhance model-steering practices. However, such workflows reveal two challenges: (1) scalability, where labor demands increase with data size, and (2) actionability, where human insights do not readily transform into model revisions. Based on our formative study (S1) on the current challenges faced by CV professionals, we developed De-Decay, an end-to-end Human-Data Alignment system offering scalable label-less assessment and actionable insight transformation . This enables engineers to investigate degradation and auto-retrain models with AI support, such as image clustering and regeneration. Our summative study (S2) showed that De-Decay helped engineers effectively identify and address CV degradation. We discuss how future research can enhance scalability and actionability in AI evaluation systems for aligning AI behaviors with human mental models. Tong Steven Sun, Huining Feng, Jinwei Ye, Sangdoo Yun, Young-Ho Kim, Sungsoo Ray Hong |
ACM Trans. Interact. Intell. Syst. | 3 |
| 2026 | Pixel Cube: Diffusion-based Portrait Video Relighting Through Realistic Lighting ReproductionabstractWe present a diffusion-based method for relighting dynamic portrait videos with photorealism and temporal consistency. Our method is fueled by a hybrid training dataset that consists of real-captured and rendered dynamic portrait videos with diverse subject appearances, facial motions, head poses, and known lighting conditions. Specifically, we construct an LED-based lighting system for realistic lighting emulation and high-speed video relighting data acquisition. By leveraging the image priors embedded in pre-trained video diffusion models, and using per-frame high dynamic range (HDR) environment map as lighting control, we train a high-performance generative model for realistic and identity-preserving dynamic portrait video relighting. In addition to the environment map control, our model uses a synthesized background image to enable control on the camera's exposure level and color tone. Our model can produce temporally consistent relit portrait video that looks realistic and harmonious under a provided new environment and faithfully preserve the subject's expression and fine facial features, including skin tone, wrinkles, and facial hair. Our model generalizes well to unseen data, in terms of the subject appearance, motion, and lighting condition. We perform extensive experiments on relighting in-the-wild videos with various environment maps and demonstrate practical applications on portrait photography. Results show that our method achieves state-of-the-art performance in photorealism, lighting harmony, and temporal consistency. Our project page: https://yufanzhang82.github.io/PixelCube/. Yufan Zhang 0001, Yu Ji 0001, Ayo Ajiboye, Rundi Wu, Yu Guo 0007, Changxi Zheng, Jinwei Ye |
ACM Trans. Graph. | 7 |
| 2025 | Seeing A 3D World in A Grain of SandabstractWe present a snapshot imaging technique for recovering 3D surrounding views of miniature scenes. Due to their intricacy, miniature scenes with objects sized in millimeters are difficult to reconstruct, yet miniatures are common in life and their 3D digitalization is desirable. We design a catadioptric imaging system with a single camera and eight pairs of planar mirrors for snapshot 3D reconstruction from a dollhouse perspective. We place paired mirrors on nested pyramid surfaces for capturing surrounding multi-view images in a single shot. Our mirror design is customizable based on the size of the scene for optimized view coverage. We use the 3D Gaussian Splatting (3DGS) representation for scene reconstruction and novel view synthesis. We over-come the challenge posed by our sparse view input by integrating visual hull-derived depth constraint. Our method demonstrates state-of-the-art performance on a variety of synthetic and real miniature scenes. Yufan Zhang 0001, Yu Ji 0001, Yu Guo 0007, Jinwei Ye |
CVPR | 4 |
| 2025 | MedLite: A Lightweight Medical Multimodal Model Based on Knowledge Distillation and Inference Optimization
Jinwei Ye, Yaokang Wang, Weibin Kong, Gengshen Wu |
ICIC (24) | 1 |
| 2025 | LLaMed: An Efficient Adaptation Framework for Medical Large Language Models Based on Low-Rank Adaptation and Dynamic Evidence RetrievalabstractIn the medical domain, where knowledge evolves rapidly and large language models (LLMs) demand high resources, we propose LLaMed, an efficient adaptation framework based on low-rank adaptation (LoRA) and dynamic evidence retrieval. Using LLaMA-2 7B as the base model, LLaMed freezes most parameters and fine-tunes only 0.1 % for medical knowledge injection. It features a dual-engine evidence augmentation module fusing BM25 for fast recall and PubMedBERT for semantic re-ranking, with real-time API integration (e.g., ClinicalTrials.gov) for timely updates. Additionally, it pioneers a PPObased adaptive feedback mechanism-the first expert-driven reinforcement learning in medical RAG-for dynamic output calibration and knowledge obsolescence prevention. On MedQAUSMLE, LLaMed achieves 68.2 % accuracy (1.6 % below full fine-tuning) with 89 % reduced VRAM (15.7 GB vs. 192.1 GB) and expert consistency$\mathbf{K} \boldsymbol{=} \mathbf{0. 8 5}$. Compared to methods like MedCoTRAG and efficient medical LLMs, LLaMed excels in parameter efficiency (0.1 % vs.$3-100 \%$), real-time capability, and robustness, supporting resource-constrained clinical decisions. Weibin Kong, Jinwei Ye, Chonglin Zhao, Yunqing Ma, Yanrong Chen |
ICPADS | 2 |
| 2025 | M2P2: A Multi-Modal Passive Perception Dataset for Off-Road Mobility in Extreme Low-Light ConditionsabstractLong-duration, off-road, autonomous missions require robots to continuously perceive their surroundings regardless of the ambient lighting conditions. Most existing autonomy systems heavily rely on active sensing, e.g., LiDAR, RADAR, and Time-of-Flight sensors, or use (stereo) visible light imaging sensors, e.g., color cameras, to perceive environment geometry and semantics. In scenarios where fully passive perception is required and lighting conditions are degraded to an extent that visible light cameras fail to perceive, most downstream mobility tasks such as obstacle avoidance become impossible. To address such a challenge, this paper presents a Multi-Modal Passive Perception dataset, M2P2, to enable off-road mobility in low-light to no-light conditions. We design a multi-modal sensor suite including thermal, event, and stereo RGB cameras, GPS, two Inertia Measurement Units (IMUs), as well as a high-resolution LiDAR for ground truth, with a multi-sensor calibration procedure that can efficiently transform multi-modal perceptual streams into a common coordinate system. Our 10-hour, 32 km dataset also includes mobility data such as robot odometry and actions and covers well-lit, low-light, and no-light conditions, along with paved, on-trail, and off-trail terrain. Our results demonstrate that off-road mobility and scene understanding under degraded visual environments is possible through only passive perception in extreme low-light conditions. The project website can be found at https://cs.gmu.edu/˜xiao/Research/M2P2/. Aniket Datar, Anuj Pokhrel, Mohammad Nazeri, Madhan B. Rao, Harsh Rangwala, Chenhui Pan, Yufan Zhang 0001, Andre Harrison, Maggie B. Wigness, Philip R. Osteen, Jinwei Ye, Xuesu Xiao |
IROS | 11 |
| 2025 | DAATSim: Depth-Aware Atmospheric Turbulence Simulation for Fast Image RenderingabstractAbstract Simulating the effects of atmospheric turbulence for imaging systems operating over long distances is a significant challenge for optical and computer graphics models. Physically‐based ray tracing over kilometers of distance is difficult due to the need to define a spatio‐temporal volume of varying refractive index. Even if such a volume can be defined, Monte Carlo rendering approximations for light refraction through the environment would not yield real‐time solutions needed for video game engines or online dataset augmentation for machine learning. While existing simulators based on procedurally‐generated noise or textures have been proposed in these settings, these simulators often neglect the significant impact of scene depth, leading to unrealistic degradations for scenes with substantial foreground‐background separation. This paper introduces a novel, physically‐based atmospheric turbulence simulator that explicitly models depth‐dependent effects while rendering frames at interactive/near real‐time (> 10 FPS) rates for image resolutions up to 1024 × 1024 (real‐time 35 FPS at 256× 256 resolution with depth or 512×512 at 33 FPS without depth). Our hybrid approach combines spatially‐varying wavefront aberrations using Zernike polynomials with pixel‐wise depth modulation of both blur (via Point Spread Function interpolation) and geometric distortion or tilt. Our approach includes a novel fusion technique that integrates complementary strengths of leading monocular depth estimators to generate metrically accurate depth maps with enhanced edge fidelity. DAATSim is implemented efficiently on GPUs using Py‐Torch incorporating optimizations like mixed‐precision computation and caching to achieve efficient performance. We present quantitative and qualitative validation demonstrating the simulator's physical plausibility for generating turbulent video. DAAT‐Sim is made publicly available and open‐source to the community: https://github.com/Riponcs/DAATSim . Ripon K. Saha, Yufan Zhang 0001, Jinwei Ye, Suren Jayasuriya |
Comput. Graph. Forum | 3 |
| 2025 | Textureless Deformable Object Tracking With Invisible MarkersabstractTracking and reconstructing deformable objects with little texture is challenging due to the lack of features. Here we introduce "invisible markers" for accurate and robust correspondence matching and tracking. Our markers are visible only under ultraviolet (UV) light. We build a novel imaging system for capturing videos of deformed objects under their original untouched appearance (which may have little texture) and, simultaneously, with our markers. We develop an algorithm that first establishes accurate correspondences using video frames with markers, and then transfers them to the untouched views as ground-truth labels. In this way, we are able to generate high-quality labeled data for training learning-based algorithms. We contribute a large real-world dataset, DOT, for tracking deformable objects with little or no texture. Our dataset has about one million video frames of various types of deformable objects. We provide ground truth tracked correspondences in both 2D and 3D. We benchmark state-of-the-art methods on optical flow and deformable object reconstruction using our dataset, which poses great challenges. By training on DOT, their performance significantly improves, not only on our dataset, but also on other unseen data. Yu Guo 0007, Yubei Tu, Yu Ji 0001, Jinwei Ye, Changxi Zheng |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2024 | Turb-Seg-Res: A Segment-then-Restore Pipeline for Dynamic Videos with Atmospheric TurbulenceabstractTackling image degradation due to atmospheric turbu-lence, particularly in dynamic environments, remains a challenge for long-range imaging systems. Existing techniques have been primarily designed for static scenes or scenes with small motion. This paper presents the first segment-then-restore pipeline for restoring the videos of dy-namic scenes in turbulent environments. We leverage mean optical flow with an unsupervised motion segmentation method to separate dynamic and static scene components prior to restoration. After camera shake compensation and segmentation, we introduce foreground/background en-hancement leveraging the statistics of turbulence strength and a transformer model trained on a novel noise-based procedural turbulence generator for fast dataset augmen-tation. Benchmarked against existing restoration meth-ods, our approach restores most of the geometric distortion and enhances the sharpness of videos. We make our code, simulator, and data publicly available to ad-vance the field of video restoration from turbulence: riponcs.github.io/TurbSegRes Ripon K. Saha, Dehao Qin, Nianyi Li, Jinwei Ye, Suren Jayasuriya |
CVPR | 4 |
| 2024 | Unsupervised Moving Object Segmentation with Atmospheric Turbulence
Dehao Qin, Ripon K. Saha, Woojeh Chung, Suren Jayasuriya, Jinwei Ye, Nianyi Li |
ECCV (6) | 5 |
| 2024 | Polarimetric Helmholtz StereopsisabstractHelmholtz stereopsis (HS) exploits the reciprocity principle of light propagation (i.e., the Helmholtz reciprocity) for 3D reconstruction of surfaces with arbitrary reflectance. In this paper, we present the polarimetric Helmholtz stereopsis (polar-HS), which extends the classical HS by considering the polarization state of light in the reciprocal paths. With the additional phase information from polarization, polar-HS requires only one reciprocal image pair. We derive the reciprocity relationship of Mueller matrix and formulate new reciprocity constraint that takes polarization state into account. We also utilize polarimetric constraints and extend them to the case of perspective projection. For the recovery of surface depths and normals, we incorporate reciprocity constraint with diffuse/specular polarimetric constraints in a unified optimization framework. For depth estimation, we further propose to utilize the consistency of diffuse angle of polarization. For normal estimation, we develop a normal refinement strategy based on degree of linear polarization. Using a hardware prototype, we show that our approach produces high-quality 3D reconstruction for different types of surfaces, ranging from diffuse to highly specular. Yuqi Ding, Yu Ji 0001, Mingyuan Zhou, Sing Bing Kang, Jinwei Ye |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | Full-Volume 3D Fluid Flow Reconstruction With Light Field PIVabstractParticle Imaging Velocimetry (PIV) is a classical method that estimates fluid flow by analyzing the motion of injected particles. To reconstruct and track the swirling particles is a difficult computer vision problem, as the particles are dense in the fluid volume and have similar appearances. Further, tracking a large number of particles is particularly challenging due to heavy occlusion. Here we present a low-cost PIV solution that uses compact lenslet-based light field cameras as imaging device. We develop novel optimization algorithms for dense particle 3D reconstruction and tracking. As a single light field camera has limited capacity in resolving depth (z-dimension measurement), the resolution of 3D reconstruction on the x-y plane is much higher than along the z-axis. To compensate for the imbalanced resolution in 3D, we use two light field cameras positioned at an orthogonal angle to capture particle images. In this way, we can achieve high-resolution 3D particle reconstruction in the full fluid volume. For each time frame, we first estimate particle depths under a single viewpoint by exploiting the focal stack symmetry of light field. We then fuse the recovered 3D particles in two views by solving a linear assignment problem (LAP). Specifically, we propose an anisotropic point-to-ray distance as matching cost to handle the resolution mismatch. Finally, given a sequence of 3D particle reconstructions over time, we recover the full-volume 3D fluid flow with a physically-constrained optical flow, which enforces local motion rigidity and fluid incompressibility. We perform comprehensive experiments on synthetic and real data for ablation and evaluation. We show that our method recovers full-volume 3D fluid flows of various types. Two-view reconstruction results achieves higher accuracy than those with one view only. Yuqi Ding, Zhong Li 0007, Yu Ji 0001, Jingyi Yu 0001, Jinwei Ye |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2022 | Polar-Photometric Stereo Under Natural IlluminationabstractWe present a 3D shape reconstruction method that leverages both photometric and polarimetric cues. Unlike many active methods that require controlled lighting condition, our method can be used under unknown and uncontrolled natural illumination (both indoor and outdoor). We use two circularly polarized spotlights to boost the polarization cues corrupted by the environment lighting, as well as to provide photometric cues. We solve surface normals with two polarization images by combining the polarimetric and photometric constraints. To mitigate the effect of uncontrolled environment light in photometric constraints, we es-timate a lighting proxy map and iteratively refine the normal and lighting estimation. We perform experiments under various natural illumination conditions and compare our results with state-of-the-arts photometric stereo and shape from polarization methods. Our method achieves good accuracy and can be used in flexible environment. Yuqi Ding, Yu Ji 0001, Jinwei Ye |
3DV | 3 |
| 2021 | Polarimetric Helmholtz StereopsisabstractHelmholtz stereopsis (HS) exploits the reciprocity principle of light propagation (i.e., the Helmholtz reciprocity) for 3D reconstruction of surfaces with arbitrary reflectance. In this paper, we present the polarimetric Helmholtz stereopsis (polar-HS), which extends the classical HS by considering the polarization state of light in the reciprocal paths. With the additional phase information from polarization, polar-HS requires only one reciprocal image pair. We formulate new reciprocity and diffuse/specular polarimetric constraints to recover surface depths and normals using an optimization framework. Using a hardware prototype, we show that our approach produces high-quality 3D reconstruction for different types of surfaces, ranging from diffuse to highly specular. Yuqi Ding, Yu Ji 0001, Mingyuan Zhou, Sing Bing Kang, Jinwei Ye |
ICCV | 5 |
| 2021 | Unsupervised Non-Rigid Image Distortion Removal via Grid DeformationabstractMany computer vision problems face difficulties when imaging through turbulent refractive media (e.g., air and water) due to the refraction and scattering of light. These effects cause geometric distortion that requires either handcrafted physical priors or supervised learning methods to remove. In this paper, we present a novel unsupervised network to recover the latent distortion-free image. The key idea is to model non-rigid distortions as deformable grids. Our network consists of a grid deformer that estimates the distortion field and an image generator that outputs the distortion-free image. By leveraging the positional encoding operator, we can simplify the network structure while maintaining fine spatial details in the recovered images. Our method doesn't need to be trained on labeled data and has good transferability across various turbulent image datasets with different types of distortions. Extensive experiments on both simulated and real-captured turbulent images demonstrate that our method can remove both air and water distortions without much customization. Nianyi Li, Simron Thapa, Cameron Whyte, Albert W. Reed, Suren Jayasuriya, Jinwei Ye |
ICCV | 6 |
| 2021 | Learning to Remove Refractive Distortions from Underwater ImagesabstractThe fluctuation of the water surface causes refractive distortions that severely downgrade the image of an underwater scene. Here, we present the distortion-guided network (DG-Net) for restoring distortion-free underwater images. The key idea is to use a distortion map to guide network training. The distortion map models the pixel displacement caused by water refraction. We first use a physically constrained convolutional network to estimate the distortion map from the refracted image. We then use a generative adversarial network guided by the distortion map to restore the sharp distortion-free image. Since the distortion map indicates correspondences between the distorted image and the distortion-free one, it guides the network to make better predictions. We evaluate our network on several real and synthetic underwater image datasets and show that it out-performs the state-of-the-art algorithms, especially in presence of large distortions. We also show results of complex scenarios, including outdoor swimming pool images captured by drone and indoor aquarium images taken by cellphone camera. Simron Thapa, Nianyi Li, Jinwei Ye |
ICCV | 3 |
| 2021 | Structure From Motion on XSlit CamerasabstractWe present a structure-from-motion (SfM) framework based on a special type of multi-perspective camera called the cross-slit or XSlit camera. Traditional perspective camera based SfM suffers from the scale ambiguity which is inherent to the pinhole camera geometry. In contrast, an XSlit camera captures rays passing through two oblique lines in 3D space and we show such ray geometry directly resolves the scale ambiguity when employed for SfM. To accommodate the XSlit cameras, we develop tailored feature matching, camera pose estimation, triangulation, and bundle adjustment techniques. Specifically, we devise a SIFT feature variant using non-uniform Gaussian kernels to handle the distortions in XSlit images for reliable feature matching. Moreover, we demonstrate that the XSlit camera exhibits ambiguities in pose estimation process which can not be handled by existing work. Consequently, we propose a 14 point algorithm to properly handle the XSlit degeneracy and estimate the relative pose between XSlit cameras from feature correspondences. We further exploit the unique depth-dependent aspect ratio (DDAR) property to improve the bundle adjustment for the XSlit camera. Synthetic and real experiments demonstrate that the proposed XSlit SfM can conduct reliable and high fidelity 3D reconstruction at an absolute scale. Wei Yang 0034, Yingliang Zhang, Jinwei Ye, Yu Ji 0001, Zhong Li 0007, Mingyuan Zhou, Jingyi Yu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2020 | Dynamic Fluid Surface Reconstruction Using Deep Neural NetworkabstractRecovering the dynamic fluid surface is a long-standing challenging problem in computer vision. Most existing image-based methods require multiple views or a dedicated imaging system. Here we present a learning-based single-image approach for 3D fluid surface reconstruction. Specifically, we design a deep neural network that estimates the depth and normal maps of a fluid surface by analyzing the refractive distortion of a reference background image. Due to the dynamic nature of fluid surfaces, our network uses recurrent layers that carry temporal information from previous frames to achieve spatio-temporally consistent reconstruction given a video input. Due to the lack of fluid data, we synthesize a large fluid dataset using physics-based fluid modeling and rendering techniques for network training and validation. Through experiments on simulated and real captured fluid images, we demonstrate that our proposed deep neural network trained on our fluid dataset can recover dynamic 3D fluid surfaces with high accuracy. Simron Thapa, Nianyi Li, Jinwei Ye |
CVPR | 3 |
| 2020 | 3D Fluid Flow Reconstruction Using Compact Light Field PIV
Zhong Li 0007, Yu Ji 0001, Jingyi Yu 0001, Jinwei Ye |
ECCV (16) | 4 |
| 2020 | Shape and Reflectance Reconstruction Using Concentric Multi-Spectral Light FieldabstractRecovering the shape and reflectance of non-Lambertian surfaces remains a challenging problem in computer vision since the view-dependent appearance invalidates traditional photo-consistency constraint. In this paper, we introduce a novel concentric multi-spectral light field (CMSLF) design that is able to recover the shape and reflectance of surfaces of various materials in one shot. Our CMSLF system consists of an array of cameras arranged on concentric circles where each ring captures a specific spectrum. Coupled with a multi-spectral ring light, we are able to sample viewpoint and lighting variations in a single shot via spectral multiplexing. We further show that our concentric camera and light source setting results in a unique single-peak pattern in specularity variations across viewpoints. This property enables robust depth estimation for specular points. To estimate depth and multi-spectral reflectance map, we formulate a physics-based reflectance model for the CMSLF under the surface camera (S-Cam) representation. Extensive synthetic and real experiments show that our method outperforms the state-of-the-art shape reconstruction methods, especially for non-Lambertian surfaces. Mingyuan Zhou, Yuqi Ding, Yu Ji 0001, S. Susan Young, Jingyi Yu 0001, Jinwei Ye |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2019 | Mirror Surface Reconstruction Using Polarization FieldabstractMirror surfaces are notoriously difficult to reconstruct. In this paper, we present a novel computational imaging approach for reconstructing complex mirror surfaces using a dense illumination field with angularly varying polarization states, which we call the polarization field. Specifically, we generate the polarization field using a commercial LCD with the top polarizer removed. We mathematically model the liquid crystals as polarization rotators using Jones calculus and show that the rotated polarization states of outgoing rays encode angular information (e.g., ray directions). To model reflection under the polarization field, we derive a reflection image formation model based on the Fresnel's equations and estimate incident ray positions and directions by coding the polarization field. Finally, we triangulate the incident rays with the camera rays to recover normals/depths of the mirror surface. Comprehensive simulations and real experiments demonstrate the effectiveness of our approach. Yu Ji 0001, Jingyi Yu 0001, Jinwei Ye |
ICCP | 4 |
| 2019 | Content Aware Image Pre-CompensationabstractThe goal of image pre-compensation is to process an image such that after being convolved with a known kernel, will appear close to the sharp reference image. In a practical setting, the pre-compensated image has significantly higher dynamic range than the latent image. As a result, some form of tone mapping is needed. In this paper, we show how global tone mapping functions affect contrast and ringing in image pre-compensation. We further enhance contrast and reduce ringing by considering the visual saliency. Specifically, we prioritize contrast preservation in salient regions while tolerating more blurriness elsewhere. For quantitative analysis, we design new metrics to measure the contrast of an image with ringing. Specifically, we set out to find its "equivalent ringing-free" image that matches its intensity histogram and uses its contrast as the measure. We illustrate our approach on projector defocus compensation and visual acuity enhancement. Compared with the state-of-the-art, our approach significantly improves the contrast. We also perform user studies to demonstrate that our method can effectively improve the viewing experience for users with impaired vision. Jinwei Ye, Yu Ji 0001, Mingyuan Zhou, Sing Bing Kang, Jingyi Yu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2018 | Learning to Dodge A Bullet: Concyclic View Morphing via Deep Learning
Ruiyang Liu, Yu Ji 0001, Jinwei Ye, Jingyi Yu 0001 |
ECCV (14) | 4 |
| 2017 | Robust 3D Human Motion Reconstruction via Dynamic Template ConstructionabstractIn multi-view human body capture systems, the recovered 3D geometry or even the acquired imagery data can be heavily corrupted due to occlusions, noise, limited fieldof- view, etc. Direct estimation of 3D pose, body shape or motion on these low-quality data has been traditionally challenging.In this paper, we present a graph-based non-rigid shape registration framework that can simultaneously recover 3D human body geometry and estimate pose/motion at high fidelity.Our approach first generates a global full-body template by registering all poses in the acquired motion sequence.We then construct a deformable graph by utilizing the rigid components in the global template. We directly warp the global template graph back to each motion frame in order to fill in missing geometry. Specifically,we combine local rigidity and temporal coherence constraints to maintain geometry and motion consistencies. Comprehensive experiments on various scenes show that our method is accurate and robust even in the presence of drastic motions. Zhong Li 0007, Yu Ji 0001, Wei Yang 0034, Jinwei Ye, Jingyi Yu 0001 |
3DV | 4 |
| 2017 | Saliency Detection on Light FieldabstractExisting saliency detection approaches use images as inputs and are sensitive to foreground/background similarities, complex background textures, and occlusions. We explore the problem of using light fields as input for saliency detection. Our technique is enabled by the availability of commercial plenoptic cameras that capture the light field of a scene in a single shot. We show that the unique refocusing capability of light fields provides useful focusness, depths, and objectness cues. We further develop a new saliency detection algorithm tailored for light fields. To validate our approach, we acquire a light field database of a range of indoor and outdoor scenes and generate the ground truth saliency map. Experiments show that our saliency detection scheme can robustly handle challenging scenarios such as similar foreground and background, cluttered background, complex occlusions, etc., and achieve high accuracy and robustness. Nianyi Li, Jinwei Ye, Yu Ji 0001, Haibin Ling, Jingyi Yu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2016 | 3D reconstruction of mirror-type objects using efficient ray codingabstractMirror-type specular objects are difficult to reconstruct: they do not possess their own appearance and the reflections from environment are view-dependent. In this paper, we present a novel computational imaging solution for reconstructing the mirror-type specular objects. Specifically, we adopt a two-layer liquid crystal display (LCD) setup to encode the illumination directions. We devise an efficient ray coding scheme by only considering the useful rays. To recover the mirror-type surface, we derive a normal integration scheme under the perspective camera model. Since the resulting surface is determined up to a scale, we develop a single view approach to resolve the scale ambiguity. To acquire the object surface as completely as possible, we further develop a multiple-surface fusion algorithm to combine the surfaces recovered from different viewpoints. Both synthetic and real experiments demonstrate that our approach is reliable on recovering small to medium scale mirror-type objects. Siu-Kei Tin, Jinwei Ye, Mahdi Nezamabadi |
ICCP | 2 |
| 2014 | Image Pre-compensation: Balancing Contrast and RingingabstractThe goal of image pre-compensation is to process an image such that after being convolved with a known kernel, will appear close to the sharp reference image. In a practical setting, the pre-compensated image has significantly higher dynamic range than the latent image. As a result, some form of tone mapping is needed. In this paper, we show how global tone mapping functions affect contrast and ringing in image pre-compensation. In particular, we show that linear tone mapping eliminates ringing but incurs severe contrast loss, while non-linear tone mapping functions such as Gamma curves slightly enhances contrast but introduces ringing. To enable quantitative analysis, we design new metrics to measure the contrast of an image with ringing. Specifically, we set out to find its "equivalent ringing-free" image that matches its intensity histogram and uses its contrast as the measure. We illustrate our approach on projector defocus compensation and visual acuity enhancement. Compared with the state-of-the-art, our approach significantly improves the contrast. We believe our technique is the first to analytically trade-off between contrast and ringing. Yu Ji 0001, Jinwei Ye, Sing Bing Kang, Jingyi Yu 0001 |
CVPR | 2 |
| 2014 | Saliency Detection on Light FieldabstractExisting saliency detection approaches use images as inputs and are sensitive to foreground/background similarities, complex background textures, and occlusions. We explore the problem of using light fields as input for saliency detection. Our technique is enabled by the availability of commercial plenoptic cameras that capture the light field of a scene in a single shot. We show that the unique refocusing capability of light fields provides useful focusness, depths, and objectness cues. We further develop a new saliency detection algorithm tailored for light fields. To validate our approach, we acquire a light field database of a range of indoor and outdoor scenes and generate the ground truth saliency map. Experiments show that our saliency detection scheme can robustly handle challenging scenarios such as similar foreground and background, cluttered background, complex occlusions, etc, and achieve high accuracy and robustness. Nianyi Li, Jinwei Ye, Yu Ji 0001, Haibin Ling, Jingyi Yu 0001 |
CVPR | 2 |
| 2014 | Coplanar Common Points in Non-centric Cameras
Wei Yang 0034, Yu Ji 0001, Jinwei Ye, S. Susan Young, Jingyi Yu 0001 |
ECCV (1) | 3 |
| 2014 | Depth-of-Field and Coded Aperture Imaging on XSlit Lens
Jinwei Ye, Yu Ji 0001, Wei Yang 0034, Jingyi Yu 0001 |
ECCV (3) | 1 |
| 2014 | Ray geometry in non-pinhole cameras: a survey
Jinwei Ye, Jingyi Yu 0001 |
Vis. Comput. | 1 |
| 2013 | Reconstructing Gas Flows Using Light-Path ApproximationabstractTransparent gas flows are difficult to reconstruct: the refractive index field (RIF) within the gas volume is uneven and rapidly evolving, and correspondence matching under distortions is challenging. We present a novel computational imaging solution by exploiting the light field probe (LF-Probe). A LF-probe resembles a view-dependent pattern where each pixel on the pattern maps to a unique ray. By observing the LF-probe through the gas flow, we acquire a dense set of ray-ray correspondences and then reconstruct their light paths. To recover the RIF, we use Fermat's Principle to correlate each light path with the RIF via a Partial Differential Equation (PDE). We then develop an iterative optimization scheme to solve for all light-path PDEs in conjunction. Specifically, we initialize the light paths by fitting Hermite splines to ray-ray correspondences, discretize their PDEs onto voxels, and solve a large, over-determined PDE system for the RIF. The RIF can then be used to refine the light paths. Finally, we alternate the RIF and light-path estimations to improve the reconstruction. Experiments on synthetic and real data show that our approach can reliably reconstruct small to medium scale gas flows. In particular, when the flow is acquired by a small number of cameras, the use of ray-ray correspondences can greatly improve the reconstruction. Yu Ji 0001, Jinwei Ye, Jingyi Yu 0001 |
CVPR | 2 |
| 2013 | Manhattan Scene Understanding via XSlit ImagingabstractA Manhattan World (MW) [3] is composed of planar surfaces and parallel lines aligned with three mutually orthogonal principal axes. Traditional MW understanding algorithms rely on geometry priors such as the vanishing points and reference (ground) planes for grouping coplanar structures. In this paper, we present a novel single-image MW reconstruction algorithm from the perspective of non-pinhole cameras. We show that by acquiring the MW using an XSlit camera, we can instantly resolve co planarity ambiguities. Specifically, we prove that parallel 3D lines map to 2D curves in an XSlit image and they converge at an XSlit Vanishing Point (XVP). In addition, if the lines are coplanar, their curved images will intersect at a second common pixel that we call Coplanar Common Point (CCP). CCP is a unique image feature in XSlit cameras that does not exist in pinholes. We present a comprehensive theory to analyze XVPs and CCPs in a MW scene and study how to recover 3D geometry in a complex MW scene from XVPs and CCPs. Finally, we build a prototype XSlit camera by using two layers of cylindrical lenses. Experimental results on both synthetic and real data show that our new XSlit-camera-based solution provides an effective and reliable solution for MW understanding. Jinwei Ye, Yu Ji 0001, Jingyi Yu 0001 |
CVPR | 1 |
| 2013 | A Rotational Stereo Model Based on XSlit ImagingabstractTraditional stereo matching assumes perspective viewing cameras under a translational motion: the second camera is translated away from the first one to create parallax. In this paper, we investigate a different, rotational stereo model on a special multi-perspective camera, the XSlit camera. We show that rotational XSlit (R-XSlit) stereo can be effectively created by fixing the sensor and slit locations but switching the two slits' directions. We first derive the epipolar geometry of R-XSlit in the 4D light field ray space. Our derivation leads to a simple but effective scheme for locating corresponding epipolar "curves". To conduct stereo matching, we further derive a new disparity term in our model and develop a patch-based graph-cut solution. To validate our theory, we assemble an XSlit lens by using a pair of cylindrical lenses coupled with slit-shaped apertures. The XSlit lens can be mounted on commodity cameras where the slit directions are adjustable to form desirable R-XSlit pairs. We show through experiments that R-XSlit provides a potentially advantageous imaging system for conducting fixed-location, dynamic baseline stereo. Jinwei Ye, Yu Ji 0001, Jingyi Yu 0001 |
ICCV | 1 |
| 2012 | Angular domain reconstruction of dynamic 3D fluid surfacesabstractWe present a novel and simple computational imaging solution to robustly and accurately recover 3D dynamic fluid surfaces. Traditional specular surface reconstruction schemes place special patterns (checkerboard or color patterns) beneath the fluid surface to establish point-pixel correspondences. However, point-pixel correspondences alone are insufficient to recover surface normal or height and they rely on additional constraints to resolve the ambiguity. In this paper, we exploit using Bokode - a computational optical device that emulates a pinhole projector - for capturing ray-ray correspondences which can then be used to directly recover the surface normals. We further develop a robust feature matching algorithm based on the Active-Appearance Model to robustly establishing ray-ray correspondences. Our solution results in an angularly sampled normal field and we derive a new angular-domain surface integration scheme to recover the surface from the normal fields. Specifically, we reformulate the problem as an over-constrained linear system under spherical coordinate and solve it using Singular Value Decomposition. Experiments results on real and synthetic surfaces demonstrate that our approach is robust and accurate, and is easier to implement than state-of-the-art multi-camera based approaches. Jinwei Ye, Yu Ji 0001, Feng Li 0005, Jingyi Yu 0001 |
CVPR | 1 |