Yanghai Tsin

dblp:17/4227 · DBLP profile ↗
← Back
30ranked-venue papers
10as first author
7since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 26 · 9 first-author · 6 since 2021Artificial intelligence and machine learning · 25 · 8 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author
YearPublicationVenuePosition
2025 Matrix3D: Large Photogrammetry Model All-in-One
abstract
We present Matrix3D, a unified model that performs several photogrammetry subtasks, including pose estimation, depth prediction, and novel view synthesis using just the same model. Matrix3D utilizes a multi-modal diffusion transformer (DiT) to integrate transformations across several modalities, such as images, camera parameters, and depth maps. The key to Matrix3D’s large-scale multi-modal training lies in the incorporation of a mask learning strategy. This enables full-modality model training even with partially complete data, such as bi-modality data of image-pose and image-depth pairs, thus significantly increases the pool of available training data. Matrix3D demonstrates state-of-the-art performance in pose estimation and novel view synthesis tasks. Additionally, it offers fine-grained control through multi-round interactions, making it an innovative tool for 3D content creation. Project page: https://nju-3dv.github.io/projects/matrix3d.
Yuanxun Lu, Jingyang Zhang, Tian Fang, Jean-Daniel Nahmias, Yanghai Tsin, Long Quan, Xun Cao, Yao Yao 0008, Shiwei Li 0001
CVPR5
2024 Direct2.5: Diverse Text-to-3D Generation via Multi-view 2.5D Diffusion
abstract
Recent advances in generative AI have unveiled significant potential for the creation of 3D content. However, current methods either apply a pre-trained 2D diffusion model with the time-consuming score distillation sampling (SDS), or a direct 3D diffusion model trained on limited 3D data losing generation diversity. In this work, we approach the problem by employing a multi-view 2.5D diffusion fine-tuned from a pre-trained 2D diffusion model. The multi-view 2.5D diffusion directly models the structural distribution of 3D data, while still maintaining the strong generalization ability of the original 2D diffusion model, filling the gap between 2D diffusion-based and direct 3D diffusion-based methods for 3D content generation. During inference, multi-view normal maps are generated using the 2.5D diffusion, and a novel differentiable rasterization scheme is introduced to fuse the almost consistent multi-view normal maps into a consistent 3D model. We further design a normal-conditioned multi-view image generation module for fast appearance generation given the 3D geometry. Our method is a one-pass diffusion process and does not require any SDS optimization as post-processing. We demonstrate through extensive experiments that, our direct 2.5D generation with the specially-designed fusion scheme can achieve diverse, mode-seeking-free, and high-fidelity 3D content generation in only 10 seconds. Project page: https://nju-3dv.github.io/projects/direct25.
Yuanxun Lu, Jingyang Zhang, Shiwei Li 0001, Tian Fang, David McKinnon, Yanghai Tsin, Long Quan, Xun Cao, Yao Yao 0008
CVPR6
2024 JointNet: Extending Text-to-Image Diffusion for Dense Distribution Modeling
abstract
We introduce JointNet, a novel neural network architecture for modeling the joint distribution of images and an additional dense modality (e.g., depth maps). JointNet is extended from a pre-trained text-to-image diffusion model, where a copy of the original network is created for the new dense modality branch and is densely connected with the RGB branch. The RGB branch is locked during network fine-tuning, which enables efficient learning of the new modality distribution while maintaining the strong generalization ability of the large-scale pre-trained diffusion model. We demonstrate the effectiveness of JointNet by using the RGB-D diffusion as an example and through extensive experiments, showcasing its applicability in a variety of applications, including joint RGB-D generation, dense depth prediction, depth-conditioned image generation, and high-resolution 3D panorama generation.
Jingyang Zhang, Shiwei Li 0001, Yuanxun Lu, Tian Fang, David McKinnon, Yanghai Tsin, Long Quan, Yao Yao 0008
ICLR6
2023 NeILF++: Inter-Reflectable Light Fields for Geometry and Material Estimation
abstract
We present a novel differentiable rendering framework for joint geometry, material, and lighting estimation from multi-view images. In contrast to previous methods which assume a simplified environment map or co-located flashlights, in this work, we formulate the lighting of a static scene as one neural incident light field (NeILF) and one outgoing neural radiance field (NeRF). The key insight of the proposed method is the union of the incident and outgoing light fields through physically-based rendering and inter-reflections between surfaces, making it possible to disentangle the scene geometry, material, and lighting from image observations in a physically-based manner. The proposed incident light and inter-reflection framework can be easily applied to other NeRF systems. We show that our method can not only decompose the outgoing radiance into incident lights and surface materials, but also serve as a surface refinement module that further improves the reconstruction detail of the neural surface. We demonstrate on several datasets that the proposed method is able to achieve state-of-the-art results in terms of geometry reconstruction quality, material estimation accuracy, and the fidelity of novel view rendering.
Jingyang Zhang, Yao Yao 0008, Shiwei Li 0001, Tian Fang, David McKinnon, Yanghai Tsin, Long Quan
ICCV7
2022 Critical Regularizations for Neural Surface Reconstruction in the Wild
abstract
Neural implicit functions have recently shown promising results on surface reconstructions from multiple views. However, current methods still suffer from excessive time complexity and poor robustness when reconstructing unbounded or complex scenes. In this paper, we present RegSDF, which shows that proper point cloud supervisions and geometry regularizations are sufficient to produce high-quality and robust reconstruction results. Specifically, RegSDF takes an additional oriented point cloud as input, and optimizes a signed distance field and a surface light field within a differentiable rendering framework. We also introduce the two critical regularizations for this optimization. The first one is the Hessian regularization that smoothly diffuses the signed distance values to the entire distance field given noisy and incomplete input. And the second one is the minimal surface regularization that compactly interpolates and extrapolates the missing geometry. Extensive experiments are conducted on DTU, Blended-MVS, and Tanks and Temples datasets. Compared with recent neural surface reconstruction approaches, RegSDF is able to reconstruct surfaces with fine details even for open scenes with complex topologies and unstructured camera trajectories.
Jingyang Zhang, Yao Yao 0008, Shiwei Li 0001, Tian Fang, David McKinnon, Yanghai Tsin, Long Quan
CVPR6
2022 ASpanFormer: Detector-Free Image Matching with Adaptive Span Transformer
Zixin Luo, Lei Zhou 0011, Yurun Tian, Mingmin Zhen, Tian Fang, David McKinnon, Yanghai Tsin, Long Quan
ECCV (32)8
2022 NeILF: Neural Incident Light Field for Physically-based Material Estimation
Yao Yao 0008, Jingyang Zhang, Yihang Qu, Tian Fang, David McKinnon, Yanghai Tsin, Long Quan
ECCV (31)7
2010 Image-Based Respiratory Motion Compensation for Fluoroscopic Coronary Roadmapping
Yanghai Tsin, Hari Sundar, Frank Sauer
MICCAI (3)2
2010 Layout Consistent Segmentation of 3-D Meshes via Conditional Random Fields and Spatial Ordering Constraints
Alexander Zouhar, Sajjad Baloch, Yanghai Tsin, Tong Fang, Siegfried Fuchs
MICCAI (3)3
2009 Explicit 3D Modeling for Vehicle Monitoring in Non-overlapping Cameras
abstract
Vehicles are indispensable in modern life. The capability of monitoring them over a long range can play significant roles in many surveillance applications. However, due to high mobility of vehicles, tracking them is difficult and we need to utilize a large network of cameras and reason on discrete sets of observations made from non-overlapping cameras. In this paper, we introduce enabling techniques for such a surveillance need. Specifically, we build explicit 3D models and use them for vehicle signature extraction and matching. The algorithm uses a single active shape model (ASM) for all consumer vehicles. After detecting presence of a vehicle, \eg, by background subtraction, our algorithm then reconstructs a texture mapped 3D model. 3D car models enable us to monitor vehicles in many novel ways otherwise impossible. Two use cases are provided.
Yanghai Tsin, Yakup Genc, Visvanathan Ramesh
AVSS1
2009 Globally optimal affine epipolar geometry from apparent contours
abstract
We study the problem of estimating the epipolar geometry from apparent contours of smooth curved surfaces with affine camera models. Since apparent contours are viewpoint dependent, the only true image correspondences are projections of the frontier points, i.e., surface points whose tangent planes are also their epipolar planes. However, frontier points are unknown a priori and must be estimated simultaneously with epipolar geometry. Previous approaches to this problem adopt local greedy search methods which are sensitive to initialization, and may get trapped in local minima. We propose the first algorithm that guarantees global optimality for this problem. We first reformulate the problem using a separable form that allows us to search effectively in a 2D space, instead of on a 5D hypersphere in the classical formulation. Next, in a branch-and-bound algorithm we introduce a novel lower bounding function through interval matrix analysis. Experimental results on both synthetic and real scenes demonstrate that the proposed method is able to quickly obtain the optimal solution.
Yanghai Tsin
ICCV2
2009 A Deformation Tracking Approach to 4D Coronary Artery Tree Reconstruction
Yanghai Tsin, Klaus J. Kirchberg, Günter Lauritsch, Chenyang Xu 0001
MICCAI (1)1
2008 Flexible Edge Arrangement Templates for Object Detection
abstract
We present a novel feature representation for categorical object detection. Unlike previous approaches that have concentrated on generic interest-point detectors, we construct object-specific features directly from the training images. Our feature is represented by a collection of Flexible Edge Arrangement Templates (FEATs). We propose a two-stage semi-supervised learning approach to feature selection. A subset of frequent templates are first selected from a large template pool. In the second stage, we formulate feature selection as a regression problem and use LASSO method to find the most discriminative templates from the preselected ones. FEATs adaptively capture the image structure and naturally accommodate local shape variations. We show that this feature can be complemented by the traditional holistic patch method, thus achieving both efficiency and accuracy. We evaluate our method on three well-known car datasets, showing performance competitive with existing methods.
Yanghai Tsin, Yakup Genc, Takeo Kanade
WACV2
2007 Exploiting Occluding Contours for Real-Time 3D Tracking: A Unified Approach
abstract
Model-based 3D object tracking is fast and robust using 3D edges. However, traditional edge-based approaches have difficulty handling occluding contours of curved surfaces, since they are not static model edges but change with the viewpoint. In this paper we propose a unified approach to edge-based tracking where 3D edges including occluding contours are utilized. This is achieved through an analysis of local surface differential geometry, which provides the foundation for incorporating occluding contours of curved surfaces into edge-based tracking. This approach uses a simple parametrization of both types of model edges within the same framework. The proposed method has been tested within the context of an existing edge-based tracking system. The system can track both types of model edges in a very fast and robust manner. Experimental results on both synthetic and real scenes are provided, which confirm that occluding contours improve real-time 3D tracking performance.
Yanghai Tsin, Yakup Genc
ICCV2
2007 Learn to Track Edges
abstract
Reliability of a model-based edge tracker critically depends on its ability to establish correct correspondences between points on the model edges and edge pixels in an image. This is a non-trivial problem especially in the presence of large inter-frame motions and in cluttered environments. We propose an online learning approach to solving this problem. An edge pixel is represented by a descriptor composed of a small segment of intensity patterns. From training examples the algorithm utilizes the randomized forest model to learn a posteriori distribution of correspondence given the descriptor. In a new frame, the edge pixels are classified using maximum a posteriori (MAP) estimation. The proposed method is very powerful and it enables us to apply the proposed tracker to many previously impossible scenarios with unprecedented robustness.
Yanghai Tsin, Yakup Genc, Ying Zhu 0006, Visvanathan Ramesh
ICCV1
2006 Real-Time Feature Matching using Adaptive and Spatially Distributed Classification Trees
abstract
This paper presents a method for real-time wide-baseline feature matching. The approach is based on the work of Lepetit and colleagues [9], where randomized decision trees are trained to establish correspondences between detected features in a training image, and those in input frames. Though extremely promising, their actual results can vary depending on the viewpoint and illumination conditions. We combine two approaches to alleviate its limitations. The first aims to update the trees at run-time, adapting them to the actual viewing conditions. The second consists in spatially distributing the trees, so that each of them models a certain viewing volume more precisely. The result is a more stable matching method that significantly extends detectable range and is much more robust to illumination changes, such as cast shadows or reflections. 1
Aurélien Boffy, Yanghai Tsin, Yakup Genc
BMVC2
2006 Statistical Shape Models for Object Recognition and Part Localization
abstract
This paper deals with part-based object recognition. We present a statistical shape model that characterizes both variations intrinsic to an object class and perturbations unique to each observation. Object parts, constrained by the learned shape model, are efficiently localized using the max-product algorithm on a triangulated Markov random field (TMRF). As a combination of both techniques, our system is able to recognize objects with large deformations and locate their parts accurately. We test our method on standard databases. The recognition performance is compared to the state of the art and favorable results are reported. 1
Yanghai Tsin, Yakup Genc, Takeo Kanade
BMVC2
2006 Stereo Matching with Linear Superposition of Layers
abstract
In this paper, we address stereo matching in the presence of a class of non-Lambertian effects, where image formation can be modeled as the additive superposition of layers at different depths. The presence of such effects makes it impossible for traditional stereo vision algorithms to recover depths using direct color matching-based methods. We develop several techniques to estimate both depths and colors of the component layers. Depth hypotheses are enumerated in pairs, one from each layer, in a nested plane sweep. For each pair of depth hypotheses, matching is accomplished using spatial-temporal differencing. We then use graph cut optimization to solve for the depths of both layers. This is followed by an iterative color update algorithm which we proved to be convergent. Our algorithm recovers depth and color estimates for both synthetic and real image sequences.
Yanghai Tsin, Sing Bing Kang, Richard Szeliski
IEEE Trans. Pattern Anal. Mach. Intell.1
2005 Object Detection Using 2D Spatial Ordering Constraints
abstract
Object detection is challenging partly due to the limited discriminative power of local feature descriptors. We amend this limitation by incorporating spatial constraints among neighboring features. We propose a two-step algorithm. First, a feature together with its spatial neighbors forms a flexible feature template. Two feature templates can be compared more informatively than two individual features without knowing the 3D object model. A large portion of false matches can be excluded after the first step. In a second global matching step, object detection is formulated as a graph-matching problem. A model graph is constructed by applying Delaunay triangulation on the surviving features. The best matching graph in an input image is computed by finding the maximum a posterior (MAP) estimate of a binary Markov random field with triangular maximal clique. The optimization is solved by the max-product algorithm (a.k.a. belief propagation). Experiments on both rigid and non-rigid objects demonstrate the generality and efficacy of the proposed methods.
Yanghai Tsin, Yakup Genc, Takeo Kanade
CVPR (2)2
2005 Object Detection Using 2D Spatial Ordering Constraints
abstract
Object detection is challenging partly due to the limited discriminative power of local feature descriptors. We propose a two-step algorithm. First, a feature together with its spatial neighbors forms a flexible feature template. Two feature templates can be compared more informatively than two individual features without knowing the 3D object model. A large portion of false matches can be excluded after the first step. In a second global matching step, object detection is formulated as a graph-matching problem. A model graph is constructed by applying Delaunay triangulation on the surviving features. The best matching graph in an input image is computed by finding the maximum a posterior (MAP) estimate of a binary Markov random field with triangular maximal clique. The optimization is solved by the max-product algorithm (a.k.a. belief propagation). Experiments on both rigid and nonrigid objects demonstrate the generality and efficacy of the proposed methods.
Yanghai Tsin, Yakup Genc, Takeo Kanade
CVPR (2)2
2005 The Promise and Perils of Near-Regular Texture
Yanxi Liu 0001, Yanghai Tsin, Wen-Chieh Lin
Int. J. Comput. Vis.2
2004 A Correlation-Based Model Prior for Stereo
Yanghai Tsin, Takeo Kanade
CVPR (1)1
2004 A Correlation-Based Approach to Robust Point Set Registration
Yanghai Tsin, Takeo Kanade
ECCV (3)1
2004 A Computational Model for Periodic Pattern Perception Based on Frieze and Wallpaper Groups
abstract
We present a computational model for periodic pattern perception based on the mathematical theory of crystallographic groups. In each N-dimensional Euclidean space, a finite number of symmetry groups can characterize the structures of an infinite variety of periodic patterns. In 2D space, there are seven frieze groups describing monochrome patterns that repeat along one direction and 17 wallpaper groups for patterns that repeat along two linearly independent directions to tile the plane. We develop a set of computer algorithms that "understand" a given periodic pattern by automatically finding its underlying lattice, identifying its symmetry group, and extracting its representative motifs. We also extend this computational model for near-periodic patterns using geometric AIC. Applications of such a computational model include pattern indexing, texture synthesis, image compression, and gait analysis.
Yanxi Liu 0001, Robert T. Collins, Yanghai Tsin
IEEE Trans. Pattern Anal. Mach. Intell.3
2003 Stereo Matching with Reflections and Translucency
abstract
In this paper, we address the stereo matching problem in the presence of reflections and translucency, where image formation can be modeled as the additive superposition of layers at different depth. The presence of such effects violates the Lambertian assumption underlying traditional stereo vision algorithms, making it impossible to recover component depths using direct color matching based methods. We develop several techniques to estimate both depths and colors of the component layers. Depth hypotheses are enumerated in pairs, one from each layer, in a nested plane sweep. For each pair of depth hypotheses, we compute a component-color-independent matching error per pixel, using a spatial-temporal differencing technique. We then use graph cut optimization to solve for the depths of both layers. This is followed by an iterative color update algorithm whose convergence is proven in our paper. We show convincing results of depth and color estimates for both synthetic and real image sequences.
Yanghai Tsin, Sing Bing Kang, Richard Szeliski
CVPR (1)1
2002 Gait Sequence Analysis Using Frieze Patterns
Yanxi Liu 0001, Robert T. Collins, Yanghai Tsin
ECCV (2)3
2001 Bayesian Color Constancy for Outdoor Object Recognition
abstract
Outdoor scene classification is challenging due to irregular geometry, uncontrolled illumination, and noisy reflectance distributions. This paper discusses a Bayesian approach to classifying a color image of an outdoor scene. A likelihood model factors in the physics of the image formation process, sensor noise distribution, and prior distributions over geometry, material types, and illuminant spectrum parameters. These prior distributions are learned through a training process that uses color observations of planar scene patches over time. An iterative linear algorithm estimates the maximum likelihood reflectance, spectrum, geometry, and object class labels for a new image. Experiments on images taken by outdoor surveillance cameras classify known material types and shadow regions correctly, and flag as outliers material types that were not seen previously.
Yanghai Tsin, Robert T. Collins, Visvanathan Ramesh, Takeo Kanade
CVPR (1)1
2001 Texture Replacement in Real Images
abstract
Texture replacement in real images has many applications, such as interior design, digital movie making and computer graphics. The goal is to replace some specified texture patterns in an image while preserving lighting effects, shadows and occlusions. To achieve convincing replacement results we have to detect texture patterns and estimate the lighting map from a given image. Near regular planar texture patterns are considered in this paper. Given a sample texture patch, a standard tile is computed. Candidate texture regions are determined by mutual information between the standard tile and each image patch. Regions with high mutual information scores are used to estimate the admissible lighting distributions, which is represented by cached statistics. Spatial lighting change constraints are represented by a Markov random field model. Maximum a posteriori estimation of the texture segmentation and lighting map is solved in a stochastic annealing fashion, namely, the Markov chain Monte Carlo method. Visually satisfactory result is achieved using this statistical sampling model.
Yanghai Tsin, Yanxi Liu 0001, Visvanathan Ramesh
CVPR (2)1
2001 Statistical Calibration of the CCD Imaging Process
Yanghai Tsin, Visvanathan Ramesh, Takeo Kanade
ICCV1
1999 Calibration of an Outdoor Active Camera System
abstract
A parametric camera model and calibration procedures are developed for an outdoor active camera system with pan, tilt and zoom control. Unlike traditional methods, active camera motion plays a key role in the calibration process, and no special laboratory setups are required. Intrinsic parameters are estimated automatically by fitting parametric models to the optic flow induced by rotating and zooming. No knowledge of 3D scene structure is needed. Extrinsic parameters are calculated by actively rotating the camera to sight a sparse set of surveyed landmarks over a virtual hemispherical field of view, yielding a well-conditioned pose estimation problem.
Robert T. Collins, Yanghai Tsin
CVPR2