VLDB 2026 Research / reviewers in the wild / expert
Tian Fang
dblp:06/1750
· DBLP profile ↗
61ranked-venue papers
3as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 46 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 44 · 2 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 5Computer networks · 3 · 2 since 2021Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Planogram: A Multi-dimensional Physical Location Planning System for DC NetworksabstractMeta's data centers underpin a vast array of Internet services and have faced unprecedented demand due to the rapid expansion of AI workloads. The traditional approach of building standardized data centers is increasingly challenged by the exponential growth in required capacity that is now sourced in a variety of non-standard physical environments and data center designs. This shift introduces a complex challenge: how to rapidly and repeatably design custom data center networks that balance multiple, often conflicting, objectives across diverse engineering disciplines. Richard Cziva, Alexander Mafusalov, Shrinivas Petale, Abhinav Triguna, Manikantan Kr, Susana Contrera, Jimmy Williams, Alexey Andreyev, Tian Fang, Satyajeet Ahuja, Ying Zhang 0022 |
SIGCOMM | 11 |
| 2026 | ES-PUF: A practical Physically Unclonable Function for wired networks using Ethernet physical-layer signals
Aiqun Hu, Linning Peng, Baofu Han, Tian Fang, Pan Feng |
Comput. Networks | 6 |
| 2025 | Matrix3D: Large Photogrammetry Model All-in-OneabstractWe present Matrix3D, a unified model that performs several photogrammetry subtasks, including pose estimation, depth prediction, and novel view synthesis using just the same model. Matrix3D utilizes a multi-modal diffusion transformer (DiT) to integrate transformations across several modalities, such as images, camera parameters, and depth maps. The key to Matrix3D’s large-scale multi-modal training lies in the incorporation of a mask learning strategy. This enables full-modality model training even with partially complete data, such as bi-modality data of image-pose and image-depth pairs, thus significantly increases the pool of available training data. Matrix3D demonstrates state-of-the-art performance in pose estimation and novel view synthesis tasks. Additionally, it offers fine-grained control through multi-round interactions, making it an innovative tool for 3D content creation. Project page: https://nju-3dv.github.io/projects/matrix3d. Yuanxun Lu, Jingyang Zhang, Tian Fang, Jean-Daniel Nahmias, Yanghai Tsin, Long Quan, Xun Cao, Yao Yao 0008, Shiwei Li 0001 |
CVPR | 3 |
| 2025 | Micro multi-objective genetic algorithm with information fitting strategy for low-power microprocessor
Hu Peng, Tian Fang, Jianpeng Xiong, Zhongtian Luo |
Expert Syst. Appl. | 2 |
| 2024 | Direct2.5: Diverse Text-to-3D Generation via Multi-view 2.5D DiffusionabstractRecent advances in generative AI have unveiled significant potential for the creation of 3D content. However, current methods either apply a pre-trained 2D diffusion model with the time-consuming score distillation sampling (SDS), or a direct 3D diffusion model trained on limited 3D data losing generation diversity. In this work, we approach the problem by employing a multi-view 2.5D diffusion fine-tuned from a pre-trained 2D diffusion model. The multi-view 2.5D diffusion directly models the structural distribution of 3D data, while still maintaining the strong generalization ability of the original 2D diffusion model, filling the gap between 2D diffusion-based and direct 3D diffusion-based methods for 3D content generation. During inference, multi-view normal maps are generated using the 2.5D diffusion, and a novel differentiable rasterization scheme is introduced to fuse the almost consistent multi-view normal maps into a consistent 3D model. We further design a normal-conditioned multi-view image generation module for fast appearance generation given the 3D geometry. Our method is a one-pass diffusion process and does not require any SDS optimization as post-processing. We demonstrate through extensive experiments that, our direct 2.5D generation with the specially-designed fusion scheme can achieve diverse, mode-seeking-free, and high-fidelity 3D content generation in only 10 seconds. Project page: https://nju-3dv.github.io/projects/direct25. Yuanxun Lu, Jingyang Zhang, Shiwei Li 0001, Tian Fang, David McKinnon, Yanghai Tsin, Long Quan, Xun Cao, Yao Yao 0008 |
CVPR | 4 |
| 2024 | JointNet: Extending Text-to-Image Diffusion for Dense Distribution ModelingabstractWe introduce JointNet, a novel neural network architecture for modeling the joint distribution of images and an additional dense modality (e.g., depth maps).
JointNet is extended from a pre-trained text-to-image diffusion model, where a copy of the original network is created for the new dense modality branch and is densely connected with the RGB branch.
The RGB branch is locked during network fine-tuning, which enables efficient learning of the new modality distribution while maintaining the strong generalization ability of the large-scale pre-trained diffusion model.
We demonstrate the effectiveness of JointNet by using the RGB-D diffusion as an example and through extensive experiments, showcasing its applicability in a variety of applications, including joint RGB-D generation, dense depth prediction, depth-conditioned image generation, and high-resolution 3D panorama generation. Jingyang Zhang, Shiwei Li 0001, Yuanxun Lu, Tian Fang, David McKinnon, Yanghai Tsin, Long Quan, Yao Yao 0008 |
ICLR | 4 |
| 2023 | NeILF++: Inter-Reflectable Light Fields for Geometry and Material EstimationabstractWe present a novel differentiable rendering framework for joint geometry, material, and lighting estimation from multi-view images. In contrast to previous methods which assume a simplified environment map or co-located flashlights, in this work, we formulate the lighting of a static scene as one neural incident light field (NeILF) and one outgoing neural radiance field (NeRF). The key insight of the proposed method is the union of the incident and outgoing light fields through physically-based rendering and inter-reflections between surfaces, making it possible to disentangle the scene geometry, material, and lighting from image observations in a physically-based manner. The proposed incident light and inter-reflection framework can be easily applied to other NeRF systems. We show that our method can not only decompose the outgoing radiance into incident lights and surface materials, but also serve as a surface refinement module that further improves the reconstruction detail of the neural surface. We demonstrate on several datasets that the proposed method is able to achieve state-of-the-art results in terms of geometry reconstruction quality, material estimation accuracy, and the fidelity of novel view rendering. Jingyang Zhang, Yao Yao 0008, Shiwei Li 0001, Tian Fang, David McKinnon, Yanghai Tsin, Long Quan |
ICCV | 5 |
| 2023 | Vis-MVSNet: Visibility-Aware Multi-view Stereo Network
Jingyang Zhang, Shiwei Li 0001, Zixin Luo, Tian Fang, Yao Yao 0008 |
Int. J. Comput. Vis. | 4 |
| 2022 | Critical Regularizations for Neural Surface Reconstruction in the WildabstractNeural implicit functions have recently shown promising results on surface reconstructions from multiple views. However, current methods still suffer from excessive time complexity and poor robustness when reconstructing unbounded or complex scenes. In this paper, we present RegSDF, which shows that proper point cloud supervisions and geometry regularizations are sufficient to produce high-quality and robust reconstruction results. Specifically, RegSDF takes an additional oriented point cloud as input, and optimizes a signed distance field and a surface light field within a differentiable rendering framework. We also introduce the two critical regularizations for this optimization. The first one is the Hessian regularization that smoothly diffuses the signed distance values to the entire distance field given noisy and incomplete input. And the second one is the minimal surface regularization that compactly interpolates and extrapolates the missing geometry. Extensive experiments are conducted on DTU, Blended-MVS, and Tanks and Temples datasets. Compared with recent neural surface reconstruction approaches, RegSDF is able to reconstruct surfaces with fine details even for open scenes with complex topologies and unstructured camera trajectories. Jingyang Zhang, Yao Yao 0008, Shiwei Li 0001, Tian Fang, David McKinnon, Yanghai Tsin, Long Quan |
CVPR | 4 |
| 2022 | ASpanFormer: Detector-Free Image Matching with Adaptive Span Transformer
Zixin Luo, Lei Zhou 0011, Yurun Tian, Mingmin Zhen, Tian Fang, David McKinnon, Yanghai Tsin, Long Quan |
ECCV (32) | 6 |
| 2022 | NeILF: Neural Incident Light Field for Physically-based Material Estimation
Yao Yao 0008, Jingyang Zhang, Yihang Qu, Tian Fang, David McKinnon, Yanghai Tsin, Long Quan |
ECCV (31) | 5 |
| 2021 | Non-fragile extended dissipative synchronization of Markov jump inertial neural networks: An event-triggered control strategy
Tian Fang, Shiyu Jiao, Dongmei Fu, Jing Wang 0071 |
Neurocomputing | 1 |
| 2020 | Visibility-aware Multi-view Stereo Network
Jingyang Zhang, Yao Yao 0008, Shiwei Li 0001, Zixin Luo, Tian Fang |
BMVC | 5 |
| 2020 | BlendedMVS: A Large-Scale Dataset for Generalized Multi-View Stereo NetworksabstractWhile deep learning has recently achieved great success on multi-view stereo (MVS), limited training data makes the trained model hard to be generalized to unseen scenarios. Compared with other computer vision tasks, it is rather difficult to collect a large-scale MVS dataset as it requires expensive active scanners and labor-intensive process to obtain ground truth 3D structures. In this paper, we introduce BlendedMVS, a novel large-scale dataset, to provide sufficient training ground truth for learning-based MVS. To create the dataset, we apply a 3D reconstruction pipeline to recover high-quality textured meshes from images of well-selected scenes. Then, we render these mesh models to color images and depth maps. To introduce the ambient lighting information during training, the rendered color images are further blended with the input images to generate the training input. Our dataset contains over 17k high-resolution images covering a variety of scenes, including cities, architectures, sculptures and small objects. Extensive experiments demonstrate that BlendedMVS endows the trained model with significantly better generalization ability compared with other MVS datasets. The dataset and pretrained models are available at https://github.com/YoYo000/BlendedMVS. Yao Yao 0008, Zixin Luo, Shiwei Li 0001, Jingyang Zhang, Yufan Ren, Lei Zhou 0011, Tian Fang, Long Quan |
CVPR | 7 |
| 2020 | ASLFeat: Learning Local Features of Accurate Shape and LocalizationabstractThis work focuses on mitigating two limitations in the joint learning of local feature detectors and descriptors. First, the ability to estimate the local shape (scale, orientation, etc.) of feature points is often neglected during dense feature extraction, while the shape-awareness is crucial to acquire stronger geometric invariance. Second, the localization accuracy of detected keypoints is not sufficient to reliably recover camera geometry, which has become the bottleneck in tasks such as 3D reconstruction. In this paper, we present ASLFeat, with three light-weight yet effective modifications to mitigate above issues. First, we resort to deformable convolutional networks to densely estimate and apply local transformation. Second, we take advantage of the inherent feature hierarchy to restore spatial resolution and low-level details for accurate keypoint localization. Finally, we use a peakiness measurement to relate feature responses and derive more indicative detection scores. The effect of each modification is thoroughly studied, and the evaluation is extensively conducted across a variety of practical scenarios. State-of-the-art results are reported that demonstrate the superiority of our methods. Zixin Luo, Lei Zhou 0011, Xuyang Bai, Yao Yao 0008, Shiwei Li 0001, Tian Fang, Long Quan |
CVPR | 8 |
| 2020 | Joint Semantic Segmentation and Boundary Detection Using Iterative Pyramid ContextsabstractIn this paper, we present a joint multi-task learning framework for semantic segmentation and boundary detection. The critical component in the framework is the iterative pyramid context module (PCM), which couples two tasks and stores the shared latent semantics to interact between the two tasks. For semantic boundary detection, we propose the novel spatial gradient fusion to suppress non-semantic edges. As semantic boundary detection is the dual task of semantic segmentation, we introduce a loss function with boundary consistency constraint to improve the boundary pixel accuracy for semantic segmentation. Our extensive experiments demonstrate superior performance over state-of-the-art works, not only in semantic segmentation but also in semantic boundary detection. In particular, a mean IoU score of 81.8% on Cityscapes test set is achieved without using coarse data or any external data for semantic segmentation. For semantic boundary detection, we improve over previous state-of-the-art works by 9.9% in terms of AP and 6.8% in terms of MF(ODS). Mingmin Zhen, Jinglu Wang, Lei Zhou 0011, Shiwei Li 0001, Tianwei Shen, Jiaxiang Shang, Tian Fang, Long Quan |
CVPR | 7 |
| 2020 | KFNet: Learning Temporal Camera Relocalization Using Kalman FilteringabstractTemporal camera relocalization estimates the pose with respect to each video frame in sequence, as opposed to one-shot relocalization which focuses on a still image. Even though the time dependency has been taken into account, current temporal relocalization methods still generally underperform the state-of-the-art one-shot approaches in terms of accuracy. In this work, we improve the temporal relocalization method by using a network architecture that incorporates Kalman filtering (KFNet) for online camera relocalization. In particular, KFNet extends the scene coordinate regression problem to the time domain in order to recursively establish 2D and 3D correspondences for the pose determination. The network architecture design and the loss formulation are based on Kalman filtering in the context of Bayesian learning. Extensive experiments on multiple relocalization benchmarks demonstrate the high accuracy of KFNet at the top of both one-shot and temporal relocalization approaches. Lei Zhou 0011, Zixin Luo, Tianwei Shen, Mingmin Zhen, Yao Yao 0008, Tian Fang, Long Quan |
CVPR | 7 |
| 2020 | Self-Supervised Monocular 3D Face Reconstruction by Occlusion-Aware Multi-view Geometry Consistency
Jiaxiang Shang, Tianwei Shen, Shiwei Li 0001, Lei Zhou 0011, Mingmin Zhen, Tian Fang, Long Quan |
ECCV (15) | 6 |
| 2020 | Learning Discriminative Feature with CRF for Unsupervised Video Object Segmentation
Mingmin Zhen, Shiwei Li 0001, Lei Zhou 0011, Jiaxiang Shang, Haoan Feng, Tian Fang, Long Quan |
ECCV (27) | 6 |
| 2020 | Stochastic Bundle Adjustment for Efficient and Scalable 3D Reconstruction
Lei Zhou 0011, Zixin Luo, Mingmin Zhen, Tianwei Shen, Shiwei Li 0001, Zhuofei Huang, Tian Fang, Long Quan |
ECCV (15) | 7 |
| 2020 | Learning Stereo Matchability in Disparity Regression NetworksabstractLearning-based stereo matching has recently achieved promising results, yet still suffers difficulties in establishing reliable matches in weakly matchable regions that are textureless, non-Lambertian, or occluded. In this paper, we address this challenge by proposing a stereo matching network that considers pixel-wise matchability. Specifically, the network jointly regresses disparity and matchability maps from 3D probability volume through expectation and entropy operations. Next, a learned attenuation is applied as the robust loss function to alleviate the influence of weakly matchable pixels in the training. Finally, a matchability-aware disparity refinement is introduced to improve the depth inference in weakly matchable regions. The proposed deep stereo matchability (DSM) framework can improve the matching result or accelerate the computation while still guaranteeing the quality. Moreover, the DSM framework is portable to many recent stereo networks. Extensive experiments are conducted on Scene Flow and KITTI stereo datasets to demonstrate the effectiveness of the proposed framework over the state-of-the-art learning-based stereo methods. Jingyang Zhang, Yao Yao 0008, Zixin Luo, Shiwei Li 0001, Tianwei Shen, Tian Fang, Long Quan |
ICPR | 6 |
| 2020 | Distributed Very Large Scale Bundle Adjustment by Global Camera ConsensusabstractThe increasing scale of Structure-from-Motion is fundamentally limited by the conventional optimization framework for the all-in-one global bundle adjustment. In this paper, we propose a distributed approach to coping with this global bundle adjustment for very large scale Structure-from-Motion computation. First, we derive the distributed formulation from the classical optimization algorithm ADMM, Alternating Direction Method of Multipliers, based on the global camera consensus. Then, we analyze the conditions under which the convergence of this distributed optimization would be guaranteed. In particular, we adopt over-relaxation and self-adaption schemes to improve the convergence rate. After that, we propose to split the large scale camera-point visibility graph in order to reduce the communication overheads of the distributed computing. The experiments on both public large scale SfM data-sets and our very large scale aerial photo sets demonstrate that the proposed distributed method clearly outperforms the state-of-the-art method in efficiency and accuracy. Siyu Zhu 0001, Tianwei Shen, Lei Zhou 0011, Zixin Luo, Tian Fang, Long Quan |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2019 | Learning Fully Dense Neural Networks for Image Semantic SegmentationabstractSemantic segmentation is pixel-wise classification which retains critical spatial information. The “feature map reuse” has been commonly adopted in CNN based approaches to take advantage of feature maps in the early layers for the later spatial reconstruction. Along this direction, we go a step further by proposing a fully dense neural network with an encoderdecoder structure that we abbreviate as FDNet. For each stage in the decoder module, feature maps of all the previous blocks are adaptively aggregated to feedforward as input. On the one hand, it reconstructs the spatial boundaries accurately. On the other hand, it learns more efficiently with the more efficient gradient backpropagation. In addition, we propose the boundary-aware loss function to focus more attention on the pixels near the boundary, which boosts the “hard examples” labeling. We have demonstrated the best performance of the FDNet on the two benchmark datasets: PASCAL VOC 2012, NYUDv2 over previous works when not considering training on other datasets. Mingmin Zhen, Jinglu Wang, Lei Zhou 0011, Tian Fang, Long Quan |
AAAI | 4 |
| 2019 | Recurrent MVSNet for High-Resolution Multi-View Stereo Depth InferenceabstractDeep learning has recently demonstrated its excellent performance for multi-view stereo (MVS). However, one major limitation of current learned MVS approaches is the scalability: the memory-consuming cost volume regularization makes the learned MVS hard to be applied to high-resolution scenes. In this paper, we introduce a scalable multi-view stereo framework based on the recurrent neural network. Instead of regularizing the entire 3D cost volume in one go, the proposed Recurrent Multi-view Stereo Network (R-MVSNet) sequentially regularizes the 2D cost maps along the depth direction via the gated recurrent unit (GRU). This reduces dramatically the memory consumption and makes high-resolution reconstruction feasible. We first show the state-of-the-art performance achieved by the proposed R-MVSNet on the recent MVS benchmarks. Then, we further demonstrate the scalability of the proposed method on several large-scale scenarios, where previous learned approaches often fail due to the memory constraint. Code is available at https://github.com/YoYo000/MVSNet. Yao Yao 0008, Zixin Luo, Shiwei Li 0001, Tianwei Shen, Tian Fang, Long Quan |
CVPR | 5 |
| 2019 | Cross-Atlas Convolution for Parameterization Invariant Learning on Textured Mesh SurfaceabstractWe present a convolutional network architecture for direct feature learning on mesh surfaces through their atlases of texture maps. The texture map encodes the parameterization from 3D to 2D domain, rendering not only RGB values but also rasterized geometric features if necessary. Since the parameterization of texture map is not pre-determined, and depends on the surface topologies, we therefore introduce a novel cross-atlas convolution to recover the original mesh geodesic neighborhood, so as to achieve the invariance property to arbitrary parameterization. The proposed module is integrated into classification and segmentation architectures, which takes the input texture map of a mesh, and infers the output predictions. Our method not only shows competitive performances on classification and segmentation public benchmarks, but also paves the way for the broad mesh surfaces learning. Shiwei Li 0001, Zixin Luo, Mingmin Zhen, Yao Yao 0008, Tianwei Shen, Tian Fang, Long Quan |
CVPR | 6 |
| 2019 | ContextDesc: Local Descriptor Augmentation With Cross-Modality ContextabstractMost existing studies on learning local features focus on the patch-based descriptions of individual keypoints, whereas neglecting the spatial relations established from their keypoint locations. In this paper, we go beyond the local detail representation by introducing context awareness to augment off-the-shelf local feature descriptors. Specifically, we propose a unified learning framework that leverages and aggregates the cross-modality contextual information, including (i) visual context from high-level image representation, and (ii) geometric context from 2D keypoint distribution. Moreover, we propose an effective N-pair loss that eschews the empirical hyper-parameter search and improves the convergence. The proposed augmentation scheme is lightweight compared with the raw local feature description, meanwhile improves remarkably on several large-scale benchmarks with diversified scenes, which demonstrates both strong practicality and generalization ability in geometric matching applications. Zixin Luo, Tianwei Shen, Lei Zhou 0011, Yao Yao 0008, Shiwei Li 0001, Tian Fang, Long Quan |
CVPR | 7 |
| 2019 | Beyond Photometric Loss for Self-Supervised Ego-Motion EstimationabstractAccurate relative pose is one of the key components in visual odometry (VO) and simultaneous localization and mapping (SLAM). Recently, the self-supervised learning framework that jointly optimizes the relative pose and target image depth has attracted the attention of the community. Previous works rely on the photometric error generated from depths and poses between adjacent frames, which contains large systematic error under realistic scenes due to reflective surfaces and occlusions. In this paper, we bridge the gap between geometric loss and photometric loss by introducing the matching loss constrained by epipolar geometry in a self-supervised framework. Evaluated on the KITTI dataset, our method outperforms the state-of-the-art unsupervised egomotion estimation methods by a large margin. The code and data are available at https://github.com/hlzz/DeepMatchVO. Tianwei Shen, Zixin Luo, Lei Zhou 0011, Hanyu Deng, Tian Fang, Long Quan |
ICRA | 6 |
| 2019 | Multi-view based neural network for semantic segmentation on 3D scenes
Yonghua Lu, Mingmin Zhen, Tian Fang |
Sci. China Inf. Sci. | 3 |
| 2018 | Matchable Image Retrieval by Learning from Surface Reconstruction
Tianwei Shen, Zixin Luo, Lei Zhou 0011, Siyu Zhu 0001, Tian Fang, Long Quan |
ACCV (1) | 6 |
| 2018 | Reconstructing Thin Structures of Manifold Surfaces by Integrating Spatial CurvesabstractThe manifold surface reconstruction in multi-view stereo often fails in retaining thin structures due to incomplete and noisy reconstructed point clouds. In this paper, we address this problem by leveraging spatial curves. The curve representation in nature is advantageous in modeling thin and elongated structures, implying topology and connectivity information of the underlying geometry, which exactly compensates the weakness of scattered point clouds. We present a novel surface reconstruction method using both curves and point clouds. First, we propose a 3D curve reconstruction algorithm based on the initialize-optimize-extend strategy. Then, tetrahedra are constructed from points and curves, where the volumes of thin structures are robustly preserved by the Curve-conformed Delaunay Refinement. Finally, the mesh surface is extracted from tetrahedra by a graph optimization. The method has been intensively evaluated on both synthetic and real-world datasets, showing significant improvements over state-of-the-art methods. Shiwei Li 0001, Yao Yao 0008, Tian Fang, Long Quan |
CVPR | 3 |
| 2018 | Very Large-Scale Global SfM by Distributed Motion AveragingabstractGlobal Structure-from-Motion (SfM) techniques have demonstrated superior efficiency and accuracy than the conventional incremental approach in many recent studies. This work proposes a divide-and-conquer framework to solve very large global SfM at the scale of millions of images. Specifically, we first divide all images into multiple partitions that preserve strong data association for well-posed and parallel local motion averaging. Then, we solve a global motion averaging that determines cameras at partition boundaries and a similarity transformation per partition to register all cameras in a single coordinate frame. Finally, local and global motion averaging are iterated until convergence. Since local camera poses are fixed during the global motion average, we can avoid caching the whole reconstruction in memory at once. This distributed framework significantly enhances the efficiency and robustness of large-scale motion averaging. Siyu Zhu 0001, Lei Zhou 0011, Tianwei Shen, Tian Fang, Ping Tan 0002, Long Quan |
CVPR | 5 |
| 2018 | GeoDesc: Learning Local Descriptors by Integrating Geometry Constraints
Zixin Luo, Tianwei Shen, Lei Zhou 0011, Siyu Zhu 0001, Yao Yao 0008, Tian Fang, Long Quan |
ECCV (9) | 7 |
| 2018 | MVSNet: Depth Inference for Unstructured Multi-view Stereo
Yao Yao 0008, Zixin Luo, Shiwei Li 0001, Tian Fang, Long Quan |
ECCV (8) | 4 |
| 2018 | Learning and Matching Multi-View Descriptors for Registration of Point Clouds
Lei Zhou 0011, Siyu Zhu 0001, Zixin Luo, Tianwei Shen, Mingmin Zhen, Tian Fang, Long Quan |
ECCV (15) | 7 |
| 2018 | FBOSS: building switch software at scaleabstractThe conventional software running on network devices, such as switches and routers, is typically vendor-supplied, proprietary and closed-source; as a result, it tends to contain extraneous features that a single operator will not most likely fully utilize. Furthermore, cloud-scale data center networks often times have software and operational requirements that may not be well addressed by the switch vendors. Sean Choi, Boris Burkov, Alex Eckert, Tian Fang, Saman Kazemkhani, Rob Sherwood, Ying Zhang 0022, Hongyi Zeng |
SIGCOMM | 4 |
| 2017 | Relative Camera Refinement for Accurate Dense ReconstructionabstractMulti-view stereo (MVS) depends on the pre-determined camera geometry, often from structure from motion (SfM) or simultaneous localization and mapping (SLAM). However, cameras may not be locally optimal for dense stereo matching, especially when it comes from the large scale SfM or the SLAM with multiple sensor fusion. In this paper, we propose a local camera refinement approach for accurate dense reconstruction. Firstly, we refines the relative geometry of independent camera pair using a tailored bundle adjustment. The refinement is also extended to a multi-view version for general MVS reconstructions. Then, the non-rigid dense alignment is formulated as an inverse-distortion problem to transfer point clouds from each local coordinate system to a global coordinate system. The proposed framework has been intensively validated in both SfM and SLAM based dense reconstructions. Results on different datasets show that our method can significantly improve the dense reconstruction quality. Yao Yao 0008, Shiwei Li 0001, Siyu Zhu 0001, Hanyu Deng, Tian Fang, Long Quan |
3DV | 5 |
| 2017 | Distributed Very Large Scale Bundle Adjustment by Global Camera ConsensusabstractThe increasing scale of Structure-from-Motion is fundamentally limited by the conventional optimization framework for the all-in-one global bundle adjustment. In this paper, we propose a distributed approach to coping with this global bundle adjustment for very large scale Structure-from-Motion computation. First, we derive the distributed formulation from the classical optimization algorithm ADMM, Alternating Direction Method of Multipliers, based on the global camera consensus. Then, we analyze the conditions under which the convergence of this distributed optimization would be guaranteed. In particular, we adopt over-relaxation and self-adaption schemes to improve the convergence rate. After that, we propose to split the large scale camera-point visibility graph in order to reduce the communication overheads of the distributed computing. The experiments on both public large scale SfM data-sets and our very large scale aerial photo sets demonstrate that the proposed distributed method clearly outperforms the state-of-the-art method in efficiency and accuracy. Siyu Zhu 0001, Tian Fang, Long Quan |
ICCV | 3 |
| 2017 | Progressive Large Scale-Invariant Image Matching in Scale SpaceabstractThe power of modern image matching approaches is still fundamentally limited by the abrupt scale changes in images. In this paper, we propose a scale-invariant image matching approach to tackling the very large scale variation of views. Drawing inspiration from the scale space theory, we start with encoding the image's scale space into a compact multi-scale representation. Then, rather than trying to find the exact feature matches all in one step, we propose a progressive two-stage approach. First, we determine the related scale levels in scale space, enclosing the inlier feature correspondences, based on an optimal and exhaustive matching in a limited scale space. Second, we produce both the image similarity measurement and feature correspondences simultaneously after restricting matching between the related scale levels in a robust way. The matching performance has been intensively evaluated on vision tasks including image retrieval, feature matching and Structurefrom- Motion (SfM). The successful integration of the challenging fusion of high aerial and low ground-level views with significant scale differences manifests the superiority of the proposed approach. Lei Zhou 0011, Siyu Zhu 0001, Tianwei Shen, Jinglu Wang, Tian Fang, Long Quan |
ICCV | 5 |
| 2017 | A feature extraction and similarity metric-learning framework for urban model retrievalabstractUrban model retrieval has wide applications in the geoscience field, and it is also a very challenging research topic due to the blur and background clutter in query images and the large spatial inconsistencies between query and database images. In this study, a feature extraction and similarity metric-learning framework for urban model retrieval is proposed. In the method, the selective search voting algorithm is presented to automatically localize and segment a query object from an input image with the help of the top-ranked retrieved database images. Then, the local features of object images are extracted via sparse coding, and the global features are learned using the spatial constrained convolutional neural network. We utilize a new similarity metric to match the database images with a query object image. Finally, similar 3D models are retrieved. Both qualitative and quantitative experimental results indicate that the proposed framework can localize and segment a query object from an input image precisely and that the retrieval results are better than those of other related approaches. Yuebin Wang, Liqiang Zhang 0001, Xiaohua Tong, Suhong Liu, Tian Fang |
Int. J. Geogr. Inf. Sci. | 5 |
| 2016 | Color Correction for Image-Based Modeling in the Large
Tianwei Shen, Jinglu Wang, Tian Fang, Siyu Zhu 0001, Long Quan |
ACCV (4) | 3 |
| 2016 | Efficient Multi-view Surface Refinement with Adaptive Resolution Control
Shiwei Li 0001, Sing Yu Siu, Tian Fang, Long Quan |
ECCV (1) | 3 |
| 2016 | Graph-Based Consistent Matching for Structure-from-Motion
Tianwei Shen, Siyu Zhu 0001, Tian Fang, Long Quan |
ECCV (3) | 3 |
| 2016 | A Local Structure and Direction-Aware Optimization Approach for Three-Dimensional Tree ModelingabstractModeling 3-D trees from terrestrial laser scanning (TLS) point clouds remains a challenging task for several well-known reasons, including their complex structure and severe occlusions. In order to accurately reconstruct 3-D tree models from TLS point clouds that typically suffer from significant occlusions, in this paper, a novel local structure and direction-aware approach is presented to successfully complete missing structures of trees. In this method, we first extract the coarse tree skeleton from the input point cloud, and thus, the branch dominant direction and the point density of each branch are obtained. By a skeleton-based Laplacian algorithm, the point cloud is further shrunk into a skeleton point cloud to highlight the branch dominant direction of each branch. For obtaining even more accurate point densities, a dictionary-based algorithm is utilized to learn and reconstruct the local structure. Finally, the branch dominant direction and point density are integrated into an iterative optimization process to recover the missing data. Extensive experimental results have shown that the proposed method is very robust to incomplete data sets, and it is capable of accurately reconstructing 3-D trees, which are partially, or even to a large extent, missing from the input point cloud. Zhen Wang 0032, Liqiang Zhang 0001, Tian Fang, Xiaohua Tong, P. Takis Mathiopoulos, Liang Zhang 0023, Jie Mei 0004 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2016 | Image-Based Building Regularization Using Structural Linear FeaturesabstractReconstructed building models using stereo-based methods inevitably suffer from noise, leading to the lack of regularity which is characterized by straightness of structural linear features and smoothness of homogeneous regions. We leverage the structural linear features embedded in the mesh to construct a novel surface scaffold structure for model regularization. The regularization comprises two iterative stages: (1) the linear features are semi-automatically proposed from images by exploiting photometric and geometric clues jointly; (2) the scaffold topology represented by spatial relations among the linear features is optimized according to data fidelity and topological rules, then the mesh is refined by adjusting itself to the consolidated scaffold. Our method has two advantages. First, the proposed scaffold representation is able to concisely describe semantic building structures. Second, the scaffold structure is embedded in the mesh, which can preserve the mesh connectivity and avoid stitching or intersecting surfaces in challenging cases. We demonstrate that our method can enhance structural characteristics and suppress irregularities in the building models robustly in some challenging datasets. Moreover, the regularization can significantly improve the results of general applications such as simplification and non-photorealistic rendering. Jinglu Wang, Tian Fang, Qingkun Su, Siyu Zhu 0001, Shengnan Cai, Chiew-Lan Tai, Long Quan |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2015 | Higher-Order CRF Structural Segmentation of 3D Reconstructed SurfacesabstractIn this paper, we propose a structural segmentation algorithm to partition multi-view stereo reconstructed surfaces of large-scale urban environments into structural segments. Each segment corresponds to a structural component describable by a surface primitive of up to the second order. This segmentation is for use in subsequent urban object modeling, vectorization, and recognition. To overcome the high geometrical and topological noise levels in the 3D reconstructed urban surfaces, we formulate the structural segmentation as a higher-order Conditional Random Field (CRF) labeling problem. It not only incorporates classical lower-order 2D and 3D local cues, but also encodes contextual geometric regularities to disambiguate the noisy local cues. A general higher-order CRF is difficult to solve. We develop a bottom-up progressive approach through a patch-based surface representation, which iteratively evolves from the initial mesh triangles to the final segmentation. Each iteration alternates between performing a prior discovery step, which finds the contextual regularities of the patch-based representation, and an inference step that leverages the regularities as higher-order priors to construct a more stable and regular segmentation. The efficiency and robustness of the proposed method is extensively demonstrated on real reconstruction models, yielding significantly better performance than classical mesh segmentation methods. Jinglu Wang, Tian Fang, Chiew-Lan Tai, Long Quan |
ICCV | 3 |
| 2015 | Joint Camera Clustering and Surface Segmentation for Large-Scale Multi-view StereoabstractIn this paper, we propose an optimal decomposition approach to large-scale multi-view stereo from an initial sparse reconstruction. The success of the approach depends on the introduction of surface-segmentation-based camera clustering rather than sparse-point-based camera clustering, which suffers from the problems of non-uniform reconstruction coverage ratio and high redundancy. In details, we introduce three criteria for camera clustering and surface segmentation for reconstruction, and then we formulate these criteria into an energy minimization problem under constraints. To solve this problem, we propose a joint optimization in a hierarchical framework to obtain the final surface segments and corresponding optimal camera clusters. On each level of the hierarchical framework, the camera clustering problem is formulated as a parameter estimation problem of a probability model solved by a General Expectation-Maximization algorithm and the surface segmentation problem is formulated as a Markov Random Field model based on the probability estimated by the previous camera clustering process. The experiments on several Internet datasets and aerial photo datasets demonstrate that the proposed approach method generates more uniform and complete dense reconstruction with less redundancy, resulting in more efficient multi-view stereo algorithm. Shiwei Li 0001, Tian Fang, Siyu Zhu 0001, Long Quan |
ICCV | 3 |
| 2015 | A Multiscale and Hierarchical Feature Extraction Method for Terrestrial Laser Scanning Point Cloud ClassificationabstractThe effective extraction of shape features is an important requirement for the accurate and efficient classification of terrestrial laser scanning (TLS) point clouds. However, the challenge of how to obtain robust and discriminative features from noisy and varying density TLS point clouds remains. This paper introduces a novel multiscale and hierarchical framework, which describes the classification of TLS point clouds of cluttered urban scenes. In this framework, we propose multiscale and hierarchical point clusters (MHPCs). In MHPCs, point clouds are first resampled into different scales. Then, the resampled data set of each scale is aggregated into several hierarchical point clusters, where the point cloud of all scales in each level is termed a point-cluster set. This representation not only accounts for the multiscale properties of point clouds but also well captures their hierarchical structures. Based on the MHPCs, novel features of point clusters are constructed by employing the latent Dirichlet allocation (LDA). An LDA model is trained according to a training set. The LDA model then extracts a set of latent topics, i.e., a feature of topics, for a point cluster. Finally, to apply the introduced features for point-cluster classification, we train an AdaBoost classifier in each point-cluster set and obtain the corresponding classifiers to separate the TLS point clouds with varying point density and data missing into semantic regions. Compared with other methods, our features achieve the best classification results for buildings, trees, people, and cars from TLS point clouds, particularly for small and moving objects, such as people and cars. Zhen Wang 0032, Liqiang Zhang 0001, Tian Fang, P. Takis Mathiopoulos, Xiaohua Tong, Huamin Qu, Zhiqiang Xiao 0002, Dong Chen 0009 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2014 | Multi-scale Tetrahedral Fusion of a Similarity Reconstruction and Noisy Positional Measurements
Tian Fang, Siyu Zhu 0001, Long Quan |
ACCV (2) | 2 |
| 2014 | Multi-view Geometry Compression
Siyu Zhu 0001, Tian Fang, Long Quan |
ACCV (2) | 2 |
| 2014 | Local Readjustment for High-Resolution 3D ReconstructionabstractGlobal bundle adjustment usually converges to a non-zero residual and produces sub-optimal camera poses for local areas, which leads to loss of details for high- resolution reconstruction. Instead of trying harder to optimize everything globally, we argue that we should live with the non-zero residual and adapt the camera poses to local areas. To this end, we propose a segment-based approach to readjust the camera poses locally and improve the reconstruction for fine geometry details. The key idea is to partition the globally optimized structure from motion points into well-conditioned segments for re-optimization, reconstruct their geometry individually, and fuse everything back into a consistent global model. This significantly reduces severe propagated errors and estimation biases caused by the initial global adjustment. The results on several datasets demonstrate that this approach can significantly improve the reconstruction accuracy, while maintaining the consistency of the 3D structure between segments. Siyu Zhu 0001, Tian Fang, Jianxiong Xiao, Long Quan |
CVPR | 2 |
| 2014 | A Structure-Aware Global Optimization Method for Reconstructing 3-D Tree Models From Terrestrial Laser Scanning DataabstractA 3-D tree structure plays an important role in many scientific fields, including forestry and agriculture. For example, terrestrial laser scanning (TLS) can efficiently capture high-precision 3-D spatial arrangements and structure of trees as a point cloud. In the past, several methods to reconstruct 3-D trees from the TLS point cloud were proposed. However, in general, they fail to process incomplete TLS data. To address such incomplete TLS data sets, a new method that is based on a structure-aware global optimization approach (SAGO) is proposed. The SAGO first obtains the approximate tree skeleton from a distance minimum spanning tree (DMst) and then defines the stretching directions of the branches on the tree skeleton. Based on these stretching directions, the SAGO recovers missing data in the incomplete TLS point cloud. The DMst is applied again to obtain the refined tree skeleton from the optimized data, and the tree skeleton is smoothed by employing a Laplacian function. To reconstruct 3-D tree models, the radius of each branch section is estimated, and leaves are added to form the crown geometry. The developed methodology has been extensively evaluated by employing a dozen TLS point clouds of various types of trees. Both qualitative and quantitative performance evaluation results have indicated that the SAGO is capable of effectively reconstructing 3-D tree models from grossly incomplete TLS point clouds with significant amounts of missing data. Zhen Wang 0032, Liqiang Zhang 0001, Tian Fang, P. Takis Mathiopoulos, Huamin Qu, Dong Chen 0009, Yuebin Wang |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2014 | Joint Segmentation of Images and Scanned Point Cloud in Large-Scale Street Scenes With Low-Annotation CostabstractWe propose a novel method for the parsing of images and scanned point cloud in large-scale street environment. The proposed method significantly reduces the intensive labeling cost in previous works by automatically generating training data from the input data. The automatic generation of training data begins with the initialization of training data with weak priors in the street environment, followed by a filtering scheme to remove mislabeled training samples. We formulate the filtering as a binary labeling optimization problem over a conditional random filed that we call object graph, simultaneously integrating spatial smoothness preference and label consistency between 2D and 3D. Toward the final parsing, with the automatically generated training data, a CRF-based parsing method that integrates the coordination of image appearance and 3D geometry is adopted to perform the parsing of large-scale street scenes. The proposed approach is evaluated on city-scale Google Street View data, with an encouraging parsing performance demonstrated. Honghui Zhang, Jinglu Wang, Tian Fang, Long Quan |
IEEE Trans. Image Process. | 3 |
| 2013 | Image-Based Modeling of Unwrappable FaçadesabstractIn this paper, we propose an unwrappable representation for image-based façade modeling from multiple registered images. An unwrappable façade is represented by the mutually orthogonal baseline and profile. We first reconstruct semidense 3D points from images, then the baseline and profile are extracted from the point cloud to construct the base shape and compose the textures of the building from the images. Through our unwrapping process, the reconstructed 3D points and composed textures are further mapped to an unwrapped space that is parameterized by the baseline and profile. In doing so, the unwrapped space becomes equivalent to the planar space in which planar façade modeling techniques can be used to reconstruct the details of the buildings. Finally, the augmented details can be wrapped back to the original 3D space to generate the final model. This newly introduced unwrappable representation extends the state-of-the-art modeling for planar façades to a more general class of façades. We demonstrate the power of the unwrappable representation with a few examples in which the façade is not planar. Tian Fang, Zhexi Wang, Honghui Zhang, Long Quan |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2011 | Partial similarity based nonparametric scene parsing in certain environmentabstractIn this paper we propose a novel nonparametric image parsing method for the image parsing problem in certain environment. A novel and efficient nearest neighbor matching scheme, the ANN bilateral matching scheme, is proposed. Based on the proposed matching scheme, we first retrieve some partially similar images for each given test image from the training image database. The test image can be well explained by these retrieved images, with similar regions existing in the retrieved images for each region in the test image. Then, we match the test image to the retrieved training images with the ANN bilateral matching scheme, and parse the test image by integrating multiple cues in a markov random field. Experiment on three datasets shows our method achieved promising parsing accuracy and outperformed two state-of-the-art nonparametric image parsing methods. Honghui Zhang, Tian Fang, Xiaowu Chen 0001, Qinping Zhao, Long Quan |
CVPR | 2 |
| 2011 | Real solution isolation with multiplicity of zero-dimensional triangular systems
Zhihai Zhang, Tian Fang, Bican Xia |
Sci. China Inf. Sci. | 2 |
| 2010 | Rectilinear parsing of architecture in urban environmentabstractWe propose an approach that parses registered images captured at ground level into architectural units for large-scale city modeling. Each parsed unit has a regularized shape, which can be used for further modeling purposes. In our approach, we first parse the environment into buildings, the ground, and the sky using a joint 2D-3D segmentation method. Then, we partition buildings into individual façades. The partition problem is formulated as a dynamic programming optimization for a sequence of natural vertical separating lines. Each façade is regularized by a floor line and a roof line. The floor line is the intersection line of the vertical plane of buildings and the horizontal plane of the ground. The roof line links edge points of roof region. The parsed results provide a first geometric approximation to the city environment, and can be further analyzed if necessary. The approach is demonstrated and validated on several large-scale city datasets. Tian Fang, Jianxiong Xiao, Honghui Zhang, Qinping Zhao, Long Quan |
CVPR | 2 |
| 2010 | Resampling Structure from Motion
Tian Fang, Long Quan |
ECCV (2) | 1 |
| 2009 | Image-based street-side city modelingabstractWe propose an automatic approach to generate street-side 3D photo-realistic models from images captured along the streets at ground level. We first develop a multi-view semantic segmentation method that recognizes and segments each image at pixel level into semantically meaningful areas, each labeled with a specific object class, such as building, sky, ground, vegetation and car. A partition scheme is then introduced to separate buildings into independent blocks using the major line structures of the scene. Finally, for each block, we propose an inverse patch-based orthographic composition and structure analysis method for façade modeling that efficiently regularizes the noisy and missing reconstructed 3D data. Our system has the distinct advantage of producing visually compelling results by imposing strong priors of building regularity. We demonstrate the fully automatic system on a typical city example to validate our methodology. Jianxiong Xiao, Tian Fang, Maxime Lhuillier, Long Quan |
ACM Trans. Graph. | 2 |
| 2008 | Single image tree modelingabstractIn this paper, we introduce a simple sketching method to generate a realistic 3D tree model from a single image. The user draws at least two strokes in the tree image: the first crown stroke around the tree crown to mark up the leaf region, the second branch stroke from the tree root to mark up the main trunk, and possibly few other branch strokes for refinement. The method automatically generates a 3D tree model including branches and leaves. Branches are synthesized by a growth engine from a small library of elementary subtrees that are pre-defined or built on the fly from the recovered visible branches. The visible branches are automatically traced from the drawn branch strokes according to image statistics on the strokes. Leaves are generated from the region bounded by the first crown stroke to complete the tree. We demonstrate our method on a variety of examples. Tian Fang, Jianxiong Xiao, Long Quan |
ACM Trans. Graph. | 2 |
| 2008 | Image-based façade modelingabstractWe propose in this paper a semi-automatic image-based approach to façade modeling that uses images captured along streets and relies on structure from motion to recover camera positions and point clouds automatically as the initial stage for modeling. We start by considering a building façade as a flat rectangular plane or a developable surface with an associated texture image composited from the multiple visible images. A façade is then decomposed and structured into a Directed Acyclic Graph of rectilinear elementary patches. The decomposition is carried out top-down by a recursive subdivision, and followed by a bottom-up merging with the detection of the architectural bilateral symmetry and repetitive patterns. Each subdivided patch of the flat façade is augmented with a depth optimized using the 3D points cloud. Our system also allows for an easy user feedback in the 2D image space for the proposed decomposition and augmentation. Finally, our approach is demonstrated on a large number of façades from a variety of street-side images. Jianxiong Xiao, Tian Fang, Eyal Ofek, Long Quan |
ACM Trans. Graph. | 2 |
| 2007 | High Resolution Animated Scenes from StillsabstractCurrent techniques for generating animated scenes involve either videos (whose resolution is limited) or a single image (which requires a significant amount of user interaction). In this paper, we describe a system that allows the user to quickly and easily produce a compelling-looking animation from a small collection of high resolution stills. Our system has two unique features. First, it applies an automatic partial temporal order recovery algorithm to the stills in order to approximate the original scene dynamics. The output sequence is subsequently extracted using a second-order Markov Chain model. Second, a region with large motion variation can be automatically decomposed into semiautonomous regions such that their temporal orderings are softly constrained. This is to ensure motion smoothness throughout the original region. The final animation is obtained by frame interpolation and feathering. Our system also provides a simple-to-use interface to help the user to fine-tune the motion of the animated scene. Using our system, an animated scene can be generated in minutes. We show results for a variety of scenes. Zhouchen Lin, Lifeng Wang 0001, Yunbo Wang, Sing Bing Kang, Tian Fang |
IEEE Trans. Vis. Comput. Graph. | 5 |