Haiyong Jiang

dblp:137/5642 · DBLP profile ↗
← Back
28ranked-venue papers
7as first author
19since 2021 · last 2026
0000-0001-7348-5844ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 23 · 7 first-author · 14 since 2021Artificial intelligence and machine learning · 12 · 2 first-author · 9 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 AF-BEV: Object-Aware Adaptive Frustum-based BEV Aggregation for 3D Object Detection
abstract
Multi-view 3D detection in bird’s-eye-view (BEV) has attracted significant attention from both industry and academia due to its unified, intuitive representation and low cost. By aggregating features along cast rays based on their spatial order, existing methods aim to address the false positive (FP) artifacts commonly observed in detections. However, this strategy may also suppress features in occluded regions along the path, which leads to both inaccurate localization and missed detections. In this paper, we propose AF-BEV, an Adaptive Frustum-based BEV detector, which leverages visible object information to compensate for occluded regions, thereby achieving more accurate localization. Our framework first extends each ray into a frustum and employs an offset prediction module to guide the adaptive frustum angle based on the perceived spatial range of objects. A learnable occlusion-aware aggregation module is then introduced to effectively aggregate frustum features for 3D detection. Finally, we incorporate a unique object-level supervision to guide the generation of appropriate frustum ranges. Extensive experiments on nuScenes demonstrate the effectiveness of our proposed AF-BEV, achieving consistent improvements over the baseline. Additionally, qualitative visualizations provide interpretable evidence of the contributions.
Bingyu Zhu, Dongbo Yu, Yunbiao Wang, Jun Xiao 0005, Haiyong Jiang
ICMR5
2026 HierRelTriple: Guiding Indoor Layout Generation With Hierarchical Relationship Triplet Losses
abstract
We present a hierarchical triplet-based indoor relationship learning method, coined HierRelTriple, with a focus on spatial relationship learning. Existing approaches often depend on attention-based network designs and manually defined training objectives using handcrafted spatial rules or simplified pairwise relationships. However, these methods fail to capture complex, multi-object relationships found in real scenarios, leading to overcrowded or physically implausible arrangements. We introduce HierRelTriple, a hierarchical framework for modeling relational triplets, which first partitions functional regions and then automatically extracts three levels of spatial relationships: object-to-region (O2R), object-to-object (O2O), and corner-to-corner (C2C). By representing these relationships as geometric triplets and employing approaches based on Delaunay Triangulation to establish spatial priors, we derive IoU-based losses between denoised and ground-truth triplets and integrate them seamlessly into the diffusion denoising process. The joint formulation of inter-object distances, angular orientations, and spatial relationships enhances the physical realism of the generated scenes. Extensive experiments on unconditional layout synthesis, floorplan-conditioned layout generation, and scene rearrangement demonstrate that HierRelTriple improves spatial-relation metrics by over 15% and substantially reduces collisions and boundary violations compared to state-of-the-art methods.
Kaifan Sun, Bingchen Yang, Peter Wonka, Jun Xiao 0005, Haiyong Jiang
IEEE Trans. Vis. Comput. Graph.5
2025 An Industrial Multi-machining Feature Dataset and Contrastive Learning-Based Network for Feature Recognition
Haochen He, Zhengda Lu, Haiyong Jiang, Yiqun Wang 0001, Jun Xiao 0005
CGI (2)4
2025 Empowering Vector Graphics with Consistently Arbitrary Viewing and View-dependent Visibility
abstract
This work presents a novel text-to-vector graphics generation approach, Dream3DVG, allowing for arbitrary viewpoint viewing, progressive detail optimization, and view-dependent occlusion awareness. Our approach is a dual-branch optimization framework, consisting of an auxiliary 3D Gaussian Splatting optimization branch and a 3D vector graphics optimization branch. The introduced 3DGS branch can bridge the domain gaps between text prompts and vector graphics with more consistent guidance. Moreover, 3DGS allows for progressive detail control by scheduling classifier-free guidance, facilitating guiding vector graphics with coarse shapes at the initial stages and finer details at later stages. We also improve the view-dependent occlusions by devising a visibility-awareness rendering module. Extensive results on 3D sketches and 3D iconographies, demonstrate the superiority of the method on different abstraction levels of details, cross-view consistency, and occlusion-aware stroke culling. Code is available at https://github.com/chenxinl/Dream3DVG.git.
Jun Xiao 0005, Zhengda Lu, Yiqun Wang 0001, Haiyong Jiang
CVPR5
2025 Activating Sparse Part Concepts for 3D Class Incremental Learning
abstract
This work tackles the challenge of 3D Class-Incremental Learning (CIL), where a model must learn to classify new 3D objects while retaining knowledge of previously learned classes. Existing methods often struggle with catastrophic forgetting, misclassifying old objects due to overreliance on shortcut local features. Our approach addresses this issue by learning a set of part concepts for part-aware features. Particularly, we only activate a small subset of part concepts for the feature representation of each part-aware feature. This facilitates better generalization across categories and mitigates catastrophic forgetting. We further improve the task-wise classification through a part relation-aware Transformer design. At last, we devise learnable affinities to fuse task-wise classification heads and avoid confusion among different tasks. We evaluate our method on three 3D CIL benchmarks, achieving state-of-the-art performance. Code is available at https://github.com/zhenyatian/ILPC.
Zhenya Tian, Jun Xiao 0005, Lupeng Liu, Haiyong Jiang
CVPR4
2025 D^3CTTA: Domain-Dependent Decorrelation for Continual Test-Time Adaption of 3D LiDAR Segmentation
abstract
Adapting pre-trained LiDAR segmentation models to dynamic domain shifts during testing is of paramount importance for the safety of autonomous driving. Most existing methods neglect the influence of domain changes and point density in continual test-time adaption (CTTA), relying on backpropagation and large batch sizes for stability. We approach this problem with three insights: 1) Point clouds at different distances usually have different densities resulting in distribution disparities; 2) The feature distribution of different domains varies, and domain-aware parameters can alleviate domain gaps; 3) Features are highly correlated and make segmentation of different labels confusing. To this end, this work presents D3CTTA, an online backpropagation-free framework for 3D continual test-time adaption for LiDAR segmentation. D3CTTA consists of a distance-aware prototype learning module to integrate LiDAR-based geometry prior and a domain-dependent decorrelation module to reduce feature correlations among different domains and different categories. Extensive experiments on three benchmarks showcase that our method achieves a state-of-the-art performance compared to both backpropagation-based methods and backpropagation-free methods. Code is available at https://github.com/ZhaoJichun1/D3CTTA.
Jichun Zhao, Haiyong Jiang, Haoxuan Song, Jun Xiao 0005, Dong Gong
CVPR2
2025 SegGraph: Leveraging Graphs of SAM Segments for Few-Shot 3D Part Segmentation
abstract
This work presents a novel framework for few-shot 3D part segmentation. Recent advances have demonstrated the significant potential of 2D foundation models for low-shot 3D part segmentation. However, it is still an open problem that how to effectively aggregate 2D knowledge from foundation models to 3D. Existing methods either ignore geometric structures for 3D feature learning or neglects the high-quality grouping clues from SAM, leading to under-segmentation and inconsistent part labels. We devise a novel SAM segment graph-based propagation method, named SegGraph, to explicitly learn geometric features encoded within SAM's segmentation masks. Our method encodes geometric features by modeling mutual overlap and adjacency between segments while preserving intra-segment semantic consistency. We construct a segment graph, conceptually similar to an atlas, where nodes represent segments and edges capture their spatial relationships (overlap/adjacency). Each node adaptively modulates 2D foundation model features, which are then propagated via a graph neural network to learn global geometric structures. To enforce intra-segment semantic consistency, we map segment features to 3D points with a novel view-direction-weighted fusion attenuating contributions from low-quality segments. Extensive experiments on PartNet-E demonstrate that our method outperforms all competing baselines by at least 6.9% mIoU. Further analysis reveals that SegGraph achieves particularly strong performance on small components and part boundaries, demonstrating its superior geometric understanding.
Yueyang Hu, Haiyong Jiang, Haoxuan Song, Jun Xiao 0005, Hao Pan 0001
NeurIPS2
2025 PS-CAD: Local Geometry Guidance via Prompting and Selection for CAD Reconstruction
abstract
Reverse engineering CAD models from raw geometry is a classic but challenging research problem. In particular, reconstructing the CAD modeling sequence from point clouds provides great interpretability and convenience for editing. Analyzing previous work, we observed that a CAD modeling sequence represented by tokens and processed by a generative model does not have an immediate geometric interpretation. To improve upon this problem, we introduce geometric guidance into the reconstruction network. Our proposed model, PS-CAD, reconstructs the CAD modeling sequence one step at a time as illustrated in Figure 1 . At each step, we provide three forms of geometric guidance. First, we provide the geometry of surfaces where the current reconstruction differs from the complete model as a point cloud. This helps the framework to focus on regions that still need work. Second, we use geometric analysis to extract a set of planar prompts, that correspond to candidate surfaces where a CAD extrusion step could be started. Third, we present a step-wise sampling to generate multiple complete candidate CAD modeling steps instead of single-tokens without direct geometric interpretation. Our framework has three major components. Geometric guidance computation extracts the first two types of geometric guidance. Single-step reconstruction computes a single candidate CAD modeling step for each provided prompt. Single-step selection selects among the candidate CAD modeling steps. The process continues until the reconstruction is completed. Our quantitative results show a significant improvement across all metrics. For example, on the dataset DeepCAD, PS-CAD improves upon the best published SOTA method by reducing the geometry errors (CD and HD) by 10%, and the structural error (ECD metric) by about 13%.
Bingchen Yang, Haiyong Jiang, Hao Pan 0001, Guosheng Lin, Jun Xiao 0005, Peter Wonka
ACM Trans. Graph.2
2025 Exploring Structural Lines for Interior Floorplan Segmentation
Bingchen Yang, Haiyong Jiang, Zhengda Lu, Jun Xiao 0005
Vis. Comput.2
2024 World to Code: Multi-modal Data Generation via Self-Instructed Compositional Captioning and Filtering
abstract
Recent advances in Vision-Language Models (VLMs) and the scarcity of high-quality multimodal alignment data have inspired numerous researches on synthetic VLM data generation.The conventional norm in VLM data construction uses a mixture of specialists in caption and OCR, or stronger VLM APIs and expensive human annotation.In this paper, we present World to Code (W2C), a meticulously curated multi-modal data construction pipeline that organizes the final generation output into a Python code format.The pipeline leverages the VLM itself to extract cross-modal information via different prompts and filter the generated outputs again via a consistency filtering strategy.Experiments have demonstrated the high quality of W2C by improving various existing visual question answering and visual grounding benchmarks across different VLMs.Further analysis also demonstrates that the new code parsing ability of VLMs presents better crossmodal equivalence than the commonly used detail caption ability.Our code is available at https://github.com/foundation-multimodal- models/World2Code.
Jiacong Wang, Bohong Wu, Haiyong Jiang, Haoyuan Guo, Jun Xiao 0005
EMNLP3
2024 PartCom: Part Composition Learning for 3D Open-Set Recognition
Tingyu Weng, Jun Xiao 0005, Hao Pan 0001, Haiyong Jiang
Int. J. Comput. Vis.4
2023 Unsupervised 3D Articulated Object Correspondences with Part Approximation and Shape Refinement
Junqi Diao, Haiyong Jiang, Feilong Yan, Jinhui Luan, Jun Xiao 0005
CAD/Graphics2
2023 VectorFloorSeg: Two-Stream Graph Attention Network for Vectorized Roughcast Floorplan Segmentation
abstract
Vector graphics (VG) are ubiquitous in industrial designs. In this paper, we address semantic segmentation of a typical VG, i.e., roughcast floorplans with bare wall structures, whose output can be directly used for further applications like interior furnishing and room space modeling. Previous semantic segmentation works mostly process well-decorated floorplans in raster images and usually yield aliased boundaries and outlier fragments in segmented rooms, due to pixel-level segmentation that ignores the regular elements (e.g. line segments) in vector floor-plans. To overcome these issues, we propose to fully utilize the regular elements in vector floorplans for more integral segmentation. Our pipeline predicts room segmentation from vector floorplans by dually classifying line segments as room boundaries, and regions partitioned by line segments as room segments. To fully exploit the structural relationships between lines and regions, we use two-stream graph neural networks to process the line segments and partitioned regions respectively, and devise a novel modulated graph attention layer to fuse the heterogeneous information from one stream to the other. Extensive experiments show that by directly operating on vector floorplans, we outper-form image-based methods in both mIoU and mAcc. In addition, we propose a new metric that captures room integrity and boundary regularity, which confirms that our method produces much more regular segmentations. Source code is available at https://github.com/DrZiji/VecFloorSeg.
Bingchen Yang, Haiyong Jiang, Hao Pan 0001, Jun Xiao 0005
CVPR2
2023 Decompose Novel into Known: Part Concept Learning For 3D Novel Class Discovery
abstract
In this work, we address 3D novel class discovery (NCD) that discovers novel classes from an unlabeled dataset by leveraging the knowledge of disjoint known classes. The key challenge of 3D NCD is that learned features by known class recognition are heavily biased and hinder generalization to novel classes. Since geometric parts are more generalizable across different classes, we propose to decompose novel into known parts, coined DNIK, to mitigate the above problems. DNIK learns a part concept bank encoding rich part geometric patterns from known classes so that novel 3D shapes can be represented as part concept compositions to facilitate cross-category generalization. Moreover, we formulate three constraints on part concepts to ensure diverse part concepts without collapsing. A part relation encoding module (PRE) is also developed to leverage part-wise spatial relations for better recognition. We construct three 3D NCD tasks for evaluation and extensive experiments show that our method achieves significantly superior results than SOTA baselines (+11.7%, +14.1%, and +16.3% improvements on average for three tasks, respectively). Code and data will be released.
Tingyu Weng, Jun Xiao 0005, Haiyong Jiang
NeurIPS3
2023 Combating Spurious Correlations in Loose-fitting Garment Animation Through Joint-Specific Feature Learning
abstract
Abstract We address the 3D animation of loose‐fitting garments from a sequence of body motions. State‐of‐the‐art approaches treat all body joints as a whole to encode motion features, which usually gives rise to learned spurious correlations between garment vertices and irrelevant joints as shown in Fig. 1. To cope with the issue, we encode temporal motion features in a joint‐wise manner and learn an association matrix to map human joints only to most related garment regions by encouraging its sparsity. In this way, spurious correlations are mitigated and better performance is achieved. Furthermore, we devise the joint‐specific pose space deformation (PSD) to decompose the high‐dimensional displacements as the combination of dynamic details caused by individual joint poses. Extensive experiments show that our method outperforms previous works in most indicators. Moreover, garment animations are not interfered with by artifacts caused by spurious correlations, which further validates the effectiveness of our approach. The code is available at https://github.com/qiji77/JointNet .
Junqi Diao, Jun Xiao 0005, Yihong He, Haiyong Jiang
Comput. Graph. Forum4
2023 PuzzleNet: Boundary-Aware Feature Matching for Non-Overlapping 3D Point Clouds Assembly
Jianwei Guo 0003, Haiyong Jiang, Yan-Chao Liu, Xiaopeng Zhang 0001, Dong-Ming Yan 0001
J. Comput. Sci. Technol.3
2023 Context-Aware 3D Point Cloud Semantic Segmentation With Plane Guidance
abstract
Point cloud segmentation is fundamental in under- standing 3D environments. However, most existing methods usually perform poorly on identifying boundaries of touching objects and large surfaces of objects. Planes in a scene usually act as supporting surfaces to separate touching objects and provide geometry priors to group points on a large surface as shown in Fig. 1. Besides, planes can roughly represent the structure of a scene, and are more efficient to encode holistic scene contexts than large scale point clouds. In light of the above advantages, we advise a plane-assisted module, coined3D-PAM, to enhance semantic segmentation of touching objects and large surface objects.3D-PAMconsists of a plane separation network (PS-Net) and a plane relation network (PR-Net).PS-Netfocuses on learning features that can robustly separate touching objects, e.g., a chair on a floor, as well as capture plane-based geometry priors to group points on a large plane, e.g., points of a desk.PR-Netencodes mutual plane relations as a proxy of a scene structure to capture holistic contexts.3D-PAMis designed as a plug-and-play module so that it can be easily plugged into any off-the-shelf semantic segmentation network. Extensive experiments demonstrate that the method achieves large segmentation improvements on several backbones, and accomplishes superior results on most categories when using a RandLA-Net backbone ($11/13$categories on S3DIS dataset and$15/20$categories on ScanNetv2 dataset). The project is available at GitHubhttps://github.com/windmillknight/Context-Aware-3D-Point-Cloud-Semantic-Segmentation-With-Plane-Guidance
Tingyu Weng, Jun Xiao 0005, Feilong Yan, Haiyong Jiang
IEEE Trans. Multim.4
2021 Neighborhood-based Neural Implicit Reconstruction from Point Clouds
abstract
Neural implicit reconstruction is emerging as a promising approach to constructing 3D geometry from point clouds due to its ability to model geometry with complicated topology and unrestricted resolution. Current methods in this category usually deliver smooth and good quality results, but suffer from defective details and generalization issues. The major reason is that these methods use either a global code or interpolated feature on 3D grids of limited resolution to estimate implicit surface, therefore may cause distortion in feature discretization. This paper presents a neighborhood-aware neural implicit reconstruction framework that consists of an encoder network, a feature aggregation module, and a decoder network to learn implicit surface. The method can easily incorporate an off-the-shelf 3D point-based or volume-based neural network as an encoder. At the heart of our framework is the aggregation module that fuses the learnt contextual features on neighbor inputs so that the method can directly exploit local features of neighboring inputs for geometry detail recovery as well as cross-domain generalization. Experimental results demonstrate that our method significantly outperforms the state-of-the-art methods (about 4.0 points IoU improvements in ShapeNet dataset and 9.0 points IoU improvements in DFAUST dataset). Furthermore, our method preserves finer shape details and can be successfully transferred to a novel category without fine-tuning.
Haiyong Jiang, Jianfei Cai 0001, Jianmin Zheng, Jun Xiao 0005
3DV1
2021 CSG-Stump: A Learning Friendly CSG-Like Representation for Interpretable Shape Parsing
abstract
Generating an interpretable and compact representation of 3D shapes from point clouds is an important and challenging problem. This paper presents CSG-Stump Net, an unsupervised end-to-end network for learning shapes from point clouds and discovering the underlying constituent modeling primitives and operations as well. At the core is a three-level structure called CSG-Stump, consisting of a complement layer at the bottom, an intersection layer in the middle, and a union layer at the top. CSG-Stump is proven to be equivalent to CSG in terms of representation, therefore inheriting the interpretable, compact and editable nature of CSG while freeing from CSG’s complex tree structures. Particularly, the CSG-Stump has a simple and regular structure, allowing neural networks to give outputs of a constant dimensionality, which makes itself deep-learning friendly. Due to these characteristics of CSG-Stump, CSG-Stump Net achieves superior results compared to previous CSG-based methods and generates much more appealing shapes, as confirmed by extensive experiments.
Daxuan Ren, Jianmin Zheng, Jianfei Cai 0001, Haiyong Jiang, Zhongang Cai, Junzhe Zhang 0002, Liang Pan, Haiyu Zhao, Shuai Yi
ICCV5
2020 End-to-End 3D Point Cloud Instance Segmentation Without Detection
abstract
3D instance segmentation plays a predominant role in environment perception of robotics and augmented reality. Many deep learning based methods have been presented recently for this task. These methods rely on either a detection branch to propose objects or a grouping step to assemble same-instance points. However, detection based methods do not ensure a consistent instance label for each point, while the grouping step requires parameter-tuning and is computationally expensive. In this paper, we introduce a novel framework to enable end-to-end instance segmentation without detection and a separate step of grouping. The core idea is to convert instance segmentation to a candidate assignment problem. At first, a set of instance candidates is sampled. Then we propose an assignment module for candidate assignment and a suppression module to eliminate redundant candidates. A mapping between instance labels and instance candidates is further sought to construct an instance grouping loss for the network training. Experimental results demonstrate that our method is more effective and efficient than previous approaches.
Haiyong Jiang, Feilong Yan, Jianfei Cai 0001, Jianmin Zheng, Jun Xiao 0005
CVPR1
2020 Inverse Procedural Modeling of Branching Structures by Inferring L-Systems
abstract
We introduce an inverse procedural modeling approach that learns L-system representations of pixel images with branching structures. Our fully automatic model generates a compact set of textual rewriting rules that describe the input. We use deep learning to discover atomic structures such as line segments or branchings. Orientation and scaling of these structures are determined and the detected structures are combined into a tree. The initial representation is analyzed, and repeating parts are encoded into a small grammar by using greedy optimization while the user can control the size of the detected rules. The output is an L-system that represents the input image as a simple text and a set of terminal symbols. We apply our approach to a variety of examples, demonstrate its robustness against noise and blur, and we show that it can detect user sketches and complex input structures.
Jianwei Guo 0003, Haiyong Jiang, Bedrich Benes, Oliver Deussen, Xiaopeng Zhang 0001, Dani Lischinski, Hui Huang 0004
ACM Trans. Graph.2
2020 Selection Expressions for Procedural Modeling
abstract
We introduce a new approach for procedural modeling. Our main idea is to select shapes using selection-expressions instead of simple string matching used in current state-of-the-art grammars like CGA shape and CGA++. A selection-expression specifies how to select a potentially complex subset of shapes from a shape hierarchy, e.g., "select all tall windows in the second floor of the main building facade". This new way of modeling enables us to express modeling ideas in their global context rather than traditional rules that operate only locally. To facilitate selection-based procedural modeling we introduce the procedural modeling language SelEx. An important implication of our work is that enforcing important constraints, such as alignment and same size constraints can be done by construction. Therefore, our procedural descriptions can generate facade and building variations without violating alignment and sizing constraints that plague the current state of the art. While the procedural modeling of architecture is our main application domain, we also demonstrate that our approach nicely extends to other man-made objects.
Haiyong Jiang, Dong-Ming Yan 0001, Xiaopeng Zhang 0001, Peter Wonka
IEEE Trans. Vis. Comput. Graph.1
2019 Skeleton-Aware 3D Human Shape Reconstruction From Point Clouds
abstract
This work addresses the problem of 3D human shape reconstruction from point clouds. Considering that human shapes are of high dimensions and with large articulations, we adopt the state-of-the-art parametric human body model, SMPL, to reduce the dimension of learning space and generate smooth and valid reconstruction. However, SMPL parameters, especially pose parameters, are not easy to learn because of ambiguity and locality of the pose representation. Thus, we propose to incorporate skeleton awareness into the deep learning based regression of SMPL parameters for 3D human shape reconstruction. Our basic idea is to use the state-of-the-art technique PointNet++ to extract point features, and then map point features to skeleton joint features and finally to SMPL parameters for the reconstruction from point clouds. Particularly, we develop an end-to-end framework, where we propose a graph aggregation module to augment PointNet++ by extracting better point features, an attention module to better map unordered point features into ordered skeleton joint features, and a skeleton graph module to extract better joint features for SMPL parameter regression. The entire framework network is first trained in an end-to-end manner on synthesized dataset, and then online fine-tuned on unseen dataset with unsupervised loss to bridges gaps between training and testing. The experiments on multiple datasets show that our method is on par with the state-of-the-art solution.
Haiyong Jiang, Jianfei Cai 0001, Jianmin Zheng
ICCV1
2019 Context-Aware Feature and Label Fusion for Facial Action Unit Intensity Estimation With Partially Labeled Data
abstract
Facial action unit (AU) intensity estimation is a fundamental task for facial behaviour analysis. Most previous methods use a whole face image as input for intensity prediction. Considering that AUs are defined according to their corresponding local appearance, a few patch-based methods utilize image features of local patches. However, fusion of local features is always performed via straightforward feature concatenation or summation. Besides, these methods require fully annotated databases for model learning, which is expensive to acquire. In this paper, we propose a novel weakly supervised patch-based deep model on basis of two types of attention mechanisms for joint intensity estimation of multiple AUs. The model consists of a feature fusion module and a label fusion module. And we augment attention mechanisms of these two modules with a learnable task-related context, as one patch may play different roles in analyzing different AUs and each AU has its own temporal evolution rule. The context-aware feature fusion module is used to capture spatial relationships among local patches while the context-aware label fusion module is used to capture the temporal dynamics of AUs. The latter enables the model to be trained on a partially annotated database. Experimental evaluations on two benchmark expression databases demonstrate the superior performance of the proposed method.
Yong Zhang 0034, Haiyong Jiang, Baoyuan Wu, Yanbo Fan
ICCV2
2019 Unsupervised Dense Light Field Reconstruction with Occlusion Awareness
abstract
Abstract Light field (LF) reconstruction is a fundamental technique in light field imaging and has applications in both software and hardware aspects. This paper presents an unsupervised learning method for LF‐oriented view synthesis, which provides a simple solution for generating quality light fields from a sparse set of views. The method is built on disparity estimation and image warping. Specifically, we first use per‐view disparity as a geometry proxy to warp input views to novel views. Then we compensate the occlusion with a network by a forward‐backward warping process. Cycle‐consistency between different views are explored to enable unsupervised learning and accurate synthesis. The method overcomes the drawbacks of fully supervised learning methods that require large labeled training dataset and epipolar plane image based interpolation methods that do not make full use of geometry consistency in LFs. Experimental results demonstrate that the proposed method can generate high quality views for LF, which outperforms unsupervised approaches and is comparable to fully‐supervised approaches.
Lixia Ni, Haiyong Jiang, Jianfei Cai 0001, Jianmin Zheng, Haifeng Li 0002, Xu Liu 0022
Comput. Graph. Forum2
2016 Symmetrization of facade layouts
Haiyong Jiang, Dong-Ming Yan 0001, Weiming Dong, Fuzhang Wu, Liangliang Nan, Xiaopeng Zhang 0001
Graph. Model.1
2016 Automatic Constraint Detection for 2D Layout Regularization
abstract
In this paper, we address the problem of constraint detection for layout regularization. The layout we consider is a set of two-dimensional elements where each element is represented by its bounding box. Layout regularization is important in digitizing plans or images, such as floor plans and facade images, and in the improvement of user-created contents, such as architectural drawings and slide layouts. To regularize a layout, we aim to improve the input by detecting and subsequently enforcing alignment, size, and distance constraints between layout elements. Similar to previous work, we formulate layout regularization as a quadratic programming problem. In addition, we propose a novel optimization algorithm that automatically detects constraints. We evaluate the proposed framework using a variety of input layouts from different applications. Our results demonstrate that our method has superior performance to the state of the art.
Haiyong Jiang, Liangliang Nan, Dong-Ming Yan 0001, Weiming Dong, Xiaopeng Zhang 0001, Peter Wonka
IEEE Trans. Vis. Comput. Graph.1
2015 Facade Layout Symmetrization
abstract
We present an automatic algorithm for symmetrizing facade layouts. Our method symmetrizes a given facade layout while minimally modifying the original layout. Based on the principles of symmetry in urban design, we formulate the problem of facade layout symmetrization as an optimization problem. Our system further enhances the regularity of the final layout by redistributing and aligning boxes in the layout. We demonstrate that the proposed solution can generate symmetric facade layouts efficiently.
Haiyong Jiang, Weiming Dong, Dong-Ming Yan 0001, Xiaopeng Zhang 0001
CAD/Graphics1