VLDB 2026 Research / reviewers in the wild / expert
Haiyong Jiang
dblp:137/5642
· DBLP profile ↗
28ranked-venue papers
7as first author
19since 2021 · last 2026
0000-0001-7348-5844ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 23 · 7 first-author · 14 since 2021Artificial intelligence and machine learning · 12 · 2 first-author · 9 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AF-BEV: Object-Aware Adaptive Frustum-based BEV Aggregation for 3D Object DetectionabstractMulti-view 3D detection in bird’s-eye-view (BEV) has attracted significant attention from both industry and academia due to its unified, intuitive representation and low cost. By aggregating features along cast rays based on their spatial order, existing methods aim to address the false positive (FP) artifacts commonly observed in detections. However, this strategy may also suppress features in occluded regions along the path, which leads to both inaccurate localization and missed detections. In this paper, we propose AF-BEV, an Adaptive Frustum-based BEV detector, which leverages visible object information to compensate for occluded regions, thereby achieving more accurate localization. Our framework first extends each ray into a frustum and employs an offset prediction module to guide the adaptive frustum angle based on the perceived spatial range of objects. A learnable occlusion-aware aggregation module is then introduced to effectively aggregate frustum features for 3D detection. Finally, we incorporate a unique object-level supervision to guide the generation of appropriate frustum ranges. Extensive experiments on nuScenes demonstrate the effectiveness of our proposed AF-BEV, achieving consistent improvements over the baseline. Additionally, qualitative visualizations provide interpretable evidence of the contributions. Bingyu Zhu, Dongbo Yu, Yunbiao Wang, Jun Xiao 0005, Haiyong Jiang |
ICMR | 5 |
| 2026 | HierRelTriple: Guiding Indoor Layout Generation With Hierarchical Relationship Triplet LossesabstractWe present a hierarchical triplet-based indoor relationship learning method, coined HierRelTriple, with a focus on spatial relationship learning. Existing approaches often depend on attention-based network designs and manually defined training objectives using handcrafted spatial rules or simplified pairwise relationships. However, these methods fail to capture complex, multi-object relationships found in real scenarios, leading to overcrowded or physically implausible arrangements. We introduce HierRelTriple, a hierarchical framework for modeling relational triplets, which first partitions functional regions and then automatically extracts three levels of spatial relationships: object-to-region (O2R), object-to-object (O2O), and corner-to-corner (C2C). By representing these relationships as geometric triplets and employing approaches based on Delaunay Triangulation to establish spatial priors, we derive IoU-based losses between denoised and ground-truth triplets and integrate them seamlessly into the diffusion denoising process. The joint formulation of inter-object distances, angular orientations, and spatial relationships enhances the physical realism of the generated scenes. Extensive experiments on unconditional layout synthesis, floorplan-conditioned layout generation, and scene rearrangement demonstrate that HierRelTriple improves spatial-relation metrics by over 15% and substantially reduces collisions and boundary violations compared to state-of-the-art methods. Kaifan Sun, Bingchen Yang, Peter Wonka, Jun Xiao 0005, Haiyong Jiang |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2025 | An Industrial Multi-machining Feature Dataset and Contrastive Learning-Based Network for Feature Recognition
Haochen He, Zhengda Lu, Haiyong Jiang, Yiqun Wang 0001, Jun Xiao 0005 |
CGI (2) | 4 |
| 2025 | Empowering Vector Graphics with Consistently Arbitrary Viewing and View-dependent VisibilityabstractThis work presents a novel text-to-vector graphics generation approach, Dream3DVG, allowing for arbitrary viewpoint viewing, progressive detail optimization, and view-dependent occlusion awareness. Our approach is a dual-branch optimization framework, consisting of an auxiliary 3D Gaussian Splatting optimization branch and a 3D vector graphics optimization branch. The introduced 3DGS branch can bridge the domain gaps between text prompts and vector graphics with more consistent guidance. Moreover, 3DGS allows for progressive detail control by scheduling classifier-free guidance, facilitating guiding vector graphics with coarse shapes at the initial stages and finer details at later stages. We also improve the view-dependent occlusions by devising a visibility-awareness rendering module. Extensive results on 3D sketches and 3D iconographies, demonstrate the superiority of the method on different abstraction levels of details, cross-view consistency, and occlusion-aware stroke culling. Code is available at https://github.com/chenxinl/Dream3DVG.git. Jun Xiao 0005, Zhengda Lu, Yiqun Wang 0001, Haiyong Jiang |
CVPR | 5 |
| 2025 | Activating Sparse Part Concepts for 3D Class Incremental LearningabstractThis work tackles the challenge of 3D Class-Incremental Learning (CIL), where a model must learn to classify new 3D objects while retaining knowledge of previously learned classes. Existing methods often struggle with catastrophic forgetting, misclassifying old objects due to overreliance on shortcut local features. Our approach addresses this issue by learning a set of part concepts for part-aware features. Particularly, we only activate a small subset of part concepts for the feature representation of each part-aware feature. This facilitates better generalization across categories and mitigates catastrophic forgetting. We further improve the task-wise classification through a part relation-aware Transformer design. At last, we devise learnable affinities to fuse task-wise classification heads and avoid confusion among different tasks. We evaluate our method on three 3D CIL benchmarks, achieving state-of-the-art performance. Code is available at https://github.com/zhenyatian/ILPC. Zhenya Tian, Jun Xiao 0005, Lupeng Liu, Haiyong Jiang |
CVPR | 4 |
| 2025 | D^3CTTA: Domain-Dependent Decorrelation for Continual Test-Time Adaption of 3D LiDAR SegmentationabstractAdapting pre-trained LiDAR segmentation models to dynamic domain shifts during testing is of paramount importance for the safety of autonomous driving. Most existing methods neglect the influence of domain changes and point density in continual test-time adaption (CTTA), relying on backpropagation and large batch sizes for stability. We approach this problem with three insights: 1) Point clouds at different distances usually have different densities resulting in distribution disparities; 2) The feature distribution of different domains varies, and domain-aware parameters can alleviate domain gaps; 3) Features are highly correlated and make segmentation of different labels confusing. To this end, this work presents D3CTTA, an online backpropagation-free framework for 3D continual test-time adaption for LiDAR segmentation. D3CTTA consists of a distance-aware prototype learning module to integrate LiDAR-based geometry prior and a domain-dependent decorrelation module to reduce feature correlations among different domains and different categories. Extensive experiments on three benchmarks showcase that our method achieves a state-of-the-art performance compared to both backpropagation-based methods and backpropagation-free methods. Code is available at https://github.com/ZhaoJichun1/D3CTTA. Jichun Zhao, Haiyong Jiang, Haoxuan Song, Jun Xiao 0005, Dong Gong |
CVPR | 2 |
| 2025 | SegGraph: Leveraging Graphs of SAM Segments for Few-Shot 3D Part SegmentationabstractThis work presents a novel framework for few-shot 3D part segmentation. Recent advances have demonstrated the significant potential of 2D foundation models for low-shot 3D part segmentation. However, it is still an open problem that how to effectively aggregate 2D knowledge from foundation models to 3D. Existing methods either ignore geometric structures for 3D feature learning or neglects the high-quality grouping clues from SAM, leading to under-segmentation and inconsistent part labels. We devise a novel SAM segment graph-based propagation method, named SegGraph, to explicitly learn geometric features encoded within SAM's segmentation masks. Our method encodes geometric features by modeling mutual overlap and adjacency between segments while preserving intra-segment semantic consistency. We construct a segment graph, conceptually similar to an atlas, where nodes represent segments and edges capture their spatial relationships (overlap/adjacency). Each node adaptively modulates 2D foundation model features, which are then propagated via a graph neural network to learn global geometric structures. To enforce intra-segment semantic consistency, we map segment features to 3D points with a novel view-direction-weighted fusion attenuating contributions from low-quality segments. Extensive experiments on PartNet-E demonstrate that our method outperforms all competing baselines by at least 6.9% mIoU. Further analysis reveals that SegGraph achieves particularly strong performance on small components and part boundaries, demonstrating its superior geometric understanding. Yueyang Hu, Haiyong Jiang, Haoxuan Song, Jun Xiao 0005, Hao Pan 0001 |
NeurIPS | 2 |
| 2025 | PS-CAD: Local Geometry Guidance via Prompting and Selection for CAD ReconstructionabstractReverse engineering CAD models from raw geometry is a classic but challenging research problem. In particular, reconstructing the CAD modeling sequence from point clouds provides great interpretability and convenience for editing. Analyzing previous work, we observed that a CAD modeling sequence represented by tokens and processed by a generative model does not have an immediate geometric interpretation. To improve upon this problem, we introduce geometric guidance into the reconstruction network. Our proposed model, PS-CAD, reconstructs the CAD modeling sequence one step at a time as illustrated in Figure 1 . At each step, we provide three forms of geometric guidance. First, we provide the geometry of surfaces where the current reconstruction differs from the complete model as a point cloud. This helps the framework to focus on regions that still need work. Second, we use geometric analysis to extract a set of planar prompts, that correspond to candidate surfaces where a CAD extrusion step could be started. Third, we present a step-wise sampling to generate multiple complete candidate CAD modeling steps instead of single-tokens without direct geometric interpretation. Our framework has three major components. Geometric guidance computation extracts the first two types of geometric guidance. Single-step reconstruction computes a single candidate CAD modeling step for each provided prompt. Single-step selection selects among the candidate CAD modeling steps. The process continues until the reconstruction is completed. Our quantitative results show a significant improvement across all metrics. For example, on the dataset DeepCAD, PS-CAD improves upon the best published SOTA method by reducing the geometry errors (CD and HD) by 10%, and the structural error (ECD metric) by about 13%. Bingchen Yang, Haiyong Jiang, Hao Pan 0001, Guosheng Lin, Jun Xiao 0005, Peter Wonka |
ACM Trans. Graph. | 2 |
| 2025 | Exploring Structural Lines for Interior Floorplan Segmentation
Bingchen Yang, Haiyong Jiang, Zhengda Lu, Jun Xiao 0005 |
Vis. Comput. | 2 |
| 2024 | World to Code: Multi-modal Data Generation via Self-Instructed Compositional Captioning and FilteringabstractRecent advances in Vision-Language Models (VLMs) and the scarcity of high-quality multimodal alignment data have inspired numerous researches on synthetic VLM data generation.The conventional norm in VLM data construction uses a mixture of specialists in caption and OCR, or stronger VLM APIs and expensive human annotation.In this paper, we present World to Code (W2C), a meticulously curated multi-modal data construction pipeline that organizes the final generation output into a Python code format.The pipeline leverages the VLM itself to extract cross-modal information via different prompts and filter the generated outputs again via a consistency filtering strategy.Experiments have demonstrated the high quality of W2C by improving various existing visual question answering and visual grounding benchmarks across different VLMs.Further analysis also demonstrates that the new code parsing ability of VLMs presents better crossmodal equivalence than the commonly used detail caption ability.Our code is available at https://github.com/foundation-multimodal- models/World2Code. Jiacong Wang, Bohong Wu, Haiyong Jiang, Haoyuan Guo, Jun Xiao 0005 |
EMNLP | 3 |
| 2024 | PartCom: Part Composition Learning for 3D Open-Set Recognition
Tingyu Weng, Jun Xiao 0005, Hao Pan 0001, Haiyong Jiang |
Int. J. Comput. Vis. | 4 |
| 2023 | Unsupervised 3D Articulated Object Correspondences with Part Approximation and Shape Refinement
Junqi Diao, Haiyong Jiang, Feilong Yan, Jinhui Luan, Jun Xiao 0005 |
CAD/Graphics | 2 |
| 2023 | VectorFloorSeg: Two-Stream Graph Attention Network for Vectorized Roughcast Floorplan SegmentationabstractVector graphics (VG) are ubiquitous in industrial designs. In this paper, we address semantic segmentation of a typical VG, i.e., roughcast floorplans with bare wall structures, whose output can be directly used for further applications like interior furnishing and room space modeling. Previous semantic segmentation works mostly process well-decorated floorplans in raster images and usually yield aliased boundaries and outlier fragments in segmented rooms, due to pixel-level segmentation that ignores the regular elements (e.g. line segments) in vector floor-plans. To overcome these issues, we propose to fully utilize the regular elements in vector floorplans for more integral segmentation. Our pipeline predicts room segmentation from vector floorplans by dually classifying line segments as room boundaries, and regions partitioned by line segments as room segments. To fully exploit the structural relationships between lines and regions, we use two-stream graph neural networks to process the line segments and partitioned regions respectively, and devise a novel modulated graph attention layer to fuse the heterogeneous information from one stream to the other. Extensive experiments show that by directly operating on vector floorplans, we outper-form image-based methods in both mIoU and mAcc. In addition, we propose a new metric that captures room integrity and boundary regularity, which confirms that our method produces much more regular segmentations. Source code is available at https://github.com/DrZiji/VecFloorSeg. Bingchen Yang, Haiyong Jiang, Hao Pan 0001, Jun Xiao 0005 |
CVPR | 2 |
| 2023 | Decompose Novel into Known: Part Concept Learning For 3D Novel Class DiscoveryabstractIn this work, we address 3D novel class discovery (NCD) that discovers novel classes from an unlabeled dataset by leveraging the knowledge of disjoint known classes. The key challenge of 3D NCD is that learned features by known class recognition are heavily biased and hinder generalization to novel classes. Since geometric parts are more generalizable across different classes, we propose to decompose novel into known parts, coined DNIK, to mitigate the above problems. DNIK learns a part concept bank encoding rich part geometric patterns from known classes so that novel 3D shapes can be represented as part concept compositions to facilitate cross-category generalization. Moreover, we formulate three constraints on part concepts to ensure diverse part concepts without collapsing. A part relation encoding module (PRE) is also developed to leverage part-wise spatial relations for better recognition. We construct three 3D NCD tasks for evaluation and extensive experiments show that our method achieves significantly superior results than SOTA baselines (+11.7%, +14.1%, and +16.3% improvements on average for three tasks, respectively). Code and data will be released. Tingyu Weng, Jun Xiao 0005, Haiyong Jiang |
NeurIPS | 3 |
| 2023 | Combating Spurious Correlations in Loose-fitting Garment Animation Through Joint-Specific Feature LearningabstractAbstract We address the 3D animation of loose‐fitting garments from a sequence of body motions. State‐of‐the‐art approaches treat all body joints as a whole to encode motion features, which usually gives rise to learned spurious correlations between garment vertices and irrelevant joints as shown in Fig. 1. To cope with the issue, we encode temporal motion features in a joint‐wise manner and learn an association matrix to map human joints only to most related garment regions by encouraging its sparsity. In this way, spurious correlations are mitigated and better performance is achieved. Furthermore, we devise the joint‐specific pose space deformation (PSD) to decompose the high‐dimensional displacements as the combination of dynamic details caused by individual joint poses. Extensive experiments show that our method outperforms previous works in most indicators. Moreover, garment animations are not interfered with by artifacts caused by spurious correlations, which further validates the effectiveness of our approach. The code is available at https://github.com/qiji77/JointNet . Junqi Diao, Jun Xiao 0005, Yihong He, Haiyong Jiang |
Comput. Graph. Forum | 4 |
| 2023 | PuzzleNet: Boundary-Aware Feature Matching for Non-Overlapping 3D Point Clouds Assembly
Jianwei Guo 0003, Haiyong Jiang, Yan-Chao Liu, Xiaopeng Zhang 0001, Dong-Ming Yan 0001 |
J. Comput. Sci. Technol. | 3 |
| 2023 | Context-Aware 3D Point Cloud Semantic Segmentation With Plane GuidanceabstractPoint cloud segmentation is fundamental in under- standing 3D environments. However, most existing methods usually perform poorly on identifying boundaries of touching objects and large surfaces of objects. Planes in a scene usually act as supporting surfaces to separate touching objects and provide geometry priors to group points on a large surface as shown in Fig. 1. Besides, planes can roughly represent the structure of a scene, and are more efficient to encode holistic scene contexts than large scale point clouds. In light of the above advantages, we advise a plane-assisted module, coined3D-PAM, to enhance semantic segmentation of touching objects and large surface objects.3D-PAMconsists of a plane separation network (PS-Net) and a plane relation network (PR-Net).PS-Netfocuses on learning features that can robustly separate touching objects, e.g., a chair on a floor, as well as capture plane-based geometry priors to group points on a large plane, e.g., points of a desk.PR-Netencodes mutual plane relations as a proxy of a scene structure to capture holistic contexts.3D-PAMis designed as a plug-and-play module so that it can be easily plugged into any off-the-shelf semantic segmentation network. Extensive experiments demonstrate that the method achieves large segmentation improvements on several backbones, and accomplishes superior results on most categories when using a RandLA-Net backbone ($11/13$categories on S3DIS dataset and$15/20$categories on ScanNetv2 dataset). The project is available at GitHubhttps://github.com/windmillknight/Context-Aware-3D-Point-Cloud-Semantic-Segmentation-With-Plane-Guidance Tingyu Weng, Jun Xiao 0005, Feilong Yan, Haiyong Jiang |
IEEE Trans. Multim. | 4 |
| 2021 | Neighborhood-based Neural Implicit Reconstruction from Point CloudsabstractNeural implicit reconstruction is emerging as a promising approach to constructing 3D geometry from point clouds due to its ability to model geometry with complicated topology and unrestricted resolution. Current methods in this category usually deliver smooth and good quality results, but suffer from defective details and generalization issues. The major reason is that these methods use either a global code or interpolated feature on 3D grids of limited resolution to estimate implicit surface, therefore may cause distortion in feature discretization. This paper presents a neighborhood-aware neural implicit reconstruction framework that consists of an encoder network, a feature aggregation module, and a decoder network to learn implicit surface. The method can easily incorporate an off-the-shelf 3D point-based or volume-based neural network as an encoder. At the heart of our framework is the aggregation module that fuses the learnt contextual features on neighbor inputs so that the method can directly exploit local features of neighboring inputs for geometry detail recovery as well as cross-domain generalization. Experimental results demonstrate that our method significantly outperforms the state-of-the-art methods (about 4.0 points IoU improvements in ShapeNet dataset and 9.0 points IoU improvements in DFAUST dataset). Furthermore, our method preserves finer shape details and can be successfully transferred to a novel category without fine-tuning. Haiyong Jiang, Jianfei Cai 0001, Jianmin Zheng, Jun Xiao 0005 |
3DV | 1 |
| 2021 | CSG-Stump: A Learning Friendly CSG-Like Representation for Interpretable Shape ParsingabstractGenerating an interpretable and compact representation of 3D shapes from point clouds is an important and challenging problem. This paper presents CSG-Stump Net, an unsupervised end-to-end network for learning shapes from point clouds and discovering the underlying constituent modeling primitives and operations as well. At the core is a three-level structure called CSG-Stump, consisting of a complement layer at the bottom, an intersection layer in the middle, and a union layer at the top. CSG-Stump is proven to be equivalent to CSG in terms of representation, therefore inheriting the interpretable, compact and editable nature of CSG while freeing from CSG’s complex tree structures. Particularly, the CSG-Stump has a simple and regular structure, allowing neural networks to give outputs of a constant dimensionality, which makes itself deep-learning friendly. Due to these characteristics of CSG-Stump, CSG-Stump Net achieves superior results compared to previous CSG-based methods and generates much more appealing shapes, as confirmed by extensive experiments. Daxuan Ren, Jianmin Zheng, Jianfei Cai 0001, Haiyong Jiang, Zhongang Cai, Junzhe Zhang 0002, Liang Pan, Haiyu Zhao, Shuai Yi |
ICCV | 5 |
| 2020 | End-to-End 3D Point Cloud Instance Segmentation Without Detectionabstract3D instance segmentation plays a predominant role in environment perception of robotics and augmented reality. Many deep learning based methods have been presented recently for this task. These methods rely on either a detection branch to propose objects or a grouping step to assemble same-instance points. However, detection based methods do not ensure a consistent instance label for each point, while the grouping step requires parameter-tuning and is computationally expensive. In this paper, we introduce a novel framework to enable end-to-end instance segmentation without detection and a separate step of grouping. The core idea is to convert instance segmentation to a candidate assignment problem. At first, a set of instance candidates is sampled. Then we propose an assignment module for candidate assignment and a suppression module to eliminate redundant candidates. A mapping between instance labels and instance candidates is further sought to construct an instance grouping loss for the network training. Experimental results demonstrate that our method is more effective and efficient than previous approaches. Haiyong Jiang, Feilong Yan, Jianfei Cai 0001, Jianmin Zheng, Jun Xiao 0005 |
CVPR | 1 |
| 2020 | Inverse Procedural Modeling of Branching Structures by Inferring L-SystemsabstractWe introduce an inverse procedural modeling approach that learns L-system representations of pixel images with branching structures. Our fully automatic model generates a compact set of textual rewriting rules that describe the input. We use deep learning to discover atomic structures such as line segments or branchings. Orientation and scaling of these structures are determined and the detected structures are combined into a tree. The initial representation is analyzed, and repeating parts are encoded into a small grammar by using greedy optimization while the user can control the size of the detected rules. The output is an L-system that represents the input image as a simple text and a set of terminal symbols. We apply our approach to a variety of examples, demonstrate its robustness against noise and blur, and we show that it can detect user sketches and complex input structures. Jianwei Guo 0003, Haiyong Jiang, Bedrich Benes, Oliver Deussen, Xiaopeng Zhang 0001, Dani Lischinski, Hui Huang 0004 |
ACM Trans. Graph. | 2 |
| 2020 | Selection Expressions for Procedural ModelingabstractWe introduce a new approach for procedural modeling. Our main idea is to select shapes using selection-expressions instead of simple string matching used in current state-of-the-art grammars like CGA shape and CGA++. A selection-expression specifies how to select a potentially complex subset of shapes from a shape hierarchy, e.g., "select all tall windows in the second floor of the main building facade". This new way of modeling enables us to express modeling ideas in their global context rather than traditional rules that operate only locally. To facilitate selection-based procedural modeling we introduce the procedural modeling language SelEx. An important implication of our work is that enforcing important constraints, such as alignment and same size constraints can be done by construction. Therefore, our procedural descriptions can generate facade and building variations without violating alignment and sizing constraints that plague the current state of the art. While the procedural modeling of architecture is our main application domain, we also demonstrate that our approach nicely extends to other man-made objects. Haiyong Jiang, Dong-Ming Yan 0001, Xiaopeng Zhang 0001, Peter Wonka |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2019 | Skeleton-Aware 3D Human Shape Reconstruction From Point CloudsabstractThis work addresses the problem of 3D human shape reconstruction from point clouds. Considering that human shapes are of high dimensions and with large articulations, we adopt the state-of-the-art parametric human body model, SMPL, to reduce the dimension of learning space and generate smooth and valid reconstruction. However, SMPL parameters, especially pose parameters, are not easy to learn because of ambiguity and locality of the pose representation. Thus, we propose to incorporate skeleton awareness into the deep learning based regression of SMPL parameters for 3D human shape reconstruction. Our basic idea is to use the state-of-the-art technique PointNet++ to extract point features, and then map point features to skeleton joint features and finally to SMPL parameters for the reconstruction from point clouds. Particularly, we develop an end-to-end framework, where we propose a graph aggregation module to augment PointNet++ by extracting better point features, an attention module to better map unordered point features into ordered skeleton joint features, and a skeleton graph module to extract better joint features for SMPL parameter regression. The entire framework network is first trained in an end-to-end manner on synthesized dataset, and then online fine-tuned on unseen dataset with unsupervised loss to bridges gaps between training and testing. The experiments on multiple datasets show that our method is on par with the state-of-the-art solution. Haiyong Jiang, Jianfei Cai 0001, Jianmin Zheng |
ICCV | 1 |
| 2019 | Context-Aware Feature and Label Fusion for Facial Action Unit Intensity Estimation With Partially Labeled DataabstractFacial action unit (AU) intensity estimation is a fundamental task for facial behaviour analysis. Most previous methods use a whole face image as input for intensity prediction. Considering that AUs are defined according to their corresponding local appearance, a few patch-based methods utilize image features of local patches. However, fusion of local features is always performed via straightforward feature concatenation or summation. Besides, these methods require fully annotated databases for model learning, which is expensive to acquire. In this paper, we propose a novel weakly supervised patch-based deep model on basis of two types of attention mechanisms for joint intensity estimation of multiple AUs. The model consists of a feature fusion module and a label fusion module. And we augment attention mechanisms of these two modules with a learnable task-related context, as one patch may play different roles in analyzing different AUs and each AU has its own temporal evolution rule. The context-aware feature fusion module is used to capture spatial relationships among local patches while the context-aware label fusion module is used to capture the temporal dynamics of AUs. The latter enables the model to be trained on a partially annotated database. Experimental evaluations on two benchmark expression databases demonstrate the superior performance of the proposed method. Yong Zhang 0034, Haiyong Jiang, Baoyuan Wu, Yanbo Fan |
ICCV | 2 |
| 2019 | Unsupervised Dense Light Field Reconstruction with Occlusion AwarenessabstractAbstract Light field (LF) reconstruction is a fundamental technique in light field imaging and has applications in both software and hardware aspects. This paper presents an unsupervised learning method for LF‐oriented view synthesis, which provides a simple solution for generating quality light fields from a sparse set of views. The method is built on disparity estimation and image warping. Specifically, we first use per‐view disparity as a geometry proxy to warp input views to novel views. Then we compensate the occlusion with a network by a forward‐backward warping process. Cycle‐consistency between different views are explored to enable unsupervised learning and accurate synthesis. The method overcomes the drawbacks of fully supervised learning methods that require large labeled training dataset and epipolar plane image based interpolation methods that do not make full use of geometry consistency in LFs. Experimental results demonstrate that the proposed method can generate high quality views for LF, which outperforms unsupervised approaches and is comparable to fully‐supervised approaches. Lixia Ni, Haiyong Jiang, Jianfei Cai 0001, Jianmin Zheng, Haifeng Li 0002, Xu Liu 0022 |
Comput. Graph. Forum | 2 |
| 2016 | Symmetrization of facade layouts
Haiyong Jiang, Dong-Ming Yan 0001, Weiming Dong, Fuzhang Wu, Liangliang Nan, Xiaopeng Zhang 0001 |
Graph. Model. | 1 |
| 2016 | Automatic Constraint Detection for 2D Layout RegularizationabstractIn this paper, we address the problem of constraint detection for layout regularization. The layout we consider is a set of two-dimensional elements where each element is represented by its bounding box. Layout regularization is important in digitizing plans or images, such as floor plans and facade images, and in the improvement of user-created contents, such as architectural drawings and slide layouts. To regularize a layout, we aim to improve the input by detecting and subsequently enforcing alignment, size, and distance constraints between layout elements. Similar to previous work, we formulate layout regularization as a quadratic programming problem. In addition, we propose a novel optimization algorithm that automatically detects constraints. We evaluate the proposed framework using a variety of input layouts from different applications. Our results demonstrate that our method has superior performance to the state of the art. Haiyong Jiang, Liangliang Nan, Dong-Ming Yan 0001, Weiming Dong, Xiaopeng Zhang 0001, Peter Wonka |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2015 | Facade Layout SymmetrizationabstractWe present an automatic algorithm for symmetrizing facade layouts. Our method symmetrizes a given facade layout while minimally modifying the original layout. Based on the principles of symmetry in urban design, we formulate the problem of facade layout symmetrization as an optimization problem. Our system further enhances the regularity of the final layout by redistributing and aligning boxes in the layout. We demonstrate that the proposed solution can generate symmetric facade layouts efficiently. Haiyong Jiang, Weiming Dong, Dong-Ming Yan 0001, Xiaopeng Zhang 0001 |
CAD/Graphics | 1 |