VLDB 2026 Research / reviewers in the wild / expert
Xin Li 0003
dblp:09/1365-3 · also Xin Shane Li
· DBLP profile ↗
118ranked-venue papers
15as first author
59since 2021 · last 2026
0000-0002-0144-9489ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 89 · 11 first-author · 37 since 2021Artificial intelligence and machine learning · 42 · 2 first-author · 30 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Common Pattern Prior-Driven Semi-Supervised Medical Image SegmentationabstractSemi-supervised learning (SSL) has emerged as a promising paradigm for medical image segmentation, aiming to alleviate the scarcity of high-quality annotations by combining limited labeled and abundant unlabeled data. However, existing SSL methods suffer from inherent limitations: 1) consistency regularization overly relies on enforcing prediction consistency under different perturbations, neglecting deep exploration of semantic and discriminative features; 2) pseudo-labeling methods are prone to introducing noise, which in turn undermines the stability of model training. To enable high-quality and more stable model learning, we propose a common pattern prior-driven network (CPP-Net) for semi-supervised medical image segmentation. To improve training quality, CPP-Net proposes a pattern learning mechanism that extracts each class's core semantic information for high-quality feature learning. At its core, it is a dynamically updated common pattern bank (CP-Bank), which stores class-specific patterns learned throughout training and serves as high-quality prior knowledge for the model. By reusing CP-Bank patterns, CPP-Net reconstructs current-stage features, reduces redundant learning of shared patterns, and boosts feature robustness and discriminability. Furthermore, an information gain-driven update strategy is proposed to ensure that the CP-Bank is aligned with the historical mean of pattern distributions, preventing excessive bias toward transient local patterns. To enhance training stability, a dynamic regulation function is designed to adaptively modulate the impact of pseudo-labels according to their confidence, thereby mitigating the adverse effects of low-confidence data. Through extensive experiments on various 2D/3D medical image segmentation datasets, CPP-Net demonstrates its effectiveness and generalizability, and achieves 7.5% mean Dice improvement over SOTA. Lexin Fang, Yunyang Xu, Anxin Zhang, Xin Li 0003, Xuemei Li 0001, Caiming Zhang 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2026 | NeuBase: Spline Surfaces with Neural Basis FunctionsabstractWe introduce NeuBase , a neural parametric surface representation that both accurately fits target surfaces with fine geometric detail and supports intuitive real time surface deformation. NeuBase consists of a Catmull-Clark subdivision base surface and an offset field defined by a set of neural basis functions encoded via a neural map. By construction, NeuBase surfaces exhibit four fundamental geometric properties, i.e., linearity, locality, smoothness, and affine equivariance, enabling real-time, direct manipulation without retraining the neural network. In addition, we propose a scalable neural map that maintains memory efficiency even for complex shapes with dense control meshes. Experiments on a large-scale dataset demonstrate that our method achieves better fitting accuracy than state-of-the-art neural parametric surface representations. Anshul Mendiratta, Lei Yang 0048, Xin Li 0003, John Keyser, Scott Schaefer, Wenping Wang 0001 |
ACM Trans. Graph. | 3 |
| 2026 | NeuPPS: Neural Piecewise Parametric SurfacesabstractPiecewise parametric surfaces have long been established as prevalent geometric representations; however, they often require surface refinement or sophisticated quadrangulation to accurately represent complex geometries. Geometric deep learning has shown that neural networks can provide greater representational power than conventional methods. Nevertheless, approaches using a single parametric surface for shape fitting struggle to capture fine-grained geometric details, while multi-patch methods fail to ensure seamless connections between adjacent patches. We present Neural Piecewise Parametric Surfaces ( NeuPPS ), the first piecewise neural surface representation that allows for coarse patch layouts composed of arbitrary n -sided surface patches to model complex surface geometries with high precision, offering enhanced flexibility compared with traditional parametric surfaces. This new surface representation guarantees, by construction, the continuity between adjacent patches, a property that other neural patch-based approaches cannot ensure. Two novel components are introduced: a learnable feature complex and a continuous mapping function approximated by multi-layer perceptrons (MLPs). We apply the proposed NeuPPS to surface fitting and shape space learning tasks. Extensive experiments demonstrate the advantages of NeuPPS over traditional parametric representations and existing patch-based learning approaches. Lei Yang 0048, Yongqing Liang 0001, Xin Li 0003, Congyi Zhang 0001, Guying Lin, Cheng Lin 0001, Alla Sheffer, Scott Schaefer, John Keyser, Wenping Wang 0001 |
ACM Trans. Graph. | 3 |
| 2026 | DecoRec: Decomposed 3D Scene Reconstruction From Single-View Images via Object-Level DiffusionabstractIn this paper, we introduce DecoRec, a novel system designed to elevate single-view 2D images to a decomposed 3D scene mesh. Current methods for single-view scene reconstruction typically rely on object retrieval or the regression of coarse 3D voxels or surfaces, leading to inaccuracies in capturing the appearance and geometry of the input image. The lack of high-quality large-scale scene-level datasets further complicates direct 3D scene generation from single-view images. To achieve high-quality 3D scene generation from a single-view image, DecoRec takes advantage of recent diffusion-based single-view object reconstruction methods to reconstruct individual objects separately. Subsequently, a refinement pipeline is proposed to effectively merge these reconstructed objects, enhancing appearance and geometry through a differentiable rendering technique and diffusion-guided refinement. Our results demonstrate that DecoRec facilitates high-quality single-view scene reconstruction in both geometry and novel synthesis, offering significant benefits for downstream applications like room interior design. Yuhan Ping, Yuan Liu 0025, Xiaoxiao Long, Peng Wang 0099, Junhui Hou, Jianyi Zheng, Jia Pan 0001, Xin Li 0003, Cheng Lin 0001 |
IEEE Trans. Vis. Comput. Graph. | 8 |
| 2026 | SVGS: Enhancing Gaussian Splatting Using Primitives With Spatially Varying ColorsabstractGaussian Splatting demonstrates impressive results in multi-view reconstruction based on Gaussian explicit representations. However, the current Gaussian primitives only have a single view-dependent color and an opacity to represent the appearance and geometry of the scene, resulting in a non-compact representation. In this paper, we introduce a new method called SVGS (Spatially Varying Gaussian Splatting) that utilizes spatially varying colors and opacity in a single Gaussian primitive to improve its representation ability. We have implemented bilinear interpolation, movable kernels, and tiny neural networks as spatially varying functions. SVGS employs 2D Gaussian surfels as primitives, which significantly enhances novel-view synthesis while maintaining high-quality geometric reconstruction. This approach is particularly effective in practical applications, as scenes combining complex textures with relatively simple geometry occur frequently in real-world environments. Quantitative and qualitative experimental results demonstrate that all three functions outperform the baseline, with the best movable kernels achieving superior novel view synthesis performance on multiple datasets, highlighting the strong potential of spatially varying functions. Rui Xu 0016, Wenyue Chen, Jiepeng Wang 0001, Yuan Liu 0025, Peng Wang 0099, Cheng Lin 0001, Shi-Qing Xin, Xin Li 0003, Wenping Wang 0001, Taku Komura |
IEEE Trans. Vis. Comput. Graph. | 8 |
| 2025 | V2X-R: Cooperative LiDAR-4D Radar Fusion with Denoising Diffusion for 3D Object DetectionabstractCurrent Vehicle-to-Everything (V2X) systems have significantly enhanced 3D object detection using LiDAR and camera data. However, they face performance degradation in adverse weather. Weather-robust 4D radar, with Doppler velocity and additional geometric information, offers a promising solution to this challenge. To this end, we present V2X-R, the first simulated V2X dataset incorporating LiDAR, camera, and 4D radar modalities. V2X-R contains 12,079 scenarios with 37,727 frames of LiDAR and 4D radar point clouds, 150,908 images, and 170,859 annotated 3D vehicle bounding boxes. Subsequently, we propose a novel cooperative LiDAR-4D radar fusion pipeline for 3D object detection and implement it with multiple fusion strategies. To achieve weather-robust detection, we additionally propose a Multi-modal Denoising Diffusion (MDD) module in our fusion pipeline. MDD utilizes weather-robust 4D radar feature as a condition to guide the diffusion model in denoising noisy LiDAR features. Experiments show that our LiDAR-4D radar fusion pipeline demonstrates superior performance in the V2X-R dataset. Over and above this, our MDD module further improved the foggy/snowy performance of the basic fusion model by up to 5.73%/6.70% and barely disrupting normal performance. The dataset and code will be publicly available at: https://github.com/ylwhxht/V2X-R. Xun Huang 0003, Qiming Xia, Siheng Chen, Bisheng Yang, Xin Li 0003, Cheng Wang 0003, Chenglu Wen |
CVPR | 6 |
| 2025 | CADDreamer: CAD Object Generation from Single-view ImagesabstractDiffusion-based 3D generation has made remarkable progress in recent years. However, existing 3D generative models often produce overly dense and unstructured meshes, which stand in stark contrast to the compact, structured, and sharply-edged Computer-Aided Design (CAD) models crafted by human designers. To address this gap, we introduce CADDreamer, a novel approach for generating boundary representations (B-rep) of CAD objects from a single image. CADDreamer employs a primitive-aware multi-view diffusion model that captures both local geometric details and high-level structural semantics during the generation process. By encoding primitive semantics into the color domain, the method leverages the strong priors of pre-trained diffusion models to align with well-defined primitives. This enables the inference of multi-view normal maps and semantic maps from a single image, facilitating the reconstruction of a mesh with primitive labels. Furthermore, we introduce geometric optimization techniques and topology-preserving extraction methods to mitigate noise and distortion in the generated primitives. These enhancements result in a complete and seamless B-rep of the CAD model. Experimental results demonstrate that our method effectively recovers high-quality CAD objects from single-view images. Compared to existing 3D generation techniques, the B-rep models produced by CADDreamer are compact in representation, clear in structure, sharp in edges, and watertight in topology. Cheng Lin 0001, Yuan Liu 0025, Xiaoxiao Long, Ningna Wang, Xin Li 0003, Wenping Wang 0001, Xiaohu Guo |
CVPR | 7 |
| 2025 | Motal: Unsupervised 3D Object Detection by Modality and Task-Specific Knowledge Transfer
Xusheng Guo, Xin Li 0003, Cheng Wang 0003, Chenglu Wen |
ICCV | 4 |
| 2025 | PolaFormer: Polarity-aware Linear Attention for Vision TransformersabstractLinear attention has emerged as a promising alternative to softmax-based attention, leveraging kernelized feature maps to reduce complexity from quadratic to linear in sequence length. However, the non-negative constraint on feature maps and the relaxed exponential function used in approximation lead to significant information loss compared to the original query-key dot products, resulting in less discriminative attention maps with higher entropy. To address the missing interactions driven by negative values in query-key pairs, we propose a polarity-aware linear attention mechanism that explicitly models both same-signed and opposite-signed query-key interactions, ensuring comprehensive coverage of relational information. Furthermore, to restore the spiky properties of attention maps, we provide a theoretical analysis proving the existence of a class of element-wise functions (with positive first and second derivatives) that can reduce entropy in the attention distribution. For simplicity, and recognizing the distinct contributions of each dimension, we employ a learnable power function for rescaling, allowing strong and weak attention signals to be effectively separated. Extensive experiments demonstrate that the proposed PolaFormer improves performance on various vision tasks, enhancing both expressiveness and efficiency by up to 4.6%. Weikang Meng, Yadan Luo, Xin Li 0003, Dongmei Jiang, Zheng Zhang 0006 |
ICLR | 3 |
| 2025 | Multi-Task Label Discovery via Hierarchical Task Tokens for Partially Annotated Dense PredictionsabstractIn recent years, simultaneous learning of multiple dense prediction tasks with partially annotated label data has emerged as an important research area. Previous works primarily focus on leveraging cross-task relations or conducting adversarial training for extra regularization, which achieve promising performance improvements, while still suffering from the lack of direct pixel-wise supervision and extra training of heavy mapping networks. To effectively tackle this challenge, we propose a novel approach to optimize a set of compact learnable hierarchical task tokens, including global and fine-grained ones, to discover consistent pixel-wise supervision signals in both feature and prediction levels. Specifically, the global task tokens are designed for effective cross-task feature interactions in a global context. Then, a group of fine-grained task-specific spatial tokens for each task is learned from the corresponding global task tokens. It is embedded to have dense interactions with each task-specific feature map. The learned global and local fine-grained task tokens are further used to discover pseudo task-specific dense labels at different levels of granularity, and they can be utilized to directly supervise the learning of the multi-task dense prediction framework. Extensive experimental results on challenging NYUD-v2, Cityscapes, and PASCAL Context datasets demonstrate significant improvements over existing state-of-the-art methods for partially annotated multi-task dense prediction. Jingdong Zhang 0003, Hanrong Ye, Xin Li 0003, Wenping Wang 0001, Dan Xu 0002 |
ACM Multimedia | 3 |
| 2025 | SPGen: Spherical Projection as Consistent and Flexible Representation for Single Image 3D Shape GenerationabstractExisting single-view 3D generative models typically adopt multiview diffusion priors to reconstruct object surfaces, yet they remain prone to inter-view inconsistencies and are unable to faithfully represent complex internal structure or nontrivial topologies. In particular, we encode geometry information by projecting it onto a bounding sphere and unwrapping it into a compact and structural multi-layer 2D Spherical Projection (SP) representation. Operating solely in the image domain, SPGen offers three key advantages simultaneously: (1) Consistency. The injective SP mapping encodes surface geometry with a single viewpoint which naturally eliminates view inconsistency and ambiguity; (2) Flexibility. Multi-layer SP maps represent nested internal structures and support direct lifting to watertight or open 3D surfaces; (3) Efficiency. The image-domain formulation allows the direct inheritance of powerful 2D diffusion priors and enables efficient finetuning with limited computational resources. Extensive experiments demonstrate that SPGen significantly outperforms existing baselines in geometric quality and computational efficiency. Jingdong Zhang 0003, Weikai Chen 0001, Yuan Liu 0025, Jionghao Wang, Zhengming Yu, Zhuowen Shen, Bo Yang 0070, Wenping Wang 0001, Xin Li 0003 |
SIGGRAPH Asia | 9 |
| 2025 | A Survey on Computational Solutions for Reconstructing Complete Objects by Reassembling Their Fractured PartsabstractAbstract Reconstructing a complete object from its parts is a fundamental problem in many scientific domains. The purpose of this article is to provide a systematic survey on this topic. This reassembly problem requires understanding the attributes of individual pieces and establishing matches between different pieces. Many approaches also model priors of the underlying complete object. Existing approaches are tightly connected problems of shape segmentation, shape matching, and learning shape priors. We provide existing algorithms in this context and emphasize their similarities and differences to general‐purpose approaches. We also survey the trends from early procedural approaches to more recent deep learning approaches. In addition to algorithms, this survey will also describe existing datasets, open‐source software packages, and applications. To the best of our knowledge, this is the first comprehensive survey on this topic in computer graphics. Jiaxin Lu 0001, Yongqing Liang 0001, Huijun Han, Jiacheng Hua, Xin Li 0003, Qixing Huang |
Comput. Graph. Forum | 6 |
| 2025 | SEArch: A self-evolving framework for network architecture optimizationabstractThis paper studies a fundamental network optimization problem that finds a network architecture with optimal performance (low loss) under given resource budgets (small number of parameters and/or fast inference). Unlike existing network optimization approaches such as network pruning, knowledge distillation (KD), and network architecture search (NAS), in this work we introduce a self-evolving pipeline to perform network optimization. In this framework, a simple network iteratively and adaptively modifies its structures by using the guidance from a teacher network, until it reaches the resource budget. An attention module is introduced to transfer the knowledge from the teacher network to the student network. A splitting edge scheme is designed to help the student model find an optimal macro architecture. The proposed framework combines the advantages of pruning, KD, and NAS, and hence, can efficiently generate networks with flexible structure and desirable performance. Extensive experiments on CIFAR-10, CIFAR-100, and ImageNet demonstrated that our framework achieves great performance in this network architecture optimization task. Yongqing Liang 0001, Dawei Xiang, Xin Li 0003 |
Neurocomputing | 3 |
| 2025 | Spectrum-guided Spatial Feature Enhancement Network for event-based lip-reading
Yi Zhang 0100, Xiuping Liu, Hongchen Tan, Xin Li 0003 |
Neurocomputing | 4 |
| 2025 | Hierarchical Event-RGB Interaction Network for single-eye expression recognition
Runduo Han, Xiuping Liu, Yi Zhang 0100, Hongchen Tan, Xin Li 0003 |
Inf. Sci. | 6 |
| 2025 | Unsupervised 3D Object Detection by Commonsense ClueabstractTraditional 3D object detectors, whether fully-, semi-, or weakly-supervised, rely heavily on extensive human annotations. In contrast, this paper introduces an unsupervised 3D object detector that automatically discerns object patterns without such annotations. To achieve this, we propose a Commonsense Prototype-based Detector (CPD) for unsupervised 3D object detection. CPD first constructs Commonsense Prototypes (CProto) to represent the geometric center and size of objects. It then generates high-quality pseudo-labels and guides detector convergence using size and geometry priors from CProto. Building on CPD, we further introduce CPD++, an enhanced version that improves performance by leveraging motion cues. CPD++ learns localization from stationary objects and recognition from moving objects, facilitating the mutual transfer of localization and recognition knowledge between these two object types. Both CPD and CPD++ outperform existing state-of-the-art unsupervised 3D detectors. Furthermore, when trained on Waymo Open Dataset (WOD) and tested on KITTI, CPD++ achieves 89.25% 3D Average Precision (AP) on the moderate car class at a 0.5 IoU threshold, reaching 95.3% of the performance attained by fully supervised counterparts. These results underscore the significant advancements brought by our method. Shijia Zhao, Xun Huang 0003, Qiming Xia, Chenglu Wen, Li Jiang 0009, Xin Li 0003, Cheng Wang 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2025 | SC3EF: A Joint Self-Correlation and Cross-Correspondence Estimation Framework for Visible and Thermal Image RegistrationabstractMultispectral imaging plays a critical role in a range of intelligent transportation applications, including advanced driver assistance systems (ADAS), traffic monitoring, and night vision. However, accurate visible and thermal (RGB-T) image registration poses a significant challenge due to the considerable modality differences. In this paper, we present a novel joint Self-Correlation and Cross-Correspondence Estimation Framework (SC3EF), leveraging both local representative features and global contextual cues to effectively generate RGB-T correspondences. For this purpose, we design a convolution-transformer-based pipeline to extract local representative features and encode global correlations of intra-modality for inter-modality correspondence estimation between unaligned visible and thermal images. After merging the local and global correspondence estimation results, we further employ a hierarchical optical flow estimation decoder to progressively refine the estimated dense correspondence maps. Extensive experiments demonstrate the effectiveness of our proposed method, outperforming the current state-of-the-art (SOTA) methods on representative RGB-T datasets. Furthermore, it also shows competitive generalization capabilities across challenging scenarios, including large parallax, severe occlusions, adverse weather, and other cross-modal datasets (e.g., RGB-N and RGB-D). Xi Tong, Jiangxin Yang, Xin Li 0003, Yanpeng Cao |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2025 | CrossGen: Learning and Generating Cross Fields for Quad MeshingabstractCross fields play a critical role in various geometry processing tasks, especially for quad mesh generation. Existing methods for cross field generation often struggle to balance computational efficiency with generation quality, using slow per-shape optimization. We introduce CrossGen , a novel framework that supports both feed-forward prediction and latent generative modeling of cross fields for quad meshing by unifying geometry and cross field representations within a joint latent space. Our method enables extremely fast computation of high-quality cross fields of general input shapes, typically within one second without per-shape optimization. Our method assumes a point-sampled surface, also called a point-cloud surface , as input, so we can accommodate various surface representations by a straightforward point sampling process. Using an auto-encoder network architecture, we encode input point-cloud surfaces into a sparse voxel grid with fine-grained latent spaces, which are decoded into both SDF-based surface geometry and cross fields (see the teaser figure). We also contribute a dataset of models with both high-quality signed distance fields (SDFs) representations and their corresponding cross fields, and use it to train our network. Once trained, the network is capable of computing a cross field of an input surface in a feed-forward manner, ensuring high geometric fidelity, noise resilience, and rapid inference. Furthermore, leveraging the same unified latent representation, we incorporate a diffusion model for computing cross fields of new shapes generated from partial input, such as sketches. To demonstrate its practical applications, we validate CrossGen on the quad mesh generation task for a large variety of surface shapes. Experimental results demonstrate that CrossGen generalizes well across diverse shapes and consistently yields high-fidelity cross fields, thus facilitating the generation of high-quality quad meshes. Qiujie Dong, Jiepeng Wang 0001, Rui Xu 0016, Cheng Lin 0001, Yuan Liu 0025, Shi-Qing Xin, Zichun Zhong, Xin Li 0003, Changhe Tu, Taku Komura, Leif Kobbelt, Scott Schaefer, Wenping Wang 0001 |
ACM Trans. Graph. | 8 |
| 2025 | NeuVAS: Neural Implicit Surfaces for Variational Shape ModelingabstractNeural implicit shape representation has drawn significant attention in recent years due to its smoothness, differentiability, and topological flexibility. However, directly modeling the shape of a neural implicit surface, especially as the zero-level set of a neural signed distance function (SDF), with sparse geometric control is still a challenging task. Sparse input shape control typically includes 3D curve networks or, more generally, 3D curve sketches, which are unstructured and cannot be connected to form a curve network, and therefore more difficult to deal with. While 3D curve networks or curve sketches provide intuitive shape control, their sparsity and varied topology pose challenges in generating high-quality surfaces to meet such curve constraints. In this paper, we propose NeuVAS, a variational approach to shape modeling using neural implicit surfaces constrained under sparse input shape control, including unstructured 3D curve sketches as well as connected 3D curve networks. Specifically, we introduce a smoothness term based on a functional of surface curvatures to minimize shape variation of the zero-level set surface of a neural SDF. We also develop a new technique to faithfully model G 0 sharp feature curves as specified in the input curve sketches. Comprehensive comparisons with the state-of-the-art methods demonstrate the significant advantages of our method. Qiujie Dong, Fangtian Liang, Hao Pan 0001, Lei Yang 0048, Congyi Zhang 0001, Guying Lin, Caiming Zhang 0001, Yuanfeng Zhou, Changhe Tu, Shi-Qing Xin, Alla Sheffer, Xin Li 0003, Wenping Wang 0001 |
ACM Trans. Graph. | 13 |
| 2025 | Skull-to-Face: Anatomy-Guided 3D Facial Reconstruction and EditingabstractDeducing the 3D face from a skull is a challenging task in forensic science and archaeology. This article proposes an end-to-end 3D face reconstruction pipeline and an exploration method that can conveniently create textured, realistic faces that match the given skull. To this end, we propose a tissue-guided face creation and adaptation scheme. With the help of the state-of-the-art text-to-image diffusion model and parametric face model, we first generate an initial reference 3D face, whose biological profile aligns with the given skull. Then, with the help of tissue thickness distribution, we modify these initial faces to match the skull through a latent optimization process. The joint distribution of tissue thickness is learned on a set of skull landmarks using a collection of scanned skull-face pairs. We also develop an efficient face adaptation tool to allow users to interactively adjust tissue thickness either globally or at local regions to explore different plausible faces. Experiments conducted on a real skull-face dataset demonstrated the effectiveness of our proposed pipeline in terms of reconstruction accuracy, diversity, and stability. Yongqing Liang 0001, Congyi Zhang 0001, Junli Zhao, Wenping Wang 0001, Xin Li 0003 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2025 | WonderHuman: Hallucinating Unseen Parts in Dynamic 3D Human ReconstructionabstractIn this paper, we present WonderHuman to reconstruct dynamic human avatars from a monocular video for high-fidelity novel view synthesis. Previous dynamic human avatar reconstruction methods typically require the input video to have full coverage of the observed human body. However, in daily practice, one typically has access to limited viewpoints, such as monocular front-view videos, making it a cumbersome task for previous methods to reconstruct the unseen parts of the human avatar. To tackle the issue, we present WonderHuman, which leverages 2D generative diffusion model priors to achieve high-quality, photorealistic reconstructions of dynamic human avatars from monocular videos, including accurate rendering of unseen body parts. Our approach introduces a Dual-Space Optimization technique, applying Score Distillation Sampling (SDS) in both canonical and observation spaces to ensure visual consistency and enhance realism in dynamic human reconstruction. Additionally, we present a View Selection strategy and Pose Feature Injection to enforce the consistency between SDS predictions and observed data, ensuring pose-dependent effects and higher fidelity in the reconstructed avatar. In the experiments, our method achieves SOTA performance in producing photorealistic renderings from the given monocular video, particularly for those challenging unseen parts. Zilong Wang 0013, Zhiyang Dou, Yuan Liu 0025, Cheng Lin 0001, Yunhui Guo, Xin Li 0003, Wenping Wang 0001, Xiaohu Guo |
IEEE Trans. Vis. Comput. Graph. | 8 |
| 2024 | Sunshine to Rainstorm: Cross-Weather Knowledge Distillation for Robust 3D Object DetectionabstractLiDAR-based 3D object detection models inevitably struggle under rainy conditions due to the degraded and noisy scanning signals. Previous research has attempted to address this by simulating the noise from rain to improve the robustness of detection models. However, significant disparities exist between simulated and actual rain-impacted data points. In this work, we propose a novel rain simulation method, termed DRET, that unifies Dynamics and Rainy Environment Theory to provide a cost-effective means of expanding the available realistic rain data for 3D detection training. Furthermore, we present a Sunny-to-Rainy Knowledge Distillation (SRKD) approach to enhance 3D detection under rainy conditions. Extensive experiments on the Waymo-Open-Dataset show that, when combined with the state-of-the-art DSVT model and other classical 3D detectors, our proposed framework demonstrates significant detection accuracy improvements, without losing efficiency. Remarkably, our framework also improves detection capabilities under sunny conditions, therefore offering a robust solution for 3D detection regardless of whether the weather is rainy or sunny. Xun Huang 0003, Xin Li 0003, Xiaoliang Fan, Chenglu Wen, Cheng Wang 0003 |
AAAI | 3 |
| 2024 | Commonsense Prototype for Outdoor Unsupervised 3D Object DetectionabstractThe prevalent approaches of unsupervised 3D object de-tection follow cluster-based pseudo-label generation and iterative self-training processes. However, the challenge arises due to the sparsity of LiDAR scans, which leads to pseudo-labels with erroneous size and position, resulting in subpar detection performance. To tackle this problem, this paper introduces a Commonsense Prototype-based Detector, termed CPD, for unsupervised 3D object de-tection. CPD first constructs Commonsense Prototype (CProto) characterized by high-quality bounding box and dense points, based on commonsense intuition. Subse-quently, CPD refines the low-quality pseudo-labels by lever-aging the size prior from CProto. Furthermore, CPD en-hances the detection accuracy of sparsely scanned objects by the geometric knowledge from CProto. CPD outper-forms state-of-the-art unsupervised 3D detectors on Waymo Open Dataset (WOD), PandaSet, and KITTI datasets by a large margin. Besides, by training CPD on WOD and testing on KITTI, CPD attains 90.85% and 81.01% 3D Aver-age Precision on easy and moderate car classes, respectively. These achievements position CPD in close prox-imity to fully supervised detectors, highlighting the sig-nificance of our method. The code will be available at https://github.com/hailanyi/CPD. Shijia Zhao, Xun Huang 0003, Chenglu Wen, Xin Li 0003, Cheng Wang 0003 |
CVPR | 5 |
| 2024 | HINTED: Hard Instance Enhanced Detector with Mixed-Density Feature Fusion for Sparsely-Supervised 3D Object DetectionabstractCurrent sparsely-supervised object detection methods largely depend on high threshold settings to derive high-quality pseudo labels from detector predictions. However, hard instances within point clouds frequently display incomplete structures, causing decreased confidence scores in their assigned pseudo-labels. Previous methods inevitably result in inadequate positive supervision for these instances. To address this problem, we propose a novel Hard INsTance Enhanced Detector (HINTED), for sparsely-supervised 3D object detection. Firstly, we design a self-boosting teacher (SBT) model to generate more potential pseudo-labels, enhancing the effectiveness of information transfer. Then, we introduce a mixed-density student (MDS) model to concentrate on hard instances during the training phase, thereby improving detection accuracy. Our extensive experiments on the KITTI dataset validate our method's superior performance. Compared with leading sparsely-supervised methods, HINTED significantly improves the detection performance on hard instances, no-tably outperforming fully-supervised methods in detecting challenging categories like cyclists. HINTED also significantly outperforms the state-of-the-art semi-supervised method on challenging categories. The code is available at https://github.com/xmuqimingxia/HINTED. Qiming Xia, Shijia Zhao, Leyuan Xing, Xun Huang 0003, Jinhao Deng, Xin Li 0003, Chenglu Wen, Cheng Wang 0003 |
CVPR | 8 |
| 2024 | CMD: A Cross Mechanism Domain Adaptation Dataset for 3D Object Detection
Jinhao Deng, Xun Huang 0003, Qiming Xia, Xin Li 0003, Wei Li 0111, Chenglu Wen, Cheng Wang 0003 |
ECCV (57) | 6 |
| 2024 | Disentangled Clothed Avatar Generation from Text Descriptions
Jionghao Wang, Yuan Liu 0025, Zhiyang Dou, Zhengming Yu, Yongqing Liang 0001, Cheng Lin 0001, Rong Xie 0004, Li Song 0001, Xin Li 0003, Wenping Wang 0001 |
ECCV (52) | 9 |
| 2024 | Surf-D: Generating High-Quality Surfaces of Arbitrary Topologies Using Diffusion Models
Zhengming Yu, Zhiyang Dou, Xiaoxiao Long, Cheng Lin 0001, Zekun Li 0002, Yuan Liu 0025, Norman Müller, Taku Komura, Marc Habermann, Christian Theobalt, Xin Li 0003, Wenping Wang 0001 |
ECCV (39) | 11 |
| 2024 | PointSmile: point self-supervised learning via curriculum mutual information
Xin Li 0003, Mingqiang Wei, Songcan Chen |
Sci. China Inf. Sci. | 1 |
| 2024 | Deep video representation learning: a survey
Elham Ravanbakhsh, Yongqing Liang 0001, J. Ramanujam, Xin Li 0003 |
Multim. Tools Appl. | 4 |
| 2024 | Attention-Bridged Modal Interaction for Text-to-Image GenerationabstractWe propose a novel Text-to-Image Generation Network, Attention-bridged Modal Interaction Generative Adversarial Network (AMI-GAN), to better explore modal interaction and perception for high-quality image synthesis. The AMI-GAN contains two novel designs: an Attention-bridged Modal Interaction (AMI) module and a Residual Perception Discriminator (RPD). In AMI, we mainly design a multi-scale attention mechanism to exploit semantics alignment, fusion, and enhancement between text and image, to better refine details and context semantics of the synthesized image. In RPD, we design a multi-scale information perception mechanism with our proposed novel information adjustment function, to encourage the discriminator to better perceive visual differences between the real and synthesized image. Consequently, the discriminator will drive the generator to improve the visual quality of the synthesized image. Besides, based on these novel designs, we can design two versions, a single-stage generation framework (AMI-GAN-S), and a multi-stage generation framework (AMI-GAN-M), respectively. The former can synthesize high-resolution images because of its low computational cost; the latter can synthesize images with realistic detail. Experimental results on two widely used T2I datasets showed that our AMI-GANs achieve competitive performance in T2I task. Hongchen Tan, Kaiqiang Xu, Huasheng Wang, Xiuping Liu, Xin Li 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2023 | Transformation-Equivariant 3D Object Detection for Autonomous Drivingabstract3D object detection received increasing attention in autonomous driving recently. Objects in 3D scenes are distributed with diverse orientations. Ordinary detectors do not explicitly model the variations of rotation and reflection transformations. Consequently, large networks and extensive data augmentation are required for robust detection. Recent equivariant networks explicitly model the transformation variations by applying shared networks on multiple transformed point clouds, showing great potential in object geometry modeling. However, it is difficult to apply such networks to 3D object detection in autonomous driving due to its large computation cost and slow reasoning speed. In this work, we present TED, an efficient Transformation-Equivariant 3D Detector to overcome the computation cost and speed issues. TED first applies a sparse convolution backbone to extract multi-channel transformation-equivariant voxel features; and then aligns and aggregates these equivariant features into lightweight and compact representations for high-performance 3D object detection. On the highly competitive KITTI 3D car detection leaderboard, TED ranked 1st among all submissions with competitive efficiency. Code is available at https://github.com/hailanyi/TED. Chenglu Wen, Wei Li 0111, Xin Li 0003, Ruigang Yang, Cheng Wang 0003 |
AAAI | 4 |
| 2023 | Virtual Sparse Convolution for Multimodal 3D Object DetectionabstractRecently, virtuall pseudo-point-based 3D object detection that seamlessly fuses RGB images and LiDAR data by depth completion has gained great attention. However, virtual points generated from an image are very dense, introducing a huge amount of redundant computation during detection. Meanwhile, noises brought by inaccurate depth completion significantly degrade detection precision. This paper proposes a fast yet effective backbone, termed Vir-ConvNet, based on a new operator VirConv (Virtual Sparse Convolution), for virtual-point-based 3D object detection. VirConv consists of two key designs: (1) StVD (Stochastic Voxel Discard) and (2) NRConv (Noise-Resistant Sub-manifold Convolution). StVD alleviates the computation problem by discarding large amounts of nearby redundant voxels. NRConv tackles the noise problem by encoding voxel features in both 2D image and 3D LiDAR space. By integrating VirConv, we first develop an efficient pipeline VirConv-L based on an early fusion design. Then, we build a high-precision pipeline Vir Conv-T based on a transformed refinement scheme. Finally, we develop a semi-supervised pipeline VirConv-S based on a pseudo-label framework. On the KITTI car 3D detection test leader-board, our VirConv-L achieves 85% AP with a fast running speed of 56ms. Our VirConv-T and VirConv-S attains a high-precision of 86.3% and 87.2% AP, and currently rank 2ndand 1st11On the date of CVPR deadline, i. e., Nov.11, 2022, respectively. The code is available at https://github.com/hailanyi/VirConv. Chenglu Wen, Shaoshuai Shi, Xin Li 0003, Cheng Wang 0003 |
CVPR | 4 |
| 2023 | Batch-based Model Registration for Fast 3D Sherd Reconstructionabstract3D reconstruction techniques have widely been used for digital documentation of archaeological fragments. However, efficient digital capture of fragments remains as a challenge. In this work, we aim to develop a portable, high-throughput, and accurate reconstruction system for efficient digitization of fragments excavated in archaeological sites. To realize high-throughput digitization of large numbers of objects, an effective strategy is to perform scanning and reconstruction in batches. However, effective batch-based scanning and reconstruction face two key challenges: 1) how to correlate partial scans of the same object from multiple batch scans, and 2) how to register and reconstruct complete models from partial scans that exhibit only small overlaps. To tackle these two challenges, we develop a new batch-based matching algorithm that pairs the front and back sides of the fragments, and a new Bilateral Boundary ICP algorithm that can register partial scans sharing very narrow overlapping regions. Extensive validation in labs and testing in excavation sites demonstrate that these designs enable efficient batch-based scanning for fragments. We show that such a batch-based scanning and reconstruction pipeline can have immediate applications on digitizing sherds in archaeological excavations. Our project page: https://jiepengwang.github.io/FIRES/. Jiepeng Wang 0001, Congyi Zhang 0001, Peng Wang 0099, Xin Li 0003, Peter J. Cobb, Christian Theobalt, Wenping Wang 0001 |
ICCV | 4 |
| 2023 | CoIn: Contrastive Instance Feature Mining for Outdoor 3D Object Detection with Very Limited AnnotationsabstractRecently, 3D object detection with sparse annotations has received great attention. However, current detectors usually perform poorly under very limited annotations. To address this problem, we propose a novel Contrastive Instance feature mining method, named CoIn. To better identify indistinguishable features learned through limited supervision, we design a Multi-Class contrastive learning module (MCcont) to enhance feature discrimination. Meanwhile, we propose a feature-level pseudo-label mining framework consisting of an instance feature mining module (InF-Mining) and a Labeled-to-Pseudo contrastive learning module (LPcont). These two modules exploit latent instances in feature space to supervise the training of detectors with limited annotations. Extensive experiments with KITTI dataset, Waymo open dataset, and nuScenes dataset show that under limited annotations, our method greatly improves the performance of baseline detectors: CenterPoint, Voxel-RCNN, and CasA. Combining CoIn with an iterative training strategy, we propose a CoIn++ pipeline, which requires only 2% annotations in the KITTI dataset to achieve performance comparable to the fully supervised methods. The code is available at https://github.com/xmuqimingxia/CoIn. Qiming Xia, Jinhao Deng, Chenglu Wen, Shaoshuai Shi, Xin Li 0003, Cheng Wang 0003 |
ICCV | 6 |
| 2023 | Surface Extraction from Neural Unsigned Distance FieldsabstractWe propose a method, named DualMesh-UDF, to extract a surface from unsigned distance functions (UDFs), encoded by neural networks, or neural UDFs. Neural UDFs are becoming increasingly popular for surface representation because of their versatility in presenting surfaces with arbitrary topologies, as opposed to the signed distance function that is limited to representing a closed surface. However, the applications of neural UDFs are hindered by the notorious difficulty in extracting the target surfaces they represent. Recent methods for surface extraction from a neural UDF suffer from significant geometric errors or topological artifacts due to two main difficulties: (1) A UDF does not exhibit sign changes; and (2) A neural UDF typically has substantial approximation errors.DualMesh-UDF addresses these two difficulties. Specifically, given a neural UDF encoding a target surface $\bar S$ to be recovered, we first estimate the tangent planes of $\bar S$ at a set of sample points close to $\bar S$. Next, we organize these sample points into local clusters, and for each local cluster, solve a linear least squares problem to determine a final surface point. These surface points are then connected to create the output mesh surface, which approximates the target surface. The robust estimation of the tangent planes of the target surface and the subsequent minimization problem constitute our core strategy, which contributes to the favorable performance of DualMesh-UDF over other competing methods. To efficiently implement this strategy, we employ an adaptive Octree. Within this framework, we estimate the location of a surface point in each of the octree cells identified as containing part of the target surface. Extensive experiments show that our method outperforms existing methods in terms of surface reconstruction quality while maintaining comparable computational efficiency. Congyi Zhang 0001, Guying Lin, Lei Yang 0048, Xin Li 0003, Taku Komura, Scott Schaefer, John Keyser, Wenping Wang 0001 |
ICCV | 4 |
| 2023 | Robust 3D Craniofacial Landmarks Localization by An End-to-End Regression NetworkabstractLandmark localization plays a significant role in craniofacial registration, reconstruction, and authentication. The key challenges for localizing landmarks on point cloud craniofacial models include irregular structures, non-uniform densities, and uncertain local regions. In this paper, we propose an end-to-end regression network that can directly estimate craniofacial landmarks on point cloud models. The proposed network utilizes edge convolution to extract local features and pooling layers to aggregate global features. It realizes the end-to-end regression for landmark localization. Experimental results demonstrate that our method is robust on point clouds with sparse and unevenly distributed sampling. It can produce accurate, controllable, and efficient 3D landmarks. Xianhe Jiao, Junli Zhao, Chenlei Lv, Fuqing Duan, Zhenkuan Pan 0001, Xin Li 0003 |
ICME | 6 |
| 2023 | Single image super-resolution based on progressive fusion of orientation-aware features
Zewei He, Yanpeng Cao, Jiangxin Yang, Yanlong Cao, Xin Li 0003, Siliang Tang, Yueting Zhuang, Zheming Lu 0001 |
Pattern Recognit. | 6 |
| 2023 | Semantics-enhanced early action detection using dynamic dilated convolution
Matthew Korban, Xin Li 0003 |
Pattern Recognit. | 2 |
| 2023 | DeepSIR: Deep semantic iterative registration for LiDAR point clouds
Qing Li 0032, Cheng Wang 0003, Chenglu Wen, Xin Li 0003 |
Pattern Recognit. | 4 |
| 2023 | A Detector-Oblivious Multi-Arm Network for Keypoint MatchingabstractThis paper presents a matching network to establish point correspondence between images. We propose a Multi-Arm Network (MAN) capable of learning region overlap and depth, which can greatly improve keypoint matching robustness while bringing an extra 50% of computational time during the inference stage. By adopting a different design from the state-of-the-art learning based pipeline SuperGlue framework, which requires retraining when a different keypoint detector is adopted, our network can directly work with different keypoint detectors without time-consuming retraining processes. Comprehensive experiments conducted on four public benchmarks involving both outdoor and indoor scenarios demonstrate that our proposed MAN outperforms state-of-the-art methods. Xuelun Shen, Xin Li 0003, Cheng Wang 0003 |
IEEE Trans. Image Process. | 3 |
| 2023 | ALR-GAN: Adaptive Layout Refinement for Text-to-Image SynthesisabstractWe propose a novel Text-to-Image Generation Network, Adaptive Layout Refinement Generative Adversarial Network (ALR-GAN), to adaptively refine the layout of synthesized images without any auxiliary information. The ALR-GAN includes an Adaptive Layout Refinement (ALR) module and a Layout Visual Refinement (LVR) loss. The ALR module aligns the layout structure (which refers to locations of objects and background) of a synthesized image with that of its corresponding real image. In ALR module, we proposed an Adaptive Layout Refinement (ALR) loss to balance the matching of hard and easy features, for more efficient layout structure matching. Based on the refined layout structure, the LVR loss further refines the visual representation within the layout area. Experimental results on two widely-used datasets show that ALR-GAN performs competitively at the Text-to-Image generation task. Hongchen Tan, Xiuping Liu, Xin Li 0003 |
IEEE Trans. Multim. | 5 |
| 2023 | MHSA-Net: Multihead Self-Attention Network for Occluded Person Re-IdentificationabstractThis article presents a novel person reidentification model, named multihead self-attention network (MHSA-Net), to prune unimportant information and capture key local information from person images. MHSA-Net contains two main novel components: multihead self-attention branch (MHSAB) and attention competition mechanism (ACM). The MHSAB adaptively captures key local person information and then produces effective diversity embeddings of an image for the person matching. The ACM further helps filter out attention noise and nonkey information. Through extensive ablation studies, we verified that the MHSAB and ACM both contribute to the performance improvement of the MHSA-Net. Our MHSA-Net achieves competitive performance in the standard and occluded person Re-ID tasks. Hongchen Tan, Xiuping Liu, Xin Li 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | DR-GAN: Distribution Regularization for Text-to-Image GenerationabstractThis article presents a new text-to-image (T2I) generation model, named distribution regularization generative adversarial network (DR-GAN), to generate images from text descriptions from improved distribution learning. In DR-GAN, we introduce two novel modules: a semantic disentangling module (SDM) and a distribution normalization module (DNM). SDM combines the spatial self-attention mechanism (SSAM) and a new semantic disentangling loss (SDL) to help the generator distill key semantic information for the image generation. DNM uses a variational auto-encoder (VAE) to normalize and denoise the image latent distribution, which can help the discriminator better distinguish synthesized images from real images. DNM also adopts a distribution adversarial loss (DAL) to guide the generator to align with normalized real image distributions in the latent space. Extensive experiments on two public datasets demonstrated that our DR-GAN achieved a competitive performance in the T2I task. The code link: https://github.com/Tan-H-C/DR-GAN-Distribution-Regularization-for-Text-to-Image-Generation. Hongchen Tan, Xiuping Liu, Xin Li 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2022 | BLNet: Bidirectional learning network for point cloudsabstractThe key challenge in processing point clouds lies in the inherent lack of ordering and irregularity of the 3D points. By relying on perpoint multi-layer perceptions (MLPs), most existing point-based approaches only address the first issue yet ignore the second one. Directly convolving kernels with irregular points will result in loss of shape information. This paper introduces a novel point-based bidirectional learning network (BLNet) to analyze irregular 3D points. BLNet optimizes the learning of 3D points through two iterative operations: feature-guided point shifting and feature learning from shifted points, so as to minimise intra-class variances, leading to a more regular distribution. On the other hand, explicitly modeling point positions leads to a new feature encoding with increased structure-awareness. Then, an attention pooling unit selectively combines important features. This bidirectional learning alternately regularizes the point cloud and learns its geometric features, with these two procedures iteratively promoting each other for more effective feature learning. Experiments show that BLNet is able to learn deep point features robustly and efficiently, and outperforms the prior state-of-the-art on multiple challenging tasks. Wenkai Han, Chenglu Wen, Cheng Wang 0003, Xin Li 0003 |
Comput. Vis. Media | 5 |
| 2022 | Learning scale awareness in keypoint extraction and description
Xuelun Shen, Cheng Wang 0003, Xin Li 0003, Yifan Peng 0001, Chenglu Wen, Ming Cheng 0002 |
Pattern Recognit. | 3 |
| 2022 | LiDAR-based localization using universal encoding and memory-aware regression
Shangshu Yu, Cheng Wang 0003, Chenglu Wen, Ming Cheng 0002, Minghao Liu 0007, Zhihong Zhang 0001, Xin Li 0003 |
Pattern Recognit. | 7 |
| 2022 | Corrigendum to "LiDAR-based localization using universal encoding and memory-aware regression" Pattern Recognition Volume 128 (2022) 108685
Shangshu Yu, Cheng Wang 0003, Chenglu Wen, Ming Cheng 0002, Minghao Liu 0007, Zhihong Zhang 0001, Xin Li 0003 |
Pattern Recognit. | 7 |
| 2022 | CasA: A Cascade Attention Network for 3-D Object Detection From LiDAR Point Cloudsabstract3D object detection from LiDAR point clouds has gained great attention in recent years due to its wide applications in smart cities and autonomous driving. Cascade framework shows its advancement in 2D object detection but is less investigated in 3D space. Conventional cascade structures use multipleseparatesub-networks to sequentially refine region proposals. Such methods, however, have limited ability to measure proposal quality in all stages, and hard to achieve a desirable performance improvement in 3D space. This paper proposes a new cascade framework, termed CasA, for 3D object detection from LiDAR point clouds. CasA consists of a Region Proposal Network (RPN) and a Cascade Refinement Network (CRN). In CRN, we designed a new Cascade Attention Module that uses multiple sub-networks and attention modules to aggregate the object features from different stages and progressively refine region proposals. CasA can be integrated into various two-stage 3D detectors and improve their performance. Extensive experiments on KITTI and Waymo datasets with various baseline detectors demonstrate the universality and superiority of our CasA. In particular, based on one variant of Voxel-RCNN, we achieve state-of-the-art results on the KITTI dataset. On the KITTI online 3D object detection leaderboard, we achieve a high detection performance of 83.06%, 47.09%, and 73.47% Average Precision (AP) in the moderate Car, Pedestrian, and Cyclist classes, respectively. Code is available at https://github.com/hailanyi/CasA. Jinhao Deng, Chenglu Wen, Xin Li 0003, Cheng Wang 0003, Jonathan Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Spatio-Temporal 3-D Residual Networks for Simultaneous Detection and Depth Estimation of CFRP Subsurface Defects in Lock-In ThermographyabstractNondestructive thermography is a high-speed, low-cost, and safe solution for subsurface defects detection of carbon fiber reinforced polymer (CFRP) materials, providing essential quality control in aerospace, automobile, and sports industries. In this article, we build a reflective lock-in thermography system and construct a dataset that contains real-captured thermal image sequences of CFRP samples with various simulated internal defects under different excitation frequencies. Then, we present a novel 3-D convolutional neural network (CNN) model incorporating a combination of spatial and temporal convolutional filters and batch-size independent group normalization (GN) as a unified framework to process thermal image sequences captured by lock-in thermography for simultaneous subsurface defect detection and depth estimation. Finally, we define a multitask loss function to perform end-to-end training of both defect detection and depth estimation tasks based on the real-captured infrared sequences. Comparative experiments are carried out on CFRP specimens with artificial defects of various sizes/shapes and at different depths. Qualitative and quantitative results illustrate that our 3-D CNN model is capable of predicting accurate locations and depths of subsurface defects and performs favorably against the hand-crafted and CNN-based methods in lock-in thermography for individual defect detection and depth estimation tasks. The captured dataset and the source codes will be made publicly available. Yafei Dong, Chenjie Xia, Jiangxin Yang, Yanlong Cao, Yanpeng Cao, Xin Li 0003 |
IEEE Trans. Ind. Informatics | 6 |
| 2022 | 3D Multi-Object Tracking in Point Clouds Based on Prediction Confidence-Guided Data AssociationabstractThis paper proposes a new 3D multi-object tracker to more robustly track objects that are temporarily missed by detectors. Our tracker can better leverage object features for 3D Multi-Object Tracking (MOT) in point clouds. The proposed tracker is based on a novel data association scheme guided by prediction confidence, and it consists of two key parts. First, we design a new predictor that employs a constant acceleration (CA) motion model to estimate future positions, and outputs a prediction confidence to guide data association through increased awareness of detection quality. Second, we introduce a new aggregated pairwise cost to exploit features of objects in point clouds for faster and more accurate data association. The proposed cost consists of geometry, appearance and motion components. Specifically, we formulate the geometry cost using resolutions (lengths, widths and heights), centroids, and orientations of 3D bounding boxes (BBs), the appearance cost using appearance features from the deep learning-based detector backbone network, and the motion cost by associating different motion vectors. Extensive multi-object tracking experiments on the KITTI tracking benchmark demonstrated that our method outperforms, by a large margin, the state-of-the-art methods in both tracking accuracy and speed. Wenkai Han, Chenglu Wen, Xin Li 0003, Cheng Wang 0003 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2022 | LAG-Net: Multi-Granularity Network for Person Re-Identification via Local Attention SystemabstractPerson re-identification (Re-ID) is a challenging research topic which aims to retrieve the pedestrian images of the same person that captured by non-overlapping cameras. Existing methods either assume the body parts of the same person are well-aligned, or use attention selection mechanisms to constrain the effective region of feature learning. But these methods concentrate only on coarse feature representation and cannot model complex real scenes effectively. We propose a novel Local Attention Guided Network (LAG-Net) to not only exploit the most salient area among different people, but also extract important local detail through a Local Attention System (LAS). LAS is an attention selection unit that could extract approximate semantic local features of human body parts without extra supervision. To learn discriminative attention feature representation, we explore an attention feature regularization scheme to enhance the relevance of body part features that belong to same personal identity. Considering the effectiveness of feature augmentation in the Re-ID task and the defect of the existing methods, we propose a Batch Attention DropBlock (BA-DropBlock) to further improve DropBlock by combining the attention selection mechanism. Results on mainstream datasets demonstrate the superiority of our model over the state-of-the-art. Especially, our approach exceeds the current best method by a large margin of 4.6${\%}$on the most challenging dataset CUHK03. Xun Gong 0002, Zu Yao, Xin Li 0003, Yueqiao Fan, Jianfeng Fan, Boji Lao |
IEEE Trans. Multim. | 3 |
| 2022 | Cross-Modal Semantic Matching Generative Adversarial Networks for Text-to-Image SynthesisabstractSynthesizing photo-realistic images based on text descriptions is a challenging image generation problem. Although many recent approaches have significantly advanced the performance of text-to-image generation, to guarantee semantic matchings between the text description and synthesized image remains very challenging. In this paper, we propose a new model, Cross-modal Semantic Matching Generative Adversarial Networks (CSM-GAN), to improve the semantic consistency between text description and synthesized image for a fine-grained text-to-image generation. Two new modules are proposed in CSM-GAN: Text Encoder Module (TEM) and Textual-Visual Semantic Matching Module (TVSMM). TVSMM is aimed at making the distance of the pairs of synthesized image and its corresponding text description closer, in global semantic embedding space, than those of mismatched pairs. This improves the semantic consistency and consequently, the generalizability of CSM-GAN. In TEM, we introduce Text Convolutional Neural Networks (Text_CNNs) to capture and highlight local visual features in textual descriptions. Thorough experiments on two public benchmark datasets demonstrated the superiority of CSM-GAN over other representative state-of-the-art methods. Hongchen Tan, Xiuping Liu, Xin Li 0003 |
IEEE Trans. Multim. | 4 |
| 2021 | Tracklet Proposal Network for Multi-Object Tracking on Point CloudsabstractThis paper proposes the first tracklet proposal network, named PC-TCNN, for Multi-Object Tracking (MOT) on point clouds. Our pipeline first generates tracklet proposals, then refines these tracklets and associates them to generate long trajectories. Specifically, object proposal generation and motion regression are first performed on a point cloud sequence to generate tracklet candidates. Then, spatial-temporal features of each tracklet are exploited and their consistency is used to refine the tracklet proposal. Finally, the refined tracklets across multiple frames are associated to perform MOT on the point cloud sequence. The PC-TCNN significantly improves the MOT performance by introducing the tracklet proposal design. On the KITTI tracking benchmark, it attains an MOTA of 91.75%, outperforming all submitted results on the online leaderboard. Qing Li 0032, Chenglu Wen, Xin Li 0003, Xiaoliang Fan, Cheng Wang 0003 |
IJCAI | 4 |
| 2021 | A Hardware-adaptive Deep Feature Matching Pipeline for Real-time 3D Reconstruction
Yabin Wang 0001, Baotong Li, Xin Li 0003 |
Comput. Aided Des. | 4 |
| 2021 | Detection of pancreatic cancer by convolutional-neural-network-assisted spontaneous Raman spectroscopy with critical feature visualization
Zhongqiang Li, Zheng Li 0036, Alexandra Ramos, J. Philip Boudreaux, Ramcharan Thiagarajan, Yvette Bren-Mattison, Michael E. Dunham, Andrew J. McWhorter, Xin Li 0003, Ji-Ming Feng, Shaomian Yao, Jian Xu 0022 |
Neural Networks | 11 |
| 2021 | Spatial context-aware network for salient object detection
Yuqiu Kong, Mengyang Feng, Xin Li 0003, Huchuan Lu, Xiuping Liu |
Pattern Recognit. | 3 |
| 2021 | FeatFlow: Learning geometric features for 3D motion estimation
Qing Li 0032, Cheng Wang 0003, Xin Li 0003, Chenglu Wen |
Pattern Recognit. | 3 |
| 2021 | ESKN: Enhanced selective kernel network for single image super-resolution
Zewei He, Guizhong Fu, Yanpeng Cao, Yanlong Cao, Jiangxin Yang, Xin Li 0003 |
Signal Process. | 6 |
| 2021 | KT-GAN: Knowledge-Transfer Generative Adversarial Network for Text-to-Image SynthesisabstractThis paper presents a new framework, Knowledge-Transfer Generative Adversarial Network (KT-GAN), for fine-grained text-to-image generation. We introduce two novel mechanisms: an Alternate Attention-Transfer Mechanism (AATM) and a Semantic Distillation Mechanism (SDM), to help generator better bridge the cross-domain gap between text and image. The AATM updates word attention weights and attention weights of image sub-regions alternately, to progressively highlight important word information and enrich details of synthesized images. The SDM uses the image encoder trained in the Image-to-Image task to guide training of the text encoder in the Text-to-Image task, for generating better text features and higher-quality images. With extensive experimental validation on two public datasets, our KT-GAN outperforms the baseline method significantly, and also achieves the competive results over different evaluation metrics. Hongchen Tan, Xiuping Liu, Meng Liu 0006, Xin Li 0003 |
IEEE Trans. Image Process. | 5 |
| 2020 | Point2Node: Correlation Learning of Dynamic-Node for Point Cloud Feature ModelingabstractFully exploring correlation among points in point clouds is essential for their feature modeling. This paper presents a novel end-to-end graph model, named Point2Node, to represent a given point cloud. Point2Node can dynamically explore correlation among all graph nodes from different levels, and adaptively aggregate the learned features. Specifically, first, to fully explore the spatial correlation among points for enhanced feature description, in a high-dimensional node graph, we dynamically integrate the node's correlation with self, local, and non-local nodes. Second, to more effectively integrate learned features, we design a data-aware gate mechanism to self-adaptively aggregate features at the channel level. Extensive experiments on various point cloud benchmarks demonstrate that our method outperforms the state-of-the-art. Wenkai Han, Chenglu Wen, Cheng Wang 0003, Xin Li 0003, Qing Li 0032 |
AAAI | 4 |
| 2020 | DDGCN: A Dynamic Directed Graph Convolutional Network for Action Recognition
Matthew Korban, Xin Li 0003 |
ECCV (20) | 2 |
| 2020 | Video Object Segmentation with Adaptive Feature Bank and Uncertain-Region RefinementabstractThis paper presents a new matching-based framework for semi-supervised video object segmentation (VOS). Recently, state-of-the-art VOS performance has been achieved by matching-based algorithms, in which feature banks are created to store features for region matching and classification. However, how to effectively organize information in the continuously growing feature bank remains under-explored, and this leads to an inefficient design of the bank. We introduced an adaptive feature bank update scheme to dynamically absorb new features and discard obsolete features. We also designed a new confidence loss and a fine-grained segmentation module to enhance the segmentation accuracy in uncertain regions. On public benchmarks, our algorithm outperforms existing state-of-the-arts. Yongqing Liang 0001, Xin Li 0003, Navid H. Jafari, Jim Chen |
NeurIPS | 2 |
| 2020 | SPM 2020 Editorial
Xin Li 0003, Michael Barton 0002, Saigopal Nelaturi |
Comput. Aided Des. | 1 |
| 2020 | A Cross-Dimension Annotations Method for 3D Structural Facial Landmark ExtractionabstractAbstract Recent methods for 2D facial landmark localization perform well on close‐to‐frontal faces, but 2D landmarks are insufficient to represent 3D structure of a facial shape. For applications that require better accuracy, such as facial motion capture and 3D shape recovery, 3DA‐2D (2D Projections of 3D Facial Annotations) is preferred. Inferring the 3D structure from a single image is an ill‐posed problem whose accuracy and robustness are not always guaranteed. This paper aims to solve accurate 2D facial landmark localization and the transformation between 2D and 3DA‐2D landmarks. One way to increase the accuracy is to input more precisely annotated facial images. The traditional cascaded regressions cannot effectively handle large or noisy training data sets. In this paper, we propose a Mini‐Batch Cascaded Regressions (MBCR) method that can iteratively train a robust model from a large data set. Benefiting from the incremental learning strategy and a small learning rate, MBCR is robust to noise in training data. We also propose a new Cross‐Dimension Annotations Conversion (CDAC) method to map facial landmarks from 2D to 3DA‐2D coordinates and vice versa. The experimental results showed that CDAC combined with MBCR outperforms the‐state‐of‐the‐art methods in 3DA‐2D facial landmark localization. Moreover, CDAC can run efficiently at up to 110 fps on a 3.4 GHz‐CPU workstation. Thus, CDAC provides a solution to transform existing 2D alignment methods into 3DA‐2D ones without slowing down the speed. Training and testing code as well as the data set can be downloaded from https://github.com/SWJTU‐3DVision/CDAC. Xun Gong 0002, Zhemin Zhang, Yue Xiang, Xin Li 0003 |
Comput. Graph. Forum | 6 |
| 2020 | A Discriminative Multi-Channel Facial Shape (MCFS) Representation and Feature Extraction for 3D Human FacesabstractAbstract Building an effective representation for 3D face geometry is essential for face analysis tasks, that is, landmark detection, face recognition and reconstruction. This paper proposes to use a Multi‐Channel Facial Shape (MCFS) representation that consists of depth, hand‐engineered feature and attention maps to construct a 3D facial descriptor. And, a multi‐channel adjustment mechanism, named filtered squeeze and reversed excitation (FSRE), is proposed to re‐organize MCFS data. To assign a suitable weight for each channel, FSRE is able to learn the importance of each layer automatically in the training phase. MCFS and FSRE blocks collaborate together effectively to build a robust 3D facial shape representation, which has an excellent discriminative ability. Extensive experimental results, testing on both high‐resolution and low‐resolution face datasets, show that facial features extracted by our framework outperform existing methods. This representation is stable against occlusions, data corruptions, expressions and pose variations. Also, unlike traditional methods of 3D face feature extraction, which always take minutes to create 3D features, our system can run in real time. Xun Gong 0002, Xin Li 0003, Tianrui Li 0001, Yongqing Liang 0001 |
Comput. Graph. Forum | 2 |
| 2020 | WaterNet: An adaptive matching pipeline for segmenting water with volatile appearanceabstractWe develop a novel network to segment water with significant appearance variation in videos. Unlike existing state-of-the-art video segmentation approaches that use a pre-trained feature recognition network and several previous frames to guide segmentation, we accommodate the object’s appearance variation by considering features observed from the current frame. When dealing with segmentation of objects such as water, whose appearance is non-uniform and changing dynamically, our pipeline can produce more reliable and accurate segmentation results than existing algorithms. Yongqing Liang 0001, Navid H. Jafari, Qin Chen 0003, Yanpeng Cao, Xin Li 0003 |
Comput. Vis. Media | 6 |
| 2020 | Reassembling Shredded Document Stripes Using Word-Path Metric and Greedy Composition Optimal Matching SolverabstractThis paper develops a shredded document reassembly algorithm based on character/word detection. A new word compatibility estimation metric and a searching strategy called Greedy Composition and Optimal Matching (GCOM) are proposed to compose documents from their vertically shredded stripes. We reduce the stripe puzzle reassembly problem to the traveling salesman problem (TSP) on a sparse graph. The word-path compatibility metric takes advantages of the optical character recognition (OCR) to compute the compatibility score among a group of stripes. The global composition strategy, based on an integration of greedy composition and optimal matching, is proposed to search for a maximal Hamiltonian path and the final global reassembly. We demonstrate that our solver outperforms the state-of-the-art puzzle solvers on reassembling stripe shredded documents. Yongqing Liang 0001, Xin Li 0003 |
IEEE Trans. Multim. | 2 |
| 2019 | LO-Net: Deep Real-Time Lidar OdometryabstractWe present a novel deep convolutional network pipeline, LO-Net, for real-time lidar odometry estimation. Unlike most existing lidar odometry (LO) estimations that go through individually designed feature selection, feature matching, and pose estimation pipeline, LO-Net can be trained in an end-to-end manner. With a new mask-weighted geometric constraint loss, LO-Net can effectively learn feature representation for LO estimation, and can implicitly exploit the sequential dependencies and dynamics in the data. We also design a scan-to-map module, which uses the geometric and semantic information learned in LO-Net, to improve the estimation accuracy. Experiments on benchmark datasets demonstrate that LO-Net outperforms existing learning based approaches and has similar accuracy with the state-of-the-art geometry-based approach, LOAM. Qing Li 0032, Shaoyang Chen, Cheng Wang 0003, Xin Li 0003, Chenglu Wen, Ming Cheng 0002, Jonathan Li 0001 |
CVPR | 4 |
| 2019 | RF-Net: An End-To-End Image Matching Network Based on Receptive FieldabstractThis paper proposes a new end-to-end trainable matching network based on receptive field, RF-Net, to compute sparse correspondence between images. Building end-to-end trainable matching framework is desirable and challenging. The very recent approach, LF-Net, successfully embeds the entire feature extraction pipeline into a jointly trainable pipeline, and produces the state-of-the-art matching results. This paper introduces two modifications to the structure of LF-Net. First, we propose to construct receptive feature maps, which lead to more effective keypoint detection. Second, we introduce a general loss function term, neighbor mask, to facilitate training patch selection. This results in improved stability in descriptor training. We trained RF-Net on the open dataset HPatches, and compared it with other methods on multiple benchmark datasets. Experiments show that RF-Net outperforms existing state-of-the-art methods. Xuelun Shen, Cheng Wang 0003, Xin Li 0003, Zenglei Yu, Jonathan Li 0001, Chenglu Wen, Ming Cheng 0002 |
CVPR | 3 |
| 2019 | Semantics-Enhanced Adversarial Nets for Text-to-Image SynthesisabstractThis paper presents a new model, Semantics-enhanced Generative Adversarial Network (SEGAN), for fine-grained text-to-image generation. We introduce two modules, a Semantic Consistency Module (SCM) and an Attention Competition Module (ACM), to our SEGAN. The SCM incorporates image-level semantic consistency into the training of the Generative Adversarial Network (GAN), and can diversify the generated images and improve their structural coherence. A Siamese network and two types of semantic similarities are designed to map the synthesized image and the groundtruth image to nearby points in the latent semantic feature space. The ACM constructs adaptive attention weights to differentiate keywords from unimportant words, and improves the stability and accuracy of SEGAN. Extensive experiments demonstrate that our SEGAN significantly outperforms existing state-of-the-art methods in generating photo-realistic images. All source codes and models will be released for comparative study. Hongchen Tan, Xiuping Liu, Xin Li 0003, Yi Zhang 0100 |
ICCV | 3 |
| 2019 | Non-iterative structural topology optimization using deep learning
Baotong Li, Congjia Huang, Xin Li 0003, Jun Hong 0001 |
Comput. Aided Des. | 3 |
| 2019 | Hierarchical fragmented image reassembly using a bundle-of-superpixel representation
Xin Li 0003, Kang Xie, Wenxing Hong, Celong Liu |
Comput. Aided Geom. Des. | 1 |
| 2019 | Automatic craniofacial registration based on radial curves
Ruikun Huang, Junli Zhao, Fuqing Duan, Xin Li 0003, Celong Liu, Xiaodan Deng, Zhenkuan Pan 0001, Zhongke Wu |
Comput. Graph. | 4 |
| 2019 | Social emotion classification based on noise-aware training
Xin Li 0003, Yanghui Rao, Haoran Xie 0001, Xuebo Liu 0004, Tak-Lam Wong, Fu Lee Wang |
Data Knowl. Eng. | 1 |
| 2019 | Real-Time Avatar Pose Transfer and Motion Generation Using Locally Encoded Laplacian Offsets
Masoud Zadghorban Lifkooee, Celong Liu, Yongqing Liang 0001, Yimin Zhu 0004, Xin Li 0003 |
J. Comput. Sci. Technol. | 5 |
| 2019 | Robust procedural model fitting with a new geometric similarity estimator
Zongliang Zhang, Jonathan Li 0001, Yulan Guo, Xin Li 0003, Yangbin Lin, Guobao Xiao, Cheng Wang 0003 |
Pattern Recognit. | 4 |
| 2019 | Fast and accurate single image super-resolution via an energy-aware improved deep residual network
Yanpeng Cao, Zewei He, Zhangyu Ye, Xin Li 0003, Yanlong Cao, Jiangxin Yang |
Signal Process. | 4 |
| 2019 | JigsawNet: Shredded Image Reassembly Using Convolutional Neural Network and Loop-Based CompositionabstractThis paper proposes a novel algorithm to reassemble an arbitrarily shredded image to its original status. Existing reassembly pipelines commonly consist of a local matching stage and a global compositions stage. In the local stage, a key challenge is to reliably compute correct pairwise matching, for which most existing algorithms use handcrafted features, and cannot reliably handle complicated puzzles. We build a deep convolutional neural network (CNN) to detect the compatibility of pairwise stitching, and use it to prune computed pairwise matches. To improve the network efficiency and accuracy, we transfer the calculation of CNN to the stitching region and apply a boost training strategy. In the global composition stage, instead of using the widely adopted greedy edge selection strategies, we propose two new loop closure-based searching algorithms. Extensive experiments show that our algorithm significantly outperforms existing methods on solving various puzzles, especially challenging ones with many fragment pieces. Canyu Le, Xin Li 0003 |
IEEE Trans. Image Process. | 2 |
| 2018 | Sparse3D: A new global model for matching sparse RGB-D dataset with small inter-frame overlap
Canyu Le, Xin Li 0003 |
Comput. Aided Des. | 2 |
| 2017 | Geometry-aware partitioning of complex domains for parallel quad meshing
Xin Li 0003, Wuyi Yu, Celong Liu |
Comput. Aided Des. | 1 |
| 2017 | Distributed poly-square mapping for large-scale semi-structured quad mesh generation
Celong Liu, Wuyi Yu, Zhonggui Chen, Xin Li 0003 |
Comput. Aided Des. | 4 |
| 2016 | Social Emotion Classification via Reader Perspective Weighted ModelabstractWith the development of Web 2.0, many users express their opinions online. This paper is concerned with the classification of social emotions on varied-scale datasets. Different from traditional models which weight training documents equally, the concept of emotional entropy is proposed to estimate the weight and tackle the issue of noisy documents. The topic assignment is also used to distinguish different emotional senses of the same word. Experimental evaluations using different data sets validate the effectiveness of the proposed social emotion classification model. Xin Li 0003, Yanghui Rao, Yanjia Chen, Xuebo Liu 0004 |
AAAI | 1 |
| 2016 | A multi-frame graph matching algorithm for low-bandwidth RGB-D SLAM
Jun Hong 0001, Kang Zhang 0002, Baotong Li, Xin Li 0003 |
Comput. Aided Des. | 5 |
| 2016 | B-spline surface fitting with knot position optimization
Yuhua Zhang, Juan Cao 0002, Zhonggui Chen, Xin Li 0003, Xiaoming Zeng |
Comput. Graph. | 4 |
| 2016 | Segmenting a surface mesh into pants using Morse theory
Mustafa Hajij, Tamal K. Dey, Xin Li 0003 |
Graph. Model. | 3 |
| 2015 | 3D Fragment Reassembly Using Integrated Template Guidance and Fracture-Region MatchingabstractThis paper studies matching of fragmented objects to recompose their original geometry. Solving this geometric reassembly problem has direct applications in archaeology and forensic investigation in the computer-aided restoration of damaged artifacts and evidence. We develop a new algorithm to effectively integrate both guidance from a template and from matching of adjacent pieces' fracture-regions. First, we compute partial matchings between fragments and a template, and pairwise matchings among fragments. Many potential matches are obtained and then selected/refined in a multi-piece matching stage to maximize global groupwise matching consistency. This pipeline is effective in composing fragmented thin-shell objects containing small pieces, whose pairwise matching is usually unreliable and ambiguous and hence their reassembly remains challenging to the existing algorithms. Kang Zhang 0002, Wuyi Yu, Mary Manhein, Warren N. Waggenspack, Xin Li 0003 |
ICCV | 5 |
| 2015 | Revised spectral matching algorithm for scenes with mutually inconsistent local transformationsabstractSpectral matching (SM) is an efficient and effective greedy algorithm for solving the graph matching problem in feature correspondence in computer vision and graphics. However, the classic SM algorithm cannot extract correspondences well when the affinity matrix is sparse and reducible (i.e. its corresponding graph is not connected). This case often happens when the geometric deformations consist of transformations with local inconsistency. The authors analyse this problem and show how the original SM could fail in this scenario. Then, the authors propose a revised two‐step pipeline to tackle this issue: (1) decompose the mutually inconsistent local deformations into several consistent transformations which can be solved by individual SM; (2) filter out incorrect correspondences through an automatic thresholding. The authors perform experiments to demonstrate that this modification can effectively handle the coarse correspondence computation in shape or image registration where the global transformation consists of multiple inconsistent local transformations. Peizhi Chen, Xin Li 0003 |
IET Image Process. | 2 |
| 2014 | Optimizing polycube domain construction for hexahedral remeshing
Wuyi Yu, Kang Zhang 0002, Shenghua Wan, Xin Li 0003 |
Comput. Aided Des. | 4 |
| 2014 | A graph-based optimization algorithm for fragmented image reassembly
Kang Zhang 0002, Xin Li 0003 |
Graph. Model. | 2 |
| 2013 | A Symmetric 4D Registration Algorithm for Respiratory Motion Modeling
Huanhuan Xu, Xin Li 0003 |
MICCAI (2) | 2 |
| 2013 | An efficient spherical mapping algorithm and its application on spherical harmonics
Shenghua Wan, Tengfei Ye, Maoqing Li, Xin Li 0003 |
Sci. China Inf. Sci. | 5 |
| 2013 | Surface Mesh to Volumetric Spline Conversion with Generalized PolycubesabstractThis paper develops a novel volumetric parameterization and spline construction framework, which is an effective modeling tool for converting surface meshes to volumetric splines. Our new splines are defined upon a novel parametric domain called generalized polycubes (GPCs). A GPC comprises a set of regular cube domains topologically glued together. Compared with conventional polycubes (CPCs), the GPC is much more powerful and flexible and has improved numerical accuracy and computational efficiency when serving as a parametric domain. We design an automatic algorithm to construct the GPC domain while also permitting the user to improve shape abstraction via interactive intervention. We then parameterize the input model on the GPC domain. Finally, we devise a new volumetric spline scheme based on this seamless volumetric parameterization. With a hierarchical fitting scheme, the proposed splines can fit data accurately using reduced number of superfluous control points. Our volumetric modeling scheme has great potential in shape modeling, engineering analysis, and reverse engineering applications. Bo Li 0014, Xin Li 0003, Kexiang Wang, Hong Qin 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2013 | Biharmonic Volumetric Mapping Using Fundamental SolutionsabstractWe propose a biharmonic model for cross-object volumetric mapping. This new computational model aims to facilitate the mapping of solid models with complicated geometry or heterogeneous inner structures. In order to solve cross-shape mapping between such models through divide and conquer, solid models can be decomposed into subparts upon which mappings is computed individually. The biharmonic volumetric mapping can be performed in each subregion separately. Unlike the widely used harmonic mapping which only allows C0 continuity along the segmentation boundary interfaces, this biharmonic model can provide C1 smoothness. We demonstrate the efficacy of our mapping framework on various geometric models with complex geometry (which are decomposed into subparts with simpler and solvable geometry) or heterogeneous interior structures (whose different material layers can be segmented and processed separately). Huanhuan Xu, Wuyi Yu, Shiyuan Gu, Xin Li 0003 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2012 | Efficient Spherical Parametrization Using Progressive Optimization
Shenghua Wan, Tengfei Ye, Maoqing Li, Xin Li 0003 |
CVM | 5 |
| 2012 | Feature-aligned 4D spatiotemporal image registration
Huanhuan Xu, Peizhi Chen, Wuyi Yu, Amit Sawant, S. Sitharama Iyengar, Xin Li 0003 |
ICPR | 6 |
| 2012 | Fragmented skull modeling using heat kernels
Wei Yu 0012, Maoqing Li, Xin Li 0003 |
Graph. Model. | 3 |
| 2012 | On Optimizing Autonomous Pipeline InspectionabstractThis paper studies the optimal inspection of autonomous robots in a complex pipeline system. We solve a 3-D region-guarding problem to suggest the necessary inspection spots. The proposed hierarchical integer linear programming optimization algorithm seeks the fewest spots necessary to cover the entire given 3-D region. Unlike most existing pipeline inspection systems that focus on designing mobility and control of the explore robots, this paper focuses on global planning of the thorough and automatic inspection of a complex environment. We demonstrate the efficacy of the computation framework using a simulated environment, where scanned pipelines and existing leaks, clogs, and deformation can be thoroughly detected by an autonomous prototype robot. Xin Li 0003, Wuyi Yu, S. Sitharama Iyengar |
IEEE Trans. Robotics | 1 |
| 2012 | Spherical DCB-Spline Surfaces with Hierarchical and Adaptive Knot InsertionabstractThis paper develops a novel surface fitting scheme for automatically reconstructing a genus-0 object into a continuous parametric spline surface. A key contribution for making such a fitting method both practical and accurate is our spherical generalization of the Delaunay configuration B-spline (DCB-spline), a new non-tensor-product spline. In this framework, we efficiently compute Delaunay configurations on sphere by the union of two planar Delaunay configurations. Also, we develop a hierarchical and adaptive method that progressively improves the fitting quality by new knot-insertion strategies guided by surface geometry and fitting error. Within our framework, a genus-0 model can be converted to a single spherical spline representation whose root mean square error is tightly bounded within a user-specified tolerance. The reconstructed continuous representation has many attractive properties such as global smoothness and no auxiliary knots. We conduct several experiments to demonstrate the efficacy of our new approach for reverse engineering and shape modeling. Juan Cao 0002, Xin Li 0003, Zhonggui Chen, Hong Qin 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2012 | Restricted Trivariate Polycube Splines for Volumetric Data ModelingabstractThis paper presents a volumetric modeling framework to construct a novel spline scheme called restricted trivariate polycube splines (RTP-splines). The RTP-spline aims to generalize both trivariate T-splines and tensor-product B-splines; it uses solid polycube structure as underlying parametric domains and strictly bounds blending functions within such domains. We construct volumetric RTP-splines in a top-down fashion in four steps: 1) Extending the polycube domain to its bounding volume via space filling; 2) building the B-spline volume over the extended domain with restricted boundaries; 3) inserting duplicate knots by adding anchor points and performing local refinement; and 4) removing exterior cells and anchors. Besides local refinement inherited from general T-splines, the RTP-splines have a few attractive properties as follows: 1) They naturally model solid objects with complicated topologies/bifurcations using a one-piece continuous representation without domain trimming/patching/merging. 2) They have guaranteed semistandardness so that the functions and derivatives evaluation is very efficient. 3) Their restricted support regions of blending functions prevent control points from influencing other nearby domain regions that stay opposite to the immediate boundaries. These features are highly desirable for certain applications such as isogeometric analysis. We conduct extensive experiments on converting complicated solid models into RTP-splines, and demonstrate the proposed spline to be a powerful and promising tool for volumetric modeling and other scientific/engineering applications where data sets with multiattributes are prevalent. Kexiang Wang, Xin Li 0003, Bo Li 0014, Huanhuan Xu, Hong Qin 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2011 | An automatic assembly and completion framework for fragmented skullsabstractWe develop a completion pipeline for fragmented and damaged skulls. The goal of this work is to convert scanned incomplete skull fragments to a complete skull model for subsequent forensic or archeological tasks such as facial reconstruction. The proposed assembly and completion algorithms can also be used to repair other fragmented objects with inherent symmetry. A two-step assembly framework is proposed: (1) rough assembly by an ICP-like template matching algorithm integrated with the slippage features and spin-image descriptors; (2) assembly refinement by a global optimization on least square transformation error (LSTE) of break-curves. The assembled skull is finally repaired by a symmetry-based completion algorithm. Experiments on repairing scanned skull fragments demonstrate the efficacy and robustness of this framework. Zhao Yin, Mary Manhein, Xin Li 0003 |
ICCV | 4 |
| 2011 | Efficient 3D region guarding for multimedia data processingabstractWith the advance of scanning devices, 3-d geometric models have been captured and widely used in animation, video, interactive virtual environment design nowadays. Their effective analysis, integration, and retrieval are important research topics in multimedia. This paper studies a geometric modeling problem called 3D region guarding. The 3D region guarding is a well known NP-hard problem; we present an efficient hierarchical integer linear programming (HILP) optimization algorithm to solve it on massive data sets. We show the effectiveness of our algorithm and briefly illustrate its applications in multimedia data processing and computer graphics such as shape analysis and retrieval, and morphing animation. Wuyi Yu, Maoqing Li, S. Sitharama Iyengar, Xin Li 0003 |
ICME | 4 |
| 2011 | Symmetry and template guided completion of damaged skulls
Xin Li 0003, Zhao Yin, Shenghua Wan, Wei Yu 0012, Maoqing Li |
Comput. Graph. | 1 |
| 2011 | A topology-preserving optimization algorithm for polycube mapping
Shenghua Wan, Zhao Yin, Kang Zhang 0002, Xin Li 0003 |
Comput. Graph. | 5 |
| 2011 | Computing 3D Shape Guarding and Star DecompositionabstractAbstract This paper proposes an effective framework to compute the visibility guarding and star decomposition of 3D solid shapes. We propose a progressive integer linear programming algorithm to solve the guarding points that can visibility cover the entire shape; we also develop a constrained region growing scheme seeded on these guarding points to get the star decomposition. We demonstrate this guarding/decomposition framework can benefit graphics tasks such as shape interpolation and shape matching/retrieval. Wuyi Yu, Xin Li 0003 |
Comput. Graph. Forum | 2 |
| 2010 | Generalized PolyCube Trivariate SplinesabstractThis paper develops a new trivariate hierarchical spline scheme for volumetric data representation. Unlike conventional spline formulations and techniques, our new framework is built upon a novel parametric domain called Generalized PolyCube (GPC), comprising a set of regular cubes being glued together. Compared with the conventional PolyCube (PC) that could serve as a "one-piece'' 3-manifold domain, GPC has more powerful and flexible representation ability. We develop an effective framework that parameterizes a solid model onto a topologically equivalent GPC domain, and design a hierarchical fitting scheme based on trivariate T-splines. The entire data-spline-conversion modeling framework provides high-accuracy data fitting and greatly reduce the number of superfluous control points. It is a powerful toolkit with broader application appeal in shape modeling, engineering analysis, and reverse engineering. Bo Li 0014, Xin Li 0003, Kexiang Wang, Hong Qin 0001 |
Shape Modeling International | 2 |
| 2010 | Feature-aligned harmonic volumetric mapping using MFS
Xin Li 0003, Huanhuan Xu, Shenghua Wan, Zhao Yin, Wuyi Yu |
Comput. Graph. | 1 |
| 2009 | Surface reconstruction using bivariate simplex splines on Delaunay configurations
Juan Cao 0002, Xin Li 0003, Guozhao Wang, Hong Qin 0001 |
Comput. Graph. | 2 |
| 2009 | Geometry-aware domain decomposition for T-spline-based manifold modeling
Hongyu Wang 0002, Ying He 0001, Xin Li 0003, Xianfeng Gu, Hong Qin 0001 |
Comput. Graph. | 3 |
| 2009 | Meshless Harmonic Volumetric Mapping Using Fundamental Solution MethodsabstractHarmonic volumetric mapping aims to establish a smooth bijective correspondence between two solid shapes with the same topology. In this paper, we develop an automatic meshless method for creating such a mapping between two given objects. With the shell surface mapping as the boundary condition, we first solve a linear system constructed by a boundary method called themethodoffundamentalsolution, and then represent the mapping using a set of points with different weights in the vicinity of the shell of the given model. Our algorithm is a true meshless method (without the need of any specific meshing structure within the solid interior) and the behavior of the interior region is directly determined by the boundary, which can improve the computational efficiency and robustness significantly. Therefore, our algorithm can be applied to massive volume data sets with various geometric primitives and topological types. We demonstrate the utility and efficacy of our algorithm in information transfer, shape registration, deformation sequence analysis, tetrahedral remeshing, and solid texture synthesis. Xin Li 0003, Xiaohu Guo, Hongyu Wang 0002, Ying He 0001, Xianfeng Gu, Hong Qin 0001 |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2009 | Surface Mapping Using Consistent Pants DecompositionabstractSurface mapping is fundamental to shape computing and various downstream applications. This paper develops a pants decomposition framework for computing maps between surfaces with arbitrary topologies. The framework first conducts pants decomposition on both surfaces to segment them into consistent sets of pants patches (a pants patch is intuitively defined as a genus-0 surface with three boundaries), then composes global mapping between two surfaces by using harmonic maps of corresponding patches. This framework has several key advantages over existing techniques. First, it is automatic. It can automatically construct mappings for surfaces with complicated topology, guaranteeing the one-to-one continuity. Second, it is general and powerful. It flexibly handles mapping computation between surfaces with different topologies. Third, it is flexible. Despite topology and geometry, it can also integrate semantics requirements from users. Through a simple and intuitive human-computer interaction mechanism, the user can flexibly control the mapping behavior by enforcing point/curve constraints. Compared with traditional user-guided, piecewise surface mapping techniques, our new method is less labor intensive, more intuitive, and requires no user's expertise in computing complicated surface maps between arbitrary shapes. We conduct various experiments to demonstrate its modeling potential and effectiveness. Xin Li 0003, Xianfeng Gu, Hong Qin 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2008 | Surface matching using consistent pants decompositionabstractSurface matching is fundamental to shape computing and various downstream applications. This paper develops a powerful pants decomposition framework for computing maps between surfaces with arbitrary topologies. We first conduct pants decomposition on both surfaces to segment them into consistent sets of pants patches (here a pants patch is intuitively defined as a genus-zero surface with three boundaries). Then we compose global mapping between two surfaces by harmonic maps of corresponding patches. This framework has several key advantages over other state-of-the-art techniques. First, the surface decomposition is automatic and general. It can automatically construct mappings for surfaces with same but complicated topology, and the result is guaranteed to be one-to-one continuous. Second, the mapping framework is very flexible and powerful. Not only topology and geometry, but also the semantics can be easily integrated into this framework with a little user involvement. Specifically, it provides an easy and intuitive human-computer interaction mechanism so that mapping between surfaces with different topologies, or with additional point/curve constraints, can be properly obtained within our framework. Compared with previous user-guided, piecewise surface mapping techniques, our new method is more intuitive, less labor-intensive, and requires no user's expertise in computing complicated surface map between arbitrary shapes. We conduct various experiments to demonstrate its modeling potential and effectiveness. © 2008 ACM. Xin Li 0003, Xianfeng Gu, Hong Qin 0001 |
Symposium on Solid and Physical Modeling | 1 |
| 2008 | Polycube splines
Hongyu Wang 0002, Ying He 0001, Xin Li 0003, Xianfeng Gu, Hong Qin 0001 |
Comput. Aided Des. | 3 |
| 2008 | Globally Optimal Surface Mapping for Surfaces with Arbitrary TopologyabstractComputing smooth and optimal one-to-one maps between surfaces of same topology is a fundamental problem in computer graphics and such a method provides us a ubiquitous tool for geometric modeling and data visualization. Its vast variety of applications includes shape registration/matching, shape blending, material/data transfer, data fusion, information reuse, etc. The mapping quality is typically measured in terms of angular distortions among different shapes. This paper proposes and develops a novel quasi-conformal surface mapping framework to globally minimize the stretching energy inevitably introduced between two different shapes. The existing state-of-the-art inter-surface mapping techniques only afford local optimization either on surface patches via boundary cutting or on the simplified base domain, lacking rigorous mathematical foundation and analysis. We design and articulate an automatic variational algorithm that can reach the global distortion minimum for surface mapping between shapes of arbitrary topology, and our algorithm is sorely founded upon the intrinsic geometry structure of surfaces. To our best knowledge, this is the first attempt towards numerically computing globally optimal maps. Consequently, our mapping framework offers a powerful computational tool for graphics and visualization tasks such as data and texture transfer, shape morphing, and shape matching. Xin Li 0003, Yunfan Bao, Xiaohu Guo, Miao Jin, Xianfeng Gu, Hong Qin 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2007 | Harmonic volumetric mapping for solid modeling applicationsabstractHarmonic volumetric mapping for two solid objects establishes a one-to-one smooth correspondence between them. It finds its applications in shape registration and analysis, shape retrieval, information reuse, and material/texture transplant. In sharp contrast to harmonic surface mapping techniques, little research has been conducted for designing volumetric mapping algorithms due to its technical challenges. In this paper, we develop an automatic and effective algorithm for computing harmonic volumetric mapping between two models of the same topology. Given a boundary mapping between two models, the volumetric (interior) mapping is derived by solving a linear system constructed from a boundary method called the fundamental solution method. The mapping is represented as a set of points with different weights in the vicinity of the solid boundary. In a nutshell, our algorithm is a true meshless method (with no need of specific connectivity) and the behavior of the interior region is directly determined by the boundary. These two properties help improve the computational efficiency and robustness. Therefore, our algorithm can be applied to massive volume data sets with various geometric primitives and topological types. We demonstrate the utility and efficacy of our algorithm in shape registration, information reuse, deformation sequence analysis, tetrahedral remeshing and solid texture synthesis. Xin Li 0003, Xiaohu Guo, Hongyu Wang 0002, Ying He 0001, Xianfeng Gu, Hong Qin 0001 |
Symposium on Solid and Physical Modeling | 1 |
| 2007 | Polycube splinesabstractThis paper proposes a new concept of polycube splines and develops novel modeling techniques for using the polycube splines in solid modeling and shape computing. Polycube splines are essentially a novel variant of manifold splines which are built upon the polycube map, serving as its parametric domain. Our rationale for defining spline surfaces over polycubes is that polycubes have rectangular structures everywhere over their domains except a very small number of corner points. The boundary of polycubes can be naturally decomposed into a set of regular structures, which facilitate tensor-product surface definition, GPU-centric geometric computing, and image-based geometric processing. We develop algorithms to construct polycube maps, and show that the introduced polycube map naturally induces the affine structure with a finite number of extraordinary points. Besides its intrinsic rectangular structure, the polycube map may approximate any original scanned data-set with a very low geometric distortion, so our method for building polycube splines is both natural and necessary, as its parametric domain can mimic the geometry of modeled objects in a topologically correct and geometrically meaningful manner. We design a new data structure that facilitates the intuitive and rapid construction of polycube splines in this paper. We demonstrate the polycube splines with applications in surface reconstruction and shape computing. Hongyu Wang 0002, Ying He 0001, Xin Li 0003, Xianfeng Gu, Hong Qin 0001 |
Symposium on Solid and Physical Modeling | 3 |
| 2006 | Curves-on-Surface: A General Shape Comparison FrameworkabstractWe develop a new surface matching framework to handle surface comparisons based on the mathematical analysis of curves on surfaces, and propose a unique signature for any closed curve on a surface. The signature describes not only the shape of the curve, but also the intrinsic relationship between the curve and its embedding surface; and furthermore, the signature metric is stable across surfaces sharing similar Riemannian geometry metrics. Based on this theoretical advance, we analyze and align features defined as closed curves on surfaces using their signatures. These curves segment a surface into different regions which are mapped onto canonical domains for the matching purpose. The experimental results are very promising, demonstrating that the curve signatures and the comparison framework are robust and discriminative for the effective shape comparison. Besides its utility in our current framework, we believe the curve signature will also serve as a powerful shape segmentation/mapping tool and can be used to aid in many existing techniques towards effective shape analysis Xin Li 0003, Ying He 0001, Xianfeng Gu, Hong Qin 0001 |
SMI | 1 |
| 2006 | Meshless Thin-Shell Simulation Based on Global Conformal ParameterizationabstractThis paper presents a new approach to the physically-based thin-shell simulation of point-sampled geometry via explicit, global conformal point-surface parameterization and meshless dynamics. The point-based global parameterization is founded upon the rigorous mathematics of Riemann surface theory and Hodge theory. The parameterization is globally conformal everywhere except for a minimum number of zero points. Within our parameterization framework, any well-sampled point surface is functionally equivalent to a manifold, enabling popular and powerful surface-based modeling and physically-based simulation tools to be readily adapted for point geometry processing and animation. In addition, we propose a meshless surface computational paradigm in which the partial differential equations (for dynamic physical simulation) can be applied and solved directly over point samples via Moving Least Squares (MLS) shape functions defined on the global parametric domain without explicit connectivity information. The global conformal parameterization provides a common domain to facilitate accurate meshless simulation and efficient discontinuity modeling for complex branching cracks. Through our experiments on thin-shell elastic deformation and fracture simulation, we demonstrate that our integrative method is very natural, and that it has great potential to further broaden the application scope of point-sampled geometry in graphics and relevant fields. Xiaohu Guo, Xin Li 0003, Yunfan Bao, Xianfeng Gu, Hong Qin 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2006 | Spatial Analysis of News SourcesabstractPeople in different places talk about different things. This interest distribution is reflected by the newspaper articles circulated in a particular area. We use data from our large-scale newspaper analysis system (Lydia) to make entity datamaps, a spatial visualization of the interest in a given named entity. Our goal is to identify entities which display regional biases. We develop a model of estimating the frequency of reference of an entity in any given city from the reference frequency centered in surrounding cities, and techniques for evaluating the spatial significance of this distribution. Andrew Mehler, Yunfan Bao, Xin Li 0003, Steven Skiena |
IEEE Trans. Vis. Comput. Graph. | 3 |