EDBT 2026 Demo / reviewers in the wild / expert
Jinxian Liu
dblp:230/1167
· DBLP profile ↗
14ranked-venue papers
6as first author
8since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 5 first-author · 7 since 2021Artificial intelligence and machine learning · 10 · 5 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Mesh2Animation: Unsupervised Animating for Quadruped 3D ObjectsabstractAnimating quadruped 3D objects, such as chairs and tables, typically involves three steps in the traditional computer graphics pipeline: Rigging, Skinning, and Retargeting. Commonly, prevailing methods for each specific step are conceived in isolation. For rigging and skinning steps, optimization-based methods are typically used, but these approaches tend to be slow and susceptible to variations in 3D mesh surfaces. For the retargeting step, the obtained results often fall short of expectations, especially when dealing with dissimilar source and target skeletons, leading to issues like joint twisting. The devised procedure is also time-intensive, resulting in a complex final pipeline. To this end, we present a unified framework, termed Mesh2Animation, providing an end-to-end solution to these challenges. In Mesh2Animation, a learning-based method is proposed for quadruped 3D skeleton estimation. We introduce both skeleton-level and mesh-level loss, allowing the rigging, skinning, and retargeting steps to be optimized simultaneously. Specifically, a general predicted estimation from the rigging step initializes the skeleton, making the skinning step faster and more accurate, which in turn leads to better results in the retargeting step. Finally, the rigging, skinning and retargeting processes are optimized simultaneously under static and temporal constraints. Additionally, we can construct a novel animating dataset termed ShapeNet2Animation (SN2Animation) based on the proposed method, which shows potential application for pose transfer. Qualitative and quantitative results on SN2Animation, ShapeNet, Object3D and ModelNet10 datasets for animation demonstrate that our method achieves competitive performance and shows promising generalization ability on quadruped 3D objects. Our project is available athttps://sites.google.com/view/mesh2animation. Zhenbo Yu, Jinxian Liu, Zefan Li, Bingbing Ni, Wenjun Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | AudioEar: Single-View Ear Reconstruction for Personalized Spatial AudioabstractSpatial audio, which focuses on immersive 3D sound rendering, is widely applied in the acoustic industry. One of the key problems of current spatial audio rendering methods is the lack of personalization based on different anatomies of individuals, which is essential to produce accurate sound source positions. In this work, we address this problem from an interdisciplinary perspective. The rendering of spatial audio is strongly correlated with the 3D shape of human bodies, particularly ears. To this end, we propose to achieve personalized spatial audio by reconstructing 3D human ears with single-view images. First, to benchmark the ear reconstruction task, we introduce AudioEar3D, a high-quality 3D ear dataset consisting of 112 point cloud ear scans with RGB images. To self-supervisedly train a reconstruction model, we further collect a 2D ear dataset composed of 2,000 images, each one with manual annotation of occlusion and 55 landmarks, named AudioEar2D. To our knowledge, both datasets have the largest scale and best quality of their kinds for public use. Further, we propose AudioEarM, a reconstruction method guided by a depth estimation network that is trained on synthetic data, with two loss functions tailored for ear data. Lastly, to fill the gap between the vision and acoustics community, we develop a pipeline to integrate the reconstructed ear mesh with an off-the-shelf 3D human body and simulate a personalized Head-Related Transfer Function (HRTF), which is the core of spatial audio rendering. Code and data are publicly available in https://github.com/seanywang0408/AudioEar. Bingbing Ni, Wenjun Zhang 0001, Jinxian Liu, Teng Li 0001 |
AAAI | 6 |
| 2023 | Fast Fluid Simulation via Dynamic Multi-Scale GriddingabstractRecent works on learning-based frameworks for Lagrangian (i.e., particle-based) fluid simulation, though bypassing iterative pressure projection via efficient convolution operators, are still time-consuming due to excessive amount of particles. To address this challenge, we propose a dynamic multi-scale gridding method to reduce the magnitude of elements that have to be processed, by observing repeated particle motion patterns within certain consistent regions. Specifically, we hierarchically generate multi-scale micelles in Euclidean space by grouping particles that share similar motion patterns/characteristics based on super-light motion and scale estimation modules. With little internal motion variation, each micelle is modeled as a single rigid body with convolution only applied to a single representative particle. In addition, a distance-based interpolation is conducted to propagate relative motion message among micelles. With our efficient design, the network produces high visual fidelity fluid simulations with the inference time to be only 4.24 ms/frame (with 6K fluid particles), hence enables real-time human-computer interaction and animation. Experimental results on multiple datasets show that our work achieves great simulation acceleration with negligible prediction error increase. Jinxian Liu, Ye Chen 0006, Bingbing Ni, Zhenbo Yu |
AAAI | 1 |
| 2023 | Learning by Restoring Broken 3D GeometryabstractThe key point for an experienced craftsman to repair broken objects effectively is that he must know about them deeply. Similarly, we believe that a model can capture rich geometry information from a shape/scene and generate discriminative representations if it is able to find distorted parts of shapes/scenes and restore them. Inspired by this observation, we propose a novel self-supervised 3D learning paradigm named learning by restoring broken shapes/scenes (collectively called 3D geometry). We first develop a destroy-method cluster, from which we sample methods to break some local parts of an object. Then the destroyed object and the normal object are both sent into a point cloud network to get representations, which are employed to segment points that belong to distorted parts and further reconstruct/restore them to normal. To perform better in these two associated pretext tasks, the model is constrained to capture useful object features, such as rich geometric and contextual information. The object representations learned by this self-supervised paradigm transfer well to different datasets and perform well on downstream classification, segmentation and detection tasks. Experimental results on shape datasets and scene datasets demonstrate that our method achieves state-of-the-art performance among unsupervised methods. We also show experimentally that pre-training with our framework significantly boosts the performance of supervised models. Jinxian Liu, Bingbing Ni, Ye Chen 0006, Zhenbo Yu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Joint Global and Dynamic Pseudo Labeling for Semi-Supervised Point Cloud Sequence SegmentationabstractSupervised learning is a mainstay for large discriminative models in 3D computer vision, while large amounts of human-annotated data are the key to achieve state-of-the-art performance. This limitation is particularly notable for large-scale point cloud sequence segmentation tasks, because point-level annotations are very time-consuming and especially expensive. To overcome this challenge, we develop a novel semi-supervised framework for point cloud sequences segmentation. Specifically, we develop two kinds of pseudo labeling methods with extracting global semantic information from labeled frames and dynamic information from each sequence respectively. Then the two kinds of generated labels are combined as more robust pseudo labels (GD-Pseudo labels) for unlabeled frames. We finally apply an efficient iterative learning scheme to train a model with a small quantity of human-annotated data and large-scale pseudo-labeled data. Equipped with our framework, the model achieves significant performance improvement (+12—25 mIoU) on SemanticKITTI and Synthia when compared with frameworks that do not utilize large amounts of unlabeled data. Moreover, our method achieves comparable performance with only 20% annotated frames on SemanticKITTI to state-of-the-art models trained with 100% human-annotated frames. Jinxian Liu, Ye Chen 0006, Bingbing Ni, Zhenbo Yu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | OCR-Pose: Occlusion-aware Contrastive Representation for Unsupervised 3D Human Pose EstimationabstractOcclusion is a significant problem in 3D human pose estimation from the 2D counterpart. On one hand, without explicit annotation, the 3D skeleton is hard to be accurately estimated from the occluded 2D pose. On the other hand, one occluded 2D pose might correspond to multiple 3D skeletons with low confidence parts. To address these issues, we decouple the 3D representation feature into view-invariant part termed occlusion-aware feature and view-dependent part termed rotation feature to facilitate subsequent optimization of the former. Then we propose an occlusion-aware contrastive representation based scheme (OCR-Pose) consisting of Topology Invariant Contrastive Learning module (TiCLR) and View Equivariant Contrastive Learning module (VeCLR). Specifically, TiCLR drives invariance to topology transformation, i.e., bridging the gap between an occluded 2D pose and the unoccluded one. While VeCLR encourages equivariance to view transformation, i.e., capturing the geometric similarity of the 3D skeleton in two views. Both modules optimize occlusion-aware constrastive representation with pose filling and lifting networks via an iterative training strategy in an end-to-end manner. OCR-Pose not only achieves superior performance against state-of-the-art unsupervised methods on unoccluded benchmarks, but also obtains significant improvements when occlusion is involved. Our project is available at https://sites.google.com/view/ocr-pose. Zhenbo Yu, Zhengyan Tong, Jinxian Liu, Wenjun Zhang 0001 |
ACM Multimedia | 5 |
| 2021 | Shape Self-Correction for Unsupervised Point Cloud UnderstandingabstractWe develop a novel self-supervised learning method named Shape Self-Correction for point cloud analysis. Our method is motivated by the principle that a good shape representation should be able to find distorted parts of a shape and correct them. To learn strong shape representations in an unsupervised manner, we first design a shape-disorganizing module to destroy certain local shape parts of an object. Then the destroyed shape and the normal shape are sent into a point cloud network to get representations, which are employed to segment points that belong to distorted parts and further reconstruct them to restore the shape to normal. To perform better in these two associated pretext tasks, the network is constrained to capture useful shape features from the object, which indicates that the point cloud network encodes rich geometric and contextual information. The learned feature extractor transfers well to downstream classification and segmentation tasks. Experimental results on ModelNet, ScanNet and ShapeNetPart demonstrate that our method achieves state-of-the-art performance among unsupervised methods. Our framework can be applied to a wide range of deep learning networks for point cloud analysis and we show experimentally that pre-training with our framework significantly boosts the performance of supervised models. Ye Chen 0006, Jinxian Liu, Bingbing Ni, Jiancheng Yang, Teng Li 0001, Qi Tian 0001 |
ICCV | 2 |
| 2021 | Geometric Granularity Aware Pixel-to-MeshabstractPixel-to-mesh has wide applications, especially in virtual or augmented reality, animation and game industry. However, existing mesh reconstruction models perform unsatisfactorily in local geometry details due to ignoring mesh topology information during learning. Besides, most methods are constrained by the initial template, which cannot reconstruct meshes of various genus. In this work, we propose a geometric granularity-aware pixel-to-mesh framework with a fidelity-selection-and-guarantee strategy, which explicitly addresses both challenges. First, a geometry structure extractor is proposed for detecting local high structured parts and capturing local spatial feature. Second, we apply it to facilitate pixel-to-mesh mapping and resolve coarse details problem caused by the neglect of structural information in previous practices. Finally, a mesh edit module is proposed to encourage non-zero genus topology to emergence by fine-grained topology modification and a patching algorithm is introduced to repair the non-closed boundaries. Extensive experimental results, both quantitatively and visually have demonstrated the high reconstruction fidelity achieved by the proposed framework. Bingbing Ni, Jinxian Liu, Dingyi Rong, Ye Qian, Wenjun Zhang 0001 |
ICCV | 3 |
| 2020 | Two-Stage Relation Constraint for Semantic Segmentation of Point Clouds
Minghui Yu, Jinxian Liu, Bingbing Ni, Caiyuan Li |
3DV | 2 |
| 2020 | Self-Prediction for Joint Instance and Semantic Segmentation of Point Clouds
Jinxian Liu, Minghui Yu, Bingbing Ni, Ye Chen 0006 |
ECCV (22) | 1 |
| 2019 | Modeling Point Clouds With Self-Attention and Gumbel Subset SamplingabstractGeometric deep learning is increasingly important thanks to the popularity of 3D sensors. Inspired by the recent advances in NLP domain, the self-attention transformer is introduced to consume the point clouds. We develop Point Attention Transformers (PATs), using a parameter-efficient Group Shuffle Attention (GSA) to replace the costly Multi-Head Attention. We demonstrate its ability to process size-varying inputs, and prove its permutation equivariance. Besides, prior work uses heuristics dependence on the input data (e.g., Furthest Point Sampling) to hierarchically select subsets of input points. Thereby, we for the first time propose an end-to-end learnable and task-agnostic sampling operation, named Gumbel Subset Sampling (GSS), to select a representative subset of input points. Equipped with Gumbel-Softmax, it produces a "soft" continuous subset in training phase, and a "hard" discrete subset in test phase. By selecting representative subsets in a hierarchical fashion, the networks learn a stronger representation of the input sets with lower computation cost. Experiments on classification and segmentation benchmarks show the effectiveness and efficiency of our methods. Furthermore, we propose a novel application, to process event camera stream as point clouds, and achieve a state-of-the-art performance on DVS128 Gesture Dataset. Jiancheng Yang, Bingbing Ni, Linguo Li, Jinxian Liu, Mengdie Zhou, Qi Tian 0001 |
CVPR | 5 |
| 2019 | Dynamic Points Agglomeration for Hierarchical Point Sets LearningabstractMany previous works on point sets learning achieve excellent performance with hierarchical architecture. Their strategies towards points agglomeration, however, only perform points sampling and grouping in original Euclidean space in a fixed way. These heuristic and task-irrelevant strategies severely limit their ability to adapt to more varied scenarios. To this end, we develop a novel hierarchical point sets learning architecture, with dynamic points agglomeration. By exploiting the relation of points in semantic space, a module based on graph convolution network is designed to learn a soft points cluster agglomeration. We construct a hierarchical architecture that gradually agglomerates points by stacking this learnable and lightweight module. In contrast to fixed points agglomeration strategy, our method can handle more diverse situations robustly and efficiently. Moreover, we propose a parameter sharing scheme for reducing memory usage and computational burden induced by the agglomeration module. Extensive experimental results on several point cloud analytic tasks, including classification and segmentation, well demonstrate the superior performance of our dynamic hierarchical learning framework over current state-of-the-art methods. Jinxian Liu, Bingbing Ni, Caiyuan Li, Jiancheng Yang, Qi Tian 0001 |
ICCV | 1 |
| 2019 | Multi-level attention model for person re-identification
Yichao Yan, Bingbing Ni, Jinxian Liu, Xiaokang Yang 0001 |
Pattern Recognit. Lett. | 3 |
| 2018 | Pose Transferrable Person Re-IdentificationabstractPerson re-identification (ReID) is an important task in the field of intelligent security. A key challenge is how to capture human pose variations, while existing benchmarks (i.e., Market1501, DukeMTMC-reID, CUHK03, etc.) do NOT provide sufficient pose coverage to train a robust ReID system. To address this issue, we propose a pose-transferrable person ReID framework which utilizes pose-transferred sample augmentations (i.e., with ID supervision) to enhance ReID model training. On one hand, novel training samples with rich pose variations are generated via transferring pose instances from MARS dataset, and they are added into the target dataset to facilitate robust training. On the other hand, in addition to the conventional discriminator of GAN (i.e., to distinguish between REAL/FAKE samples), we propose a novel guider sub-network which encourages the generated sample (i.e., with novel pose) towards better satisfying the ReID loss (i.e., cross-entropy ReID loss, triplet ReID loss). In the meantime, an alternative optimization procedure is proposed to train the proposed Generator-Guider-Discriminator network. Experimental results on Market-1501, DukeMTMC-reID and CUHK03 show that our method achieves great performance improvement, and outperforms most state-of-the-art methods without elaborate designing the ReID model. Jinxian Liu, Bingbing Ni, Yichao Yan, Peng Zhou 0010, Jianguo Hu |
CVPR | 1 |