Xuejin Chen

dblp:17/4378 · DBLP profile ↗
← Back
92ranked-venue papers
5as first author
49since 2021 · last 2026
0000-0003-0478-7018ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 73 · 3 first-author · 33 since 2021Artificial intelligence and machine learning · 22 · 1 first-author · 19 since 2021Applied, interdisciplinary, general and emerging computing · 18 · 1 first-author · 12 since 2021Systems, architecture and hardware · 2 · 1 since 2021Computer networks · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Neuron Segment Connectivity Prediction With Multimodal Features for Connectomics
abstract
Reconstructing neurons from large electron microscopy (EM) datasets for connectomic analysis presents a significant challenge, particularly in segmenting neurons of complex morphologies. Previous deep learning-based neuron segmentation methods often rely on pixel-level image context and produce extensive oversegmented fragments. Detecting these split errors and merging the split neuron segments are non-trivial for various neurons in a large-scale EM data volume. In this work, we exploit multimodal features in the full workflow of automatic neuron proofreading. We propose a novel connection point detection network that utilizes both global 3D morphological features and high-resolution local image context to extract candidate segment pairs from massive adjacent segments. To effectively fuse the 3D morphological feature and the dense image features from very different scales, we design a proposal-based image feature sampling to improve the efficiency of multimodal cross-attentions. Integrating the connection point detection network with our connectivity prediction network which also utilizes multimodal features, we make a fully automatic neuron segment merging pipeline, closely imitating human proofreading. Comprehensive experimental results verify the effectiveness of the proposed modules and demonstrate the robustness of the entire pipeline in large-scale neuron reconstruction. The code and data are available at https://github.com/Levishery/Neuron-Segment-Connection-Prediction.
Qihua Chen, Xuejin Chen, Chenxuan Wang, Zhiwei Xiong, Feng Wu 0005
IEEE Trans. Medical Imaging2
2026 Long Video Understanding With Learnable Retrieval in Video-Language Models
abstract
The remarkable natural language understanding, reasoning, and generation capabilities of large language models (LLMs) have made them attractive for application to video understanding, utilizing video tokens as contextual input. However, employing LLMs for long video understanding presents significant challenges. The extensive number of video tokens leads to considerable computational costs for LLMs while using aggregated tokens results in loss of vision details. Moreover, the presence of abundant question-irrelevant tokens introduces noise to the video reasoning process. To address these issues, we introduce a simple yet effective learnable retrieval-based video-language model (R-VLM) for efficient long video understanding. Specifically, given a question and a long video, our model identifies the most relevant$K$video chunks and uses their associated visual tokens to serve as context for the LLM inference. This effectively reduces the number of video tokens, eliminates noise interference, and enhances system performance. We achieve this by incorporating a learnable lightweight MLP block to facilitate the efficient retrieval of question-relevant chunks, through the end-to-end training of our video-language model with a proposed soft matching loss. Experimental results on multiple zero-shot video question answering datasets validate the effectiveness of our framework for comprehending long videos.
Cuiling Lan, Wenxuan Xie, Xuejin Chen, Yan Lu 0001
IEEE Trans. Multim.4
2026 Hie4DGS: Hierarchical 4D Gaussian Splatting From Monocular Dynamic Video
abstract
Monocular dynamic video reconstruction is a typical ill-posed problem due to the limited observations and complex 3D motions. Despite the recent advances in dynamic 3D Gaussian splatting techniques, most of them still struggle with the monocular setting, since they heavily rely on geometric cues from multiple cameras or ignore the structural coherence among the optimized 3D Gaussains. To address this, we propose Hie4DGS, a novel hierarchical structure representation to model the complex dynamic motions from monocular dynamic videos. Specifically, we decompose the motions of a dynamic scene into groups of multiple structure granularities and progressively compose them to derive the motion of each 3D Gaussian. Building on this representation, we leverage hierarchical semantic segmentation to group Gaussians and initialize their motion using depth and tracking priors within each group. Additionally, we introduce a structure rendering loss that enforces consistency between the learned motion structure and semantic priors, further reducing motion ambiguity. Compared to the state-of-the-art dynamic Gaussian methods, we achieve significant improvement in rendering quality on monocular video datasets featuring complex real-world motions.
Kaizhi Yang, Xiaoxiao Long, Xuejin Chen
IEEE Trans. Vis. Comput. Graph.5
2025 Efficient Modeling of Long-Range Morphology for 3D Neuron Reconstruction from Electron Microscopy Images
Chenxuan Wang, Qihua Chen, Xuejin Chen
CGI (2)4
2025 MC-Gaussian: 3D Gaussian Splatting with Multi-camera Setup in Autonomous Driving
Xuejin Chen
CGI (2)3
2025 Resgs: Residual Densification of 3D Gaussian for Efficient Detail Recovery
abstract
Recently, 3D Gaussian Splatting (3D-GS) has prevailed in novel view synthesis, achieving high fidelity and efficiency. However, it often struggles to capture rich details and complete geometry. Our analysis reveals that the 3D-GS densification operation lacks adaptiveness and faces a dilemma between geometry coverage and detail recovery. To address this, we introduce a novel densification operation, residual split, which adds a downscaled Gaussian as a residual. Our approach is capable of adaptively retrieving details and complementing missing geometry. To further support this method, we propose a pipeline named ResGS. Specifically, we integrate a Gaussian image pyramid for progressive supervision and implement a selection scheme that prioritizes the densification of coarse Gaussians over time. Extensive experiments demonstrate that our method achieves SOTA rendering quality. Consistent performance improvements can be achieved by applying our residual split on various 3D-GS variants, underscoring its versatility and potential for broader application in 3D-GS-based applications.
Yanzhe Lyu, Xuejin Chen
ICCV4
2025 Hierarchical Part-Based Generative Model for Realistic 3D Blood Vessel
Jiahao Lai, Bingzhi Shen, Sihong Zhang, Caixia Dong, Xuejin Chen, Yang Li 0104
MICCAI (3)7
2025 CausalCtrl: Causality-Aware Control Framework for Text-Guided Visual Editing
abstract
Text-guided visual editing aims to modify visual content according to a target prompt while faithfully preserving the structure and identity of the source image or video. However, existing methods ignore confounding effects brought from the pretrained model, i.e., harmful biases learned from the pretraining datasets, leading to spurious correlations during the editing processing. To address this issue, we introduce CausalCtrl, a novel training-free framework that reformulates text-guided visual editing from a causal inference perspective. The core idea is to leverage frontdoor adjustment to estimate the interventional distribution of the output, effectively blocking the influence of hidden confounders introduced by the pretrained model. Specifically, we first design a dual-branch inversion mechanism that disentangles the source content and target semantics into two separate latent embeddings to simplify the sampling space of interventional operation, and perform unbiased denoising through their controlled interaction. Besides, we propose a Structured Attention Injection Module (SAIM) that adaptively identifies and amplifies dominant attention heads using a lightweight SVD-based top-K selection strategy. Extensive experiments on several challenging image and video editing benchmarks demonstrate that CausalCtrl consistently outperforms existing methods in both target semantic alignment and source content preservation, validating the effectiveness of causal intervention in this task.
Haoxiang Cao, Chaoqun Wang 0011, Yongwen Lai, Shaobo Min, Xuejin Chen
ACM Multimedia5
2025 Robust visual place recognition with adaptive deformable token aggregation
Chaoqun Wang 0011, Shaobo Min, Xuejin Chen
Comput. Graph.4
2025 BioSAM: Generating SAM Prompts From Superpixel Graph for Biological Instance Segmentation
abstract
Proposal-free instance segmentation methods have significantly advanced the field of biological image analysis. Recently, the Segment Anything Model (SAM) has shown an extraordinary ability to handle challenging instance boundaries. However, directly applying SAM to biological images that contain instances with complex morphologies and dense distributions fails to yield satisfactory results. In this work, we propose BioSAM, a new biological instance segmentation framework generating SAM prompts from a superpixel graph. Specifically, to avoid over-merging, we first generate sufficient superpixels as graph nodes and construct an initialized graph. We then generate initial prompts from each superpixel and aggregate them through a graph neural network (GNN) by predicting the relationship of superpixels to avoid over-segmentation. We employ the SAM encoder embeddings and the SAM-assisted superpixel similarity as new features for the graph to enhance its discrimination capability. With the graph-based prompt aggregation, we utilize the aggregated prompts in SAM to refine the segmentation and generate more accurate instance boundaries. Comprehensive experiments on four representative biological datasets demonstrate that our proposed method outperforms state-of-the-art methods.
Xiaoyu Liu 0006, Zhiwei Xiong, Xuejin Chen
IEEE J. Biomed. Health Informatics4
2025 Region-Enhanced Feature Learning for Scene Semantic Segmentation
abstract
Semantic segmentation in complex scenes relies not only on object appearance but also on object location and the surrounding environment. Nonetheless, it is difficult to model long-range context in the format of pairwise point correlations due to the huge computational cost for large-scale point clouds. In this paper, we propose using regions as the intermediate representation of point clouds instead of fine-grained points or voxels to reduce the computational burden. We introduce a novel Region-Enhanced Feature Learning Network (REFL-Net) that leverages region correlations to enhance point feature learning. We design a region-based feature enhancement (RFE) module, which consists of a Semantic-Spatial Region Extraction stage and a Region Dependency Modeling stage. In the first stage, the input points are grouped into a set of regions based on their semantic and spatial proximity. In the second stage, we explore inter-region semantic and spatial relationships by employing a self-attention block on region features and then fuse point features with the region features to obtain more discriminative representations. Our proposed RFE module is plug-and-play and can be integrated with common semantic segmentation backbones. We conduct extensive experiments on ScanNetV2 and S3DIS datasets and evaluate our RFE module with different segmentation backbones. Our REFL-Net achieves 1.8% mIoU gain on ScanNetV2 and 1.7% mIoU gain on S3DIS with negligible computational cost compared with backbone models. Both quantitative and qualitative results show the powerful long-range context modeling ability and strong generalization ability of our REFL-Net.
Chaoqun Wang 0011, Xuejin Chen
IEEE Trans. Multim.3
2025 IMLS-Splatting: Efficient Mesh Reconstruction from Multi-view Images via Point Representation
abstract
Multi-view mesh reconstruction has long been a challenging problem in graphics and computer vision. In contrast to recent volumetric rendering methods that generate meshes through post-processing, we propose an end-to-end mesh optimization approach called IMLS-Splatting. Our method leverages the sparsity and flexibility of point clouds to efficiently represent the underlying surface. To achieve this, we introduce a splatting-based differentiable Implicit Moving-Least Squares (IMLS) algorithm that enables the fast conversion of point clouds into SDFs and texture fields, optimizing both mesh reconstruction and rasterization. Additionally, the IMLS representation ensures that the reconstructed SDF and mesh maintain continuity and smoothness without the need for extra regularization. With this efficient pipeline, our method enables the reconstruction of highly detailed meshes in approximately 11 minutes, supporting high-quality rendering and achieving state-of-the-art reconstruction performance. Our code is available at https://github.com/SilenKZYoung/IMLS-Splatting.
Kaizhi Yang, Liu Dai, Isabella Liu, Xiaoshuai Zhang, Xiaoyan Sun 0001, Xuejin Chen, Zexiang Xu, Hao Su 0001
ACM Trans. Graph.6
2024 Self-Supervised Learning of Skeleton-Aware Morphological Representation for 3D Neuron Segments
abstract
Effective morphological analysis of large-scale 3D neural data plays a crucial role in neuroscience research. However, the ultra-scale data volume from high-resolution microscopy imaging makes manual analysis significantly challenging for 3D rendering, morphological analysis, and morphology-based neuron classification. In this paper, we propose a self-supervised approach to learn skeleton-aware morphological representations from ultra-scale 3D segments to support efficient rendering and morphological analysis. Our approach, named ConSkeletonNet, connects skeleton-aware shape simplification and morphology-based neuron classification to enhance the discriminability of learned morphological representations through multi-task joint training. Through experiments on the neuron segments in a full fly brain EM FAFB-FFN1 and a data volume in the cerebral cortex of a human H01, our ConSkeletonNet shows superiority in learning skeletal-aware morphological representation for both neuron segments skeleton extraction and neuron classification. We also apply our ConSkeleton-Net to the 3D model dataset of man-made objects ShapeNet and achieve state-of-the-art performance in skeleton extraction and shape abstraction.
Daiyi Zhu, Qihua Chen, Xuejin Chen
3DV3
2024 Learning Multimodal Volumetric Features for Large-Scale Neuron Tracing
abstract
The current neuron reconstruction pipeline for electron microscopy (EM) data usually includes automatic image segmentation followed by extensive human expert proofreading. In this work, we aim to reduce human workload by predicting connectivity between over-segmented neuron pieces, taking both microscopy image and 3D morphology features into account, similar to human proofreading workflow. To this end, we first construct a dataset, named FlyTracing, that contains millions of pairwise connections of segments expanding the whole fly brain, which is three orders of magnitude larger than existing datasets for neuron segment connection. To learn sophisticated biological imaging features from the connectivity annotations, we propose a novel connectivity-aware contrastive learning method to generate dense volumetric EM image embedding. The learned embeddings can be easily incorporated with any point or voxel-based morphological representations for automatic neuron tracing. Extensive comparisons of different combination schemes of image and morphological representation in identifying split errors across the whole fly brain demonstrate the superiority of the proposed approach, especially for the locations that contain severe imaging artifacts, such as section missing and misalignment. The dataset and code are available at https://github.com/Levishery/Flywire-Neuron-Tracing.
Qihua Chen, Xuejin Chen, Chenxuan Wang, Yixiong Liu, Zhiwei Xiong, Feng Wu 0001
AAAI2
2024 Hierarchical Intra-Modal Correlation Learning for Label-Free 3D Semantic Segmentation
abstract
Recent methods for label-free 3D semantic segmentation aim to assist 3D model training by leveraging the open-world recognition ability of pre-trained vision language models. However, these methods usually suffer from in-consistent and noisy pseudo-labels provided by the vision language models. To address this issue, we present a hierarchical intra-modal correlation learning framework that captures visual and geometric correlations in 3D scenes at three levels: intra-set, intra-scene, and inter-scene, to help learn more compact 3D representations. We refine pseudo-labels using intra-set correlations within each geometric consistency set and align features of visually and geometrically similar points using intra-scene and inter-scene correlation learning. We also introduce a feedback mechanism to distill the correlation learning capability into the 3D model. Experiments on both indoor and outdoor datasets show the superiority of our method. We achieve a state-of-the-art 36.6% mIoU on the ScanNet dataset, and a 23.0% mIoU on the nuScenes dataset, with improvements of 7.8% mIoU and 2.2% mIoU compared with previous SOTA. We also provide theoretical analysis and qualitative visualization results to discuss the mechanism and conduct thorough ablation studies to support the effectiveness of our framework.
Jiahao Li 0001, Xuejin Chen, Yan Lu 0001
CVPR4
2024 Cross-dimension Affinity Distillation for 3D EM Neuron Segmentation
abstract
Accurate 3D neuron segmentation from electron mi-croscopy (EM) volumes is crucial for neuroscience re-search. However, the complex neuron morphology often leads to over-merge and over-segmentation results. Recent advancements utilize 3D CNNs to predict a 3D affinity map with improved accuracy but suffer from two challenges: high computational cost and limited input size, especially for practical deployment for large-scale EM volumes. To address these challenges, we propose a novel method to leverage lightweight 2D CNNs for efficient neuron segmen-tation. Our method employs a 2D Y-shape network to generate two embedding maps from adjacent 2D sections, which are then converted into an affinity map by measuring their embedding distance. While the 2D network better captures pixel dependencies inside sections with larger in-put sizes, it overlooks inter-section dependencies. To over-come this, we introduce a cross-dimension affinity distillation (CAD) strategy that transfers inter-section dependency knowledge from a 3D teacher network to the 2D student network by ensuring consistency between their output affin-ity maps. Additionally, we design a feature grafting in-teraction (FGI) module to enhance knowledge transfer by grafting embedding maps from the 2D student onto those from the 3D teacher. Extensive experiments on multiple EM neuron segmentation datasets, including a newly built one by ourselves, demonstrate that our method achieves supe-rior performance over state-of-the-art methods with only 1/20 inference latency. We release our code and dataset at https://github.com/liuxyll03/CAD.
Xiaoyu Liu 0006, Yinda Chen, Yueyi Zhang 0001, Te Shi 0003, Ruobing Zhang, Xuejin Chen, Zhiwei Xiong
CVPR7
2024 UC-NERF: Neural Radiance Field for Under-Calibrated Multi-View Cameras in Autonomous Driving
abstract
Multi-camera setups find widespread use across various applications, such as autonomous driving, as they greatly expand sensing capabilities. Despite the fast development of Neural radiance field (NeRF) techniques and their wide applications in both indoor and outdoor scenes, applying NeRF to multi-camera systems remains very challenging. This is primarily due to the inherent under-calibration issues in multi-camera setup, including inconsistent imaging effects stemming from separately calibrated image signal processing units in diverse cameras, and system errors arising from mechanical vibrations during driving that affect relative camera poses. In this paper, we present UC-NeRF, a novel method tailored for novel view synthesis in under-calibrated multi-view camera systems. Firstly, we propose a layer-based color correction to rectify the color inconsistency in different image regions. Second, we propose virtual warping to generate more viewpoint-diverse but color-consistent virtual views for color correction and 3D recovery. Finally, a spatiotemporally constrained pose refinement is designed for more robust and accurate pose calibration in multi-camera systems. Our method not only achieves state-of-the-art performance of novel view synthesis in multi-camera setups, but also effectively facilitates depth estimation in large-scale outdoor scenes with the synthesized novel views.
Xiaoxiao Long, Wei Yin 0006, Jin Wang 0001, Zhiqiang Wu 0001, Yuexin Ma, Xiaozhi Chen, Xuejin Chen
ICLR9
2024 MovingParts: Motion-based 3D Part Discovery in Dynamic Radiance Field
abstract
We present MovingParts, a NeRF-based method for dynamic scene reconstruction and part discovery. We consider motion as an important cue for identifying parts, that all particles on the same part share the common motion pattern. From the perspective of fluid simulation, existing deformation-based methods for dynamic NeRF can be seen as parameterizing the scene motion under the Eulerian view, i.e., focusing on specific locations in space through which the fluid flows as time passes. However, it is intractable to extract the motion of constituting objects or parts using the Eulerian view representation. In this work, we introduce the dual Lagrangian view and enforce representations under the Eulerian/Lagrangian views to be cycle-consistent. Under the Lagrangian view, we parameterize the scene motion by tracking the trajectory of particles on objects. The Lagrangian view makes it convenient to discover parts by factorizing the scene motion as a composition of part-level rigid motions. Experimentally, our method can achieve fast and high-quality dynamic scene reconstruction from even a single moving camera, and the induced part-based representation allows direct applications of part tracking, animation, 3D scene editing, etc.
Kaizhi Yang, Xiaoshuai Zhang, Zhiao Huang, Xuejin Chen, Zexiang Xu, Hao Su 0001
ICLR4
2024 GaussianPro: 3D Gaussian Splatting with Progressive Propagation
abstract
3D Gaussian Splatting (3DGS) has recently revolutionized the field of neural rendering with its high fidelity and efficiency. However, 3DGS heavily depends on the initialized point cloud produced by Structure-from-Motion (SfM) techniques. When tackling large-scale scenes that unavoidably contain texture-less surfaces, SfM techniques fail to produce enough points in these surfaces and cannot provide good initialization for 3DGS. As a result, 3DGS suffers from difficult optimization and low-quality renderings. In this paper, inspired by classic multi-view stereo (MVS) techniques, we propose GaussianPro, a novel method that applies a progressive propagation strategy to guide the densification of the 3D Gaussians. Compared to the simple split and clone strategies used in 3DGS, our method leverages the priors of the existing reconstructed geometries of the scene and utilizes patch matching to produce new Gaussians with accurate positions and orientations. Experiments on both large-scale and small-scale scenes validate the effectiveness of our method. Our method significantly surpasses 3DGS on the Waymo dataset, exhibiting an improvement of 1.15dB in terms of PSNR. Codes and data are available at https://github.com/kcheng1021/GaussianPro.
Xiaoxiao Long, Kaizhi Yang, Yao Yao 0008, Wei Yin 0006, Yuexin Ma, Wenping Wang 0001, Xuejin Chen
ICML8
2024 Slot-VLM: Object-Event Slots for Video-Language Modeling
abstract
Video-Language Models (VLMs), powered by the advancements in Large Language Models (LLMs), are charting new frontiers in video understanding. A pivotal challenge is the development of an effective method to encapsulate video content into a set of representative tokens to align with LLMs. In this work, we introduce Slot-VLM, a new framework designed to generate semantically decomposed video tokens, in terms of object-wise and event-wise visual representations, to facilitate LLM inference. Particularly, we design an Object-Event Slots module, i.e., OE-Slots, that adaptively aggregates the dense video tokens from the vision encoder to a set of representative slots. In order to take into account both the spatial object details and the varied temporal dynamics, we build OE-Slots with two branches: the Object-Slots branch and the Event-Slots branch. The Object-Slots branch focuses on extracting object-centric slots from features of high spatial resolution but low frame sample rate, emphasizing detailed object information. The Event-Slots branch is engineered to learn event-centric slots from high temporal sample rate but low spatial resolution features. These complementary slots are combined to form the vision context, serving as the input to the LLM for effective video reasoning. Our experimental results demonstrate the effectiveness of our Slot-VLM, which achieves the state-of-the-art performance on video question-answering.
Cuiling Lan, Wenxuan Xie, Xuejin Chen, Yan Lu 0001
NeurIPS4
2024 PCLC-Net: Point Cloud Completion in Arbitrary Poses with Learnable Canonical Space
abstract
Abstract Recovering the complete structure from partial point clouds in arbitrary poses is challenging. Recently, many efforts have been made to address this problem by developing SO(3)‐equivariant completion networks or aligning the partial point clouds with a predefined canonical space before completion. However, these approaches are limited to random rotations only or demand costly pose annotation for model training. In this paper, we present a novel Network for Point cloud Completion with Learnable Canonical space (PCLC‐Net) to reduce the need for pose annotations and extract SE(3)‐invariant geometry features to improve the completion quality in arbitrary poses. Without pose annotations, our PCLC‐Net utilizes self‐supervised pose estimation to align the input partial point clouds to a canonical space that is learnable for an object category and subsequently performs shape completion in the learned canonical space. Our PCLC‐Net can complete partial point clouds with arbitrary SE(3) poses without requiring pose annotations for supervision. Our PCLC‐Net achieves state‐of‐the‐art results on shape completion with arbitrary SE(3) poses on both synthetic and real scanned data. To the best of our knowledge, our method is the first to achieve shape completion in arbitrary poses without pose annotations during network training.
Hanmo Xu, Qingyao Shuai, Xuejin Chen
Comput. Graph. Forum3
2024 Context-Aware Proposal-Boundary Network With Structural Consistency for Audiovisual Event Localization
abstract
Audiovisual event localization aims to localize the event that is both visible and audible in a video. Previous works focus on segment-level audio and visual feature sequence encoding and neglect the event proposals and boundaries, which are crucial for this task. The event proposal features provide event internal consistency between several consecutive segments constructing one proposal, while the event boundary features offer event boundary consistency to make segments located at boundaries be aware of the event occurrence. In this article, we explore the proposal-level feature encoding and propose a novel context-aware proposal-boundary (CAPB) network to address audiovisual event localization. In particular, we design a local-global context encoder (LGCE) to aggregate local-global temporal context information for visual sequence, audio sequence, event proposals, and event boundaries, respectively. The local context from temporally adjacent segments or proposals contributes to event discrimination, while the global context from the entire video provides semantic guidance of temporal relationship. Furthermore, we enhance the structural consistency between segments by exploiting the above-encoded proposal and boundary representations. CAPB leverages the context information and structural consistency to obtain context-aware event-consistent cross-modal representation for accurate event localization. Extensive experiments conducted on the audiovisual event (AVE) dataset show that our approach outperforms the state-of-the-art methods by clear margins in both supervised event localization and cross-modality localization.
Hao Wang 0161, Zhengjun Zha, Liang Li 0003, Xuejin Chen, Jiebo Luo 0001
IEEE Trans. Neural Networks Learn. Syst.4
2024 Graph Representation Learning for Large-Scale Neuronal Morphological Analysis
abstract
The analysis of neuronal morphological data is essential to investigate the neuronal properties and brain mechanisms. The complex morphologies, absence of annotations, and sheer volume of these data pose significant challenges in neuronal morphological analysis, such as identifying neuron types and large-scale neuron retrieval, all of which require accurate measuring and efficient matching algorithms. Recently, many studies have been conducted to describe neuronal morphologies quantitatively using predefined measurements. However, hand-crafted features are usually inadequate for distinguishing fine-grained differences among massive neurons. In this article, we propose a novel morphology-aware contrastive graph neural network (MACGNN) for unsupervised neuronal morphological representation learning. To improve the retrieval efficiency in large-scale neuronal morphological datasets, we further propose Hash-MACGNN by introducing an improved deep hash algorithm to train the network end-to-end to learn binary hash representations of neurons. We conduct extensive experiments on the largest dataset, NeuroMorpho, which contains more than 100 000 neurons. The experimental results demonstrate the effectiveness and superiority of our MACGNN and Hash-MACGNN for large-scale neuronal morphological analysis.
Jie Zhao 0020, Xuejin Chen, Zhiwei Xiong, Zhengjun Zha, Feng Wu 0001
IEEE Trans. Neural Networks Learn. Syst.2
2023 Paint by Example: Exemplar-based Image Editing with Diffusion Models
abstract
Language-guided image editing has achieved great success recently. In this paper, we investigate exemplar-guided image editing for more precise control. We achieve this goal by leveraging self-supervised training to disentangle and re-organize the source image and the exemplar. However, the naive approach will cause obvious fusing artifacts. We carefully analyze it and propose a content bottleneck and strong augmentations to avoid the trivial solution of directly copying and pasting the exemplar image. Meanwhile, to ensure the controllability of the editing process, we design an arbitrary shape mask for the exemplar image and leverage the classifier-free guidance to increase the similarity to the exemplar image. The whole framework involves a single forward of the diffusion model without any iterative optimization. We demonstrate that our method achieves an impressive performance and enables controllable editing on in-the-wild images with high fidelity. The code and pretrained models are available at https://github.com/Fantasy-Studio/Paint-by-Example.
Binxin Yang, Shuyang Gu, Bo Zhang 0025, Ting Zhang 0002, Xuejin Chen, Xiaoyan Sun 0001, Dong Chen 0003, Fang Wen 0001
CVPR5
2023 Learning Cross-Representation Affinity Consistency for Sparsely Supervised Biomedical Instance Segmentation
abstract
Sparse instance-level supervision has recently been explored to address insufficient annotation in biomedical instance segmentation, which is easier to annotate crowded instances and better preserves instance completeness for 3D volumetric datasets compared to common semi-supervision. In this paper, we propose a sparsely supervised biomedical instance segmentation framework via cross-representation affinity consistency regularization. Specifically, we adopt two individual networks to enforce the perturbation consistency between an explicit affinity map and an implicit affinity map to capture both feature-level instance discrimination and pixel-level instance boundary structure. We then select the highly confident region of each affinity map as the pseudo label to supervise the other one for affinity consistency learning. To obtain the highly confident region, we propose a pseudo-label noise filtering scheme by integrating two entropy-based decision strategies. Extensive experiments on four biomedical datasets with sparse instance annotations show the state-of-the-art performance of our proposed framework. For the first time, we demonstrate the superiority of sparse instance-level supervision on 3D volumetric datasets, compared to common semi-supervision under the same annotation cost. Code is available at https://github.com/liuxy1103/CRAC.
Xiaoyu Liu 0006, Wei Huang 0036, Zhiwei Xiong, Shenglong Zhou 0002, Yueyi Zhang 0001, Xuejin Chen, Zhengjun Zha, Feng Wu 0001
ICCV6
2023 DPF-Net: Combining Explicit Shape Priors in Deformable Primitive Field for Unsupervised Structural Reconstruction of 3D Objects
abstract
Unsupervised methods for reconstructing structures face significant challenges in capturing the geometric details with consistent structures among diverse shapes of the same category. To address this issue, we present a novel unsupervised structural reconstruction method, named DPF-Net, based on a new Deformable Primitive Field (DPF) representation, which allows for high-quality shape reconstruction using parameterized geometric primitives. We design a two-stage shape reconstruction pipeline which consists of a primitive generation module and a primitive deformation module to approximate the target shape of each part progressively. The primitive generation module estimates the explicit orientation, position, and size parameters of parameterized geometric primitives, while the primitive deformation module predicts a dense deformation field based on a parameterized primitive field to recover shape details. The strong shape prior encoded in parameterized geometric primitives enables our DPF-Net to extract high-level structures and recover fine-grained shape details consistently. The experimental results on three categories of objects in diverse shapes demonstrate the effectiveness and generalization ability of our DPF-Net on structural reconstruction and shape segmentation.
Qingyao Shuai, Chi Zhang 0044, Kaizhi Yang, Xuejin Chen
ICCV4
2023 Structure-Aware Point Cloud Completion
Zhihua Cheng, Xuejin Chen
ICIG (3)2
2023 Cascade-LogoNet: Eliminating Classification Ambiguity with Cascaded Logo Detection
abstract
Due to a large number of brand logos and similar text appearances, existing deep detection networks are prone to generating multiple overlapping logos at a single location. In this paper, we propose a two-stage approach, Cascade-LogoNet, by adding a verification stage to eliminate the classification ambiguity. We group overlapping logo proposals and verify the class predictions inside each group by learning more discriminative features with strong inter-class discrepancy. We also design a lightweight detector to improve computational efficiency and prevent over-fitting given limited training data. In addition, we propose a new data augmentation strategy, named small logo preserving stitching (SLPS), to improve the detection accuracy of small logos. When stitching training images to increase the loss ratio of small logos, we keep the original small logos on the stitched image to avoid extremely undistinguishable small ones. Our Cascade-LogoNet achieves 63.5 mAP and 64.4 mAP on the dataset QMUL-OpenLogo [1] when using ResNet-50 and ResNet-50-DCN [2] as backbone respectively, which surpasses previous methods by a large margin.
Teqiang Zou, Xuejin Chen, Yan Chen 0007, Qibin Sun
ISCAS2
2023 Class-Aware Feature Alignment for Domain Adaptative Mitochondria Segmentation
Dan Yin, Wei Huang 0036, Zhiwei Xiong, Xuejin Chen
MICCAI (4)4
2023 Neural message-passing for objective-based uncertainty quantification and optimal experimental design
Qihua Chen, Xuejin Chen, Hyun-Myung Woo, Byung-Jun Yoon
Eng. Appl. Artif. Intell.2
2023 Semantic and Relation Modulation for Audio-Visual Event Localization
abstract
We study the problem of localizing audio-visual events that are both audible and visible in a video. Existing works focus on encoding and aligning audio and visual features at the segment level while neglecting informative correlation between segments of the two modalities and between multi-scale event proposals. We propose a novel Semantic and Relation Modulation Network (SRMN) to learn the above correlation and leverage it to modulate the related auditory, visual, and fused features. In particular, for semantic modulation, we propose intra-modal normalization and cross-modal normalization. The former modulates features of a single modality with the event-relevant semantic guidance of the same modality. The latter modulates features of two modalities by establishing and exploiting the cross-modal relationship. For relation modulation, we propose a multi-scale proposal modulating module and a multi-alignment segment modulating module to introduce multi-scale event proposals and enable dense matching between cross-modal segments, which strengthen correlations between successive segments within one proposal and between all segments. With the features modulated by the correlation information regarding audio-visual events, SRMN performs accurate event localization. Extensive experiments conducted on the public AVE dataset demonstrate that our method outperforms the state-of-the-art methods in both supervised event localization and cross-modality localization tasks.
Hao Wang 0161, Zhengjun Zha, Liang Li 0003, Xuejin Chen, Jiebo Luo 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 Current Progress and Challenges in Large-Scale 3D Mitochondria Instance Segmentation
abstract
In this paper, we present the results of the MitoEM challenge on mitochondria 3D instance segmentation from electron microscopy images, organized in conjunction with the IEEE-ISBI 2021 conference. Our benchmark dataset consists of two large-scale 3D volumes, one from human and one from rat cortex tissue, which are 1,986 times larger than previously used datasets. At the time of paper submission, 257 participants had registered for the challenge, 14 teams had submitted their results, and six teams participated in the challenge workshop. Here, we present eight top-performing approaches from the challenge participants, along with our own baseline strategies. Posterior to the challenge, annotation errors in the ground truth were corrected without altering the final ranking. Additionally, we present a retrospective evaluation of the scoring system which revealed that: 1) challenge metric was permissive with the false positive predictions; and 2) size-based grouping of instances did not correctly categorize mitochondria of interest. Thus, we propose a new scoring system that better reflects the correctness of the segmentation results. Although several of the top methods are compared favorably to our own baselines, substantial errors remain unsolved for mitochondria with challenging morphologies. Thus, the challenge remains open for submission and automatic evaluation, with all volumes available for download.
Daniel Franco-Barranco, Zudi Lin, Won-Dong Jang, Xueying Wang 0002, Qijia Shen, Yutian Fan, Mingxing Li 0003, Chang Chen 0004, Zhiwei Xiong, Rui Xin 0003, Huai Chen, Zhili Li, Jie Zhao 0020, Xuejin Chen, Constantin Pape, Ryan Conrad, Luke Nightingale, Joost de Folter, Martin L. Jones, Dorsa Ziaei, Stephan Huschauer, Ignacio Arganda-Carreras, Hanspeter Pfister, Donglai Wei 0001
IEEE Trans. Medical Imaging16
2023 Semantics-Preserving Sketch Embedding for Face Generation
abstract
With recent advances in image-to-image translation tasks, remarkable progress has been witnessed in generating face images from sketches. However, existing methods frequently fail to generate images with details that are semantically and geometrically consistent with the input sketch, especially when various decoration strokes are drawn. To address this issue, we introduce a novel$\mathcal {W}$-$\mathcal {W^+}$encoder architecture to take advantage of the high expressive power of$\mathcal {W^+}$space and semantic controllability of$\mathcal {W}$space. We introduce an explicit intermediate representation for sketch semantic embedding. With a semantic feature matching loss for effective semantic supervision, our sketch embedding precisely conveys the semantics in the input sketches to the synthesized images. Moreover, a novel sketch semantic interpretation approach is designed to automatically extract semantics from vectorized sketches. We conduct extensive experiments on both synthesized sketches and hand-drawn sketches, and the results demonstrate the superiority of our method over existing approaches on both semantics-preserving and generalization ability.
Binxin Yang, Xuejin Chen, Chaoqun Wang 0011, Chi Zhang 0044, Xiaoyan Sun 0001
IEEE Trans. Multim.2
2022 Self-supervised Learning of Morphological Representation for 3D EM Segments with Cluster-Instance Correlations
Chi Zhang 0044, Qihua Chen, Xuejin Chen
MICCAI (8)3
2022 Semi-Supervised Neuron Segmentation via Reinforced Consistency Learning
abstract
Emerging deep learning-based methods have enabled great progress in automatic neuron segmentation from Electron Microscopy (EM) volumes. However, the success of existing methods is heavily reliant upon a large number of annotations that are often expensive and time-consuming to collect due to dense distributions and complex structures of neurons. If the required quantity of manual annotations for learning cannot be reached, these methods turn out to be fragile. To address this issue, in this article, we propose a two-stage, semi-supervised learning method for neuron segmentation to fully extract useful information from unlabeled data. First, we devise a proxy task to enable network pre-training by reconstructing original volumes from their perturbed counterparts. This pre-training strategy implicitly extracts meaningful information on neuron structures from unlabeled data to facilitate the next stage of learning. Second, we regularize the supervised learning process with the pixel-level prediction consistencies between unlabeled samples and their perturbed counterparts. This improves the generalizability of the learned model to adapt diverse data distributions in EM volumes, especially when the number of labels is limited. Extensive experiments on representative EM datasets demonstrate the superior performance of our reinforced consistency learning compared to supervised learning, i.e., up to 400% gain on the VOI metric with only a few available labels. This is on par with a model trained on ten times the amount of labeled data in a supervised manner. Code is available at https://github.com/weih527/SSNS-Net.
Wei Huang 0036, Chang Chen 0004, Zhiwei Xiong, Yueyi Zhang 0001, Xuejin Chen, Xiaoyan Sun 0001, Feng Wu 0001
IEEE Trans. Medical Imaging5
2021 Learning Scale-Adaptive Representations for Point-Level LiDAR Semantic Segmentation
abstract
A large number of objects with various scales and categories in autonomous driving scenes pose a great challenge to LiDAR semantic segmentation. Voxel-based 3D convolutional networks have been widely employed by existing state-of-the-art methods to extract features with different spatial scales. However, the voxel network architecture limits its effectiveness in combining multi-scale features for point-level discrimination. In this paper, we propose point-wise prediction by taking the geometric structure of the original point cloud into account. We propose a Scale-Adaptive Fusion (SAF) module that progressively and selectively fuses multi-scale features to deal with scale variations across objects adaptively. Moreover, we propose a novel Local Point Refinement (LPR) module to address the quantization loss problem of voxel-based methods. Our approach achieves state-of-the-art performance on three public datasets, i.e., Semantic-KITTI, Semantic-POSS, and nuScenes dataset, while greatly improving the computational and memory efficiency.
Tongfeng Zhang, Kaizhi Yang, Xuejin Chen
3DV3
2021 Task-Independent Knowledge Makes for Transferable Representations for Generalized Zero-Shot Learning
abstract
Generalized Zero-Shot Learning (GZSL) targets recognizing new categories by learning transferable image representations. Existing methods find that, by aligning image representations with corresponding semantic labels, the semantic-aligned representations can be transferred to unseen categories. However, supervised by only seen category labels, the learned semantic knowledge is highly task-specific, which makes image representations biased towards seen categories. In this paper, we propose a novel Dual-Contrastive Embedding Network (DCEN) that simultaneously learns task-specific and task-independent knowledge via semantic alignment and instance discrimination. First, DCEN leverages task labels to cluster representations of the same semantic category by cross-modal contrastive learning and exploring semantic-visual complementarity. Besides task-specific knowledge, DCEN then introduces task-independent knowledge by attracting representations of different views of the same image and repelling representations of different images. Compared to high-level seen category supervision, this instance discrimination supervision encourages DCEN to capture low-level visual knowledge, which is less biased toward seen categories and alleviates the representation bias. Consequently, the task-specific and task-independent knowledge jointly make for transferable representations of DCEN, which obtains averaged 4.1% improvement on four public benchmarks.
Chaoqun Wang 0011, Xuejin Chen, Shaobo Min, Xiaoyan Sun 0001, Houqiang Li
AAAI2
2021 S2R-DepthNet: Learning a Generalizable Depth-Specific Structural Representation
abstract
Human can infer the 3D geometry of a scene from a sketch instead of a realistic image, which indicates that the spatial structure plays a fundamental role in understanding the depth of scenes. We are the first to explore the learning of a depth-specific structural representation, which captures the essential feature for depth estimation and ignores irrelevant style information. Our S2R-DepthNet (Synthetic to Real DepthNet) can be well generalized to un-seen real-world data directly even though it is only trained on synthetic data. S2R-DepthNet consists of: a) a Structure Extraction (STE) module which extracts a domain-invariant structural representation from an image by dis-entangling the image into domain-invariant structure and domain-specific style components, b) a Depth-specific Attention (DSA) module, which learns task-specific knowledge to suppress depth-irrelevant structures for better depth estimation and generalization, and c) a depth prediction module (DP) to predict depth from the depth-specific representation. Without access of any real-world images, our method even outperforms the state-of-the-art unsupervised domain adaptation methods which use real-world images of the tar-get domain for training. In addition, when using a small amount of labeled real-world data, we achieve the state-of-the-art performance under the semi-supervised setting.
Xiaotian Chen, Yuwang Wang, Xuejin Chen, Wenjun Zeng 0001
CVPR3
2021 Depth Estimation for Colonoscopy Images with Self-supervised Learning from Videos
Yiting Ma, Xuejin Chen
MICCAI (6)5
2021 Learning Neuron Stitching for Connectomics
Xiaoyu Liu 0006, Yueyi Zhang 0001, Zhiwei Xiong, Chang Chen 0004, Wei Huang 0036, Xuejin Chen, Feng Wu 0001
MICCAI (8)6
2021 LDPolypVideo Benchmark: A Large-Scale Colonoscopy Video Dataset of Diverse Polyps
Yiting Ma, Xuejin Chen
MICCAI (5)2
2021 Uncertainty-Aware Label Rectification for Domain Adaptive Mitochondria Segmentation
Chang Chen 0004, Zhiwei Xiong, Xuejin Chen, Xiaoyan Sun 0001
MICCAI (3)4
2021 Deep 3D Modeling of Human Bodies from Freehand Sketching
Kaizhi Yang, Jintao Lu, Siyu Hu, Xuejin Chen
MMM (2)4
2021 Dual Progressive Prototype Network for Generalized Zero-Shot Learning
abstract
Generalized Zero-Shot Learning (GZSL) aims to recognize new categories with auxiliary semantic information, e.g., category attributes. In this paper, we handle the critical issue of domain shift problem, i.e., confusion between seen and unseen categories, by progressively improving cross-domain transferability and category discriminability of visual representations. Our approach, named Dual Progressive Prototype Network (DPPN), constructs two types of prototypes that record prototypical visual patterns for attributes and categories, respectively. With attribute prototypes, DPPN alternately searches attribute-related local regions and updates corresponding attribute prototypes to progressively explore accurate attribute-region correspondence. This enables DPPN to produce visual representations with accurate attribute localization ability, which benefits the semantic-visual alignment and representation transferability. Besides, along with progressive attribute localization, DPPN further projects category prototypes into multiple spaces to progressively repel visual representations from different categories, which boosts category discriminability. Both attribute and category prototypes are collaboratively learned in a unified framework, which makes visual representations of DPPN transferable and distinctive.Experiments on four benchmarks prove that DPPN effectively alleviates the domain shift problem in GZSL.
Chaoqun Wang 0011, Shaobo Min, Xuejin Chen, Xiaoyan Sun 0001, Houqiang Li
NeurIPS3
2021 Structure-Guided Deep Video Inpainting
abstract
A fundamental challenge in video inpainting is the difficulty of generating video contents with fine details, while keeping spatio-temporal coherence in the missing region. Recent studies focus on synthesizing temporally smooth pixels by exploiting the flow information, while ignoring maintaining the semantic structural coherence between frames. This makes them suffer from over-smoothing and blurry contours, which significantly reduce the visual quality of inpainting results. To address this issue, we present a novel structure-guided video inpainting approach that enhances temporal structure coherence to improve video inpainting results. In contrast to directly synthesizing the missing pixel colors, we first complete edges in the missing regions to depict scene structures and object shapes via an edge inpainting network with 3D convolutions. Then, we replenish textures using a coarse-to-fine synthesis network with a structure attention module (SAM), under the guidance of the synthesized edges. Specifically, our SAM is designed to model the semantic correlation between video textures and structural edges to generate more realistic content. Besides, motion flows between neighboring frames are employed to enhance temporal consistency for self-supervision during training the edge inpainting and texture inpainting modules. Consequently, the inpainting results using our approach are visually pleasing with fine details and temporal coherence. Experiments on the YouTubeVOS, DAVIS, and 300VW datasets show that our method obtains state-of-the-art performance under diverse video inpainting settings.
Chaoqun Wang 0011, Xuejin Chen, Shaobo Min, Jiaping Wang, Zhengjun Zha
IEEE Trans. Circuits Syst. Video Technol.2
2021 Weakly Supervised Neuron Reconstruction From Optical Microscopy Images With Morphological Priors
abstract
Manually labeling neurons from high-resolution but noisy and low-contrast optical microscopy (OM) images is tedious. As a result, the lack of annotated data poses a key challenge when applying deep learning techniques for reconstructing neurons from noisy and low-contrast OM images. While traditional tracing methods provide a possible way to efficiently generate labels for supervised network training, the generated pseudo-labels contain many noisy and incorrect labels, which lead to severe performance degradation. On the other hand, the publicly available dataset, BigNeuron, provides a large number of single 3D neurons that are reconstructed using various imaging paradigms and tracing methods. Though the raw OM images are not fully available for these neurons, they convey essential morphological priors for complex 3D neuron structures. In this paper, we propose a new approach to exploit morphological priors from neurons that have been reconstructed for training a deep neural network to extract neuron signals from OM images. We integrate a deep segmentation network in a generative adversarial network (GAN), expecting the segmentation network to be weakly supervised by pseudo-labels at the pixel level while utilizing the supervision of previously reconstructed neurons at the morphology level. In our morphological-prior-guided neuron reconstruction GAN, named MP-NRGAN, the segmentation network extracts neuron signals from raw images, and the discriminator network encourages the extracted neurons to follow the morphology distribution of reconstructed neurons. Comprehensive experiments on the public VISoR-40 dataset and BigNeuron dataset demonstrate that our proposed MP-NRGAN outperforms state-of-the-art approaches with less training effort.
Xuejin Chen, Chi Zhang 0044, Jie Zhao 0020, Zhiwei Xiong, Zhengjun Zha, Feng Wu 0001
IEEE Trans. Medical Imaging1
2021 A Mutually Attentive Co-Training Framework for Semi-Supervised Recognition
abstract
Self-training plays an important role in practical recognition applications where sufficient clean labels are unavailable. Existing methods focus on generating reliable pseudo labels to retrain a model, while ignoring the importance of improving model reliability to those inevitably mislabeled data. In this paper, we propose a novel Mutually Attentive Co-training Framework (MACF) that can effectively alleviate the negative impacts of incorrect labels on model retraining by exploring deep model disagreements. Specifically, MACF trains two symmetrical sub-networks that have the same input and are connected by several attention modules at different layers. Each attention module analyzes the inferred features from two sub-networks for the same input and feedback attention maps for them to indicate noisy gradients. This is realized by exploring the back-propagation process of incorrect labels at different layers to design attention modules. By multi-layer interception, the noisy gradients caused by incorrect labels can be effectively reduced for both sub-networks, leading to robust training to potential incorrect labels. In addition, a hierarchical distillation strategy is developed to improve the pseudo labels by aggregating the predictions from multi-models and data transformations. The experiments on six general benchmarks, including classification and biomedical segmentation, demonstrate that MACF is much robust to noisy labels than previous methods.
Shaobo Min, Xuejin Chen, Hongtao Xie 0001, Zhengjun Zha, Yongdong Zhang 0001
IEEE Trans. Multim.2
2021 Laplacian Pyramid Neural Network for Dense Continuous-Value Regression for Complex Scenes
abstract
Many computer vision tasks, such as monocular depth estimation and height estimation from a satellite orthophoto, have a common underlying goal, which is regression of dense continuous values for the pixels given a single image. We define them as dense continuous-value regression (DCR) tasks. Recent approaches based on deep convolutional neural networks significantly improve the performance of DCR tasks, particularly on pixelwise regression accuracy. However, it still remains challenging to simultaneously preserve the global structure and fine object details in complex scenes. In this article, we take advantage of the efficiency of Laplacian pyramid on representing multiscale contents to reconstruct high-quality signals for complex scenes. We design a Laplacian pyramid neural network (LAPNet), which consists of a Laplacian pyramid decoder (LPD) for signal reconstruction and an adaptive dense feature fusion (ADFF) module to fuse features from the input image. More specifically, we build an LPD to effectively express both global and local scene structures. In our LPD, the upper and lower levels, respectively, represent scene layouts and shape details. We introduce a residual refinement module to progressively complement high-frequency details for signal prediction at each level. To recover the signals at each individual level in the pyramid, an ADFF module is proposed to adaptively fuse multiscale image features for accurate prediction. We conduct comprehensive experiments to evaluate a number of variants of our model on three important DCR tasks, i.e., monocular depth estimation, single-image height estimation, and density map estimation for crowd counting. Experiments demonstrate that our method achieves new state-of-the-art performance in both qualitative and quantitative evaluation on the NYU-D V2 and KITTI for monocular depth estimation, the challenging Urban Semantic 3D (US3D) for satellite height estimation, and four challenging benchmarks for crowd counting. These results demonstrate that the proposed LAPNet is a universal and effective architecture for DCR problems.
Xuejin Chen, Xiaotian Chen, Xueyang Fu, Zhengjun Zha
IEEE Trans. Neural Networks Learn. Syst.1
2021 Unsupervised learning for cuboid shape abstraction via joint segmentation from point clouds
abstract
Representing complex 3D objects as simple geometric primitives, known as shape abstraction, is important for geometric modeling, structural analysis, and shape synthesis. In this paper, we propose an unsupervised shape abstraction method to map a point cloud into a compact cuboid representation. We jointly predict cuboid allocation as part segmentation and cuboid shapes and enforce the consistency between the segmentation and shape abstraction for self-learning. For the cuboid abstraction task, we transform the input point cloud into a set of parametric cuboids using a variational auto-encoder network. The segmentation network allocates each point into a cuboid considering the point-cuboid affinity. Without manual annotations of parts in point clouds, we design four novel losses to jointly supervise the two branches in terms of geometric similarity and cuboid compactness. We evaluate our method on multiple shape collections and demonstrate its superiority over existing shape abstraction methods. Moreover, based on our network architecture and learned representations, our approach supports various applications including structured shape generation, shape interpolation, and structural shape clustering.
Kaizhi Yang, Xuejin Chen
ACM Trans. Graph.2
2020 Isotropic Reconstruction of 3D EM Images with Unsupervised Degradation Learning
Shiyu Deng, Xueyang Fu, Zhiwei Xiong, Chang Chen 0004, Dong Liu 0002, Xuejin Chen, Qing Ling 0001, Feng Wu 0001
MICCAI (5)6
2020 Towards Neuron Segmentation from Macaque Brain Images: A Weakly Supervised Approach
Meng Dong, Dong Liu 0002, Zhiwei Xiong, Xuejin Chen, Yueyi Zhang 0001, Zhengjun Zha, Guoqiang Bi, Feng Wu 0001
MICCAI (5)4
2020 Dual Path Interaction Network for Video Moment Localization
abstract
Video moment localization aims to localize a specific moment in a video by a natural language query. Previous works either use alignment information to find out the best-matching candidate (i.e., top-down approach) or use discrimination information to predict the temporal boundaries of the match (i.e., bottom-up approach). Little research has taken both the candidate-level alignment information and frame-level boundary information together and considers the complementarity between them. In this paper, we propose a unified top-down and bottom-up approach called Dual Path Interaction Network (DPIN), where the alignment and discrimination information are closely connected to jointly make the prediction. Our model includes a boundary prediction pathway encoding the frame-level representation and an alignment pathway extracting the candidate-level representation. The two branches of our network predict two complementary but different representations for moment localization. To enforce the consistency and strengthen the connection between the two representations, we propose a semantically conditioned interaction module. The experimental results on three popular benchmarks (i.e., TACoS, Charades-STA, and Activity-Caption) demonstrate that the proposed approach effectively localizes the relevant moment and outperforms the state-of-the-art approaches.
Hao Wang 0050, Zhengjun Zha, Xuejin Chen, Zhiwei Xiong, Jiebo Luo 0001
ACM Multimedia3
2020 DeepFacePencil: Creating Face Images from Freehand Sketches
abstract
In this paper, we explore the task of generating photo-realistic face images from hand-drawn sketches. Existing image-to-image translation methods require a large-scale dataset of paired sketches and images for supervision. They typically utilize synthesized edge maps of face images as training data. However, these synthesized edge maps strictly align with the edges of the corresponding face images, which limit their generalization ability to real hand-drawn sketches with vast stroke diversity. To address this problem, we propose DeepFacePencil, an effective tool that is able to generate photo-realistic face images from hand-drawn sketches, based on a novel dual generator image translation network during training. A novel spatial attention pooling (SAP) is designed to adaptively handle stroke distortions which are spatially varying to support various stroke styles and different level of details. We conduct extensive experiments and the results demonstrate the superiority of our model over existing methods on both image quality and model generalization to hand-drawn sketches.
Xuejin Chen, Binxin Yang, Zhihua Cheng, Zhengjun Zha
ACM Multimedia2
2020 Semantic Image Analogy with a Conditional Single-Image GAN
abstract
Recent image-specific Generative Adversarial Networks (GANs) provide a way to learn generative models from a single image instead of a large dataset. However, the semantic meaning of patches inside a single image is less explored. In this work, we first define the task of Semantic Image Analogy: given a source image and its segmentation map, along with another target segmentation map, synthesizing a new image that matches the appearance of the source image as well as the semantic layout of the target segmentation. To accomplish this task, we propose a novel method to model the patch-level correspondence between semantic layout and appearance of a single image by training a single-image GAN that takes semantic labels as conditional input. Once trained, a controllable redistribution of patches from the training image can be obtained by providing the expected semantic layout as spatial guidance. The proposed method contains three essential parts: 1) a self-supervised training framework, with a progressive data augmentation strategy and an alternating optimization procedure; 2) a semantic feature translation module that predicts transformation parameters in the image domain from the segmentation domain; and 3) a semantics-aware patch-wise loss that explicitly measures the similarity of two images in terms of patch distribution. Compared with existing solutions, our method generates much more realistic results given arbitrary semantic labels as conditional input.
Jiacheng Li 0004, Zhiwei Xiong, Dong Liu 0002, Xuejin Chen, Zhengjun Zha
ACM Multimedia4
2020 Joint Sketch-Attribute Learning for Fine-Grained Face Synthesis
Binxin Yang, Xuejin Chen, Richang Hong, Zhengjun Zha
MMM (1)2
2020 Neuronal Population Reconstruction From Ultra-Scale Optical Microscopy Images via Progressive Learning
abstract
Reconstruction of neuronal populations from ultra-scale optical microscopy (OM) images is essential to investigate neuronal circuits and brain mechanisms. The noises, low contrast, huge memory requirement, and high computational cost pose significant challenges in the neuronal population reconstruction. Recently, many studies have been conducted to extract neuron signals using deep neural networks (DNNs). However, training such DNNs usually relies on a huge amount of voxel-wise annotations in OM images, which are expensive in terms of both finance and labor. In this paper, we propose a novel framework for dense neuronal population reconstruction from ultra-scale images. To solve the problem of high cost in obtaining manual annotations for training DNNs, we propose a progressive learning scheme for neuronal population reconstruction (PLNPR) which does not require any manual annotations. Our PLNPR scheme consists of a traditional neuron tracing module and a deep segmentation network that mutually complement and progressively promote each other. To reconstruct dense neuronal populations from a terabyte-sized ultra-scale image, we introduce an automatic framework which adaptively traces neurons block by block and fuses fragmented neurites in overlapped regions continuously and smoothly. We build a dataset "VISoR-40" which consists of 40 large-scale OM image blocks from cortical regions of a mouse. Extensive experimental results on our VISoR-40 dataset and the public BigNeuron dataset demonstrate the effectiveness and superiority of our method on neuronal population reconstruction and single neuron reconstruction. Furthermore, we successfully apply our method to reconstruct dense neuronal populations from an ultra-scale mouse brain slice. The proposed adaptive block propagation and fusion strategies greatly improve the completeness of neurites in dense neuronal population reconstruction.
Jie Zhao 0020, Xuejin Chen, Zhiwei Xiong, Dong Liu 0002, Chaoyu Xie, Yueyi Zhang 0001, Zhengjun Zha, Guoqiang Bi, Feng Wu 0001
IEEE Trans. Medical Imaging2
2019 A Two-Stream Mutual Attention Network for Semi-Supervised Biomedical Segmentation with Noisy Labels
abstract
Learning-based methods suffer from a deficiency of clean annotations, especially in biomedical segmentation. Although many semi-supervised methods have been proposed to provide extra training data, automatically generated labels are usually too noisy to retrain models effectively. In this paper, we propose a Two-Stream Mutual Attention Network (TSMAN) that weakens the influence of back-propagated gradients caused by incorrect labels, thereby rendering the network robust to unclean data. The proposed TSMAN consists of two sub-networks that are connected by three types of attention models in different layers. The target of each attention model is to indicate potentially incorrect gradients in a certain layer for both sub-networks by analyzing their inferred features using the same input. In order to achieve this purpose, the attention models are designed based on the propagation analysis of noisy gradients at different layers. This allows the attention models to effectively discover incorrect labels and weaken their influence during parameter updating process. By exchanging multi-level features within two-stream architecture, the effects of noisy labels in each sub-network are reduced by decreasing the noisy gradients. Furthermore, a hierarchical distillation is developed to provide reliable pseudo labels for unlabelded data, which further boosts the performance of TSMAN. The experiments using both HVSMR 2016 and BRATS 2015 benchmarks demonstrate that our semi-supervised learning framework surpasses the state-of-the-art fully-supervised results.
Shaobo Min, Xuejin Chen, Zhengjun Zha, Feng Wu 0001, Yongdong Zhang 0001
AAAI2
2019 Accurate Segmentation of Synaptic Cleft with Contour Growing Concatenated with a Convnet
abstract
Synaptic cleft is an important area for neuroscientists to analyze the macromolecular complexes related to neurotransmitter transmission. However, the large amount of noise and low signal-to-noise ratio in raw electron micrographs make it challenging to extract this region automatically. In this paper, we propose a simple but effective framework to automatically extract accurate boundaries of synaptic cleft regions. Our approach concatenates a novel contour growing algorithm to a fully convolutional network (FCN), so that it takes both advantages of large receptive field of FCNs and fine-level localization of contour evolution. The contour growing algorithm is based on the flexible evolving tension and synchronous growing controlling to localize the opening contour of clef region. With consideration of both global localization and local segmentation, our approach is more robust to noisy electron micrographs and outperforms all existing single-model FCNs on accurate segmentation of synaptic clefts.
Shaobo Min, Xuejin Chen, Hongtao Xie 0001, Zhengjun Zha, Guoqiang Bi, Feng Wu 0001, Yongdong Zhang 0001
ICIP2
2019 Structure Generation and Guidance Network for Unsupervised Monocular Depth Estimation
abstract
Structure information is important to unsupervised depth learning from monocular videos. However, most existing methods focus on depth smoothing on planar regions, while other structure information, such as object shape and surface curvature, is ignored. In this work, we propose SGGN, a novel Structure Generation and Guidance Network to refine depth estimation under the guidance of extracted image structure. We introduce second-order Domain Transform filtering, which explores spatial depth variation by gradient propagation, to capture long-range dependence in the extracted structure for depth refinement. Then, several structure-aware constraints, as well as an attention mechanism, are applied to guide the training of SGGN, which leads to better depth estimation with structural guidance. Notably, our structure-aware constraints are designed in terms of different characteristics. Experiments on three benchmarks demonstrate the effectiveness of our structure-guided model and its state-of-the-art performance for unsupervised depth estimation.
Chaoqun Wang 0011, Xuejin Chen, Shaobo Min, Feng Wu 0001
ICME2
2019 Structure-Aware Residual Pyramid Network for Monocular Depth Estimation
abstract
Monocular depth estimation is an essential task for scene understanding. The underlying structure of objects and stuff in a complex scene is critical to recovering accurate and visually-pleasing depth maps. Global structure conveys scene layouts, while local structure reflects shape details. Recently developed approaches based on convolutional neural networks (CNNs) significantly improve the performance of depth estimation. However, few of them take into account multi-scale structures in complex scenes. In this paper, we propose a Structure-Aware Residual Pyramid Network (SARPN) to exploit multi-scale structures for accurate depth prediction. We propose a Residual Pyramid Decoder (RPD) which expresses global scene structure in upper levels to represent layouts, and local structure in lower levels to present shape details. At each level, we propose Residual Refinement Modules (RRM) that predict residual maps to progressively add finer structures on the coarser structure predicted at the upper level. In order to fully exploit multi-scale image features, an Adaptive Dense Feature Fusion (ADFF) module, which adaptively fuses effective features from all scales for inferring structures of each scale, is introduced. Experiment results on the challenging NYU-Depth v2 dataset demonstrate that our proposed approach achieves state-of-the-art performance in both qualitative and quantitative evaluation. The code is available at https://github.com/Xt-Chen/SARPN.
Xiaotian Chen, Xuejin Chen, Zhengjun Zha
IJCAI2
2019 Instance Segmentation from Volumetric Biomedical Images Without Voxel-Wise Labeling
Meng Dong, Dong Liu 0002, Zhiwei Xiong, Xuejin Chen, Yueyi Zhang 0001, Zhengjun Zha, Guoqiang Bi, Feng Wu 0001
MICCAI (2)4
2019 Progressive Learning for Neuronal Population Reconstruction from Optical Microscopy Images
Jie Zhao 0020, Xuejin Chen, Zhiwei Xiong, Dong Liu 0002, Yueyi Zhang 0001, Zhengjun Zha, Guoqiang Bi, Feng Wu 0001
MICCAI (1)2
2019 Fast and Accurate Electron Microscopy Image Registration with 3D Convolution
Shenglong Zhou 0002, Zhiwei Xiong, Chang Chen 0004, Xuejin Chen, Dong Liu 0002, Yueyi Zhang 0001, Zhengjun Zha, Feng Wu 0001
MICCAI (1)4
2019 Cross-Fiber Spatial-Temporal Co-enhanced Networks for Video Action Recognition
abstract
The 3D convolutional neural networks recently have been applied to explore spatial-temporal content for video action recognition. However, they either suffer from high computational cost by spatial-temporal feature extraction or ignore the correlation between appearance and motion. In this work, we propose a novel Cross-Fiber Spatial-Temporal Co-enhanced (CFST) architecture aiming to reduce the number of parameters tremendously while achieve accurate recognition of actions. We slice the complex 3D convolutional network into a group of lightweight fibers that run through the whole network. Crossing separated fibers, we introduce the Cross-Fiber Recalibration unit which shares extracted features from each fiber and measures the interaction between fibers to emphasize informative ones. Within each fiber, the Spatial-Temporal Co-enhanced unit is put forward to co-enhance the learning of spatial and temporal features, leading to more discriminative spatial-temporal representation. An end-to-end deep network, CFST-Net, is also presented based on the proposed CFST architecture for video action recognition. Extensive experimental results show that our CFST-Net significantly boosts the performance of existing convolution networks and achieves state-of-the-art accuracy on three challenging benchmarks, i.e., UCF-101, HMDB-51 and Kinetics-400, with much fewer parameters and FLOPs.
Haoze Wu 0003, Zhengjun Zha, Zhenzhong Chen 0001, Dong Liu 0002, Xuejin Chen
ACM Multimedia6
2019 LinesToFacePhoto: Face Photo Generation From Lines With Conditional Self-Attention Generative Adversarial Networks
abstract
In this paper, we explore the task of generating photo-realistic face images from lines. Previous methods based on conditional generative adversarial networks (cGANs) have shown their power to generate visually plausible images when a conditional image and an output image share well-aligned structures. However, these models fail to synthesize face images with a whole set of well-defined structures, e.g. eyes, noses, mouths, etc., especially when the conditional line map lacks one or several parts. To address this problem, we propose a conditional self-attention generative adversarial network (CSAGAN). We introduce a conditional self-attention mechanism to cGANs to capture long-range dependencies between different regions in faces. We also build a multi-scale discriminator. The large-scale discriminator enforces the completeness of global structures and the small-scale discriminator encourages fine details, thereby enhancing the realism of generated face images. We evaluate the proposed model on the CelebA-HD dataset by two perceptual user studies and three quantitative metrics. The experiment results demonstrate that our method generates high-quality facial images while preserving facial structures. Our results outperform state-of-the-art methods both quantitatively and qualitatively.
Xuejin Chen, Feng Wu 0001, Zhengjun Zha
ACM Multimedia2
2019 Ground-Aware Point Cloud Semantic Segmentation for Autonomous Driving
abstract
Semantic understanding of 3D scenes is essential for autonomous driving. Although a number of efforts have been devoted to semantic segmentation of dense point clouds, the great sparsity of 3D LiDAR data poses significant challenges in autonomous driving. In this paper, we work on the semantic segmentation problem of extremely sparse LiDAR point clouds with specific consideration of the ground as reference. In particular, we propose a ground-aware framework that well solves the ambiguity caused by data sparsity. We employ a multi-section plane fitting approach to roughly extract ground points to assist segmentation of objects on the ground. Based on the roughly extracted ground points, our approach implicitly integrates the ground information in a weakly-supervised manner and utilizes ground-aware features with a new ground-aware attention module. The proposed ground-aware attention module captures long-range dependence between ground and objects, which significantly facilitates the segmentation of small objects that only consist of a few points in extremely sparse point clouds. Extensive experiments on two large-scale LiDAR point cloud datasets for autonomous driving demonstrate that the proposed method achieves state-of-the-art performance both quantitatively and qualitatively.
Jianbo Jiao, Qingxiong Yang, Zhengjun Zha, Xuejin Chen
ACM Multimedia5
2019 Preventing self-intersection with cycle regularization in neural networks for mesh reconstruction from a single RGB image
Siyu Hu, Xuejin Chen
Comput. Aided Geom. Des.2
2019 Automatic Generation of Vivid LEGO Architectural Sculptures
abstract
Abstract Brick elements are very popular and have been widely used in many areas, such as toy design and architectural fields. Designing a vivid brick sculpture to represent a three‐dimensional (3D) model is a very challenging task, which requires professional skills and experience to convey unique visual characteristics. We introduce an automatic system to convert an architectural model into a LEGO sculpture while preserving the original model's shape features. Unlike previous legolization techniques that generate a LEGO sculpture exactly based on the input model's voxel representation, we extract the model's visual features, including repeating components, shape details and planarity. Then, we translate these visual features into the final LEGO sculpture by employing various brick types. We propose a deformation algorithm in order to resolve discrepancies between an input mesh's continuous 3D shape and the discrete positions of bricks in a LEGO sculpture. We evaluate our system on various architectural models and compare our method with previous voxelization‐based methods. The results demonstrate that our approach successfully conveys important visual features from digital models and generates vivid LEGO sculptures.
Jie Zhou 0010, Xuejin Chen, Ying-Qing Xu
Comput. Graph. Forum2
2019 Reconstructing piecewise planar scenes with multi-view regularization
abstract
Reconstruction of man-made scenes from multi-view images is an important problem in computer vision and computer graphics. Observing that manmade scenes are usually composed of planar surfaces, we encode plane shape prior in reconstructing man-made scenes. Recent approaches for single-view reconstruction employ multi-branch neural networks to simultaneously segment planes and recover 3D plane parameters. However, the scale of available annotated data heavily limits the generalizability and accuracy of these supervised methods. In this paper, we propose multi-view regularization to enhance the capability of piecewise planar reconstruction during the training phase, without demanding extra annotated data. Our multi-view regularization enables the consistency among multiple views by making the feature embedding more robust against view change and lighting variations. Thus, the neural network trained by multi-view regularization performs better on a wide range of views and lightings in the test phase. Based on more consistent prediction results, we merge the recovered models from multiple views to reconstruct scenes. Our approach achieves state-of-the-art reconstruction performance compared to previous approaches on the public ScanNet dataset.
Weijie Xi, Xuejin Chen
Comput. Vis. Media2
2019 Designing deployable 3D scissor structures with ball-and-socket joints
abstract
Abstract Scissor structures, which transform from a compact state to an expanded state, are widely used in various fields, ranging from architectural design to aerospace applications. We focus on a challenging problem that breaks through the restriction of planarity: to design scissor structures that expand from one given 3D shape to another 3D shape. To achieve this purpose, we propose a three‐step algorithm to construct a 3D scissor structure that realizes non‐uniform concentration between two different 3D curves. First, the input shapes are divided into scissor segments, which are composed of a sequence of planar scissor units based on the shape correspondence. Secondly, we compute the scissor unit geometry of each segment in a suggestive manner. Finally, the ball‐and‐socket joints with parameterized guide slits are integrated in order to connect the scissor segments, thereby completing the 3D deployment. Judging from a series of simulation and fabrication results, we demonstrate that our approach generates deployable structures for a wide range of 3D shape pairs.
Xuejin Chen, Haoming Jiang, Tingting Xuan, Lihan Huang, Ligang Liu 0001
Comput. Animat. Virtual Worlds1
2019 Dense 3D-Convolutional Neural Network for Person Re-Identification in Videos
abstract
Person re-identification aims at identifying a certain pedestrian across non-overlapping multi-camera networks in different time and places. Existing person re-identification approaches mainly focus on matching pedestrians on images; however, little attention has been paid to re-identify pedestrians in videos. Compared to images, video clips contain motion patterns of pedestrians, which is crucial to person re-identification. Moreover, consecutive video frames present pedestrian appearance with different body poses and from different viewpoints, providing valuable information toward addressing the challenge of pose variation, occlusion, and viewpoint change, and so on. In this article, we propose a Dense 3D-Convolutional Network (D3DNet) to jointly learn spatio-temporal and appearance representation for person re-identification in videos. The D3DNet consists of multiple three-dimensional (3D) dense blocks and transition layers. The 3D dense blocks enlarge the receptive fields of visual neurons in both spatial and temporal dimensions, leading to discriminative appearance representation as well as short-term and long-term motion patterns of pedestrians without the requirement of an additional motion estimation module. Moreover, we formulate a loss function consisting of an identification loss and a center loss to minimize intra-class variance and maximize inter-class variance simultaneously, toward addressing the challenge of large intra-class variance and small inter-class variance. Extensive experiments on two real-world video datasets of person identification, i.e., MARS and iLIDS-VID, have shown the effectiveness of the proposed approach.
Jiawei Liu 0001, Zhengjun Zha, Xuejin Chen, Zilei Wang, Yongdong Zhang 0001
ACM Trans. Multim. Comput. Commun. Appl.3
2019 Point sets joint registration and co-segmentation
Siyu Hu, Xuejin Chen, Xin Tong 0001
Vis. Comput.2
2018 Unsupervised Depth Estimation from Light Field Using a Convolutional Neural Network
abstract
This paper proposes an unsupervised CNN-based method for explicit depth estimation from light field, which learns an end-to-end mapping from a 4D light field to the corresponding disparity map without the supervision of groundtruth depth. Specifically, we design a combined loss function imposing both compliance and divergence constraints on the warped sub-aperture images to the central view, which guarantees our network to generate an accurate and robust disparity map. Furthermore, we find that increasing the number of referenced views in depth feature extraction and complementing missing information caused by warping greatly boost the performance of our network. Due to the difficulty of obtaining groundtruth depth of real-world scenes in practice, the proposed method is much more feasible than supervised learning. On the other hand, compared with traditional non-learning methods, the proposed method better exploits the correlations in the 4D light field and generates superior depth results both quantitatively and qualitatively. Also, the proposed method helps improve the performance of subsequent applications based on the estimated depth, e.g., spatial super-resolution of light field.
Jiayong Peng, Zhiwei Xiong, Dong Liu 0002, Xuejin Chen
3DV4
2018 Robust Room Layout Estimation from a Single Image with Geometric Hints
abstract
Estimation of room layout suffers from heavy occlusions and clutters in indoor scenes. In this paper, we propose a deep network that combines textures and geometric hints to predict the surface layout from a single image. Our method consists of three steps. First, depths and normals are extracted from the input RGB image. Secondly, a multi-channel FCN (MC-FCN) is presented to integrate these geometric hints for semantic surface segmentation. Thirdly, an optimization framework is adopted to refine the layout estimation. The results on two commonly used benchmark datasets demonstrate the robustness of our method on complex scenes.
Ruifeng Deng, Xuejin Chen
ICIP2
2018 3D Cnn-Based Soma Segmentation from Brain Images at Single-Neuron Resolution
abstract
Neuron segmentation is an important task for automatic analyses of brain images that are of huge volume. Previous methods for neuron segmentation rely on handcrafted image features, and have difficulty in coping with high-resolution, low signal-to-noise-ratio brain images. Convolutional neural network (CNN) has achieved remarkable success in natural image segmentation, but CNN requires accurately labeled data for training that are difficult to achieve on brain images of huge volume. In this paper, we present a weakly supervised learning strategy to deal with the inaccurate training data problem, and thus adopt 3D CNN to perform automatic soma segmentation from brain images. We test our method on our own collected mouse brain images that are of single-neuron resolution, and results show that 3D CNN-based method outperforms the traditional methods by a significant margin.
Meng Dong, Dong Liu 0002, Zhiwei Xiong, Chaoyu Yang, Xuejin Chen, Zhengjun Zha, Guoqiang Bi, Feng Wu 0001
ICIP5
2018 Sketchpointnet: A Compact Network for Robust Sketch Recognition
abstract
Sketch recognition is a challenging image processing task. In this paper, we propose a novel point-based network with a compact architecture, named SketchPointNet, for robust sketch recognition. Sketch features are hierarchically learned from three mini PointNets, by successively sampling and grouping 2D points in a bottom-up fashion. SketchPointNet exploits both temporal and spatial context in strokes during point sampling and grouping. By directly consuming the sparse points, SketchPointN et is very compact and efficient. Compared with state-of-the-art techniques, SketchPointNet achieves comparable performance on the challenging TU-Berlin dataset while it significantly reduces the network size.
Xuejin Chen, Zhengjun Zha
ICIP2
2018 LA-Net: Layout-Aware Dense Network for Monocular Depth Estimation
abstract
Depth estimation from monocular images is an ill-posed and inherently ambiguous problem. Recently, deep learning technique has been applied for monocular depth estimation seeking data-driven solutions. However, most existing methods focus on pursuing the minimization of average depth regression error at pixel level and neglect to encode the global layout of scene, resulting in layout-inconsistent depth map. This paper proposes a novel Layout-Aware Convolutional Neural Network (LA-Net) for accurate monocular depth estimation by simultaneously perceiving scene layout and local depth details. Specifically, a Spatial Layout Network (SL-Net) is proposed to learn a layout map representing the depth ordering between local patches. A Layout-Aware Depth Estimation Network (LDE-Net) is proposed to estimate pixel-level depth details using multi-scale layout maps as structural guidance, leading to layout-consistent depth map. A dense network module is used as the base network to learn effective visual details resorting to dense feed-forward connections. Moreover, we formulate an order-sensitive softmax loss to well constrain the ill-posed depth inferring problem. Extensive experiments on both indoor scene (NYUD-v2) and outdoor scene (Make3D) datasets have demonstrated that the proposed LA-Net outperforms the state-of-the-art methods and leads to faithful 3D projections.
Kecheng Zheng, Zhengjun Zha, Yang Cao 0010, Xuejin Chen, Feng Wu 0001
ACM Multimedia4
2018 Real-Time Object Tracking with Motion Information
abstract
Motion is a vital information for object tracking. However, most existing methods, including the classic Siamese FC network [1], only consider the object appearance, and ignore the vital motion feature. In this paper, we design a dual-network object tracker, which is called DOT for short, to effectively combine the appearance and motion information. Our method employs two branches, S-net and M-net, to exploit the appearance and motion information respectively. Moreover, an attention fusion module is also introduced to effectively integrate these two aspects. The experiments carried out on OTB-2013 demonstrate the improvement on object tracking by the integration of motion information with our dual-network and attention fusion.
Chaoqun Wang 0011, Xiaoyan Sun 0001, Xuejin Chen, Wenjun Zeng 0001
VCIP3
2018 Folding cartons: Interactive manipulation of cartons from 2D layouts
Shuang Shan, Yiting Ma, Chengcheng Tang, Xuejin Chen
Comput. Aided Geom. Des.4
2018 Convertible furniture design
Jie Zhou 0010, Xuejin Chen
Comput. Graph.2
2017 An interactive system for efficient 3D furniture arrangement
abstract
We present an interactive example-based system for non-expert users to generate 3D indoor scenes intuitively. From a set of examples of an interior scene, we extract furniture layout constraints including pairwise and group relationships, ergonomic factors and user habits. Instead of inserting a single furniture object to a visual pleasing and functional position independently, we take advantages of manipulating a group of functionality-related furniture objects for smart editing. To deal with the arrangement problem which requires jointly optimizing a variety of functional and visual criteria, we introduce an efficient two-step pipeline consisting of rough arrangement and layout refinement. In the rough arrangement step, furniture objects are roughly placed according to pairwise relationships for accelerating the further optimization process. In the layout refinement step, we optimize the rough furniture layout to a semantical and functional layout considering spatial relationships, and ergonomic factors and avoiding collisions. User interactions are also allowed to move a specific object or a structural group while our system automatically refines the entire layout. The experimental results demonstrate that our system for scene generation measurably increases the visual quality and time efficiency compared with the completely manual scene editing or modeling.
Meng Yan 0007, Xuejin Chen, Jie Zhou 0010
CGI2
2017 An interactive system for low-poly illustration generation from images using adaptive thinning
abstract
Low-poly style illustrations, which have 3D abstract appearance, have become a popular stylish recently. Most previous methods require special knowledges in 3D modeling and need tedious interactions. We present an interactive system for non-expert users to easily manipulate the low-poly style illustration. Our system consists of two parts: vertex sampling and mesh rendering. In the vertex sampling stage, we extract a set of candidate points from the image and rank them according to their importance of structure preserving using adaptive thinning. Based on the pre-ranked point list, the user can select an arbitrary number of vertices for the triangle mesh construction. In the mesh rendering stage, we optimize triangle colors to create stereo-looking low-polys. We also provide three tools for exible modication of vertex numbers, color contrast, and local region emphasis. The experiment results demonstrate that our system outperforms state-of-the-art method via simple user interactions.
Yiting Ma, Xuejin Chen
ICME2
2016 HF-FCN: Hierarchically Fused Fully Convolutional Network for Robust Building Extraction
Tongchun Zuo, Juntao Feng, Xuejin Chen
ACCV (1)3
2016 Efficient structure-preserving superpixel segmentation based on minimum spanning tree
abstract
We propose a novel superpixel algorithm based on Minimum Spanning Tree (MST), to generate superpixels efficiently while strictly adhere to object boundaries. The MST, which built by gradually removing strong edges of the image graph extracted from the image, is more sensitive to image local structures. Therefore, an efficient hierarchical clustering strategy is basically employed in our algorithm to segment the input image into superpixels based on the tree distance. To gradually merge the image pixels and remove texture noises, a multi-layer scheme with different resolutions of superpixels is proposed. In each layer, the graph is constructed from the lower layer and segmented into superpixels in a linear complexity with the node number in the graph. Because the node number in each layer is exponentially reduced, the computational time of our method mainly concentrates on the first few layers, which is linear with the number of image pixels. The experimental results conducted on the Berkeley Segmentation Dataset demonstrate that our method outperforms state-of-the-art methods both in terms of structure preservation and computational efficiency.
Xuejin Chen
ICME2
2016 Robust Sketch-Based Image Retrieval by Saliency Detection
Xuejin Chen
MMM (1)2
2016 Designing Planar Deployable Objects via Scissor Structures
abstract
Scissor structure is used to generate deployable objects for space-saving in a variety of applications, from architecture to aerospace science. While deployment from a small, regular shape to a larger one is easy to design, we focus on a more challenging task: designing a planar scissor structure that deploys from a given source shape into a specific target shape. We propose a two-step constructive method to generate a scissor structure from a high-dimensional parameter space. Topology construction of the scissor structure is first performed to approximate the two given shapes, as well as to guarantee the deployment. Then the geometry of the scissor structure is optimized in order to minimize the connection deflections and maximize the shape approximation. With the optimized parameters, the deployment can be simulated by controlling an anchor scissor unit. Physical deployable objects are fabricated according to the designed scissor structures by using 3D printing or manual assembly. We show a number of results for different shapes to demonstrate that even with fabrication errors, our designed structures can deform fluently between the source and target shapes.
Shiwei Wang 0004, Xuejin Chen, Luo Jiang, Jie Zhou 0010, Ligang Liu 0001
IEEE Trans. Vis. Comput. Graph.3
2015 An efficient volumetric method for non-rigid registration
Xuejin Chen, Takaaki Shiratori, Xin Tong 0001, Ligang Liu 0001
Graph. Model.2
2013 LSGP: Line-SIFT Geometric Pattern for wide-baseline image matching
abstract
In this paper, a novel descriptor - Line-SIFT Geometric Pattern (LSGP) is proposed for wide-baseline image matching. In this descriptor, the geometry relationship between a line segment and its neighboring SIFT features is used to describe the line segments. By measuring the similarities between line segments, we could find line correspondences between multiple images which embed more intuitive information for further application such as 3D reconstruction. The experiment results have proved that our LSGP descriptor has good performance under various conditions.
Xuejin Chen, Zhefu Tu
ISCAS2
2010 Printed Patterns for Enhanced Shape Perception of Papercraft Models
abstract
Abstract Papercraft models can serve as inexpensive prototypes in shape design applications. However, in making the models some geometric detail is necessarily lost, and artificial creases may be visible, thereby limiting the utility of these models. To compensate for these practical limitations, we introduce the use of printed patterns on papercraft models to enhance the perception of the shape they are intended to represent. We propose pattern generation schemes that modulate the sizes, directions, and densities of glyphs of patterns based on geometric attributes. We present a psychophysical experiment designed to explore the effect that printed patterns have on the perception of the papercraft model shapes. We find that models with printed patterns are perceived to represent the intended shape more accurately, and, further, that the type of printed pattern has an impact on the perceived shape.
Su Xue, Xuejin Chen, Julie Dorsey, Holly E. Rushmeier
Comput. Graph. Forum2
2008 Sketching reality: Realistic interpretation of architectural designs
abstract
In this article, we introduce sketching reality , the process of converting a freehand sketch into a realistic-looking model. We apply this concept to architectural designs. As the sketch is being drawn, our system periodically interprets its 2.5D-geometry by identifying new junctions, edges, and faces, and then analyzing the extracted topology. The user can add detailed geometry and textures through sketches as well. This is possible through the use of databases that match partial sketches to models of detailed geometry and textures. The final product is a realistic texture-mapped 2.5D-model of the building. We show a variety of buildings that have been created using this system.
Xuejin Chen, Sing Bing Kang, Ying-Qing Xu, Julie Dorsey, Harry Shum
ACM Trans. Graph.1
2008 Sketch-based tree modeling using Markov random field
abstract
In this paper, we describe a new system for converting a user's freehand sketch of a tree into a full 3D model that is both complex and realistic-looking. Our system does this by probabilistic optimization based on parameters obtained from a database of tree models. The best matching model is selected by comparing its 2D projections with the sketch. Branch interaction is modeled by a Markov random field, subject to the constraint of 3D projection to sketch. Our system then uses the notion of self-similarity to add new branches before finally populating all branches with leaves of the user's choice. We show a variety of natural-looking tree models generated from freehand sketches with only a few strokes.
Xuejin Chen, Boris Neubert, Ying-Qing Xu, Oliver Deussen, Sing Bing Kang
ACM Trans. Graph.1
2005 Retargeting vector animation for small displays
abstract
We present a method that preserves the recognizability of key object interactions in a vector animation. The method allows an artist to author an animation once, and then output it to any display device. We specifically target mobile devices with small screen sizes. In order to adapt an animation, the author specifies an importance value for objects in the animation. The algorithm then identifies and categorizes the vector graphics objects that comprise the animation, leveraging the implicit relationship between extensible Markup Language (XML) and scalable vector graphics (SVG). Based on importance, the animation can then be automatically retargeted for any display using artistically motivated resizing and grouping algorithms that budget size and spatial detail for each object.
Vidya Setlur, Ying-Qing Xu, Xuejin Chen, Bruce Gooch
MUM3