Xiaokang Chen

dblp:163/6632 · DBLP profile ↗
← Back
31ranked-venue papers
7as first author
24since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 20 · 5 first-author · 15 since 2021Artificial intelligence and machine learning · 19 · 5 first-author · 17 since 2021Systems, architecture and hardware · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation
abstract
We present JanusFlow, a powerful framework that unifies image understanding and generation in a single model. JanusFlow introduces a minimalist architecture that integrates autoregressive language models with rectified flow, a state-of-the-art method in generative modeling. Our key finding demonstrates that rectified flow can be straightforwardly trained within the large language model framework, eliminating the need for complex architectural modifications. To further improve the performance of our unified model, we adopt two key strategies: (i) decoupling the understanding and generation encoders, and (ii) aligning their representations during unified training. Extensive experiments show that JanusFlow achieves comparable or superior performance to specialized models in their respective domains, while significantly outperforming existing unified approaches. This work represents a step toward more efficient and versatile vision-language models.
Yiyang Ma, Xingchao Liu, Xiaokang Chen, Wen Liu 0011, Chengyue Wu, Zhiyu Wu, Zizheng Pan, Zhenda Xie, Xingkai Yu, Liang Zhao 0026, Jiaying Liu 0001, Chong Ruan
CVPR3
2025 Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
abstract
We introduce Janus, an autoregressive framework that unifies multimodal understanding and generation. Prior research often relies on a single visual encoder for both tasks, such as Chameleon. However, due to the differing levels of information granularity required by multimodal understanding and generation, this approach can lead to suboptimal performance, particularly in multimodal understanding. To address this issue, we decouple visual encoding into separate pathways, while still leveraging a single, unified transformer architecture for processing. The decoupling not only alleviates the conflict between the visual encoder’s roles in understanding and generation, but also enhances the framework’s flexibility. For instance, both the multi-modal understanding and generation components can independently select their most suitable encoding methods. Experiments show that Janus surpasses previous unified model and matches or exceeds the performance of task-specific models. The simplicity, high flexibility, and effectiveness of Janus make it a strong candidate for next-generation unified multimodal models. The code will be made available.
Chengyue Wu, Xiaokang Chen, Zhiyu Wu, Yiyang Ma, Xingchao Liu, Zizheng Pan, Wen Liu 0011, Zhenda Xie, Xingkai Yu, Chong Ruan, Ping Luo 0002
CVPR2
2025 ConstFS: Controlled and Stable Face Stylization with High Identity-Preserved
abstract
Facial style transfer often faces challenges, including content loss, style degradation, and difficulty balancing style and content consistency. Existing methods, including those based on StyleGAN and Stable Diffusion, encounter issues such as facial feature distortion, style degradation, and color leakage. We propose ConstFS (Controlled and Stable Face Stylization) to overcome these challenges, which enhances traditional diffusion models by incorporating advanced style extraction, identity preservation, and dynamic control mechanisms. The proposed ConstFS achieves superior image control, enabling it to adaptively match the style image. The experimental results demonstrate that our method significantly enhances style integrity and content consistency, outperforming existing techniques in both qualitative and quantitative experiments.
Feichi Chen, Xiaokang Chen, Xuanhao Lou, Naye Ji
VINCI2
2025 Real-Time Neural Radiance Talking Portrait Synthesis via Audio-Spatial Decomposition
Jiaxiang Tang, Kaisiyuan Wang, Hang Zhou 0009, Xiaokang Chen, Dongliang He, Tianshu Hu, Jingtuo Liu, Ziwei Liu 0002, Jingdong Wang 0001
Int. J. Comput. Vis.4
2025 SST: Self-training with self-adaptive thresholding for semi-supervised learning
abstract
Neural networks have demonstrated exceptional performance in supervised learning, benefiting from abundant high-quality annotated data. However, obtaining such data in real-world scenarios is costly and labor-intensive. Semi-supervised learning (SSL) offers a solution to this problem by utilizing a small amount of labeled data along with a large volume of unlabeled data . Recent studies, such as Semi-ViT and Noisy Student, which employ consistency regularization or pseudo-labeling, have demonstrated significant achievements. However, they still face challenges, particularly in accurately selecting sufficient high-quality pseudo-labels due to their reliance on fixed thresholds. Recent methods such as FlexMatch and FreeMatch have introduced flexible or self-adaptive thresholding techniques, greatly advancing SSL research. Nonetheless, their process of updating thresholds at each iteration is deemed time-consuming, computationally intensive, and potentially unnecessary. To address these issues, we propose Self-training with Self-adaptive Thresholding (SST), a novel, effective, and efficient SSL framework. SST integrates with both supervised (Super-SST) and semi-supervised (Semi-SST) learning. SST introduces an innovative Self-Adaptive Thresholding (SAT) mechanism that adaptively adjusts class-specific thresholds based on the model’s learning progress. SAT ensures the selection of high-quality pseudo-labeled data, mitigating the risks of inaccurate pseudo-labels and confirmation bias (where models reinforce their own mistakes during training). Specifically, SAT prevents the model from prematurely incorporating low-confidence pseudo-labels, reducing error reinforcement and enhancing model performance. Extensive experiments demonstrate that SST achieves state-of-the-art performance with remarkable efficiency, generalization, and scalability across various architectures and datasets. Notably, Semi-SST-ViT-Huge achieves the best results on competitive ImageNet-1K SSL benchmarks (no external data), with 80.7%/84.9% Top-1 accuracy using only 1%/10% labeled data. Compared to the fully-supervised DeiT-III-ViT-Huge, which achieves 84.8% Top-1 accuracy using 100% labeled data, our method demonstrates superior performance using only 10% labeled data. This indicates a tenfold reduction in human annotation costs, significantly narrowing the performance disparity between semi-supervised and fully-supervised methods. These advancements pave the way for further innovations in SSL and practical applications where obtaining labeled data is either challenging or costly.
Heyan Huang, Xiaokang Chen, Rui Wang 0043
Inf. Process. Manag.4
2024 LGM: Large Multi-view Gaussian Model for High-Resolution 3D Content Creation
Jiaxiang Tang, Zhaoxi Chen 0009, Xiaokang Chen, Tengfei Wang 0002, Ziwei Liu 0002
ECCV (4)3
2024 Improving Long Text Understanding with Knowledge Distilled from Summarization Model
abstract
Long text understanding is important yet challenging for natural language processing. A long article or document usually contains many redundant words that are not pertinent to its gist and sometimes can be regarded as noise. With recent advances of abstractive summarization, we propose our Gist Detector to leverage the gist detection ability of a summarization model and integrate the extracted gist into downstream models to enhance their long text understanding ability. Specifically, Gist Detector first learns the gist detection knowledge distilled from a summarization model, and then produces gist-aware representations to augment downstream models. We evaluate our method on three different tasks: long document classification, distantly supervised open-domain question answering, and non-parallel text style transfer. The experimental results show that our method can significantly improve the performance of baseline models on all tasks.
Yazheng Yang, Xiaokang Chen
ICASSP3
2024 The Devil is in the Neurons: Interpreting and Mitigating Social Biases in Language Models
abstract
Pre-trained Language models (PLMs) have been acknowledged to contain harmful information, such as social biases, which may cause negative social impacts or even bring catastrophic results in application. Previous works on this problem mainly focused on using black-box methods such as probing to detect and quantify social biases in PLMs by observing model outputs. As a result, previous debiasing methods mainly finetune or even pre-train PLMs on newly constructed anti-stereotypical datasets, which are high-cost. In this work, we try to unveil the mystery of social bias inside language models by introducing the concept of {\sc Social Bias Neurons}. Specifically, we propose {\sc Integrated Gap Gradients (IG$^2$)} to accurately pinpoint units (i.e., neurons) in a language model that can be attributed to undesirable behavior, such as social bias. By formalizing undesirable behavior as a distributional property of language, we employ sentiment-bearing prompts to elicit classes of sensitive words (demographics) correlated with such sentiments. Our IG$^2$ thus attributes the uneven distribution for different demographics to specific Social Bias Neurons, which track the trail of unwanted behavior inside PLM units to achieve interoperability. Moreover, derived from our interpretable technique, {\sc Bias Neuron Suppression (BNS)} is further proposed to mitigate social biases. By studying BERT, RoBERTa, and their attributable differences from debiased FairBERTa, IG$^2$ allows us to locate and suppress identified neurons, and further mitigate undesired behaviors. As measured by prior metrics from StereoSet, our model achieves a higher degree of fairness while maintaining language modeling ability with low cost\footnote{This work contains examples that potentially implicate stereotypes, associations, and other harms that could be offensive to individuals in certain social groups.}.
Yan Liu 0002, Xiaokang Chen, Daoguang Zan, Min-Yen Kan, Tsung-Yi Ho
ICLR3
2024 Open-set Hierarchical Semantic Segmentation for 3D Scene
abstract
The Segment-Anything Model (SAM) shows exceptional zero-shot capabilities for 2D images. Developing a similar model for 3D, however, is challenging due to limited datasets. In this paper, we introduce a zero-shot algorithm to segment a 3D scene into elements at various levels of detail, and further organize the results in a hierarchical tree structure. We propose a tree quality metric to evaluate the algorithm’s performance. Notably, our algorithm eliminates the need for 3D annotations. It uses robust 2D models to generate a 2D segmentation tree for each rendered image. Then, using graph neural networks, it aggregates these 2D trees to form a unified 3D segmentation tree. Extensive experiments on the PartNet dataset and complex 3D scenes validate the algorithm’s effectiveness. We release the source code at https://github.com/dnvtmf/OTS.
Diwen Wan, Jiaxiang Tang, Jingbo Wang 0003, Xiaokang Chen, Lingyun Gan
ICME4
2024 D3ETR: Decoder Distillation for Detection Transformer
Xiaokang Chen, Jiaxiang Tang
IJCAI1
2024 Context Autoencoder for Self-supervised Representation Learning
Xiaokang Chen, Mingyu Ding, Shentong Mo, Shumin Han, Ping Luo 0002, Jingdong Wang 0001
Int. J. Comput. Vis.1
2023 Uncovering and Categorizing Social Biases in Text-to-SQL
abstract
Content Warning: This work contains examples that potentially implicate stereotypes, associations, and other harms that could be offensive to individuals in certain social groups.Large pre-trained language models are acknowledged to carry social biases towards different demographics, which can further amplify existing stereotypes in our society and cause even more harm.Text-to-SQL is an important task, models of which are mainly adopted by authoritative institutions, where unfair decisions may lead to catastrophic consequences.However, existing Text-to-SQL models are trained on clean, neutral datasets, such as Spider and WikiSQL.This, to some extent, cover up social bias in models under ideal conditions, which nevertheless may emerge in real application scenarios.In this work, we aim to uncover and categorize social biases in Text-to-SQL models.We summarize the categories of social biases that may occur in structured data for Text-to-SQL models.We build test benchmarks and reveal that models with similar task accuracy can contain social biases at very different rates.We show how to take advantage of our methodology to uncover and assess social biases in the downstream Text-to-SQL task 1 .
Yan Liu 0002, Yan Gao 0002, Xiaokang Chen, Elliott Ash, Jian-Guang Lou
ACL (1)4
2023 Parallel Sentence-Level Explanation Generation for Real-World Low-Resource Scenarios
abstract
In order to reveal the rationale behind model predictions, many works have exploited providing explanations in various forms. Recently, to further guarantee readability, more and more works turn to generate sentence-level human language explanations. However, current works pursuing sentence- level explanations rely heavily on annotated training data, which limits the development of interpretability to only a few tasks. As far as we know, this paper is the first to explore this problem smoothly from weak-supervised learning to unsupervised learning. Besides, we also notice the high latency of autoregressive sentence-level explanation generation, which leads to asynchronous interpretability after prediction. Therefore, we propose a non-autoregressive interpretable model to facilitate parallel explanation generation and simultaneous prediction. Through extensive experiments on Natural Language Inference task and Spouse Prediction task, we find that users are able to train classifiers with comparable performance 10 − 15× faster with parallel explanation generation using only a few or no annotated training data.
Xiaokang Chen, Qi Dai 0001
ICASSP2
2023 Group DETR: Fast DETR Training with Group-Wise One-to-Many Assignment
abstract
Detection transformer (DETR) relies on one-to-one assignment, assigning one ground-truth object to one prediction, for end-to-end detection without NMS post-processing. It is known that one-to-many assignment, assigning one ground-truth object to multiple predictions, succeeds in detection methods such as Faster R-CNN and FCOS. While the naive one-to-many assignment does not work for DETR, and it remains challenging to apply one-to-many assignment for DETR training. In this paper, we introduce Group DETR, a simple yet efficient DETR training approach that introduces a group-wise way for one-to-many assignment. This approach involves using multiple groups of object queries, conducting one-to-one assignment within each group, and performing decoder self-attention separately. It resembles data augmentation with automatically-learned object query augmentation. It is also equivalent to simultaneously training parameter-sharing networks of the same architecture, introducing more supervision and thus improving DETR training. The inference process is the same as DETR trained normally and only needs one group of queries without any architecture modification. Group DETR is versatile and is applicable to various DETR variants. The experiments show that Group DETR signifi-cantly speeds up the training convergence and improves the performance of various DETR-based models. Code will be available at https://github.com/Atten4Vis/GroupDETR.
Qiang Chen 0007, Xiaokang Chen, Jian Wang 0066, Shan Zhang 0002, Haocheng Feng, Junyu Han, Errui Ding, Jingdong Wang 0001
ICCV2
2023 Delicate Textured Mesh Recovery from NeRF via Adaptive Surface Refinement
abstract
Neural Radiance Fields (NeRF) have constituted a remarkable breakthrough in image-based 3D reconstruction. However, their implicit volumetric representations differ significantly from the widely-adopted polygonal meshes and lack support from common 3D software and hardware, making their rendering and manipulation inefficient. To overcome this limitation, we present a novel framework that generates textured surface meshes from images. Our approach begins by efficiently initializing the geometry and view-dependency decomposed appearance with a NeRF. Subsequently, a coarse mesh is extracted, and an iterative surface refinement algorithm is developed to adaptively adjust both vertex positions and face density based on reprojected rendering errors. We jointly refine the appearance with geometry and bake it into texture images for real-time rendering. Extensive experiments demonstrate that our method achieves superior mesh quality and competitive rendering quality.
Jiaxiang Tang, Hang Zhou 0009, Xiaokang Chen, Tianshu Hu, Errui Ding, Jingdong Wang 0001
ICCV3
2023 Uncovering and Quantifying Social Biases in Code Generation
abstract
With the popularity of automatic code generation tools, such as Copilot, the study of the potential hazards of these tools is gaining importance. In this work, we explore the social bias problem in pre-trained code generation models. We propose a new paradigm to construct code prompts and successfully uncover social biases in code generation models. To quantify the severity of social biases in generated code, we develop a dataset along with three metrics to evaluate the overall social bias and fine-grained unfairness across different demographics. Experimental results on three pre-trained code generation models (Codex, InCoder, and CodeGen) with varying sizes, reveal severe social biases. Moreover, we conduct analysis to provide useful insights for further choice of code generation models with low social bias.
Yan Liu 0002, Xiaokang Chen, Yan Gao 0002, Fengji Zhang, Daoguang Zan, Jian-Guang Lou, Tsung-Yi Ho
NeurIPS2
2023 VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks
abstract
Large language models (LLMs) have notably accelerated progress towards artificial general intelligence (AGI), with their impressive zero-shot capacity for user-tailored tasks, endowing them with immense potential across a range of applications. However, in the field of computer vision, despite the availability of numerous powerful vision foundation models (VFMs), they are still restricted to tasks in a pre-defined form, struggling to match the open-ended task capabilities of LLMs. In this work, we present an LLM-based framework for vision-centric tasks, termed VisionLLM. This framework provides a unified perspective for vision and language tasks by treating images as a foreign language and aligning vision-centric tasks with language tasks that can be flexibly defined and managed using language instructions. An LLM-based decoder can then make appropriate predictions based on these instructions for open-ended tasks. Extensive experiments show that the proposed VisionLLM can achieve different levels of task customization through language instructions, from fine-grained object-level to coarse-grained task-level customization, all with good results. It's noteworthy that, with a generalist LLM-based framework, our model can achieve over 60% mAP on COCO, on par with detection-specific models. We hope this model can set a new baseline for generalist vision and language models. The code shall be released.
Wenhai Wang, Zhe Chen 0017, Xiaokang Chen, Jiannan Wu, Xizhou Zhu, Ping Luo 0002, Tong Lu 0002, Jie Zhou 0001, Yu Qiao 0001, Jifeng Dai
NeurIPS3
2022 Not All Voxels Are Equal: Semantic Scene Completion from the Point-Voxel Perspective
abstract
We revisit Semantic Scene Completion (SSC), a useful task to predict the semantic and occupancy representation of 3D scenes, in this paper. A number of methods for this task are always based on voxelized scene representations. Although voxel representations keep local structures of the scene, these methods suffer from heavy computation redundancy due to the existence of visible empty voxels when the network goes deeper. To address this dilemma, we propose our novel point-voxel aggregation network for this task. We first transfer the voxelized scenes to point clouds by removing these visible empty voxels and adopt a deep point stream to capture semantic information from the scene efficiently. Meanwhile, a light-weight voxel stream containing only two 3D convolution layers preserves local structures of the voxelized scenes. Furthermore, we design an anisotropic voxel aggregation operator to fuse the structure details from the voxel stream into the point stream, and a semantic-aware propagation module to enhance the up-sampling process in the point stream by semantic labels. We demonstrate that our model surpasses state-of-the-arts on two benchmarks by a large margin, with only the depth images as input.
Jiaxiang Tang, Xiaokang Chen, Jingbo Wang 0003
AAAI2
2022 Point Scene Understanding via Disentangled Instance Mesh Reconstruction
Jiaxiang Tang, Xiaokang Chen, Jingbo Wang 0003
ECCV (32)2
2022 MaskGroup: Hierarchical Point Grouping and Masking for 3D Instance Segmentation
abstract
This paper studies the 3D instance segmentation problem, which has a variety of real-world applications such as robotics and augmented reality. Since the surroundings of 3D objects are of high complexity, the separating of different objects is very difficult. To address this challenging problem, we propose a novel framework to group and refine the 3D instances. In practice, we first learn an offset vector for each point and shift it to its predicted instance center. To better group these points, we propose a Hierarchical Point Grouping algorithm to merge the centrally aggregated points progressively. All points are grouped into small clusters, which further gradually undergo another clustering procedure to merge into larger groups. These multi-scale groups are exploited for instance prediction, which is beneficial for predicting instances with different scales. In addition, a novel MaskScoreNet is developed to produce binary point masks of these groups for further refining the segmentation results. Extensive experiments conducted on the ScanNetV2 and S3DIS benchmarks demonstrate the effectiveness of the proposed method. For instance, our MaskGroup achieves a 66.4% mAP with the 0.5 IoU threshold on the ScanNetV2 test set, which is 1.9% higher than the state-of-the-art method.
Xinghao Chen 0001, Xiaokang Chen, Yunhe Wang 0001
ICME3
2022 Compressible-composable NeRF via Rank-residual Decomposition
abstract
Neural Radiance Field (NeRF) has emerged as a compelling method to represent 3D objects and scenes for photo-realistic rendering. However, its implicit representation causes difficulty in manipulating the models like the explicit mesh representation.Several recent advances in NeRF manipulation are usually restricted by a shared renderer network, or suffer from large model size. To circumvent the hurdle, in this paper, we present a neural field representation that enables efficient and convenient manipulation of models.To achieve this goal, we learn a hybrid tensor rank decomposition of the scene without neural networks. Motivated by the low-rank approximation property of the SVD algorithm, we propose a rank-residual learning strategy to encourage the preservation of primary information in lower ranks. The model size can then be dynamically adjusted by rank truncation to control the levels of detail, achieving near-optimal compression without extra optimization.Furthermore, different models can be arbitrarily transformed and composed into one scene by concatenating along the rank dimension.The growth of storage cost can also be mitigated by compressing the unimportant objects in the composed scene. We demonstrate that our method is able to achieve comparable rendering quality to state-of-the-art methods, while enabling extra capability of compression and composition.Code is available at https://github.com/ashawkey/CCNeRF.
Jiaxiang Tang, Xiaokang Chen, Jingbo Wang 0003
NeurIPS2
2021 Semi-Supervised Semantic Segmentation With Cross Pseudo Supervision
abstract
In this paper, we study the semi-supervised semantic segmentation problem via exploring both labeled data and extra unlabeled data. We propose a novel consistency regularization approach, called cross pseudo supervision (CPS). Our approach imposes the consistency on two segmentation networks perturbed with different initialization for the same input image. The pseudo one-hot label map, output from one perturbed segmentation network, is used to supervise the other segmentation network with the standard cross-entropy loss, and vice versa. The CPS consistency has two roles: encourage high similarity between the predictions of two perturbed networks for the same input image, and expand training data by using the unlabeled data with pseudo labels. Experiment results show that our approach achieves the state-of-the-art semi-supervised segmentation performance on Cityscapes and PASCAL VOC 2012.
Xiaokang Chen, Yuhui Yuan, Jingdong Wang 0001
CVPR1
2021 Conditional DETR for Fast Training Convergence
abstract
The recently-developed DETR approach applies the transformer encoder and decoder architecture to object detection and achieves promising performance. In this paper, we handle the critical issue, slow training convergence, and present a conditional cross-attention mechanism for fast DETR training. Our approach is motivated by that the cross-attention in DETR relies highly on the content embeddings for localizing the four extremities and predicting the box, which increases the need for high-quality content embeddings and thus the training difficulty.Our approach, named conditional DETR, learns a conditional spatial query from the decoder embedding for decoder multi-head cross-attention. The benefit is that through the conditional spatial query, each cross-attention head is able to attend to a band containing a distinct region, e.g., one object extremity or a region inside the object box. This narrows down the spatial range for localizing the distinct regions for object classification and box regression, thus relaxing the dependence on the content embeddings and easing the training. Empirical results show that conditional DETR converges 6.7× faster for the backbones R50 and R101 and 10× faster for stronger backbones DC5-R50 and DC5-R101. Code is available at https://github.com/Atten4Vis/ConditionalDETR.
Depu Meng, Xiaokang Chen, Zejia Fan, Houqiang Li, Yuhui Yuan, Jingdong Wang 0001
ICCV2
2021 Joint Implicit Image Function for Guided Depth Super-Resolution
abstract
Guided depth super-resolution is a practical task where a low-resolution and noisy input depth map is restored to a high-resolution version, with the help of a high-resolution RGB guide image. Existing methods usually view this task as a generalized guided filtering problem that relies on designing explicit filters and objective functions, or a dense regression problem that directly predicts the target image via deep neural networks. These methods suffer from either model capability or interpretability. Inspired by the recent progress in implicit neural representation, we propose to formulate the guided super-resolution as a neural implicit image interpolation problem, where we take the form of a general image interpolation but use a novel Joint Implicit Image Function (JIIF) representation to learn both the interpolation weights and values. JIIF represents the target image domain with spatially distributed local latent codes extracted from the input image and the guide image, and uses a graph attention mechanism to learn the interpolation weights at the same time in one unified deep implicit function. We demonstrate the effectiveness of our JIIF representation on guided depth super-resolution task, significantly outperforming state-of-the-art methods on three public benchmarks. Code can be found at https://git.io/JC2sU
Jiaxiang Tang, Xiaokang Chen
ACM Multimedia2
2020 3D Sketch-Aware Semantic Scene Completion via Semi-Supervised Structure Prior
abstract
The goal of the Semantic Scene Completion (SSC) task is to simultaneously predict a completed 3D voxel representation of volumetric occupancy and semantic labels of objects in the scene from a single-view observation. Since the computational cost generally increases explosively along with the growth of voxel resolution, most current state-of-the-arts have to tailor their framework into a low-resolution representation with the sacrifice of detail prediction. Thus, voxel resolution becomes one of the crucial difficulties that lead to the performance bottleneck. In this paper, we propose to devise a new geometry-based strategy to embed depth information with low-resolution voxel representation, which could still be able to encode sufficient geometric information, e.g., room layout, object's sizes and shapes, to infer the invisible areas of the scene with well structure-preserving details. To this end, we first propose a novel 3D sketch-aware feature embedding to explicitly encode geometric information effectively and efficiently. With the 3D sketch in hand, we further devise a simple yet effective semantic scene completion framework that incorporates a light-weight 3D Sketch Hallucination module to guide the inference of occupancy and the semantic labels via a semi-supervised structure prior learning strategy. We demonstrate that our proposed geometric embedding works better than the depth feature learning from habitual SSC frameworks. Our final model surpasses state- of-the-arts consistently on three public benchmarks, which only requires 3D volumes of 60 × 36 × 60 resolution for both input and output.
Xiaokang Chen, Kwan-Yee Lin, Chen Qian 0006, Hongsheng Li 0001
CVPR1
2020 Bi-directional Cross-Modality Feature Propagation with Separation-and-Aggregation Gate for RGB-D Semantic Segmentation
Xiaokang Chen, Kwan-Yee Lin, Jingbo Wang 0003, Wayne Wu, Chen Qian 0006, Hongsheng Li 0001
ECCV (11)1
2020 Real-Time Semantic Scene Completion Via Feature Aggregation And Conditioned Prediction
abstract
Semantic Scene Completion (SSC) aims to simultaneously predict the volumetric occupancy and semantic category of a 3D scene. In this paper, we propose a real-time semantic scene completion method with a feature aggregation strategy and conditioned prediction module. Feature aggregation fuses feature with different receptive fields and gathers context to improve scene completion performance. And the conditioned prediction module adopts a two-step prediction scheme that takes volumetric occupancy as a condition to enhance semantic completion prediction. We conduct experiments on three recognized benchmarks NYU, NYUCAD, and SUNCG. Our method achieves competitive performance at a speed of 110 FPS on one GTX 1080 Ti GPU.
Xiaokang Chen, Yajie Xing
ICIP1
2019 Research on the Method and System of Word Segmentation and POS Tagging for Ancient Chinese Medicine Literature
abstract
Objectives: To better explore the valuable experience and knowledge in traditional Chinese medicine (TCM) literature and promote the academic heritage and innovation of TCM, this paper tries to provide methods and a system of word segmentation and POS tagging for ancient Chinese medicine literature. Method: Based on HMM and Java language for POS tagging, this paper develops a word segmentation system of Chinese ancient books by constructing a Thesaurus of Chinese medicine terminology and a special POS tagging method for Chinese medicine, and adopting Ansj open source code as the core word segmentation algorithm. Results: Thesaurus of Chinese medicine terminology involving 155,343 Chinese medicine words are constructed. These words were divided into 14 categories and 7 levels of 891 parts of speech (POS). Online website of Chinese word segmentation and POS tagging proprietary system were established (http://www.zhongyifenci.org). And the F measure reaches 90.65%, which is much higher than the system of Ansj. Clinical or Biological Impact: This system, with high precision and recall rate, will be very helpful to knowledge discovery of Chinese medicine and be beneficial to giving full play to the original advantages of Chinese medicine.
Xianjun Fu, Xuebo Li, Fangning Ju, Jintong Li, Xiaokang Chen, Sang Xiaoming
BIBM8
2019 2.5D Convolution for RGB-D Semantic Segmentation
abstract
Convolutional neural networks (CNN) have achieved great success in RGB semantic segmentation. RGB-D images provide additional depth information, which can improve segmentation performance. To take full advantages of the 3D geometry relations provided by RGB-D images, in this paper, we propose 2.5D convolution, which mimics one 3D convolution kernel by several masked 2D convolution kernels. Our 2.5D convolution can effectively process spatial relations between pixels in a manner similar to 3D convolution while still sampling pixels on 2D plane, and thus saves computational cost. And it can be seamlessly incorporated into pretrained CNNs. Experiments on two challenging RGB-D semantic segmentation benchmarks NYUDv2 and SUN-RGBD validate the effectiveness of our approach.
Yajie Xing, Jingbo Wang 0003, Xiaokang Chen
ICIP3
2019 Coupling Two-Stream RGB-D Semantic Segmentation Network by Idempotent Mappings
abstract
In RGB-D semantic segmentation tasks, it has been shown that HHA embeddings effectively encode rich depth features and using HHA together with RGB images can improve segmentation performance. In this paper, we propose a novel method to effectively integrate RGB and HHA features. By replacing identity mappings in ResNet-based two-stream network with idempotent mappings, we can couple the originally separated two branches to mix features from two modalities, while still keep the good information flow nature of ResNet. Moreover, our method does not bring any additional network blocks or parameters, and only needs very small modification on basic two-stream networks. We conduct experiments on two challenging RGB-D semantic segmentation datasets NYUDv2 and SUN-RGBD. The experiment results show that our method can significantly improve segmentation performance and our method achieves the state-of-the-art on these two datasets.
Yajie Xing, Jingbo Wang 0003, Xiaokang Chen
ICIP3
2017 Extending Blockchain Functionality with Statechain
abstract
Blockchain has attracted great attention as the basis of cryptocurrencies such as Bitcoin, but its capabilities extend far beyond that. Blockchain has the potential to revolutionize applications because it provides a transparent, immutable, and append-only ledger that can be used to build new decentralized applications. However, it is hard to implement changes to Bitcoin without forking blockchain. As a result, Bitcoin has difficulty in adapting to new demands and extending new functionality. The traditional way to introduce new functionality is to modify Bitcoin's code base, but which will lead to some problems such as (a) the security of blockchains, (b) the allocation of currency, and (c) the waste of resources. We present a novel design, statechain, which uses Bitcoin blockchain to propagate application log. It can add new functionality without requiring blockchain fork changes from Bitcoin. Statechain enables application nodes to efficiently query the log as well as transfer log between blockchains. We detail how the application nodes achieve application-level consensus at each block. As far as we know, this technology has not been exploited. However, it has a great practical significance that enables the introduction of new functionality safer and makes the development of blockchain-based applications easier. We have used statechain to explore the field of transferring bitcoins and other cryptocurrencies directly between multiple blockchains.
Xiaokang Chen, Kunlong Zhang
ICPADS1