EDBT 2026 Demo / reviewers in the wild / expert
Xin Kong
dblp:126/5666
· DBLP profile ↗
13ranked-venue papers
6as first author
9since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 5 first-author · 8 since 2021Systems, architecture and hardware · 7 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | EscherNet: A Generative Model for Scalable View SynthesisabstractWe introduce EscherNet, a multi-view conditioned diffusion model for view synthesis. EscherNet learns implicit and generative 3D representations coupled with a specialised camera positional encoding, allowing precise and continuous relative control of the camera transformation between an arbitrary number of reference and target views. EscherNet offers exceptional generality, flexibility, and scalability in view synthesis ─it can generate more than 100 consistent target views simultaneously on a single consumer-grade GPU, despite being trained with a fixed number of 3 reference views to 3 target views. As a result, EscherNet not only addresses zero-shot novel view synthesis, but also naturally unifies single- and multi-image 3D reconstruction, combining these diverse tasks into a single, cohesive framework. Our extensive experiments demonstrate that EscherNet achieves state-of-the-art performance in multiple benchmarks, even when compared to methods specifically tailored for each individual problem. This remarkable versatility opens up new directions for designing scalable neural architectures for 3D vision. Project page: https://kxhit.github.io/EscherNet. Xin Kong, Shikun Liu, Xiaoyang Lyu, Marwan Taher, Xiaojuan Qi 0001, Andrew J. Davison |
CVPR | 1 |
| 2023 | vMAP: Vectorised Object Mapping for Neural Field SLAMabstractWe present vMAP, an object-level dense SLAM system using neural field representations. Each object is repre-sented by a small MLP, enabling efficient, watertight object modelling without the needfor 3D priors. As an RGB-D camera browses a scene with no prior in-formation, vMAP detects object instances on-the-fly, and dynamically adds them to its map. Specifically, thanks to the power of vectorised training, vMAP can optimise as many as 50 individual objects in a single scene, with an extremely efficient training speed of 5Hz map update. We experimentally demonstrate significantly improved scene-level and object-level reconstruction quality compared to prior neural field SLAM systems. Project page: https://kxhit.github.io/vMAP. Xin Kong, Shikun Liu, Marwan Taher, Andrew J. Davison |
CVPR | 1 |
| 2023 | GADA-SegNet: gated attentive domain adaptation network for semantic segmentation of LiDAR point clouds
Xin Kong, Shifeng Xia, Ningzhong Liu, Mingqiang Wei |
Vis. Comput. | 1 |
| 2022 | A novel ConvLSTM with multifeature fusion for financial intelligent tradingabstractHigh fluctuation and self-similarity are typical characteristics of financial time series. Furthermore, affected by market environment, such as regular announcement of important economic data, time/date-sensitive fluctuations commonly exist in financial time series. However, the existing learning models were usually lack the consideration of essential characteristics of financial data, where both the fusion learning of multiple temporal features and the necessary attention to time-sensitive fluctuations were ignored. Inspired by this, to represent the temporal characteristics of self-similarity and reveal intrinsic feature details, in this article, time series and its features are converted into visibility graphs using the technique of Gramian Angular Fields, based on which convolutional long short-term memory (ConvLSTM) is applied to implement multifeature fusion learning. Moreover, to capture the time/date-sensitive fluctuation existing in financial time series, a subspace decomposition composed of the fuzzy control mechanism is first introduced into the ConvLSTM model, which considerably improves the prediction performance. On the basis of the proposed learning model, a concise intelligent trading strategy is designed. By using real foreign exchange data, various experiments are implemented to show the effectiveness of the proposed model. Xin Kong, Chao Luo 0001 |
Int. J. Intell. Syst. | 1 |
| 2021 | HR-Depth: High Resolution Self-Supervised Monocular Depth EstimationabstractSelf-supervised learning shows great potential in monocular depth estimation, using image sequences as the only source of supervision. Although people try to use the high-resolution image for depth estimation, the accuracy of prediction has not been significantly improved. In this work, we find the core reason comes from the inaccurate depth estimation in large gradient regions, making the bilinear interpolation error gradually disappear as the resolution increases. To obtain more accurate depth estimation in large gradient regions, it is necessary to obtain high-resolution features with spatial and semantic information. Therefore, we present an improved DepthNet, HR-Depth, with two effective strategies: (1) re-design the skip-connection in DepthNet to get better high-resolution features and (2) propose feature fusion Squeeze-and-Excitation(fSE) module to fuse feature more efficiently. Using Resnet-18 as the encoder, HR-Depth surpasses all previous state-of-the-art(SoTA) methods with the least parameters at both high and low resolution. Moreover, previous SoTA methods are based on fairly complex and deep networks with a mass of parameters which limits their real applications. Thus we also construct a lightweight network which uses MobileNetV3 as encoder. Experiments show that the lightweight network can perform on par with many large models like Monodepth2 at high-resolution with only20%parameters. All codes and models will be available at https://github.com/shawLyu/HR-Depth. Xiaoyang Lyu, Liang Liu 0007, Mengmeng Wang 0005, Xin Kong, Lina Liu 0010, Yong Liu 0007, Xinxin Chen, Yi Yuan 0002 |
AAAI | 4 |
| 2021 | PocoNet: SLAM-oriented 3D LiDAR Point Cloud Online Compression NetworkabstractIn this paper, we present PocoNet: Point cloud Online COmpression NETwork to address the task of SLAM-oriented compression. The aim of this task is to select a compact subset of points with high priority to maintain localization accuracy. The key insight is that points with high priority have similar geometric features in SLAM scenarios. Hence, we tackle this task as point cloud segmentation to capture complex geometric information. We calculate observation counts by matching between maps and point clouds and divide them into different priority levels. Trained by labels annotated with such observation counts, the proposed network could evaluate the point-wise priority. Experiments are conducted by integrating our compression module into an existing SLAM system to evaluate compression ratios and localization performances. Experimental results on two different datasets verify the feasibility and generalization of our approach. Jinhao Cui, Xin Kong, Xuemeng Yang, Xiangrui Zhao, Yong Liu 0007, Wanlong Li, Hongbo Zhang 0004 |
ICRA | 3 |
| 2021 | SA-LOAM: Semantic-aided LiDAR SLAM with Loop ClosureabstractLiDAR-based SLAM system is admittedly more accurate and stable than others, while its loop closure detection is still an open issue. With the development of 3D semantic segmentation for point cloud, semantic information can be obtained conveniently and steadily, essential for high-level intelligence and conductive to SLAM. In this paper, we present a novel semantic-aided LiDAR SLAM with loop closure based on LOAM, named SA-LOAM, which leverages semantics in odometry as well as loop closure detection. Specifically, we propose a semantic-assisted ICP, including semantically matching, downsampling and plane constraint, and integrates a semantic graph-based place recognition method in our loop closure detection module. Benefitting from semantics, we can improve the localization accuracy, detect loop closures effectively, and construct a global consistent semantic map even in large-scale scenes. Extensive experiments on KITTI and Ford Campus dataset show that our system significantly improves baseline performance, has generalization ability to unseen data and achieves competitive results compared with state-of-the-art methods. Lin Li 0091, Xin Kong, Xiangrui Zhao, Wanlong Li, Hongbo Zhang 0004, Yong Liu 0007 |
ICRA | 2 |
| 2021 | SSC: Semantic Scan Context for Large-Scale Place RecognitionabstractPlace recognition gives a SLAM system the ability to correct cumulative errors. Unlike images that contain rich texture features, point clouds are almost pure geometric information which makes place recognition based on point clouds challenging. Existing works usually encode low-level features such as coordinate, normal, reflection intensity, etc., as local or global descriptors to represent scenes. Besides, they often ignore the translation between point clouds when matching descriptors. Different from most existing methods, we explore the use of high-level features, namely semantics, to improve the descriptor’s representation ability. Also, when matching descriptors, we try to correct the translation between point clouds to improve accuracy. Concretely, we propose a novel global descriptor, Semantic Scan Context, which explores semantic information to represent scenes more effectively. We also present a two-step global semantic ICP to obtain the 3D pose (x, y, yaw) used to align the point cloud to improve matching performance. Our experiments on the KITTI dataset show that our approach outperforms the state-of-the- art methods with a large margin. Our code is available at: https://github.com/lilin-hitcrt/SSC. Lin Li 0091, Xin Kong, Xiangrui Zhao, Tianxin Huang, Wanlong Li, Hongbo Zhang 0004, Yong Liu 0007 |
IROS | 2 |
| 2021 | Semantic Segmentation-assisted Scene Completion for LiDAR Point CloudsabstractOutdoor scene completion is a challenging issue in 3D scene understanding, which plays an important role in intelligent robotics and autonomous driving. Due to the sparsity of LiDAR acquisition, it is far more complex for 3D scene completion and semantic segmentation. Since semantic features can provide constraints and semantic priors for completion tasks, the relationship between them is worth exploring. Therefore, we propose an end-to-end semantic segmentation-assisted scene completion network, including a 2D completion branch and a 3D semantic segmentation branch. Specifically, the network takes a raw point cloud as input, and merges the features from the segmentation branch into the completion branch hierarchically to provide semantic information. By adopting BEV representation and 3D sparse convolution, we can benefit from the lower operand while maintaining effective expression. Besides, the decoder of the segmentation branch is used as an auxiliary, which can be discarded in the inference stage to save computational consumption. Extensive experiments demonstrate that our method achieves competitive performance on SemanticKITTI dataset with low latency. Code and models will be released at https://github.com/jokester-zzz/SSA-SC. Xuemeng Yang, Xin Kong, Tianxin Huang, Yong Liu 0007, Wanlong Li, Hongbo Zhang 0004 |
IROS | 3 |
| 2020 | Semantic Graph Based Place Recognition for 3D Point CloudsabstractDue to the difficulty in generating the effective descriptors which are robust to occlusion and viewpoint changes, place recognition for 3D point cloud remains an open issue. Unlike most of the existing methods that focus on extracting local, global, and statistical features of raw point clouds, our method aims at the semantic level that can be superior in terms of robustness to environmental changes. Inspired by the perspective of humans, who recognize scenes through identifying semantic objects and capturing their relations, this paper presents a novel semantic graph based approach for place recognition. First, we propose a novel semantic graph representation for the point cloud scenes by reserving the semantic and topological information of the raw point cloud. Thus, place recognition is modeled as a graph matching problem. Then we design a fast and effective graph similarity network to compute the similarity. Exhaustive evaluations on the KITTI dataset show that our approach is robust to the occlusion as well as viewpoint changes and outperforms the state-of-the-art methods with a large margin. Our code is available at: https://github.com/kxhit/SG_PR. Xin Kong, Xuemeng Yang, Guangyao Zhai, Xiangrui Zhao, Xianfang Zeng, Mengmeng Wang 0005, Yong Liu 0007, Wanlong Li |
IROS | 1 |
| 2020 | F-Siamese Tracker: A Frustum-based Double Siamese Network for 3D Single Object TrackingabstractThis paper presents F-Siamese Tracker, a novel approach for single object tracking prominently characterized by more robustly integrating 2D and 3D information to reduce redundant search space. A main challenge in 3D single object tracking is how to reduce search space for generating appropriate 3D candidates. Instead of solely relying on 3D proposals, firstly, our method leverages the Siamese network applied on RGB images to produce 2D region proposals which are then extruded into 3D viewing frustums. Besides, we perform an on-line accuracy validation on the 3D frustum to generate refined point cloud searching space, which can be embedded directly into the existing 3D tracking backbone. For efficiency, our approach gains better performance with fewer candidates by reducing search space. In addition, benefited from introducing the online accuracy validation, for occasional cases with strong occlusions or very sparse points, our approach can still achieve high precision, even when the 2D Siamese tracker loses the target. This approach allows us to set a new state-of-the-art in 3D single object tracking by a significant margin on a sparse outdoor dataset (KITTI tracking). Moreover, experiments on 2D single object tracking show that our framework boosts 2D tracking performance as well. Jinhao Cui, Xin Kong, Chujuan Zhang, Yong Liu 0007, Wanlong Li |
IROS | 3 |
| 2020 | Learning to Compensate for the Drift and Error of Gyroscope in Vehicle LocalizationabstractSelf-localization is an essential technology for autonomous vehicles. Building robust odometry in a GPS-denied environment is still challenging, especially when LiDAR and camera are uninformative. In this paper, We propose a learning-based approach to cure the drift of gyroscope for vehicle localization. For consumer-level MEMS gyroscope (stability ~10° /h), our GyroNet can estimate the error of each measurement. For high-precision Fiber optics Gyroscope (stability ~0.05° /h), we build a FoGNet which can obtain its drift by observing data in a long time window. We perform comparative experiments on publicly available datasets. The results demonstrate that our GyroNet can get higher precision angular velocity than traditional digital filters and static initialization methods. In the vehicle localization, the FoGNet can effectively correct the small drift of the Fiber optics Gyroscope (FoG) and can achieve better results than the state-of-the-art method. Xiangrui Zhao, Chunfang Deng, Xin Kong, Jinhong Xu, Yong Liu 0007 |
IV | 3 |
| 2019 | PASS3D: Precise and Accelerated Semantic Segmentation for 3D Point CloudabstractIn this paper, we propose PASS3D to achieve point-wise semantic segmentation for 3D point cloud. Our framework combines the efficiency of traditional geometric methods with robustness of deep learning methods, consisting of two stages: At stage -1, our accelerated cluster proposal algorithm will generate refined cluster proposals by segmenting point clouds without ground, capable of generating less redundant proposals with higher recall in an extremely short time; stage -2 we will amplify and further process these proposals by a neural network to estimate semantic label for each point and meanwhile propose a novel data augmentation method to enhance the network's recognition capability for all categories especially for non-rigid objects. Evaluated on KITTI raw dataset, PASS3D stands out against the state-of-the-art on some results, making itself competent to 3D perception in autonomous driving system. Our source code will be open-sourced. A video demonstration is available at https://www.youtube.com/watch?v=cukEqDuP_Qw. Xin Kong, Guangyao Zhai, Baoquan Zhong, Yong Liu 0007 |
IROS | 1 |