EDBT 2026 Demo / reviewers in the wild / expert
Kang Li 0005
dblp:181/2763-0005
· DBLP profile ↗
17ranked-venue papers
0as first author
14since 2021 · last 2026
0000-0001-6218-5715ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 9 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Computer networks · 2 · 2 since 2021Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Point-DPA: Unifying contrastive and generative learning for 3D point cloud understanding via dynamic prototypes
Xin Cao 0004, Xinmeng Hu, Kang Li 0005, Linzhi Su, Fengjun Zhao |
Inf. Sci. | 4 |
| 2026 | Attention-guided multi-scale local reconstruction for point clouds via masked autoencoder self-supervised learning
Xin Cao 0004, Jiaxu Shi, Linzhi Su, Xinda Liu, Kang Li 0005 |
Multim. Syst. | 6 |
| 2025 | One Shot Learning for Edge Detection on Point CloudsabstractEach scanner possesses its unique characteristics and exhibits its distinct sampling error distribution. Training a network on a dataset that includes data collected from different scanners is less effective than training it on data specific to a single scanner. Therefore, we present a novel one-shot learning method allowing for edge extraction on point clouds, by learning the specific data distribution of the target point cloud, and thus achieve superior results compared to networks that were trained on general data distributions. More specifically, we present how to train a lightweight network named OSFENet (One-Shot edge Feature Extraction Network), by designing a filtered-KNN-based surface patch representation that supports a one-shot learning framework. Additionally, we introduce an RBF_DoS module, which integrates Radial Basis Function-based Descriptor of the Surface patch, highly beneficial for the edge extraction on point clouds. The advantage of the proposed OSFENet is demonstrated through comparative analyses against 7 baselines on the ABC dataset, and its practical utility is validated by results across diverse real-scanned datasets, including indoor scenes like S3DIS dataset, and outdoor scenes such as the Semantic3D dataset and UrbanBIS dataset. Zhikun Tu, Yiou Jia, Kang Li 0005, Daniel Cohen-Or |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2024 | CReStyler: Text-Guided Single Image Style Transfer Method Based on CNN and RestormerabstractText-guided image style transfer methods have gradually become a research hotspot. However, existing text-guided style transfer method suffers from content information missing and artifacts in the generated stylized images. Therefore, we propose CReStyler, a text-guided image style method based on the dual-branch structure of CNN and Restormer. In this work, in the first branch, we introduce a novel convolutional structure called FcasNet, which composes Frequency-domain Channel Attention Mechanism (FcaNet) and Cross Convolutional Block Attention Module (CRCBAM). It can generate rough style images lfcaaccording to the input target text. In the second branch, the use of pixel-based Restormer constrains the phenomenon of fake images and content information missing due to excessive convolution. It can generate stylized images lreswith complete content information based on the target text. Finally, we combine lfcaand lresthrough weighted fusion to obtain refined stylized images. During the training process, we utilize directional CLIP loss to constrain text-image alignment. Experimental results show that our method produces better results compared with existing methods such as CLIPStyler, LDAST, Text2LIVE, InstructPix2Pix. Long Feng, Guohua Geng, Kang Li 0005 |
ICASSP | 6 |
| 2024 | Neighborhood Multi-Compound Transformer for Point Cloud RegistrationabstractPoint cloud registration is a critical issue in 3D reconstruction and computer vision, particularly challenging in cases of low overlap and different datasets, where algorithm generalization and robustness are pressing challenges. In this paper, we propose a point cloud registration algorithm called Neighborhood Multi-compound Transformer (NMCT). To capture local information, we introduce Neighborhood Position Encoding for the first time. By employing a nearest neighbor approach to select spatial points, this encoding enhances the algorithm’s ability to extract relevant local feature information and local coordinate information from dispersed points within the point cloud. Furthermore, NMCT utilizes the Multi-compound Transformer as the interaction module for point cloud information. In this module, the Spatial Transformer phase engages in local-global fusion learning based on Neighborhood Position Encoding, facilitating the extraction of internal features within the point cloud. The Temporal Transformer phase, based on Neighborhood Position Encoding, performs local position-local feature interaction, achieving local and global interaction between two point cloud. The combination of these two phases enables NMCT to better address the complexity and diversity of point cloud data. The algorithm is extensively tested on different datasets (3DMatch, ModelNet, KITTI, MVP-RG), demonstrating outstanding generalization and robustness. Yong Wang 0057, Pengbo Zhou, Guohua Geng, Kang Li 0005, Ruoxue Li |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | PointCluster: Deep Clustering of 3-D Point Clouds With Semantic Pseudo-LabelingabstractPoint cloud classification is a fundamental problem in 3-D point cloud analysis. However, most existing methods are supervised, which requires costly and laborious annotations of large-scale point cloud datasets. This severely limits the practical applicability of point clouds. Therefore, exploring point cloud clustering methods, which can group point clouds into semantically meaningful clusters in an unsupervised manner, is of great importance. However, this remains a formidable challenge for humans. Here, we present PointCluster, a novel framework for deep clustering of 3-D point clouds. To enable accurate and reliable self-supervision for the clustering process, the framework introduces two semantic pseudo-labeling algorithms: prototype pseudo-labeling and reliable pseudo-labeling. We devise a three-step training process for the clustering network. First, we adopt a cross-modal representation learning approach to optimize the feature model. Second, we freeze the network parameters of the feature model and apply the prototype pseudo-labeling algorithm to optimize the clustering heads separately. Third, we use the reliable pseudo-labeling algorithm to jointly train the feature model and the clustering head in a semi-supervised manner, which enhances the overall clustering performance. The experimental results demonstrate that PointCluster achieves the state-of-the-art clustering results on public datasets such as ShapeNet. Moreover, our method narrows the gap between unsupervised point cloud clustering and supervised point cloud classification, offering a new perspective for the point cloud classification task. Xinxin Han, Huan Xia, Kang Li 0005, Gang Zhen, Linzhi Su, Fengjun Zhao, Xin Cao 0004 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | MATR: Multicompound Adaptive Transformer for Point Cloud RegistrationabstractPoint cloud registration plays a key role in the fields of computer vision, particularly in scenarios with low overlap, large scenes, different datasets, where difficulties, such as difficulty matching, scale changes and geometric deformations, local feature loss are commonly encountered. In this article, we propose a point cloud registration algorithm named multicompound adaptive transformer, which introduces adaptive position encoding, dynamically adjusting the local coordinates and feature information of scattered points within the point cloud through an adaptive threshold enhancement mechanism. Simultaneously, the multicompound transformer is introduced. In the spatial transformer stage, it accomplishes the local position-local feature interaction of individual point clouds through adaptive position encoding. Then, in the temporal transformer stage, it achieves local–local interaction and local–global information interaction between two point clouds through a dual-branch multiscale transformer. Through experiments on different datasets, we validate the algorithm's superior generalization performance in scenarios with low overlap, large scenes, and different datasets. Yong Wang 0057, Pengbo Zhou, Guohua Geng, Kang Li 0005 |
IEEE Trans. Ind. Informatics | 6 |
| 2024 | : Towards Collaborative and Cross-Domain Wi-Fi Sensing: A Case Study for Human Activity RecognitionabstractThe quality of a learning-based Wi-Fi sensing system is bounded by the quantity and quality of training data. However, obtaining sufficient and high-quality data across different domains is difficult due to extensive user involvement. We present CARING, a federated-learning-based framework to support collaborative and cross-domain Wi-Fi sensing. A key challenge of CARING is to allow the effective exchange and learning of knowledge across local models that are derived from heterogeneous data sources with uneven data distributions. We overcome this challenge by first extracting the activity-related representation to train local models. The shared global model aggregates received local model parameters and sends them back to individual devices for fine-tuning locally in the deployed environment. By leveraging the crowdsourced knowledge, CARING allows local models to quickly adapt to domain changes using just a few samples seen at test time. We demonstrate the benefit of CARING by applying it to activity recognition across three public datasets collected from 5 environments, 7 deployments, 31 users, and 29 activities. Experimental results show that CARING is highly effective and robust, improving the alternative approach for using single-sourced training data by up to 47%, giving an accuracy of over 80% (up to 100%) for various cross-domain scenarios. Xinyi Li 0005, Fengyi Song, Mina Luo, Kang Li 0005, Liqiong Chang, Xiaojiang Chen, Zheng Wang 0001 |
IEEE Trans. Mob. Comput. | 4 |
| 2024 | AQMon: A Fine-grained Air Quality Monitoring System Based on UAV Images for Smart CitiesabstractAir quality monitoring is important to the green development of smart cities. Several technical challenges exist for intelligent, high-precision monitoring, such as computing overhead, area division, and monitoring granularity. In this article, we propose a fine-grained air quality monitoring system based on visual inspection analysis embedded in unmanned aerial vehicle (UAV), referred to as AQMon . This system employs a lightweight neural network to obtain an accurate estimate of atmospheric transmittance in visual information while reducing computation and transmission overhead. Considering that air quality is affected by multiple factors, we design a dynamic fitting approach to model the relationship between scattering coefficients and PM2.5 concentration in real time. The proposed system is evaluated using public datasets and the results show that AQMon outperforms four existing methods with a processing time of 13.8 ms. Shuangqing Xia, Tianzhang Xing, Chase Qishi Wu, Jiadi Yang, Kang Li 0005 |
ACM Trans. Sens. Networks | 6 |
| 2023 | Gender-Cartoon: Image Cartoonization Method Based on Gender ClassificationabstractQin Opera art is one of China’s intangible cultural heritage, and its influence is gradually declining. The cartoonization of Qin Opera is one of the feasible methods. However, current cartoonization methods suffer from the inability to classify and accurately cartoonize Qinqiang portraits by gender. Therefore, we propose Gender-Cartoon, which can achieve different gender portrait cartoons. The proposed method consists of four modules: gender classification, content feature extraction, gender identification style feature extraction and cartoonization. The gender classification module is used to obtain the gender labels of cartoon images, content feature module extracts content features from portraits by stacking convolutional blocks, and gender identification style extraction module uses the gender labels of cartoon images and image semantic features to obtain the corresponding gender style feature. The cartoonization module fuses content features and style features as the input of the adaptive residual block to obtain the corresponding cartoonization results. The experimental results show that the model can transform both texture and shape on the our collected Qinq Cartoon dataset Face2QinqCartoon and the public dataset Selfie2anime. Long Feng, Guohua Geng, Longquan Yan, Xingrui Ma, Kang Li 0005 |
ICASSP | 7 |
| 2023 | Delta Path Tracing for Real-Time Global Illumination in Mixed RealityabstractVisual coherence between real and virtual objects is important in mixed reality (MR), and illumination consistency is one of the key aspects to achieve coherence. Apart from matching the illumination of the virtual objects with the real environments, the change of illumination on the real scenes produced by the inserted virtual objects should also be considered but is difficult to compute in real-time due to the heavy computation demands of global illumination. In this work, we propose delta path tracing (DPT), which only computes the radiance blocked by the virtual objects from the light sources at the primary hit points of Monte Carlo path tracing, then combines the blocked radiance and multi-bounce indirect illumination with the image of the real scene. Multiple importance sampling (MIS) between BRDF and environment map is performed to handle all-frequency environment maps captured by a panorama camera. Compared to conventional differential rendering methods, our method can remarkably reduce the number of times required to access the environment map and avoid rendering scenes twice. Therefore, the performance can be significantly improved. We implement our method using hardware-accelerated ray tracing on modern GPUs, and the results demonstrate that our method can render global illumination at real-time frame rates and produce plausible visual coherence between real and virtual objects in MR environments. Yang Xu 0092, Yuanfa Jiang, Kang Li 0005, Guohua Geng |
VR | 4 |
| 2023 | PuzzleFixer: A Visual Reassembly System for Immersive Fragments RestorationabstractWe present PuzzleFixer, an immersive interactive system for experts to rectify defective reassembled 3D objects. Reassembling the fragments of a broken object to restore its original state is the prerequisite of many analytical tasks such as cultural relics analysis and forensics reasoning. While existing computer-aided methods can automatically reassemble fragments, they often derive incorrect objects due to the complex and ambiguous fragment shapes. Thus, experts usually need to refine the object manually. Prior advances in immersive technologies provide benefits for realistic perception and direct interactions to visualize and interact with 3D fragments. However, few studies have investigated the reassembled object refinement. The specific challenges include: 1) the fragment combination set is too large to determine the correct matches, and 2) the geometry of the fragments is too complex to align them properly. To tackle the first challenge, PuzzleFixer leverages dimensionality reduction and clustering techniques, allowing users to review possible match categories, select the matches with reasonable shapes, and drill down to shapes to correct the corresponding faces. For the second challenge, PuzzleFixer embeds the object with node-link networks to augment the perception of match relations. Specifically, it instantly visualizes matches with graph edges and provides force feedback to facilitate the efficiency of alignment interactions. To demonstrate the effectiveness of PuzzleFixer, we conducted an expert evaluation based on two cases on real-world artifacts and collected feedback through post-study interviews. The results suggest that our system is suitable and efficient for experts to refine incorrect reassembled objects. Shuainan Ye, Chen Zhu-Tian, Xiangtong Chu, Kang Li 0005, Juntong Luo 0002, Guohua Geng, Yingcai Wu |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2022 | Precomputed Discrete Visibility Fields for Real-Time Ray-Traced Environment Lighting
Yang Xu 0092, Yuanfa Jiang, Kang Li 0005, Pengbo Zhou, Guohua Geng |
EGSR (ST) | 4 |
| 2022 | A novel compression framework of the dense point-cloud model for cultural heritage artifacts
Kang Li 0005, Jiaojiao Kou, Xiaoxue Chen, Linqi Hai, Guohua Geng, Shunli Zhang 0002 |
Multim. Tools Appl. | 2 |
| 2020 | Ancestry Estimation of Skull in Chinese Population Based on Improved Convolutional Neural NetworkabstractThe estimation of ancestry is an essential benchmark for positive identification of heavily decomposed bodies that are recovered in a variety of death and crime scenes. Aiming at the problem of skull ancestry estimation, this paper proposes an improved convolutional neural network method to realize ancestry estimation. We use the six-angle images of the skull as the input of the network. By improving the basic model LeNet5 of the convolutional neural network, we preserve the depth semantics and content information of the image, reduce the number of parameters, and ensure the learning ability of network features. In the experiment, 156 yellow skulls from northern China and 178 white skulls from Xinjiang were used as subjects, 80% of skull samples were used as training sets and 20% as test sets. Experiments on the training set and test set show that the improved CNN network architecture achieves 95.88% accuracy on the training set and 95.52% accuracy on the test set. In addition, we also designed experiments on the contribution of various parts of the skull to ancestor identification. The experimental results show that each region of the skull is useful for ancestor identification, but the effect is different. Compared with other networks, the network structure of this paper has the highest accuracy and better performance. Wen Yang 0003, Pengyue Lin, Guohua Geng, Xiaoning Liu 0001, Kang Li 0005 |
BIBM | 6 |
| 2020 | Stroke controllable style transfer based on dilated convolutionsabstractTransferring a photo to a stylised image with beautiful texture has become one of the most popular topics in computer vision and the application of image processing. Controlling the stroke size of the texture is one of the challenging problems in this task. Recent representative methods for such problem introduce a pyramid model to regulate receptive fields in the network. Meanwhile, dilated convolutions are proved to be a very efficient way to adjust receptive fields without losing resolution. By combining the advantages of both approaches and making special optimisation for VGG19 model for style transfer tasks, the authors propose to exploit dilated convolutions to extract texture information endowing the network with stroke controllable. Several sets of contrast experiments were conducted and results show that their algorithm can generate more attractive stylisation images and control stroke size flexibly. It demonstrates the superiority of applying dilated convolutions as a texture extraction method for maintaining more texture information and controlling stroke size. Zhaopan Xu, Yu Zhang 0040, Kang Li 0005, Shengling Geng |
IET Comput. Vis. | 5 |
| 2020 | A computerized craniofacial reconstruction method for an unidentified skull based on statistical shape models
Wuyang Shui, Steve C. Maddock, Qingqiong Deng, Kang Li 0005, Yachun Fan, Xiujie Wu |
Multim. Tools Appl. | 6 |