EDBT 2026 Demo / reviewers in the wild / expert
Guohua Geng
dblp:25/3718
· DBLP profile ↗
49ranked-venue papers
0as first author
38since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 23 · 18 since 2021Artificial intelligence and machine learning · 17 · 13 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 6 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Walking in the Wild: Safe and Natural Redirected Walking in Open Physical SpacesabstractRedirected Walking (RDW) enables continuous locomotion in virtual environments (VEs) within limited physical spaces. However, classic RDW methods rely on static settings with tight boundaries, which can reduce their applicability in large shared physical spaces where boundary constraints are not dominant, especially under dynamic and multi-user conditions. To overcome this, we introduce a new task: Redirected Walking in Open Physical Spaces (OPSRDW), which allows users to navigate expansive VEs safely and naturally despite dynamic obstacles and without requiring a fixed boundary model. We further propose Dynamic Control and Redirection with Safety Constraints (DyCoRe), which formulates OPS-RDW as a constrained optimization problem. DyCoRe uses Dynamic Control Barrier Functions to model real-time collision avoidance constraints and solves an online Quadratic Programming problem to compute optimal velocities that minimize path deviation while ensuring safety. These velocities are mapped into real-time redirection gains, guiding users along natural paths while reducing collision risk. A relaxation mechanism is incorporated to handle infeasible scenarios. Extensive simulations suggest consistent improvements in safety and obstacle clearance over state-of-theart methods. User studies demonstrate that DyCoRe significantly improves navigation continuity, reduces the number of resets, and tends to reduce perceived discomfort compared to classic RDW strategies. DyCoRe provides an efficient and learning-free solution for safe and natural VR locomotion in open physical spaces with dynamic obstacles and multiple users. Xinda Liu, Guoqiang Yang, Yunchen Li, Jian Wu 0033, Guohua Geng, Lili Wang 0006 |
VR | 6 |
| 2026 | Diffusion-based lossy geometry compression for three dimensional point clouds
Qicheng Wang, Guohua Geng |
Eng. Appl. Artif. Intell. | 7 |
| 2026 | A lightweight vision transformer for efficient craniofacial reconstruction
Xizhi Wang, Qinghe Tao, Qingrong Liu, Ruoxue Li, Guohua Geng, Wuyang Shui |
Eng. Appl. Artif. Intell. | 7 |
| 2026 | POFNet: Self-supervised 3D pyramid optical flow network with 3D swapping and ConvLSTM for point cloud prediction
Zhaoying Ye, Muhammad Shahroz Ajmal, Guohua Geng |
Expert Syst. Appl. | 5 |
| 2026 | FPM-GAN: Craniofacial reconstruction method based on frequency domain perception and multi-scale attention
Wen Yang 0003, Longqian Ma, Dengwei Yan, Zhengran Cao, Wenfeng Zhang, Guohua Geng |
Expert Syst. Appl. | 7 |
| 2026 | From SVG to DSVG: Leveraging Foveal Visual Cues to Mitigate Cybersickness in Redirected WalkingabstractABSTRACT Redirected Walking (RDW) effectively extends the navigable area of virtual environments but frequently induces cybersickness due to vestibular‐visual conflicts. To mitigate this, this study proposes a hierarchical visual guidance framework. We first introduce a Steering Visual Guidance (SVG) model, which employs a grid‐based pattern to provide a stable visual reference frame. While effective, our evaluation revealed that the static, full‐field nature of SVG can occlude the user's view and reduce visual clarity. To address this limitation, we optimized the model into Dynamic Steering Visual Guidance (DSVG). Grounded in the physiological principles of central and peripheral vision, DSVG dynamically renders stability cues exclusively within the foveal region, fading them in the periphery based on retinal sensitivity. A user study demonstrates that both SVG and DSVG significantly reduce subjective discomfort and physiological markers of sickness compared to a control condition without cues. Crucially, DSVG achieves sickness mitigation comparable to SVG while reducing visual occlusion by approximately 70%. These findings suggest that DSVG offers a robust, occlusion‐minimizing solution for deploying RDW in constrained physical spaces. Xinda Liu, Yongbo Tang, Xiaoning Liu 0001, Pengbo Zhou, Guohua Geng |
Comput. Animat. Virtual Worlds | 5 |
| 2026 | PeHNet: Closed-loop prototype enhancement for weakly supervised few-shot segmentation
Muhammad Shahroz Ajmal, Guohua Geng, Pengbo Zhou, Mohsin Ashraf |
Knowl. Based Syst. | 2 |
| 2026 | Knowledge-injected prompt tuning with semantic regularization for fine-grained image recognition
Xinda Liu, Pengbo Zhou, Guohua Geng |
Knowl. Based Syst. | 5 |
| 2026 | Structured-condensed prompt tuning in vision-language models for fine-grained image recognition
Xinda Liu, Weiqing Min, Guohua Geng, Shuqiang Jiang |
Pattern Recognit. | 4 |
| 2026 | Dual-Branch Feature Fusion for Sparse-View X-Ray 3D ReconstructionabstractWith the rapid development of medical imaging, sparse-view X-ray 3D reconstruction has become an essential technique for addressing low-dose X-ray imaging challenges. However, due to sparse angular sampling, traditional reconstruction methods often face challenges in handling complex bone and soft tissue structures, leading to information loss and insufficient detail capture. To address these issues, this paper proposes a sparse-view X-ray 3D reconstruction method based on Neural Radiance Fields (NeRF) with a dual-branch feature fusion framework. By synergistically extracting local and global features, this approach enhances the reconstruction of intricate bone and soft tissue structures. For specific applications in regions like the pelvis and aneurism, the method employs depthwise separable convolutions in the local branch to efficiently capture X-ray image details, enhancing the reconstruction of complex bone structures. In the global branch, a pooling Transformer with window mechanisms and hybrid positional encoding is introduced to capture the global features of soft tissue structures like aneurism. Experimental results demonstrate the superiority of this method on multiple medical imaging datasets, particularly in reconstructing complex bone regions and recovering details of soft tissue structures, compared to traditional methods and existing deep learning models. Yong Wang 0057, Guohua Geng, Wen Tang 0004 |
IEEE Signal Process. Lett. | 3 |
| 2026 | From Structure to Semantics: Hypergraph-Based AR Assembly Guidance with LLM-Mediated NarrationabstractEffective Augmented Reality (AR) guidance for complex assembly faces a dual challenge: the inability of conventional liaison graphs to represent procedural logic, and the cognitive burden imposed by visual instructions. We argue that the solution requires a more expressive structure to overcome these representational deficits and a narration approach to mediate instruction complexity. Our method first employs an assembly hypergraph to capture the task's hierarchical information, from which an A* search algorithm generates an optimal assembly path. Then a Large Language Model (LLM)-mediated narration workflow is designed to address the ergonomic deficiencies of the machine-centric path. It employs an optimizer to improve fluency, followed by a narrator that crafts the steps into an intuitive instruction narration. A within-subjects user study (N = 24) revealed a progressive enhancement from our method's components. The transition from a liaison-graph baseline to the hypergraph alone improved objective outcomes by reducing task time and errors and improving subjective ratings (SUS, NASA-TLX, TAM, and ARI). Subsequently, augmenting the LLM-mediated narration maintained these gains while lowering cognitive load and elevating user experience and usability. Our findings indicate the value of our AR assembly design and discuss the opportunities of using LLM as a mediation layer for better user interaction. Xinda Liu, Jiaju Xu, Jian Wu 0033, Guohua Geng, Lili Wang 0006 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2025 | Real-Time Physically-Based Relighting and Composition of Radiance Fields with Proxy MeshesabstractRadiance fields, such as neural radiance fields (NeRFs) and 3D Gaussian splatting (3DGS), are the new primitives to represent 3D scenes. Relighting and composition of radiance fields are critical for modeling the complex 3D world in computer graphics. However, it is difficult to relight and composite radiance fields because traditional physically-based rendering techniques, such as path tracing, cannot be directly applied to radiance fields. We propose a physically-based relighting and composition method for radiance fields with proxy meshes. A unified framework is presented to enable us to use radiance fields as the traditional assets in computer graphics. We generate proxy meshes of the radiance fields by reconstructing the geometries of the scenes using Gaussian-based surface reconstruction and the materials using physically-based differentiable rendering. We leverage differential rendering, which is previously used in augmented reality (AR) and mixed reality (MR), to evaluate the radiance change on the proxy meshes introduced by the changing lighting condition, the inserted radiance fields, or the inserted mesh models. Proxy meshes can help us utilize hardwareaccelerated ray tracing to perform real-time path tracing. Experimental results show that our method outperforms the baselines in terms of relighting performance and can achieve photorealistic relighting and composition of radiance fields in real-time. Jinyang Bo, Guohua Geng |
ISMAR | 7 |
| 2025 | Self-Supervised Image Segmentation Using Meta-Learning and Multi-Backbone Feature FusionabstractFew-shot segmentation (FSS) aims to reduce the need for manual annotation, which is both expensive and time-consuming. While FSS enhances model generalization to new concepts with only limited test samples, it still relies on a substantial amount of labeled training data for base classes. To address these issues, we propose a multi-backbone few shot segmentation (MBFSS) method. This self-supervised FSS technique utilizes unsupervised saliency for pseudo-labeling, allowing the model to be trained on unlabeled data. In addition, it integrates features from multiple backbones (ResNet, ResNeXt, and PVT v2) to generate a richer feature representation than a single backbone. Through extensive experimentation on PASCAL-5i and COCO-20i, our method achieves 54.3% and 25.1% on one-shot segmentation, exceeding the baseline methods by 13.5% and 4%, respectively. These improvements significantly enhance the model's performance in real-world applications with negligible labeling effort. Muhammad Shahroz Ajmal, Guohua Geng, Mohsin Ashraf |
Int. J. Neural Syst. | 2 |
| 2025 | IOPCNet: inner and outer point classification based low overlap rate local-to-global point cloud registration
Pengbo Zhou, Wen Tang 0004, Wuyang Shui, Guohua Geng |
Multim. Syst. | 8 |
| 2025 | CR-DM: A novel craniofacial reconstruction framework based on diffusion model
Xizhi Wang, Yanan Jin, Ruoxue Li, Guohua Geng |
Multim. Syst. | 7 |
| 2025 | CMFF: Cross-modal feature fusion network for robust point cloud completion
Pengbo Zhou, Xinda Liu, Longquan Yan, Guohua Geng |
Neural Networks | 7 |
| 2025 | A novel deep neural network for identification of sex and ethnicity based on unknown skulls
Qianhong Li, Xizhi Wang, Qianyi Wu, Chaohui Ma, Guohua Geng |
Pattern Recognit. | 7 |
| 2024 | CReStyler: Text-Guided Single Image Style Transfer Method Based on CNN and RestormerabstractText-guided image style transfer methods have gradually become a research hotspot. However, existing text-guided style transfer method suffers from content information missing and artifacts in the generated stylized images. Therefore, we propose CReStyler, a text-guided image style method based on the dual-branch structure of CNN and Restormer. In this work, in the first branch, we introduce a novel convolutional structure called FcasNet, which composes Frequency-domain Channel Attention Mechanism (FcaNet) and Cross Convolutional Block Attention Module (CRCBAM). It can generate rough style images lfcaaccording to the input target text. In the second branch, the use of pixel-based Restormer constrains the phenomenon of fake images and content information missing due to excessive convolution. It can generate stylized images lreswith complete content information based on the target text. Finally, we combine lfcaand lresthrough weighted fusion to obtain refined stylized images. During the training process, we utilize directional CLIP loss to constrain text-image alignment. Experimental results show that our method produces better results compared with existing methods such as CLIPStyler, LDAST, Text2LIVE, InstructPix2Pix. Long Feng, Guohua Geng, Kang Li 0005 |
ICASSP | 2 |
| 2024 | Where Should a Virtual Guide Stand in a VR Museum?
Xinda Liu, Jian Wu 0033, Lili Wang 0006, Guohua Geng |
ICXR | 5 |
| 2024 | One-stop multiscale reconciliation attention network with scribble supervision for salient object detection in optical remote sensing images
Ruixiang Yan, Longquan Yan, Yufei Cao, Guohua Geng, Pengbo Zhou |
Appl. Intell. | 4 |
| 2024 | Global-Local Semantic Interaction Network for Salient Object Detection in Optical Remote Sensing Images With Scribble SupervisionabstractSalient object detection in optical remote sensing images (RSI-SOD) is critical in remote sensing, yet it faces challenges such as dependency on intensive pixel-level annotations and limited research on low-cost, weakly supervised methods. These challenges are compounded by difficulties in handling complex backgrounds and varying salient object features with existing CNN-based methods. We introduce the Global-Local Semantic Interaction Network (GLSIN), a high-performance, cost-effective RSI-SOD approach based on scribble supervision. GLSIN employs an encoder-decoder framework, blending a Transformer and CNN to create a Dual Branch Encoder that effectively captures both global and local features of images. The Global-Local Affinity Block (GLAB) and Feature Shrinkage Decoder with the Global-Local Fusion Block (GLFB) are integrated to enhance feature interaction and precision in saliency map generation. Experimental results on two public datasets show that our method achievesFmaxβ,Emaxξ,Sα, andMscores of 86.6%, 96.5%, 91.8%, and 0.7% on the EORSSD dataset, and 90.1%, 97.2%, 91.7%, and 1.1% on the ORSSD dataset, respectively. The performance surpasses existing weakly-supervised or unsupervised SOD methods and even some fully-supervised models. Ruixiang Yan, Longquan Yan, Yufei Cao, Guohua Geng, Pengbo Zhou, Yongle Meng |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2024 | GTGMM: geometry transformer and Gaussian Mixture Models for robust point cloud registration
Linqi Hai, Ruoxue Li, Guohua Geng |
Multim. Tools Appl. | 6 |
| 2024 | STNet: Structure and texture-guided network for image inpainting
Yingfei Du, Yongqin Zhang, Guohua Geng |
Pattern Recognit. | 7 |
| 2024 | DM-GAN: CNN hybrid vits for training GANs under limited data
Longquan Yan, Ruixiang Yan, Bosong Chai, Guohua Geng, Pengbo Zhou |
Pattern Recognit. | 4 |
| 2024 | Low-Overlap Point Cloud Registration With TransformerabstractIn real-world scenarios, due to factors like sensor noise, point cloud data often exhibits low overlap, posing challenges for traditional registration methods. To address this issue, we propose a low-overlap point cloud registration with Transformer. This algorithm employs a dynamic positional encoding strategy that adaptively computes position encodings for each point based on its distribution. This enables better capturing of richer spatial relationships between point clouds and facilitates adaptation to diverse point cloud distributions across various scenes. Furthermore, we combine the mechanisms of self-attention and graph convolutions. The self-attention mechanism captures global dependencies among points, while the graph convolutions capture local neighborhood information between points. Lastly, in the context of cross-attention, adaptive weights are introduced during the attention calculation process. This involves multiplying attention scores by adaptive weights, enhancing the model's ability to focus on crucial registration areas. In scenarios with low overlap, this algorithm significantly enhances the success rate of successful registrations. It achieves notable improvements and attains a new state-of-the-art performance in the 3DLoMatch benchmark test, reaching a registration recall rate of 71.7%. Yong Wang 0057, Pengbo Zhou, Guohua Geng, Qi Zhang 0091 |
IEEE Signal Process. Lett. | 3 |
| 2024 | Neighborhood Multi-Compound Transformer for Point Cloud RegistrationabstractPoint cloud registration is a critical issue in 3D reconstruction and computer vision, particularly challenging in cases of low overlap and different datasets, where algorithm generalization and robustness are pressing challenges. In this paper, we propose a point cloud registration algorithm called Neighborhood Multi-compound Transformer (NMCT). To capture local information, we introduce Neighborhood Position Encoding for the first time. By employing a nearest neighbor approach to select spatial points, this encoding enhances the algorithm’s ability to extract relevant local feature information and local coordinate information from dispersed points within the point cloud. Furthermore, NMCT utilizes the Multi-compound Transformer as the interaction module for point cloud information. In this module, the Spatial Transformer phase engages in local-global fusion learning based on Neighborhood Position Encoding, facilitating the extraction of internal features within the point cloud. The Temporal Transformer phase, based on Neighborhood Position Encoding, performs local position-local feature interaction, achieving local and global interaction between two point cloud. The combination of these two phases enables NMCT to better address the complexity and diversity of point cloud data. The algorithm is extensively tested on different datasets (3DMatch, ModelNet, KITTI, MVP-RG), demonstrating outstanding generalization and robustness. Yong Wang 0057, Pengbo Zhou, Guohua Geng, Kang Li 0005, Ruoxue Li |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | ASNet: Adaptive Semantic Network Based on Transformer-CNN for Salient Object Detection in Optical Remote Sensing ImagesabstractSalient object detection in optical remote sensing images (RSI-SOD) has recently become a key area of research, driven by the unique challenges posed by the variability in remote sensing imagery. Traditional approaches, largely based on Convolutional Neural Networks (CNNs), are limited in handling the diverse scenarios of remote sensing due to their static network construction and reliance on local feature extraction. To tackle these limitations, we present the Adaptive Semantic Network (ASNet), a novel framework specifically designed for RSI-SOD. ASNet innovatively integrates Transformer and CNN technologies in a Dual Branch Encoder, which captures both global dependencies and local fine-grained image details. The network also features an Adaptive Semantic Matching Module (ASMM) for dynamically harmonizing filter responses to global and local contexts, an Adaptive Feature Enhancement Module (AFEM) that effectively enhances salient region features while restoring image resolution, and a Multi-scale Fine-grained Inference Module (MFIM) which refines high-level semantic features by integrating detailed low-level information, leading to the generation of precise, high-quality saliency maps. These components work in concert to adaptively respond to the complex nature of remote sensing images. Extensive experimental evaluations confirm that ASNet substantially outperforms existing models in the RSI-SOD task. Ruixiang Yan, Longquan Yan, Guohua Geng, Yufei Cao, Pengbo Zhou, Yongle Meng |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | MATR: Multicompound Adaptive Transformer for Point Cloud RegistrationabstractPoint cloud registration plays a key role in the fields of computer vision, particularly in scenarios with low overlap, large scenes, different datasets, where difficulties, such as difficulty matching, scale changes and geometric deformations, local feature loss are commonly encountered. In this article, we propose a point cloud registration algorithm named multicompound adaptive transformer, which introduces adaptive position encoding, dynamically adjusting the local coordinates and feature information of scattered points within the point cloud through an adaptive threshold enhancement mechanism. Simultaneously, the multicompound transformer is introduced. In the spatial transformer stage, it accomplishes the local position-local feature interaction of individual point clouds through adaptive position encoding. Then, in the temporal transformer stage, it achieves local–local interaction and local–global information interaction between two point clouds through a dual-branch multiscale transformer. Through experiments on different datasets, we validate the algorithm's superior generalization performance in scenarios with low overlap, large scenes, and different datasets. Yong Wang 0057, Pengbo Zhou, Guohua Geng, Kang Li 0005 |
IEEE Trans. Ind. Informatics | 3 |
| 2023 | Stature estimation using measurements of the skull in Chinese populationsabstractEstimating stature based on skeleton plays an important role in forensic investigation and individual identification. The skull consists of hard tissue and is the best-preserved part of an individual’s postmortem skeleton. Therefore, in many cases, it is the only one available. The purpose of this study was to evaluate the correlation between stature and skull measurements in a contemporary Chinese population using a 3D skull model reconstructed from CT scans. A total of 320 skull models were sampled in this study. These data were collected from populations in northern and southern China. Of these, 256 were used to establish regression equations and 64 were used for blind testing. We performed twelve skull measurements for each skull and used the Pearson correlation coefficient to estimate the correlation between stature and each measurement. Simple linear regression equations and multiple regression equations are established based on the measurement data and the correlation between stature and measurement variables. The results showed that all skull measurements except for Frontal sinus height were significantly correlated with stature. The standard errors of estimation (SEE) of the multivariate regression equation are smaller than that of the simple linear regression equation, which indicates that the multivariate regression equation is more reliable in estimating the stature of the Chinese population. Compared with other researchers, the regression equation established in this paper is used to estimate the stature of the Chinese population is effective. The results of this study are meaningful to forensics, anthropologists, and archaeologists, and can help them solve complex medical-legal problems. Wen Yang 0003, Guohua Geng, Dengwei Yan, Xiaoning Liu 0001 |
BIBM | 2 |
| 2023 | Gender-Cartoon: Image Cartoonization Method Based on Gender ClassificationabstractQin Opera art is one of China’s intangible cultural heritage, and its influence is gradually declining. The cartoonization of Qin Opera is one of the feasible methods. However, current cartoonization methods suffer from the inability to classify and accurately cartoonize Qinqiang portraits by gender. Therefore, we propose Gender-Cartoon, which can achieve different gender portrait cartoons. The proposed method consists of four modules: gender classification, content feature extraction, gender identification style feature extraction and cartoonization. The gender classification module is used to obtain the gender labels of cartoon images, content feature module extracts content features from portraits by stacking convolutional blocks, and gender identification style extraction module uses the gender labels of cartoon images and image semantic features to obtain the corresponding gender style feature. The cartoonization module fuses content features and style features as the input of the adaptive residual block to obtain the corresponding cartoonization results. The experimental results show that the model can transform both texture and shape on the our collected Qinq Cartoon dataset Face2QinqCartoon and the public dataset Selfie2anime. Long Feng, Guohua Geng, Longquan Yan, Xingrui Ma, Kang Li 0005 |
ICASSP | 2 |
| 2023 | Delta Path Tracing for Real-Time Global Illumination in Mixed RealityabstractVisual coherence between real and virtual objects is important in mixed reality (MR), and illumination consistency is one of the key aspects to achieve coherence. Apart from matching the illumination of the virtual objects with the real environments, the change of illumination on the real scenes produced by the inserted virtual objects should also be considered but is difficult to compute in real-time due to the heavy computation demands of global illumination. In this work, we propose delta path tracing (DPT), which only computes the radiance blocked by the virtual objects from the light sources at the primary hit points of Monte Carlo path tracing, then combines the blocked radiance and multi-bounce indirect illumination with the image of the real scene. Multiple importance sampling (MIS) between BRDF and environment map is performed to handle all-frequency environment maps captured by a panorama camera. Compared to conventional differential rendering methods, our method can remarkably reduce the number of times required to access the environment map and avoid rendering scenes twice. Therefore, the performance can be significantly improved. We implement our method using hardware-accelerated ray tracing on modern GPUs, and the results demonstrate that our method can render global illumination at real-time frame rates and produce plausible visual coherence between real and virtual objects in MR environments. Yang Xu 0092, Yuanfa Jiang, Kang Li 0005, Guohua Geng |
VR | 5 |
| 2023 | EPCS: Endpoint-based part-aware curve skeleton extraction for low-quality point clouds
Guohua Geng |
Comput. Graph. | 3 |
| 2023 | SparseFormer: Sparse transformer network for point cloud classification
Yong Wang 0057, Pengbo Zhou, Guohua Geng, Qi Zhang 0091 |
Comput. Graph. | 4 |
| 2023 | PuzzleFixer: A Visual Reassembly System for Immersive Fragments RestorationabstractWe present PuzzleFixer, an immersive interactive system for experts to rectify defective reassembled 3D objects. Reassembling the fragments of a broken object to restore its original state is the prerequisite of many analytical tasks such as cultural relics analysis and forensics reasoning. While existing computer-aided methods can automatically reassemble fragments, they often derive incorrect objects due to the complex and ambiguous fragment shapes. Thus, experts usually need to refine the object manually. Prior advances in immersive technologies provide benefits for realistic perception and direct interactions to visualize and interact with 3D fragments. However, few studies have investigated the reassembled object refinement. The specific challenges include: 1) the fragment combination set is too large to determine the correct matches, and 2) the geometry of the fragments is too complex to align them properly. To tackle the first challenge, PuzzleFixer leverages dimensionality reduction and clustering techniques, allowing users to review possible match categories, select the matches with reasonable shapes, and drill down to shapes to correct the corresponding faces. For the second challenge, PuzzleFixer embeds the object with node-link networks to augment the perception of match relations. Specifically, it instantly visualizes matches with graph edges and provides force feedback to facilitate the efficiency of alignment interactions. To demonstrate the effectiveness of PuzzleFixer, we conducted an expert evaluation based on two cases on real-world artifacts and collected feedback through post-study interviews. The results suggest that our system is suitable and efficient for experts to refine incorrect reassembled objects. Shuainan Ye, Chen Zhu-Tian, Xiangtong Chu, Kang Li 0005, Juntong Luo 0002, Guohua Geng, Yingcai Wu |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2022 | Precomputed Discrete Visibility Fields for Real-Time Ray-Traced Environment Lighting
Yang Xu 0092, Yuanfa Jiang, Kang Li 0005, Pengbo Zhou, Guohua Geng |
EGSR (ST) | 6 |
| 2022 | A novel compression framework of the dense point-cloud model for cultural heritage artifacts
Kang Li 0005, Jiaojiao Kou, Xiaoxue Chen, Linqi Hai, Guohua Geng, Shunli Zhang 0002 |
Multim. Tools Appl. | 8 |
| 2021 | CFR-GAN: A Generative Model for Craniofacial ReconstructionabstractCraniofacial reconstruction is to reconstruct the face from the skull based on the relationship between the skull and the face to help recognition. This paper proposes a deep generative model for craniofacial reconstruction: CFR-GAN, which avoids the disadvantages of traditional methods of insufficient deep information learning ability of craniofacial data and insufficient ability to express specific features of the dataset. The model is divided into two steps: rough reconstruction and refinement reconstruction. Rough reconstruction rebuilds the overall structural content of the corresponding human head through the skull, and refinement reconstruction restorate facial feature contours. This paper constructs a dataset of 2210 two-dimensional images with craniofacial depth information, which is used to train a CFRGAN model to realize facial reconstruction of skull images. Experiments are conducted from the perspectives of qualitative analysis and quantitative analysis. The results show that CFRGAN generated image retains more identity information, and the similarity between the reconstructed face image and the real face image reaches 94%, which is better than the existing methods. In summary, CFR-GAN proposed in this paper has ability to generate high-fidelity images and is efficient at craniofacial reconstructing. Pengyue Lin, Wen Yang 0003, Siyuan Xia, Xiaoning Liu 0001, Guohua Geng |
BIBM | 6 |
| 2021 | ANINet: a deep neural network for skull ancestry estimationabstractBACKGROUND: Ancestry estimation of skulls is under a wide range of applications in forensic science, anthropology, and facial reconstruction. This study aims to avoid defects in traditional skull ancestry estimation methods, such as time-consuming and labor-intensive manual calibration of feature points, and subjective results. RESULTS: This paper uses the skull depth image as input, based on AlexNet, introduces the Wide module and SE-block to improve the network, designs and proposes ANINet, and realizes the ancestry classification. Such a unified model architecture of ANINet overcomes the subjectivity of manually calibrating feature points, of which the accuracy and efficiency are improved. We use depth projection to obtain the local depth image and the global depth image of the skull, take the skull depth image as the object, use global, local, and local + global methods respectively to experiment on the 95 cases of Han skull and 110 cases of Uyghur skull data sets, and perform cross-validation. The experimental results show that the accuracies of the three methods for skull ancestry estimation reached 98.21%, 98.04% and 99.03%, respectively. Compared with the classic networks AlexNet, Vgg-16, GoogLenet, ResNet-50, DenseNet-121, and SqueezeNet, the network proposed in this paper has the advantages of high accuracy and small parameters; compared with state-of-the-art methods, the method in this paper has a higher learning rate and better ability to estimate. CONCLUSIONS: In summary, skull depth images have an excellent performance in estimation, and ANINet is an effective approach for skull ancestry estimation. Pengyue Lin, Siyuan Xia, Jiang Yi, Wen Yang 0003, Xiaoning Liu 0001, Guohua Geng |
BMC Bioinform. | 6 |
| 2020 | Ancestry Estimation of Skull in Chinese Population Based on Improved Convolutional Neural NetworkabstractThe estimation of ancestry is an essential benchmark for positive identification of heavily decomposed bodies that are recovered in a variety of death and crime scenes. Aiming at the problem of skull ancestry estimation, this paper proposes an improved convolutional neural network method to realize ancestry estimation. We use the six-angle images of the skull as the input of the network. By improving the basic model LeNet5 of the convolutional neural network, we preserve the depth semantics and content information of the image, reduce the number of parameters, and ensure the learning ability of network features. In the experiment, 156 yellow skulls from northern China and 178 white skulls from Xinjiang were used as subjects, 80% of skull samples were used as training sets and 20% as test sets. Experiments on the training set and test set show that the improved CNN network architecture achieves 95.88% accuracy on the training set and 95.52% accuracy on the test set. In addition, we also designed experiments on the contribution of various parts of the skull to ancestor identification. The experimental results show that each region of the skull is useful for ancestor identification, but the effect is different. Compared with other networks, the network structure of this paper has the highest accuracy and better performance. Wen Yang 0003, Pengyue Lin, Guohua Geng, Xiaoning Liu 0001, Kang Li 0005 |
BIBM | 4 |
| 2019 | Fast parallel image reconstruction for cone-beam FDK algorithmabstractSummary FDK algorithm is a popular analytical reconstruction method for practical cone‐beam CT scanners. Compared with iterative methods, the FDK algorithm is computationally efficient. However, the reconstruction speed remains a limitation for its application when dealing with high resolution images. In this paper, we propose a fast method for parallel implementation of the FDK algorithm by the use of multi‐GPU. First, we optimize the backprojection operation of FDK according to the property of geometric symmetry and the correlation between adjacent slices. Then, we utilize the multi‐thread technology to realize the parallel implementation of the optimized FDK algorithm on multi‐GPU. Finally, we implement the proposed method on a multi‐GPU platform. Numerical experiment shows that the proposed multi‐GPU‐based approach can reconstruct a 512 cubed volume in 1.9 seconds from 360 projections of resolution 512 ×512, which is 511 times faster than a traditional CPU–based approach and 5 times faster than a single GPU–based approach. In addition, the reconstruction results also indicate that the proposed method can maintain the same precision with traditional method. Shunli Zhang 0002, Guohua Geng, Jian Zhao 0002 |
Concurr. Comput. Pract. Exp. | 2 |
| 2018 | Rigid blocks matching method based on contour curves and feature regionsabstractThis study proposes a blocks matching method based on contour curves and feature regions that improve the matching precision and speed with which rigid blocks with a specified thickness in point clouds are matched. The method comprises two steps: coarse matching and fine matching. In the coarse matching step, the rigid blocks are first segmented into a series of surfaces and the fracture surfaces are distinguished. Then, the contour curves of the fracture surfaces are extracted using an improved boundary growth method and the rigid blocks are coarsely matched with them. In the fine matching step, feature regions are first extracted from the fracture surfaces. Then, the centroid of each feature region is calculated and the fine matching of rigid blocks with the centroid sets is completed using an improved iterative closest point (ICP) algorithm. The improved ICP algorithm integrates the rotation angle constraint and dynamic iteration coefficient into a probability ICP algorithm, which significantly improves matching precision and speed. Experiments conducted using public blocks and Terracotta Warriors blocks indicate that the proposed method carries out rigid blocks matching more accurately and rapidly than various conventional methods. Fuqun Zhao, Guohua Geng, Lipin Zhu |
IET Comput. Vis. | 3 |
| 2017 | Sex Determination of Incomplete Skull of Han Ethnic in China
Xiaoning Liu 0001, Xiongle Liu, Lipin Zhu, Qianna Zhao, Guohua Geng |
ICIC (3) | 5 |
| 2016 | A statistical approach for extraction of feature lines from point clouds
Guohua Geng, Xiaoran Wei, Shunli Zhang 0002 |
Comput. Graph. | 2 |
| 2015 | Video recommendation based on multi-modal information and multiple kernel
Guohua Geng, Xiaojiang Chen, Pan-Pan Zheng |
Multim. Tools Appl. | 3 |
| 2014 | Multiple instance learning based on positive instance selection and bag structure construction
Guohua Geng, Junli Liang |
Pattern Recognit. Lett. | 2 |
| 2013 | A hierarchical dense deformable model for 3D face reconstruction from skull
Yongli Hu, Fuqing Duan, Zhongke Wu, Guohua Geng |
Multim. Tools Appl. | 7 |
| 2011 | A Spherical Cover Algorithm to Reconstruct 3D Face ModelabstractIn this paper, we propose a face reconstruction algorithm from CT images based on a spherical cover algorithm. First we present an algorithm to simplify the data points of CT data. Then we use the simplified data to reconstruct 3D face model based on a spherical cover algorithm. We define a function to determine the radii of the covering spheres and then connect the centers of the spheres to reconstruct one 3D face model, which is an entry of our face database. Finally Principal Component Analysis is used to generate a statistical face model. By changing the parameters of the statistical model, a specific face can be deformed from the statistical model. The experimental results have showed great effects compared to other algorithm. Guohua Geng |
ICIG | 2 |
| 2010 | Face Appearance Reconstruction Based on a Regional Statistical Craniofacial Model (RCSM)abstractThe reconstruction of facial soft tissue is an essential processing phase in a few of fields. In this paper, we propose a face appearance reconstruction algorithm based on a Regional Statistical Craniofacial model called RSCM. Specifically, the shape of the craniofacial model is decomposed into a few of segments, such as the eyes, the nose and the mouth regions, then the joint statistical models of different regions are constructed independently to address the small sample size problem. The face reconstruction task is formulated as a miss data problem, and is also fulfilled region by region respectively. Finally, the recovered regions are assembled together to achieve a completed face model. The experimental results show that the proposed reconstruction scheme achieves less error rate than a state of the art method. Yan-Fei Zhang, Guohua Geng |
ICPR | 3 |
| 2005 | Content Based Retrieval and Classification of Cultural Relic Images
M. Emre Celebi 0001, Guohua Geng |
ISNN (2) | 3 |