EDBT 2026 Demo / reviewers in the wild / expert
Shanshan Gao 0003
dblp:43/5926-3
· DBLP profile ↗
24ranked-venue papers
0as first author
16since 2021 · last 2026
0000-0003-3920-3713ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 9 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Gender-independent kinship verification network via fuzzy disentangling and multi-metric inferenceabstractKinship verification aims to determine whether two individuals share a familial relationship based on facial information. Cross-gender relationships (i.e., Father-Daughter and Mother-Son) continue to face formidable challenges due to the diversity and uncertainty of genetic inheritance. Existing studies primarily focus on extracting robust features and measuring similarity, with limited attention given to the fuzziness of gender differences. To address this issue, this paper proposes a kinship verification framework based on a fuzzy neural network, which adaptively extracts gender-independent kinship features and handles relationship fuzziness to improve cross-gender verification performance. Specifically, the Swin Transformer, which has demonstrated excellent performance in facial analysis, is employed to extract initial features. A fuzzy neural network is then designed to disentangle gender and kinship features, with a gender recognition task introduced to further enhance this disentanglement and improve the gender independence of kinship features. Subsequently, a multi-metric fuzzy reasoning module is adopted to integrate kinship features, extract latent kinship cues, and leverage a contrastive loss function to effectively mine potential negative sample information, thereby significantly enhancing the model's robustness. Experimental results on three publicly available datasets demonstrate that the proposed method achieves state-of-the-art performance. Lei Li 0008, Shanshan Gao 0003, Chaoran Cui, Zhaoqiang Xia |
Neural Networks | 3 |
| 2026 | HDFLStyler: Hierarchical domain-invariant feature learning for source-free domain generalization
Deqian Mao, Shanshan Gao 0003, Faqiang Huang, Caiming Zhang 0001, Yuanfeng Zhou |
Neural Networks | 2 |
| 2026 | Mean teacher based on class prototype contrast for domain adaptive object detection
Fukang Zhang, Shanshan Gao 0003, Honghao Dai, Yuanfeng Zhou |
Neural Networks | 2 |
| 2026 | Face Presentation Attack Detection by Exploiting Prior Knowledge of Region RelationshipsabstractFace recognition systems have been widely deployed in mobile devices for user authentication and payment applications. However, these biometric systems remain vulnerable to face presentation attacks, posing significant security risks. In recent years, numerous countermeasures have been proposed, with analysis of differences between bona fide and attack presentations being a commonly adopted strategy. Nevertheless, the variations in image attributes and region movements have not been thoroughly explored. In attack images, the textures of the facial region and the background tend to be more similar, while local regions often exhibit more consistent directions of movement compared to bona fide presentations. Motivated by this observation, we propose a novel face presentation attack detection method that leverages prior knowledge of region relationships. Specifically, each input face image sequence is first divided into small patches, which are then processed by a pre-trained$TimeSformer$network utilizing divided time and space attention mechanisms to extract deep features. Two metrics—$Cosine$similarity and mean squared error ($MSE$)—are subsequently employed to measure the texture similarity and movement relationships of the regions of interest. During the inference phase, these measurements are fused to distinguish bona fide from attack presentations. Extensive ablation and comparison experiments, conducted on six face presentation attack detection (PAD) databases (i.e., Idiap Replay-Attack, CASIA-MFSD, OULU-NPU, MSU-MFSD, 3DMAD, and HKBU-MARs V1+), demonstrate that our method achieves superior detection performance, significantly improving precision over state-of-the-art approaches in most experimental settings. Lei Li 0008, Shanshan Gao 0003, Zhaoqiang Xia, Fabio Roli, Yuanfeng Zhou |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2025 | An improved algorithm for full-mouth lesion detection based on YOLOv8abstractIn medical imaging detection of oral Cone Beam Computed Tomography (CBCT), there exist tiny lesions that are challenging to detect with low accuracy. The existing detection models are relatively complex. To address this, this paper presents a dual-stage YOLO detection method improved based on YOLOv8. Specifically, we first reconstruct the backbone network based on MobileNetV3 to enhance computational speed and efficiency. Second, we improve detection accuracy from three aspects: we design a composite feature fusion network to enhance the model’s feature extraction capability, addressing the issue of decreased detection accuracy for small lesions due to the loss of shallow information during the fusion process; we further combine spatial and channel information to design the C2f-SCSA module, which delves deeper into the lesion information. To tackle the problem of limited types and insufficient samples of lesions in existing CBCT images, our team collaborated with a professional dental hospital to establish a high-quality dataset, which includes 15 types of lesions and over 2000 accurately labeled oral CBCT images, providing solid data support for model training. Experimental results indicate that the improved method enhances the accuracy of the original algorithm by 3.5 percentage points, increases the recall rate by 4.7 percentage points, and raises the mean Average Precision (mAP) by 3.3 percentage points, a computational load of only 7.6 GFLOPs. This demonstrates a significant advantage in intelligent diagnosis of full-mouth lesions while improving accuracy and reducing computational load. Xinchen Jiao, Shanshan Gao 0003, Faqiang Huang, Wenhan Dou, Yuanfeng Zhou, Caiming Zhang 0001 |
Graph. Model. | 2 |
| 2024 | Low-Quality Ultrasound Image Enhancement Using Spatially Variable Neural Implicit NetworkabstractUltrasound images are widely used because they are easy to use and non-invasive. However, differences in acquisition equipment and frequency settings can result in varying image quality. To address this issue, we propose an ultrasound image enhancement method based on spatially variable neural implicit networks. Our approach consists of several steps: First, the encoder of an arbitrary-scale super-resolution network processes low-quality ultrasound images to generate a feature matrix. Concurrently, the parameter generation network decomposes the low-quality image into a parameter matrix. The pixel prediction module of the arbitrary-scale super-resolution network then uses the parameter matrix to map encoded features to pixel values. To further align the predicted pixel values with the distribution of high-quality ultrasound images and reduce interference from artifacts, we design a spectrum discriminator network based on the fast Fourier transform. This network is highly effective at capturing high-quality information and minimizing the impact of redundant information. In both qualitative and quantitative comparisons with current state-of-the-art image enhancement methods and arbitrary magnification techniques, our approach demonstrates superior performance. Shanshan Gao 0003, Yuanfeng Zhou |
BIBM | 4 |
| 2024 | Face anti-spoofing via jointly modeling local texture and constructed depth
Lei Li 0008, Zhihao Yao 0007, Shanshan Gao 0003, Huijian Han, Zhaoqiang Xia |
Eng. Appl. Artif. Intell. | 3 |
| 2024 | Aggregating Global and Local Representations via Hybrid Transformer for Video DerainingabstractAlthough video deraining technology has achieved great success in recent years, extracting spatiotemporal feature representations across the domains of spatial and temporal in successive frames, then performing spatial and temporal modeling, and restoring high-quality deraining videos with rich details are still challenging tasks. In this paper, we use the hybrid Transformer for the first attempt in video rain removal tasks, and propose a novel video deraining network based on hybrid transformer (VDN-HT) to aggregate global and local representations to accomplish video deraining. In the feature extraction process, we propose to use a U-shaped structure based on serial Transformer blocks to extract shallow local features, deep global features and global dependencies, and then adaptively aggregate them to obtain rainy video features with rain streaks of different directions and densities. In order to better model spatiotemporal relationships, the VDN-HT uses the Transformer’s long-range and relational modeling abilities to obtain the features of spatial and the correlations of temporal between continuous video frames to achieve multi-frame alignment. For ensuring the global-local consistency of the reconstructed frames, we design a global-local reconstruction module composed of Transformer and convolutional neural network (CNN) in parallel to aggregate global and local information to better reconstruct each frame. In addition, the proposed gating-based refinement module and color loss effectively retain the details and color information after removing rain streaks. Extensive experiments on NTURain, RainSynLight25 and RainSynHeavy25 datasets have shown that the VDN-HT can handle many types of rainy videos and perform better than previous methods. Deqian Mao, Shanshan Gao 0003, Honghao Dai, Yunfeng Zhang 0001, Yuanfeng Zhou |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | An Adaptive Sample Assignment Network for Tiny Object DetectionabstractTiny objects often have a small proportion of pixels in the image, leading to significant differences in the number of positive and negative samples and the lack of feature information. Accurately determining the position and category of tiny objects remains a huge challenge for object detection research. Therefore, we design an Adaptive Sample Assignment Strategy(ASAS) and tiny object focusing enhancement module to solve the above two problems. Specifically, starting from the study of positive and negative sample selection and balance strategies for tiny objects, we construct a lightweight Object Existence Probability Determination Network (OEPD/Net) to focus on the areas where tiny objects exist, and achieve adaptive assignment and balance of samples. A top/down, layer by layer focusing enhancement module is designed to effectively enhance the propagation ability of high/level semantic information for tiny objects. The above two solutions have excellent generalization and migration capabilities and can be applied to any stage and two-stage object detection network, effectively enhancing TOD performance. Finally, this article provides a performance analysis of detection performance the detection network based on the OEPD/Net output results, and demonstrates the effectiveness of the proposed OEPD-Net and focusing enhancement module through extensive experiments on a public dataset. Honghao Dai, Shanshan Gao 0003, Deqian Mao, Chenhao Zhang 0001, Yuanfeng Zhou |
IEEE Trans. Multim. | 2 |
| 2024 | Deep Plug-and-Play Non-Iterative Cluster for 3D Global Feature ExtractionabstractEfficient and accurate point cloud feature extraction is crucial for critical tasks such as 3D recognition and semantic segmentation. However, existing global feature extraction methods for 3D data often require designing different models for different input types (point clouds, voxels, and maps). This article proposes an efficient plug-and-play non-iterative clustering method (NICM) to establish a unified point cloud global feature extraction paradigm suitable for any input type to solve the above problems. The core idea of the NICM is to construct the connection between a single point and other global points based only on the cosine similarity between center points to achieve global feature extraction, which has linear complexity characteristics and can be combined with any existing feature extraction model. Additionally, to better integrate the features extracted by NICM and the original model, this article designs an adaptive feature fusion module is designed based on the gate unit, which retains similar features and effectively fuses dissimilar features based on their importance to downstream tasks. We have applied our method to downstream tasks such as point cloud recognition, part segmentation and scene segmentation. Sufficient experiments have proven that our method can provide comprehensive and robust features for the original model, and effectively improve the performance of downstream tasks. Shanshan Gao 0003, Deqian Mao, Shouwen Song, Lei Li 0008, Yuanfeng Zhou |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2023 | Deep Residual Fourier and Self-Attention for Arbitrary Scale MRI Super-ResolutionabstractThe excellent inherent contrast between biological tissues afforded by MR imaging is one of the foremost characteristics of this technique, but it depends on the scan time and hardware devices. Recent studies have shown significant progress in deep learning-based single-image super-resolution (SISR) algorithms based on MR images. Many researchers have employed implicit functions for super-resolution tasks, achieving arbitrary resolution upsampling. Nevertheless, challenges like loss of generated image texture and weak high-frequency features persist. We propose an arbitrary multiple super-resolution model based on deep residual Fourier transform and a selfattention mechanism to tackle these issues. We first use the deep Fourier residual block to construct the high and low-frequency difference of the image, compensating for the spectral bias of the multiple perceptron (MLP) network. Building upon the substantial similarity in medical image tissue structures, we add a vertical and horizontal self-attention mechanism to capture the internal correlation of features in the vertical and horizontal directions. Finally, we learn a continuous functional representation to execute super-resolution tasks at any scale. Experiment results show the effectiveness of our method on the T2 sequence of public dataset SIMON and T1, T2, and Proton Density (PD) weighted scan sequences of public dataset IXI. We conduct qualitative and quantitative comparisons of the three contrasts, illustrating the superiority of our approach. Shanshan Gao 0003, Xingwei Hao, Yuanfeng Zhou |
BIBM | 3 |
| 2023 | Metaballs-Based Real-Time Elastic Object Simulation via Projective Dynamics
Shanshan Gao 0003, Yuanfeng Zhou |
CAD/Graphics | 4 |
| 2022 | Tooth instance segmentation based on capturing dependencies and receptive field adjustment in cone beam computed tomographyabstractAbstract Automatic and accurate instance segmentation of teeth can provide important support for computer‐aided orthodontic work. Traditional methods for tooth segmentation studies often ignore the rich structural features of teeth. Capturing the complete and accurate geometry as well as morphological details of a single tooth remains a challenge for current tooth segmentation studies. In this article, a new tooth segmentation deeplearning network based on capturing dependencies and receptive field adjustment in cone beam computed tomography (CBCT) is proposed to achieve automatic and accurate instance segmentation of dental CBCT data. The method acquires coarse‐level features of tooth and accurate tooth centroids in the first stage, and acquires the instance information and spatial position localization of the tooth. The encoding process in the second stage of the network introduces a guidance module for obtaining tooth geometry information based on a 3D self‐attention mechanism to capture dependencies in CBCT. The proposed tooth feature integration module is based on multiscale fusion of dilated convolutions to capture tooth detailed information at multiple scales, and the network receptive field was adjusted. Extensive evaluation, ablation, and comparison experiments demonstrate that our method exhibits state‐of‐the‐art segmentation performance and accurate instance segmentation results, reflecting their potential applicability in clinical medicine. Wenhan Dou, Shanshan Gao 0003, Deqian Mao, Honghao Dai, Chenhao Zhang 0001, Yuanfeng Zhou |
Comput. Animat. Virtual Worlds | 2 |
| 2022 | DHNet: Salient Object Detection With Dynamic Scale-Aware Learning and Hard-Sample RefinementabstractDuring the annotation procedure of salient object detection, researchers usually locate the approximate location of the salient objects first and then process the pixels that need to be finely annotated. Following this idea, we find that the existing methods have limited exploration for solving the problem of positioning salient objects. Furthermore, no effective solution has been proposed for the hard-sample problem related to this task. Therefore, we propose dynamic scale-aware learning to learn dynamic scale weights that vary with different images to solve the first problem. Second, we design a dense sampling strategy for hard samples to construct a graph representation with samples from different classes and different confidence levels. Then, we achieve targeted feature aggregation based on the constructed graph with the help of the graph attention mechanism. We conduct extensive experiments on five benchmark datasets using comprehensive evaluation metrics. The results show that our method outperforms the current state-of-the-art approaches. Chenhao Zhang 0001, Shanshan Gao 0003, Deqian Mao, Yuanfeng Zhou |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Probability driven approach for point cloud registration of indoor scene
Shanshan Gao 0003, Shi-Qing Xin, Yuanfeng Zhou |
Vis. Comput. | 2 |
| 2021 | Blind motion deblurring via L0 sparse representation
Menghang Li, Shanshan Gao 0003, Chenhao Zhang 0001, Minfeng Xu, Caiming Zhang 0001 |
Comput. Graph. | 2 |
| 2020 | Coarse to Fine: Weak Feature Boosting Network for Salient Object DetectionabstractAbstract Salient object detection is to identify objects or regions with maximum visual recognition in an image, which brings significant help and improvement to many computer visual processing tasks. Although lots of methods have occurred for salient object detection, the problem is still not perfectly solved especially when the background scene is complex or the salient object is small. In this paper, we propose a novel Weak Feature Boosting Network (WFBNet) for the salient object detection task. In the WFBNet, we extract the unpredictable regions (low confidence regions) of the image via a polynomial function and enhance the features of these regions through a well‐designed weak feature boosting module (WFBM). Starting from a coarse saliency map, we gradually refine it according to the boosted features to obtain the final saliency map, and our network does not need any post‐processing step. We conduct extensive experiments on five benchmark datasets using comprehensive evaluation metrics. The results show that our algorithm has considerable advantages over the existing state‐of‐the‐art methods. Chenhao Zhang 0001, Shanshan Gao 0003, Yuanfeng Zhou |
Comput. Graph. Forum | 2 |
| 2020 | Skeletal saliency map computation based on projection symmetry analysis
Shi-Qing Xin, Shanshan Gao 0003, Yuanfeng Zhou |
Graph. Model. | 3 |
| 2018 | 2D skeleton extraction based on heat equation
Fengyi Gao, Guangshun Wei, Shi-Qing Xin, Shanshan Gao 0003, Yuanfeng Zhou |
Comput. Graph. | 4 |
| 2018 | Patch-Based Image Inpainting via Two-Stage Low Rank ApproximationabstractTo recover the corrupted pixels, traditional inpainting methods based on low-rank priors generally need to solve a convex optimization problem by an iterative singular value shrinkage algorithm. In this paper, we propose a simple method for image inpainting using low rank approximation, which avoids the time-consuming iterative shrinkage. Specifically, if similar patches of a corrupted image are identified and reshaped as vectors, then a patch matrix can be constructed by collecting these similar patch-vectors. Due to its columns being highly linearly correlated, this patch matrix is low-rank. Instead of using an iterative singular value shrinkage scheme, the proposed method utilizes low rank approximation with truncated singular values to derive a closed-form estimate for each patch matrix. Depending upon an observation that there exists a distinct gap in the singular spectrum of patch matrix, the rank of each patch matrix is empirically determined by a heuristic procedure. Inspired by the inpainting algorithms with component decomposition, a two-stage low rank approximation (TSLRA) scheme is designed to recover image structures and refine texture details of corrupted images. Experimental results on various inpainting tasks demonstrate that the proposed method is comparable and even superior to some state-of-the-art inpainting algorithms. Qiang Guo 0003, Shanshan Gao 0003, Xiaofeng Zhang 0003, Yilong Yin, Caiming Zhang 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2017 | Interactive facial expression editing based on spatio-temporal coherencyabstractWe present a novel approach for interactively and intuitively editing 3D facial animation in this paper. It determines a new expression by combining the user-specified constraints with the priors contained in a pre-recorded facial expression set, which effectively overcomes the generation of an unnatural expression caused by only user-constraints. The approach is based on the framework of example-based linear interpolation. It adaptively segments the face model into soft regions based on user-interaction. In dependently modeling each region, we propose a new function to estimate the blending weight of each example that matches the user-constraints as well as the spatio-temporal properties of the face set. In blending the regions into a single expression, we present a new criterion that fully exploits the spatial proximity and the spatio-temporal motion consistency over the face set to measure the coherency between vertices and use the coherency to reasonably propagate the influence of each region to the entire face model. Experiments show that our approach, even with inappropriate user’s edits, can create a natural expression that optimally satisfies the user-desired goal. Jing Chi, Shanshan Gao 0003, Caiming Zhang 0001 |
Vis. Comput. | 2 |
| 2015 | A Solution of Collaboration and Interoperability for Networked Enterprises
Shanshan Gao 0003 |
CDVE | 2 |
| 2015 | A region-based expression tracking algorithm for spacetime faces
Jing Chi, Shanshan Gao 0003, Yunfeng Zhang 0001, Caiming Zhang 0001 |
Comput. Graph. | 2 |
| 2007 | A Quasi-Laplacian Smoothing Approach on Arbitrary Triangular MeshesabstractA new method for smoothing triangular meshes is presented. Mean curvature normal is used to define a Quasi-Laplacian for smoothing inner vertices at a local region. Vertices are moved along the normal direction in a more appropriate velocity which can make mesh smoothing and shape preserving harmonizing well. For the boundary vertices, a new method for estimating the mean curvature normal is presented, so that for an arbitrary triangular mesh, the inner and the boundary vertices can be smoothed by the same smoothing process. Features of the original mesh can be preserved by the weighted mean curvature normal restriction of the neighbors of one vertex effectively. Experiments of comparison between the new method and previous methods are included in this paper. Yuanfeng Zhou, Caiming Zhang 0001, Shanshan Gao 0003 |
CAD/Graphics | 3 |