VLDB 2026 Research / reviewers in the wild / expert
Shaofan Wang 0001
dblp:23/8874-1
· DBLP profile ↗
70ranked-venue papers
9as first author
60since 2021 · last 2026
0000-0002-3045-624XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 32 · 7 first-author · 24 since 2021Artificial intelligence and machine learning · 22 · 22 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 1 first-author · 11 since 2021Databases, data management, data science and information retrieval · 6 · 5 since 2021Computer networks · 4 · 1 first-author · 4 since 2021Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Deep Residual Discriminative Dictionary Learning for image classification
Lichun Wang 0002, Jianjia Xin, Kai Xu 0012, Huiyong Zhang, Shaofan Wang 0001, Dehui Kong |
Knowl. Based Syst. | 6 |
| 2026 | Hybrid graph attention learning with pseudo-label guided adaptive evolution
Jinlu Wang, Junbin Gao, Shaofan Wang 0001, Qi Zhang 0095, Yachao Yang, Jipeng Guo 0001 |
Neural Networks | 4 |
| 2026 | Quadruplet Augmentation With Attribute and Structure Invariance for Online Continual LearningabstractOnline Continual Learning (OCL) learns from non-independently and identically distributed streaming data with unknown task boundaries during training and testing. Previous methods suffer from the shortcut feature trap and limited plasticity, leading to two requirements: attribute invariance and structure invariance. The former requires to capture the attributes of objects which maintain invariance during all sessions of OCL, while the latter requires to capture the relation of different attributes during OCL. From the causal invariant representation perspective, we propose Quadruplet Augmentation (QuadAug) by preserving attribute and structure invariance via data and channel augmentation with four types of augmentation strategies. First, we build a fine-grained causal graph of OCL to isolate the session-invariant attributes from confounders. Then, by observing different roles of amplitude and phase components of Fourier domain during knowledge transfer, QuadAug preserves attribute invariance by an Amplitude-Phase augmentation (AP-aug) module via a bidirectional data augmentation strategy, to intervene subtle confounders: the single-session class factor and the class-irrelevant factor. Finally, by decomposing the structure invariance into two necessary conditions: channel independence and channel sufficiency, QuadAug preserves structure invariance by an Independence-Sufficiency augmentation (IS-aug) module, which preserves the channel independence property with an inter-channel discrepancy constraint, and the channel sufficiency property with an adversarial augmentation constraint. QuadAug produces significant improvement on four sequential datasets and three blurry datasets for OCL. Jialu Wu, Shaofan Wang 0001, Boyue Wang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | ComGRL: Comprehensive Graph Representation Learning From Local to Global Bridged by MixupabstractGraph neural networks (GNNs) have demonstrated remarkable effectiveness in various graph representation learning tasks. However, most of existing methods achieve the information extraction by the simple fusion of local and global information, which is a coarse-grained and static way. This may hinder the establishment of a dynamic and collaborative interaction between global and local information, which is crucial for comprehensively understanding graph data. To address this challenge, we propose a novel framework called comprehensive graph representation learning (ComGRL). ComGRL integrates local information into global information to derive powerful representations. It achieves this by implicitly smoothing local information through flexible graph contrastive learning, ensuring reliable representations for subsequent global exploration. Then ComGRL transfers the locally derived representations to a multihead self-attention module, enhancing their discriminative ability by uncovering diverse and rich global correlations. To achieve dynamic transformation between global and local information under self-supervision with pseudo-labels, ComGRL employs a triple sampling strategy to construct mixed node pairs and applies reliable Mixup augmentation across attributes and structure for the main frameworks. This approach broadens the receptive field and facilitates coordination between local and global representation learning, enabling them to reinforce each other. Experimental results across six widely used graph datasets demonstrate that ComGRL achieves excellent performance in node classification tasks. The code could be available athttps://github.com/JinluWang1002/ComGRL. Jinlu Wang, Jiapu Wang, Junbin Gao, Shaofan Wang 0001, Jipeng Guo 0001 |
IEEE Trans. Comput. Soc. Syst. | 5 |
| 2026 | Filtering and Alternating Calibration: Spatiotemporal Context Alternating Fusion for Event-Based Monocular Depth EstimationabstractEvent cameras capture asynchronous pixel-level intensity changes, leading to wide applications to monocular depth estimation under high-speed and low-light environment. Existing event-based depth estimation suffers from two issues: (a) event sparsity and pollution, due to the large amount of spike noise and light sensitivity; (b) spatiotemporal relation deviation, due to the unevenly distribution of spikes along temporal axis. Inspired from the attention calibration in nature language processing, we propose a Filtering-and-Alternating-Calibration network (FAC), using a U-Net architecture with swin transformer based encoders and decoders. The key components of FAC are filtering-based temporal context fusion (FTF) modules and alternating-calibration-based spatiotemporal context fusion (ACSF) modules, serving as the skip connection between pairwise encoded and decoded feature maps. Towards the issue (a), each FTF module learns the cross attention between current encoded feature maps of current event frame and previous event frame, where the latter is filtered with a low-pass filter, reducing the redundancy and irrelevant noise. Towards the issue (b), each ACSF module utilizes the alternating attention calibration between the temporal context fusion map (the output of FTF module) and the decoded feature map, which facilitates the spatial context interaction and calibrates long-range spatiotemporal relation. Thanks to the alternating attention calibration, the encoded and decoded feature maps calibrate each other with motion-corrected temporal contexts and deeper spatial contexts, respectively. Experiments onMVSECandDENSEdatasets show that, FAC outperforms several state-of-the-art depth estimation approaches in terms of several metrics. The code is available at https://github.com/wangsfan/FAC. Shaofan Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2026 | Text-Prompted Prompt Generator with Uncertainty Regularization for Rehearsal-Free Class-Incremental LearningabstractPrompt learning is a kind of popular methods for continual learning, by training a tiny collection of parameters based on frozen networks pre-trained on large-scale datasets, to adapt the model to sequential tasks. Most of prompt methods encounter serious dependence on the delicately designed prompt pools and lead to two shortcomings: prompt inconsistency between training and inference, and prompt selection mismatch during inference. Benefiting from the uniqueness of language description and powerful vision transformers, we propose a T ext- P rompted P rompt Generator Net work (TPPNet), which designs a Text-Prompted Prompt (TPP) generator by encapsulating the pre-trained text embeddings into the visual class token, yielding versatile TPP for resisting the shortcomings of previous methods. Typically, the versatile TPP exhibits three properties: (a) Expressiveness : TPPNet generates prompts with are prompted by text prompts, leveraging the uniqueness of semantics of language and improving the intra-task expressiveness of TPP; (b) Inter-task compatibleness : TPP equips the visual image token with the text embeddings of both old and current tasks and absorbs old knowledge to shrink the prompt inconsistency and improve its anti-forgetting ability; (c) Prompt-query avoidance : TPPNet avoids the prompt query process by generating instance-level prompts and effectively handles the prompt selection mismatch issue during inference. We conduct experiments on four datasets, and the results show that TPPNet outperforms or is comparable with the state-of-the-art-methods for rehearsal-free class-incremental learning tasks. Shaofan Wang 0001, Fuhao Wei |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2025 | Magnetic Framelet-Based Graph Contrastive Learning for Signed-Directed GraphabstractSigned-Directed graphs, known for the ability to express complex relationships, has been the most valuable among several graph types. However, real-world data often contains severe noise and complexity, making it challenging to analyze. Graph contrastive learning, a powerful technique for learning discriminative representations, has become an essential tool for various graph mining tasks. Therefore, we explore graph contrastive learning with signed-directed graphs. To provide multi-scale discriminative representations, we apply framelet transforms to create more intricate filtering representations and propose a Magnetic Framelet-Based Graph Contrastive Learning for Signed-Directed graphs (Framelet-Gcl). This approach offers a robust model for complex graph data analysis by employing structure and Laplacian perturbations to construct node-level contrastive loss. The overall architecture performs framelet-based convolution in both real and complex domains for each view, enhancing the basis for signal processing. Experimental results demonstrate that the proposed approach outperforms existing state-of-the-art methods across various evaluation metrics on four real-world datasets. Yuting Chu, Fujiao Ju, Junbin Gao, Shaofan Wang 0001 |
ICME | 5 |
| 2025 | Uncertainty-guided recurrent prototype distillation for graph few-shot class-incremental learning
Shaofan Wang 0001 |
Multim. Syst. | 2 |
| 2025 | Multi-scale signed graph convolutional network based on framelet
Yuting Chu, Fujiao Ju, Shaofan Wang 0001, Junbin Gao |
Neural Networks | 4 |
| 2025 | BiFormer: A Bipartite-stream Information Fusion framework for large-scale graph representation learning
Qi Zhang 0095, Shaofan Wang 0001, Junbin Gao |
Neural Networks | 3 |
| 2025 | United diverse subgraph for graph incremental learning
Qi Zhang 0095, Shaofan Wang 0001 |
Pattern Recognit. Lett. | 4 |
| 2025 | DGNN: Decoupled Graph Neural Networks With Structural Consistency Between Attribute and Graph Embedding RepresentationsabstractGraph neural networks (GNNs) exhibit a robust capability for representation learning on graphs with complex structures, demonstrating superior performance across various applications. Most existing GNNs utilize graph convolution operations that integrate both attribute and structural information through coupled way. And these GNNs, from an optimization perspective, seek to learn a consensus and compromised embedding representation that balances attribute and graph information, selectively exploring and retaining valid information in essence. To obtain a more comprehensive embedding representation, a novel GNN framework, dubbed Decoupled Graph Neural Networks (DGNN), is introduced. DGNN separately explores distinctive embedding representations from the attribute and graph spaces by decoupled terms. Considering that the semantic graph, derived from attribute feature space, contains different node connection information and provides enhancement for the topological graph, both topological and semantic graphs are integrated by DGNN for powerful embedding representation learning. Further, structural consistency between the attribute embedding and the graph embedding is promoted to effectively eliminate redundant information and establish soft connection. This process involves facilitating factor sharing for adjacency matrices reconstruction, which aims at exploring consensus and high-level correlations. Finally, a more powerful and comprehensive representation is achieved through the concatenation of these embeddings. Experimental results conducted on several graph benchmark datasets demonstrate its superiority in node classification tasks. Jinlu Wang, Jipeng Guo 0001, Junbin Gao, Shaofan Wang 0001, Yachao Yang |
IEEE Trans. Big Data | 5 |
| 2025 | Training Large-Scale Graph Neural Networks via Graph Partial PoolingabstractGraph Neural Networks (GNNs) are powerful tools for graph representation learning, but they face challenges when applied to large-scale graphs due to substantial computational costs and memory requirements. To address scalability limitations, various methods have been proposed, including samplingbased and decoupling-based methods. However, these methods have their limitations: sampling-based methods inevitably discard some link information during the sampling process, while decoupling-based methods require alterations to the model's structure, reducing their adaptability to various GNNs. This paper proposes a novel graph pooling method, Graph Partial Pooling (GPPool), for scaling GNNs to large-scale graphs. GPPool is a versatile and straightforward technique that enhances training efficiency while simultaneously reducing memory requirements. GPPool constructs small-scale pooled graphs by pooling partial nodes into supernodes. Each pooled graph consists of supernodes and unpooled nodes, preserving valuable local and global information. Training GNNs on these graphs reduces memory demands and enhances their performance. Additionally, this paper provides a theoretical analysis of training GNNs using GPPool-constructed graphs from a graph diffusion perspective. It shows that a GNN can be transformed from a large-scale graph into pooled graphs with minimal approximation error. A series of experiments on datasets of varying scales demonstrates the effectiveness of GPPool. Qi Zhang 0095, Shaofan Wang 0001, Junbin Gao, Yongli Hu |
IEEE Trans. Big Data | 3 |
| 2025 | Dual-Domain Division Multiplexer for General Continual Learning: A Pseudo Causal Intervention StrategyabstractAs a continual learning paradigm where non-stationary data arrive in the form of streams and training occurs whenever a small batch of samples is accumulated, general continual learning (GCL) suffers from both inter-task bias and intra-task bias. Existing GCL methods can hardly simultaneously handle two issues since it requires models to avoid from lying into the spurious correlation trap of GCL. From a causal perspective, we formalize a structural causality model of GCL and conclude that spurious correlation exists not only between confounders and input, but also within multiple causal variables. Inspired by frequency transformation techniques which harbor intricate patterns of image comprehension, we propose a plug-and-play module: the Dual-Domain Division Multiplex (D3M) unit, which intervenes confounders and multiple causal factors over frequency and spatial domains with a two-stage pseudo causal intervention strategy. Typically, D3M consists of a frequency division multiplexer (FDM) module and a spatial division multiplexer (SDM) module, each of which prioritizes target-relevant causal features by dividing and multiplexing features over frequency domain and spatial domain, respectively. As a lightweight and model-agonistic unit, D3M can be seamlessly integrated into most current GCL methods. Extensive experiments on four popular datasets demonstrate that D3M significantly enhances accuracy and diminishes catastrophic forgetting compared to current methods. The code is available at https://github.com/wangsfan/D3M. Jialu Wu, Shaofan Wang 0001, Qingming Huang |
IEEE Trans. Image Process. | 2 |
| 2025 | Heterogeneous Pedestrian Simulation in Commercial Complexes: When Attractive Potential Meets Social ForceabstractCommercial complex is a compositional scenario that involves pedestrians with different travel purposes such as commuting, shopping and business. Simulation of pedestrian behavior in a commercial complex is difficult due to the heterogeneity of pedestrians and different attractions from various kinds of shops. Inspired from the law of universal gravitation and the pedestrian dynamics theory, we propose an attractive potential based social force framework for pedestrian simulation in commercial complexes. Our framework consists of an attractive potential model and an attractive potential based social force model. The former model evaluates the attractive force of different types of shops towards heterogeneous pedestrians, by associating the attributes of pedestrians with the types of shops. The latter model effectively evaluates the travel direction of pedestrians by incorporating the attractive force with the social force model via a synthesis force criterion. We conduct the simulation experiment based on 600 groups of pedestrian tracking data collected in the China World Mall, Beijing, a typical commercial complex. The simulation results show that the model can effectively simulate not only the trajectory of real pedestrians, but also the behavior of entering shops. Our research significantly improves the authenticity of pedestrian simulation and provides support for travel behavior modeling of heterogeneous pedestrians in the commercial complex. Jingxuan Peng, Zhonghua Wei, Shaofan Wang 0001, Yongxing Li, Taku Fujiyama |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2025 | GFformer: A Graph Transformer for Extracting All Frequency Information from Large-scale GraphsabstractGraph Transformers have demonstrated outstanding performance across various graph-based applications. Despite their success, applying them to large-scale graphs presents significant scalability challenges, limiting their practical use in industrial environments. Recent studies have attempted to overcome this challenge by focusing on the spatial domain of graphs, leading to the development of various scalable models. However, these approaches neglect the spectral characteristics of graphs, which are crucial for adaptively extracting information from full-frequency bands based on the graph’s inherent properties. As a result, existing scalable Graph Transformers tend to rely heavily on low-frequency features, overlooking valuable mid- and high-frequency information. This article proposes the Graph Filter Transformer (GFformer), a framework designed to effectively extract full-frequency information from large-scale graphs. Unlike existing Graph Transformers, GFformer integrates graph filters into the Transformer architecture, thereby enhancing its ability to model both structural and frequency-related properties. Utilizing the proposed Spectral Token Converter (ST-converter), GFformer generates a unique spectral token sequence for each node by incorporating features from diverse frequencies that act as tokens. This design enables the independent learning of node representations in parallel and supports mini-batch training with flexible batch sizes, making GFformer highly scalable. ST-converter employs spectral graph filters, including low-, mid-, and high-pass filters, to extract features serving as tokens. Consequently, each sequence encompasses features from various frequencies, enabling GFformer to capture comprehensive frequency information effectively. Extensive experiments on datasets of varying scales, including both homophilic and heterophilic graphs, consistently demonstrate that GFformer outperforms existing representative methods. Qi Zhang 0095, Mengmeng Si, Shaofan Wang 0001, Junbin Gao |
ACM Trans. Knowl. Discov. Data | 4 |
| 2025 | Dual-Attention Transformers for Class-Incremental Learning: A Tale of Two MemoriesabstractClass-incremental learning (Class-IL) aims to continuously learn a model from a sequence of tasks, which suffers from the issue of catastrophic forgetting. Recently, a few transformer based methods are proposed to address this issue by transferring self-attention into task-specific attention. However, these methods utilize shared task-specific attention modules across the whole incremental learning process, and are unable to achieve the balance between consolidation and plasticity, i.e., to remember the knowledge learned from previous tasks and absorb the knowledge from the current task simultaneously. Motivated by the mechanism of LSTM and hippocampus memory, we point out that dual attention on long and short-term memories can handle the consolidation-plasticity dilemma of Class-IL. Typically, we propose Dual-Attention Transformers (DAFormer) to learn external attention and internal attention. The former utilizes sample-dependent keys which exclusively focused on the new tasks, while the latter consolidates the knowledge from previous tasks by using sample-agnostic keys. We present two editions of DAFormer: DAFormer-S and DAFormer-M: the former utilizes shared external keys and maintains a small parameter size, while the latter utilizes multiple external keys and enhances the long-term memory. Furthermore, we propose the$K$-nearest neighbor invariant based distillation scheme, which distills knowledge from previous tasks to current task by maintaining the same neighborhood relationship of each sample over old and new models. Experimental results onCIFAR-100,ImageNet-subsetandImageNet-fulldemonstrate that DAFormer significantly outperforms all the state-of-the-art parameter-static and parameter-growing methods. Shaofan Wang 0001, Zhiyong Wang 0001, Boyue Wang |
IEEE Trans. Multim. | 1 |
| 2025 | Redundancy is Not What You Need: An Embedding Fusion Graph Auto-Encoder for Self-Supervised Graph Representation LearningabstractAttribute graphs are a crucial data structure for graph communities. However, the presence of redundancy and noise in the attribute graph can impair the aggregation effect of integrating two different heterogeneous distributions of attribute and structural features, resulting in inconsistent and distorted data that ultimately compromises the accuracy and reliability of attribute graph learning. For instance, redundant or irrelevant attributes can result in overfitting, while noisy attributes can lead to underfitting. Similarly, redundant or noisy structural features can affect the accuracy of graph representations, making it challenging to distinguish between different nodes or communities. To address these issues, we propose the embedded fusion graph auto-encoder framework for self-supervised learning (SSL), which leverages multitask learning to fuse node features across different tasks to reduce redundancy. The embedding fusion graph auto-encoder (EFGAE) framework comprises two phases: pretraining (PT) and downstream task learning (DTL). During the PT phase, EFGAE uses a graph auto-encoder (GAE) based on adversarial contrastive learning to learn structural and attribute embeddings separately and then fuses these embeddings to obtain a representation of the entire graph. During the DTL phase, we introduce an adaptive graph convolutional network (AGCN), which is applied to graph neural network (GNN) classifiers to enhance recognition for downstream tasks. The experimental results demonstrate that our approach outperforms state-of-the-art (SOTA) techniques in terms of accuracy, generalization ability, and robustness. Mengran Li 0001, Yong Zhang 0029, Shaofan Wang 0001, Yongli Hu |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Graph Neural Networks with Soft Association between Topology and AttributeabstractGraph Neural Networks (GNNs) have shown great performance in learning representations for graph-structured data. However, recent studies have found that the interference between topology and attribute can lead to distorted node representations. Most GNNs are designed based on homophily assumptions, thus they cannot be applied to graphs with heterophily. This research critically analyzes the propagation principles of various GNNs and the corresponding challenges from an optimization perspective. A novel GNN called Graph Neural Networks with Soft Association between Topology and Attribute (GNN-SATA) is proposed. Different embeddings are utilized to gain insights into attributes and structures while establishing their interconnections through soft association. Further as integral components of the soft association, a Graph Pruning Module (GPM) and Graph Augmentation Module (GAM) are developed. These modules dynamically remove or add edges to the adjacency relationships to make the model better fit with graphs with homophily or heterophily. Experimental results on homophilic and heterophilic graph datasets convincingly demonstrate that the proposed GNN-SATA effectively captures more accurate adjacency relationships and outperforms state-of-the-art approaches. Especially on the heterophilic graph dataset Squirrel, GNN-SATA achieves a 2.81% improvement in accuracy, utilizing merely 27.19% of the original number of adjacency relationships. Our code is released at https://github.com/wwwfadecom/GNN-SATA. Yachao Yang, Shaofan Wang 0001, Jipeng Guo 0001, Junbin Gao, Fujiao Ju |
AAAI | 3 |
| 2024 | Multi-graph Fusion and Virtual Node Enhanced Graph Neural Networks
Yachao Yang, Jipeng Guo 0001, Shaofan Wang 0001 |
ICANN (5) | 4 |
| 2024 | AutoFGNN: A Framework for Extracting All Frequency Information from Large-Scale GraphsabstractAs a powerful model for deep learning on graph-structured data, the scalability limitation of Graph Neural Networks (GNNs) are receiving increasing attention. To tackle this limitation, two categories of scalable GNNs have been proposed: sampling-based and model simplification methods. However, sampling-based methods suffer from high communication costs and poor performance due to the sampling process. Conversely, existing model simplification methods only rely on parameter-free feature propagation, disregarding its spectral properties. Consequently, these methods can only capture low-frequency information while disregarding valuable middle- and high-frequency information. This paper proposes Automatic Filtering Graph Neural Networks (AutoFGNN), a framework that can extract all frequency information from large-scale graphs. AutoFGNN employs parameter-free low-, middle-, and high-pass filters, which extract the corresponding information for all nodes without introducing parameters. To merge the extracted features, a trainable transformer-based information fusion module is utilized, enabling AutoFGNN to be trained in a mini-batch manner and ensuring scalability for large-scale graphs. Experimental results show that AutoFGNN outperforms existing methods on various scale graphs. Qi Zhang 0095, Jipeng Guo 0001, Shaofan Wang 0001, Junbin Gao |
ICASSP | 4 |
| 2024 | UCloudNet: A Residual U-Net with Deep Supervision for Cloud Image SegmentationabstractRecent advancements in meteorology involve the use of ground-based sky cameras for cloud observation. Analyzing images from these cameras helps in calculating cloud coverage and understanding atmospheric phenomena. Traditionally, cloud image segmentation relied on conventional computer vision techniques. However, with the advent of deep learning, convolutional neural networks (CNNs) are increasingly applied for this purpose. Despite their effectiveness, CNNs often require many epochs to converge, posing challenges for real-time processing in sky camera systems. In this paper, we introduce a residual U-Net with deep supervision for cloud segmentation which provides better accuracy than previous approaches, and with less training consumption. By utilizing residual connection in encoders of UCloudNet, the feature extraction ability is further improved. In the spirit of reproducible research, the model code, dataset, and results of the experiments in this paper are available at: https://github.com/Att100/UCloudNet. Yijie Li 0003, Hewei Wang 0001, Shaofan Wang 0001, Yee Hui Lee, Muhammad Salman Pathan, Soumyabrata Dev |
IGARSS | 3 |
| 2024 | Nia-GNNs: neighbor-imbalanced aware graph neural networks for imbalanced node classification
Shaofan Wang 0001 |
Appl. Intell. | 3 |
| 2024 | A dual attentional skip connection based Swin-UNet for real-time cloud segmentationabstractAbstract Developing real‐time cloud segmentation technology is urgent for many remote sensing based applications such as weather forecasting. Existing deep learning based cloud segmentation methods involve two shortcomings. (a): They tend to produce discontinuous boundaries and fail to capture less salient feature, which corresponds to thin cloud pixels; (b): they are unrobust towards different scenarios. Those issues are circumvented by integrating U‐Net and the swin transformer together, with an efficiently designed dual attention mechanism based skip connection. Typically, a swin transformer based encoder‐decoder network, by incorporating a dual attentional skip connection with Swin‐UNet (DASUNet) is proposed. DASUNet captures the global relationship of image patches based on its window attention mechanism, which fits the real‐time requirement. Moreover, DASUNet characterizes the less salient features by equipping with token dual attention modules among the skip connection, which compensates the ignorance of less salient features incurred from traditional attention mechanism during the stacking of transformer layers. Experiments on ground‐based images ( SWINySeg ) and remote sensing images ( HRC‐WHU , 38‐Cloud ) show that, DASUNet achieves the state‐of‐the‐art or competitive results for cloud segmentation (six top‐1 positions of six metrics among 11 methods on SWINySeg , two top‐1 positions of five metrics among 10 methods on HRC‐WHU , two top‐1 positions of four metrics among 12 methods with ParaNum on 38‐Cloud ), with 100FPS implementation speed averagely for each image. Fuhao Wei, Shaofan Wang 0001 |
IET Image Process. | 2 |
| 2024 | ADOSMNet: a novel visual affordance detection network with object shape mask guided feature encoders
Dongpan Chen, Dehui Kong, Shaofan Wang 0001 |
Multim. Tools Appl. | 4 |
| 2024 | Beyond low-pass filtering on large-scale graphs via Adaptive Filtering Graph Neural Networks
Qi Zhang 0095, Shaofan Wang 0001, Junbin Gao |
Neural Networks | 4 |
| 2024 | Self-Attention Graph Convolution Imputation Network for Spatio-Temporal Traffic DataabstractMissing data in time series is a pervasive problem that serves as obstacles for subsequent traffic data analysis. Consequently, extensive research works have been conducted on traffic missing data imputation tasks. The state-of-the-art traffic data imputation models are mostly based on recurrent neural networks. However, these methods belong to autoregressive models which are highly susceptible to error propagation. The attention-based methods are non-autoregressive models that can avoid compounding errors and help achieve better imputation quality. Moreover, the attention-based methods in now widely applied and have achieved remarkable results, whereas their application on traffic data imputation is still limited. Thus, this paper proposes Self-Attention Graph Convolution Imputation Network (SAGCIN) for spatio-temporal traffic data. To ensure the accuracy of data imputation, it is necessary to fully capture the spatio-temporal contextual information of traffic data to impute missing values. To this end, the SAGCIN model incorporates self-attention mechanism with diffusion graph convolution network. The SAGCIN model consists of two spatio-temporal blocks with a spatio-temporal encoder and an imputation decoder. The encoder learns spatio-temporal representations specialized for traffic data imputation tasks. Based on the learned representation, the decoder performs two stages of imputation operator for missing data. A joint-optimization training approach of imputation and reconstruction is introduced for SAGCIN to perform missing value imputation for traffic data. Empirical results demonstrate that SAGCIN model outperforms state-of-the-art methods in imputation tasks on relevant real-world benchmarks. Xiulan Wei, Yong Zhang 0029, Shaofan Wang 0001, Xia Zhao 0003, Yongli Hu |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | A Dual-Masked Deep Structural Clustering Network With Adaptive Bidirectional Information DeliveryabstractStructured clustering networks, which alleviate the oversmoothing issue by delivering hidden features from autoencoder (AE) to graph convolutional networks (GCNs), involve two shortcomings for the clustering task. For one thing, they used vanilla structure to learn clustering representations without considering feature and structure corruption; for another thing, they exhibit network degradation and vanishing gradient issues after stacking multilayer GCNs. In this article, we propose a clustering method called dual-masked deep structural clustering network (DMDSC) with adaptive bidirectional information delivery (ABID). Specifically, DMDSC enables generative self-supervised learning to mine deeper interstructure and interfeature correlations by simultaneously reconstructing corrupted structures and features. Furthermore, DMDSC develops an ABID module to establish an information transfer channel between each pairwise layer of AE and GCNs to alleviate the oversmoothing and vanishing gradient problems. Numerous experiments on six benchmark datasets have shown that the proposed DMDSC outperforms the most advanced deep clustering algorithms. Yachao Yang, Shaofan Wang 0001, Junbin Gao, Fujiao Ju |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Semi-supervised Video Object Segmentation Via an Edge Attention Gated Graph Convolutional NetworkabstractVideo object segmentation (VOS) exhibits heavy occlusions, large deformation, and severe motion blur. While many remarkable convolutional neural networks are devoted to the VOS task, they often mis-identify background noise as the target or output coarse object boundaries, due to the failure of mining detail information and high-order correlations of pixels within the whole video. In this work, we propose an edge attention gated graph convolutional network (GCN) for VOS. The seed point initialization and graph construction stages construct a spatio-temporal graph of the video by exploring the spatial intra-frame correlation and the temporal inter-frame correlation of superpixels. The node classification stage identifies foreground superpixels by using an edge attention gated GCN which mines higher-order correlations between superpixels and propagates features among different nodes. The segmentation optimization stage optimizes the classification of foreground superpixels and reduces segmentation errors by using a global appearance model which captures the long-term stable feature of objects. In summary, the key contribution of our framework is twofold: (a) the spatio-temporal graph representation can propagate the seed points of the first frame to subsequent frames and facilitate our framework for the semi-supervised VOS task; and (b) the edge attention gated GCN can learn the importance of each node with respect to both the neighboring nodes and the whole task with a small number of layers. Experiments on Davis 2016 and Davis 2017 datasets show that our framework achieves the excellent performance with only small training samples (45 video sequences). Yong Zhang 0029, Shaofan Wang 0001, Yun Liang 0003 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2023 | Learning visual-and-semantic knowledge embedding for zero-shot image classification
Dehui Kong, Xiliang Li, Shaofan Wang 0001 |
Appl. Intell. | 3 |
| 2023 | PASIFTNet: Scale-and-Directional-Aware Semantic Segmentation of Point Clouds
Shaofan Wang 0001, Lichun Wang 0002 |
Comput. Aided Des. | 1 |
| 2023 | SACNet: Shuffling atrous convolutional U-Net for medical image segmentationabstractAbstract Medical images exhibit multi‐granularity and high obscurity along boundaries. As representative work, the U‐Net and its variants exhibit two shortcomings on medical image segmentation: (a) they expand the range of reception fields by applying addition or concatenate operators to features with different reception fields, which disrupts the distribution of the essential feature of objects; (b) they utilize the downsampling or atrous convolution to characterize multi‐granular features of objects, which can obtain a large range of reception fields but leads to blur boundaries of objects. A Shuffling Atrous Convolutional U‐Net (SACNet) for circumventing those issues is proposed. The significant component of SACNet is the Shuffling Atrous Convolution (SAC) module, which fuses different atrous convolutional layers together by using a shuffle concatenate operation, so that the features from the same channel (which correspond to the same attribute of objects) are merged together. Besides the SAC modules, SACNet utilizes an EP module during the fine and medium levels to enhance the boundaries of objects, and utilizes a Transformer module during the coarse level to capture an overall correlation of pixels. Experiments on three medical image segmentation tasks: abdominal organ, cardiac, and skin lesion segmentation demonstrate that, SACNet outperforms several state‐of‐the‐art methods and facilitates easy transplant to other semantic segmentation tasks. Shaofan Wang 0001 |
IET Image Process. | 1 |
| 2023 | CIGNet: Category-and-Intrinsic-Geometry Guided Network for 3D coarse-to-fine reconstruction
Junna Gao, Dehui Kong, Shaofan Wang 0001 |
Neurocomputing | 3 |
| 2023 | Multi-scale latent feature-aware network for logical partition based 3D voxel reconstruction
Dehui Kong, Shaofan Wang 0001, Qianxing Li |
Neurocomputing | 3 |
| 2023 | A tensorial weighted Schatten-p norm model with neighbor regularization for traffic data completion and traffic system correlation exploration
Yongbo Zhao 0003, Shaofan Wang 0001 |
Neurocomputing | 3 |
| 2023 | A subgraph sampling method for training large-scale graph convolutional network
Qi Zhang 0095, Yongli Hu, Shaofan Wang 0001 |
Inf. Sci. | 4 |
| 2023 | Multi-graph Fusion Graph Convolutional Networks with pseudo-label supervision
Yachao Yang, Fujiao Ju, Shaofan Wang 0001, Junbin Gao |
Neural Networks | 4 |
| 2023 | A Survey of Visual Affordance Recognition Based on Deep LearningabstractVisual affordance recognition is an important research topic in robotics, human-computer interaction, and other computer vision tasks. In recent years, deep learning-based affordance recognition methods have achieved remarkable performance. However, there is no unified and intensive survey of these methods up to now. Therefore, this article reviews and investigates existing deep learning-based affordance recognition methods from a comprehensive perspective, hoping to pursue greater acceleration in this research domain. Specifically, this article first classifies affordance recognition into five tasks, delves into the methodologies of each task, and explores their rationales and essential relations. Second, several representative affordance recognition datasets are investigated carefully. Third, based on these datasets, this article provides a comprehensive performance comparison and analysis of the current affordance recognition methods, reporting the results of different methods on the same datasets and the results of each method on different datasets. Finally, this article summarizes the progress of affordance recognition, outlines the existing difficulties and provides corresponding solutions, and discusses its future application trends. Dongpan Chen, Dehui Kong, Shaofan Wang 0001 |
IEEE Trans. Big Data | 4 |
| 2023 | Hierarchical Coupled Discriminative Dictionary Learning for Zero-Shot LearningabstractZero-shot learning (ZSL) aims to recognize images of novel classes, but does not use any images belonging to the novel classes during model training, which is realized by exploiting the auxiliary semantic information. Recently, most ZSL methods focus on learning visual-semantic embeddings to transfer knowledge from the seen classes to the novel classes. Visual-semantic embedding is usually established based on the visual features of images and the semantic information of classes, i.e., class attributes. However, image features are extracted at the individual level, while class attributes are obtained at the group level, so the granularity of these features is different, which makes it difficult to match the two kinds of features. To tackle such problem, we propose hierarchical coupled discriminative dictionary learning (HCDDL) method to hierarchically establish visual-semantic embedding at class-level and image-level with a coarse-to-fine way. Firstly, a class-level coupled dictionary is trained to build basic and coarse-grained connection between visual space and semantic space. Using the class-level coupled dictionary, image attributes are generated. Based on the fine-grained image attributes and images features, an image-level coupled dictionary is learned. In addition, during the learning of hierarchical coupled dictionaries, the discriminative losses are adopted to ensure dictionaries learn more accurate representation, which is beneficial to the recognition task. Recognition of unseen images is performed through searching the class nearest to the unseen image in multiple spaces. Experiments on four widely used benchmark datasets show the effectiveness of the proposed method, and sufficient ablation experiments demonstrate that the coarse-to-fine way leads to good performances. Lichun Wang 0002, Shaofan Wang 0001, Dehui Kong |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | DASI: Learning Domain Adaptive Shape Impression for 3D Object ReconstructionabstractPrevious 3D object reconstruction methods from 2D images involve two issues: the lack of in-depth exploration of the prior knowledge of 3D shapes, and the difficulty of dealing with the serious occluded parts. Inspired by human’s perception on real-world objects which is composed of an overall impression (known asshape impression) and an enhanced cognition, we propose a deep network (denoted by DASI) to learn the Domain Adaptive Shape Impression for 3D reconstruction from arbitrary view images. DASI consists of two modules: shape reconstruction module and shape refinement module. The former module reconstructs a coarse volume by learning a domain adaptive shape impression as embedding in image-based reconstruction. We first leverage 3D objects to learn a shape impression being associated with prior knowledge of 3D objects. To attain consensus on shape impression from 2D images, we regard the 3D shape and the 2D image as two different domains. By adapting the two domains, the shape impression learned from 3D objects is transferred to 2D images and guides the images-based reconstruction. The latter module refines the objects by modeling the whole 3D volume to local 3D patches and exploring their intrinsic geometry relationships. Quantitative and qualitative experimental results on two benchmark datasets demonstrate that DASI outperforms several state-of-the-arts for 3D reconstruction from single and multi-view 2D images. Junna Gao, Dehui Kong, Shaofan Wang 0001 |
IEEE Trans. Multim. | 3 |
| 2022 | SYGNet: A SVD-YOLO based GhostNet for Real-time Driving Scene ParsingabstractIn this paper, we propose SYGNet to strengthen the scene parsing ability of autonomous driving under complicated road conditions. The SYGNet includes feature extraction component and SVD-YOLO GhostNet component. The SVD-YOLO GhostNet component combines Singular Value Decomposition (SVD), You Only Look Once (YOLO) and GhostNet. In the feature extraction component, we propose an algorithm based on VoxelNet to extract point cloud features and image features. In SVD-YOLO GhostNet component, the image data is decomposed by SVD, and we obtain data with stronger spatial and environmental characteristics. YOLOv3 is used to obtain the future map, then convert to GhostNet, which is used to realize the real-time scene parsing. We use KITTI data set to perform our experiments and the results show that the SYGNet is more robust and can further enhance the accuracy of real-time driving scene parsing. The model code, data set, and results of the experiments in this paper are available at: https://github.com/WangHewei16/SYGNet-for-Real-time-Driving-Scene-Parsing. Hewei Wang 0001, Bolun Zhu, Yijie Li 0003, Kaiwen Gong, Ziyuan Wen, Shaofan Wang 0001, Soumyabrata Dev |
ICIP | 6 |
| 2022 | HPGCN: Hierarchical poselet-guided graph convolutional network for 3D pose estimation
Yongpeng Wu 0002, Dehui Kong, Shaofan Wang 0001 |
Neurocomputing | 3 |
| 2022 | Grassmannian graph-attentional landmark selection for domain adaptation
Shaofan Wang 0001, Dehui Kong |
Multim. Tools Appl. | 2 |
| 2022 | Adaptive graph convolutional clustering network with optimal probabilistic graph
Jipeng Guo 0001, Junbin Gao, Shaofan Wang 0001 |
Neural Networks | 5 |
| 2022 | GAN for vision, KG for relation: A two-stage network for zero-shot action recognition
Dehui Kong, Shaofan Wang 0001 |
Pattern Recognit. | 3 |
| 2022 | Real-Time Human Action Recognition Using Locally Aggregated Kinematic-Guided Skeletonlet and Supervised Hashing-by-Analysis Modelabstract3-D action recognition is referred to as the classification of action sequences which consist of 3-D skeleton joints. While many research works are devoted to 3-D action recognition, it mainly suffers from three problems: 1) highly complicated articulation; 2) a great amount of noise; and 3) low implementation efficiency. To tackle all these problems, we propose a real-time 3-D action-recognition framework by integrating the locally aggregated kinematic-guided skeletonlet (LAKS) with a supervised hashing-by-analysis (SHA) model. We first define the skeletonlet as a few combinations of joint offsets grouped in terms of the kinematic principle and then represent an action sequence using LAKS, which consists of a denoising phase and a locally aggregating phase. The denoising phase detects the noisy action data and adjusts it by replacing all the features within it with the features of the corresponding previous frame, while the locally aggregating phase sums the difference between an offset feature of the skeletonlet and its cluster center together over all the offset features of the sequence. Finally, the SHA model combines sparse representation with a hashing model, aiming at promoting the recognition accuracy while maintaining high efficiency. Experimental results on MSRAction3D, UTKinectAction3D, and Florence3DAction datasets demonstrate that the proposed method outperforms state-of-the-art methods in both recognition accuracy and implementation efficiency. Shaofan Wang 0001, Dehui Kong, Lichun Wang 0002 |
IEEE Trans. Cybern. | 2 |
| 2022 | What Size of Aisle Is Necessary? a System Dynamics Model for Mitigating Bottleneck Congestion in Entrance Halls of Metro StationsabstractAs large passenger flow commonly arises in metropolitan metro stations during peak hours, most of bottleneck congestion appears in the entrance halls of metro stations. Previous research work on mitigating metro congestion mostly focus on optimizing single facility and lack a comprehensive analysis on how congestion arises. We propose a system dynamics based framework for optimizing the aisle length and mitigating bottleneck congestion in the entrance halls. We first define the key bottleneck region (KBR) of metro stations consisting of security check area, automatic fare gate area, and the area between them. Then we propose a system dynamics model for simulating passenger flows in KBR by determining the boundary and using a causal analysis of a KBR system. Finally we conduct a case study on the optimization of aisle length, based on the field observation data from three metro stations of Beijing, China. The results reveal that the minimum aisle length of commuter metro stations for acceptable system’s passing efficiency is 6m, and suggest the optimized aisle length for three different commuter metro stations. Our study provides a feasible and effective approach to alleviate passenger flow bottleneck without changing the number of facilities, and can improve the efficiency of metro stations effectively. Jingxuan Peng, Zhonghua Wei, Shi Qiu 0005, Shaofan Wang 0001 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2022 | How Many Facilities are Needed? Evaluating Configurations of Subway Security Check Systems via a Hybrid Queueing ModelabstractSubway security check systems (SSCS), which play an important role for smooth operation of subways of metropolises, exhibit a high complexity and randomness due to their multiple queues, mixed services, and heterogeneous passengers. Traditional approaches for evaluating configurations of SSCS either assume a single-queueing model or ignore various attributes of passengers for the sake of simplicity of models, leading to unsatisfactory results. In this paper, we propose a robust framework for evaluating the performance of configurations of SSCS. By considering passenger attributes (e.g., with or without bags) and special services (e.g., special check, additional search), we regard SSCS as two subsystems each of which is modeled as an iterative probabilistic formula, and introduce a hybrid queueing model to evaluate the waiting time of each passenger in SSCS. The parameters of the model are calibrated and validated using the field observation data collected in 24 subway stations of Beijing, which serve as the input of simulation studies. We conduct two simulation studies to evaluate three indices of SSCS: passengers’ waiting time in SSCS,$K$-systematic time of SSCS, and density flow map, and draw three conclusions. (a) 1X2D, 2X3D, 2X4D configurations are suitable for small, medium and large passenger flows, respectively, where$m\text{X}n\text{D}$denotes the configuration of$m$X-ray machines and$n$detector doors; (b) acceptable efficiency of SSCS is achieved when the ratio between the numbers of X-ray machines and detector doors is in the range of 1:2-1:1, and the efficiency achieves highest at the ratio 1:2; (c) Increasing additional detector doors for no-bag channels can improve the efficiency of SSCS. However, switching existing doors to no-bag channels leads to opposite results. Our findings are beneficial to improving both the capacity and efficiency of SSCS, as well as allocating security check facilities reasonably. Zhonghua Wei, Jingxuan Liang, Shi Qiu 0005, Shaofan Wang 0001 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2022 | A Spatial Relationship Preserving Adversarial Network for 3D Reconstruction from a Single Depth ViewabstractRecovering the geometry of an object from a single depth image is an interesting yet challenging problem. While previous learning based approaches have demonstrated promising performance, they don’t fully explore spatial relationships of objects, which leads to unfaithful and incomplete 3D reconstruction. To address these issues, we propose a Spatial Relationship Preserving Adversarial Network (SRPAN) consisting of 3D Capsule Attention Generative Adversarial Network (3DCAGAN) and 2D Generative Adversarial Network (2DGAN) for coarse-to-fine 3D reconstruction from a single depth view of an object. Firstly, 3DCAGAN predicts the coarse geometry using an encoder-decoder based generator and a discriminator. The generator encodes the input as latent capsules represented as stacked activity vectors with local-to-global relationships (i.e., the contribution of components to the whole shape), and then decodes the capsules by modeling local-to-local relationships (i.e., the relationships among components) in an attention mechanism. Afterwards, 2DGAN refines the local geometry slice-by-slice, by using a generator learning a global structure prior as guidance, and stacked discriminators enforcing local geometric constraints. Experimental results show that SRPAN not only outperforms several state-of-the-art methods by a large margin on both synthetic datasets and real-world datasets, but also reconstructs unseen object categories with a higher accuracy. Dehui Kong, Shaofan Wang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2021 | Latent Feature-Aware and Local Structure-Preserving Network for 3D Completion from a Single Depth View
Dehui Kong, Shaofan Wang 0001 |
ICANN (2) | 3 |
| 2021 | Zero-shot Recognition with Image Attributes Generation using Hierarchical Coupled Dictionary LearningabstractZero-shot learning (ZSL) aims to recognize images from unseen (novel) classes with the training images from seen classes. The attributes of each class is exploited as auxiliary semantic information. Recently most ZSL approaches focus on learning visual-semantic embeddings to transfer knowledge from the seen classes to the unseen classes. However, few works study whether the auxiliary semantic information in the class-level is extensive enough or not for the ZSL task. To tackle such problem, we propose a hierarchical coupled dictionary learning (HCDL) approach to hierarchically align the visual-semantic structures in both the class-level and the image-level. Firstly, the class-level coupled dictionary is trained to establish a basic connection between visual space and semantic space. Then, the image attributes are generated based on the basic connection. Finally, the fine-grained information can be embedded by training the image-level coupled dictionary. Zero-shot recognition is performed in multiple spaces by searching the nearest neighbor class of the unseen image. Experiments on two widely used benchmark datasets show the effectiveness of the proposed approach. Lichun Wang 0002, Shaofan Wang 0001, Dehui Kong |
MMAsia | 3 |
| 2021 | A Local-Global Commutative Preserving Functional Map for Shape CorrespondenceabstractExisting non-rigid shape matching methods mainly involve two disadvantages. (a) Local details and global features of shapes can not be carefully explored. (b) A satisfactory trade-off between the matching accuracy and computational efficiency can be hardly achieved. To address these issues, we propose a local-global commutative preserving functional map (LGCP) for shape correspondence. The core of LGCP involves an intra-segment geometric submodel and a local-global commutative preserving submodel, which accomplishes the segment-to-segment matching and the point-to-point matching tasks, respectively. The first submodel consists of an ICP similarity term and two geometric similarity terms which guarantee the correct correspondence of segments of two shapes, while the second submodel guarantees the bijectivity of the correspondence on both the shape level and the segment level. Experimental results on both segment-to-segment matching and point-to-point matching show that, LGCP not only generate quite accurate matching results, but also exhibit a satisfactory portability and a high efficiency. Qianxing Li, Shaofan Wang 0001, Dehui Kong |
MMAsia | 2 |
| 2021 | Deep3D reconstruction: methods, data, and challengesabstractThree-dimensional (3D) reconstruction of shapes is an important research topic in the fields of computer vision, computer graphics, pattern recognition, and virtual reality. Existing 3D reconstruction methods usually suffer from two bottlenecks: (1) they involve multiple manually designed states which can lead to cumulative errors, but can hardly learn semantic features of 3D shapes automatically; (2) they depend heavily on the content and quality of images, as well as precisely calibrated cameras. As a result, it is difficult to improve the reconstruction accuracy of those methods. 3D reconstruction methods based on deep learning overcome both of these bottlenecks by automatically learning semantic features of 3D shapes from low-quality images using deep networks. However, while these methods have various architectures, in-depth analysis and comparisons of them are unavailable so far. We present a comprehensive survey of 3D reconstruction methods based on deep learning. First, based on different deep learning model architectures, we divide 3D reconstruction methods based on deep learning into four types, recurrent neural network, deep autoencoder, generative adversarial network, and convolutional neural network based methods, and analyze the corresponding methodologies carefully. Second, we investigate four representative databases that are commonly used by the above methods in detail. Third, we give a comprehensive comparison of 3D reconstruction methods based on deep learning, which consists of the results of different methods with respect to the same database, the results of each method with respect to different databases, and the robustness of each method with respect to the number of views. Finally, we discuss future development of 3D reconstruction methods based on deep learning. Dehui Kong, Shaofan Wang 0001, Zhiyong Wang 0001 |
Frontiers Inf. Technol. Electron. Eng. | 3 |
| 2021 | An improved ℓ 1 median model for extracting 3D human body curve-skeleton
Yong Zhang 0029, Lufei Chen, Shaofan Wang 0001 |
Multim. Tools Appl. | 4 |
| 2021 | A Low Rank Dynamic Mode Decomposition Model for Short-Term Traffic Flow PredictionabstractTraffic flow data has three main characteristics: large amount of noise and incompleteness, temporal and spatial correlation, and dynamic sequential property. Problems of noise, loss and incompleteness could decrease the prediction performance and make it difficult for transportation system management. Inspired by recent work on low rank representation (LRR) and dynamic mode decomposition (DMD), we propose a Low Rank Dynamic Mode Decomposition (LRDMD) model which solves the aforementioned problems simultaneously. LRDMD predicts traffic flow by using a state transition matrix which characterizes the relationship between temporally neighboring fragments of traffic flow with low rank regularization. We conduct experiments of traffic flow prediction of different time intervals using loop coil detector data of Qingdao, and the results show that LRDMD outperforms state-of-the-art methods. Yadong Yu, Yong Zhang 0029, Sean Qian, Shaofan Wang 0001, Yongli Hu |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2021 | Joint Transferable Dictionary Learning and View Adaptation for Multi-view Human Action RecognitionabstractMulti-view human action recognition remains a challenging problem due to large view changes. In this article, we propose a transfer learning-based framework called transferable dictionary learning and view adaptation (TDVA) model for multi-view human action recognition. In the transferable dictionary learning phase, TDVA learns a set of view-specific transferable dictionaries enabling the same actions from different views to share the same sparse representations, which can transfer features of actions from different views to an intermediate domain. In the view adaptation phase, TDVA comprehensively analyzes global, local, and individual characteristics of samples, and jointly learns balanced distribution adaptation, locality preservation, and discrimination preservation, aiming at transferring sparse features of actions of different views from the intermediate domain to a common domain. In other words, TDVA progressively bridges the distribution gap among actions from various views by these two phases. Experimental results on IXMAS, ACT4 2 , and NUCLA action datasets demonstrate that TDVA outperforms state-of-the-art methods. Dehui Kong, Shaofan Wang 0001, Lichun Wang 0002 |
ACM Trans. Knowl. Discov. Data | 3 |
| 2021 | DLGAN: Depth-Preserving Latent Generative Adversarial Network for 3D ReconstructionabstractAlthough deep networks based methods outperform traditional 3D reconstruction methods which require multiocular images or class labels to recover the full 3D geometry, they may produce incomplete recovery and unfaithful reconstruction when facing occluded parts of 3D objects. To address these issues, we propose Depth-preserving Latent Generative Adversarial Network (DLGAN) which consists of 3D Encoder-Decoder based GAN (EDGAN, serving as a generator and a discriminator) and Extreme Learning Machine (ELM, serving as a classifier) for 3D reconstruction from a monocular depth image of an object. Firstly, EDGAN decodes a latent vector from the 2.5D voxel grid representation of an input image, and generates the initial 3D occupancy grid under common GAN losses, a latent vector loss and a depth loss. For the latent vector loss, we design 3D deep AutoEncoder (AE) to learn a target latent vector from ground truth 3D voxel grid and utilize the vector to penalize the latent vector encoded from the input 2.5D data. For the depth loss, we utilize the input 2.5D data to penalize the initial 3D voxel grid from 2.5D views. Afterwards, ELM transforms float values of the initial 3D voxel grid to binary values under a binary reconstruction loss. Experimental results show that DLGAN not only outperforms several state-of-the-art methods by a large margin on both a synthetic dataset and a real-world dataset, but also predicts more occluded parts of 3D objects accurately without class labels. Dehui Kong, Shaofan Wang 0001 |
IEEE Trans. Multim. | 3 |
| 2021 | Hardness-Aware Dictionary Learning: Boosting Dictionary for RecognitionabstractSparse representation is a powerful tool in many visual applications since images can be represented effectively and efficiently with a dictionary. Conventional dictionary learning methods usually treat each training sample equally, which would lead to the degradation of recognition performance when the samples from same category distribute dispersedly. This is because the dictionary focuses more on easy samples (known as highly clustered samples), and those hard samples (known as widely distributed samples) are easily ignored. As a result, the test samples which exhibit high dissimilarities to most of intra-category samples tend to be misclassified. To circumvent this issue, this paper proposes a simple and effective hardness-aware dictionary learning (HADL) method, which considers training samples discriminatively based on the AdaBoost mechanism. Different from learning one optimal dictionary, HADL learns a set of dictionaries and corresponding sub-classifiers jointly in an iterative fashion. In each iteration, HADL learns a dictionary and a sub-classifier, and updates the weights based on the classification errors given by current sub-classifier. Those correctly classified samples are assigned with small weights while those incorrectly classified samples are assigned with large weights. Through the iterated learning procedure, the hard samples are associated with different dictionaries. Finally, HADL combines the learned sub-classifiers linearly to form a strong classifier, which improves the overall recognition accuracy effectively. Experiments on well-known benchmarks show that HADL achieves promising classification results. Lichun Wang 0002, Shaofan Wang 0001, Dehui Kong |
IEEE Trans. Multim. | 3 |
| 2021 | 3D human body skeleton extraction from consecutive surfaces using a spatial-temporal consistency model
Yong Zhang 0029, Shaofan Wang 0001 |
Vis. Comput. | 3 |
| 2021 | Discriminative matrix-variate restricted Boltzmann machine classification model
Pengyu Tian, Dehui Kong, Lichun Wang 0002, Shaofan Wang 0001 |
Wirel. Networks | 5 |
| 2020 | Matrix-variate variational auto-encoder with applications to image process
Huixia Yan, Junbin Gao, Dehui Kong, Lichun Wang 0002, Shaofan Wang 0001 |
J. Vis. Commun. Image Represent. | 6 |
| 2020 | An Unsupervised Real-Time Framework of Human Pose Tracking From Range Image SequencesabstractPose tracking from range image sequences remains a difficult task due to strong noise and serious self-occlusion of human body. Existing work either rely on extremely large and precisely annotated datasets, or rely on accurate human mesh model and GPU acceleration. In this paper, we propose an unsupervised real-time framework of pose tracking from range image sequences. Our framework consists of a visible hybrid model (VHM), a componentwise correspondence optimization (CCO) and a dynamic database lookup (DDL). VHM consists of component sphere sets and component visible spherical point sets which exhibits both simplicity and high accuracy. CCO converts the matching between VHM and input point cloud into several subproblems regarding local rotations of components and a global translation of body abdominal joint, each of which has an efficient closed form solution. DDL is designed to recover correct pose when tracking fails, which effectively mitigates accumulative error during tracking. Experiments on SMMC, PDT, EVAL datasets indicate that our framework not only achieves better or competitive precision compared with state-of-the-art methods, but also produces real-time efficiency in personal computers without GPU acceleration. Yongpeng Wu 0002, Dehui Kong, Shaofan Wang 0001 |
IEEE Trans. Multim. | 3 |
| 2019 | 3D human pose estimation from range images with depth difference and geodesic distance
Dehui Kong, Shaofan Wang 0001, Zhiyong Wang 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2019 | Effective human action recognition using global and local offsets of skeleton joints
Dehui Kong, Shaofan Wang 0001, Lichun Wang 0002 |
Multim. Tools Appl. | 3 |
| 2019 | Unsupervised Learning of Human Pose Distance Metric via Sparsity Locality Preserving ProjectionsabstractHuman poses admit complicated articulations and multigranular similarity. Previous works on learning human pose metric utilize sparse models, which concentrate large weights on highly similar poses and fail to depict an overall structure of poses with multigranular similarity. Moreover, previous works require a large number of similar/dissimilar annotated pairwise poses, which is an tedious task and remains inaccurate due to different subjective judgments of experts. Motivated by graph-based neighbor assignment techniques, we propose an unsupervised model called sparsity locality preserving projection with adaptive neighbors (SLPPAN), for learning human pose distance metric. By using a property of the graph Laplacian, SLPPAN introduces a fixed-rank constraint to enforce an adaptive graph structure of poses and learns the neighbor assignment, the similarity measurement, and pose metric simultaneously. Experiments on pose retrieval of the CMU Mocap database demonstrate that SLPPAN outperforms traditional pose metric learning methods by capturing viewpoint variations of human poses. Experiments on keyframe extraction of the MSRAction3D database demonstrate that SLPPAN outperforms current methods by precisely detecting important frames of action sequences. Shaofan Wang 0001, Yongjia Xin, Dehui Kong |
IEEE Trans. Multim. | 1 |
| 2016 | Realistic 3D Mesh Compression Based on Predicted Angle-Normal ImagesabstractIn this paper, we propose angle-normal images to reduce the number of normal component channels from three to two and present predicting the angle-normal images by the reconstructed geometry images. We implement the scheme on realistic meshes. Experimental results verify effectiveness of the proposed scheme. For geometry image codec, the proposed scheme outperforms up 1.78 dB PSNR gains, and average 0.57 dB PSNR gains. For normal image codec, the proposed scheme outperforms up 2.03 dB PSNR gains, and average 1.31 dB PSNR gains. Yunhui Shi, Shaofan Wang 0001, Wenpeng Ding, Jin Wang 0023 |
DCC | 3 |
| 2016 | Extracting hand articulations from monocular depth images using curvature scale space descriptorsabstractWe propose a framework of hand articulation detection from a monocular depth image using curvature scale space (CSS) descriptors. We extract the hand contour from an input depth image, and obtain the fingertips and finger-valleys of the contour using the local extrema of a modified CSS map of the contour. Then we recover the undetected fingertips according to the local change of depths of points in the interior of the contour. Compared with traditional appearance-based approaches using either angle detectors or convex hull detectors, the modified CSS descriptor extracts the fingertips and finger-valleys more precisely since it is more robust to noisy or corrupted data; moreover, the local extrema of depths recover the fingertips of bending fingers well while traditional appearance-based approaches hardly work without matching models of hands. Experimental results show that our method captures the hand articulations more precisely compared with three state-of-the-art appearance-based approaches. Shaofan Wang 0001, Dehui Kong |
Frontiers Inf. Technol. Electron. Eng. | 1 |
| 2016 | Sparse Pose Regression via Componentwise Clustering Feature Point RepresentationabstractWe propose two-dimensional pose estimation from a single range image of the human body, using sparse regression with a componentwise clustering feature point representation (CCFPR) model. CCFPR includes primary feature points and secondary feature points. The primary feature points consist of the torso center and five extremal points of human body, and further serve to classify all body pixels as the points of six body components. The secondary feature points are given by the cluster centers of each of the five components other than the torso, using K-means cluster. The human pose is obtained by learning a sparse projection matrix, which maps CCFPR to the skeleton points of human body, based on the assumption that each skeleton point be represented by a combination of a few feature points of associated body components. Experimental results on both virtual data and real data show that, under the sparse regression model with a suitably selected cluster number, CCFPR outperforms the random decision forest approach and prediction results of Kinect sensor v2 . Dehui Kong, Shaofan Wang 0001 |
IEEE Trans. Multim. | 3 |
| 2015 | Connectivity-preserving geometry images
Shaofan Wang 0001, Dehui Kong, Juan Xue, Weijia Zhu, Hubert Roth |
Vis. Comput. | 1 |
| 2011 | Estimate of The Bézout Number For Linear Piecewise Algebraic Curves Over Arbitrary TriangulationsabstractA piecewise algebraic curve is a curve determined by the zero set of a bivariate spline function. This paper gives an upper bound of the Bézout number, the maximum number of intersections between two linear piecewise algebraic curves whose intersections are finite, over arbitrary triangulations. Shaofan Wang 0001, Renhong Wang 0001, Dehui Kong |
SIAM J. Discret. Math. | 1 |