VLDB 2026 Research / reviewers in the wild / expert
Shunzhou Wang
dblp:208/4915
· DBLP profile ↗
40ranked-venue papers
6as first author
29since 2021 · last 2026
0000-0001-5401-043XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 25 · 3 first-author · 19 since 2021Artificial intelligence and machine learning · 13 · 3 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HDFNet:Hybrid-domain fusion network for medical image restoration
Liqun Lin, Shunzhou Wang, Si Chen 0002, Chao Zeng 0005, Nanfeng Jiang, Dahan Wang |
Expert Syst. Appl. | 3 |
| 2026 | Learning domain-agnostic spatial-angular feature for light field image super-resolution
Shunzhou Wang, Lingwen Xu |
Pattern Recognit. | 3 |
| 2026 | AFH-Net: An adaptive feature harmonization network for document image De-warping
Xinyue Zhou, Nanfeng Jiang, Wang Man, Xu-Yao Zhang, Shunzhou Wang, Dahan Wang |
Pattern Recognit. | 6 |
| 2026 | Com-PCQA: No-Reference Point Cloud Quality Assessment via Complex-Valued Feature LearningabstractThe visual quality of point clouds is critical for perception-centric immersive media. Point Cloud Quality Assessment (PCQA) is crucial for reducing costs associated with human evaluation, optimizing compression pipeline and enhancing human visual perception. However, real-valued PCQA methods often struggle to capture the coupled geometric and perceptual cues that govern quality. Com-PCQA, a novel no-reference PCQA framework leveraging complex-valued feature learning, is proposed. First, a Hilbert dual-stream module transforms multi-modal inputs of point clouds and images into analytic signals in the complex domain, enabling joint modeling of global structure and local texture with efficient tensor operations. Second, a complex amplitude-phase attention (CAPA) module explicitly decomposes and fuses amplitude features that describe geometric structure and phase features that capture fine-grained details, and it can be seamlessly integrated into other PCQA frameworks to enhance performance. Third, an adversarial joint scoring module integrates adversarial training with collaborative learning to calibrate multi-modal, multi-scale representations and enhance robustness. Extensive experiments on three public databases show that Com-PCQA achieves state-of-the-art correlations with subjective scores and consistently outperforms recent PCQA methods, demonstrating its effectiveness and robustness. The code will be available at https://openi.pcl.ac.cn/OpenPointCloud and https://github.com/LareinaSu/Com-PCQA. Jingxuan Su, Ge Li 0002, Shunzhou Wang, Honglei Su, Weisi Lin, Wei Gao 0003 |
IEEE Trans. Image Process. | 3 |
| 2025 | Zero-shot Quantization for Large-kernels via Shape-based Distribution and Diversity Self-distillationabstractZero-shot quantization (ZSQ) has emerged as an effective method to reduce model complexity and memory footprint without using original training data, thereby mitigating data privacy and security concerns during model deployment. Recently, Large-Kernel Convolutional Neural Networks (LKCNNs) have achieved state-of-the-art performance on various vision tasks, which introduce challenges in terms of increased parameters and network complexity, making them difficult to deploy on resource-constrained edge devices. Despite the success of ZSQ, existing methods fail to apply to LKCNNs due to architectural differences such as Batch Normalization (BN) layers in models and thus result in significant performance declines. In this paper, we propose a novel ZSQ framework tailored specifically for LKCNNs, considering their two key characteristics: the large receptive field and the reliance on shape bias. Correspondingly, we first employ an edge detection-based loss to optimize synthetic images that closely mimic the distribution of real images, and a diversity self-distillation loss to maintain consistency in feature representation to enable the generation of synthetic images. Afterward, we use these synthetic images to fine-tune the quantization parameters with a shape-enhance data augmentation strategy. Experiment results demonstrate the superiority of the proposed framework over existing methods, with significant improvements in maintaining accuracy after quantization across various quantization configurations on the ImageNet dataset. Zhuozhen Yu, Xinrui Chen 0001, Shunzhou Wang, Wei Gao 0003 |
ICASSP | 4 |
| 2025 | Zero-shot Quantization of Vision Transformers: Leveraging Multi-model Ensembles and Attention MixupabstractZero-shot quantization (ZSQ) shows promise in compressing and accelerating deep neural networks in scenarios where the original training data is inaccessible. Recently, ZSQ for vision transformers (ViTs) has been proposed to synthesize samples for ViT network quantization, the quality of which significantly impacts the performance of quantized models. Nonetheless, we observe that the synthetic samples produced by current ZSQ techniques exhibit severe bias and insufficient optimization, which deviate from real data and lead to substantial performance declines. On the one hand, the synthetic samples generated by a single ViT network are inaccurate and biased to the specific model. On the other hand, unlike convolutional neural networks (CNNs) with BatchNorm layers, ViTs do not store any training set statistics in the networks, hindering the generation of high-quality calibration samples. To address the above issues, we propose leveraging Multi-model Ensembles and Attention Mixup in ZSQ for ViTs (MMA-ViT). Specifically, MMA-ViT employs an ensemble of diverse pre-trained proxy CNN models to narrow the sample synthesizing space, utilizing their predictive capabilities and BatchNorm statistics to generate exact synthetic images that enhance generality. Additionally, MMA-ViT integrates a unique attention-driven mixup technique for accurate data augmentation during the sample synthesis process, avoiding over-fitting to the networks. The efficacy of MMA-ViT has been demonstrated through extensive experiments and ablation studies on the ImageNet dataset. For example, when Swin-B is quantized to W3/A4, our method achieves a 11.89% top-1 accuracy increase on ImageNet compared to state-of-the-art methods. Xinrui Chen 0001, Zhuozhen Yu, Shunzhou Wang, Wei Gao 0003 |
ICME | 4 |
| 2025 | Contextual and orientation correction modules enhance weakly-supervised aerial object detection in remote sensing images
Le Yang 0008, Shunzhou Wang, Xuerong Wang, Shutong Wang, Binglu Wang |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | Hierarchical spatial-angular integration for lightweight light field image super-resolution
Shunzhou Wang |
Knowl. Based Syst. | 3 |
| 2025 | Adaptive Fusion Learning for Compositional Zero-Shot RecognitionabstractCompositional Zero-Shot Learning (CZSL) aims to learn visual concepts (i.e., attributes and objects) from seen compositions and combine them to predict unseen compositions. Existing visual encoders in CZSL typically use traditional visual encoders (i.e., CNN and Transformer) or image encoders from Visual-Language Models (VLMs) to encode image features. However, traditional visual encoders need more multi-modal textual information, and image encoders of VLMs exhibit dependence on pre-training data, making them less effective when used independently for predicting unseen compositions. To overcome this limitation, we propose a novel approach based on the joint modeling of traditional visual encoders and VLMs visual encoders to enhance the prediction ability for uncommon and unseen compositions. Specifically, we design an adaptive fusion module that automatically adjusts the weighted parameters of similarity scores between traditional and VLMs methods during training, and these weighted parameters are inherited during the inference process. Given the significance of disentangling attributes and objects, we design a Multi-Attribute Object Module that, during the training phase, incorporates multiple pairs of attributes and objects as prior knowledge, leveraging this rich prior knowledge to facilitate the disentanglement of attributes and objects. Building upon this, we select the text encoder from VLMs to construct the Adaptive Fusion Network. We conduct extensive experiments on the Clothing16 K, UT-Zappos50 K, and C-GQA datasets, achieving excellent performance on the Clothing16 K and UT-Zappos50 K datasets. Lingtong Min, Ziman Fan, Shunzhou Wang, Feiyang Dou, Xin Li 0042, Binglu Wang |
IEEE Trans. Multim. | 3 |
| 2025 | Multimodal Large Models are Effective Action AnticipatorsabstractThe task of long-term action anticipation demands solutions that can effectively model temporal dynamics over extended periods while deeply understanding the inherent semantics of actions. Traditional approaches, which primarily rely on recurrent units or Transformer layers to capture long-term dependencies, often fall short in addressing these challenges. Large Language Models (LLMs), with their robust sequential modeling capabilities and extensive commonsense knowledge, present new opportunities for long-term action anticipation. In this work, we introduce the ActionLLM framework, a novel approach that treats video sequences as successive tokens, leveraging LLMs to anticipate future actions. Our baseline model simplifies the LLM architecture by setting future tokens, incorporating an action tuning module, and reducing the textual decoder layer to a linear layer, enabling straightforward action prediction without the need for complex instructions or redundant descriptions. To further harness the commonsense reasoning of LLMs, we predict action categories for observed frames and use sequential textual clues to guide semantic understanding. In addition, we introduce a Cross-Modality Interaction Block, designed to explore the specificity within each modality and capture interactions between vision and textual modalities, thereby enhancing multimodal tuning. Extensive experiments on benchmark datasets demonstrate the superiority of the proposed ActionLLM framework, encouraging a promising direction to explore LLMs in the context of action anticipation. Binglu Wang, Shunzhou Wang, Le Yang 0008 |
IEEE Trans. Multim. | 3 |
| 2025 | Light field angular super-resolution by view-specific queries
Shunzhou Wang, Peiqi Xia, Wei Gao 0003 |
Vis. Comput. | 1 |
| 2024 | Trident Transformer for Light Field Image Super-ResolutionabstractLight Field (LF) image Super-Resolution (SR) requires leveraging the spatial-angular relationship to super-resolve low-resolution LF images into corresponding high-resolution counterparts. Recently, many Transformer-based methods have been proposed for LFSR. However, these methods struggle to recover sharp edges and intricate structures due to the Self-Attention (SA) mechanism’s intrinsic defects of capturing high-frequency information. Additionally, most of them fail to excavate the global spatial-angular information across all views hindered by the expensive computational cost of SA on 4D LF data. To tackle these issues, we introduce Trident Transformer (TriFormer) with three parallel branches: the high-frequency branch, which utilizes convolution and max-pooling for recovering fine-grained textures; the low-frequency branch, which adopts vanilla SA to preserve the low-frequency component; and the interactive-frequency branch, which interacts the frequency information and enhances full-frequency feature, aiding in capturing global information across all angular views. A progressive feature fusion approach is then applied to integrate all distinct information. Experimental results demonstrate our TriFormer’s superiority over leading LFSR methods on five benchmarks, while maintaining a compact model size and computational efficiency. The code is publicly available at https://github.com/wziqi/TriFormer. Shunzhou Wang, Peiqi Xia |
ICME | 3 |
| 2024 | Revisiting Large Kernel Convolution for Light Field Image Angular Super-ResolutionabstractLight field (LF) image angular super-resolution (LFASR) aims to reconstruct densely-sampled LF images from a sparsely-sampled sub-aperture image array. However, mainstream LFASR methods based on convolutional neural networks (CNNs) usually suffer from long-range modeling capability due to the limited receptive field of standard convolutions. This leads to insufficient utilization of non-local spatial and angular information under large disparity situations. More recently, some Transformer-based methods have emerged to tackle this issue and achieve promising performance, however, the self-attention mechanism is suboptimal to recover detailed textures and computationally expensively. To tackle these issues, we propose a novel architecture Large Kernel Convolution Attention (LKCA) which combines large kernel convolutions with the transformer architecture to enhance the ability to capture long-range dependencies of LF images. We design a Muti-scale Hybrid Attention Block (MHAB) based on LKCA to fully utilize non-local spatial-angular information to exploit spatial, angular, and epipolar plane image features. Extensive experiments are carried out to show that our method can achieve more accurate reconstruction effects and generate more evident texture compared with the state-of-the-art methods. Peiqi Xia, Yao Lu 0001, Shunzhou Wang |
ICME | 4 |
| 2024 | Omni Spatial-Angular Correlations Exploration for Light Field Image Super-ResolutionabstractRecently, many deep neural network based methods have been proposed for light field (LF) image super-resolution (SR). Although these methods have shown consistent improvement and yield visually pleasing outcomes, the diverse spatial-angular correlations embedded in light fields are still underexploited, which is crucial for LF image SR. In this paper, we exploit the Omni Spatial-Angular Correlations (OSAC) of LFs. OSAC has three types of correlations: Intra-SAC, Inter-SAC, and Geometry-SAC. Intra-SAC learns separate spatial and angular representations inside a sub-aperture image (SAI) and macro-pixel image (MacPI). Inter-SAC learns complementary contextual information and depth variation information on SAI arrays and macro-pixel arrays. Geometry-SAC learns sub-pixel shift information on epipolar plane images (EPI). To achieve efficient OSAC learning for LF image SR, we develop a network named OSANet equipped with Omni Spatial-Angular Correlations Exploration blocks, which can fully incorporate the comprehensive SACs of LFs. Extensive experiments are carried out on five LF benchmarks, and the results show our methods’ superiority both qualitatively and quantitatively. Code is available at https://github.com/stanley-313/OSANet Shunzhou Wang, Peiqi Xia |
ICME | 3 |
| 2024 | Focal Aggregation Transformer for Light Field Image Super-Resolution
Shunzhou Wang, Yao Lu 0001 |
PRCV (8) | 1 |
| 2024 | Adaptive LPU Decision for Dynamic Point Cloud CompressionabstractWith the rapid development of point cloud applications, dynamic point cloud compression has become a hot topic. A fast and accurate motion estimation scheme is the focus of dynamic point cloud compression, and the size of prediction units (PUs) affects the result of motion estimation. Improper size may even make the performance of inter-frame coding worse than that of intra-frame coding. However, the size of PUs is not fully explored. In this letter, we explore the impact of PUs size on inter-frame coding and propose an adaptive largest prediction unit (LPU) decision strategy. We first downsample original point clouds and obtain the features of adjacent frames. Then, the relationship between the optimal size of LPU and the features of adjacent frames is built. Finally, the optimal size of LPU is used to guide the inter-frame coding. Experimental results show that the better and faster coding performance is achieved by our algorithm, where the bitrate is saved by 2.48%, and the encoding time is saved by 33.80% for point cloud lossless geometry compression. Moreover, our method ensures that inter-frame coding performance of G-PCC is superior to intra-frame coding in all the sequences. Xingming Mu, Wei Gao 0003, Shunzhou Wang, Ge Li 0002 |
IEEE Signal Process. Lett. | 4 |
| 2024 | Scale-Aware Backprojection Transformer for Single Remote Sensing Image Super-ResolutionabstractBackprojection networks have achieved promising super-resolution performance for nature images but not well be explored in the remote sensing image super-resolution (RSISR) field due to the high computation costs. In this article, we propose a scale-aware backprojection Transformer termed SPT for RSISR. SPT incorporates the backprojection learning strategy into a Transformer framework. It consists of scale-aware backprojection-based self-attention layers (SPALs) for scale-aware low-resolution feature learning and scale-aware backprojection-based Transformer blocks (SPTBs) for hierarchical feature learning. A backprojection-based reconstruction module (PRM) is also introduced to enhance the hierarchical features for image reconstruction. SPT stands out by efficiently learning low-resolution features without excessive modules for high-resolution processing, resulting in lower computational resources. Experimental results on UCMerced and AID datasets demonstrate that SPT obtains state-of-the-art results compared to other leading RSISR methods. Jinglei Hao, Wukai Li, Yongqiang Zhao 0001, Shunzhou Wang, Binglu Wang |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | Two-Stage Spatial-Frequency Joint Learning for Large-Factor Remote Sensing Image Super-ResolutionabstractSuper-resolution neural networks have recently achieved great progress in restoring high-quality remote sensing images at low zoom-in magnitude. However, these networks often struggle with challenges like shape distortion and blurring effects due to the severe absence of structure and texture details in large-factor remote sensing image super-resolution. Addressing these challenges, we propose a novel Two-Stage Spatial-Frequency Joint Learning Network (TSFNet). TSFNet innovatively merges insights from both spatial and frequency domains, enabling a progressive refinement of super-resolution results from coarse to fine. Specifically, different from existing frequency feature extraction approaches, we design a novel amplitude-guided-phase adaptive filter module to explicitly disentangle and sequentially recover both the global common image degradation and specific structural degradation in the frequency domain. Additionally, we introduce the cross-stage feature fusion design to enhance feature representation and selectively propagate useful information from stage one to stage two. Quantitative and qualitative experimental results demonstrate that our proposed method surpasses state-of-the-art techniques in large-factor remote sensing image super-resolution. Our code is available at https://github.com/likakakaka/TSFNet_RSISR. Shunzhou Wang, Binglu Wang, Teng Long 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Detail-Preserving Transformer for Light Field Image Super-resolutionabstractRecently, numerous algorithms have been developed to tackle the problem of light field super-resolution (LFSR), i.e., super-resolving low-resolution light fields to gain high-resolution views. Despite delivering encouraging results, these approaches are all convolution-based, and are naturally weak in global relation modeling of sub-aperture images necessarily to characterize the inherent structure of light fields. In this paper, we put forth a novel formulation built upon Transformers, by treating LFSR as a sequence-to-sequence reconstruction task. In particular, our model regards sub-aperture images of each vertical or horizontal angular view as a sequence, and establishes long-range geometric dependencies within each sequence via a spatial-angular locally-enhanced self-attention layer, which maintains the locality of each sub-aperture image as well. Additionally, to better recover image details, we propose a detail-preserving Transformer (termed as DPT), by leveraging gradient maps of light field to guide the sequence learning. DPT consists of two branches, with each associated with a Transformer for learning from an original or gradient image sequence. The two branches are finally fused to obtain comprehensive feature representations for reconstruction. Evaluations are conducted on a number of light field datasets, including real-world scenes and synthetic data. The proposed method achieves superior performance comparing with other state-of-the-art schemes. Our code is publicly available at: https://github.com/BITszwang/DPT. Shunzhou Wang, Tianfei Zhou, Yao Lu 0001, Huijun Di |
AAAI | 1 |
| 2022 | Local-Global Feature Aggregation for Light Field Image Super-ResolutionabstractDeep convolutional neural networks (CNNs) have been widely explored in light field (LF) image super-resolution (SR) to achieve remarkable progress. However, most of the existing CNNs-based methods ignore the similarity of local neighbor views in the 4D LF data. Besides, due to the limitations of CNNs, these methods can’t fully model the global spatial properties of the whole LF images. In this paper, we propose a network with Local-Global Feature Aggregation (LF-LGFA) to handle these problems for LF image SR. Specifically, the Local Aggregation Module is designed to incorporate the local angular information by utilizing the similarity of the local neighbor views’ features in LF images. Moreover, the Global Aggregation Module is designed to capture long-range spatial information via row-wise and column-wise self-attention. Extensive experimental results on five public LF datasets demonstrate that our method achieves comparable results against state-of-the-art techniques. Yao Lu 0001, Shunzhou Wang, Zijian Wang 0007 |
ICASSP | 3 |
| 2022 | Multimodal Unsupervised Image-to-Image Translation Without Independent Style Encoder
Yanbei Sun, Yao Lu 0001, Haowei Lu 0002, Qingjie Zhao, Shunzhou Wang |
MMM (1) | 5 |
| 2022 | Cascade Scale-Aware Distillation Network for Lightweight Remote Sensing Image Super-Resolution
Haowei Ji, Huijun Di, Shunzhou Wang, Qingxuan Shi |
PRCV (4) | 3 |
| 2022 | Distillation Remote Sensing Object Counting via Multi-Scale Context Feature AggregationabstractRemote sensing object counting is an important issue in remote sensing analysis. Remote sensing object counting has many challenges, such as large-scale variations and complex backgrounds. The previous counting methods have many shortboards, such as only focusing on local appearance features of target scenes and ignoring the self-supervision ability of the network itself. To remedy the above problems, in this article, we propose a novel remote sensing object counting method, which contains the adaptive multi-scale context aggregation module (AMCAM) and the self-context distillation module (SCDM). The AMCAM can model and fuse context information from different receptive fields effectively. It also keeps detailed information through multiple pixel attention (PA) modules step by step. The SCDM can improve the representation learning without adding any additional supervision information. SCDM uses feature maps from the deeper layer of the network to supervise feature maps from the earlier layer of the network. Our method has achieved good performance on the remote sensing object counting dataset, RSOC, and mainstream crowd counting datasets, such as ShanghaiTech and UCF-QNRF datasets. Zuodong Duan, Shunzhou Wang, Huijun Di, Jiahao Deng |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Contextual Transformation Network for Lightweight Remote-Sensing Image Super-ResolutionabstractCurrent super-resolution networks typically reduce network parameters and multiadds operations by designing lightweight structures, but lightening the convolution layer is often ignored. In this work, we observe that$3 \times 3$convolutions occupy a high percentage of network parameters in most lightweight super-resolution networks. This motivates us to consider lightening super-resolution networks by replacing$3 \times 3$convolutions with lightweight convolutions, while maintaining the performance. To achieve this, we propose a lightweight convolution layer named contextual transformation layer (CTL). It can yield efficient contextual features through a context feature extraction module and enrich extracted contextual features through a context feature transformation module. Based on CTLs, we build a lightweight super-resolution network called contextual transformation network (CTN) for remote-sensing image super-resolution. Specifically, we use two CTLs to construct a contextual transformation block (CTB) for hierarchical feature learning. Interleaved with a CTB, a context enhancement module (CEM) is employed to enhance the extracted feature representations. All extracted features are processed by a contextual feature aggregation module for final remote-sensing image super-resolution. Extensive experiments are performed on a remote-sensing image super-resolution benchmark named UC Merced. Our method achieves superior results to the other state-of-the-art methods. To demonstrate the generalization ability of our CTL, we extend our CTN to two relevant tasks: natural image super-resolution and natural image denoising. Experimental results on natural image super-resolution benchmarks (i.e., Set5, Set14, B100, Urban100, and Manga109) and natural image denoising benchmarks (i.e., SIDD and DND) further prove the superiority of our method. Our code is publicly available athttps://github.com/BITszwang/CTNet. Shunzhou Wang, Tianfei Zhou, Yao Lu 0001, Huijun Di |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2021 | Learning Multi-level Interaction Relations and Feature Representations for Group Activity Recognition
Yao Lu 0001, Shunzhou Wang |
MMM (1) | 3 |
| 2021 | SA-InterNet: Scale-Aware Interaction Network for Joint Crowd Counting and Localization
Xiuqi Chen, Huijun Di, Shunzhou Wang |
PRCV (1) | 4 |
| 2021 | Scale-Aware Distillation Network for Lightweight Image Super-Resolution
Haowei Lu 0002, Yao Lu 0001, Gongping Li, Yanbei Sun, Shunzhou Wang, Yugang Li |
PRCV (3) | 5 |
| 2021 | LF-MAGNet: Learning Mutual Attention Guidance of Sub-Aperture Images for Light Field Image Super-Resolution
Zijian Wang 0007, Yao Lu 0001, Haowei Lu 0002, Shunzhou Wang, Binglu Wang |
PRCV (3) | 5 |
| 2021 | Single image super-resolution with attention-based densely connected module
Zijian Wang 0007, Yao Lu 0001, Shunzhou Wang, Xuebo Wang, Xiaozhen Chen |
Neurocomputing | 4 |
| 2020 | Motion-Attentive Transition for Zero-Shot Video Object SegmentationabstractIn this paper, we present a novel Motion-Attentive Transition Network (MATNet) for zero-shot video object segmentation, which provides a new way of leveraging motion information to reinforce spatio-temporal object representation. An asymmetric attention block, called Motion-Attentive Transition (MAT), is designed within a two-stream encoder, which transforms appearance features into motion-attentive representations at each convolutional stage. In this way, the encoder becomes deeply interleaved, allowing for closely hierarchical interactions between object motion and appearance. This is superior to the typical two-stream architecture, which treats motion and appearance separately in each stream and often suffers from overfitting to appearance information. Additionally, a bridge network is proposed to obtain a compact, discriminative and scale-sensitive representation for multi-level encoder features, which is further fed into a decoder to achieve segmentation results. Extensive experiments on three challenging public benchmarks (i.e., DAVIS-16, FBMS and Youtube-Objects) show that our model achieves compelling performance against the state-of-the-arts. Code is available at: https://github.com/tfzhou/MATNet. Tianfei Zhou, Shunzhou Wang, Yi Zhou 0007, Yazhou Yao, Jianwu Li, Ling Shao 0001 |
AAAI | 2 |
| 2020 | RSANet: Deep Recurrent Scale-Aware Network for Crowd CountingabstractMost recent works have made significant progress in crowd counting by fusing multi-scale features directly with weighted sum or concatenation to handle large scale variation problems. Meanwhile, there is very little attention paid on the prediction of high-resolution density maps and predicted low-resolution density maps lead to inaccurate counting results. In this paper, we present a novel recurrent scale-aware network(RSANet) to generate a high-resolution density map with scale-aware feature fusion approach. Within this network, we introduce a coarse-to-fine scheme restoring the high-resolution feature map from a low-resolution feature map progressively with stacked dilated convolution blocks. Then, we incorporate recurrent modules to capture dynamic scale-aware information and to benefit the restoration of high-resolution feature maps through multi-scale feature fusion to generate a high-resolution density map. We also use a multi-resolution supervision strategy for training to improve the performance of our network. Extensive experiments on three challenging crowd counting datasets demonstrate the effectiveness of the proposed method. Yujun Xie 0002, Yao Lu 0001, Shunzhou Wang |
ICIP | 3 |
| 2020 | Blind Super-Resolution with Kernel-Aware Feature Refinement
Yao Lu 0001, Gongping Li, Shunzhou Wang, Xuebo Wang, Zijian Wang 0007 |
PRCV (1) | 4 |
| 2020 | SCLNet: Spatial context learning network for congested crowd counting
Shunzhou Wang, Yao Lu 0001, Tianfei Zhou, Huijun Di, Lin Zhang 0033 |
Neurocomputing | 1 |
| 2020 | MATNet: Motion-Attentive Transition Network for Zero-Shot Video Object SegmentationabstractIn this paper, we present a novel end-to-end learning neural network, i.e., MATNet, for zero-shot video object segmentation (ZVOS). Motivated by the human visual attention behavior, MATNet leverages motion cues as a bottom-up signal to guide the perception of object appearance. To achieve this, an asymmetric attention block, named Motion-Attentive Transition (MAT), is proposed within a two-stream encoder network to firstly identify moving regions and then attend appearance learning to capture the full extent of objects. Putting MATs in different convolutional layers, our encoder becomes deeply interleaved, allowing for close hierarchical interactions between object apperance and motion. Such a biologically-inspired design is proven to be superb to conventional two-stream structures, which treat motion and appearance independently in separate streams and often suffer severe overfitting to object appearance. Moreover, we introduce a bridge network to modulate multi-scale spatiotemporal features into more compact, discriminative and scale-sensitive representations, which are subsequently fed into a boundary-aware decoder network to produce accurate segmentation with crisp boundaries. We perform extensive quantitative and qualitative experiments on four challenging public benchmarks, i.e., DAVIS16, DAVIS17, FBMS and YouTube-Objects. Results show that our method achieves compelling performance against current state-of-the-art ZVOS methods. To further demonstrate the generalization ability of our spatiotemporal learning framework, we extend MATNet to another relevant task: dynamic visual attention prediction (DVAP). The experiments on two popular datasets (i.e., Hollywood-2 and UCF-Sports) further verify the superiority of our model. Our implementations have been made publicly available at https://github.com/tfzhou/MATNet. Tianfei Zhou, Jianwu Li, Shunzhou Wang, Ran Tao 0003, Jianbing Shen |
IEEE Trans. Image Process. | 3 |
| 2020 | GAIM: Graph Attention Interaction Model for Collective Activity RecognitionabstractUnbalanced interaction relationships at personal and group levels play a pivotal role in collective activity recognition, which has not been adaptively and jointly explored by previous approaches. In this paper, we propose a graph attention interaction model (GAIM) embedded with the graph attention block (GAB) to explicitly and adaptively infer unbalanced interaction relations at personal and group levels in a unified architecture, and further to learn the spatial and temporal evolutions of the collective activity from these interactions to predict the activity labels. We first design the spatiotemporal graphs tailored to the collective activity where the concurrent person and group nodes, respectively, represent individuals' actions and the collective activity. The graphs provide both spatial structures and semantic appearance features for the collective activity. Then, GAB performs convolution-like filters on the graphs to infer unequal and two-level interaction relations in the collective activity by implementing graph convolutional networks with a shared attention mechanism. At the personal level, the GAB learns different levels of interactions for each person node from its neighbor person nodes under the guidance from the group node. At the group level, the GAB assesses various degrees of interactions to the group node contributed by person nodes. Equipped with the GRUs network, the GAIM learns the spatial and temporal evolutions of individuals' actions as well as the collective activity from the captured interactions, and finally predicts the label of the collective activity. Experiments on four publicly available datasets and ablation studies are conducted to evaluate the performance of our GAIM, and the improved performance demonstrates the effectiveness of our model. Yao Lu 0001, Ruizhe Yu, Huijun Di, Lin Zhang 0033, Shunzhou Wang |
IEEE Trans. Multim. | 6 |
| 2019 | Cross-Domain Scene Text Detection via Pixel and Image-Level Adaptation
Danlu Chen, Yao Lu 0001, Ruizhe Yu, Shunzhou Wang, Lin Zhang 0033, Tingxi Liu |
ICONIP (5) | 5 |
| 2019 | SAF: Semantic Attention Fusion Mechanism for Pedestrian Detection
Ruizhe Yu, Shunzhou Wang, Yao Lu 0001, Huijun Di, Lin Zhang 0033 |
PRICAI (2) | 2 |
| 2019 | Spatio-temporal attention mechanisms based model for collective activity recognition
Huijun Di, Yao Lu 0001, Lin Zhang 0033, Shunzhou Wang |
Signal Process. Image Commun. | 5 |
| 2018 | A two-level attention-based interaction model for multi-person activity recognition
Huijun Di, Yao Lu 0001, Lin Zhang 0033, Shunzhou Wang |
Neurocomputing | 5 |
| 2017 | Improving Deep Crowd Density Estimation via Pre-classification of Density
Shunzhou Wang, Huailin Zhao, Weiren Wang, Huijun Di |
ICONIP (3) | 1 |