Songtao Wu

dblp:36/1593 · DBLP profile ↗
← Back
22ranked-venue papers
4as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 9 · 7 since 2021Databases, data management, data science and information retrieval · 2Systems, architecture and hardware · 1 · 1 first-authorComputer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 ACID-Style: An Adaptive Condition Injection Diffusion Model for Arbitrary Style Transfer
abstract
Arbitrary style transfer (AST), a popular AI-powered photo editing function, aims to strike an optimal balance between content and style injection from two images in order to generate a novel high-fidelity stylised image. Recently, diffusion models have been applied to AST due to their high generation quality as well as flexibility to embed conditions. However, these models are still not satisfactory and may exhibit inferior performance compared to non-diffusion based methods. This is due to the diffusion process not being purposely designed for AST, leading to suboptimal solutions to trade-off content preservation and style embedding. In this paper, we propose ACID-Style, a novel adaptive condition injection diffusion-based AST framework for improved content/style feature injection to address this research challenge. Using two lightweight adapters, a content and a style injection module, and an adaptive injection mechanism, our approach is able to fully exploit a pre-trained stable diffusion model for AST-specific adaptation and our diffusion model thus learns the most effective timing for content and style injection in the diffusion sampling process. Comprehensive evaluations demonstrate that our method achieves superior style transfer performance, both quantitatively and qualitatively, compared to other state-of-the-art style transfer methods.
Ting Yang 0009, Siyu Yang 0005, Xiyao Liu 0001, Songtao Wu, Gerald Schaefer, Kuanhong Xu, Hui Fang 0003
AAAI4
2026 UniStyleDiff: A unified diffusion-driven framework for image and video style transfer
Siyu Yang 0005, Chunchen Ke, Jian Zhang 0048, Chunwei Miao, Xiyao Liu 0001, Songtao Wu, Kuanhong Xu, Da Huang 0002, Hui Fang 0003
Expert Syst. Appl.6
2025 ProDehaze: Prompting Diffusion Models Toward Faithful Image Dehazing
abstract
Recent approaches using large-scale pretrained diffusion models for image dehazing improve perceptual quality but often suffer from hallucination issues, producing unfaithful dehazed image to the original one. To mitigate this, we propose ProDehaze, a framework that employs internal image priors to direct external priors encoded in pretrained models. We introduce two types of selective internal priors that prompt the model to concentrate on critical image areas: a Structure-Prompted Restorer in the latent space that emphasizes structure-rich regions, and a Haze-Aware Self-Correcting Refiner in the decoding process to align distributions between clearer input regions and the output. Extensive experiments on real-world datasets demonstrate that ProDehaze achieves high-fidelity results in image dehazing, particularly in reducing color shifts. Our code is at https://github.com/TianwenZhou/ProDehaze.
Tianwen Zhou, Songtao Wu, Kuanhong Xu
ICME3
2025 Pseudo-Label Guided Incomplete Partial View-Aligned Clustering
abstract
Addressing the challenges of incomplete and misaligned data in multi-view learning is critical, giventhe inherent uncertainties and complexities of real-world data collection. These challenges often result in significant discrepancies in the quantity, quality, and completeness of data across different views. However, previous research has predominantly focused on addressing either incompleteness or misalignment in isolation. To address both incompleteness and misalignment simultaneously, we propose a novel model, Pseudo-Label Guided Incomplete Partial View-aligned Clustering (PGIPVC). Specifically, A pseudo-label acquisition module based on Cauchy divergence is proposed to preliminarily train the clustering structure of data in a single view, thereby obtaining pseudo-labels for each sample in the view. Subsequently, an incomplete partial alignment clustering module is designed to obtain discriminative latent representations through contrastive learning with positive-negative pairs selected based on KNN and pseudo-labeled samples. Extensive experiments on benchmark datasets demonstrate the superiority of our method compared to other state-of-the-art approaches.
Shubin Ma, Liang Zhao 0005, Songtao Wu, Bo Xu 0008
IEEE Signal Process. Lett.3
2025 Learning Mutual Excitation for Hand-to-Hand and Human-to-Human Interaction Recognition
abstract
Recognizing interactive actions, including hand-to-hand interaction and human-to-human interaction, has attracted increasing attention for various applications in the field of video analysis and human–robot interaction. Considering the success of graph convolution in modeling topology-aware features from skeleton data, recent methods commonly operate graph convolution on separate entities and use late fusion for interactive action recognition, which can barely model the mutual semantic relationships between pairwise entities. To this end, we propose a mutual excitation graph convolutional network (me-GCN) by stacking mutual excitation graph convolution (me-GC) layers. Specifically, me-GC uses a mutual topology excitation module to firstly extract adjacency matrices from individual entities and then adaptively model the mutual constraints between them. Moreover, me-GC extends the above idea and further uses a mutual feature excitation module to extract and merge deep features from pairwise entities. Compared with graph convolution, our proposed me-GC gradually learns mutual information in each layer and each stage of graph convolution operations. Extensive experiments on a challenging hand-to-hand interaction dataset, i.e., the Assembely101 dataset, and two large-scale human-to-human interaction datasets, i.e., NTU60-Interaction and NTU120-Interaction consistently verify the superiority of our proposed method, which outperforms the state-of-the-art GCN-based and Transformer-based methods.
Mengyuan Liu 0001, Chen Chen 0015, Songtao Wu, Fanyang Meng, Hong Liu 0008
IEEE Trans. Hum. Mach. Syst.3
2025 Dynamic Graph Guided Progressive Partial View-Aligned Clustering
abstract
In recent years, there has been a growing focus on multiview data, driven by its rich complementary and consistent information, which has the potential to significantly enhance the performance of downstream tasks. Although many multiview clustering (MVC) methods have achieved promising results by integrating the information of multiple views to learn the consistent representation or consistent graph, these methods typically require complete and entirely accurate correspondences between multiview data, which is challenging to fulfill in practice leading to the problem of partially view-aligned clustering (PVC). To tackle it, we propose a novel method, called dynamic graph guided progressive partial view-aligned clustering (DGPPVC) in this article. To the best of our knowledge, this could be the first work to employ graph convolutional network (GCN) to address the problem of PVC, which explores GCN with dynamic adjacency matrix to reduce unreliable alignments and locate the feature representation with consistent graph structure. In particular, DGPPVC develops an end-to-end framework that encompasses graph construction, feature representation learning, and alignment relationships learning, in which the three parts mutually influence and benefit each other. Moreover, DGPPVC adopts a novel alignment learning strategy that progresses from simplicity to complexity, enabling the step-by-step acquisition of unknown correspondences between different modalities. By giving priority to simple instance pairs, a variant of Jaccard similarities is designed to identify more reliable and complex alignments progressively. During the gradual learning process of alignment relationships, the graph structure matrix is continually and dynamically optimized, thus acquiring a greater variety of graph information between different views. Experiments on several real-world datasets show our promising performance compared with the state-of-the-art methods in partially view-aligned clustering.
Liang Zhao 0005, Qiongjie Xie, Zhengtao Li, Songtao Wu, Yi Yang 0006
IEEE Trans. Neural Networks Learn. Syst.4
2024 Towards Compact Reversible Image Representations for Neural Style Transfer
Xiyao Liu 0001, Siyu Yang 0005, Jian Zhang 0048, Gerald Schaefer, Jiya Li, Xunli Fan, Songtao Wu, Hui Fang 0003
ECCV (66)7
2024 Federated Learning with CSMA Based User Selection for IoT Applications
abstract
User selection has became crucial for improving energy efficiency in communication of federated learning (FL) over wireless networks. However, centralized user selection causes additional system complexity. This study proposes a network intrinsic approach of distributed user selection that leverages the radio resource competition mechanism in random access. Taking the carrier sensing multiple access (CSMA) mechanism as an example of random access, we manipulate the contention window (CW) size to prioritize certain users for obtaining radio resources in each round of training. Training data bias is used as a target scenario for FL with user selection. Prioritization is based on the distance between the newly trained local model and the global model of the previous round. To avoid “excessive contribution” by certain users, a counting mechanism is used to ensure fairness. Simulations with various datasets demonstrate that the proposed method can rapidly achieve convergence similar to that of the centralized user selection approach.
Chen Sun 0006, Shiyao Ma, Songtao Wu, Qiang Tong 0002, Wenqi Zhang 0002
ICC4
2024 CHASE: Learning Convex Hull Adaptive Shift for Skeleton-based Multi-Entity Action Recognition
abstract
Skeleton-based multi-entity action recognition is a challenging task aiming to identify interactive actions or group activities involving multiple diverse entities. Existing models for individuals often fall short in this task due to the inherent distribution discrepancies among entity skeletons, leading to suboptimal backbone optimization. To this end, we introduce a Convex Hull Adaptive Shift based multi-Entity action recognition method (CHASE), which mitigates inter-entity distribution gaps and unbiases subsequent backbones. Specifically, CHASE comprises a learnable parameterized network and an auxiliary objective. The parameterized network achieves plausible, sample-adaptive repositioning of skeleton sequences through two key components. First, the Implicit Convex Hull Constrained Adaptive Shift ensures that the new origin of the coordinate system is within the skeleton convex hull. Second, the Coefficient Learning Block provides a lightweight parameterization of the mapping from skeleton sequences to their specific coefficients in convex combinations. Moreover, to guide the optimization of this network for discrepancy minimization, we propose the Mini-batch Pair-wise Maximum Mean Discrepancy as the additional objective. CHASE operates as a sample-adaptive normalization method to mitigate inter-entity distribution discrepancies, thereby reducing data bias and improving the subsequent classifier's multi-entity action recognition performance. Extensive experiments on six datasets, including NTU Mutual 11/26, H2O, Assembly101, Collective Activity and Volleyball, consistently verify our approach by seamlessly adapting to single-entity backbones and boosting their performance in multi-entity scenarios. Our code is publicly available at https://github.com/Necolizer/CHASE .
Yuhang Wen 0001, Mengyuan Liu 0001, Songtao Wu, Beichen Ding
NeurIPS3
2024 Frequency compensated diffusion model for real-scene dehazing
Songtao Wu, Qiang Tong 0002, Kuanhong Xu
Neural Networks2
2023 Novel Motion Patterns Matter for Practical Skeleton-Based Action Recognition
abstract
Most skeleton-based action recognition methods assume that the same type of action samples in the training set and the test set share similar motion patterns. However, action samples in real scenarios usually contain novel motion patterns which are not involved in the training set. As it is laborious to collect sufficient training samples to enumerate various types of novel motion patterns, this paper presents a practical skeleton-based action recognition task where the training set contains common motion patterns of action samples and the test set contains action samples that suffer from novel motion patterns. For this task, we present a Mask Graph Convolutional Network (Mask-GCN) to focus on learning action-specific skeleton joints that mainly convey action information meanwhile masking action-agnostic skeleton joints that convey rare action information and suffer more from novel motion patterns. Specifically, we design a policy network to learn layer-wise body masks to construct masked adjacency matrices, which guide a GCN-based backbone to learn stable yet informative action features from dynamic graph structure. Extensive experiments on our newly collected dataset verify that Mask-GCN outperforms most GCN-based methods when testing with various novel motion patterns.
Mengyuan Liu 0001, Fanyang Meng, Chen Chen 0001, Songtao Wu
AAAI4
2023 Dynamic Compositional Graph Convolutional Network for Efficient Composite Human Motion Prediction
abstract
With potential applications in fields including intelligent surveillance and human-robot interaction, the human motion prediction task has become a hot research topic and also has achieved high success, especially using the recent Graph Convolutional Network (GCN). Current human motion prediction task usually focuses on predicting human motions for atomic actions. Observing that atomic actions can happen at the same time and thus formulating the composite actions, we propose the composite human motion prediction task. To handle this task, we first present a Composite Action Generation (CAG) module to generate synthetic composite actions for training, thus avoiding the laborious work of collecting composite action samples. Moreover, we alleviate the effect of composite actions on demand for a more complicated model by presenting a Dynamic Compositional Graph Convolutional Network (DC-GCN). Extensive experiments on the Human3.6M dataset and our newly collected CHAMP dataset consistently verify the efficiency of our DC-GCN method, which achieves state-of-the-art motion prediction accuracies and meanwhile needs few extra computational costs than traditional GCN-based human motion methods.
Fanyang Meng, Songtao Wu, Mengyuan Liu 0001
ACM Multimedia4
2021 Rethinking Noise Modeling in Extreme Low-Light Environments
abstract
Recent research has shown Convolutional Neural Networks (CNNs) trained in a fully-supervised fashion achieve promising performance on extreme low-light image denoising task. However, a large amount of "noisy-clean" image pairs are required to train a network, which are difficult to obtain. In this paper, we propose a compact yet effective noise model to generate synthetic noisy images for training. Especially, we address the severe color distortion problem in low-light images by identifying a novel noise component, black calibration error, as its physical origin. We prove that a small error at the sensing stage will strongly affect the following in-camera signal processing (ISP) pipeline and eventually lead to color bias. Experiment results demonstrate that the proposed model is superior in preserving perceptual quality and achieves state-of-the-art performance among existing noise synthesis methods.
Yitong Yu, Songtao Wu, Chang Lei, Kuanhong Xu
ICME3
2020 A Novel Convolutional Neural Network for Image Steganalysis With Shared Normalization
abstract
Image steganalysis is to discriminate innocent images (cover images) and those suspected images (stego images) with hidden messages. The task is challenging since modifications to cover images due to message hiding are extremely small. To handle this difficulty, modern approaches proposed using convolutional neural network (CNN) models to detect steganography with paired learning, i.e., cover images and their stegos are both in training set. In this paper, we explore an important technique in CNN models, the batch normalization (BN), for the task of image steganalysis in the paired learning framework. Our theoretical analysis shows that a CNN model with multiple batch normalization layers is difficult to be generalized to new data in the test set when it is well trained with paired learning. To address this problem, we propose a novel normalization technique called shared normalization (SN) in this paper. Unlike the BN layer utilizing the mini-batch mean and standard deviation to normalize each input batch, SN shares consistent statistics for training samples. Based on the proposed SN layer, we further propose a novel neural network model for image steganalysis. Extensive experiments demonstrate that the proposed network with SN layers is stable and can detect the state-of-the-art steganography with better performances than previous methods.
Songtao Wu, Shenghua Zhong, Yan Liu 0004
IEEE Trans. Multim.1
2019 Joint Dynamic Pose Image and Space Time Reversal for Human Action Recognition from Videos
abstract
Human action recognition aims to classify a given video according to which type of action it contains. Disturbance brought by clutter background and unrelated motions makes the task challenging for video frame-based methods. To solve this problem, this paper takes advantage of pose estimation to enhance the performances of video frame features. First, we present a pose feature called dynamic pose image (DPI), which describes human action as the aggregation of a sequence of joint estimation maps. Different from traditional pose features using sole joints, DPI suffers less from disturbance and provides richer information about human body shape and movements. Second, we present attention-based dynamic texture images (att-DTIs) as pose-guided video frame feature. Specifically, a video is treated as a space-time volume, and DTIs are obtained by observing the volume from different views. To alleviate the effect of disturbance on DTIs, we accumulate joint estimation maps as attention map, and extend DTIs to attention-based DTIs (att-DTIs). Finally, we fuse DPI and att-DTIs with multi-stream deep neural networks and late fusion scheme for action recognition. Experiments on NTU RGB+D, UTD-MHAD, and Penn-Action datasets show the effectiveness of DPI and att-DTIs, as well as the complementary property between them.
Mengyuan Liu 0001, Fanyang Meng, Chen Chen 0001, Songtao Wu
AAAI4
2019 Content-adaptive selective steganographer detection via embedding probability estimation deep networks
Mingjie Zheng 0002, Jianmin Jiang, Songtao Wu, Shenghua Zhong, Yan Liu 0004
Neurocomputing3
2018 Steganographer Detection based on Multiclass Dilated Residual Networks
abstract
Steganographer detection task is to identify criminal users, who attempt to conceal confidential information by steganography methods, among a large number of innocent users. The significant challenge of the task is how to collect the evidences to identify the guilty user with suspicious images, which are embedded with secret messages generating by unknown steganography and payload. Unfortunately, existing methods for steganalysis were served for the binary classification. It makes them harder to classify the images with different kinds of payloads, especially when the payloads of images in test dataset have not been provided in advance. In this paper, we propose a novel steganographer detection method based on multiclass deep neural networks. In the training stage, the networks are trained to classify the images with six types of payloads. The networks can preserve even strengthen the weak stego signals from secret messages in much larger receptive filed by virtue of residual and dilated residual learning. In the inference stage, the learnt model is used to extract the discriminative features, which can capture the difference between guilty users and innocent users. A series of empirical experimental results demonstrate that the proposed method achieves good performance in spatial and frequency domains even though the embedding payload is low. The proposed method achieves a higher level of robustness of inter-steganographic algorithms and can provide a possible solution to address the payload mismatch problem
Mingjie Zheng 0002, Shenghua Zhong, Songtao Wu, Jianmin Jiang
ICMR3
2018 Deep residual learning for image steganalysis
Songtao Wu, Shenghua Zhong, Yan Liu 0004
Multim. Tools Appl.1
2017 Residual convolution network based steganalysis with adaptive content suppression
abstract
Image steganalysis is to discriminate innocent images and those suspected images with hidden messages. In this paper, we propose a unified Convolutional Neural Network (CNN) model for this task. In order to reliably detect modern steganographic algorithms, we design the proposed model from two aspects. For the first, different from existing CNN based steganalytic algorithms that use a predefined highpass kernel to suppress image content, we integrate the highpass filtering operation into the proposed network by building a content suppression subnetwork. For the second, we propose a novel sub-network to actively preserve the weak stego signal generated by secret messages based on residual learning, making the successive network capture the difference between cover images and stego images. Extensive experiments demonstrate that the proposed model can detect states-of-the-art steganography with much lower detection error rates than previous methods.
Songtao Wu, Shenghua Zhong, Yan Liu 0004
ICME1
2017 Steganographer detection via deep residual network
abstract
Steganographer detection problem is to identify culprit actors, who try to hide confidential information with steganography, among many innocent actors. This task has significant challenges, including various embedding steganographic algorithms and payloads, which are usually avoided in steganalysis under laboratory conditions. In this paper, we propose a novel steganographer detection model based on deep residual network. The proposed method strengthens the signal coming from secret messages, which is beneficial for the discrimination between guilty actors and innocent actors. Comprehensive experiments demonstrate that the proposed model achieves very low detection error rates in steganographer detection task. It also outperforms the classical rich model method and other CNN based method. Moreover, the model shows the robustness of inter-steganographic algorithms and inter-payloads.
Mingjie Zheng 0002, Shenghua Zhong, Songtao Wu, Jianmin Jiang
ICME3
2017 Implicit Visual Learning: Image Recognition via Dissipative Learning Model
abstract
According to consciousness involvement, human’s learning can be roughly classified into explicit learning and implicit learning. Contrasting strongly to explicit learning with clear targets and rules, such as our school study of mathematics, learning is implicit when we acquire new information without intending to do so. Research from psychology indicates that implicit learning is ubiquitous in our daily life. Moreover, implicit learning plays an important role in human visual perception. But in the past 60 years, most of the well-known machine-learning models aimed to simulate explicit learning while the work of modeling implicit learning was relatively limited, especially for computer vision applications. This article proposes a novel unsupervised computational model for implicit visual learning by exploring dissipative system, which provides a unifying macroscopic theory to connect biology with physics. We test the proposed Dissipative Implicit Learning Model (DILM) on various datasets. The experiments show that DILM not only provides a good match to human behavior but also improves the explicit machine-learning performance obviously on image classification tasks.
Yan Liu 0004, Yang Liu 0007, Shenghua Zhong, Songtao Wu
ACM Trans. Intell. Syst. Technol.4
2016 Steganalysis via Deep Residual Network
abstract
Recent studies have demonstrated that a well designed deep convolutional neural network (CNN) model achieves competitive performances on detecting the presence of secret message in digital images, compared with the classical rich model based steganalysis. In this paper, we propose to investigate a category of very deep CNN model-the deep residual network (DRN), for steganalysis. DRN is suitable for steganalysis from two aspects. For the first, the DRN model usually contains a large number of network layers, which proves to be effective to capture the complex statistics of digital images. For the second, DRN's residual learning (ResL) method actively strengthens the signal coming from secret messages, which is extremely beneficial for the discrimination between cover images and stego images. Comprehensive experiments on standard dataset show that the DRN model achieves very low detection error rates for the state of arts steganographic algorithms. It also outperforms the classical rich model method and several recently proposed CNN based methods.
Songtao Wu, Shenghua Zhong, Yan Liu 0004
ICPADS1