EDBT 2026 Demo / reviewers in the wild / expert
Wei Wang 0335
dblp:35/7092-335
· DBLP profile ↗
56ranked-venue papers
12as first author
52since 2021 · last 2026
0000-0001-8676-1190ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 30 · 7 first-author · 28 since 2021Graphics, computer vision, multimedia, augmented reality and games · 29 · 6 first-author · 27 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Decision-Driven Orthogonal Learning with Complementary Feature Mining for Robust Synthetic Image DetectionabstractThe widespread and inconsistent compression applied by Online Social Networks severely degrades the performance of synthetic image detectors. We attribute this degradation to two main issues: 1) the model confuses forgery artifacts with compression artifacts, and 2) compression erodes crucial discriminative high-frequency details. Existing methods suppress compression features during training but overlook the overlap between compression features and forgery-related features, leading to the unintended removal of forgery traces. To address artifact confusion, we introduce a Decision-Driven Orthogonal Constraint, which defines a classification decision axis pointing from the real class centroid to the forged class centroid. This constraint enforces compression artifacts to be orthogonal to the decision axis, mitigating their interference with forgery detection without entirely removing them, thus preventing the suppression of forgery-related features. To mitigate the erosion of high-frequency details, we propose to mine complementary forgery cues from both low-frequency information and compressed high-frequency components. A bidirectional update strategy and an adaptive global-local modulator are proposed to facilitate the utilization of forgery cues. Extensive experiments demonstrate that our method achieves state-of-the-art generalization performance in challenging open-world detection scenarios. Wei Wang 0335, Linchao Zhang, Wenqi Ren |
AAAI | 2 |
| 2026 | Neural Discrimination-Prompted Transformers for Efficient UHD Image Restoration and Enhancement
Cong Wang 0018, Jinshan Pan, Wei Wang 0335, Yang Yang 0009 |
Int. J. Comput. Vis. | 4 |
| 2026 | Multimodal backdoor attack on VLMs for autonomous driving via graffiti and cross-lingual triggers
Lidan Liang, Zengzhen Su, Haifeng Xia, Yuan-Ting Yan, Wei Wang 0335 |
Pattern Anal. Appl. | 7 |
| 2026 | Bidomain multi-order modeling for image dehazing
Chenxu Wu, Junling Li, Wei Wang 0335, Wenqi Ren |
Pattern Recognit. | 5 |
| 2026 | Understanding multimodal sentiment with deep modality interaction learning
Jie Mu, Jing Zhang 0037, Zhizheng Sun, Wei Wang 0335 |
Pattern Recognit. | 6 |
| 2026 | Phrase Grounding-Based Style Transfer for Single-Domain Generalized Object DetectionabstractSingle-domain generalized object detection aims to enhance a model’s generalization to multiple unseen target domains using only data from a single source domain during training. This is a practical yet challenging scenario, as it requires the model to address domain shift without incorporating target domain data into the training process. In this paper, we propose a novel phrase-grounding-based style transfer (PGST) approach for the task. Specifically, we first define textual prompts to describe objects for potential unseen target domains. Then, we leverage the grounded language-image pre-training (GLIP) model to capture the styles of these target domains and perform style transfer from the source to the target domains. The style-transferred visual features from the source domain are semantically rich and closely approximate those of their hypothetical counterparts in the target domain. Finally, we employ these style-transferred visual features to fine-tune GLIP. By introducing these imaginary counterparts, the detector can be effectively generalized to unseen target domains using only a single source domain during training. Our method significantly improves mean average precision (mAP), with an average increase of 8.8% across five diverse weather-driving benchmarks. Notably, our approach outperforms or matches the performance of domain-adaptive object detection methods, which require target domain data for training, in several challenging scenarios. Wei Wang 0335, Cong Wang 0018, Mengzhu Wang, Xiang Zhang 0008, Long Lan, Xinwang Liu 0002, Kenli Li 0001, Xiaochun Cao |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | Ultra-High-Definition Image Restoration: New Benchmarks and a Dual Interaction Prior-Driven SolutionabstractUltra-High-Definition (UHD) image restoration has acquired remarkable attention due to its practical demand. In this paper, we construct UHD snow and rain benchmarks, named UHD-Snow and UHD-Rain, to remedy the deficiency in this field. The UHD-Snow/UHD-Rain is established by simulating the physics process of rain/snow into consideration and each benchmark contains 3200 degraded/clear image pairs of 4K resolution. Furthermore, we propose an effective UHD image restoration solution by considering gradient and normal priors in model design, thanks to these priors’ spatial and detail contributions. Specifically, our method contains two branches: (a) feature fusion and reconstruction branch in high-resolution space and (b) prior feature interaction branch in low-resolution space. The former learns high-resolution features and fuses prior-guided low-resolution features to reconstruct clear images, while the latter utilizes normal and gradient priors to mine useful spatial features and detail features to guide high-resolution recovery better. To better utilize these priors, we introduce single prior feature interaction and dual prior feature interaction, where the former respectively fuses normal and gradient priors with high-resolution features to enhance prior ones, while the latter calculates the similarity between enhanced prior ones and further exploits dual guided filtering to boost the feature interaction of dual priors. We conduct experiments on both new and existing public datasets and demonstrate the state-of-the-art performance of our method on UHD image low-light enhancement, dehazing, deblurring, desnowing, and deraining. The source codes and benchmarks are available at https://github.com/wlydlut/UHDDIP. Cong Wang 0018, Jinshan Pan, Xiaofeng Liu 0001, Weixiang Zhou, Xiaoran Sun, Wei Wang 0335, Zhixun Su |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2026 | Noise-Induced Cross-Modal Information Interaction and Dual-Prompt Learning for Medical Image SegmentationabstractAccurate medical image segmentation plays a vital role in clinical diagnostics by facilitating the precise delineation of anatomical structures and pathological regions. However, the performance of existing segmentation methods is often constrained by the scarcity of high-quality annotated datasets, as manual labeling is both labor-intensive and reliant on domain-specific expertise. To address this limitation without requiring additional annotations, we propose a novel multimodal segmentation framework that leverages medical text annotations as an auxiliary modality to complement visual information. In particular, our approach introduces a learnable encoding strategy for joint distribution modeling of image and text, which enables discriminative fusion and effectively suppresses cross-modal redundancy. Moreover, we innovatively design a frequency-domain prompt encoder based on the discrete wavelet transform (DWT) to capture multi-frequency features, thereby significantly enhancing the model's ability to delineate fine-grained boundaries. Overall, our framework integrates cross-attention for effective cross-modal interaction, employs joint distribution modeling to enable discriminative and redundancy-reduced multimodal fusion, and incorporates auxiliary supervision to strengthen the learning of task-relevant features. Extensive experiments on nine public datasets across three clinical tasks-including cell, lung infection, and polyp segmentation-demonstrate that our method achieves competitive segmentation performance while maintaining favorable computational efficiency. Comprehensive ablation studies and feature distribution visualizations further validate the effectiveness and robustness of our proposed components. The code will be made publicly available at https://github.com/chenpeng052/MDFP. Chao Huang 0008, Jie Wen 0001, Wei Wang 0335, Li Shen 0008, Wenqi Ren, Xiaochun Cao, Chengliang Liu 0003 |
IEEE Trans. Image Process. | 4 |
| 2026 | Vision Enhancing LLMs: Empowering Multimodal Knowledge Storage and Sharing in LLMsabstractRecent advancements in multimodal large language models (MLLMs) have achieved significant multimodal generation capabilities, akin to GPT-4. These models predominantly map visual information into language representation space, leveraging the vast knowledge and powerful text generation abilities of LLMs to produce multimodal instruction-following responses. We could term this method as LLMs for Vision because of its employing LLMs for visual understanding and reasoning, yet observe that these MLLMs neglect the potential of harnessing visual knowledge to enhance the overall capabilities of LLMs, which could be regarded as Vision Enhancing LLMs. In this paper, we propose an approach called MKS2, aimed at enhancing LLMs through empowering Multimodal Knowledge Storage and Sharing in LLMs. Specifically, we introduce Modular Visual Memory (MVM), a component integrated into the internal blocks of LLMs, designed to store open-world visual information efficiently. Additionally, we present a soft Mixture of Multimodal Experts (MoMEs) architecture in LLMs to invoke multimodal knowledge collaboration during text generation. Our comprehensive experiments demonstrate that MKS2 substantially augments the reasoning capabilities of LLMs in contexts necessitating physical or commonsense knowledge. It also delivers competitive results on image-text understanding multimodal benchmarks. The codes will be available at: https://github.com/HITsz-TMG/MKS2-Multimodal-Knowledge-Storage-and-Sharing. Yunxin Li, Baotian Hu, Wei Wang 0335, Xiaochun Cao, Min Zhang 0005 |
IEEE Trans. Image Process. | 4 |
| 2026 | CLIP-SENet: CLIP-Based Semantic Enhancement Network for Vehicle Re-IdentificationabstractVehicle re-identification (Re-ID) is a crucial task in intelligent transportation systems (ITS), aimed at retrieving and matching the same vehicle across different surveillance cameras. Numerous studies have explored methods to enhance vehicle Re-ID by focusing on semantic enhancement. However, these methods often rely on additional annotated information to enable models to extract effective semantic features, which brings many limitations. In this work, we propose a CLIP-based Semantic Enhancement Network (CLIP-SENet), an end-to-end framework designed to autonomously extract and refine vehicle semantic attributes, facilitating the generation of more robust semantic feature representations. Inspired by zero-shot solutions for downstream tasks presented by large-scale vision-language models, we leverage the powerful cross-modal descriptive capabilities of the CLIP image encoder to initially extract general semantic information. Instead of using a text encoder for semantic alignment, we design an adaptive fine-grained enhancement module (AFEM) to adaptively enhance this general semantic information at a fine-grained level to obtain robust semantic feature representations. These features are then fused with common Re-ID appearance features to further refine the distinctions between vehicles. Our comprehensive evaluation on three benchmark datasets demonstrates the effectiveness of CLIP-SENet. Our approach achieves new state-of-the-art performance, with 92.9% mAP and 98.7% Rank-1 on VeRi-776 dataset, 90.4% Rank-1 and 98.7% Rank-5 on VehicleID dataset, and 89.1% mAP and 97.9% Rank-1 on the more challenging VeRi-Wild dataset. Duanfeng Chu, Wei Wang 0335, Bingrong Xu |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2025 | Critical Forgetting-Based Multi-Scale Disentanglement for Deepfake DetectionabstractRecent face forgery detection methods based on disentangled representation learning utilize paired images for cross-reconstruction, aiming to extract forgery-relevant attributes and forgery-irrelevant content. However, there still exist the following issues that may comprise the detector performance: 1) using information-dense images as the decoupling targets increases the decoupling difficulty; 2) the extracted attribute features are reconstruction-irrelevant rather than forgery-relevant, and single-scale forgery representation decoupling cannot capture sufficient discriminative information; 3) the generalization performance of decoupled attribute features is poor as the detector focuses on learning specific artifact types in the training set. To address these issues, we propose a novel disentangled representation learning framework for deepfake detection. First, we extract features by partitioning the dense information within the image, focusing independently on texture, color, or edges. These features are then used as the decoupling targets rather than the images themselves, which could mitigate the decoupling difficulty. Second, we extend reconstruction loss from image-level to feature-level, thus extending the forgery representation decoupling from single-scale to multi-scale. Third, we propose a critical forgetting mechanism that forces the detector to forget the most salient features during training, which correspond to specific forgery artifact types in the training set. Extensive experimental results validate the efficacy of the proposed method. Wenqi Ren, Jianshu Li, Wei Wang 0335, Xiaochun Cao |
AAAI | 4 |
| 2025 | SdalsNet: Self-Distilled Attention Localization and Shift Network for Unsupervised Camouflaged Object DetectionabstractUnsupervised camouflaged object detection (UCOD) poses significant challenges, primarily attributed to the absence of human labels. Existing UCOD methodologies, leveraging attention mechanisms, often struggle to achieve precise localization of camouflaged objects. To overcome this limitation, we introduce a groundbreaking fully unsupervised algorithm for attention-guided camouflaged object localization, shift, and inference, termed the self-distilled attention localization and shift network (SdalsNet). In this study, we formulate an attention localization methodology aimed at accurately identifying the central coordinate of the camouflaged object. Furthermore, we propose four distinct loss functions tailored to refine the precision of attentional positioning. These loss functions effectively constrain the distances between three types of class tokens, facilitating seamless attentional shifting across the input sample. Additionally, we design a sophisticated prediction inference technique to reconstruct the binary output of an attention map, thereby providing a comprehensive understanding of the detected camouflaged objects. Experimental results on four challenging COD benchmark datasets corroborate the effectiveness of our proposed approach, demonstrating notable superiority over state-of-the-art methods. Peiyao Shou, Yixiu Liu, Wei Wang 0335, Yaoqi Sun, Zhigao Zheng 0001, Shangdong Zhu, Chenggang Yan 0001 |
AAAI | 3 |
| 2025 | Intra and Inter Parser-Prompted Transformers for Effective Image RestorationabstractWe propose Intra and Inter Parser-Prompted Transformers (PPTformer) that explore useful features from visual foundation models for image restoration. Specifically, PPTformer contains two parts: an Image Restoration Network (IRNet) for restoring images from degraded observations and a Parser-Prompted Feature Generation Network (PPFGNet) for providing IRNet with reliable parser information to boost restoration. To enhance the integration of the parser within IRNet, we propose Intra Parser-Prompted Attention (IntraPPA) and Inter Parser-Prompted Attention (InterPPA) to implicitly and explicitly learn useful parser features to facilitate restoration. The IntraPPA re-considers cross attention between parser and restoration features, enabling implicit perception of the parser from a long-range and intra-layer perspective. Conversely, the InterPPA initially fuses restoration features with those of the parser, followed by formulating these fused features within an attention mechanism to explicitly perceive parser information. Further, we propose a parser-prompted feed-forward network to guide restoration within pixel-wise gating modulation. Experimental results show that PPTformer achieves state-of-the-art performance on image deraining, defocus deblurring, desnowing, and low-light enhancement. Cong Wang 0018, Jinshan Pan, Wei Wang 0335 |
AAAI | 4 |
| 2025 | Detecting Synthetic Image by Cross-Modal Commonality InteractionabstractExisting synthetic image detection approaches can be categorized into three paradigms: spatial, frequency, and fingerprint-based methods. Our analysis reveals a fundamental commonality across these paradigms: a significant reliance on high-frequency image components. This observation highlights the discriminative power of high-frequency information for this task and provides a strong rationale for learning generalized artifact representations based on multi-modal fusion strategies. Building on this insight, we introduce a multi-modal high-frequency interactive detection framework for general synthetic image detection. This framework explicitly integrates high-frequency information from both the spatial and frequency domains. Specifically, its spatial processing branch incorporates a novel high-frequency self-enhancement module to bolster local high-frequency representations. Concurrently, the frequency processing branch utilizes a multi-scale frequency information enhancement module to capture diverse contextual cues. At the feature fusion stage, we propose a pooling-guided cross-modal high-frequency interaction module, which dynamically weights cross-modal information to further reinforce salient high-frequency representations. Extensive experiments on public datasets demonstrate that our proposed framework achieves state-of-the-art performance in real-world detection scenarios. Wenqi Ren, Wei Wang 0335, Linchao Zhang, Xiaochun Cao |
ACM Multimedia | 3 |
| 2025 | DSPF: Dual-Stage Preservation and Fusion for Source-Free Domain Adaptive Point Cloud CompletionabstractPoint cloud completion is crucial for downstream tasks in 3D visual perception. However, existing methods often struggle to generalize to real-world scans due to their heavy reliance on abundant paired point clouds for training and their neglect of the distribution shift between training and testing datasets. To address these limitations, this paper explores a practical and challenging setting: ''source-free domain adaptive point cloud completion'', where a well-trained source model must adapt to the target data distribution without access to source data, aiming to improve completion performance. To tackle this problem, we propose a novel method called ''Dual-Stage Preservation and Fusion'' (DSPF), which comprises two key training stages tailored to this new setting. In the source preservation stage, we introduce graph structural alignment and marginal feature alignment to preserve and transfer essential knowledge from the source domain. In the target fusion stage, we design a self-supervised loss to capture the geometric structure of target instances and establish a bidirectional interaction mechanism to transfer partial source knowledge to the target distribution. Extensive experiments on various cross-domain point cloud completion benchmarks demonstrate that our proposed DSPF significantly outperforms existing methods, validating its effectiveness and robustness in source-free domain adaptation scenarios. Our code is available at https://github.com/ZhiXia-SEU/DSPF. Zhiqian Xia, Haifeng Xia, Shichao Jin, Wei Wang 0335, Zhengming Ding, Xiaochun Cao |
ACM Multimedia | 4 |
| 2025 | Rethinking Joint Maximum Mean Discrepancy for Visual Domain AdaptationabstractIn domain adaption (DA), joint maximum mean discrepancy (JMMD), as a famous distribution-distance metric, aims to measure joint probability distribution difference between the source domain and target domain, while it is still not fully explored and especially hard to be applied into a subspace-learning framework as its empirical estimation involves a tensor-product operator whose partial derivative is difficult to obtain. To solve this issue, we deduce a concise JMMD based on the Representer theorem that avoids the tensor-product operator and obtains two essential findings. First, we reveal the uniformity of JMMD by proving that previous marginal, class conditional, and weighted class conditional probability distribution distances are three special cases of JMMD with different label reproducing kernels. Second, inspired by graph embedding, we observe that the similarity weights, which strengthen the intra-class compactness in the graph of Hilbert Schmidt independence criterion (HSIC), take opposite signs in the graph of JMMD, revealing why JMMD degrades the feature discrimination. This motivates us to propose a novel loss JMMD-HSIC by jointly considering JMMD and HSIC to promote discrimination of JMMD. Extensive experiments on several cross-domain datasets could demonstrate the validity of our revealed theoretical results and the effectiveness of our proposed JMMD-HSIC. Wei Wang 0335, Haifeng Xia, Chao Huang 0008, Zhengming Ding, Cong Wang 0018, Xiaochun Cao |
NeurIPS | 1 |
| 2025 | Robust Label Propagation and Graph Embedding for Cross-Domain Image ClassificationabstractCross-domain label propagation (LP) faces two main challenges: 1) learning domain-invariant and 2) discriminative feature representations and obtaining high-confidence predicted labels. The distribution differences between domains can make labels difficult to propagate across domains. Low-quality labels can distort the modeling process associated with label-induced loss, resulting in decreased performance. We propose a novel cross-domain image classification method, namely, robust LP and graph embedding (RLPGE). We introduce a nuclear norm maximization constraint in order to make the predicted labels more diverse in categories while preserving their discriminability. The graph embedding process brings two nearby same-class samples close in the embedding subspace, ensuring domain invariance and local discriminability of the embedded features. For optimal graph learning, we simultaneously optimize the cross-domain graph and two intradomain graphs using both features and labels, enhancing their local discriminability and robustness to feature noise. We conducted comprehensive experiments on four cross-domain image classification datasets. The results demonstrate that our proposed RLPGE method outperforming some state-of-the-art approaches Chengjin Yu, Wuchang Liang, Wei Wang 0335, Yuan-Ting Yan, Hua Zhang 0008 |
IEEE Internet Things J. | 5 |
| 2025 | An unsupervised medical image registration network for intelligent medical education
Jie Mu, Jing Zhang 0037, Tiantian Yan, Wei Wang 0335, Hua Zhang 0008, Wenqi Ren |
Neural Comput. Appl. | 6 |
| 2025 | Deep Label Propagation With Nuclear Norm Maximization for Visual Domain AdaptationabstractDomain adaptation aims to leverage abundant label information from a source domain to an unlabeled target domain with two different distributions. Existing methods usually rely on a classifier to generate high-quality pseudo-labels for the target domain, facilitating the learning of discriminative features. Label propagation (LP), as an effective classifier, propagates labels from the source domain to the target domain by designing a smooth function over a similarity graph, which represents structural relationships among data points in feature space. However, LP has not been thoroughly explored in deep neural network-based domain adaptation approaches. Additionally, the probability labels generated by LP are low-confident and LP is sensitive to class imbalance problem. To address these problems, we propose a novel approach for domain adaptation named deep label propagation with nuclear norm maximization (DLP-NNM). Specifically, we employ the constraint of nuclear norm maximization to enhance both label confidence and class diversity in LP and propose an efficient algorithm to solve the corresponding optimization problem. Subsequently, we utilize the proposed LP to guide the classifier layer in a deep discriminative adaptation network using the cross-entropy loss. As such, the network could produce more reliable predictions for the target domain, thereby facilitating more effective discriminative feature learning. Extensive experimental results on three cross-domain benchmark datasets demonstrate that the proposed DLP-NNM surpasses existing state-of-the-art domain adaptation approaches. Wei Wang 0335, Cong Wang 0018, Chao Huang 0008, Zhengming Ding, Feiping Nie 0001, Xiaochun Cao |
IEEE Trans. Image Process. | 1 |
| 2025 | Optimal Graph Learning-Based Label Propagation for Cross-Domain Image ClassificationabstractLabel propagation (LP) is a popular semi-supervised learning technique that propagates labels from a training dataset to a test one using a similarity graph, assuming that nearby samples should have similar labels. However, the recent cross-domain problem assumes that training (source domain) and test data sets (target domain) follow different distributions, which may unexpectedly degrade the performance of LP due to small similarity weights connecting the two domains. To address this problem, we propose optimal graph learning-based label propagation (OGL2P), which optimizes one cross-domain graph and two intra-domain graphs to connect the two domains and preserve domain-specific structures, respectively. During label propagation, the cross-domain graph draws two labels close if they are nearby in feature space and from different domains, while the intra-domain graph pulls two labels close if they are nearby in feature space and from the same domain. This makes label propagation more insensitive to cross-domain problems. During graph embedding, we optimize the three graphs using features and labels in the embedded subspace to extract locally discriminative and domain-invariant features and make the graph construction process robust to noise in the original feature space. Notably, as a more relaxed constraint, locally discriminative and domain-invariant can somewhat alleviate the contradiction between discriminability and domain-invariance. Finally, we conduct extensive experiments on five cross-domain image classification datasets to verify that OGL2P outperforms some state-of-the-art cross-domain approaches. Wei Wang 0335, Mengzhu Wang, Chao Huang 0008, Cong Wang 0018, Jie Mu, Feiping Nie 0001, Xiaochun Cao |
IEEE Trans. Image Process. | 1 |
| 2025 | Multimodal Large Language Model with LoRA Fine-Tuning for Multimodal Sentiment AnalysisabstractMultimodal sentiment analysis has become a popular research topic in recent years. However, existing methods have two unaddressed limitations: (1) they use limited supervised labels to train models, which makes it impossible for model to fully learn sentiments in different modal data; (2) they employ text and image pre-trained models trained in different unimodal tasks to extract different modal features, so that the extracted features cannot take into account the interactive information between image and text. To solve these problems, in this paper we propose a Vision-Language Contrastive Learning network (VLCLNet). First, we introduce a pre-trained Large Language Model (LLM), which is trained from vast quantities of multimodal data, has better understanding ability for image and text contents, thus being effectively applied to different tasks while requiring few amount of labelled training data. Second, we adapt a Multimodal Large Language Model (MLLM), BLIP-2 (Bootstrapping Language-Image Pre-training) network, to extract multimodal fusion feature. Such MLLM can fully consider the correlation between images and texts when extracting features. In addition, due to the discrepancy between the pre-training task and the sentiment analysis task, the pre-trained model will output the suboptimal prediction results. We use Low-Rank Adaptation (LoRA) fine-tuning strategy to update the model parameters on sentiment analysis task, which avoids the issue of inconsistent task between pre-training task and downstream task. Experiments verify that the proposed VLCLNet is superior to other strong baselines. Jie Mu, Wei Wang 0335, Tiantian Yan, Guanglu Wang |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2025 | Multimodal Evidential Learning for Open-World Weakly-Supervised Video Anomaly DetectionabstractEfforts in weakly-supervised video anomaly detection center on detecting abnormal events within videos by coarse-grained labels, which has been successfully applied to many real-world applications. However, a significant limitation of most existing methods is that they are only effective for specific objects in specific scenarios, which makes them prone to misclassification or omission when confronted with previously unseen anomalies. Relative to conventional anomaly detection tasks, Open-world Weakly-supervised Video Anomaly Detection (OWVAD) poses greater challenges due to the absence of labels and fine-grained annotations for unknown anomalies. To address the above problem, we propose a multi-scale evidential vision-language model to achieve open-world video anomaly detection. Specifically, we leverage generalized visual-language associations derived from CLIP to harness the full potential of large pre-trained models in addressing the OWVAD task. Subsequently, we integrate a multi-scale temporal modeling module with a multimodal evidence collector to achieve precise frame-level detection of both seen and unseen anomalies. Extensive experiments on two widely-utilized benchmarks have conclusively validated the effectiveness of our method. The code will be made publicly available. Chao Huang 0008, Weiliang Huang, Qiuping Jiang, Wei Wang 0335, Jie Wen 0001, Bob Zhang 0001 |
IEEE Trans. Multim. | 4 |
| 2025 | STFormer: Spatial-Temporal-Aware Transformer for Video Instance SegmentationabstractVideo instance segmentation (VIS) is a challenging task, requiring handling object classification, segmentation, and tracking in videos. Existing Transformer-based VIS approaches have shown remarkable success, combining encoded features and instance queries as decoder inputs. However, their decoder inputs are low-resolution due to computational cost, resulting in a loss of fine-grained information, sensitivity to background interference, and poor handling of small objects. Moreover, the queries are randomly initialized without location information, hindering convergence efficiency and accurate object instance localization. To address these issues, we propose a novel VIS approach, STFormer, with a spatial-temporal feature aggregation (STFA) module and spatial-temporal-aware Transformer (STT). Specifically, STFA obtains robust high-resolution masked features efficiently for the decoder, while STT's location-guided instance query (LGIQ) improves initial instance queries. STFormer preserves more fine-grained information, improves convergence efficiency, and localizes object instance features accurately. Extensive experiments on YouTube-VIS 2019, YouTube-VIS 2021, and OVIS datasets show that STFormer outperforms mainstream VIS methods. Wei Wang 0335, Mengzhu Wang, Huibin Tan, Long Lan, Zhigang Luo, Xinwang Liu 0002, Kenli Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | Smooth-Guided Implicit Data Augmentation for Domain GeneralizationabstractThe training process of a domain generalization (DG) model involves utilizing one or more interrelated source domains to attain optimal performance on an unseen target domain. Existing DG methods often use auxiliary networks or require high computational costs to improve the model's generalization ability by incorporating a diverse set of source domains. In contrast, this work proposes a method called Smooth-Guided Implicit Data Augmentation (SGIDA) that operates in the feature space to capture the diversity of source domains. To amplify the model's generalization capacity, a distance metric learning (DML) loss function is incorporated. Additionally, rather than depending on deep features, the suggested approach employs logits produced from cross entropy (CE) losses with infinite augmentations. A theoretical analysis shows that logits are effective in estimating distances defined on original features, and the proposed approach is thoroughly analyzed to provide a better understanding of why logits are beneficial for DG. Moreover, to increase the diversity of the source domain, a sampling-based method called smooth is introduced to obtain semantic directions from interclass relations. The effectiveness of the proposed approach is demonstrated through extensive experiments on widely used DG, object detection, and remote sensing datasets, where it achieves significant improvements over existing state-of-the-art methods across various backbone networks. Mengzhu Wang, Junze Liu, Ge Luo 0003, Shanshan Wang 0008, Wei Wang 0335, Long Lan, Ye Wang 0023, Feiping Nie 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Correlation Matching Transformation Transformers for UHD Image RestorationabstractThis paper proposes UHDformer, a general Transformer for Ultra-High-Definition (UHD) image restoration. UHDformer contains two learning spaces: (a) learning in high-resolution space and (b) learning in low-resolution space. The former learns multi-level high-resolution features and fuses low-high features and reconstructs the residual images, while the latter explores more representative features learning from the high-resolution ones to facilitate better restoration. To better improve feature representation in low-resolution space, we propose to build feature transformation from the high-resolution space to the low-resolution one. To that end, we propose two new modules: Dual-path Correlation Matching Transformation module (DualCMT) and Adaptive Channel Modulator (ACM). The DualCMT selects top C/r (r is greater or equal to 1 which controls the squeezing level) correlation channels from the max-pooling/mean-pooling high-resolution features to replace low-resolution ones in Transformers, which can effectively squeeze useless content to improve the feature representation in low-resolution space to facilitate better recovery. The ACM is exploited to adaptively modulate multi-level high-resolution features, enabling to provide more useful features to low-resolution space for better learning. Experimental results show that our UHDformer reduces about ninety-seven percent model sizes compared with most state-of-the-art methods while significantly improving performance under different training sets on 3 UHD image restoration tasks, including low-light image enhancement, image dehazing, and image deblurring. The source codes will be made available at https://github.com/supersupercong/UHDformer. Cong Wang 0018, Jinshan Pan, Wei Wang 0335, Mengzhu Wang, Xiao-Ming Wu 0003, Jun Liu 0036 |
AAAI | 3 |
| 2024 | SelfPromer: Self-Prompt Dehazing Transformers with Depth-ConsistencyabstractThis work presents an effective depth-consistency Self-Prompt Transformer, terms as SelfPromer, for image dehazing. It is motivated by an observation that the estimated depths of an image with haze residuals and its clear counterpart vary. Enforcing the depth consistency of dehazed images with clear ones, therefore, is essential for dehazing. For this purpose, we develop a prompt based on the features of depth differences between the hazy input images and corresponding clear counterparts that can guide dehazing models for better restoration. Specifically, we first apply deep features extracted from the input images to the depth difference features for generating the prompt that contains the haze residual information in the input. Then we propose a prompt embedding module that is designed to perceive the haze residuals, by linearly adding the prompt to the deep features. Further, we develop an effective prompt attention module to pay more attention to haze residuals for better removal. By incorporating the prompt, prompt embedding, and prompt attention into an encoder-decoder network based on VQGAN, we can achieve better perception quality. As the depths of clear images are not available at inference, and the dehazed images with one-time feed-forward execution may still contain a portion of haze residuals, we propose a new continuous self-prompt inference that can iteratively correct the dehazing model towards better haze-free image generation. Extensive experiments show that our SelfPromer performs favorably against the state-of-the-art approaches on both synthetic and real-world datasets in terms of perception metrics including NIQE, PI, and PIQE. The source codes will be made available at https://github.com/supersupercong/SelfPromer. Cong Wang 0018, Jinshan Pan, Wanyu Lin, Jiangxin Dong, Wei Wang 0335, Xiao-Ming Wu 0003 |
AAAI | 5 |
| 2024 | Improving Forest Management Efficiency: A New Metric for IoT Node DeploymentabstractThe internet of things (IoT) is revolutionizing various industries by enabling the creation of smart systems for forest management, promoting the emergence of the forestry Internet of Things (IoFT). However, existing IoT node deployment methods often overlook geographic constraints, while solar-powered IoFT nodes face limitations in forests due to tree obstruction and maintenance challenges. Furthermore, to achieve cost-effective monitoring, it is essential to consider both coverage and network costs. In this paper, we propose a new metric, the coverage benefit ratio (CBR), which balances coverage and cost, ensuring long-term stable operation of forest monitoring systems while reducing maintenance costs and environmental impact. We first formulate the optimal deployment model to find the minimum-cost IoFT. Then, we propose develop a low complexity algorithm to solve the defined NP-hard optimization problem. Simulation results demonstrate that the effectiveness and progressiveness of the proposed method. Pengju Si, Yixiu Liu, Zhigao Zheng 0001, Wei Wang 0335 |
HPCC | 6 |
| 2024 | AllWeather-Net: Unified Image Enhancement for Autonomous Driving Under Adverse Weather and Low-Light Conditions
Chenghao Qian, Mahdi Rezaei 0001, Saeed Anwar, Wenjing Li 0005, Tanveer Hussain 0001, Mohsen Azarmi, Wei Wang 0335 |
ICPR (30) | 7 |
| 2024 | Explore Internal and External Similarity for Single Image Deraining with Graph Neural Networks
Cong Wang 0018, Wei Wang 0335, Chengjin Yu, Jie Mu |
IJCAI | 2 |
| 2024 | Optimal Graph Learning and Nuclear Norm Maximization for Deep Cross-Domain Robust Label Propagation
Wei Wang 0335, Chao Huang 0008, Yang Cao 0011, Cong Wang 0018, Xiaochun Cao |
IJCAI | 1 |
| 2024 | Progressive Local and Non-Local Interactive Networks with Deeply Discriminative Training for Image DerainingabstractIn this paper, we develop a progressive local and non-local interactive network with multi-scale cross-content deeply discriminative learning to solve image deraining. The proposed model contains two key techniques: 1) Progressive Local and Non-Local Interactive Network (PLNLIN) and 2) Multi-Scale Cross-Content Deeply Discriminative Learning (MCDDL). The PLNLIN is a U-shaped encoder-decoder network, where the proposed new Progressive Local and Non-Local Interactive Module (PLNLIM) is the basic unit in the encoder-decoder framework. The PLNLIM fully explores local and non-local learning in convolution and Transformer operation respectively and the local and non-local content are further interactively learned in a progressive manner. The proposed MCDDL not only discriminates the output of the generator but also receives the deep content from the generator to distinguish real and fake features at each side layer of the discriminator in a multi-scale manner. We show that the proposed MCDDL has fast and stable convergence properties that lack in existing discriminative learning manners. Extensive experiments demonstrate that the proposed method outperforms state-of-the-art methods on five public synthetic datasets and one real-world data. The source codes will be made available at https://github.com/supersupercong/PLNLIN-MCDDL. Cong Wang 0018, Jie Mu, Chengjin Yu, Wei Wang 0335 |
ACM Multimedia | 5 |
| 2024 | PercepLIE: A New Path to Perceptual Low-Light Image EnhancementabstractWhile current CNN-based low-light image enhancement (LIE) approaches have achieved significant progress, they often fail to generate better perceptual quality which requires restoring better details and more natural colors. To address these problems, we set a new path, called PercepLIE, by presenting the VQGAN with Multi-luminance Detail Compensation (MDC) and Global Color Adjustment (GCA). Specifically, observed that latent light features of the low-light images are quite different from those captured in normal light, we utilize VQGAN to explore the latent light representation of normal-light images to help the estimation of the low-light and normal-light mapping. Furthermore, we employ Gamma correction with varying Gamma values on the gradient to create multi-luminance details, forming the basis for our MDC module to facilitate better detail estimation. To optimize the colors of low-light input images, we introduce a simple yet effective GCA module that is based on spatially-varying representation between the estimated normal-light images in this module and low-light inputs. By combining the VQGAN with MDC and GCA within a stage-wise training mechanism, our method generates images with finer details and natural colors and achieves favorable performance on both synthetic and real-world datasets in terms of perceptual quality metrics including NIQE, PI, and LPIPS. The source codes will be made available at https://github.com/supersupercong/PercepLIE. Cong Wang 0018, Chengjin Yu, Jie Mu, Wei Wang 0335 |
ACM Multimedia | 4 |
| 2024 | Coarse-to-fine mechanisms mitigate diffusion limitations on image restoration
Qinyu Yang, Cong Wang 0018, Wei Wang 0335, Zhixun Su |
Comput. Vis. Image Underst. | 4 |
| 2024 | Inter-Class and Inter-Domain Semantic Augmentation for Domain GeneralizationabstractThe domain generalization approach seeks to develop a universal model that performs well on unknown target domains with the aid of diverse source domains. Data augmentation has proven to be an effective method to enhance domain generalization in computer vision. Recently, semantic-level based data augmentation has yielded remarkable results. However, these methods focus on sampling semantic directions on feature space from intra-class and intra-domain, limiting the diversity of the source domain. To address this issue, we propose a novel approach called Inter-Class and Inter-Domain Semantic Augmentation (CDSA) for domain generalization. We first introduce a sampling-based method called CrossSmooth to obtain semantic directions from inter-class. Then, CrossVariance obtains the styles of different domains by sampling semantic directions. Our experiments on four well-known domain generalization benchmark datasets (Digits-DG, PACS, Office-Home, and DomainNet) demonstrate the effectiveness of our approach. We also validate our approach on commonly-used semantic segmentation datasets, namely GTAV, SYNTHIA, Cityscapes, Mapillary, and BDDS which also show significant improvements. Mengzhu Wang, Yuehua Liu, Jianlong Yuan, Shanshan Wang 0008, Zhibin Wang 0004, Wei Wang 0335 |
IEEE Trans. Image Process. | 6 |
| 2024 | MOCOLNet: A Momentum Contrastive Learning Network for Multimodal Aspect-Level Sentiment AnalysisabstractMultimodal aspect-level sentiment analysis has attracted increasing attention in recent years. However, existing methods have two unaddressed limitations: (1) due to the lack of labelled pre-training data of dedicated sentiment analysis, the methods with a pre-training manner produce suboptimal prediction results; (2) most existing methods employ a self-attention encoder to fuse multimodal tokens, which not only ignores the alignment relationship between different modal tokens but also makes the model unable to capture the semantic links between images and texts. In this paper, we propose a momentum contrastive learning network (MOCOLNet) to overcome above limitations. First, we merge the pre-training stage with the training stage to design an end-to-end training manner which uses less labelled data dedicated to sentiment analysis to obtain better prediction results. Second, we propose a multimodal contrastive learning method to align the different modal representations before data fusing, and design a cross-modal matching strategy to provide semantic interactive information between texts and images. Moreover, we introduce an auxiliary momentum strategy to increase the robustness of model. We also analyse the effectiveness of the proposed multimodal contrastive learning method using a mutual information theory. Experiments verify that the proposed MOCOLNet is superior to other strong baselines. Jie Mu, Feiping Nie 0001, Wei Wang 0335, Jing Zhang 0037, Han Liu 0008 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2023 | A Closer Look at Classifier in Adversarial Domain GeneralizationabstractThe task of domain generalization is to learn a classification model from multiple source domains and generalize it to unknown target domains. The key to domain generalization is learning discriminative domain-invariant features. Invariant representations are achieved using adversarial domain generalization as one of the primary techniques. For example, generative adversarial networks have been widely used, but suffer from the problem of low intra-class diversity, which can lead to poor generalization ability. To address this issue, we propose a new method called auxiliary classifier in adversarial domain generalization (CloCls). CloCls improve the diversity of the source domain by introducing auxiliary classifier. Combining typical task-related losses, e.g., cross-entropy loss for classification and adversarial loss for domain discrimination, our overall goal is to guarantee the learning of condition-invariant features for all source domains while increasing the diversity of source domains. Further, inspired by smoothing optima have improved generalization for supervised learning tasks like classification. We leverage that converging to a smooth minima with respect task loss stabilizes the adversarial training leading to better performance on unseen target domain which can effectively enhances the performance of domain adversarial methods. We have conducted extensive image classification experiments on benchmark datasets in domain generalization, and our model exhibits sufficient generalization ability and outperforms state-of-the-art DG methods. Ye Wang 0023, Junyang Chen 0001, Mengzhu Wang, Hao Li 0058, Wei Wang 0335, Houcheng Su, Zhihui Lai 0001, Wei Wang 0077, Zhenghan Chen |
ACM Multimedia | 5 |
| 2023 | PromptRestorer: A Prompting Image Restoration Method with Degradation PerceptionabstractWe show that raw degradation features can effectively guide deep restoration models, providing accurate degradation priors to facilitate better restoration. While networks that do not consider them for restoration forget gradually degradation during the learning process, model capacity is severely hindered. To address this, we propose a Prompting image Restorer, termed as PromptRestorer. Specifically, PromptRestorer contains two branches: a restoration branch and a prompting branch. The former is used to restore images, while the latter perceives degradation priors to prompt the restoration branch with reliable perceived content to guide the restoration process for better recovery. To better perceive the degradation which is extracted by a pre-trained model from given degradation observations, we propose a prompting degradation perception modulator, which adequately considers the characters of the self-attention mechanism and pixel-wise modulation, to better perceive the degradation priors from global and local perspectives. To control the propagation of the perceived content for the restoration branch, we propose gated degradation perception propagation, enabling the restoration branch to adaptively learn more useful features for better recovery. Extensive experimental results show that our PromptRestorer achieves state-of-the-art results on 4 image restoration tasks, including image deraining, deblurring, dehazing, and desnowing. Cong Wang 0018, Jinshan Pan, Wei Wang 0335, Jiangxin Dong, Mengzhu Wang, Yakun Ju, Junyang Chen 0001 |
NeurIPS | 3 |
| 2023 | mShield: Protecting In-process Sensitive Data Against Vulnerable Third-Party Libraries
Yunming Zhang, Quanwei Cai 0001, Houqiang Li, Jingqiang Lin 0001, Wei Wang 0335 |
SecureComm (1) | 5 |
| 2023 | Importance filtered soft label-based deep adaptation network
Wei Wang 0335, Mengzhu Wang, Zhihui Wang 0001 |
Knowl. Based Syst. | 1 |
| 2023 | Boosting unsupervised domain adaptation: A Fourier approach
Mengzhu Wang, Shanshan Wang 0008, Ye Wang 0023, Wei Wang 0335, Tianyi Liang 0001, Junyang Chen 0001, Zhigang Luo |
Knowl. Based Syst. | 4 |
| 2023 | Class-specific and self-learning local manifold structure for domain adaptation
Wei Wang 0335, Mengzhu Wang, Long Lan, Quannan Zu, Xiang Zhang 0008, Cong Wang 0018 |
Pattern Recognit. | 1 |
| 2023 | Reducing bi-level feature redundancy for unsupervised domain adaptation
Mengzhu Wang, Shanshan Wang 0008, Wei Wang 0335, Li Shen 0008, Xiang Zhang 0008, Long Lan, Zhigang Luo |
Pattern Recognit. | 3 |
| 2023 | Rethinking Maximum Mean Discrepancy for Visual Domain AdaptationabstractExisting domain adaptation approaches often try to reduce distribution difference between source and target domains and respect domain-specific discriminative structures by some distribution [e.g., maximum mean discrepancy (MMD)] and discriminative distances (e.g., intra-class and inter-class distances). However, they usually consider these losses together and trade off their relative importance by estimating parameters empirically. It is still under insufficient exploration so far to deeply study their relationships to each other so that we cannot manipulate them correctly and the model's performance degrades. To this end, this article theoretically proves two essential facts: 1) minimizing MMD equals to jointly minimizing their data variance with some implicit weights but, respectively, maximizing the source and target intra-class distances so that feature discriminability degrades and 2) the relationship between intra-class and inter-class distances is as one falls and another rises. Based on this, we propose a novel discriminative MMD with two parallel strategies to correctly restrain the degradation of feature discriminability or the expansion of intra-class distance; specifically: 1) we directly impose a tradeoff parameter on the intra-class distance that is implicit in the MMD according to 1) and 2) we reformulate the inter-class distance with special weights that are analogical to those implicit ones in the MMD and maximizing it can also lead to the intra-class distance falling according to 2). Notably, we do not consider the two strategies in one model due to 2). The experiments on several benchmark datasets not only prove the validity of our revealed theoretical results but also demonstrate that the proposed approach could perform better than some compared state-of-art methods substantially. Our preliminary MATLAB code will be available at https://github.com/WWLoveTransfer/. Wei Wang 0335, Zhengming Ding, Feiping Nie 0001, Junyang Chen 0001, Zhihui Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | MCascade R-CNN: A Modified Cascade R-CNN for Detection of Calcified on Coronary Artery Angiography ImagesabstractAmong cardiovascular diseases, coronary artery calcification (CAC) is a high-risk factor for worsening protopathy and increased mortality. However, the coronary artery an-giogram, which is the main approach for CAC diagnosis, suffers from plenty of photographing noise. This brings difficulties to detect calcification from the background. In this paper, a modified Cascade R-CNN (MCascade R-CNN) network is proposed to deal with the problem of calcium detection in angiograms. In the proposed network, we propose an innovative balanced aggregation pyramid structure, integrating multi-level features of every depth in the feature map, based on enhanced propagation of strong semantic features. In addition, a new convolutional attention mechanism is designed to improve the performance of the detector. Experiments show that the proposed method enjoys better performance in detecting and marking CAC in angiograms, Wei Wang 0335, Honggang Zhang 0002, Lihua Xie 0003, Bo Xu 0002 |
VCIP | 1 |
| 2022 | Informative pairs mining based adaptive metric learning for adversarial domain adaptation
Mengzhu Wang, Paul Li, Li Shen 0008, Ye Wang 0023, Shanshan Wang 0008, Wei Wang 0335, Xiang Zhang 0008, Junyang Chen 0001, Zhigang Luo |
Neural Networks | 6 |
| 2022 | Joint Adaptive Dual Graph and Feature Selection for Domain AdaptationabstractDomain adaptation aims to exploit domain-invariant features by aligning the cross-domain distributions in the manifold subspace for applying the classifier trained on the source domain to the target domain. However, two limitations may still deteriorate their performances: (1) the influences of noisy or irrelevant features in the original feature space are ignored, which may unexpectedly hurt the classification of target samples; (2) the graph constructed directly in the original data space cannot accurately capture the inherent local manifold structures of high-dimensional data due to the curse of dimensionality, which may seriously mislead the transferable features learning. In this paper, we propose a novel approach to address these problems, referred to as joint Adaptive Dual Graph and Feature Selection for domain adaptation (ADGFS). Specifically, feature selection can characterize the relative importance of different features through a scaling factor, which enables ADGFS to not only reduce the impacts of noisy or irrelevant features on knowledge transfer but also learn informative domain-invariant features. Meanwhile, ADGFS adaptively optimizes the dual graph by learning the similarity matrices of both instance-level and feature-level graphs in the projected low-dimensional manifold subspace rather than the original high-dimensional space, such that the intrinsic local manifold structures of data can be captured precisely. Moreover, ADGFS simultaneously aligns the marginal and conditional probability distributions in the nonnegative matrix factorization framework to narrow the distribution discrepancies between the two different domains, which can adequately transfer knowledge from the source domain to the target domain. Comprehensive experiments on four benchmark datasets can demonstrate that the effectiveness of the proposed approach in cross-domain image classification. Jing Sun 0012, Zhihui Wang 0001, Wei Wang 0335, Fuming Sun, Zhengming Ding |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Confidence Regularized Label Propagation Based Domain AdaptationabstractIn domain adaptation (DA), label-induced losses generally occupy a dominant position and most previous models regard hard or soft labels as their inputs. However, these two types of labels may mislead the modeling process of label-induced losses since hard label is sensitive to a wrongly-predicted sample while soft label may introduce label noise, thus they may cause negative transfer. To relieve this problem, we propose a novel label learning approach namely confidence regularized label propagation (CRLP) that regularizes the confidence of predicted soft labels with constraints of F-norm or L21-norm. It is validated that maximizing either one of these two constraints equals to minimizing entropy loss. Specially, we illustrate that L21-norm is more suitable for DA than F-norm when the dataset contain a large number of categories. Then, we leverage the regularized soft labels produced by CRLP to reformulate some popular label-induced losses that consider feature transferability and discriminability such as class-wise maximum mean discrepancy, intra-class compactness and inter-class dispersion in a probability manner to present a novel DA method (i.e., CRLP-DA). Comprehensive analysis and experiments on four cross-domain object recognition datasets verify that the proposed CRLP-DA outperforms some state-of-the-art methods, especially 59.5% for Office10+Caltech10 dataset with SURF features. For others to better reproduce, our preliminary Matlab code will be available athttps://github.com/WWLoveTransfer/CRLP-DA/. Wei Wang 0335, Baopu Li, Mengzhu Wang, Feiping Nie 0001, Zhihui Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Self-Training Enhanced: Network Embedding and Overlapping Community Detection With Adversarial LearningabstractNetwork embedding (NE) aims to encode the relations of vertices into a low-dimensional space. After NE, we can obtain the learned vectors of vertices that preserve the proximity of network structures for subsequent applications, e.g., vertex classification and link prediction. In existing NE models, they usually exploit the skip-gram with a negative sampling method to optimize their objective functions. Generally, this method learns the vertex representation only from the local connectivity of vertices (i.e., neighbors). However, there is a larger scope of vertex connectivity in real-world scenarios: a vertex may have multifaceted aspects and should belong to overlapping communities. Taking a social network as the overlapping example, a user may subscribe to the channels of politics, economy, and sports simultaneously, but the politics share more common attributes with the economy and less with the sports. In this article, we propose an adversarial learning approach (ACNE) for modeling overlapping communities of vertices. Specifically, we map the association between communities and vertices into an embedding space. Moreover, we take further research on enhancing our ACNE with the following two operations. First, in the initialization stage, we adopt a walking strategy with perception to obtain paths containing more possible boundary vertices to improve overlapping community detection. Then, after representation learning with ACNE, we use soft community assignments from a simple classifier as supervision to update the weights of ACNE. This self-training mechanism referred to as ACNE-ST can help ACNE to achieve better performance. Experimental results demonstrate that the proposed methods, including ACNE and ACNE-ST, can outperform the state-of-the-art models on the subsequent tasks of vertex classification and overlapping community detection. Junyang Chen 0001, Zhiguo Gong, Jiqian Mo, Wei Wang 0077, Wei Wang 0335, Cong Wang 0018, Weiwen Liu, Kaishun Wu |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2021 | Cost Affinity Learning Network for Stereo MatchingabstractExisting stereo matching methods mainly tend to directly aggregate features output from Convolutional Neural Network to obtain more discriminative cost features, but ignore the affinity of each element in the cost feature which also plays a key role in enhancing the cost feature. In this work, we propose a novel cost affinity learning network(CAL-Net) whose Affinity Enhanced Module(AEM) extracts the affinity of the elements in the cost feature and reconstructs a more discriminative feature. In addition, CAL-Net designs a Disparity Weight Loss(DWL) to guide training. Specifically, AEM takes the advatange of the self-attention mechanism to learn internal affinity between different elements and exploits it to reconstruct the cost feature for emphasizing informative elements. DWL calculates the adaptive weight according to disparity error. As the error decreases, the weight gradually increases and enables the network to gradually transit from the pixel level disparity to sub-pixel level. Experiments demonstrate that CAL-Net boosts the performance, especially in textureless and reflective regions, and achieves better results on Scene Flow and KITTI 2012 benchmarks than some typical related methods. Shenglun Chen, Baopu Li, Wei Wang 0335, Hong Zhang 0011, Zhihui Wang 0001 |
ICASSP | 3 |
| 2021 | InterBN: Channel Fusion for Adversarial Unsupervised Domain AdaptationabstractA classifier trained on one dataset rarely works on other datasets obtained under different conditions because of domain shifting. Such a problem is usually solved by domain adaptation methods. In this paper, we propose a novel unsupervised domain adaptation (UDA) method based on Interchangeable Batch Normalization (InterBN) to fuse different channels in deep neural networks for adversarial domain adaptation.Specifically, we first observe that the channels with small batch normalization scaling factor have less influence on the whole domain adaption, followed by a theoretical proof that the scaling factors for some channels will definitely come close to zero when imposing a sparsity regularization. Then, we replace the channels that have smaller scaling factors in the source domain with the mean of the channels which have larger scaling factors in the target domain or vice versa. Such a simple but effective channel fusion scheme can drastically increase the domain adaption ability.Extensive experimental results show that our InterBN significantly outperforms the current adversarial domain adaptation methods by a large margin on four visual benchmarks. In particular, InterBN achieves a remarkable improvement of 7.7% over the conditional adversarial adaptation networks (CDAN) on VisDA-2017 benchmark. Mengzhu Wang, Wei Wang 0335, Baopu Li, Xiang Zhang 0008, Long Lan, Huibin Tan, Tianyi Liang 0001, Wei Yu 0029, Zhigang Luo |
ACM Multimedia | 2 |
| 2021 | Domain adaptation with geometrical preservation and distribution alignment
Jing Sun 0012, Zhihui Wang 0001, Wei Wang 0335, Fuming Sun |
Neurocomputing | 3 |
| 2021 | Sparsely-labeled source assisted domain adaptation
Wei Wang 0335, Shenglun Chen, Yuankai Xiang, Jing Sun 0012, Zhihui Wang 0001, Fuming Sun, Zhengming Ding, Baopu Li |
Pattern Recognit. | 1 |
| 2020 | Adaptive Local Neighbors for Transfer Discriminative Feature LearningabstractIn Domain Adaptation (DA), how to reduce the distributional differences across domains and preserve the data structures are two critical issues to obtain domain-invariant features. Existing DA methods either preserve the Local Manifold Structure (LMS) or the Global Discriminative Consistency (GDC), while fail to take those two metrics into account simultaneously. Therefore, the extracted features are either short of discriminative ability or sensitive to the multimodally distributed data. Moreover, the local neighbored relationships among data points are mostly established in original data space, which is unreliable, especially for data with large noises. Therefore, this paper proposes a novel DA approach, i.e., Adaptive Local Neighbors for Transfer Discriminative Feature Learning, to leverage LMS and GDC into a unified transfer feature learning model, where we only focus on the GDC between the local neighbors, so that the extracted features are more discriminative and robust to the multimodally distributed data. Moreover, the data points' local neighbors are revealed adaptively in the learned subspace so that it is insensitive to the data noises. Compared with the state-of-the-art methods, the proposed approach achieves higher performance for different cross-domain image classification tasks, especially 3.0% improved for Office10+Caltech10 dataset. Wei Wang 0335, Zhihui Wang 0001, Zhengming Ding |
ECAI | 1 |
| 2020 | Inductive Document Representation Learning for Short Text Clustering
Junyang Chen 0001, Zhiguo Gong, Wei Wang 0077, Wei Wang 0335, Weiwen Liu, Cong Wang 0018 |
ECML/PKDD (3) | 5 |
| 2011 | Robust optical flow estimation based on brightness correction fieldsabstractOptical flow estimation is still an important task in computer vision with many interesting applications. However, the results obtained by most of the optical flow techniques are affected by motion discontinuities or illumination changes. In this paper, we introduce a brightness correction field combined with a gradient constancy constraint to reduce the sensibility to brightness changes between images to be estimated. The advantage of this brightness correction field is its simplicity in terms of computational complexity and implementation. By analyzing the deficiencies of the traditional total variation regularization term in weakly textured areas, we also adopt a structure-adaptive regularization based on the robust Huber norm to preserve motion discontinuities. Finally, the proposed energy functional is minimized by solving its corresponding Euler-Lagrange equation in a more effective multi-resolution scheme, which integrates the twice downsampling strategy with a support-weight median filter. Numerous experiments show that our method is more effective and produces more accurate results for optical flow estimation. Wei Wang 0335, Zhixun Su, Jinshan Pan, Ye Wang 0023, Riming Sun |
J. Zhejiang Univ. Sci. C | 1 |
| 2007 | A Multiparty Videoconferencing System Over an Application-Level Multicast ProtocolabstractIncreased speeds of PCs and networks have made media communications possible on the Internet. Today, the need for desktop videoconferencing is experiencing robust growth in both business and consumer markets. However, the synchronous delivery of high-volume media content is still a big challenge under a current heterogeneous Internet environment. In this paper, we present a multiparty videoconferencing system based on a peer-to-peer (P2P) solution. The contribution of our paper is twofold. On the one hand, we design an application-level multicast scheme which intends to tolerate the heterogeneity in videoconferencing applications. Design tradeoffs are analyzed and our decisions are made based on extensive experimentation. On the other, we design a five-layer architecture for implementing a multiparty videoconferencing system. This architecture makes a clear-cut distinction between different functional modules and therefore provides rich flexibility in feature adaptation. We believe that our work can be a helpful reference in other efforts on building desktop videoconferencing systems. Chong Luo 0001, Wei Wang 0335, Jiang Li 0008 |
IEEE Trans. Multim. | 2 |