Weining Wang 0001

dblp:97/6006-1 · DBLP profile ↗
← Back
33ranked-venue papers
3as first author
31since 2021 · last 2026
0000-0001-7299-6431ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 22 · 2 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 2 first-author · 16 since 2021Security and privacy · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 UniAlignment: Semantic Alignment for Unified Image Generation, Understanding, Manipulation and Perception
abstract
The remarkable success of diffusion models in text-to-image generation has sparked growing interest in expanding their capabilities to a variety of multi-modal tasks, including image understanding, manipulation, and perception. These tasks require advanced semantic comprehension across both visual and textual modalities, especially in scenarios involving complex semantic instructions. However, existing approaches often rely heavily on vision-language models (VLMs) or modular designs for semantic guidance, leading to fragmented architectures and computational inefficiency. To address these challenges, we propose UniAlignment, a unified multimodal generation framework within a single diffusion transformer. UniAlignment introduces a dual-stream diffusion training strategy that incorporates both intrinsic-modal semantic alignment and cross-modal semantic alignment, thereby enhancing the model's cross-modal consistency and instruction-following robustness. Additionally, we present SemGen-Bench, a new benchmark specifically designed to evaluate multimodal semantic consistency under complex textual instructions. Extensive experiments across multiple tasks and benchmarks demonstrate that UniAlignment outperforms existing baselines, underscoring the significant potential of diffusion models in unified multimodal generation.
Xinyang Song, Weining Wang 0001, Shaozhen Liu, Jingdong Chen, Qi Li 0005, Zhenan Sun
AAAI3
2026 Learning Unknown Spoof Prompts for Generalized Face Anti-Spoofing Using Only Real Face Images
Fangling Jiang, Qi Li 0005, Weining Wang 0001, Zhenan Sun
Int. J. Comput. Vis.3
2026 CAS-AIR-3D: A Large-scale Low-quality Multi-modal Face Database
Qi Li 0005, Xiaoxiao Dong, Weining Wang 0001, Zhenan Sun, Tieniu Tan, Caifeng Shan
Int. J. Comput. Vis.3
2026 Learning Knowledge-Based Prompts for Robust 3D Mask Presentation Attack Detection
abstract
3D mask presentation attack detection is crucial for protecting face recognition systems against the rising threat of 3D mask attacks. While most existing methods utilize multimodal features or remote photoplethysmography (rPPG) signals to distinguish between real faces and 3D masks, they face significant challenges, such as the high costs associated with multimodal sensors and limited generalization ability. Detection-related text descriptions offer concise, universal information and are cost-effective to obtain. However, the potential of vision-language multimodal features for 3D mask presentation attack detection remains unexplored. In this paper, we propose a novel knowledge-based prompt learning framework to explore the strong generalization capability of vision-language models for 3D mask presentation attack detection. Specifically, our approach incorporates entities and triples from knowledge graphs into the prompt learning process, generating fine-grained, task-specific explicit prompts that effectively harness the knowledge embedded in pre-trained vision-language models. Furthermore, considering different input images may emphasize distinct knowledge graph elements, we introduce a visual-specific knowledge filter based on an attention mechanism to refine relevant elements according to the visual context. Additionally, we leverage causal graph theory insights into the prompt learning process to further enhance the generalization ability of our method. During training, a spurious correlation elimination paradigm is employed, which removes category-irrelevant local image patches using guidance from knowledge-based text features, fostering the learning of generalized causal prompts that align with category-relevant local patches. Experimental results demonstrate that the proposed method achieves state-of-the-art intra- and cross-scenario detection performance on benchmark datasets.
Fangling Jiang, Qi Li 0005, Weining Wang 0001, Caifeng Shan, Zhenan Sun, Ming-Hsuan Yang 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 AR-Diffusion: Asynchronous Video Generation with Auto-Regressive Diffusion
abstract
The task of video generation requires synthesizing visually realistic and temporally coherent video frames. Existing methods primarily use asynchronous auto-regressive models or synchronous diffusion models to address this challenge. However, asynchronous auto-regressive models often suffer from inconsistencies between training and inference, leading to issues such as error accumulation, while synchronous diffusion models are limited by their reliance on rigid sequence length. To address these issues, we introduce Auto-Regressive Diffusion (AR-Diffusion), a novel model that combines the strengths of auto-regressive and diffusion models for flexible, asynchronous video generation. Specifically, our approach leverages diffusion to gradually corrupt video frames in both training and inference, reducing the discrepancy between these phases. Inspired by auto-regressive generation, we incorporate a non-decreasing constraint on the corruption timesteps of individual frames, ensuring that earlier frames remain clearer than subsequent ones. This setup, together with temporal causal attention, enables flexible generation of videos with varying lengths while preserving temporal coherence. In addition, we design two specialized timestep schedulers: the FoPP scheduler for balanced timestep sampling during training, and the AD scheduler for flexible timestep differences during inference, supporting both synchronous and asynchronous generation. Extensive experiments demonstrate the superiority of our proposed method, which achieves competitive and state-of-the-art results across four challenging benchmarks.1 2
Mingzhen Sun, Weining Wang 0001, Jiawei Liu 0001, Wanquan Feng, Shanshan Lao, SiYu Zhou 0002, Jing Liu 0001
CVPR2
2025 VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and Dataset
abstract
In this paper, we propose the Vision-Audio-Language Omni-peRception pretraining model (VALOR) for multimodal understanding and generation. Unlike widely-studied vision-language pretraining models, VALOR jointly models the relationships among vision, audio, and language in an end-to-end manner. It consists of three separate encoders for single modality representations and a decoder for multimodal conditional text generation. We design two pretext tasks to pretrain the VALOR model: Multimodal Grouping Alignment (MGA) and Multimodal Grouping Captioning (MGC). MGA projects vision, language, and audio into the same common space, simultaneously building vision-language, audio-language, and audiovisual-language alignment. MGC learns to generate text tokens under conditions of vision, audio, or both. To promote vision-audio-language pretraining research, we construct a large-scale, high-quality tri-modality dataset named VALOR-1M, containing 1 million audible videos with human-annotated audiovisual captions. Extensive experiments show that VALOR can learn strong multimodal correlations and generalize to various downstream tasks (e.g., retrieval, captioning, and question answering) with different input modalities (e.g., vision-language, audio-language, and audiovisual-language). VALOR achieves new state-of-the-art performance on a series of public cross-modality benchmarks.
Jing Liu 0001, Xingjian He, Longteng Guo, Weining Wang 0001, Jinhui Tang 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2025 CGViT: Cross-image GroupViT for zero-shot semantic segmentation
Jie Jiang 0016, Xingjian He, Weining Wang 0001, Jing Liu 0001
Pattern Recognit.4
2025 Hierarchical Contrastive Learning for Semantic Segmentation
abstract
Recently, pixel-to-pixel contrastive learning in single-scale feature space has been widely studied in semantic segmentation to learn a unified feature expression for pixels of the same category. However, the unified representation is too extreme, and the receptive field of each single-scale pixel is limited, which is insufficient to reflect the representative features of the category. To address these problems, this article extends the single-scale feature space to that of multiscale and proposes a hierarchical contrastive learning (Hi-CL) method to explore pixel-to-component semantic relationships. First, we generate multiscale candidate samples by applying several pooling windows with different sizes on a feature map, where different windows may represent different parts of the objects in the image. Then, we prune the sample set through threshold-based criteria to select appropriate samples for feature representation learning. Finally, Hi-CL is performed to learn the pixel-to-component consistency with the pruned samples. Our method is easy to be applied on existing semantic segmentation models and obtains consistent improvement. Furthermore, we achieve state-of-the-art results on three popular benchmarks, including Cityscapes, ADE20K, and COCO Stuff datasets.
Jie Jiang 0016, Xingjian He, Weining Wang 0001, Hanqing Lu, Jing Liu 0001
IEEE Trans. Neural Networks Learn. Syst.3
2024 MM-LDM: Multi-Modal Latent Diffusion Model for Sounding Video Generation
abstract
Sounding Video Generation (SVG) is an audio-video joint generation task challenged by high-dimensional signal spaces, distinct data formats, and different patterns of content information. To address these issues, we introduce a novel multi-modal latent diffusion model (MM-LDM) for the SVG task. We first unify the representation of audio and video data by converting them into a single or a couple of images. Then, we introduce a hierarchical multi-modal autoencoder that constructs a low-level perceptual latent space for each modality and a shared high-level semantic feature space. The former space is perceptually equivalent to the raw signal space of each modality but drastically reduces signal dimensions. The latter space serves to bridge the information gap between modalities and provides more insightful cross-modal guidance. Our proposed method achieves new state-of-the-art results with significant quality and efficiency gains. Specifically, our method achieves a comprehensive improvement on all evaluation metrics and a faster training and sampling speed on Landscape and AIST++ datasets. Moreover, we explore its performance on open-domain sounding video generation, long sounding video generation, audio continuation, video continuation, and conditional single-modal generation tasks for a comprehensive evaluation, where our MM-LDM demonstrates exciting adaptability and generalization ability.
Mingzhen Sun, Weining Wang 0001, Yanyuan Qiao, Longteng Guo, Jing Liu 0001
ACM Multimedia2
2024 Open-Set Single-Domain Generalization for Robust Face Anti-Spoofing
Fangling Jiang, Qi Li 0005, Weining Wang 0001, Zhenan Sun
Int. J. Comput. Vis.3
2024 Learning Disentangled Representation for One-Shot Progressive Face Swapping
abstract
Although face swapping has attracted much attention in recent years, it remains a challenging problem. Existing methods leverage a large number of data samples to explore the intrinsic properties of face swapping without considering the semantic information of face images. Moreover, the representation of the identity information tends to be fixed, leading to suboptimal face swapping. In this paper, we present a simple yet efficient method named FaceSwapper, for one-shot face swapping based on Generative Adversarial Networks. Our method consists of a disentangled representation module and a semantic-guided fusion module. The disentangled representation module comprises an attribute encoder and an identity encoder, which aims to achieve the disentanglement of the identity and attribute information. The identity encoder is more flexible, and the attribute encoder contains more attribute details than its competitors. Benefiting from the disentangled representation, FaceSwapper can swap face images progressively. In addition, semantic information is introduced into the semantic-guided fusion module to control the swapped region and model the pose and expression more accurately. Experimental results show that our method achieves state-of-the-art results on benchmark datasets with fewer training samples.
Qi Li 0005, Weining Wang 0001, Cheng-Zhong Xu 0001, Zhenan Sun, Ming-Hsuan Yang 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2024 Reparameterizing and dynamically quantizing image features for image generation
Mingzhen Sun, Weining Wang 0001, Jing Liu 0001
Pattern Recognit.2
2024 Learnable Feature Augmentation Framework for Temporal Action Localization
abstract
Temporal action localization (TAL) has drawn much attention in recent years, however, the performance of previous methods is still far from satisfactory due to the lack of annotated untrimmed video data. To deal with this issue, we propose to improve the utilization of current data through feature augmentation. Given an input video, we first extract video features with pre-trained video encoders, and then randomly mask various semantic contents of video features to consider different views of video features. To avoid damaging important action-related semantic information, we further develop a learnable feature augmentation framework to generate better views of videos. In particular, a Mask-based Feature Augmentation Module (MFAM) is proposed. The MFAM has three advantages: 1) it captures the temporal and semantic relationships of original video features, 2) it generates masked features with indispensable action-related information, and 3) it randomly recycles some masked information to ensure diversity. Finally, we input the masked features and the original features into shared action detectors respectively, and perform action classification and localization jointly for model learning. The proposed framework can improve the robustness and generalization of action detectors by learning more and better views of videos. In the testing stage, the MFAM can be removed, which does not bring extra computational costs. Extensive experiments are conducted on four TAL benchmark datasets. Our proposed framework significantly improves different TAL models and achieves the state-of-the-art performances.
Yepeng Tang, Weining Wang 0001, Chunjie Zhang 0001, Jing Liu 0001, Yao Zhao 0001
IEEE Trans. Image Process.2
2024 Sounding Video Generator: A Unified Framework for Text-Guided Sounding Video Generation
abstract
As a combination of visual and audio signals, video is inherently multi-modal. However, existing video generation methods are primarily intended for the synthesis of visual frames, whereas audio signals in realistic videos are disregarded. In this work, we concentrate on a rarely investigated problem of text-guided sounding video generation and propose the Sounding Video Generator (SVG), a unified framework for generating realistic videos along with audio signals. Specifically, we present the SVG-VQGAN to transform visual frames and audio mel-spectrograms into discrete tokens. SVG-VQGAN applies a novel hybrid contrastive learning method to model inter-modal and intra-modal consistency and improve the quantized representations. A cross-modal attention module is employed to extract associated features of visual frames and audio signals for contrastive learning. Then, a Transformer-based decoder is used to model associations between texts, visual frames, and audio signals at token level for auto-regressive sounding video generation. AudioSet-Cap, a human annotated text-video-audio paired dataset, is produced for training SVG. Experimental results demonstrate the superiority of our method when compared with existing text-to-video generation methods as well as audio generation methods on Kinetics and VAS datasets.
Jiawei Liu 0001, Weining Wang 0001, Jing Liu 0001
IEEE Trans. Multim.2
2024 Temporal Action Proposal Generation With Action Frequency Adaptive Network
abstract
As the cornerstone of human-behavior analysis in video understanding, temporal action proposal generation aims to predict the starting and ending time of human action instances in untrimmed videos. Although large achievements in temporal action proposal generation have been achieved, most previous studies ignore the variability of action frequency in raw videos, leading to unsatisfying performances on high-action-frequency videos. In fact, there exists two main issues which should be well addressed: data imbalance between high and low action-frequency videos, and inferior detection of short actions in high-action-frequency videos. To address the above issues, we propose an effective framework by adapting to the variability of action frequency, namely Action Frequency Adaptive Network (AFAN), which can be flexibly built upon any temporal action proposal generation method. AFAN consists of two modules: Learning From Experts (LFE) and Fine-Grained Processing (FGP). The LFE first trains a series of action proposal generators on different subsets of imbalanced data as experts and then teaches a unified student model via knowledge distillation. To better detect short actions, FGP first finds out high-action-frequency videos and then performs fine-grained detection. Extensive experimental results on four benchmark datasets (ActivityNet-1.3, HACS, THUMOS14 and FineAction) demonstrate the effectiveness and generalizability of the proposed AFAN, especially for high-action-frequency videos.
Yepeng Tang, Weining Wang 0001, Chunjie Zhang 0001, Jing Liu 0001, Yao Zhao 0001
IEEE Trans. Multim.2
2023 MOSO: Decomposing MOtion, Scene and Object for Video Prediction
abstract
Motion, scene and object are three primary visual components of a video. In particular, objects represent the foreground, scenes represent the background, and motion traces their dynamics. Based on this insight, we propose a two-stage MOtion, Scene and Object decomposition framework (MOSO)11Codes have been released in https://github.com/iva-mzsun/MOSO for video prediction, consisting of MOSO-VQVAE and MOSO-Transformer. In the first stage, MOSO-VQVAE decomposes a previous video clip into the motion, scene and object components, and represents them as distinct groups of discrete tokens. Then, in the second stage, MOSO-Transformer predicts the object and scene tokens of the subsequent video clip based on the previous tokens and adds dynamic motion at the token level to the generated object and scene tokens. Our framework can be easily extended to unconditional video generation and video frame interpolation tasks. Experimental results demonstrate that our method achieves new state-of-the-art performance on five challenging benchmarks for video prediction and unconditional video generation: BAIR, RoboNet, KTH, KITTI and UCF101. In addition, MOSO can produce realistic videos by combining objects and scenes from different videos.
Mingzhen Sun, Weining Wang 0001, Jing Liu 0001
CVPR2
2023 WL-MSR: Watch and Listen for Multimodal Subtitle Recognition
abstract
Video subtitles could be defined as the combination of visualized subtitles in frames and textual content recognized from speech, which play a significant role in video understanding for both humans and machines. In this paper, we propose a novel Watch and Listen for Multimodal Subtitle Recognition (WL-MSR) framework to obtain comprehensive video subtitles, by fusing the information provided by Optical Character Recognition (OCR) and Automatic Speech Recognition (ASR) models. Specifically, we build a Transformer model with mask and crop strategies and multi-level identity embeddings to aggregate both the textual results and features of the two modalities. To pre-filter out the noise items in OCR results before fusion, we adopt an OCR filter based on ASR results and confidence scores of OCR. By combining these techniques, our solution wins the 2nd place in Multimodal Subtitle Recognition Challenge on ICPR2022.
Jiawei Liu 0001, Weining Wang 0001, Xingjian He, Jing Liu 0001
ICASSP3
2023 ED-T2V: An Efficient Training Framework for Diffusion-based Text-to-Video Generation
abstract
Diffusion models have achieved remarkable performance on image generation. However, It is difficult to reproduce this success on video generation because of expensive training cost. In fact, pretrained image generation models have already acquired visual generation capabilities and could be utilized for video generation. Thus, we propose an Efficient training framework for Diffusion-based Text-to-Video generation (ED-T2V), which is built on a pretrained text-to-image generation model. To model the temporal dynamic information, we propose temporal transformer blocks with novel identity attention and temporal cross-attention. ED-T2V has the following advantages: 1) most of the parameters of pretrained model are frozen to inherit the generation capabilities and reduce the training cost; 2) the identity attention requires the currently generated frame to attend to all positions of its previous frame, thus providing an efficient way to keep main content consistent across frames and enable movement generation; 3) the temporal cross-attention is proposed to construct associations between textual descriptions and multiple video tokens in the time dimension, which could better model video movement than traditional cross-attention methods. With the aforementioned benefits, ED-T2V not only significantly reduces the training cost of video diffusion models, but also has excellent generation fidelity and controllability.
Jiawei Liu 0001, Weining Wang 0001, Jing Liu 0001
IJCNN2
2023 GLOBER: Coherent Non-autoregressive Video Generation via GLOBal Guided Video DecodER
abstract
Video generation necessitates both global coherence and local realism. This work presents a novel non-autoregressive method GLOBER, which first generates global features to obtain comprehensive global guidance and then synthesizes video frames based on the global features to generate coherent videos. Specifically, we propose a video auto-encoder, where a video encoder encodes videos into global features, and a video decoder, built on a diffusion model, decodes the global features and synthesizes video frames in a non-autoregressive manner. To achieve maximum flexibility, our video decoder perceives temporal information through normalized frame indexes, which enables it to synthesize arbitrary sub video clips with predetermined starting and ending frame indexes. Moreover, a novel adversarial loss is introduced to improve the global coherence and local realism between the synthesized video frames. Finally, we employ a diffusion-based video generator to fit the global features outputted by the video encoder for video generation. Extensive experimental results demonstrate the effectiveness and efficiency of our proposed method, and new state-of-the-art results have been achieved on multiple benchmarks.
Mingzhen Sun, Weining Wang 0001, Jing Liu 0001
NeurIPS2
2023 Anchor-free temporal action localization via Progressive Boundary-aware Boosting
Yepeng Tang, Weining Wang 0001, Chunjie Zhang 0001, Jing Liu 0001
Inf. Process. Manag.2
2023 CASIA-E: A Large Comprehensive Dataset for Gait Recognition
abstract
Gait recognition plays a special role in visual surveillance due to its unique advantage, e.g., long-distance, cross-view and non-cooperative recognition. However, it has not yet been widely applied. One reason for this awkwardness is the lack of a truly big dataset captured in practical outdoor scenarios. Here, the "big" at least means: (1) huge amount of gait videos; (2) sufficient subjects; (3) rich attributes; and (4) spatial and temporal variations. Moreover, most existing large-scale gait datasets are collected indoors, which have few challenges from real scenes, such as the dynamic and complex background clutters, illumination variations, vertical view variations, etc. In this article, we introduce a newly built big outdoor gait dataset, called CASIA-E. It contains more than one thousand people distributed over near one million videos. Each person involves 26 view angles and varied appearances caused by changes of bag carrying, dressing and walking styles. The videos are captured across five months and across three kinds of outdoor scenes. Soft biometric features are also recorded for all subjects including age, gender, height, weight, and nationality. Besides, we report an experimental benchmark and examine some meaningful problems that have not been well studied previously, e.g., the influence of million-level training videos, vertical view angles, walking styles, and the thermal infrared modality. We believe that such a big outdoor dataset and the experimental benchmark will promote the development of gait recognition in both academic research and industrial applications.
Chunfeng Song, Yongzhen Huang, Weining Wang 0001, Liang Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 Semantic-based conditional generative adversarial hashing with pairwise labels
Qi Li 0005, Weining Wang 0001, Yuan Yan Tang, Cheng-Zhong Xu 0001, Zhenan Sun
Pattern Recognit.2
2022 Super-resolution semantic segmentation with relation calibrating network
Jie Jiang 0016, Jing Liu 0001, Jun Fu 0005, Weining Wang 0001, Hanqing Lu
Pattern Recognit.4
2022 An Efficient Sampling-Based Attention Network for Semantic Segmentation
abstract
Self-attention is widely explored to model long-range dependencies in semantic segmentation. However, this operation computes pair-wise relationships between the query point and all other points, leading to prohibitive complexity. In this paper, we propose an efficient Sampling-based Attention Network which combines a novel sample method with an attention mechanism for semantic segmentation. Specifically, we design a Stochastic Sampling-based Attention Module (SSAM) to capture the relationships between the query point and a stochastic sampled representative subset from a global perspective, where the sampled subset is selected by a Stochastic Sampling Module. Compared to self-attention, our SSAM achieves comparable segmentation performance while significantly reducing computational redundancy. In addition, with the observation that not all pixels are interested in the contextual information, we design a Deterministic Sampling-based Attention Module (DSAM) to sample features from a local region for obtaining the detailed information. Extensive experiments demonstrate that our proposed method can compete or perform favorably against the state-of-the-art methods on the Cityscapes, ADE20K, COCO Stuff, and PASCAL Context datasets.
Xingjian He, Jing Liu 0001, Weining Wang 0001, Hanqing Lu
IEEE Trans. Image Process.3
2022 Semi-Supervised Temporal Action Proposal Generation via Exploiting 2-D Proposal Map
abstract
Temporal action proposal generation aims to generate temporal video segments containing human actions in untrimmed videos, which is always a preliminary for such video understanding tasks as action localization and temporally description grounding,etc. Fully-supervised solutions, though proven to be effective, suffer much from heavy data annotation overhead. To address this problem, this paper focuses on a rarely investigated yet practical problem of semi-supervised learning for temporal action proposal generation. Firstly, we propose aProposal Map oriented Mean-Teacher(PM-MT) model, which can use both labeled and unlabeled data for end-to-end model training. Secondly, aSuppression-and-Re-Generation(SRG) strategy is designed to generate high-quality pseudo labels for unlabeled data, which are then used to finetune the model. Extensive experiments demonstrate the effectiveness of our proposed method, by achieving the state-of-the-art results on two public benchmark datatsets on the task of semi-supervised action proposal generation and outperforming fully-supervised learning methods with only a portion of labeled data.
Weining Wang 0001, Dongliang He, Fu Li 0003, Shilei Wen, Liang Wang 0001, Jing Liu 0001
IEEE Trans. Multim.1
2021 CAS-AIR-3D Face: A Low-Quality, Multi-Modal and Multi-Pose 3D Face Database
abstract
Benefiting from deep learning with large scale face databases, 2D face recognition has made significant progress in recent years. However, it still highly depends on lighting conditions and human poses, and suffers from face spoofing problem. In contrast, 3D face recognition reveals a new path that can overcome the previous limitations of 2D face recognition. One of the most important problems for 3D face recognition is to construct a suitable database, which can be exploited to train different 3D face recognition algorithms. In this work, we propose a new database, CAS-AIR-3D Face, for low-quality 3D face recognition. It includes 24713 videos from 3093 individuals, which is captured by Intel RealSense SR305. The database contains three modalities: color, depth and near infrared, and is rich in pose, expression, occlusion and distance variations. To the best of our konwledge, CAS-AIR-3D Face is the largest low-quality 3D face database in terms of the number of individuals and the sample variations. Moreover, we preprocess the data via a sophisticated face alignment method, and Point Cloud Spherical Cropping Method (SCM) is leveraged to remove the background noise in the depth images. Finally, an evaluation protocol is designed for fair comparison, and extensive experiments are conducted with different backbone networks to provide different baselines on this database.
Qi Li 0005, Xiaoxiao Dong, Weining Wang 0001, Caifeng Shan
IJCB3
2021 Face Sketch Synthesis via Semantic-Driven Generative Adversarial Network
abstract
Face sketch synthesis has made significant progress with the development of deep neural networks in these years. The delicate depiction of sketch portraits facilitates a wide range of applications like digital entertainment and law enforcement. However, accurate and realistic face sketch generation is still a challenging task due to the illumination variations and complex backgrounds in the real scenes. To tackle these challenges, we propose a novel Semantic-Driven Generative Adversarial Network (SDGAN) which embeds global structure-level style injection and local class-level knowledge re-weighting. Specifically, we conduct facial saliency detection on the input face photos to provide overall facial texture structure, which could be used as a global type of prior information. In addition, we exploit face parsing layouts as the semantic-level spatial prior to enforce globally structural style injection in the generator of SDGAN. Furthermore, to enhance the realistic effect of the details, we propose a novel Adaptive Re-weighting Loss (ARLoss) which dedicates to balance the contributions of different semantic classes. Experimentally, our extensive experiments on CUFS and CUFSF datasets show that our proposed algorithm achieves state-of-the-art performance.
Xingqun Qi, Muyi Sun, Weining Wang 0001, Xiaoxiao Dong, Qi Li 0005, Caifeng Shan
IJCB3
2021 HAIR: Hierarchical Visual-Semantic Relational Reasoning for Video Question Answering
abstract
Relational reasoning is at the heart of video question answering. However, existing approaches suffer from several common limitations: (1) they only focus on either object-level or frame-level relational reasoning, and fail to integrate the both; and (2) they neglect to leverage semantic knowledge for relational reasoning. In this work, we propose a Hierarchical VisuAl-Semantic RelatIonal Reasoning (HAIR) framework to address these limitations. Specifically, we present a novel graph memory mechanism to perform relational reasoning, and further develop two types of graph memory: a) visual graph memory that leverages visual information of video for relational reasoning; b) semantic graph memory that is specifically designed to explicitly leverage semantic knowledge contained in the classes and attributes of video objects, and perform relational reasoning in the semantic space. Taking advantage of both graph memory mechanisms, we build a hierarchical framework to enable visual-semantic relational reasoning from object level to frame level. Experiments on four challenging benchmark datasets show that the proposed framework leads to state-of-the-art performance, with fewer parameters and faster inference speed. Besides, our approach also shows superior performance on other video+language task.
Fei Liu 0047, Jing Liu 0001, Weining Wang 0001, Hanqing Lu
ICCV3
2021 Keypoint Context Aggregation for Human Pose Estimation
Wenzhu Wu, Weining Wang 0001, Longteng Guo, Jing Liu 0001
ICIG (2)2
2021 Temporal Memory Attention for Video Semantic Segmentation
abstract
Video semantic segmentation requires to utilize the complex temporal relations between frames of the video sequence. Previous works usually exploit accurate optical flow to leverage the temporal relations, which suffer much from heavy computational cost. In this paper, we propose a Temporal Memory Attention Network (TMANet) to adaptively integrate the long-range temporal relations over the video sequence based on the self-attention mechanism without exhaustive optical flow prediction. Specially, we construct a memory using several past frames to store the temporal information of the current frame. We then propose a temporal memory attention module to capture the relation between the current frame and the memory to enhance the representation of the current frame. Our method achieves new state-of-the-art performances on two challenging video semantic segmentation datasets, particularly 80.3% mIoU on Cityscapes and 76.5% mIoU on CamVid with ResNet-50.
Weining Wang 0001, Jing Liu 0001
ICIP2
2021 Multi-caption Text-to-Face Synthesis: Dataset and Algorithm
abstract
Text-to-Face synthesis with multiple captions is still an important yet less addressed problem because of the lack of effective algorithms and large-scale datasets. We accordingly propose a Semantic Embedding and Attention (SEA-T2F) network that allows multiple captions as input to generate highly semantically related face images. With a novel Sentence Features Injection Module, SEA-T2F can integrate any number of captions into the network. In addition, an attention mechanism named Attention for Multiple Captions is proposed to fuse multiple word features and synthesize fine-grained details. Considering text-to-face generation is an ill-posed problem, we also introduce an attribute loss to guide the network to generate sentence-related attributes. Existing datasets for text-to-face are either too small or roughly generated according to attribute labels, which is not enough to train deep learning based methods to synthesize natural face images. Therefore, we build a large-scale dataset named CelebAText-HQ, in which each image is manually annotated with 10 captions. Extensive experiments demonstrate the effectiveness of our algorithm.
Jianxin Sun 0003, Qi Li 0005, Weining Wang 0001, Jian Zhao 0006, Zhenan Sun
ACM Multimedia3
2020 Long video question answering: A Matching-guided Attention Model
Weining Wang 0001, Yan Huang 0008, Liang Wang 0001
Pattern Recognit.1
2019 Language-Driven Temporal Activity Localization: A Semantic Matching Reinforcement Learning Model
abstract
Current studies on action detection in untrimmed videos are mostly designed for action classes, where an action is described at word level such as jumping, tumbling, swing, etc. This paper focuses on a rarely investigated problem of localizing an activity via a sentence query which would be more challenging and practical. Considering that current methods are generally time-consuming due to the dense frame-processing manner, we propose a recurrent neural network based reinforcement learning model which selectively observes a sequence of frames and associates the given sentence with video content in a matching-based manner. However, directly matching sentences with video content performs poorly due to the large visual-semantic discrepancy. Thus, we extend the method to a semantic matching reinforcement learning (SM-RL) model by extracting semantic concepts of videos and then fusing them with global context features. Extensive experiments on three benchmark datasets, TACoS, Charades-STA and DiDeMo, show that our method achieves the state-of-the-art performance with a high detection speed, demonstrating both effectiveness and efficiency of our method.
Weining Wang 0001, Yan Huang 0008, Liang Wang 0001
CVPR1