Xi Wang 0014

dblp:08/5760-14 · DBLP profile ↗
← Back
28ranked-venue papers
4as first author
21since 2021 · last 2025
0000-0002-6642-8160ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 21 · 4 first-author · 15 since 2021Artificial intelligence and machine learning · 6 · 4 since 2021Security and privacy · 3 · 3 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Modality-Agnostic Deepfakes Detection
abstract
As AI-generated content (AIGC) thrives, deepfakes have expanded from single-modality falsification to cross-modal fake content creation, where either audio or visual components can be manipulated.While using two unimodal detectors can detect audio-visual deepfakes, cross-modal forgery clues could be overlooked.Existing multimodal deepfake detectors typically establish correspondence between the audio and visual modalities for binary real/fake classification and require the co-occurrence of both modalities.However, in real-world multi-modal applications, missing modality scenarios may occur where either modality is unavailable.In such cases, audio-visual detection methods are less practical than two independent unimodal methods.Consequently, the detector can not always obtain the number or type of manipulated modalities beforehand, necessitating a fake-modality-agnostic audio-visual detector.In this work, we introduce a comprehensive framework that is agnostic to fake modalities, which facilitates the identification of multimodal deepfakes and handles situations with missing modalities, regardless of the manipulations embedded in audio, video, or even cross-modal forms.To enhance the modeling of cross-modal forgery clues, we employ audio-visual speech recognition (AVSR)
Jin Liu 0020, Jiao Dai, Xi Wang 0014, Shan Jia, Siwei Lyu, Jizhong Han
IH&MMSec6
2025 OMS: One More Step Noise Searching to Enhance Membership Inference Attacks for Diffusion Models
abstract
The data-intensive nature of Diffusion models amplifies the risks of privacy infringements and copyright disputes, particularly when training on extensive unauthorized data scraped from the Internet. Membership Inference Attacks (MIA) aim to determine whether a data sample has been utilized by the target model during training, thereby serving as a pivotal tool for privacy preservation. Current MIA employs the prediction loss to distinguish between training member samples and non-members. These methods assume that, compared to non-members, members, having been encountered by the model during training result in a smaller prediction loss. However, this assumption proves ineffective in diffusion models due to the random noise sampled during the training process. Rather than estimating the loss, our approach examines this random noise and reformulate the MIA as a noise search problem, assuming that members are more feasible to find the noise used in the training process. We formulate this noise search process as an optimization problem and employ the fixed-point iteration to solve it. We analyze current MIA methods through the lens of the noise search framework and reveal that they rely on the first residual as the discriminative metric to differentiate members and non-members. Inspired by this observation, we introduce OMS, which augments existing MIA methods by iterating One More fixed-point Step to include a further residual, i.e., the second residual. We integrate our method into various MIA methods across different diffusion models. The experimental results validate the efficacy of our proposed approach.
Xiaomeng Fu, Xi Wang 0014, Jin Liu 0020, Jiao Dai, Jizhong Han, Xingyu Gao 0001
IJCAI2
2025 Unlocking Generative Priors: A New Membership Inference Framework for Diffusion Models
abstract
Diffusion models pose risks of privacy breaches and copyright disputes, primarily stemming from the potential utilization of unauthorized data during the training phase. Membership inference is aimed to determine whether a specific sample has been used in the training process of a target model, representing a critical tool for privacy violation verification. However, the increased model complexity and stochasticity inherent in diffusion renders traditional shadow-model-based or metric-based methods ineffective when applied to diffusion models. Moreover, existing methods only yield binary classification labels which lack necessary comprehensibility in practical applications. In this paper, we explore a novel perspective for membership inference by leveraging the intrinsic generative priors within the diffusion model. Compared with unseen samples, training samples exhibit stronger generative priors within the diffusion model, enabling the successful reconstruction of substantially degraded training images. Consequently, we propose the Degrade Restore Compare (DRC) framework. In this framework, an image undergoes sequential degradation and restoration, and its membership is determined by comparing it with the restored counterpart. Experimental results verify that our approach not only significantly outperforms existing methods in terms of accuracy but also provides comprehensible decision criteria, offering evidence for potential privacy violations.
Xiaomeng Fu, Xi Wang 0014, Jin Liu 0020, Jiao Dai, Jizhong Han, Xingyu Gao 0001
IEEE Trans. Inf. Forensics Secur.2
2025 Multimodal Consistency Suppression Factor for Fake News Detection
abstract
Recent multimodal fake news detection methods often use the consistency between textual and visual contents to determine the truth or fake of news information. Higher levels of textual-visual consistency typically lead to a greater likelihood of classifying a news item as real. However, a critical observation reveals that creators of most fake news intentionally select images that align with the textual content, thereby enhancing the credibility of the news. Consequently, high consistency between textual and visual contents alone cannot guarantee the authenticity of the information. To address this problem, we introduce a novel approach termed Multimodal Consistency-based Suppression Factor to modulate the significance of textual-visual consistency in information assessment. When the textual-visual matching is high, this suppression factor reduces the influence of consistency during the judgment process. Moreover, we use contrastive language-image pre-training (CLIP) model to extract features and measure the consistency level between modalities to guide multimodal fusion. In addition, we also use a method of compressing and fusing modal information based on variational autoencoder (VAE) to reconstruct CLIP features, learning the shared representation of different modal information of CLIP. Finally, extensive experiments were conducted on three publicly datasets, Weibo, Twitter, and Weibo21, and the results confirmed that our method outperformed the state-of-the-art methods in the field and had 0.8%, 2.6%, and 4.1% effect improvement on the accuracy rate.
Zhulin Tao, Xingyu Gao 0001, Xi Wang 0014, Xianglin Huang
ACM Trans. Multim. Comput. Commun. Appl.5
2024 Generative Transferable Universal Adversarial Perturbation for Combating Deepfakes
abstract
Recently, Deepfake has posed a significant threat to our digital society. This technology allows for the modification of facial identity, expression, and attributes in facial images and videos. The misuse of Deepfake can invade personal privacy, damage individuals’ reputations, and have serious consequences. To counter this threat, researchers have proposed active defense methods using adversarial perturbation to distort Deepfake products which can hinder the dissemination of false information. However, the existing methods are primarily based on image-specific approaches, which are inefficient for large-scale data. To address these issues, we propose an end-to-end approach to generate universal perturbations for combating Deepfake. To further cope with diverse Deepfakes, we introduce an adaptive balancing strategy to combat multiple models simultaneously. Specifically, for different scenarios, we propose two types of universal perturbations. Disrupting Universal Perturbation (DUP) leads Deepfake models to generate distorted outputs. In contrast, Lapsing Universal Perturbation (LUP) tries to make the output consistent with the original image, allowing the correct information to continue propagating. Experiments demonstrate the effectiveness and better generalization of our proposed perturbation compared with state-of-the-art methods. Consequently, our proposed method offers a powerful and efficient solution for combating Deepfake, which can help preserve personal privacy and prevent reputational damage.
Xi Wang 0014, Xiaomeng Fu, Jin Liu 0020, Zhaoxing Li, Yesheng Chai, Jizhong Han
CSCWD2
2024 Real Appearance Modeling for More General Deepfake Detection
Cai Yu, Xi Wang 0014, Zihao Xiao 0002, Jiao Dai, Jizhong Han, Yesheng Chai
ECCV (52)3
2024 Generative Universal Nullifying Perturbation for Countering Deepfakes Through Combined Unsupervised Feature Aggregation
Xi Wang 0014, Xiaomeng Fu, Jin Liu 0020, Zhaoxing Li, Jizhong Han
ICANN (2)2
2024 ConfR: Conflict Resolving for Generalizable Deepfake Detection
abstract
Deepfake detectors often encounter performance degradation when tested on unseen forgery methods. Existing literature tries to capture common features among multiple source forgery domains. However, we show that conflict arises in the shared feature space when each domain expresses domain-specific bias. If left unresolved, this conflict might mislead the model to learn domain-specific features and lead to inferior generalization. In this paper, we propose a new learning approach, Conflict Resolving (ConfR), designed to minimize conflict and learn features that generalize across forgeries. ConfR incorporates two key elements: the Intra-Domain Consistency Preserving (ICP) loss ensures updating consistency within forgery types, and the Inter-Domain Conflict Resolving (ICR) Module resolves updating conflicts between different forgery types. Extensive experiments demonstrate that ConfR significantly improves upon the state-of-the-art method, highlighting its potential for more generalizable deepfake detection.
Cai Yu, Xi Wang 0014, Zhaoxing Li, Yesheng Chai, Jiao Dai, Jizhong Han
ICME4
2024 Explicit Correlation Learning for Generalizable Cross-Modal Deepfake Detection
abstract
With the rising prevalence of deepfakes, there is a growing interest in developing generalizable detection methods for various types of deepfakes. While effective in their specific modalities, traditional detection methods fall short in addressing the generalizability of detection across diverse cross-modal deepfakes. This paper aims to explicitly learn potential cross-modal correlation to enhance deepfake detection towards various generation scenarios. Our approach introduces a correlation distillation task, which models the inherent cross-modal correlation based on content information. This strategy helps to prevent the model from overfitting merely to audio-visual synchronization. Additionally, we present the Cross-Modal Deepfake Dataset (CMDFD), a comprehensive dataset with four generation methods to evaluate the detection of diverse cross-modal deepfakes. The experimental results on CMDFD and FakeAVCeleb datasets demonstrate the superior generalizability of our method over existing state-of-the-art methods. Our code and data can be found at https://github.com/ljj898/CMDFD-Dataset-and-Deepfake-Detection.
Cai Yu, Shan Jia, Xiaomeng Fu, Jin Liu 0020, Jiao Dai, Xi Wang 0014, Siwei Lyu, Jizhong Han
ICME7
2024 Unveiling Structural Memorization: Structural Membership Inference Attack for Text-to-Image Diffusion Models
abstract
With the rapid advancements of large-scale text-to-image diffusion models, various practical applications have emerged, bringing significant convenience to society. However, model developers may misuse the unauthorized data to train diffusion models. These data are at risk of being memorized by the models, thus potentially violating citizens' privacy rights. Therefore, in order to judge whether a specific image is utilized as a member of a model's training set, Membership Inference Attack (MIA) is proposed to serve as a tool for privacy protection. Current MIA methods predominantly utilize pixel-wise comparisons as distinguishing clues, considering the pixel-level memorization characteristic of diffusion models. However, it is practically impossible for text-to-image models to memorize all the pixel-level information in massive training sets. Therefore, we move to the more advanced structure-level memorization. Observations on the diffusion process show that the structures of members are better preserved compared to those of nonmembers, indicating that diffusion models possess the capability to remember the structures of member images from training sets. Drawing on these insights, we propose a simple yet effective MIA method tailored for text-to-image diffusion models. Extensive experimental results validate the efficacy of our approach. Compared to current pixel-level baselines, our approach not only achieves state-of-the-art performance but also demonstrates remarkable robustness against various distortions.
Xiaomeng Fu, Xi Wang 0014, Jin Liu 0020, Xingyu Gao 0001, Jiao Dai, Jizhong Han
ACM Multimedia3
2024 Dynamic Mixed-Prototype Model for Incremental Deepfake Detection
abstract
The rapid advancement of deepfake technology poses significant threats to social trust. Although recent deepfake detectors have exhibited promising results on deepfakes of the same type as those present in training, their effectiveness degrades significantly on novel deepfakes crafted by unseen algorithms due to the gap in forgery patterns. Some studies have enhanced detectors by adapting to the continuously emerging deepfakes through incremental learning. Despite the progress, they overlooked the scarcity of novel samples that can easily lead to insufficient learning of forgery patterns. To mitigate this issue, we introduce the Dynamic Mixed-Prototype (DMP) model, which dynamically increases prototypes to adapt to novel deepfakes efficiently. Specifically, the DMP model adopts multiple prototypes to represent both real and fake classes, enabling learning novel patterns by expanding prototypes and jointly retaining knowledge learned in previous prototypes. Furthermore, we propose the Prototype-Guided Replay strategy and Prototype Representation Distillation loss, both of which effectively prevent forgetting learned knowledge based on the prototypical representation of samples. Our method surpasses existing incremental deepfake detectors across four datasets and can generalize to novel deepfakes by learning limited deepfake samples.
Cai Yu, Xi Wang 0014, Zihao Xiao 0002, Jizhong Han, Yesheng Chai
ACM Multimedia3
2024 OSM-Net: One-to-Many One-Shot Talking Head Generation With Spontaneous Head Motions
abstract
One-shot talking head generation has no explicit head movement reference, thus it is difficult to generate talking heads with head motions. Some existing works only edit the mouth area and generate still talking heads, leading to unreal talking head performance. Other works construct one-to-one mapping between audio signal and head motion sequences, introducing ambiguity correspondences into the mapping since people can behave differently in head motions when speaking the same content. This unreasonable mapping form fails to model the diversity and produces either nearly static or even exaggerated head motions, which are unnatural and strange. Therefore, the one-shot talking head generation task is actually a one-to-many ill-posed problem and people present diverse head motions when speaking. Based on the above observation, we propose OSM-Net, aone-to-manyone-shot talking head generation network with natural head motions. OSM-Net constructs a motion space that contains rich and various clip-level head motion features. Each basis of the space represents a feature of meaningful head motion in a clip rather than just a frame, thus providing more coherent and natural motion changes in talking heads. The driving audio is mapped into the motion space, around which various motion features can be sampled within a reasonable range to achieve the one-to-many mapping. Besides, the landmark constraint and time window feature input improve the accurate expression feature extraction and video generation. Extensive experiments show that OSM-Net generates more natural realistic head motions under reasonable one-to-many mapping paradigm compared with other methods.
Jin Liu 0020, Xi Wang 0014, Xiaomeng Fu, Yesheng Chai, Cai Yu, Jiao Dai, Jizhong Han
IEEE Trans. Circuits Syst. Video Technol.2
2024 Learning to Discover Forgery Cues for Face Forgery Detection
abstract
Locating manipulation maps,i.e., pixel-level annotation of forgery cues, is crucial for providing interpretable detection results in face forgery detection. Related learning objects have also been widely adopted as auxiliary tasks to improve the classification performance of detectors whereas they require comparisons between paired real and forged faces to obtain manipulation maps as supervision. This requirement restricts their applicability to unpaired faces and contradicts real-world scenarios. Moreover, the used comparison methods annotate all changed pixels, including noise introduced by compression and upsampling. Using such maps as supervision hinders the learning of exploitable cues and makes models prone to overfitting. To address these issues, we introduce a weakly supervised model in this paper, named Forgery Cue Discovery (FoCus), to locate forgery cues in unpaired faces. Unlike some detectors that claim to locate forged regions in attention maps, FoCus is designed to sidestep their shortcomings of capturing partial and inaccurate forgery cues. Specifically, we propose a classification attentive regions proposal module to locate forgery cues during classification and a complementary learning module to facilitate the learning of richer cues. The produced manipulation maps can serve as better supervision to enhance face forgery detectors. Visualization of the manipulation maps of the proposed FoCus exhibits superior interpretability and robustness compared to existing methods. Experiments on five datasets and four multi-task models demonstrate the effectiveness of FoCus in both in-dataset and cross-dataset evaluations.
Cai Yu, Xiaomeng Fu, Xi Wang 0014, Jiao Dai, Jizhong Han
IEEE Trans. Inf. Forensics Secur.5
2024 Knowledge Enhanced Vision and Language Model for Multi-Modal Fake News Detection
abstract
The rapid dissemination of fake news and rumors through the Internet and social media platforms poses significant challenges and raises concerns in the public sphere. Automatic detection of fake news plays a crucial role in mitigating the spread of misinformation. While recent approaches have focused on leveraging neural networks to improve textual and visual representations in multi-modal fake news analysis, they often overlook the potential of incorporating knowledge information to verify facts within news articles. In this paper, we propose a knowledge enhanced vision and language model for multi-modal fake news detection. Our proposed model integrates information from large scale open knowledge graphs to augment its ability to discern the veracity of news content. Unlike previous methods that utilize separate models to extract textual and visual features, we synthesize a unified model capable of extracting both types of features simultaneously. To represent news articles, we introduce a graph structure where nodes encompass entities, relationships extracted from the textual content, and objects depicted in associated images. By utilizing the knowledge graph, we establish meaningful relationships between nodes within the news articles. Experimental evaluations on a real-world multi-modal dataset from Twitter demonstrate significant performance improvement by incorporating knowledge information.
Xingyu Gao 0001, Xi Wang 0014, Zhenyu Chen 0003, Wei Zhou 0019, Steven C. H. Hoi
IEEE Trans. Multim.2
2023 OPT: One-shot Pose-Controllable Talking Head Generation
abstract
One-shot talking head generation produces lip-sync talking heads based on arbitrary audio and one source face. To guarantee the naturalness and realness, recent methods propose to achieve free pose control instead of simply editing mouth areas. However, existing methods do not preserve accurate identity of source face when generating head motions. To solve the identity mismatch problem and achieve high-quality free pose control, we present One-shot Pose-controllable Talking head generation network (OPT). Specifically, the Audio Feature Disentanglement Module separates content features from audios, eliminating the influence of speaker-specific information contained in arbitrary driving audios. Later, the mouth expression feature is extracted from the content feature and source face, during which the landmark loss is designed to enhance the accuracy of facial structure and identity preserving quality. Finally, to achieve free pose control, controllable head pose features from reference videos are fed into the Video Generator along with the expression feature and source face to generate new talking heads. Extensive quantitative and qualitative experimental results verify that OPT generates high-quality pose-controllable talking heads with no identity mismatch problem, outperforming previous SOTA methods.
Jin Liu 0020, Xi Wang 0014, Xiaomeng Fu, Yesheng Chai, Cai Yu, Jiao Dai, Jizhong Han
ICASSP2
2023 Large Pose Friendly Face Reenactment using subtle motions
abstract
Face reenactment aims to synthesis a photo-realistic video of the source face by imitating the motion and expression of the driving video while keeping the source appearance (i.e. identity). Although good results have achieved recently, most state-of-the-art methods remain vulnerable to extreme conditions, which greatly restricts the application in the real world. Among various extreme conditions, the large pose problem is the most common one. We clarify that the large pose problem is mainly caused by the severe motion change between the source image and the current driving frame. An intuitive solution is to divide the severe motion change into a sequence of subtle motions. Therefore, we propose a new scheme that exploring the temporal coherence between previous neighbor frame and current frame. The smaller motion change between consecutive frames help to solve the large pose problem. Furthermore, a calibration net is designed to eliminate the error accumulation of the previous step. Extensive experiments demonstrate that our method performs better on large pose face reenactment than the state-of-the-art in terms of large pose cases and visual quality.
Xiaomeng Fu, Xi Wang 0014, Jin Liu 0020, Jiao Dai, Jizhong Han
ICME2
2023 FONT: Flow-guided One-shot Talking Head Generation with Natural Head Motions
abstract
One-shot talking head generation has received growing attention in recent years, with various creative and practical applications. An ideal natural and vivid generated talking head video should contain natural head pose changes. However, it is challenging to map head pose sequences from driving audio since there exists a natural gap between audio-visual modalities. In this work, we propose a Flow-guided One-shot model that achieves NaTural head motions(FONT) over generated talking heads. Specifically, we design a probabilistic CVAE-based model to predict head pose sequences from driving audio and source face. Then we develop a keypoint predictor that produces unsupervised keypoints describing the facial structure information from the source face, driving audio and pose sequences. Finally, a flow- guided occlusion-aware generator is employed to produce photo-realistic talking head videos from the estimated keypoints and source face. Extensive experimental results prove that FONT generates talking heads with natural head poses and synchronized mouth shapes, outperforming other compared methods.
Jin Liu 0020, Xi Wang 0014, Xiaomeng Fu, Yesheng Chai, Cai Yu, Jiao Dai, Jizhong Han
ICME2
2023 MFR-Net: Multi-faceted Responsive Listening Head Generation via Denoising Diffusion Model
abstract
Face-to-face communication is a common scenario including roles of speakers and listeners. Most existing research methods focus on producing speaker videos, while the generation of listener heads remains largely overlooked. Responsive listening head generation is an important task that aims to model face-to-face communication scenarios by generating a listener head video given a speaker video and a listener head image. An ideal generated responsive listening video should respond to the speaker with attitude or viewpoint expressing while maintaining diversity in interaction patterns and accuracy in listener identity information. To achieve this goal, we propose the Multi-Faceted Responsive Listening Head Generation Network (MFR-Net). Specifically, MFR-Net employs the probabilistic denoising diffusion model to predict diverse head pose and expression features. In order to perform multi-faceted response to the speaker video, while maintaining accurate listener identity preservation, we design the Feature Aggregation Module to boost listener identity features and fuse them with other speaker-related features. Finally, a renderer finetuned with identity consistency loss produces the final listening head videos. Our extensive experiments demonstrate that MFR-Net not only achieves multi-faceted responses in diversity and speaker identity information but also in attitude and viewpoint expression.
Jin Liu 0020, Xi Wang 0014, Xiaomeng Fu, Yesheng Chai, Cai Yu, Jiao Dai, Jizhong Han
ACM Multimedia2
2023 CoP: Chain-of-Pose for Image Animation in Large Pose Changes
abstract
Image animation involves generating a video of a source image imitating the pose of a driving video. Despite recent advancements in the image animation task, most state-of-the-art methods remain vulnerable to large pose changes. In cases of large pose changes, existing methods struggle to model the complex nonlinear motion and yield distorted results, which greatly restricts their application in the real world. To tackle this problem, we present a novel approach called Chain-of-Pose (CoP) that decomposes large pose changes into a sequence of intermediate pose changes. This enables us to handle simplified pose changes and improves the accuracy of pose estimation. Furthermore, to better preserve the appearance of the source object, we introduce the Appearance Refinement Module (ARM) that effectively integrates the appearance texture feature of the source image with the structural pose feature from the pose chain. Our experimental results demonstrate that our method qualitatively and quantitatively outperforms state-of-the-art approaches on four diverse datasets, comprising talking faces, human bodies, and pixel animals. Notably, our approach significantly improves video quality in the case of large object pose changes. Our code is attached to the supplementary material.
Xiaomeng Fu, Xi Wang 0014, Jin Liu 0020, Shuhui Wang, Jiao Dai, Jizhong Han
ACM Multimedia2
2022 Focus by Prior: Deepfake Detection Based on Prior-Attention
abstract
Nowadays advanced facial manipulation techniques produce deepfake videos more realistically, which makes deepfake detection more difficult. To capture subtle and intricate artifacts, recent works attempt to enhance low-level textural information by attention-based framework. However, these methods require complex simulated data or extra supervision. Highly dependent on training settings, these methods not only have high training costs but also are prone to overfitting. To address this issue, we propose a novel perspective of deepfake detection via so-called prior-attention. Specifically, we introduce prior textural information, such as edge and noise, to model the attention maps explicitly. Benefiting from these natural “attention maps”, our model significantly enhances discriminative information without additional supervision. Furthermore, we design a Feature Abstraction Block (FAB) to facilitate cross-layer features interaction and insert it into distinct layers of CNN to detect the inconsistencies at multiple spatial levels. Extensive experiments demonstrate that our method achieves performance comparable to state-of-the-art methods.
Cai Yu, Jiao Dai, Xi Wang 0014, Weibo Zhang, Jin Liu 0020, Jizhong Han
ICME4
2022 MakeItSmile: Detail-Enhanced Smiling Face Reenactment
abstract
Given a target face and a driving face, face reenactment aims to transfer attributes from the driving face to the target face. In the last decade, a great number of methods have been proposed to generate realistic reenacted faces. However, when these methods are applied to generate a smiling face, most of them can only get a mouth with blurry teeth, making the reenacted face unrealistic. This problem is mainly caused by incomplete tooth structure in the target face image under the setting of one-shot reenactment. In order to obtain smiling reenacted faces with detailed tooth structure, our method uses the tooth information from the driving face rather than the target face. Furthermore, to better represent the tooth structure and expressions of the driving face, we extract the texture with a carefully designed geometry-aware encoder. By training the encoder with tooth segmentation task and non-identity classification task, we acquire refined tooth representations and meanwhile derive the non-identity part of the driving face. We also design a specific generator to fuse the tooth texture features into the target face. Moreover, we add a mouth loss function to further ensure the high definition of the smiling reenacted face. We compare our method to existing state-of-the-art approaches. The experiments show that our method gets comparable results on non-smiling face reenactment and has superior performance on smiling face reenactment.
Xiaomeng Fu, Xi Wang 0014, Jin Liu 0020, Wantao Liu, Jiao Dai, Jizhong Han
IJCNN2
2019 Text Recognition using local correlation
Hongchao Gao, Xi Wang 0014, Jizhong Han, Ruixuan Li 0001
BMVC3
2019 Self-Representation Convolutional Neural Networks
abstract
The traditional convolutional neural networks (CNNs) learn numbers of fixed kernels (filters), which are used to obtain the representations of fixed patterns. Therefore, the knowledge representations of CNNs are limited to the number of kernels. In this paper, we present a Self-Representation Convolutional (SRC) layer to obtain richer knowledge representations of images by fully considering the self-similarity between adjacent pixels. SRC layer comprises a learnable local correlation measurement which measures the importance of adjacent pixels to the current pixel and two learnable linear parameters that perform linear projection on adjacent pixels and the weighted sum vectors, respectively. Compared with regular convolutional layers, the SRC layers can not only obtain comparable knowledge representations, but also reduce by a factor of 3× to 56× in the number of learnable parameters. Empirically, CNNs with SRC layers, called Self-Representation Convolutional Neural Networks (SRCNN), achieve strong performances on a range of visual datasets (SVHN, CIFAR-10 and CIFAR-100) while enjoying significant parameters and FLOPs savings.
Hongchao Gao, Xi Wang 0014, Jizhong Han, Songlin Hu 0001, Ruixuan Li 0001
ICME2
2019 Ensemble Attention For Text Recognition In Natural Images
abstract
Recognizing text from natural images is a challenging and hot research topic in computer vision, yet not completely solved. The recent methods regard this task as a sequence labeling problem. In this task, there is a strong correspondence between the position of the input image patches sequence and the output character sequence. However, most of the recent recognition systems rarely consider this local information of the input sequence when recognizing the current character. In contrast to this, we present a Local Restricted Attention (LRA) mechanism to encode the current vector by considering adjacent vectors of the input sequence. We propose an ensemble decoder block which combines LRA mechanism with a regular decoder mechanism. This block not only brings significant improvement of recognition results under shorter training time but also can be easily embedded in other recognition frameworks. In addition, we propose a scene text recognition network based on the ensemble decoder. The experimental performances show that the proposed model achieves the state-of-the-art on several benchmark datasets including IIIT-5K, SVT, CUTE80, SVT-Perspective and ICDARs.
Hongchao Gao, Xi Wang 0014, Jizhong Han, Ruixuan Li 0001
IJCNN3
2014 Face Distortion Recovery Based on Online Learning Database for Conversational Video
abstract
With the real-time requirement for video conversation, the coding system needs to adopt low delay and low complexity strategy to encode conversational videos, which may result in a significant decline of the video quality under the constrained bandwidth of network. In conversational videos, the face region attracts most human attentions. Therefore, recovering the distortion of face region will effectively improve the visual quality of conversational video. Actually, the participants in a conversation are usually unchanged in a relative long period, and similar facial expressions of the participants would be often repetitive. However, conventional video coding methods just consider the correlation of several neighboring frames while the long-range correlation of similar face regions in the whole conversational video has not been fully used. In this paper, we propose a face distortion recovery system to improve the visual quality of decoded conversational video by online learning an own face feature database for each user. First, at the sender side, the face feature database is established and online updated to include different facial expressions of the person. Then, at the receiver side, the low quality face regions in decoded video are recovered with the face patches in the database. Experimental results show that, under low bits rates the proposed method achieves average 5.22 dB gain with small burden to update the database.
Xi Wang 0014, Li Su 0003, Honggang Qi, Qingming Huang, Guorong Li
IEEE Trans. Multim.1
2013 Online Learning Based Face Distortion Recovery for Conversational Video Coding
abstract
In a video conversation, the participants usually remain the same. As the conversation continues, similar facial expressions of the same person would occur intermittently. However, the correlation of similar face features has not been fully used since the conventional methods only focus on independent frames. We set up a face feature database and updated it online to include new facial expressions during the whole conversation. At the receiver side, the database is used to recover the face distortion and thus improve the visual quality. Additionally, the proposed method brings small burden to update the database and is generic to various CODEC.
Xi Wang 0014, Li Su 0003, Qingming Huang, Guorong Li, Honggang Qi
DCC1
2012 Motion Based Perceptual Distortion and Rate Optimization for Video Coding
abstract
Most conventional distortion metrics regard a video frame as a static image, and seldom exploit using the motion information of video frames in succession. Moreover, these methods usually calculate the visual distortion based on the independent spatial pixels. Recently, many researches show that the way people perceive the video signals is similar to the way filters process signals in the frequency domain. Therefore, in order to achieve better visual quality, we introduce a novel distortion measurement into the video coding system, which is consistent with human visual perception, and establish a perception-based rate-distortion optimization model. In this paper, we adopt Gabor filter family to decompose the video signals into frequency domain, and combine the video motion information to measure the perceptual distortion. We call it Motion tuned Distortion metric For Video coding (MDFV). After that we set up an MDFV based rate-distortion optimization model to select the best encoding mode. The experimental results show that the proposed approach is effective.
Xi Wang 0014, Li Su 0003, Qingming Huang, Chunxi Liu, Ling-Yu Duan
ICME1
2011 Visual perception based Lagrangian rate distortion optimization for video coding
abstract
In the conventional rate distortion optimization (RDO) video coding, the measure of distortion is mainly from the perspective of signal processing, while dose not fully take into account the characteristics of visual perception. People have concerns about not only the information of independent pixels, but also the temporal and spatial correlations between them. For different video content, human visual perception has different sensitivity. In this paper, in order to establish a RDO model which is more consistent with the human visual perception, we introduce the structural similarity and the content saliency information into the distortion metric. An adaptive Lagrange multiplier selection scheme is presented to allocate the bit resources more rationally by keeping the balance of the bit-rate and the visual quality. Experimental results show that the proposed method averagely reduces 10.14% bit-rate under the similar visual quality.
Xi Wang 0014, Li Su 0003, Qingming Huang, Chunxi Liu
ICIP1