EDBT 2026 Demo / reviewers in the wild / expert
Shi-Lin Wang
dblp:29/3890 · also Shilin Wang
· DBLP profile ↗
105ranked-venue papers
10as first author
51since 2021 · last 2026
0000-0002-8214-6809ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 68 · 7 first-author · 38 since 2021Artificial intelligence and machine learning · 29 · 4 first-author · 18 since 2021Security and privacy · 13 · 2 since 2021Databases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 2Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A robust multi-source transfer classification method based on belief functions for cross-domain pattern recognition
Linqing Huang, Jinfu Fan, Gongshen Liu, Shi-Lin Wang |
Int. J. Approx. Reason. | 4 |
| 2026 | Auditing Partial Dataset Usage in Large Language Models via Fuzzy Membership AggregationabstractThe remarkable capabilities of Large Language Models (LLMs) are fueled by massive internet-scale corpora. However, scraped data owners often do not consent to its use for training, raising significant legal and ethical concerns over copyright and privacy.Data auditingtechniques seek to verify whether a protected dataset was used in training a target LLM, typically framing the task as membership inference: estimating binary sample-level membership and aggregating to a dataset-level decision. In this paper, we identify a fundamental limitation of this crisp binary paradigm: in realistic training pipelines, datasets are rarely used in full. Instead, models are trained on mixtures of partial subsets drawn from multiple sources. Existing auditing techniques, built upon anall-or-noneassumption—declaring a dataset either entirely present or absent from training—collapse inpartial dataset usagescenarios. Their predictions fluctuate unpredictably with the member ratio, causing unstable performance and high false-negative rates. Inspired byfuzzy set theory, we relax the crisp notion of binary membership to a continuousfuzzy membershipin [0,1], quantifying each sample's degree of inclusion in the model's training set. We establish a theoretical bridge between sample-level fuzzy memberships and the dataset-level usage ratio, facilitating inference of the proportion of a protected dataset used during training. Aneural network fuzzifierfirst estimates sample-level fuzzy memberships from binary labels in a reference set, then refines them using dataset-level member ratios as higher-order supervision. Finally, adefuzzificationstage aggregates calibrated memberships to determine partial usage. Across LLMs of varying scales and multiple auditing datasets, ourFuzzy Auditorsubstantially outperforms state-of-the-art crisp binary techniques in detecting partial usage, estimating member proportions, and identifying individual member samples. Hongyu Zhu 0004, Sichu Liang, Bofan Chen, Shi-Lin Wang, Zhuosheng Zhang 0001, Weiping Ding 0001 |
IEEE Trans. Fuzzy Syst. | 4 |
| 2026 | Generalizable and Adaptive Continual Learning Framework for AI-Generated Image DetectionabstractThe malicious misuse and widespread dissemination of AI-generated images pose a significant threat to the authenticity of online information. Current detection methods often struggle to generalize to unseen generative models, and the rapid evolution of generative techniques continuously exacerbates this challenge. Without adaptability, detection models risk becoming ineffective in real-world applications. To address this critical issue, we propose a novel three-stage domain continual learning framework designed for continuous adaptation to evolving generative models. In the first stage, we employ a strategic parameter-efficient fine-tuning approach to develop a transferable offline detection model with strong generalization capabilities. Building upon this foundation, the second stage integrates unseen data streams into a continual learning process. To efficiently learn from limited samples of novel generated models and mitigate overfitting, we design a data augmentation chain with progressively increasing complexity. Furthermore, we leverage the Kronecker-Factored Approximate Curvature (K-FAC) method to approximate the Hessian and alleviate catastrophic forgetting. Finally, the third stage utilizes a linear interpolation strategy based on Linear Mode Connectivity, effectively capturing commonalities across diverse generative models and further enhancing overall performance. We establish a comprehensive benchmark of 27 generative models, including GANs, deepfakes, and diffusion models, chronologically structured up to August 2024 to simulate real-world scenarios. Extensive experiments demonstrate that our initial offline detectors surpass the leading baseline by +5.51% in terms of mean average precision. Our continual learning strategy achieves an average accuracy of 92.20%, outperforming state-of-the-art methods. Jun Lan 0001, Yaoyu Kang, Huijia Zhu, Weiqiang Wang 0002, Zhuosheng Zhang 0001, Shi-Lin Wang |
IEEE Trans. Multim. | 7 |
| 2025 | Stealing Knowledge from Auditable DatasetsabstractThe success of modern deep learning hinges on vast training data, much of which is scraped from the web and may include copyrighted or private content—raising serious legal and ethical concerns when used without authorization. Dataset provenance seeks to identify whether a model has been trained on specific data collections, thus protecting copyright holders while preserving data utility. Existing techniques either watermark datasets to embed distinctive behaviors, or directly infer usage from discrepancies in model outputs between seen and unseen samples. These approaches exploit the fundamental problem of empirical risk minimization to overfit to seen features. Hence, provenance signals are considered inherently hard to erase, while the adversary’s perspective remains largely overlooked, limiting our ability to assess reliability in real-world scenarios. In this work, we present a unified framework that interprets both watermarking and inference-based provenance as manifestations of output divergence, modeling the interaction between auditor and adversary as a min-max game over such divergences. This perspective motivates DivMin, a simple yet effective learning strategy that minimizes the relevant divergence to suppress provenance cues. Experiments across diverse image datasets demonstrate that, starting from a pretrained vision-language model, DivMin retains over 93% of the full fine-tuning performance gain relative to a zero-shot baseline, while evading all six state-of-the-art auditing methods. Our findings establish divergence minimization as a direct and practical path to obfuscating provenance, offering a realistic simulation of potential adversary strategies to guide the development of more robust auditing techniques. Code and Appendix will be available at https://github.com/GradOpt/DivMin. Hongyu Zhu 0004, Sichu Liang, Fangqi Li 0001, Shi-Lin Wang, Zhuosheng Zhang 0001 |
ECAI | 6 |
| 2025 | Rethinking the Fragility and Robustness of Fingerprints of Deep Neural NetworksabstractFingerprints characterize deep neural networks that are deployed as black-boxes. To achieve copyright tracing and integrity verification, fingerprints are categorized into robust fingerprints and fragile fingerprints. Despite of their distinct motivations, we show that both kinds of neural network fingerprints can be evaluated under a modification-scalable framework, which gives rise to a duality between their key metrics. These observations lead to a simultaneous scheme that reduces the cost of netural network intellectual property protection, with a controllable false negative rate. We implemented eleven representative families of modifications to evaluate fingerprints regarding both fragility and robustness, and verified the advantage of the simultaneous solution. Codes for reproducibility are available at https://github.com/solour-lfq/Fragile-and-Robust-Curves-of-DNN-Fingerprint. Fangqi Li 0001, Shi-Lin Wang, Lei Yang 0062 |
ICASSP | 2 |
| 2025 | Efficient and Effective Model ExtractionabstractModel extraction aims to steal a functionally similar copy from a machine learning as a service (MLaaS) API with minimal overhead, typically for illicit profit or as a precursor to further attacks, posing a significant threat to the MLaaS ecosystem. However, recent studies have shown that model extraction is highly inefficient, particularly when the target task distribution is unavailable. In such cases, even substantially increasing the attack budget fails to produce a sufficiently similar replica, reducing the adversary’s motivation to pursue extraction attacks. In this paper, we revisit the elementary design choices throughout the extraction lifecycle. We propose an embarrassingly simple yet dramatically effective algorithm, Efficient and Effective Model Extraction (E3), focusing on both query preparation and training routine. E3achieves superior generalization compared to state-of-the-art methods while minimizing computational costs. For instance, with only 0.005× the query budget and less than 0.2× the runtime, E3outperforms classical generative model based data-free model extraction by an absolute accuracy improvement of over 50% on CIFAR-10. Our findings underscore the persistent threat posed by model extraction and suggest that it could serve as a valuable benchmarking algorithm for future security evaluations. Hongyu Zhu 0004, Sichu Liang, Fangqi Li 0001, Shi-Lin Wang |
ICASSP | 6 |
| 2025 | Enhancing Visual Forced Alignment with Local Context-Aware Feature Extraction and Multi-Task LearningabstractThis paper introduces a novel approach to Visual Forced Alignment (VFA), aiming to accurately synchronize utterances with corresponding lip movements, without relying on audio cues. We propose a novel VFA approach that integrates a local context-aware feature extractor and employs multitask learning to refine both global and local context features, enhancing sensitivity to subtle lip movements for precise word-level and phoneme-level alignment. Incorporating the improved Viterbi algorithm for post-processing, our method significantly reduces misalignments. Experimental results show our approach outperforms existing methods, achieving a 6% accuracy improvement at the word-level and 27% improvement at the phoneme-level in LRS2 dataset. These improvements offer new potential for applications in automatically subtitling TV shows or user-generated content platforms like TikTok and YouTube Shorts. Lei Yang 0062, Shi-Lin Wang |
ICASSP | 3 |
| 2025 | Towards A Distribution Alignment Framework for Incomplete Data ClassificationabstractMissing attribute values frequently affect data classification, reducing accuracy as most models rely on complete datasets. Imputing missing values is typically used to restore data completeness, which is essential for building models. The effectiveness of imputation significantly impacts the classification accuracy. Therefore, improving imputed values’ quality is crucial for better classification outcomes. Here, we introduce a new distribution alignment framework (DAF) to address classification issues with complete training data but incomplete test data. Initially, DAF imputes missing test data values using mean vectors from complete training data, minimizing the first-order distributional discrepancies. Next, it aligns the second-order statistical distributions, specifically covariance matrices, of both training and imputed test data to derive a feature transformation matrix. This matrix generates new feature representations for the incomplete test data. The classifier trained on the complete training data then classifies the imputed test data under this new feature representation. The experiments on several benchmark datasets show that DAF usually outperforms many advanced methods, achieving the higher classification performance. Linqing Huang, Jinfu Fan, Shi-Lin Wang, Gongshen Liu, Shouxuan Liu |
ICASSP | 3 |
| 2025 | MIFAE-Forensics: Masked Image-Frequency AutoEncoder for DeepFake DetectionabstractWith continuously evolving generative models and increasingly diverse face forgery products, there is a growing demand for DeepFake detectors with stronger generalization ability and robustness. Previous works mainly capture method-specific forgery artifacts in the training set, thus failing to generalize well to unseen manipulations. In this paper, our key insight is that exploring common characteristics of natural faces is more ideal to alleviate overfitting rather than relying on specific forgery clues, as all sorts of manipulated images have intrinsic distributional differences from those captured by cameras. Hence, we propose a two-stage method, termed MIFAE-Forensics. Specifically, it reconstructs both facial semantics and local details from masked facial regions and high-frequency components, respectively, aiming to capture natural facial consistency in spatial domain and high-frequency details in frequency domain simultaneously. This facilitates the learning of a robust and transferable facial representation specialized for DeepFake detection. Subsequently, the pre-trained model is further fine-tuned to perform binary forgery classification along with reconstructing real faces in spatial domain, which ensures that the detector can maintain the ability to model real faces and encourages it to make decisions based on reconstruction discrepancies. Extensive experiments show superior results over state-of-the-arts on a wide range of DeepFake detection benchmarks. Our code is available at https://github.com/Mark-Dou/Forensics. Shi-Lin Wang |
ICASSP | 3 |
| 2025 | Personalized Speech Enhancement without User Enrollment for Real-World Audio Replay ScenariosabstractMany speech enhancement (SE) approaches have been proposed to deal with cocktail party problem. Personalized speech enhancement (PSE) approaches improve SE performance by utilizing user enrollment speech. However, PSE requires users to record additional clean audio for registration, which can be redundant works or impractical for many real-world scenarios. For instance, in personal devices and Vloggers’ audio playback scenarios, there are already many video/audios available recorded under different noise types and SNR conditions, but without any target speaker pre-registered speech. To better utilize information from existing video/audio stock, this paper propose a novel speech enhancement approach that integrates PSE methods without requiring pre-registered user speech. With user adaptation and noise adaptation training modules, the proposed approach automatically selects high-quality speech segments to assist in denoising low-quality speech segments. Additionally, two test sets were collected to evaluate the performance in the aforementioned scenarios. Experimental results demonstrate that the proposed approach outperforms the corresponding SE methods in both objective and subjective evaluation metrics. Shi-Lin Wang, Yanhua Long |
ICASSP | 2 |
| 2025 | Membership Encoding for Black-Box Neural Network WatermarkingabstractDeep neural network watermarking is an emerging technique for protecting the copyright of models. Most existing black-box watermarking methods leverage the backdoor, making them inherently vulnerable to backdoor removal attacks. In this paper, we propose a novel watermark removal attack, Misleading Fine-tuning, which effectively eliminates backdoor-based watermarks with limited data. To counter this threat, we present a novel black-box watermarking method based on membership encoding. This method overfits the protected model on a subset of training data that serve as triggers, thereby making it resistant to backdoor removal attacks. Extensive experiments demonstrate its fidelity and robustness against adversarial modifications, whether applied to the model or the inputs. Hangwei Zhang, Fangqi Li 0001, Shi-Lin Wang |
ICASSP | 3 |
| 2025 | ROAR: Reducing Inversion Error in Generative Image Watermarking
Shi-Lin Wang, Ee-Chien Chang |
ICCV | 3 |
| 2025 | Evading Data Provenance in Deep Neural NetworksabstractModern over-parameterized deep models are highly data-dependent, with large scale general-purpose and domain-specific datasets serving as the bedrock for rapid advancements. However, many datasets are proprietary or contain sensitive information, making unrestricted model training problematic. In the open world where data thefts cannot be fully prevented, Dataset Ownership Verification (DOV) has emerged as a promising method to protect copyright by detecting unauthorized model training and tracing illicit activities. Due to its diversity and superior stealth, evading DOV is considered extremely challenging. However, this paper identifies that previous studies have relied on oversimplistic evasion attacks for evaluation, leading to a false sense of security. We introduce a unified evasion framework, in which a teacher model first learns from the copyright dataset and then transfers task-relevant yet identifier-independent domain knowledge to a surrogate student using an out-of-distribution (OOD) dataset as the intermediary. Leveraging Vision-Language Models and Large Language Models, we curate the most informative and reliable subsets from the OOD gallery set as the final transfer set, and propose selectively transferring task-oriented knowledge to achieve a better trade-off between generalization and evasion effectiveness. Experiments across diverse datasets covering eleven DOV methods demonstrate our approach simultaneously eliminates all copyright identifiers and significantly outperforms nine state-of-the-art evasion attacks in both generalization and effectiveness, with moderate computational overhead. As a proof of concept, we reveal key vulnerabilities in current DOV methods, highlighting the need for long-term development to enhance practicality. Hongyu Zhu 0004, Sichu Liang, Zhuomeng Zhang, Fangqi Li 0001, Shi-Lin Wang |
ICCV | 6 |
| 2025 | Towards a Practical Screen-Filming Resistant Image Watermarking System
Shicong Han, Fangqi Li 0001, Shi-Lin Wang |
ICIG (3) | 4 |
| 2025 | New Multi-Source Distributed Transfer Learning FrameworkabstractIn pattern recognition, where the labeled data is scarce, transfer learning (also called domain adaptation in some cases) methods frequently come into play to transfer knowledge from the source domains to bolster the construction of classification models within the target domain. The judicious fusion of information from multiple source domains typically enhances classification precision. In light of this, we introduce a new Multi-source Distributed Transfer Learning (MDTL) framework designed to adeptly integrate complementary information across various source domains through the application of belief functions. In this approach, the distributions of each source and target domain are aligned independently. Subsequently, the resultant soft classification outcomes, facilitated by different source domains, are amalgamated using belief functions. This integration incorporates novel weighting factors that consider both distribution discrepancies and classifier effectiveness. The effectiveness of MDTL was assessed against a range of related methods, and the experimental findings confirm that it markedly improves classification accuracy in the target domain. Linqing Huang, Yumei Hu, Shi-Lin Wang, Gongshen Liu, Jinfu Fan |
ICIP | 4 |
| 2025 | Revisiting Data Auditing in Large Vision-Language ModelsabstractWith the surge of large language models (LLMs), Large Vision-Language Models (VLMs)-which integrate vision encoders with LLMs for accurate visual grounding-have shown great potential in tasks like generalist agents and robotic control. However, VLMs are typically trained on massive web-scraped images, raising concerns over copyright infringement and privacy violations, and making data auditing increasingly urgent. Membership inference (MI), which determines whether a sample was used in training, has emerged as a key auditing technique, with promising results on open-source VLMs like LLaVA (AUC > 80%). In this work, we revisit these advances and uncover a critical issue: current MI benchmarks suffer from distribution shifts between member and non-member images, introducing shortcut cues that inflate MI performance. We further analyze the nature of these shifts and propose a principled metric based on optimal transport to quantify the distribution discrepancy. To evaluate MI in realistic settings, we construct new benchmarks with i.i.d. member and non-member images. Existing MI methods fail under these unbiased conditions, performing only marginally better than chance. Further, we explore the theoretical upper bound of MI by probing the Bayes Optimality within the VLM's embedding space and find the irreducible error rate remains high. Despite this pessimistic outlook, we analyze why MI for VLMs is particularly challenging and identify three practical scenarios-fine-tuning, access to ground-truth texts, and set-based inference-where auditing becomes feasible. Our study presents a systematic view of the limits and opportunities of MI for VLMs, providing guidance for future efforts in trustworthy data auditing. Code and data will be available at https://github.com/GradOpt/Revisiting-VLM-MIA\faGithub. Hongyu Zhu 0004, Sichu Liang, Boheng Li, Tongxin Yuan, Fangqi Li 0001, Shi-Lin Wang, Zhuosheng Zhang 0001 |
ACM Multimedia | 8 |
| 2025 | Boosting the Uniqueness of Neural Networks Fingerprints with Informative TriggersabstractOne prerequisite for secure and reliable artificial intelligence services is tracing the copyright of backend deep neural networks.
In the black-box scenario, the copyright of deep neural networks can be traced by their fingerprints, i.e., their outputs on a series of fingerprinting triggers.
The performance of deep neural network fingerprints is usually evaluated in robustness, leaving the accuracy of copyright tracing among a large number of models with a limited number of triggers intractable.
This fact challenges the application of deep neural network fingerprints as the cost of queries is becoming a bottleneck. This paper studies the performance of deep neural network fingerprints from an information theoretical perspective.
With this new perspective, we demonstrate that copyright tracing can be more accurate and efficient by using triggers with the largest marginal mutual information. Extensive experiments demonstrate that our method can be seamlessly incorporated into any existing fingerprinting scheme to facilitate the copyright tracing of deep neural networks. Zhuomeng Zhang, Fangqi Li 0001, Shi-Lin Wang |
NeurIPS | 4 |
| 2025 | Towards Highly Generalized Lip-Sync Deepfake Detection via Detailed Audio-Visual Inconsistency Analysis
Shi-Lin Wang |
PRCV (11) | 2 |
| 2025 | The improved mountain gazelle optimizer for spatiotemporal support vector regression: a novel method for railway subgrade settlement prediction integrating multi-source information
Guangwu Chen, Shilin Zhao, Peng Li 0070, Shi-Lin Wang, Vyacheslav V. Potekhin |
Appl. Intell. | 4 |
| 2025 | A New Multi-Target Domain Adaptation Method Based on Evidence Theory for Distribution Inconsistent Data ClassificationabstractIn the context of distribution inconsistent data classification, addressing distribution shift is crucial, typically accomplished through domain adaptation (DA) techniques. Once distributions are aligned between the source and target domains, the problem transforms into a conventional recognition task. This article introduces a new method called multitarget DA based on Evidence Theory (MET). For a given target domain, a random merger with other target domains is performed, generating distinct new target domains. Domain-invariant features corresponding to each new target domain are learned by minimizing distribution discrepancies separately between the source and different new target domains. The merging of target domains alters the distribution of the new target domain, leading to variations in the retained information within the learned domain-invariant features. For a query pattern in this target domain, multiple soft classification results (CCR) are obtained after aligning the distributions of the source and different new target domains. These soft CCR complement each other, and evidence theory is employed as a tool to represent and combine uncertain information, fusing these results. The weights for this fusion are automatically learned by minimizing the mean squared error between the combined results and the ground truth on labeled source domain data. The final class decision is determined through the weighted evidential combination of multiple pieces of soft CCR. MET is assessed on several datasets (i.e., Office+Caltech-10, VLSC, and V-RSIR) and compared to various advanced DA methods (e.g., GNN, MT, PAL, PTD, and so on) to validate its effectiveness. The experimental results demonstrate that MET usually can obtain a higher classification performance (i.e., the accuracy can be improved by 2% compared to many methods in most cases). Linqing Huang, Jinfu Fan, Shi-Lin Wang |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2024 | Revisiting the Information Capacity of Neural Network Watermarks: Upper Bound Estimation and BeyondabstractTo trace the copyright of deep neural networks, an owner can embed its identity information into its model as a watermark. The capacity of the watermark quantify the maximal volume of information that can be verified from the watermarked model. Current studies on capacity focus on the ownership verification accuracy under ordinary removal attacks and fail to capture the relationship between robustness and fidelity. This paper studies the capacity of deep neural network watermarks from an information theoretical perspective. We propose a new definition of deep neural network watermark capacity analogous to channel capacity, analyze its properties, and design an algorithm that yields a tight estimation of its upper bound under adversarial overwriting. We also propose a universal non-invasive method to secure the transmission of the identity message beyond capacity by multiple rounds of ownership verification. Our observations provide evidence for neural network owners and defenders that are curious about the tradeoff between the integrity of their ownership and the performance degradation of their products. Fangqi Li 0001, Haodong Zhao, Shi-Lin Wang |
AAAI | 4 |
| 2024 | Improve Deep Forest with Learnable Layerwise Augmentation Policy SchedulesabstractAs a modern ensemble technique, Deep Forest (DF) employs a cascading structure to construct deep models, providing stronger representational power compared to traditional decision forests. However, its greedy multi-layer learning procedure is prone to overfitting, limiting model effectiveness and generalizability. This paper presents AugDF, an optimized Deep Forest featuring learnable, layerwise data augmentation policy schedules. Specifically, We introduce the Cut Mix for Tabular data (CMT) augmentation technique to mitigate overfitting and develop a population-based search algorithm to tailor augmentation intensity for each layer. Additionally, we propose to incorporate outputs from intermediate layers into a checkpoint ensemble for more stable performance. Experimental results show that AugDF sets new state-of-the-art (SOTA) benchmarks in various tabular classification tasks, outperforming shallow tree ensembles, deep forests, deep neural network, and AutoML competitors. The learned policies also transfer effectively to Deep Forest variants, underscoring its potential for enhancing non-differentiable deep learning modules in tabular signal processing. Hongyu Zhu 0004, Sichu Liang, Fangqi Li 0001, Yali Yuan, Shi-Lin Wang, Guang Cheng 0001 |
ICASSP | 6 |
| 2024 | Speaker-Adaptive Lipreading Via Spatio-Temporal Information LearningabstractLipreading has been rapidly developed recently with the help of large-scale datasets and large models. Despite the significant progress made, the performance of lipreading models still falls short when dealing with unseen speakers. Therefore, it is necessary to utilize the speaker’s videos for fine-tuning to obtain a speaker-adaptive model. However, this approach can result in high overheads, especially for full fine-tuning. To address this problem, we propose a novel parameter-efficient fine-tuning method based on spatio-temporal information learning. In our approach, a low-rank adaptation module which can influence global spatial features and a plug-and-play temporal adaptive weight learning module are designed in the front-end and back-end network, which can adapt to the speaker’s unique features such as the shape of the lips and the style of speech, respectively. An Adapter module is added between them to further enhance the spatio-temporal learning. The final experiments on the LRW-ID and GRID datasets demonstrate that our method achieves state-of-the-art performance even with fewer parameters. Lei Yang 0062, Shi-Lin Wang |
ICASSP | 5 |
| 2024 | Data-Free Watermark for Deep Neural Networks by Truncated Adversarial DistillationabstractModel watermarking secures ownership verification and copyright protection of deep neural networks. In the black-box scenario, watermarking schemes commonly rely on injecting triggers and requiring the model's training data to maintain its performance. However, such knowledge might be unavailable in commercial settings as model transactions or copyright transfers. To tackle this challenge, we propose a novel data-free black-box watermarking scheme. Our approach modifies data-free adversarial distillation to efficiently obtain a generator that produces samples serving as a substitute for the training data so the watermark can achieve high fidelity without referring to the training data. Chao-Bo Yan, Fangqi Li 0001, Shi-Lin Wang |
ICASSP | 3 |
| 2024 | Personatalk: Preserving Personalized Dynamic Speech Style In Talking Face GenerationabstractRecent visual speaker authentication methods claimed their effectiveness against deepfake attacks. However, the success is attributed to the inadequacy of existing talking face generation methods to preserve the dynamic speech style of the speaker, which serves as the key cue for authentication methods in verification. To address this, we propose PersonaTalk, a speaker-specific method utilizing the speaker’s video data to enhance the fidelity of the speaker’s dynamic speech styles in generated videos. Our approach introduces a visual context block to integrate lip motion information into the audio features. Additionally, to enhance reading intelligibility in dubbed videos, a cross dubbing phase is incorporated during training. Experiments on the GRID dataset show the superiority of PersonaTalk over existing SOTA methods. These findings emphasize the need for enhanced defense measures in existing lip-based speaker authentication methods. Qianxi Lu, Shi-Lin Wang |
ICIP | 3 |
| 2024 | QMixCAT: Unsupervised Speech Enhancement Using Quality-guided Signal Mixing and Competitive Alternating Model Training
Shi-Lin Wang, Haixin Guan, Yanhua Long |
INTERSPEECH | 1 |
| 2024 | Reliable Model Watermarking: Defending against Theft without Compromising on EvasionabstractWith the rise of Machine Learning as a Service (MLaaS) platforms, safeguarding the intellectual property of deep learning models is becoming paramount. Among various protective measures, trigger set watermarking has emerged as a flexible and effective strategy for preventing unauthorized model distribution. However, this paper identifies an inherent flaw in the current paradigm of trigger set watermarking: evasion adversaries can readily exploit the shortcuts created by models memorizing watermark samples that deviate from the main task distribution, significantly impairing their generalization in adversarial settings. To counteract this, we leverage diffusion models to synthesize unrestricted adversarial examples as trigger sets. By learning the model to accurately recognize them, unique watermark behaviors are promoted through knowledge injection rather than error memorization, thus avoiding exploitable shortcuts. Furthermore, we uncover that the resistance of current trigger set watermarking against removal attacks primarily relies on significantly damaging the decision boundaries during embedding, intertwining unremovability with adverse impacts. By optimizing the knowledge transfer properties of protected models, our approach conveys watermark behaviors to extraction surrogates without aggressive decision boundary perturbation. Experimental results on CIFAR-10/100 and Imagenette datasets demonstrate the effectiveness of our method, showing not only improved robustness against evasion adversaries but also superior resistance to watermark removal attacks compared to state-of-the-art solutions. Hongyu Zhu 0004, Sichu Liang, Fangqi Li 0001, Ju Jia, Shi-Lin Wang |
ACM Multimedia | 6 |
| 2024 | CDAF3D: Cross-Dimensional Attention Fusion for Indoor 3D Object Detection
Shi-Lin Wang, Hai Huang 0001, Yueyan Zhu, Zhenqi Tang |
PRCV (13) | 1 |
| 2024 | Lip Feature Disentanglement for Visual Speaker Authentication in Natural ScenesabstractRecent studies have shown that lip shape and movement can be used as an effective biometric feature for speaker authentication. By using random prompt text scheme, lip-based authentication system can also achieve good liveness detection performance in laboratory scenarios. However, due to the increasingly widespread mobile application, the authentication system may face additional practical difficulties such as complex background, limited user samples, etc., which will degrade the authentication performance derived by current methods. To confront the above problems, a new deep neural network, i.e. the Triple-feature Disentanglement Network for Visual Speaker Authentication (TDVSA-Net), is proposed in this paper to extract discriminative and disentangled lip features for visual speaker authentication in the random prompt text scenario. Three decoupled lip features, including the content feature inferring the speech content, the physiological lip feature describing the static lip shape and appearance and the behavioral lip feature depicting the unique pattern in lip movements during utterance, are extracted by TDVSA-Net and fed into corresponding modules to authenticate both the prompt text and the speaker’s identity. Experiment results have demonstrated that compared with several SOTA visual speaker authentication methods, the proposed TDVSA-Net can extract more discriminative and robust lip features which boost the content recognition and identity authentication performance against both human imposters and DeepFake attacks. Lei Yang 0062, Shi-Lin Wang, Alan Wee-Chung Liew |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Fine-Grained Lip Image Segmentation Using Fuzzy Logic and Graph ReasoningabstractFine-grained lip image segmentation plays a critical role in downstream tasks, such as automatic lipreading, as it enables the accurate identification of inner mouth components, such as teeth and tongue, which are essential for comprehending spoken utterances. However, achieving accurate and robust lip image segmentation in natural scenes is still challenging due to significant variations in lighting condition, head pose, and background. This article proposes a novel deep neural network-based method for fine-grained lip image segmentation that exploits fuzzy and graph theories to handle these variations. A fuzzy learning module is designed to deal with the uncertainties in color and edge information and enhance feature maps at various scales. The fuzzy graph reasoning module with fuzzy projection models the relationship among semantics components and achieves a global receptive field. In our experiments, a fine-grained lip region segmentation dataset, i.e., FLRSeg, is built for evaluation, and experimental results have shown that the proposed method can achieve superior segmentation performance (94.36% in pixel accuracy and 74.89% in mIoU) compared with several state-of-the-art (SOTA) lip image segmentation methods. Lei Yang 0062, Shi-Lin Wang, Alan Wee-Chung Liew |
IEEE Trans. Fuzzy Syst. | 2 |
| 2023 | Content-Insensitive Dynamic Lip Feature Extraction for Visual Speaker Authentication Against Deepfake AttacksabstractRecent research has shown that lip-based speaker authentication system can achieve good authentication performance. However, with emerging deepfake technology, attackers can make high fidelity talking videos of a user, thus posing a great threat to these systems. Confronted with this threat, we propose a new deep neural network for lip-based visual speaker authentication against human imposters and deepfake attacks. One dynamic enhanced block with context modeling scheme is designed to capture a user’s unique talking habit by learning from his/her lip movement. Meanwhile, a cross-modality content-guided loss is designed to help extract discriminative features when learning from different lip movement of a user uttering different content. This loss makes the proposed method insensitive to content variation. Experiments on the GRID dataset show that the proposed method not only outperforms three state-of-the-art methods but also simplifies the training process and reduces the training cost. Shi-Lin Wang |
ICASSP | 2 |
| 2023 | Measure and Countermeasure of the Capsulation Attack Against Backdoor-Based Deep Neural Network WatermarksabstractBackdoor-based watermarking schemes were proposed to protect the intellectual property of deep neural networks under the black-box setting. However, additional security risks emerge after the schemes have been published for as forensics tools. This paper reveals the capsulation attack that can easily invalidate most established backdoor-based watermarking schemes without sacrificing the pirated model’s functionality. By encapsulating the deep neural network with a filter, an adversary can block abnormal queries and reject the ownership verification. We propose a metric to measure a backdoor-based watermarking scheme’s security against the capsulation attack, and design a new backdoor-based deep neural network watermarking scheme that is secure against the capsulation attack by inverting the encoding process. Fangqi Li 0001, Shi-Lin Wang |
ICASSP | 2 |
| 2023 | An Auto-Encoder Based Method for Camera Fingerprint CompressionabstractCamera fingerprint links a picture to its camera sensor, which is widely applied in sensor device identification, social network tracing and forgery detection. However, such fingerprints are in high dimensionality and cost substantial memory and computing resources, limiting their uses in real-time processing on embedded devices. In this paper, we introduce a new method to compress high-dimensional floating-point fingerprints to low-dimensional binary features to save storage as well as maintaining their representative abilities. Also, we present a much faster approach to sensor device matching with hamming distance, compared with the commonly used Peak to Correlation Energy (PCE) distance. Our method contains two stages. First, raw fingerprints are compressed into low-dimensional features with our proposed grouping strategy and auto-encoder based model. Then, the compressed floating-point features are further converted into more compact binary features. Experiments show that our method achieves superior performance over several competitive compression methods in both identification and verification tasks. Jiashang Hu, Shi-Lin Wang |
ICASSP | 4 |
| 2023 | Exploiting Complementary Dynamic Incoherence for DeepFake Video DetectionabstractRecently, manipulated videos based on DeepFake technology have spread widely on social media, causing concerns about the authenticity of video content and personal privacy protection. Although existing DeepFake detection methods achieve remarkable progress in some specific scenarios, their detection performance usually drops drastically when detecting unseen manipulation methods. Compared with static information such as human face, dynamic information depicting the movements of facial features is more difficult to forge without leaving visual or statistical traces. Hence, in order to achieve better generalization ability, we focus on dynamic information analysis to disclose such traces and propose a novel Complementary Dynamic Interaction Network (CDIN). Inspired by the DeepFake detection methods based on mouth region analysis, both the global (entire face) and local (mouth region) dynamics are analyzed with properly designed network branches, respectively, and their feature maps at various levels are communicated with each other using a newly proposed Complementary Cross Dynamics Fusion Module (CCDFM). With CCDFM, the global branch will pay more attention to anomalous mouth movements and the local branch will gain more information about the global context. Finally, a multi-task learning scheme is designed to optimize the network with both the global and local information. Extensive experiments have demonstrated that our approach achieves better detection results compared with several SOTA methods, especially in detecting video forgeries manipulated by unseen methods. Shi-Lin Wang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Linear Functionality Equivalence Attack Against Deep Neural Network Watermarks and a Defense Method by Neuron MappingabstractAs an ownership verification technique for deep neural networks, the white-box neural network watermark is being challenged by the functionality equivalence attack. By leveraging the structural symmetry within a deep neural network and manipulating the parameters accordingly, an adversary can invalidate almost all white-box watermarks without affecting the network’s performance. This paper introduces the linear functionality equivalence attack, which can adapt to different network architectures without requiring knowledge of either the watermark or data. We also propose NeuronMap, a framework that can efficiently neutralize linear functionality equivalence attacks and can be easily combined with existing white-box watermarks to enhance their robustness. Experiments conducted on several deep neural networks and state-of-the-art white-box watermarking schemes have demonstrated not only the destructive power of linear functionality equivalence attacks but also the defense capability of NeuronMap. Our result shows that the threat of basic linear functionality equivalence attacks against deep neural network watermarks can be effectively solved using NeuronMap. Fangqi Li 0001, Shi-Lin Wang, Alan Wee-Chung Liew |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2022 | Cross-Domain Local Characteristic Enhanced Deepfake Video Detection
Shi-Lin Wang |
ACCV (5) | 3 |
| 2022 | Few-shot Table-to-text Generation with Prefix-Controlled GeneratorabstractNeural table-to-text generation approaches are data-hungry, limiting their adaption for low-resource real-world applications. Previous works mostly resort to Pre-trained Language Models (PLMs) to generate fluent summaries of a table. However, they often contain hallucinated contents due to the uncontrolled nature of PLMs. Moreover, the topological differences between tables and sequences are rarely studied. Last but not least, fine-tuning on PLMs with a handful of instances may lead to over-fitting and catastrophic forgetting. To alleviate these problems, we propose a prompt-based approach, Prefix-Controlled Generator (i.e., PCG), for few-shot table-to-text generation. We prepend a task-specific prefix for a PLM to make the table structure better fit the pre-trained input. In addition, we generate an input-specific prefix to control the factual contents and word order of the generated text. Both automatic and human evaluations on different domains (humans, books and songs) of the Wikibio dataset prove the effectiveness of our approach. Yutao Luo, Menghua Lu, Gongshen Liu, Shi-Lin Wang |
COLING | 4 |
| 2022 | Fostering The Robustness Of White-Box Deep Neural Network Watermarks By Neuron AlignmentabstractThe wide application of deep learning techniques is boosting the regulation of deep learning models, especially deep neural networks (DNN), as commercial products. A necessary prerequisite for such regulations is identifying the owner of deep neural networks, which is usually done through the watermark. Current DNN watermarking schemes, particularly white-box ones, are uniformly fragile against a family of functionality equivalence attacks, especially the neuron permutation. This operation can effortlessly invalidate the ownership proof and escape copyright regulations. To enhance the robustness of white-box DNN watermarking schemes, this paper presents a procedure that aligns neurons into the same order as when the watermark is embedded, so the watermark can be correctly recognized. This neuron alignment process significantly facilitates the functionality of established deep neural network watermarking schemes. Fangqi Li 0001, Shi-Lin Wang |
ICASSP | 2 |
| 2022 | Multi-Task Learning Improves Synthetic Speech DetectionabstractWith the development of deep learning, synthetic speech has become more and more realistic and easier to spoof Automatic Speaker Verification (ASV) devices. Based on mining more effective hand-crafted features and proposing more powerful networks, many algorithms have been proposed to detect this malicious attack. In this paper, by observing that deepening the network impairs the performance of the network in detecting unknown attacks, we propose that the synthetic speech detection problem is an out-of-distribution (OOD) generalization problem and we enhance the robustness of networks by using multi-task learning. In our system, three auxiliary tasks are used to assist synthetic speech detection: bonafide speech reconstruction, spoofing voice conversion and speaker classification. Experimental results show that our approach can be applied to multiple architectures and can significantly improve the performance on both known attacks (development set) and unknown attacks (evaluation set). In addition, our best-performing network is quite competitive to recent state-of-the-art (SOTA) systems. It demonstrates the potential application of multi-task learning in synthetic speech detection. Yichuan Mo, Shi-Lin Wang |
ICASSP | 2 |
| 2022 | Stgat-Mad : Spatial-Temporal Graph Attention Network For Multivariate Time Series Anomaly DetectionabstractAnomaly detection in multivariate time series data is challenging due to complex temporal and feature correlations. This paper proposes a novel unsupervised multi-scale stacked spatial-temporal graph attention network for multivariate time series anomaly detection (STGAT-MAD). The core of our framework is to coherently capture the feature and temporal correlations among multivariate time-series data by stackable STGAT networks. Meanwhile, a multi-scale input network is exploited to capture the temporal correlations in different time-scales. Besides, a new dataset derived from a real-world wind farm is built and released for multivariate time series anomaly detection. Experiments on the proprietary dataset and three public datasets show that our method significantly outperforms existing baseline approaches, and provides interpretability for anomaly location. Siqi Wang 0001, Xiandong Ma, Chengkun Wu, Canqun Yang, Detian Zeng, Shi-Lin Wang |
ICASSP | 7 |
| 2022 | Chinese Mandarin Lipreading using Cascaded Transformers with Multiple Intermediate RepresentationsabstractAutomatic lipreading has attracted much research interest over the past few decades. Different from English, Chinese is a tone-based language with a large alphabet and thus the correlation between Chinese characters and lip motions is more complex. Most existing methods employed an intermediate representation (usually Pinyin), and adopted a cascaded architecture for Chinese lipreading. However, such a cascaded structure may accumulate errors, and employing Pinyin as the intermediate representation would cause the loss of visual information. Moreover, these approaches do not perform well for unseen speakers due to inter-speaker variability. In this paper, we propose a cascaded Transformer-based model with a new cross-level attention mechanism, enriching the ways of information transmission between cascading structures and reducing the accumulation of errors. Multiple intermediate representations including Chinese Pinyin and the visemes are adopted to acquire multi-perspective visual and linguistic features and to improve the generalization ability for unseen speakers. Evaluations on the public sentence-level Chinese lipreading database, i.e. CMLR, have demonstrated the advantages of the proposed method in both speaker-independent and multi-speaker scenarios over state-of-the-art approaches. Xinghua Ma, Shi-Lin Wang |
ICIP | 2 |
| 2022 | PPT: Backdoor Attacks on Pre-trained Models via Poisoned Prompt TuningabstractRecently, prompt tuning has shown remarkable performance as a new learning paradigm, which freezes pre-trained language models (PLMs) and only tunes some soft prompts. A fixed PLM only needs to be loaded with different prompts to adapt different downstream tasks. However, the prompts associated with PLMs may be added with some malicious behaviors, such as backdoors. The victim model will be implanted with a backdoor by using the poisoned prompt. In this paper, we propose to obtain the poisoned prompt for PLMs and corresponding downstream tasks by prompt tuning. We name this Poisoned Prompt Tuning method "PPT". The poisoned prompt can lead a shortcut between the specific trigger word and the target label word to be created for the PLM. So the attacker can simply manipulate the prediction of the entire model by just a small prompt. Our experiments on various text classification tasks show that PPT can achieve a 99% attack success rate with almost no accuracy sacrificed on original task. We hope this work can raise the awareness of the possible security threats hidden in the prompt. Yichun Zhao, Boqun Li, Gongshen Liu, Shi-Lin Wang |
IJCAI | 5 |
| 2022 | KLAttack: Towards Adversarial Attack and Defense on Neural Dependency Parsing ModelsabstractAlthough neural language models achieve great performance on many Natural Language Processing tasks, they suffer from various adversarial attacks. Previous works mainly focus on semantic adversarial examples, which have similar semantics to the original sentences, while syntactic adversarial attacks against the dependency parsing task are still in an early stage of research. In this paper, we propose a novel method KLAttack, crafting word-level adversarial examples to attack neural-network-based dependency parsing models. Specifically, we retrieve the class probabilities from the victim dependency parsing model and compute the KL divergence by masking every word in a sentence. Then we use pre-trained language models and reference parsers to generate candidates for substitution. Experiments on the English Penn Treebank (PTB) dataset show that our method improves the attack success rate against Deep Biaffine Parser by up to 13.04% compared with previous related studies. Based on KLAttack, we further propose Syntax-Aware Transformer for Input Reconstruction, a denoiser to recover the original sentences from the adversarial examples. Trained adversarially with successfully attacked sentences from KLAttack, we enhance the robustness of the dependency parsing models by concatenating the denoiser ahead of the victim models. Yutao Luo, Menghua Lu, Chaoqi Yang, Gongshen Liu, Shi-Lin Wang |
IJCNN | 5 |
| 2022 | Online Intrusion Detection for Internet of Things Systems With Full Bayesian Possibilistic Clustering and Ensembled Fuzzy ClassifiersabstractThe pervasive deployment of the Internet of Things (IoT) has significantly facilitated manufacturing and living. The diversity and continual updates of IoT systems make their security a crucial challenge, among which the detection of malicious network traffic turns out to be the most common yet destructive threat. Despite the efforts on feature engineering and classification backend designing, established intrusion detection systems sometimes lack robustness and are inflexible against the shift of the traffic distribution. To deal with these disadvantages, we design a fuzzy system for the online defense of IoT. Our framework incorporates a full Bayesian possibilistic clustering module for feature processing and an ensemble module motivated by reinforcement learning and adaptive boosting that dynamically fits the streaming data. The proposed clustering module overcomes the issue of determining the number of clusters and can dynamically identify new patterns. The classifier backend combines a collection of fuzzy decision trees that provide readable decision boundaries. The ensembled classifiers can accommodate the drift of data distribution to optimize the long-time performance. Our proposal is tested on settings including one dataset collected from real IoT systems and is compared to numerous competitors. Experimental results verified the advantage of our system regarding accuracy and stability. Fangqi Li 0001, Ruijie Zhao 0001, Shi-Lin Wang, Libo Chen 0001, Alan Wee-Chung Liew, Weiping Ding 0001 |
IEEE Trans. Fuzzy Syst. | 3 |
| 2021 | Spatial Aggregation for Scene Text Recognition
Yi-Li Huang, Chengyu Gu, Shi-Lin Wang, Kai Chen 0006 |
BMVC | 3 |
| 2021 | Improving the Generalization Ability of Deepfake Detection via Disentangled Representation LearningabstractDeepfake refers to a deep learning based technology which can synthesize visually realistic face images/videos. The misuse of this technology poses a great threat to the society. Although numerous approaches have been proposed to detect Deepfake forgeries, their generalization ability on unseen datasets is limited. In this paper, we propose a new approach that detects human face forgeries by automatically locating the forgery-related region to make the final decision. The proposed network contains two modules, including: the disentanglement module to extract forgery relevant information and the classification module to detect the manipulation artifacts from various regions at different scales. The experiment results on three widely used Deepfake datasets show that the proposed approach can achieve high detection accuracies and outperforms several state-of-the-arts methods especially when evaluated on the unseen datasets. Jiashang Hu, Shi-Lin Wang |
ICIP | 2 |
| 2021 | Persistent Watermark For Image Classification Neural Networks By Penetrating The AutoencoderabstractDeep neural networks for image processing, especially image classification, have become ubiquitous. To protect them as intellectual properties and standardize the commercialization of their service, watermarking schemes have been proposed to authenticate the author of models. Many black-box watermarking schemes insert a backdoor into the neural network by poisoning the training dataset. Their performance declines if the adversary who has stolen the model adds a noise reducer, in particular an autoencoder, to ruin the backdoor. To cope with this kind of piracy, we propose an enhanced watermarking scheme by using triggers that penetrates the adversary’s autoencoder. The penetrative triggers are generated from a collection of shadow models that approximate the adversary’s autoencoder, which is assumed to be hidden from the genuine host of the model. The proposed scheme is shown to be resistant to the filtering of autoencoders and significantly increase the robustness of ownership verification. Fangqi Li 0001, Shi-Lin Wang |
ICIP | 2 |
| 2021 | Speaker-Independent Lipreading By Disentangled Representation LearningabstractWith the development of the deep learning technology, automatic lipreading based on deep neural network can achieve reliable results for speakers appeared in the training dataset. However, speaker-independent lipreading, i.e. lipreading for unseen speakers, is still a challenging task, especially when the training samples are quite limited. To improve the recognition performance in the speaker-independent scenario, a new deep neural network structure, named Disentangled Visual Speech Recognition Network (DVSR-Net), is proposed in this paper. DVSR-Net is designed to disentangle the identity-related features and the content-related features from the lip image sequence. To further eliminate the identity information that remained in the content features, a content feature refinement stage is designed in network optimization. By this way, the extracted features are closely related to the content information and irrelevant to the various talking style and thus the speech recognition performance for unseen speakers can be improved. Experiments on two widely used datasets have demonstrated the effectiveness of the proposed network in the speaker-independent scenario. Shi-Lin Wang, Gongliang Chen |
ICIP | 2 |
| 2021 | Heterogeneous ensemble selection for evolving data streams
Anh Vu Luong, Tien Thanh Nguyen, Alan Wee-Chung Liew, Shi-Lin Wang |
Pattern Recognit. | 4 |
| 2021 | Large-Scale Malicious Software Classification With Fuzzified Features and Boosted Fuzzy Random ForestabstractClassification of malicious software, especially in a very large dataset, is a challenging task for machine intelligence. Malware can have highly diversified features, each of which has highly heterogeneous distributions. These factors increase the difficulties for traditional data analytic approaches to deal with them. Although deep learning based methods have reported good classification performance, the deep models usually lack interpretability and are fragile under adversarial attacks. To solve these problems, fuzzy systems have become a competitive candidate in malware analysis. In this article, a new fuzzy-based approach is proposed for malware classification. We focused on portable executable files in the Windows platform and analyzed the distributions of static features and content-oriented features. Fuzzification was used to reduce the ubiquitous impact of noise and outliers in a very large dataset. Finally, a novel boosted classifier consisted of fuzzy decision trees and support vector machine is proposed to perform the malware classification. By using fuzzy decision trees, the inner structure of the classifier can be readily interpreted as discriminative rules, whereas the novel boosting strategy provides state-of-the-art classification performance. Extensive experimental results showed that our method significantly outperformed several state-of-the-art classifiers. Fangqi Li 0001, Shi-Lin Wang, Alan Wee-Chung Liew, Weiping Ding 0001, Gongshen Liu |
IEEE Trans. Fuzzy Syst. | 2 |
| 2021 | Preventing DeepFake Attacks on Speaker Authentication by Dynamic Lip Movement AnalysisabstractRecent research has demonstrated that lip-based speaker authentication systems can not only achieve good authentication performance but also guarantee liveness. However, with modern DeepFake technology, attackers can produce the talking video of a user without leaving any visually noticeable fake traces. This can seriously compromise traditional face-based or lip-based authentication systems. To defend against sophisticated DeepFake attacks, a new visual speaker authentication scheme based on the deep convolutional neural network (DCNN) is proposed in this paper. The proposed network is composed of two functional parts, namely, the Fundamental Feature Extraction network (FFE-Net) and the Representative lip feature extraction and Classification network (RC-Net). The FFE-Net provides the fundamental information for speaker authentication. As the static lip shape and lip appearance is vulnerable to DeepFake attacks, the dynamic lip movement is emphasized in the FFE-Net. The RC-Net extracts high-level lip features that discriminate against human imposters while capturing the client's talking style. A multi-task learning scheme is designed, and the proposed network is trained end-to-end. Experiments on the GRID and MOBIO datasets have demonstrated that the proposed approach is able to achieve an accurate authentication result against human imposters and is much more robust against DeepFake attacks compared to three state-of-the-art visual speaker authentication algorithms. It is also worth noting that the proposed approach does not require any prior knowledge of the DeepFake spoofing method and thus can be applied to defend against different kinds of DeepFake attacks. Chenzhao Yang, Shi-Lin Wang, Alan Wee-Chung Liew |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2020 | Feature Extraction For Visual Speaker Authentication Against Computer-Generated Video AttacksabstractRecent research shows that the lip feature can achieve reliable authentication performance with a good liveness detection ability. However, with the development of sophisticated human face generation methods by the deepfake technology, the talking videos can be forged with high quality and the static lip information is not reliable in such case. Meeting with such challenge, in this paper, we propose a new deep neural network structure to extract robust lip features against human and Computer-Generated (CG) imposters. Two novel network units, i.e. the feature-level Difference block (Diffblock) and the pixel-level Dynamic Response block (DRblock), are proposed to reduce the influence of the static lip information and to represent the dynamic talking habit information. Experiments on the GRID dataset have demonstrated that the proposed network can extract discriminative and robust lip features and outperform two state-of-the-art visual speaker authentication approaches in both human imposter and CG imposter scenarios. Shi-Lin Wang, Aixin Zhang, Alan Wee-Chung Liew |
ICIP | 2 |
| 2020 | Speaker-Independent Lipreading With Limited DataabstractRecent researches have demonstrated that with a huge annotated training dataset, some sophisticated automatic lipreading methods perform even better than a professional human lip reader. However, when the training set is limited, i.e. containing a few number of speakers, most existing lipreading approaches cannot provide accurate recognition results for unseen speakers due to the inter-speaker variability. To improve the lipreading performance in the speaker-independent scenario, a new deep neural network (DNN) is proposed in this paper. The proposed network is composed of two parts, i.e. the Transformer-based Visual Speech Recognition Network (TVSR-Net) and the Speaker Confusion Block (SC-Block). The TVSR-Net is designed to extract lip features and recognize the speech. The SC-Block aims to achieve speaker normalization by eliminating the influence of various talking styles/habits. A Multi-Task Learning (MTL) scheme is designed for network optimization. Experiment results on the GRID dataset have demonstrated the effectiveness of the proposed network on speaker-independent recognition with limited training data. Chenzhao Yang, Shi-Lin Wang, Xingxuan Zhang |
ICIP | 2 |
| 2020 | Weakly Supervised Attention Rectification for Scene Text RecognitionabstractScene text recognition has become a hot topic in recent years due to its booming real-life applications. Attention-based encoder-decoder framework has become one of the most popular frameworks especially in the irregular text scenario. However, the “attention drift” problem hinders the recognition performance for most existing attention-based scene text recognition methods. To solve this problem, we propose an auxiliary supervision branch along with the attention-based encoder-decoder framework. A new loss function is designed to refine the feature map and to help the attention region align the target character area. Compared with existing attention rectification mechanisms, our method does not require character-level annotations or introduce any additional trainable parameter. Furthermore, our method can improve the performance for both RNN-Attention and Scaled Dot-Product Attention. The experiment results on various benchmarks have demonstrated that the proposed approach outperforms the state-of-the-art methods in both regular and irregular text recognition scenarios. Chengyu Gu, Shi-Lin Wang, Yiwei Zhu, Kai Chen 0006 |
ICPR | 2 |
| 2020 | Lip Image Segmentation Based on a Fuzzy Convolutional Neural NetworkabstractResearch has shown that the human lip and its movements are a rich source of information related to speech content and speaker's identity. Lip image segmentation, as a fundamental step in many lip-reading and visual speaker authentication systems, is of vital importance. Because of variations in lip color, lighting conditions and especially the complex appearance of an open mouth, accurate lip region segmentation is still a challenging task. To address this problem, this article proposes a new fuzzy deep neural network having an architecture that integrates fuzzy units and traditional convolutional units. The convolutional units are used to extract discriminative features at different scales to provide comprehensive information for pixel-level lip segmentation. The fuzzy logic modules are employed to handle various kinds of uncertainties and to provide a more robust segmentation result. An end-to-end training scheme is then used to learn the optimal parameters for both the fuzzy and the convolutional units. A dataset containing more than 48 000 images of various speakers, under different lighting conditions, was used to evaluate lip segmentation performance. According to the experimental results, the proposed method achieves state-of-the-art performance when compared with other algorithms. Cheng Guan, Shi-Lin Wang, Alan Wee-Chung Liew |
IEEE Trans. Fuzzy Syst. | 2 |
| 2019 | Spatio-Temporal Fusion Based Convolutional Sequence Learning for Lip ReadingabstractCurrent state-of-the-art approaches for lip reading are based on sequence-to-sequence architectures that are designed for natural machine translation and audio speech recognition. Hence, these methods do not fully exploit the characteristics of the lip dynamics, causing two main drawbacks. First, the short-range temporal dependencies, which are critical to the mapping from lip images to visemes, receives no extra attention. Second, local spatial information is discarded in the existing sequence models due to the use of global average pooling (GAP). To well solve these drawbacks, we propose a Temporal Focal block to sufficiently describe short-range dependencies and a Spatio-Temporal Fusion Module (STFM) to maintain the local spatial information and to reduce the feature dimensions as well. From the experiment results, it is demonstrated that our method achieves comparable performance with the state-of-the-art approach using much less training data and much lighter Convolutional Feature Extractor. The training time is reduced by 12 days due to the convolutional structure and the local self-attention mechanism. Xingxuan Zhang, Shi-Lin Wang |
ICCV | 3 |
| 2019 | Lip Image Segmentation in Mobile Devices Based on Alternative Knowledge DistillationabstractLip image segmentation, as the first step in many lip-related tasks (e.g. automatic lipreading), is of vital significance for the subsequent procedures. Nowadays, with the increasing computational power of the mobile devices, mobile applications become more and more popular. In this paper, a new approach is proposed, which is able to segment the lip region in natural scenes and is of acceptable computational complexity to be implemented in mobile devices. Two networks including a complex teacher network and a compact student network with the same structure are employed. With the proposed remedy loss and the alternative knowledge distillation scheme, the student network can learn useful knowledge from the teacher network effectively and efficiently, and even rectify some of its segmentation errors. A dataset containing 49 people captured under natural scenes by various cellphone cameras is adopted for evaluation and the experiment results have demonstrated that the proposed student network even outperforms the teacher network with much less computational cost. Cheng Guan, Shi-Lin Wang, Gongshen Liu, Alan Wee-Chung Liew |
ICIP | 2 |
| 2019 | Viewpoint Estimation in Images by a Key-Point Based Deep Neural NetworkabstractViewpoint estimation in a 2D image is a challenging task due to the great variations in the object's shape, appearance, visible parts, etc. To overcome the above difficulties, a new deep neural network is proposed, which employs the key-points of the object as a regularization term and a semantic bridge connecting the raw pixels with the object's viewpoint. A series of Hourglass structures are adopted for key-point extraction. With the extracted key-points, an LSTM based network is designed to model both the intrinsic relationship among the key-points and the underlying connections between the key-points and the object's viewpoint. A multitasks learning scheme is designed to optimize the key-point detection and the viewpoint estimation performance simultaneously. The experiment results on the PASCAL 3D+ dataset have demonstrated the effectiveness of the proposed approach. Jiana Yang, Shi-Lin Wang, Gongshen Liu |
ICIP | 2 |
| 2019 | Text Recognition in Images Based on Transformer with Hierarchical AttentionabstractRecognizing text in images has been a hot research topic in computer vision for decades due to its various application. However, the variations in text appearance in term of perspective distortion, text line curvature, text styles, etc., cause great trouble in text recognition. Inspired by the Transformer structure [1] that achieved outstanding performance in many natural language processing related applications, we propose a new Transformer-like structure for text recognition in images, which is referred to as the Hierarchical Attention Transformer Network (HATN). The entire network can be trained end-to-end by using only images and sentence-level annotations. A new hierarchical attention mechanism is proposed to lean the character-level, word-level and sentence-level contexts more efficiently and sufficiently. Extensive experiments on seven public datasets with regular and irregular text arrangements have demonstrated that the proposed HATN can achieve accurate recognition results with high efficiency. Yiwei Zhu, Shi-Lin Wang, Kai Chen 0006 |
ICIP | 2 |
| 2019 | Visual Speaker Authentication by a CNN-Based Scheme with Discriminative Segment Analysis
Shi-Lin Wang, Quanhai Zhang |
ICONIP (4) | 2 |
| 2019 | A Bayesian Possibilistic C-Means clustering approach for cervical cancer screening
Fangqi Li 0001, Shi-Lin Wang, Gongshen Liu |
Inf. Sci. | 2 |
| 2019 | A global and local context integration DCNN for adult image classification
Shi-Lin Wang, Alan Wee-Chung Liew, Gongshen Liu |
Pattern Recognit. | 2 |
| 2018 | Detecting Double Jpeg Compression with Same Quantization Matrix Based on Dense Cnn FeatureabstractDetection of double JPEG compression with same quantization matrix has been regarded as a challenging task in digital image forensics because there are very few modification cues in the tampered images especially when the compression quality factor is low. In order to solve this problem, a comprehensive feature representation based on the dense CNN framework is proposed, which is sensitive to the artifacts caused by double JPEG compression and is not related to the image content. With the appropriate network design and contributing to the characteristics of average pooling, dense connection and transition, the proposed network can differentiate double JPEG compression artifacts accurately. Experiment results on the two datasets have demonstrated that the proposed feature outperforms several state-of-the-art approaches investigated. Xiaosa Huang, Shi-Lin Wang, Gongshen Liu |
ICIP | 2 |
| 2018 | 3D Convolutional Neural Networks Based Speaker Identification and AuthenticationabstractResearch shows that human lips can be used as a new kind of biometrics in personal identification and authentication. In this letter, a novel end-to-end method based on 3D convolutional neural network (3DCNN) is proposed to extract discriminative spatiotemporal features from raw lip video streams. In our approach, the lip video is first divided into a series of overlapping clips. For each clip, the lip-characteristics network is proposed to characterize the minutiae of the lip region and its movement. Finally, the entire lip video is represented by a set of sub-features corresponding to each clip in it. Experiments have been performed on a dataset with 200 speakers and the proposed method achieves high identification accuracy of 99.18% and very low authentication error (HTER of 0.15%). Compared with several state-of-the-art methods, our approach achieves better performance and higher robustness against variations caused by different speaker's pose and position. Jianguo Liao, Shi-Lin Wang, Xingxuan Zhang, Gongshen Liu |
ICIP | 2 |
| 2018 | Adult Image Classification by a Local-Context Aware NetworkabstractTo build a healthy online environment, adult image recognition is a crucial and challenging task. Recent deep learning based methods have brought great advances to this task. However, the recognition accuracy and generalization ability need to be further improved. In this paper, a local-context aware network is proposed to improve the recognition accuracy and a corresponding curriculum learning strategy is proposed to guarantee a good generalization ability. The main idea is to integrate the global classification and the local sensitive region detection into one network and optimize them simulatenously. Such strategy helps the classification networks focus more on suspicious regions and thus provide better recognition performance. Two datasets containing over 150,000 images have been collected to evaluate the performance of the proposed approach. From the experiment results, it is observed that our approach can always achieve the best classification accuracy compared with several state-of-the-art approaches investigated. Shi-Lin Wang, Huanrong Sun, Gongshen Liu |
ICIP | 3 |
| 2018 | Visual speaker authentication with random prompt texts by a dual-task CNN framework
Shi-Lin Wang, Alan Wee-Chung Liew |
Pattern Recognit. | 2 |
| 2018 | Variational inference based bayes online classifiers with concept drift adaptation
Thi Thu Thuy Nguyen, Tien Thanh Nguyen, Alan Wee-Chung Liew, Shi-Lin Wang |
Pattern Recognit. | 4 |
| 2018 | Detection of Double Compression With the Same Coding Parameters Based on Quality Degradation Mechanism AnalysisabstractDetection of double compression with the same coding parameters is a very challenging problem in video forensics, since traces of recompression operations are extremely slight in this case. To solve this problem, we first analyze degradation mechanisms during recompression. It is observed that the video quality tends to become nearly unchanged after multiple recompressions with the same coding parameters. The degree of quality degradation is used to distinguish single and double compressed videos. This property can be described using the convergent tendency of video data to unchanged states after continuous recompressions. For MPEG videos, statistical features of rounding and truncation errors are extracted from the intra-coding process while macroblock-mode based features are obtained from the inter-coding process. The final feature is generated by concatenating these two sets of features to provide robust detection capability. Then, extracted features are fed to the SVM classifier to obtain the final detection result. In addition, aforementioned features are modified and extended to detect double compression on H.264 videos based on the unique coding techniques developed in the H.264 standard, such as intra-prediction. Several public available YUV sequences are used to construct double compression databases with three popular coding standards, including MPEG-2, MPEG-4, and H.264. In experiments, the proposed method outperforms several state-of-the-art methods for different compression qualities and rate control schemes. Experimental results demonstrate the proposed method has more robust detection capability of double compression under various encoding configurations. Xinghao Jiang, Peisong He, Tanfeng Sun, Shi-Lin Wang |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2017 | 3D human action recognition based on the Spatial-Temporal Moving Skeleton DescriptorabstractWith the popularization of the Kinect sensor, human actions can be recognized based on the 3D skeletal information. In this paper, the Spatial-Temporal Moving Skeleton Descriptor (STMSD) is proposed by the fusion of three complementary features which are the Relative Geometric Velocity (RGV) between body parts, Relative Joint Positions (RJP), and Joint Angles (JA). The STMSD descriptor gives a complete view of the body skeleton in space and time. Among the three features, the Relative Geometric Velocity (RGV) is first proposed in our work. Inspired by the relative geometry using the Lie group and the Lie algebra, RGV describes the variation rates of body transformations which include 3D rotations and translations. Then interpolation and normalization are applied in frame descriptors. After the temporal modeling, Principal Component Analysis (PCA) is utilized. Experimental results on three datasets show that our approach performs better than existing action recognition approaches, including skeleton-based and other types. Hongxian Yao, Xinghao Jiang, Tanfeng Sun, Shi-Lin Wang |
ICME | 4 |
| 2017 | A Novel Image Classification Method with CNN-XGBoost Model
Xu-Die Ren, Shenghong Li 0001, Shi-Lin Wang, Jianhua Li 0001 |
IWDW | 4 |
| 2017 | Detection of double compression in MPEG-4 videos based on block artifact measurement
Peisong He, Xinghao Jiang, Tanfeng Sun, Shi-Lin Wang |
Neurocomputing | 4 |
| 2017 | Frame-wise detection of relocated I-frames in double compressed H.264 videos based on convolutional neural network
Peisong He, Xinghao Jiang, Tanfeng Sun, Shi-Lin Wang, Bin Li 0011 |
J. Vis. Commun. Image Represent. | 4 |
| 2016 | Detecting double MPEG compression with the same quantiser scale based on MBM featureabstractDetecting double MPEG compression is of prime significance in video forensics. However, existing methods are effective only when the primary compression and the secondary compression have different quantiser scales (QS). There is a lack of effective methods dealing with double MPEG compression with the same QS. In this paper, a novel method based on the statistical feature of macroblock mode (MBM) which consists of macroblock type and motion vector in P-frames is proposed to detect double MPEG compression with the same QS. The MBM statistical feature is extracted during multiple decoding procedures when the video is repeatedly compressed with the same QS for several times. Finally, the proposed feature is combined with the support vector machine (SVM) to classify the single MPEG compression and double MPEG compression. Experiments have demonstrated the effectiveness of the proposed method and the robustness to a wide range of QSs and different encoders. Jieyuan Chen, Xinghao Jiang, Tanfeng Sun, Peisong He, Shi-Lin Wang |
ICASSP | 5 |
| 2016 | Visual speaker authentication by ensemble learning over static and dynamic lip detailsabstractThis paper presents a new visual speaker authentication scheme which can extract the most representative details of a speaker's lip feature. For each speaker, the entire utterance pronouncing a specific prompt text is divided into several word-level segments and a mute segment. Three kinds of lip feature details are investigated including: i) lip movements in each word segment; ii) lip movements in each word transition; and iii) lip appearance in the mute segment. An HMM with model adaptation is adopted to depict each dynamic details and a linear SVM is used to differentiate the static details. A confident measure is introduced to evaluate the discriminative ability of each details between the speaker and other speakers. Finally, an ensemble learning structure is proposed which focuses on the discriminative details and provides a reliable authentication result. Experimental results have demonstrated that our scheme achieves better performance compared with some traditional approaches investigated. Xiao-Xing Shi, Shi-Lin Wang, Jun-Yao Lai |
ICIP | 2 |
| 2016 | Detecting Double H.264 Compression Based on Analyzing Prediction Residual Distribution
Tanfeng Sun, Xinghao Jiang, Peisong He, Shi-Lin Wang, Yun Q. Shi 0001 |
IWDW | 5 |
| 2016 | Visual speaker identification and authentication by joint spatiotemporal sparse coding and hierarchical pooling
Jun-Yao Lai, Shi-Lin Wang, Alan Wee-Chung Liew, Xing-Jian Shi |
Inf. Sci. | 2 |
| 2016 | Double compression detection based on local motion vector field analysis in static-background videos
Peisong He, Xinghao Jiang, Tanfeng Sun, Shi-Lin Wang |
J. Vis. Commun. Image Represent. | 4 |
| 2015 | Double Compression Detection in MPEG-4 Videos Based on Block Artifact Measurement with Variation of Prediction Footprint
Peisong He, Tanfeng Sun, Xinghao Jiang, Shi-Lin Wang |
ICIC (3) | 4 |
| 2015 | Image-splicing forgery detection based on local binary patterns of DCT coefficientsabstractAbstract The wide use of high‐performance image acquisition devices and powerful image‐processing software has made it easy to tamper images for malicious purposes. Image splicing, which has constituted a menace to integrity and authenticity of images, is a very common and simple trick in image tampering. Therefore, image‐splicing detection is of great importance in digital forensics. In this paper, an effective framework for revealing image‐splicing forgery is proposed. First, the local binary pattern operator is used to model magnitude components of two‐dimensional arrays obtained by applying multisize block discrete cosine transform to test images. Then, all of bins of histograms computed from local binary pattern codes are served as discriminative features for image‐splicing detection. After that, kernel principal component analysis is utilized to reduce the dimensionality of the proposed features to avoid the high computational complexity, high mutual correlation among the constructed features and possible overfitting for support vector machine classifier. Finally, support vector machine classifier is employed to distinguish spliced images from authentic images by using the final dimensionality‐reduced feature set. The experiment results show that the proposed method can perform better than some state‐of‐the‐art methods in terms of the detection performance over the Columbia image‐splicing detection evaluation dataset. Copyright © 2013 John Wiley & Sons, Ltd. Chenglin Zhao, Yiming Pi, Shenghong Li 0001, Shi-Lin Wang |
Secur. Commun. Networks | 5 |
| 2015 | Passive Image-Splicing Detection by a 2-D Noncausal Markov ModelabstractIn this paper, a 2-D noncausal Markov model is proposed for passive digital image-splicing detection. Different from the traditional Markov model, the proposed approach models an image as a 2-D noncausal signal and captures the underlying dependencies between the current node and its neighbors. The model parameters are treated as the discriminative features to differentiate the spliced images from the natural ones. We apply the model in the block discrete cosine transformation domain and the discrete Meyer wavelet transform domain, and the cross-domain features are treated as the final discriminative features for classification. The support vector machine which is the most popular classifier used in the image-splicing detection is exploited in our paper for classification. To evaluate the performance of the proposed method, all the experiments are conducted on public image-splicing detection evaluation data sets, and the experimental results have shown that the proposed approach outperforms some state-of-the-art methods. Shi-Lin Wang, Shenghong Li 0001, Jianhua Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2014 | A Distributed Local Margin Learning based scheme for high-dimensional feature processing in image tampering detectionabstractWith the development of image tampering detection, more and more features are involved to improve the detection rate, and nowadays the high-dimensional features based methods get the state-of-the-art detection accuracies. However, the high dimensionality will cause excessive time cost in the classification phase, moreover it would probably introduce redundant features which will confuse the classifier. An effective scheme based on Distributed Local Margin Learning (D-LML) is proposed in this paper to solve the problems caused by the high-dimensionality of features in the image tampering detection work. Local Margin Learning algorithm distributed to different clients is employed to rank the importance of the original features, and we can get features with low dimensionality by preserving the important features while excluding the insignificant features. Experimental results show that the D-LML method could greatly reduce the dimensionality of the original features, and keep the detection rates fluctuating in a relatively small range. Jianhua Li 0001, Shi-Lin Wang, Shenghong Li 0001 |
ICME | 3 |
| 2014 | Revealing the Traces of Median Filtering Using High-Order Local Ternary PatternsabstractRecently, detecting the traces introduced by the content-preserving image manipulations has received a great deal of attention from forensic analyzers. It is well known that the median filter is a widely used nonlinear denoising operator. Therefore, the detection of median filtering is of important realistic significance in image forensics. In this letter, a novel local texture operator, named the second-order local ternary pattern (LTP), is proposed for median filtering detection. The proposed local texture operator encodes the local derivative direction variations by using a 3-valued coding function and is capable of effectively capturing the changes of local texture caused by median filtering. In addition, kernel principal component analysis (KPCA) is exploited to reduce the dimensionality of the proposed feature set, making the computational cost manageable. The experiment results have shown that the proposed scheme performs better than several state-of-the-art approaches investigated. Shenghong Li 0001, Shi-Lin Wang, Yun Q. Shi 0001 |
IEEE Signal Process. Lett. | 3 |
| 2013 | Unsupervised feature learning using Markov deep belief networkabstractRecently, deep architectures, such as Deep Belief Network (DBN), have been used to learn features from unlabeled data. However, since DBN supports bi-directional inference and the units between two layers are fully connected, it is difficult to directly apply the traditional convolutional network to DBN, or scale DBN to fit the large images (e.g. 1024×768). In this paper, a new deep learning model, named Markov DBN (MDBN), is proposed to address these problems. This model employs a new way for DBN to reduce computational burden and handle large images. Markov sub-layers are also adopted to take the neighboring relationship of the inputs into consideration. To train MDBN, we devise Block Restricted Boltzmann Machine (BRBM) which chooses non-overlapping blocks as input. Furthermore, SIFT descriptor is employed to enable this model to learn translation, scaling and rotation invariant features. Experimental results on datasets Caltech-101 and Caltech-256 have demonstrated the superiority of our model. Dongyang Cheng, Tanfeng Sun, Xinghao Jiang, Shi-Lin Wang |
ICIP | 4 |
| 2013 | Image splicing detection based on noncausal Markov modelabstractIn this paper, a noncausal Markov model is proposed for digital image splicing detection. Different from the traditional Markov model in image splicing detection, the proposed approach models an observation array as a 2-D noncausal signal and captures the underlying statistical characteristics. We give the solutions to the model and the model parameters are treated as discriminative features for classification (detection). To evaluate the generalization and effectiveness of the proposed method, we apply the model in the block DCT domain and discrete Meyer wavelet transform domain respectively and experimental results have shown that the proposed approach outperforms most of the state-of-the-art methods. Shi-Lin Wang, Shenghong Li 0001, Jianhua Li 0001, Quanqiao Yuan |
ICIP | 2 |
| 2013 | Identifying Video Forgery Process Using Optical Flow
Wan Wang, Xinghao Jiang, Shi-Lin Wang, Meng Wan, Tanfeng Sun |
IWDW | 3 |
| 2013 | A Distributed Scheme for Image Splicing Detection
Shi-Lin Wang, Shenghong Li 0001, Jianhua Li 0001 |
IWDW | 2 |
| 2013 | Estimation of the primary quantization parameter in MPEG videosabstractThe advanced technology and sophisticated software have rendered audiovisual content exposed to forgery, inspiring the emergence of multimedia forensic research. Since video tampering may involve double compression, the analysis of compression history is of significance. In this paper, we consider the processing chains of two compression steps and propose an algorithm that aims at identifying the quantization parameter used in the previous coding process. The method relies on the fact that characteristic footprints can be observed under different relationships between quantization parameters of consecutive compression operations. Features are extracted from both Discrete Cosine Transform (DCT) coefficients and their differential counterparts to capture the statistical disturbance. Experimental results demonstrate the effectiveness of our method. Wan Wang, Xinghao Jiang, Shi-Lin Wang, Tanfeng Sun |
VCIP | 3 |
| 2013 | Detection of Double Compression in MPEG-4 Videos Based on Markov StatisticsabstractWith the spread of powerful and easy-to-use video editing software, digital videos are exposed to various forms of tampering. Nowadays, a considerable proportion of surveillance systems and video cameras have built-in MPEG-4 codec. Therefore, the detection of double compression in MPEG-4 videos as a first step in video forensics research is of significance. In this paper, Markov based features are adopted to detect double compression artifacts, which imply that the original video may have been interpolated. The advantages and limitations of double MPEG-4 compression detection are analyzed. Experimental results have demonstrated that our scheme outperforms most existing methods. Xinghao Jiang, Wan Wang, Tanfeng Sun, Yun Q. Shi 0001, Shi-Lin Wang |
IEEE Signal Process. Lett. | 5 |
| 2012 | Countering Universal Image Tampering Detection with Histogram Restoration
Luyi Chen, Shi-Lin Wang, Shenghong Li 0001, Jianhua Li 0001 |
IWDW | 2 |
| 2012 | Rapid Image Splicing Detection Based on Relevance Vector Machine
Quanqiao Yuan, Mengying Zhai, Shi-Lin Wang |
IWDW | 5 |
| 2012 | Physiological and behavioral lip biometrics: A comprehensive study of their discriminative power
Shi-Lin Wang, Alan Wee-Chung Liew |
Pattern Recognit. | 1 |
| 2011 | New Feature Presentation of Transition Probability Matrix for Image Tampering Detection
Luyi Chen, Shi-Lin Wang, Shenghong Li 0001, Jianhua Li 0001 |
IWDW | 2 |
| 2011 | A Comprehensive Study on Third Order Statistical Features for Image Splicing Detection
Shi-Lin Wang, Shenghong Li 0001, Jianhua Li 0001 |
IWDW | 2 |
| 2011 | An automatic video content classification scheme based on combined visual features model with modified DAGSVM
Xinghao Jiang, Tanfeng Sun, Shi-Lin Wang |
Multim. Tools Appl. | 3 |
| 2010 | Detecting Digital Image Splicing in Chroma Spaces
Jianhua Li 0001, Shenghong Li 0001, Shi-Lin Wang |
IWDW | 4 |
| 2008 | A Semi-fragile Watermark Scheme Based on the Logistic Chaos Sequence and Singular Value Decomposition
Shenghong Li 0001, Shi-Lin Wang, Danhong Yao |
IDEAL | 4 |
| 2008 | An Automatic Lipreading System for Spoken Digits With Limited Training DataabstractIt is well known that visual cues of lip movement contain important speech relevant information. This paper presents an automatic lipreading system for small vocabulary speech recognition tasks. Using the lip segmentation and modeling techniques we developed earlier, we obtain a visual feature vector composed of outer and inner mouth features from the lip image sequence for recognition. A spline representation is employed to transform the discrete-time sampled features from the video frames into the continuous domain. The spline coefficients in the same word class are constrained to have similar expression and are estimated from the training data by the EM algorithm. For the multiple-speaker/speaker-independent recognition task, an adaptive multimodel approach is proposed to handle the variations caused by various talking styles. After building the appropriate word models from the spline coefficients, a maximum likelihood classification approach is taken for the recognition. Lip image sequences of English digits from 0 to 9 have been collected for the recognition test. Two widely used classification methods, HMM and RDA, have been adopted for comparison and the results demonstrate that the proposed algorithm deliver the best performance among these methods. Shi-Lin Wang, Alan Wee-Chung Liew, Wing Hong Lau, Shu Hung Leung |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2007 | Information-Based Color Feature Representation for Image ClassificationabstractFor the image classification task, the color histogram is widely used as an important color feature indicating the content of the image. However, the high-resolution color histograms are usually of high dimension and contain much redundant information which does not relate to the image content, while the low-resolution histograms cannot provide adequate discriminative information for image classification. In this paper, a new color feature representation is proposed which not only takes the correlation among neighbouring components of the conventional color histogram into account but removes the redundant information as well. A high-resolution, uniform quantized color histogram is first obtained from the image. Then the redundant bins are removed and some neighbouring bins are combined together to generate a new feature component to maximize the discriminative ability. The mutual information is adopted to evaluate the discriminative power of a specific feature set and an iterative algorithm is performed to derive the histogram quantization and their corresponding feature generation. To illustrate the effectiveness of the proposed feature representation, an application of detecting adult images, i.e., image classification between erotic and benign images, is carried out. Two widely used classification techniques, SVM and Adaboost, are employed as the classifier. Experimental results show the superior performance of our color representation compared with the conventional color histogram in image classification. Shi-Lin Wang, Alan Wee-Chung Liew |
ICIP (6) | 1 |
| 2007 | Robust lip region segmentation for lip images with complex background
Shi-Lin Wang, Wing Hong Lau, Alan Wee-Chung Liew, Shu Hung Leung |
Pattern Recognit. | 1 |
| 2004 | Lip features selection with application to person authenticationabstractAuthentication system solely based on visual lip information is of advantage since the uttering characteristics/manner is unique to individual. This paper presents the results of studying for the most appropriate features extracted from the lip motion for person authentication application. The geometric-based, shape-based and inner lip features are derived from a 14-point active shape model (ASM) lip model. The dynamic features reflecting the differential changes of the feature parameters are also considered in the experiment. A database consists of 40 speakers and each of their utterance last for 3 seconds. Various combinations of features obtained by analyzing this database are fed into a hidden Markov model (HMM) classifier for processing. It is observed that the best result has been obtained when both shape-based features and inner lip features are used. Lin Leung Mok, Wing Hong Lau, Shu Hung Leung, Shi-Lin Wang |
ICASSP (3) | 4 |
| 2004 | Lip segmentation with the presence of beardsabstractLip image analysis has attracted much interest in recent years because some important speech information is contained in the shape and movement of the lip. To extract such information from the images, accurate and robust lip region segmentation is of vital importance. However, most of the current lip segmentation methods fail to provide accurate results if the person has a beard. We propose a "one object, multiple background" clustering method to solve the problem. Since the non-lip region becomes inhomogeneous in the presence of a beard, multiple background clusters can produce better fitting to a rather complex background region than a single cluster. Spatial information in terms of the physical distance towards the lip center is incorporated to enhance the differentiation between the lip and background region. Experimental results demonstrate that our algorithm provides accurate lip segmentation results for images with beards. Shi-Lin Wang, Wing Hong Lau, Shu Hung Leung, Alan Wee-Chung Liew |
ICASSP (3) | 1 |
| 2004 | Automatic lip contour extraction from color images
Shi-Lin Wang, Wing Hong Lau, Shu Hung Leung |
Pattern Recognit. | 1 |
| 2004 | Lip image segmentation using fuzzy clustering incorporating an elliptic shape functionabstractRecently, lip image analysis has received much attention because its visual information is shown to provide improvement for speech recognition and speaker authentication. Lip image segmentation plays an important role in lip image analysis. In this paper, a new fuzzy clustering method for lip image segmentation is presented. This clustering method takes both the color information and the spatial distance into account while most of the current clustering methods only deal with the former. In this method, a new dissimilarity measure, which integrates the color dissimilarity and the spatial distance in terms of an elliptic shape function, is introduced. Because of the presence of the elliptic shape function, the new measure is able to differentiate the pixels having similar color information but located in different regions. A new iterative algorithm for the determination of the membership and centroid for each class is derived, which is shown to provide good differentiation between the lip region and the nonlip region. Experimental results show that the new algorithm yields better membership distribution and lip shape than the standard fuzzy c-means algorithm and four other methods investigated in the paper. Shu Hung Leung, Shi-Lin Wang, Wing Hong Lau |
IEEE Trans. Image Process. | 2 |
| 2003 | A new real-time lip contour extraction algorithmabstractA new lip contour extraction algorithm that combines the merits of the point-based lip model and the parametric lip model is presented. A 16-point lip model is used to describe the lip contour. With the aid of the FCMS (fuzzy clustering method incorporating shape function), a robust probability map is generated and a region-based cost function can be established. An iterative optimization procedure has been developed to fit the lip model to the probability map. In each iteration, the adjustment of the 16 lip points is governed by three pieces of quadratic curves which constrain the points to form a physical lip shape. Experimental results show that the proposed approach provides satisfactory results for 5,000 lip images of over 20 individuals. A real-time lip contour extraction system has also been implemented. Shi-Lin Wang, Wing Hong Lau, Shu Hung Leung |
ICASSP (3) | 1 |
| 2002 | Lip segmentation by fuzzy clustering incorporating with shape functionabstractIn lip segmentation, most of the clustering methods are based on color information. In this paper, a new fuzzy clustering method that combines the color information and the shape information in the measure is proposed. An elliptic shape function is introduced in the measure that modifies the measure according to the distance of the pixel from the lip center with which the differentiation between the lip and non-lip regions is enhanced. Compared with the standard fuzzy c-means algorithm, experimental results show that the new algorithm has better membership distribution and shape. This method is applicable to other images with a well-predicted shape. Shi-Lin Wang, Shu Hung Leung, Wing Hong Lau |
ICASSP | 1 |