VLDB 2026 Research / reviewers in the wild / expert
Yongjian Hu
dblp:39/737
· DBLP profile ↗
49ranked-venue papers
18as first author
22since 2021 · last 2025
0000-0002-7775-3786ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 19 · 7 first-author · 7 since 2021Security and privacy · 13 · 6 first-author · 7 since 2021Artificial intelligence and machine learning · 12 · 9 since 2021Software engineering, systems software and programming languages · 5 · 4 first-authorSystems, architecture and hardware · 2 · 1 first-authorComputer networks · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Deepfake detection with domain generalization and mask-guided supervision
Yongjian Hu, Huimin She, Chang-Tsun Li |
Pattern Recognit. | 2 |
| 2025 | Deep learning techniques for Video Instance Segmentation: A surveyabstractVideo Instance Segmentation (VIS), also known as multi-object tracking and segmentation, represents a fundamental challenge in computer vision that requires simultaneous detection, segmentation, and tracking of object instances across video frames. This complex task has gained significant attention due to its crucial role in various real-world applications. The advent of deep learning has promoted VIS approaches, leading to numerous architectural innovations and performance improvements. This survey presents a systematic review of deep learning-based VIS methods, introducing a novel categorization based on temporal modeling strategies: frame-by-frame, clip-based, in-memory feature propagation, and in-memory object query propagation. Comprehensive quantitative comparisons of existing work across three major VIS benchmark datasets are also provided. Additionally, emerging challenges in the field are explored, with several promising research directions identified, aiming to provide valuable insights for researchers and practitioners interested in VIS, while further advancing deep learning techniques for VIS. • Categorization of VIS approaches based on their temporal modeling strategies. • Comprehensive quantitative comparison of current VIS methods. • Analysis of the challenges and potential future research directions in VIS. Chenhao Xu 0003, Chang-Tsun Li, Yongjian Hu, Chee Peng Lim, Douglas C. Creighton |
Pattern Recognit. | 3 |
| 2025 | Clean-Label Attack on Face Authentication Systems Through Rolling Shutter MechanismabstractWe introduce a novel clean-label black-box face presentation attack on face authentication systems, i.e., face recognition and verification systems, under mild conditions. Different from other clean-label attacks which require inserting complicated or intensity patterns after the image-capturing phase, our designed pattern can be automatically inserted during the exposure by utilizing the rolling shutter mechanism and modulating environment LEDs in a specialized waveform. This method provides a potential way to conduct backdoor attacks in the physical domain. Additionally, we propose an optimization strategy based on evolutionary computing to optimize the parameters of the stripe patterns, enhancing the attack success rate. The experimental results on several face recognition models and face verification services provided by the leading technology companies demonstrate the effectiveness of our attack method. Our study reveals a new attack applicable in the physical world, highlighting significant security concerns for existing face recognition, verification, and face anti-spoofing techniques. Yufei Wang 0006, Haoliang Li, Liepiao Zhang, Yongjian Hu, Alex Chichung Kot |
IEEE Signal Process. Lett. | 4 |
| 2024 | Rethinking Vision Transformer and Masked Autoencoder in Multimodal Face Anti-SpoofingabstractAbstract Recently, vision transformer (ViT) based multimodal learning methods have been proposed to improve the robustness of face anti-spoofing (FAS) systems. However, there are still no works to explore the fundamental natures (e.g., modality-aware inputs, suitable multimodal pre-training, and efficient finetuning) in vanilla ViT for multimodal FAS. In this paper, we investigate three key factors (i.e., inputs, pre-training, and finetuning) in ViT for multimodal FAS with RGB, Infrared (IR), and Depth. First, in terms of the ViT inputs, we find that leveraging local feature descriptors (such as histograms of oriented gradients) benefits the ViT on IR modality but not RGB or Depth modalities. Second, in consideration of the task (FAS vs. generic object classification) and modality (multimodal vs. unimodal) gaps, ImageNet pre-trained models might be sub-optimal for the multimodal FAS task. Finally, in observation of the inefficiency on direct finetuning the whole or partial ViT, we design an adaptive multimodal adapter (AMA), which can efficiently aggregate local multimodal features while freezing majority of ViT parameters. To bridge these gaps, we propose the modality-asymmetric masked autoencoder (M $$^{2}$$ 2 A $$^{2}$$ 2 E) for multimodal FAS self-supervised pre-training without costly annotated labels. Compared with the previous modality-symmetric autoencoder, the proposed M $$^{2}$$ 2 A $$^{2}$$ 2 E is able to learn more intrinsic task-aware representation and compatible with modality-agnostic (e.g., unimodal, bimodal, and trimodal) downstream settings. Extensive experiments with both unimodal (RGB, Depth, IR) and multimodal (RGB+Depth, RGB+IR, Depth+IR, RGB+Depth+IR) settings conducted on multimodal FAS benchmarks demonstrate the superior performance of the proposed methods. One highlight is that the proposed method is robust under various missing-modality cases where previous multimodal FAS models suffer serious performance drops. We hope these findings and solutions can facilitate the future research for ViT-based multimodal FAS. Zitong Yu, Rizhao Cai, Yawen Cui, Xin Liu 0012, Yongjian Hu, Alex Chichung Kot |
Int. J. Comput. Vis. | 5 |
| 2024 | Self-Supervised 3D Action Representation Learning With Skeleton Cloud Colorizationabstract3D Skeleton-based human action recognition has attracted increasing attention in recent years. Most of the existing work focuses on supervised learning which requires a large number of labeled action sequences that are often expensive and time-consuming to annotate. In this paper, we address self-supervised 3D action representation learning for skeleton-based action recognition. We investigate self-supervised representation learning and design a novel skeleton cloud colorization technique that is capable of learning spatial and temporal skeleton representations from unlabeled skeleton sequence data. We represent a skeleton action sequence as a 3D skeleton cloud and colorize each point in the cloud according to its temporal and spatial orders in the original (unannotated) skeleton sequence. Leveraging the colorized skeleton point cloud, we design an auto-encoder framework that can learn spatial-temporal features from the artificial color labels of skeleton joints effectively. Specifically, we design a two-steam pretraining network that leverages fine-grained and coarse-grained colorization to learn multi-scale spatial-temporal features. In addition, we design a Masked Skeleton Cloud Repainting task that can pretrain the designed auto-encoder framework to learn informative representations. We evaluate our skeleton cloud colorization approach with linear classifiers trained under different configurations, including unsupervised, semi-supervised, fully-supervised, and transfer learning settings. Extensive experiments on NTU RGB+D, NTU RGB+D 120, PKU-MMD, NW-UCLA, and UWA3D datasets show that the proposed method outperforms existing unsupervised and semi-supervised 3D action recognition methods by large margins and achieves competitive performance in supervised 3D action recognition as well. Siyuan Yang 0001, Jun Liu 0036, Shijian Lu, Meng Hwa Er, Yongjian Hu, Alex Chichung Kot |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | Boosting Deepfake Feature Extractors Using Unsupervised Domain AdaptationabstractTo make deepfake detectors generalizable to different target domains, one effective way is to let the source domain for training be similar to the target domain for detection. This letter tackles the problem from the perspective of domain adaptation and achieves both image-level and feature-level domains alignments. The proposed unsupervised domain adapter accomplishes the image-level domain alignment relying on the combination of cross-domain style feature mixing and diffusion model, and the feature-level domain alignment relying on prototypical consistency guided supervision and adversarial learning. The style is transferred from source to target for the generation of target data proxy in the form of stylized images. The features of the stylized images are further aligned with the target prototypical features. We apply the domain adapter as a feature booster to four current deepfake detectors. Experimental results show that all the detectors get a significant increase in AUC values on cross-dataset testings. We further propose a deepfake detector based on the Xception backbone with our booster. Compared with five state-of-the-art detectors, the proposed detector performs best in all experiments. Yongjian Hu, Zhaolong Gong, Xiangui Kang |
IEEE Signal Process. Lett. | 2 |
| 2024 | S-Adapter: Generalizing Vision Transformer for Face Anti-Spoofing With Statistical TokensabstractFace Anti-Spoofing (FAS) aims to detect malicious attempts to invade a face recognition system by presenting spoofed faces. State-of-the-art FAS techniques predominantly rely on deep learning models but their cross-domain generalization capabilities are often hindered by the domain shift problem, which arises due to different distributions between training and testing data. In this study, we develop a generalized FAS method under the Efficient Parameter Transfer Learning (EPTL) paradigm, where we adapt the pre-trained Vision Transformer models for the FAS task. During training, the adapter modules are inserted into the pre-trained ViT model, and the adapters are updated while other pre-trained parameters remain fixed. We find the limitations of previous vanilla adapters in that they are based on linear layers, which lack a spoofing-aware inductive bias and thus restrict the cross-domain generalization. To address this limitation and achieve cross-domain generalized FAS, we propose a novel Statistical Adapter (S-Adapter) that gathers local discriminative and statistical information from localized token histograms. To further improve the generalization of the statistical tokens, we propose a novel Token Style Regularization (TSR), which aims to reduce domain style variance by regularizing Gram matrices extracted from tokens across different domains. Our experimental results demonstrate that our proposed S-Adapter and TSR provide significant benefits in both zero-shot and few-shot cross-domain testing, outperforming state-of-the-art methods on several benchmark tests. We will release the source code upon acceptance. Rizhao Cai, Zitong Yu, Chenqi Kong, Haoliang Li, Changsheng Chen 0001, Yongjian Hu, Alex Chichung Kot |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2024 | Beyond the Prior Forgery Knowledge: Mining Critical Clues for General Face Forgery DetectionabstractFace forgery detection is essential in combating malicious digital face attacks. Previous methods mainly rely on prior expert knowledge to capture specific forgery clues, such as noise patterns, blending boundaries, and frequency artifacts. However, these methods tend to get trapped in local optima, resulting in limited robustness and generalization capability. To address these issues, we propose a novel Critical Forgery Mining (CFM) framework, which can be flexibly assembled with various backbones to boost their generalization and robustness performance. Specifically, we first build a fine-grained triplet and suppress specific forgery traces through prior knowledge-agnostic data augmentation. Subsequently, we propose a fine-grained relation learning prototype to mine critical information in forgeries through instance and local similarity-aware losses. Moreover, we design a novel progressive learning controller to guide the model to focus on principal feature components, enabling it to learn critical forgery features in a coarse-to-fine manner. The proposed method achieves state-of-the-art forgery detection performance under various challenging evaluation settings. The source code is available at:https://github.com/LoveSiameseCat/CFM. Anwei Luo, Chenqi Kong, Jiwu Huang, Yongjian Hu, Xiangui Kang, Alex Chichung Kot |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2024 | Using Graph Neural Networks to Improve Generalization Capability of the Models for Deepfake DetectionabstractDeepfake detection plays a key role in preventing the misuses of artificial intelligence in video editing. Current deep learning-based deepfake detection methods often perform quite well in intra-dataset testing, but they may lose good performance in cross-dataset testing. In other words, generalization capability is still a crucial problem to be resolved. In this paper, we address deepfake detection by treating an image as non-Euclidean data and representing it as a graph so as to infer the informative connections between image patches/nodes to improve the detector’s generalization capability. Specifically, we propose a graph neural network-based paradigm that casts deepfake detection as a graph binary classification problem. First, we propose a dual-branch network to extract node features from both RGB images and their color difference images (CDIs) via the Transformer-based trainable node encoder module (TNEM). Second, we adopt the adjacency matrix to establish the connections of the nodes and further optimize the graph representation by applying the adaptive threshold to the adjacency matrix. Third, multi-head graph convolutional neural networks are carried out for node feature extraction. RGB node features and CDI node features are concatenated and separately fed into the graph classifier and node classifier for forgery detection and forgery localization. Experimental results demonstrate that our method can overall outperform other state-of-the-art methods on 7 popular benchmark datasets. Notably, our model achieves the highest AUC values of 96.19%, 80.99% and 87.68% on Celeb-DF-V2, DFDC and DFDCP in turn when trained on FF++ (C23). The visualization of node classification results also provides good interpretability of our proposed approach. Huimin She, Yongjian Hu, Chang-Tsun Li |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2023 | Rehearsal-Free Domain Continual Face Anti-Spoofing: Generalize More and Forget LessabstractFace Anti-Spoofing (FAS) is recently studied under the continual learning setting, where the FAS models are expected to evolve after encountering data from new domains. However, existing methods need extra replay buffers to store previous data for rehearsal, which becomes infeasible when previous data is unavailable because of privacy issues. In this paper, we propose the first rehearsal-free method for Domain Continual Learning (DCL) of FAS, which deals with catastrophic forgetting and unseen domain generalization problems simultaneously. For better generalization to unseen domains, we design the Dynamic Central Difference Convolutional Adapter (DCDCA) to adapt Vision Transformer (ViT) models during the continual learning sessions. To alleviate the forgetting of previous domains without using previous data, we propose the Proxy Prototype Contrastive Regularization (PPCR) to constrain the continual learning with previous domain knowledge from the proxy prototypes. Simulating practical DCL scenarios, we devise two new protocols which evaluate both generalization and anti-forgetting performance. Extensive experimental results show that our proposed method can improve the generalization performance in unseen domains and alleviate the catastrophic forgetting of previous knowledge. The code and protocol files are released on https://github.com/RizhaoCai/DCL-FAS-ICCV2023. Rizhao Cai, Yawen Cui, Zitong Yu, Haoliang Li, Yongjian Hu, Alex Chichung Kot |
ICCV | 6 |
| 2023 | Hybrid Domain Meta-Learning Network for Face Forgery Detection and Localization in DeepfakesabstractExisting face forgery detection methods often consider the manipulation detection problem as a binary classification problem, which are easily prone to overfitting. The detection performance degrades for datasets not appearing in the training stage. To solve this problem, we utilize the idea of domain generalization and design a hybrid domain meta-learning network (HDMNet) for face forgery detection and manipulation localization. The network allows the reveal of more essential face-swapping traces through multi-domain off-sets. Specifically, HDMNet consists of three modules, namely the feature extraction module (FEM), the face mask representation module (FMRM) and the meta-learning module (MLM). The FEM is composed of a comprehensive feature representation of image color features and high-frequency noise features. The FMRM utilizes graph convolution to predict face masks, providing supervision of domain knowledge. The MLM extracts domain-invariant features through a hybrid domain meta-learning strategy. Extensive experiments on several benchmark databases validate the effectiveness and good generalization ability of our method compared with several state-of-the-art methods. Hongjie Zhao, Yongjian Hu, Chang-Tsun Li |
IJCNN | 3 |
| 2023 | Self-Supervised Leaf Segmentation under Complex Lighting Conditions
Xufeng Lin, Chang-Tsun Li, Scott D. Adams, Abbas Z. Kouzani, Richard Jiang 0001, Ligang He, Yongjian Hu, Michael Vernon, Egan H. Doeven, Lawrence Webb, Todd Mcclellan, Adam Guskic |
Pattern Recognit. | 7 |
| 2023 | Learnable Information-Preserving Image Resizer for Face Forgery DetectionabstractResizing input face images of arbitrary sizes to a uniform size is an essential preprocessing to satisfy the architectural requirements of face forgery detectors. In this letter, we reveal an important observation that traditional resizing methods degrade the performance of face forgery detectors due to the loss of high-frequency information. To address this issue, we propose a simple yet effective learnable information-preserving resizer to replace its lossy traditional counterparts. Specifically, we use Haar transform to separate low-and high-frequency components, and then perform learnable resizing on the high-frequency subbands. We conduct experiments to compare our learnable resizer with other methods and evaluate three existing detectors with and without incorporating our resizer. Experimental results show that our resizer outperforms other resizers and consistently enhances the detection performance of tested detectors, confirming the effectiveness of our proposed resizer Huimin She, Yongjian Hu, Chang-Tsun Li |
IEEE Signal Process. Lett. | 2 |
| 2023 | Asymmetric Modality Translation for Face Presentation Attack DetectionabstractFace presentation attack detection (PAD) is an essentialmeasure to protect face recognition systems from being spoofed by malicious users and has attracted great attention from both academia and industry. Although most of the existing methods can achieve desired performance to some extent, the generalization issue of face presentation attack detection under cross-domain settings (e.g., the setting of unseen attacks and varying illumination) remains to be solved. In this paper, we propose a novel framework based on asymmetric modality translation for face presentation attack detection in bi-modality scenarios. Under the framework, we establish connections between two modality images of genuine faces. Specifically, a novel modality fusion scheme is presented that the image of one modality is translated to the other one through an asymmetric modality translator, then fused with its corresponding paired image. The fusion result is fed as the input to a discriminator for inference. The training of the translator is supervised by an asymmetric modality translation loss. Besides, an illumination normalization module based on Pattern of Local Gravitational Force (PLGF) representation is used to reduce the impact of illumination variation. We conduct extensive experiments on three public datasets, which validate that our method is effective in detecting various types of attacks and achieves state-of-the-art performance under different evaluation protocols. Zhi Li 0054, Haoliang Li, Yongjian Hu, Kwok-Yan Lam, Alex Chichung Kot |
IEEE Trans. Multim. | 4 |
| 2022 | Learning Meta Pattern for Face Anti-SpoofingabstractFace Anti-Spoofing (FAS) is essential to secure face recognition systems and has been extensively studied in recent years. Although deep neural networks (DNNs) for the FAS task have achieved promising results in intra-dataset experiments with similar distributions of training and testing data, the DNNs’ generalization ability is limited under the cross-domain scenarios with different distributions of training and testing data. To improve the generalization ability, recent hybrid methods have been explored to extract task-aware handcrafted features (e.g., Local Binary Pattern) as discriminative information for the input of DNNs. However, the handcrafted feature extraction relies on experts’ domain knowledge, and how to choose appropriate handcrafted features is underexplored. To this end, we propose a learnable network to extract Meta Pattern (MP) in our learning-to-learn framework. By replacing handcrafted features with the MP, the discriminative information from MP is capable of learning a more generalized model. Moreover, we devise a two-stream network to hierarchically fuse the input RGB image and the extracted MP by using our proposed Hierarchical Fusion Module (HFM). We conduct comprehensive experiments and show that our MP outperforms the compared handcrafted features. Also, our proposed method with HFM and the MP can achieve state-of-the-art performance on two different domain generalization evaluation benchmarks. Rizhao Cai, Zhi Li 0054, Renjie Wan, Haoliang Li, Yongjian Hu, Alex Chichung Kot |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2022 | One-Class Knowledge Distillation for Face Presentation Attack DetectionabstractFace presentation attack detection (PAD) has been extensively studied by research communities to enhance the security of face recognition systems. Although existing methods have achieved good performance on testing data with similar distribution as the training data, their performance degrades severely in application scenarios with data of unseen distributions. In situations where the training and testing data are drawn from different domains, a typical approach is to apply domain adaptation techniques to improve face PAD performance with the help of target domain data. However, it has always been a non-trivial challenge to collect sufficient data samples in the target domain, especially for attack samples. This paper introduces a teacher-student framework to improve the cross-domain performance of face PAD with one-class domain adaptation. In addition to the source domain data, the framework utilizes only a few genuine face samples of the target domain. Under this framework, a teacher network is trained with source domain samples to provide discriminative feature representations for face PAD. Student networks are trained to mimic the teacher network and learn similar representations for genuine face samples of the target domain. In the test phase, the similarity score between the representations of the teacher and student networks is used to distinguish attacks from genuine ones. To evaluate the proposed framework under one-class domain adaptation settings, we devised two new protocols and conducted extensive experiments. The experimental results show that our method outperforms baselines under one-class domain adaptation settings and even state-of-the-art methods with unsupervised domain adaptation. Zhi Li 0054, Rizhao Cai, Haoliang Li, Kwok-Yan Lam, Yongjian Hu, Alex Chichung Kot |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2021 | Violence Detection in Videos Based on Fusing Visual and Audio InformationabstractDetermining whether given video frames contain violent content is a basic problem in violence detection. Visual and audio information are useful for detecting violence included in a video, and are usually complementary; however, violence detection studies focusing on fusing visual and audio information are relatively rare. Therefore, we explored methods for fusing visual and audio information. We proposed a neural network containing three modules for fusing multimodal information: 1) attention module for utilizing weighted features to generate effective features based on the mutual guidance between visual and audio information; 2) fusion module for integrating features by fusing visual and audio information based on the bilinear pooling mechanism; and 3) mutual Learning module for enabling the model to learn visual information from another neural network with a different architecture. Experimental results indicated that the proposed neural network outperforms existing state-of-the-art methods on the XD-Violence dataset. Wen-Feng Pang, Qianhua He, Yongjian Hu, Yanxiong Li |
ICASSP | 3 |
| 2021 | Exposing Deepfake Videos with Spatial, Frequency and Multi-scale Temporal Artifacts
Yongjian Hu, Hongjie Zhao, Zeqiong Yu, Xiangyu Yu |
IWDW | 1 |
| 2021 | Double-Stream Segmentation Network with Temporal Self-attention for Deepfake Video Detection
Yongjian Hu |
IWDW | 3 |
| 2021 | LIS-Net: An end-to-end light interior search network for speech command recognition
Nguyen Tuan Anh, Yongjian Hu, Qianhua He, Tran Thi Ngoc Linh, Hoang Thi Kim Dung, Chen Guang |
Comput. Speech Lang. | 2 |
| 2021 | An HEVC steganalytic approach against motion vector modification using local optimality in candidate list
Shuowei Liu, Yongjian Hu, Chang-Tsun Li |
Pattern Recognit. Lett. | 2 |
| 2021 | Non-Degraded Adaptive HEVC Steganography by Advanced Motion Vector PredictionabstractCurrent video steganography operates with either the decoded frame images or the compression coding parameters, which could cause quality degradation of the reconstructed frames. In this letter, by exploiting the advanced motion vector prediction (AMVP) technique of High Efficiency Video Coding (HEVC) standard, we propose a non-degraded adaptive steganographic approach for H.265/HEVC videos. The index value in the candidate list of the prediction unit (PU) is used for embedding. Experimental results demonstrate the superiority of the proposed steganographic approach against both hand-crafted feature-based and deep learning network-based steganalytic detectors. Our work explores a new embedding space that is not previously studied. It is a significant development in finding new ways to escape from video quality change-based steganalysis. Shuowei Liu, Yongjian Hu, Xianfeng Zhao |
IEEE Signal Process. Lett. | 3 |
| 2020 | Face Anti-spoofing with a Noise-Attention Network Using Color-Channel Difference Images
Yuanyuan Ren, Yongjian Hu, Yixiang Xie |
ICANN (1) | 2 |
| 2020 | A New Efficient Finger-Vein Verification Based on Lightweight Neural Network Using Multiple Schemes
Haocong Zheng, Yongjian Hu, Alex Chichung Kot |
ICANN (1) | 2 |
| 2020 | Exposing Deepfake Videos by Tracking Eye MovementsabstractIt has recently become a major threat to the public media that fake videos are rapidly spreading over the Internet. The advent of Deepfake, a deep-learning based toolkit, has facilitated a massive abuse of improper synthesized videos, which may influence the media credibility and human rights. A worldwide alert has been set off that finding ways to detect such fake videos is not only crucial but also urgent. This paper reports a novel approach to expose deepfake videos. We found that most fake videos are markedly different from the real ones in the way the eyes move. We are thus motivated to define four features that could well capture such differences. The features are then fed to SVM for classification. It is shown to be a promising approach that without high dimensional features and complicated neural networks, we are able to achieve competitive results on several public datasets. Moreover, the proposed features could well participate with other existing methods in the confrontation with deepfakes. Yongjian Hu |
ICPR | 3 |
| 2019 | Variational Representation Learning for Vehicle Re-IdentificationabstractVehicle Re-identification is attracting more and more attention in recent years. One of the most challenging problems is to learn an efficient representation for a vehicle from its multi-viewpoint images. Existing methods tend to derive features of dimensions ranging from thousands to tens of thousands. In this work we proposed a deep learning based framework that can lead to an efficient representation of vehicles. While the dimension of the learned features can be as low as 256, experiments on different datasets show that the Top-1 and Top-5 retrieval accuracies exceed multiple state-of-the-art methods. The key to our framework is two-fold. Firstly, variational feature learning is employed to generate variational features which are more discriminating. Secondly, long short-term memory (LSTM) is used to learn the relationship among different viewpoints of a vehicle. The LSTM also plays as an encoder to downsize the features. Saghir Ahmed Saghir Alfasly, Yongjian Hu, Tiancai Liang, Xiaofeng Jin, Qingli Zhao |
ICIP | 2 |
| 2019 | Elix: Path-Selective Taint Analysis for Extracting Mobile App LinksabstractApp links, also known as mobile deep links, are URIs that point to specific pages in an app. App links are essential to many mobile experiences: Google and Bing use them to link search results directly to relevant pages in an app and apps use them for cross-app navigation. However, app links are hard to discover and, since they must be explicitly built into apps by developers, only exist for a small fraction of apps. To address these two problems, we propose Elix, an automated app link extractor. We define link extraction as a static information flow problem where a link, with its scheme and parameters, is synthesized by analyzing the data flow between subsequent pages in an app. As static analysis is prone to false positives, Elix adopts a novel, path-selective taint analysis that leverages symbolic execution to reason about path constraints and abandon infeasible paths. Elix can automatically and correctly discover links that are exposed by an app, and many others that are not explicitly exposed, thus increasing coverage of both link-enabled apps and link-enabled pages in an app. Elix also simplifies the scheme of extracted links by reducing complex types to a minimal set of primitive types. We have implemented Elix on Android and applied it to 1007 popular Android apps. Elix can extract 80-90% of an app's links, and above 80% of the extracted links are stable. Yongjian Hu, Oriana Riva, Suman Nath, Iulian Neamtiu |
MobiSys | 1 |
| 2018 | Static Detection of Event-based Races in Android AppsabstractEvent-based races are the main source of concurrency errors in Android apps. Prior approaches for scalable detection of event-based races have been dynamic. Due to their dynamic nature, these approaches suffer from coverage and false negative issues. We introduce a precise and scalable static approach and tool, named SIERRA, for detecting Android event-based races. SIERRA is centered around a new concept of "concurrency action" (that reifies threads, events/messages, system and user actions) and statically-derived order (happens-before relation) between actions. Establishing action order is complicated in Android, and event-based systems in general, because of externally-orchestrated control flow, use of callbacks, asynchronous tasks, and ad-hoc synchronization. We introduce several novel approaches that enable us to infer order relations statically: auto-generated code models which impose order among lifecycle and GUI events; a novel context abstraction for event-driven programs named action-sensitivity and finally, on-demand path sensitivity via backward symbolic execution to further rule out false positives. We have evaluated SIERRA on 194 Android apps. Of these, we chose 20 apps for manual analysis and comparison with a state-of-the-art dynamic race detector. Experimental results show that SIERRA is effective and efficient, typically taking 960 seconds to analyze an app and revealing 43 potential races. Compared with the dynamic race detector, SIERRA discovered an average 29.5 true races with 3.5 false positives, where the dynamic detector only discovered 4 races (hence missing 25.5 races per app) -- this demonstrates the advantage of a precise static approach. We believe that our approach opens the way for precise analysis and static event race detection in other event-driven systems beyond Android. Yongjian Hu, Iulian Neamtiu |
ASPLOS | 1 |
| 2018 | Modelling and developing conflict-aware scheduling on large-scale data centres
Chao Chen 0011, Ligang He, Bo Gao 0001, Jiadong Ren, Zhangjie Fu 0001, Songling Fu, Yongjian Hu, Chang-Tsun Li |
Future Gener. Comput. Syst. | 8 |
| 2017 | Defining and Detecting Environment Discrimination in Android Apps
Yunfeng Hong, Yongjian Hu, Chun-Ming Lai, Shyhtsun Felix Wu, Iulian Neamtiu, Patrick D. McDaniel, Paul L. Yu, Hasan Çam, Gail-Joon Ahn |
SecureComm | 2 |
| 2016 | Automatically verifying and reproducing event-based races in Android appsabstractConcurrency has been a perpetual problem in Android apps, mainly due to event-based races. Several event-based race detectors have been proposed, but they produce false positives, cannot reproduce races, and cannot distinguish be- tween benign and harmful races. To address these issues, we introduce a race verification and reproduction approach named ERVA. Given a race report produced by a race detector, ERVA uses event dependency graphs, event flipping, and replay to verify the race and determine whether it is a false positive, or a true positive; for true positives, ERVA uses state comparison to distinguish benign races from harmful races. ERVA automatically produces an event schedule that can be used to deterministically reproduce the race, so developers can fix it. Experiments on 16 apps indicate that only 3% of the races reported by race detectors are harmful, and that ERVA can verify an app in 20 minutes on average. Yongjian Hu, Iulian Neamtiu, Arash Alavi 0001 |
ISSTA | 1 |
| 2015 | Versatile yet lightweight record-and-replay for AndroidabstractRecording and replaying the execution of smartphone apps is useful in a variety of contexts, from reproducing bugs to profiling and testing. Achieving effective record-and-replay is a balancing act between accuracy and overhead. On smartphones, the act is particularly complicated, because smartphone apps receive a high-bandwidth stream of input (e.g., network, GPS, camera, microphone, touchscreen) and concurrency events, but the stream has to be recorded and replayed with minimal overhead, to avoid interfering with app execution. Prior record-and-replay approaches have focused on replaying machine instructions or system calls, which is not a good fit on smartphones. We propose a novel, stream-oriented record-and-replay approach which achieves high-accuracy and low-overhead by aiming at a sweet spot: recording and replaying sensor and network input, event schedules, and inter-app communication via intents. To demonstrate the versatility of our approach, we have constructed a tool named VALERA that supports record-and-replay on the Android platform. VALERA works with apps running directly on the phone, and does not require access to the app source code. Through an evaluation on 50 popular Android apps, we show that: VALERA's replay fidelity far exceeds current record-and-replay approaches for Android; VALERA's precise timing control and low overhead (about 1% for either record or replay) allows it to replay high-throughput, timing-sensitive apps such as video/audio capture and recognition; and VALERA's support for event schedule replay enables the construction of useful analyses, such as reproducing event-driven race bugs. Yongjian Hu, Tanzirul Azim, Iulian Neamtiu |
OOPSLA | 1 |
| 2015 | Fast source camera identification using matching signs between query and reference fingerprintsabstractFast camera fingerprint search is an important issue for source camera identification in real-world applications. So far there has been little work done in this area. In this paper, we propose a novel fast search algorithm. We use global information derived from the relationship between the query fingerprint/digest and the reference fingerprints/digests in the database to guide fast search. This information can provide more accurate and robust clues for the selection of candidate matching database fingerprints. Because the quality of query fingerprints may degrade or vary in realistic applications, the construction of robust search clues is significant. To speed up the search process, we adopt a lookup table that is built on the separate-chaining hash table. The proposed algorithm has been tested using query images from real-world photos. Experiments demonstrate that our algorithm can well adapt to query fingerprints with different quality. It can achieve higher detection rates with lower computational cost than the traditional brute-force search algorithm and a pioneering fast search algorithm in literature. Yongjian Hu, Chang-Tsun Li, Zhimao Lai |
Multim. Tools Appl. | 1 |
| 2014 | An efficient reversible data hiding scheme using prediction and optimal side information selection
Jiangqun Ni, Yongjian Hu |
J. Vis. Commun. Image Represent. | 3 |
| 2013 | Exposing image forgery through the detection of contrast enhancementabstractIn this paper, a novel forensic method of exposing cut-and-paste image forgery through detecting contrast enhancement is proposed. We reveal the inter-channel correlation introduced by color image interpolation, and show how a linear or nonlinear contrast enhancement can disturb this natural inter-channel dependency. We then construct a metric to measure these correlations, which are useful in distinguishing the original and contrast enhanced images. The effectiveness of the proposed algorithm is experimentally validated on natural color images captured by commercial cameras. Finally, its robustness against some anti-forensic algorithms is also discussed. Xufeng Lin, Chang-Tsun Li, Yongjian Hu |
ICIP | 3 |
| 2013 | An improved VLC-based lossless data hiding scheme for JPEG images
Yongjian Hu, Zheming Lu 0001 |
J. Syst. Softw. | 1 |
| 2013 | A high capacity lossless data hiding scheme for JPEG images
Zheming Lu 0001, Yongjian Hu |
J. Syst. Softw. | 3 |
| 2012 | Audio Forgery Detection Based on Max Offsets for Cross Correlation between ENF and Reference Signal
Yongjian Hu, Chang-Tsun Li, Zhisheng Lv |
IWDW | 1 |
| 2010 | Source camera identification from significant noise residual regionsabstractThis paper investigates the digital forensic problem of determining whether an image has been produced by a specific digital camera. We employ the binary hypothesis testing scheme to detect the presence of photo-response non-uniformity( PRNU) in the image. The main challenge of this scheme is the extremely weak amount of PRNU in the observed noise residual. We propose to extract from the noise residual the significant regions with higher signal quality and discard those regions heavily deteriorated by irrelevant noises. Experimental results demonstrate that the proposed algorithm can improve the identification performance in the sense of decreasing the false rejection rate, which is a critical measure in practical applications. Yongjian Hu, Heung-Kyu Lee |
ICIP | 2 |
| 2010 | Using improved imaging sensor pattern noise for source camera identificationabstractThe imaging sensor (e.g., CCD) pattern noise is a noise-like spread-spectrum signal inherently cast onto every digital image by each imaging device. It has been recognized as a reliable device fingerprint for source camera/scanner identification and image origin verification. However, one weakness of current sensor pattern noise-based camera fingerprints is that their construction only uses information from single color channel images. The resultant fingerprint can not comprehensively reflect characteristics of the camera sensor array. Taking into account the characteristics of the CFA (color filter array) structure, this work proposes a new way to construct camera fingerprints using information from all of the three single color channel images. Good detection results have been achieved on different camera models. Yongjian Hu, Chao Jian, Chang-Tsun Li |
ICME | 1 |
| 2009 | DE-Based Reversible Data Hiding With Improved Overflow Location MapabstractFor difference-expansion (DE)-based reversible data hiding, the embedded bit-stream mainly consists of two parts: one part that conveys the secret message and the other part that contains embedding information, including the 2-D binary (overflow) location map and the header file. The first part is the payload while the second part is the auxiliary information package for blind detection. To increase embedding capacity, we have to make the size of the second part as small as possible. Tian's classical DE method has a large auxiliary information package. Thodi mitigated the problem by using a payload-independent overflow location map. However, the compressibility of the overflow location map is still undesirable in some image types. In this paper, we focus on improving the overflow location map. We design a new embedding scheme that helps us construct an efficient payload-dependent overflow location map. Such an overflow location map has good compressibility. Our accurate capacity control capability also reduces unnecessary alteration to the image. Under the same image quality, the proposed algorithm often has larger embedding capacity. It performs well in different types of images, including those where other algorithms often have difficulty in acquiring good embedding capacity and high image quality. Yongjian Hu, Heung-Kyu Lee |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2008 | Difference Expansion Based Reversible Data Hiding Using Two Embedding DirectionsabstractCurrent difference-expansion (DE) embedding techniques perform one layer embedding in a difference image. They do not turn to the next difference image for another layer embedding unless the current difference image has no expandable differences left. The obvious disadvantage of these techniques is that image quality may have been severely degraded even before the later layer embedding begins because the previous layer embedding has used up all expandable differences, including those with large magnitude. Based on integer Haar wavelet transform, we propose a new DE embedding algorithm, which utilizes the horizontal as well as vertical difference images for data hiding. We introduce a dynamical expandable difference search and selection mechanism. This mechanism gives even chances to small differences in two difference images and effectively avoids the situation that the largest differences in the first difference image are used up while there is almost no chance to embed in small differences of the second difference image. We also present an improved histogram-based difference selection and shifting scheme, which refines our algorithm and makes it resilient to different types of images. Compared with current algorithms, the proposed algorithm often has better embedding capacity versus image quality performance. The advantage of our algorithm is more obvious near the embedding rate of 0.5 bpp. Yongjian Hu, Heung-Kyu Lee, Kaiying Chen |
IEEE Trans. Multim. | 1 |
| 2007 | Reversible Data Hiding Using Prediction Error Values Embedding
Yongjian Hu, Heung-Kyu Lee, Kaiying Chen |
IWDW | 1 |
| 2006 | Reversible Visible Watermarking Technique for ImagesabstractThis paper proposes a reversible visible watermarking algorithm to satisfy a new application scenario where the visible watermark serves as a tag or ownership identifier, but can be completely removed at the receiver end to resume the original image data. To achieve lossless recovery of the image, the proposed algorithm consists of two processing procedures: data hiding and watermark embedding. In the first procedure, we preserve the information of image region to be covered by the visible watermark in the uncovered image portion. In the second procedure, we embed the watermark based on a user-key-controlled embedding mechanism. The two distinct procedures are integrated into a secure watermarking system by a specially designed user key. With correct user keys, authorized users can completely remove the watermark and losslessly recover the original image. Yongjian Hu, Byeungwoo Jeon |
ICIP | 1 |
| 2006 | Analysis and Comparison of Typical Reversible Watermarking Methods
Yongjian Hu, Byeungwoo Jeon, Zhiquan Lin |
IWDW | 1 |
| 2006 | Reversible Visible Watermarking and Lossless Recovery of Original ImagesabstractIn this paper, we propose a reversible visible watermarking algorithm to satisfy a new application scenario where the visible watermark serves as a tag or ownership identifier, but can be completely removed to resume the original image data. It includes two procedures: data hiding and visible watermark embedding. In order to losslessly recover both the watermark-covered and nonwatermark-covered image contents at the receiver end, the payload consists of two reconstruction data packets, one for recovering the watermark-covered region, and the other for the nonwatermark-covered region. The data hiding technique reversibly hides the payload in the image region not covered by the visible watermark. To satisfy the requirements of large capacity and high image quality, our hiding technique is based on data compression and uses a payload-adaptive scheme. It further adopts error diffusion for improving subjective image quality and arithmetic compression using a character-based model for increasing computational efficiency. The visible watermark is securely embedded based on a user-key-controlled embedding mechanism. The data hiding and the visible watermark embedding procedures are integrated into a secure watermarking system by a specially designed user key Yongjian Hu, Byeungwoo Jeon |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2006 | An algorithm for removable visible watermarkingabstractA visible watermark may convey ownership information that identifies the originator of image and video. A potential application scenario for visible watermarks was proposed by IBM where an image is originally embedded with a visible watermark before posting on the web for free observation and download. The watermarked image which serves as a "teaser." The watermark can be removed to recreate the unmarked image by request of interested buyers. Before we can design an algorithm for satisfying this application, three basic problems should be solved. First, we need to find a strategy suitable for producing large amount of visually same but numerically different watermarked versions of the image for different users. Second, the algorithm should let the embedding parameters reachable for any legal user to make the embedding process invertible. Third, an unauthorized user should be prevented from removing the embedded watermark pattern. In this letter, we propose a user-key-dependent removable visible watermarking system (RVWS). The user key structure decides both the embedded subset of watermark and the host information adopted for adaptive embedding. The neighbor-dependent embedder adjusts the marking strength to host features and makes unauthorized removal very difficult. With correct user keys, watermark removal can be accomplished in "informed detection" and the high quality unmarked image can be restored. In contrast, unauthorized operation either overly or insufficiently removes the watermark due to wrong estimation of embedding parameters, and thus, the resulting image has apparent defect. Yongjian Hu, Sam Kwong, Jiwu Huang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2003 | Image Fusion Based Visible Watermarking Using Dual-Tree Complex Wavelet Transform
Yongjian Hu, Jiwu Huang, Sam Kwong, Yiu-Keung Chan |
IWDW | 1 |
| 2002 | The Design and Application of DWT-Domain Optimum Decoders
Yongjian Hu, Sam Kwong, Yiu-Keung Chan |
IWDW | 1 |