EDBT 2026 Demo / reviewers in the wild / expert
Rui Wang 0032
dblp:06/2293-32
· DBLP profile ↗
50ranked-venue papers
3as first author
29since 2021 · last 2026
0000-0002-4792-1945ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 35 · 3 first-author · 24 since 2021Artificial intelligence and machine learning · 24 · 16 since 2021Computer networks · 5 · 1 first-author · 3 since 2021Security and privacy · 4Databases, data management, data science and information retrieval · 3Applied, interdisciplinary, general and emerging computing · 2Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | False Positives Matter: Multidimensional Localization Evaluation and Training-Free Explainable Adversarial Patch DefenseabstractAdversarial patch attacks pose a significant threat to visual systems. While current patch purification-based defense methods enhance core metrics of visual perception models, they overlook the critical issue of false positive patches, severely compromising image usability. This paper reveals the inadequacy of existing evaluations for adversarial patch defenses, and pioneers a multidimensional adversarial patch localization evaluation framework, which comprehensively quantifies false positives, recall capability, and overall localization accuracy, providing a novel perspective for comparative analysis within the field. Furthermore, building upon the observation that false positives stem from a lack of semantic understanding, we propose a Semantic-Aware Training-free Explainable Defense method (SATED). SATED achieves zero-shot patch localization, false detection correction, and decision explanation by constructing a patch reasoning chain, while simultaneously performing integrated text-guided patch inpainting. Extensive experiments across digital and physical scenarios, detection and segmentation tasks, and diverse adversarial patches, demonstrate that our method significantly reduces false positives and doubles the overall patch localization accuracy, boosting both the generalizability and explainability of the defense. Lihua Jing, Rui Wang 0032, Jinwen Zhong, Runbo Li, Zixuan Zhu 0002 |
AAAI | 2 |
| 2025 | Revisiting Change Captioning from Self-supervised Global-Part AlignmentabstractThe goal of image change captioning is to capture the content differences between two images and describe them in natural language. The key is how to learn stable content changes from noise such as viewpoint and image structure. However, current work mostly focuses on identifying changes, and the influence of global noise leads to unstable recognition of global features. In order to tackle this problem, we propose a Self-supervised Global-Part Alignment (SSGPA) network and revisit the image change captioning task by enhancing the construction process of overall image global features, enabling the model to integrate global changes such as viewpoint into local changes, and to detect and describe changes in the image through alignment. Concretely, we first design a Global-Part Transport Alignment mechanism to enhance global features and learn stable content changes through a self-supervised method of optimal transport. Further, we design a Change Fusion Adapter with pre-trained vision-language model to enhance the similar parts features of paired images, thereby enhancing global features, and expanding content changes. Extensive experiments show our method achieves the state-of-the-art results on four datasets. Feixiao Lv, Rui Wang 0032, Lihua Jing |
AAAI | 2 |
| 2025 | EntropyMark: Towards More Harmless Backdoor Watermark via Entropy-based Constraint for Open-source Dataset Copyright ProtectionabstractHigh-quality open-source datasets are essential for advancing deep neural networks. However, the unauthorized commercial use of these datasets has raised significant concerns about copyright protection. One promising approach is backdoor watermark-based dataset ownership verification (BW-DOV), in which dataset protectors implant specific backdoors into illicit models through dataset watermarking, enabling the tracing of these models through abnormal prediction behaviors. Unfortunately, the targeted nature of these BW-DOV methods can be maliciously exploited, potentially leading to harmful side effects. While existing harmless methods attempt to mitigate these risks, watermarked datasets can still negatively affect prediction results, partially compromising dataset functionality. In this paper, we propose a more harmless backdoor watermark, called EntropyMark, which improves prediction confidence without altering the final prediction results. For this purpose, an entropy-based constraint is introduced to regulate the probability distribution. Specifically, we design an iterative clean-label dataset watermarking framework. Our framework employs gradient matching and adaptive data selection to optimize backdoor injection. In parallel, we introduce a hypothesis test method grounded in entropy inconsistency to verify dataset ownership. Extensive experiments on benchmark datasets demonstrate the effectiveness, transferability, and defense resistance of our approach. Our code is available at https://github.com/AaronSun2000/EntropyMark. Ming Sun 0010, Rui Wang 0032, Zixuan Zhu 0002, Lihua Jing, Yuanfang Guo |
CVPR | 2 |
| 2025 | GPT-C: Generative PrompT CompressionabstractLarge language models (LLMs) increasingly rely on lengthy prompts to achieve complex tasks. However, longer prompts not only increase inference costs but pose significant challenges for limited context windows of LLMs. Prompt compression aims to accelerate inference in long-text scenarios while preserving performance in various tasks. Conventional extractive compression methods often lead to semantic incoherence and loss of primitive key information. To this end, we propose a Generative PrompT Compression (GPT-C) paradigm. In the absence of appropriate datasets, we design a Collaborative Ordered Agent (Co-Agent) framework to distill high-quality datasets. We then develop a task-agnostic, length-adaptive reinforcement learning strategy that controls compression length and enhances the generalization across diverse tasks. Extensive experiments demonstrate that GPT-C achieves performance comparable to the original prompts while utilizing only 20% of the tokens. Rui Wang 0032, Lihua Jing, Feixiao Lv, Zixuan Zhu 0002 |
ICASSP | 2 |
| 2025 | Multi-Task Robustness Enhancement Framework against Various Adversarial PatchesabstractAutonomous systems leveraging visual perception face a rising threat from adversarial patches, jeopardizing their robustness. Existing defense methods adaptable to various pre-trained models typically rely on observed patch characteristics or prior attack data, having difficulty adapting to new threats. This study innovatively focuses on modeling patch attack behavior instead of existing patches, proposing a unified robustness enhancement framework against various adversarial patches. Through self-supervised learning, we accurately locate diverse adversarial patches without prior attack knowledge. Furthermore, we introduce an efficient adaptive patch inpainting method to mitigate patch impact while maintaining visual coherence. Experiments show that our methods effectively boost the robustness of visual perception models against various adversarial patches across different tasks. Lihua Jing, Rui Wang 0032, Runbo Li, Zixuan Zhu 0002, Xingxing Wei 0001 |
ICRA | 2 |
| 2025 | Preventing Latent Diffusion Model-Based Image Mimicry via Angle Shifting and Ensemble LearningabstractThe remarkable progress of Latent Diffusion Models (LDMs) in image generation has raised concerns about the potential for unauthorized image mimicry. To address these concerns, studies on adversarial attacks against LDMs have gained increasing attention in recent years. However, existing methods face bottlenecks when attacking the denoising module. In this work, we reveal that the robustness of the denoising module stems from two key factors: the cancellation effect between adversarial perturbations and estimated noise, and unstable gradients caused by randomly sampled timesteps and Gaussian noise. Based on these insights, we introduce a cosine similarity adversarial loss to prevent the generation of perturbations that are easily impaired and develop a more stable optimization strategy by ensembling gradients and fixing the noise in the latent space. Additionally, we propose an alternating iterative framework to reduce memory usage by mathematically dividing the optimization process into two spaces: latent space and pixel space. Compared to previous strategies, our proposed framework reduces video memory demands without sacrificing attack effectiveness. Extensive experiments demonstrate that the alternating iterative framework and the stable optimization strategy on cosine similarity loss are more efficient and more effective. Code is available at https://github.com/MinghaoLi01/cosattack. Rui Wang 0032, Ming Sun 0010, Lihua Jing |
IJCAI | 2 |
| 2025 | McGE '25: The 3rd International Workshop on Multimedia Content Generation and Evaluation: New Methods and PracticeabstractThis workshop addresses next-generation methods in multimedia research, with a focus on content generation, quality assessment, and dataset development. These three areas are foundational for advancing multimedia technologies and applications. Emerging approaches in multimedia content generation, powered by generative AI and multimodal learning, are reshaping domains such as entertainment, advertising, education, and healthcare. At the same time, robust quality assessment is essential to ensure that generated content achieves high standards of perceptual fidelity, semantic consistency, and user satisfaction, thereby determining the real-world impact of multimedia systems. Datasets remain indispensable for training and evaluating algorithms, and innovative strategies in dataset construction-ranging from augmentation and annotation to addressing issues of bias and small-sample imbalance-are driving the development of more reliable and ethical multimedia applications. By convening leading researchers and practitioners, this workshop provides a platform to explore state-of-the-art methods, share best practices, and discuss open challenges in next-generation multimedia research. The goal is to foster interdisciplinary collaboration and inspire innovative solutions that advance the creation, evaluation, and application of multimedia content, setting new benchmarks for the field and shaping the future of multimedia technologies. Cheng Jin 0001, Mingli Song, Rui Wang 0032, Xingjiao Wu |
ACM Multimedia | 3 |
| 2025 | AED-PADA: Improving Generalizability of Adversarial Example Detection via Principal Adversarial Domain AdaptationabstractAdversarial example detection, which can be conveniently applied in many scenarios, is important in the area of adversarial defense. Unfortunately, existing detection methods suffer from poor generalization performance because their training process usually relies on the examples generated from a single known adversarial attack and there exists a large discrepancy between the training and unseen testing adversarial examples. To address this issue, we propose a novel method, named Adversarial Example Detection via Principal Adversarial Domain Adaptation (AED-PADA). Specifically, our approach identifies the Principal Adversarial Domains (PADs), i.e., a combination of features of the adversarial examples generated by different attacks, which possesses a large portion of the entire adversarial feature space. Subsequently, we pioneer to exploit Multi-source Unsupervised Domain Adaptation in adversarial example detection, with PADs as the source domains. Experimental results demonstrate the superior generalization ability of our proposed AED-PADA. Note that this superiority is particularly achieved in challenging scenarios characterized by employing the minimal magnitude constraint for the perturbations. Heqi Peng, Yunhong Wang 0001, Ruijie Yang, Beichen Li 0001, Rui Wang 0032, Yuanfang Guo |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2024 | Frequency Shuffling and Enhancement for Open Set RecognitionabstractOpen-Set Recognition (OSR) aims to accurately identify known classes while effectively rejecting unknown classes to guarantee reliability. Most existing OSR methods focus on learning in the spatial domain, where subtle texture and global structure are potentially intertwined. Empirical studies have shown that DNNs trained in the original spatial domain are inclined to over-perceive subtle texture. The biased semantic perception could lead to catastrophic over-confidence when predicting both known and unknown classes. To this end, we propose an innovative approach by decomposing the spatial domain to the frequency domain to separately consider global (low-frequency) and subtle (high-frequency) information, named Frequency Shuffling and Enhancement (FreSH). To alleviate the overfitting of subtle texture, we introduce the High-Frequency Shuffling (HFS) strategy that generates diverse high-frequency information and promotes the capture of low-frequency invariance. Moreover, to enhance the perception of global structure, we propose the Low-Frequency Residual (LFR) learning procedure that constructs a composite feature space, integrating low-frequency and original spatial features. Experiments on various benchmarks demonstrate that the proposed FreSH consistently trumps the state-of-the-arts by a considerable margin. Rui Wang 0032, Lihua Jing, Chuan Wang 0002 |
AAAI | 2 |
| 2024 | PAD: Patch-Agnostic Defense against Adversarial Patch AttacksabstractAdversarial patch attacks present a significant threat to real-world object detectors due to their practical feasibil-ity. Existing defense methods, which rely on attack data or prior knowledge, struggle to effectively address a wide range of adversarial patches. In this paper, we show two inherent characteristics of adversarial patches, semantic in-dependence and spatial heterogeneity, independent of their appearance, shape, size, quantity, and location. Seman-tic independence indicates that adversarial patches oper-ate autonomously within their semantic context, while spatial heterogeneity manifests as distinct image quality of the patch area that differs from original clean image due to the independent generation process. Based on these observations, we propose PAD, a novel adversarial patch localization and removal method that does not require prior knowledge or additional training. PAD offers patch-agnostic de-fense against various adversarial patches, compatible with any pretrained object detectors. Our comprehensive digital and physical experiments involving diverse patch types, such as localized noise, printable, and naturalistic patches, ex-hibit notable improvements over state-of-the-art works. Our code is available at https://github.com/Lihua-Jing/PAD. Lihua Jing, Rui Wang 0032, Wenqi Ren, Xin Dong 0015, Cong Zou |
CVPR | 2 |
| 2024 | Logit Standardization in Knowledge DistillationabstractKnowledge distillation involves transferring soft labels from a teacher to a student using a shared temperature-based softmax function. However, the assumption of a shared temperature between teacher and student implies a mandatory exact match between their logits in terms of logit range and variance. This side-effect limits the performance of student, considering the capacity discrepancy between them and the finding that the innate logit relations of teacher are sufficient for student to learn. To address this issue, we propose setting the temperature as the weighted standard deviation of logit and performing a plug-and-play Z-score pre-process of logit standardization before applying softmax and Kullback-Leibler divergence. Our pre-process enables student to focus on essential logit relationsfrom teacher rather than requiring a magnitude match, and can improve the performance of existing logit-based distillation methods. We also show a typical case where the conventional setting of sharing temperature between teacher and student cannot reliably yield the authentic dis-tillation evaluation; nonetheless, this challenge is success-fully alleviated by our Z-score. We extensively evaluate our method for various student and teacher models on CIFAR-100 and ImageNet, showing its significant superiority. The vanilla knowledge distillation powered by our pre-process can achieve favorable performance against state-of-the-art methods, and other distillation variants can obtain considerable gain with the assistance of our pre-process. The codes, pre-trained models and logs are released on Github. Shangquan Sun, Wenqi Ren, Jingzhi Li 0002, Rui Wang 0032, Xiaochun Cao |
CVPR | 4 |
| 2024 | MakeupAttack: Feature Space Black-Box Backdoor Attack on Face Recognition via Makeup TransferabstractBackdoor attacks pose a significant threat to the training process of deep neural networks (DNNs). As a widely-used DNN-based application in real-world scenarios, face recognition systems once implanted into the backdoor, may cause serious consequences. Backdoor research on face recognition is still in its early stages, and the existing backdoor triggers are relatively simple and visible. Furthermore, due to the perceptibility, diversity, and similarity of facial datasets, many state-of-the-art backdoor attacks lose effectiveness on face recognition tasks. In this work, we propose a novel feature space backdoor attack against face recognition via makeup transfer, dubbed MakeupAttack. In contrast to many feature space attacks that demand full access to target models, our method only requires model queries, adhering to black-box attack principles. In our attack, we design an iterative training paradigm to learn the subtle features of the proposed makeup-style trigger. Additionally, MakeupAttack promotes trigger diversity using the adaptive selection method, dispersing the feature distribution of malicious samples to bypass existing defense methods. Extensive experiments were conducted on two widely-used facial datasets targeting multiple models. The results demonstrate that our proposed attack method can bypass existing state-of-the-art defenses while maintaining effectiveness, robustness, naturalness, and stealthiness, without compromising model performance. Our code is available at https://github.com/AaronSun2000/MakeupAttack. Ming Sun 0010, Lihua Jing, Zixuan Zhu 0002, Rui Wang 0032 |
ECAI | 4 |
| 2024 | Restoring Images in Adverse Weather Conditions via Histogram Transformer
Shangquan Sun, Wenqi Ren, Xinwei Gao, Rui Wang 0032, Xiaochun Cao |
ECCV (22) | 4 |
| 2024 | HPattack: An Effective Adversarial Attack for Human Parsing
Xin Dong 0015, Rui Wang 0032, Sanyi Zhang, Lihua Jing |
MMM (2) | 2 |
| 2024 | EnsIR: An Ensemble Algorithm for Image Restoration via Gaussian Mixture ModelsabstractImage restoration has experienced significant advancements due to the development of deep learning. Nevertheless, it encounters challenges related to ill-posed problems, resulting in deviations between single model predictions and ground-truths. Ensemble learning, as a powerful machine learning technique, aims to address these deviations by combining the predictions of multiple base models. Most existing works adopt ensemble learning during the design of restoration models, while only limited research focuses on the inference-stage ensemble of pre-trained restoration models. Regression-based methods fail to enable efficient inference, leading researchers in academia and industry to prefer averaging as their choice for post-training ensemble. To address this, we reformulate the ensemble problem of image restoration into Gaussian mixture models (GMMs) and employ an expectation maximization (EM)-based algorithm to estimate ensemble weights for aggregating prediction candidates. We estimate the range-wise ensemble weights on a reference set and store them in a lookup table (LUT) for efficient ensemble inference on the test set. Our algorithm is model-agnostic and training-free, allowing seamless integration and enhancement of various pre-trained image restoration models. It consistently outperforms regression-based methods and averaging ensemble approaches on 14 benchmarks across 3 image restoration tasks, including super-resolution, deblurring and deraining. The codes and all estimated weights have been released in Github. Shangquan Sun, Wenqi Ren, Zikun Liu 0001, Hyunhee Park, Rui Wang 0032, Xiaochun Cao |
NeurIPS | 5 |
| 2024 | HIST: Hierarchical and sequential transformer for image captioningabstractAbstract Image captioning aims to automatically generate a natural language description of a given image, and most state‐of‐the‐art models have adopted an encoder–decoder transformer framework. Such transformer structures, however, show two main limitations in the task of image captioning. Firstly, the traditional transformer obtains high‐level fusion features to decode while ignoring other‐level features, resulting in losses of image content. Secondly, the transformer is weak in modelling the natural order characteristics of language. To address theseissues, the authors propose a HI erarchical and S equential T ransformer ( HIST ) structure, which forces each layer of the encoder and decoder to focus on features of different granularities, and strengthen the sequentially semantic information. Specifically, to capture the details of different levels of features in the image, the authors combine the visual features of multiple regions and divide them into multiple levels differently. In addition, to enhance the sequential information, the sequential enhancement module in each decoder layer block extracts different levels of features for sequentially semantic extraction and expression. Extensive experiments on the public datasets MS‐COCO and Flickr30k have demonstrated the effectiveness of our proposed method, and show that the authors’ method outperforms most of previous state of the arts. Feixiao Lv, Rui Wang 0032, Lihua Jing, Pengwen Dai |
IET Comput. Vis. | 2 |
| 2024 | Granularity-Aware Single-Point Scene Text Spotting With Sequential Recurrence Self-AttentionabstractScene text spotting, a unified framework between text detection and text recognition, has made great progress in recent years. Existing methods usually adopt the fully-supervised learning strategy, which relies on time-consuming location annotations, particularly for scene texts with arbitrary shapes. In this paper, we propose a weakly-supervised scene text spotting method via the location labels of single points with the corresponding text transcriptions. Due to the weak location annotations for challenging scene texts, previous weakly-supervised methods adopting the convolution neural network structure make it hard to model the different-scale text feature representations under blurring or nosing scenarios. In addition, as the single-point location can only cover part of the text instance, it will burden the confusion of sequential-like scene text recognition. To address these issues, we present a novel sequential recurrence self-attention for granularity-aware single-point scene text spotting. Specifically, we first enhance the scene text feature representations with different scales by integrating the global intra-interaction of high-level features with the low-level local features. Then, based on the granularity-aware text features, we decode them into text transcriptions in the sequential recurrence self-attention manner to capture the sequence-dependent relation in character-level semantics and locations. Extensive experiments show that our proposed method outperforms existing state-of-the-art weakly-supervised scene text spotters by a large margin. Xunquan Tong, Pengwen Dai, Xugong Qin, Rui Wang 0032, Wenqi Ren |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | S2CL-Leaf Net: Recognizing Leaf Images Like Human BotanistsabstractAutomatically classifying plant leaves is a challenging fine-grained classification task because of the diversity in leaf morphology, including size, texture, shape, and venation. Although powerful deep learning-based methods have achieved great improvement in leaf classification, these methods still require a large number of well-labeled samples for supervised training, which is difficult to get. In contrast, relying on the specific coarse-to-fine classification strategy, human botanists only require a small number of samples for accurate leaf recognition. Inspired by the classification strategy of human botanists, we propose a novel S 2 CL-Leaf Net , which exploits multi-granularity clues with a hierarchical attention mechanism and boosts the learning ability with the supervised sampling contrastive learning with limited training samples to classify plant leaves as human botanists do. Specifically, to fully explore and exploit the subtle details of the leaves, a novel sampling transformation mechanism is combined with the supervised contrastive learning to enhance the network’s perception of details by amplifying the discriminative regions with a weighted sampling of different regions. Furthermore, we construct the hierarchical attention mechanism to produce attention maps of different granularity, which helps to discover details in leaves that are important for classification. Experiments are conducted on the open-access leaf datasets, including Flavia, Swedish, and LeafSnap, which prove the effectiveness of the proposed S 2 CL-Leaf Net . Cong Zou, Rui Wang 0032, Cheng Jin 0001, Sanyi Zhang, Xin Wang 0019 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2023 | The Victim and The Beneficiary: Exploiting a Poisoned Model to Train a Clean Model on Poisoned DataabstractRecently, backdoor attacks have posed a serious security threat to the training process of deep neural networks (DNNs). The attacked model behaves normally on benign samples but outputs a specific result when the trigger is present. However, compared with the rocketing progress of backdoor attacks, existing defenses are difficult to deal with these threats effectively or require benign samples to work, which may be unavailable in real scenarios. In this paper, we find that the poisoned samples and benign samples can be distinguished with prediction entropy. This inspires us to propose a novel dual-network training framework: The Victim and The Beneficiary (V&B), which exploits a poisoned model to train a clean model without extra benign samples. Firstly, we sacrifice the Victim network to be a powerful poisoned sample detector by training on suspicious samples. Secondly, we train the Beneficiary network on the credible samples selected by the Victim to inhibit backdoor injection. Thirdly, a semi-supervised suppression strategy is adopted for erasing potential backdoors and improving model performance. Furthermore, to better inhibit missed poisoned samples, we propose a strong data augmentation method, AttentionMix, which works well with our proposed V&B framework. Extensive experiments on two widely used datasets against 6 state-of-the-art attacks demonstrate that our framework is effective in preventing backdoor injection and robust to various attacks while maintaining the performance on benign samples. Our code is available at https://github.com/Zixuan-Zhu/VaB. Zixuan Zhu 0002, Rui Wang 0032, Cong Zou, Lihua Jing |
ICCV | 2 |
| 2023 | Face Encryption via Frequency-Restricted Identity-Agnostic AttacksabstractBillions of people are sharing their daily live images on social media everyday. However, malicious collectors use deep face recognition systems to easily steal their biometric information (e.g., faces) from these images. Some studies are being conducted to generate encrypted face photos using adversarial attacks by introducing imperceptible perturbations to reduce face information leakage. However, existing studies need stronger black-box scenario feasibility and more natural visual appearances, which challenge the feasibility of privacy protection. To address these problems, we propose a frequency-restricted identity-agnostic (FRIA) framework to encrypt face images from unauthorized face recognition without access to personal information. As for the weak black-box scenario feasibility, we obverse that representations of the average feature in multiple face recognition models are similar, thus we propose to utilize the average feature via the crawled dataset from the Internet as the target to guide the generation, which is also agnostic to identities of unknown face recognition systems; in nature, the low-frequency perturbations are more visually perceptible by the human vision system. Inspired by this, we restrict the perturbation in the low-frequency facial regions by discrete cosine transform to achieve the visual naturalness guarantee. Extensive experiments on several face recognition models demonstrate that our FRIA outperforms other state-of-the-art methods in generating more natural encrypted faces while attaining high black-box attack success rates of 96%. In addition, we validate the efficacy of FRIA using real-world black-box commercial API, which reveals the potential of FRIA in practice. Our codes can be found in https://github.com/XinDong10/FRIA. Xin Dong 0015, Rui Wang 0032, Siyuan Liang 0004, Aishan Liu, Lihua Jing |
ACM Multimedia | 2 |
| 2023 | McGE '23: 1st International Workshop on Multimedia Content Generation and Evaluation: New Methods and PracticeabstractThe proposed workshop's topics, focusing on multimedia content generation, quality assessment, datasets and construction, is crucial due to its direct impact on the growth and success of the multimedia field. Multimedia content generation is essential for various applications, such as entertainment, advertising, and education. Quality assessment ensures the overall value and effectiveness of multimedia content, directly influencing user satisfaction and application success. Datasets are indispensable for training and evaluating multimedia algorithms, driving innovation, and fostering progress in the field. Finally, effective dataset construction methods set new benchmarks for the research community, stimulating innovation and unlocking new opportunities for leveraging multimedia data in various applications. The goal of this workshop is to bring together leading researchers in the field in a joint forum for advancing multimedia content generation and evaluation. Cheng Jin 0001, Liang He 0001, Mingli Song, Rui Wang 0032 |
ACM Multimedia | 4 |
| 2023 | Transferring CLIP's Knowledge into Zero-Shot Point Cloud Semantic SegmentationabstractTraditional 3D segmentation methods can only recognize a fixed range of classes that appear in the training set, which limits their application in real-world scenarios due to the lack of generalization ability. Large-scale visual-language pre-trained models, such as CLIP, have shown their generalization ability in the zero-shot 2D vision tasks, but are still unable to be applied to 3D semantic segmentation directly. In this work, we focus on zero-shot point cloud semantic segmentation and propose a simple yet effective baseline to transfer the visual-linguistic knowledge implied in CLIP to point cloud encoder at both feature and output levels. Both feature-level and output-level alignments are conducted between 2D and 3D encoders for effective knowledge transfer. Concretely, a Multi-granularity Cross-modal Feature Alignment (MCFA) module is proposed to align 2D and 3D features from global semantic and local position perspectives for feature-level alignment. For the output level, per-pixel pseudo labels of unseen classes are extracted using the pre-trained CLIP model as supervision for the 3D segmentation model to mimic the behavior of the CLIP image encoder. Extensive experiments are conducted on two popular benchmarks of point cloud segmentation. Our method outperforms significantly previous state-of-the-art methods under zero-shot setting (+29.2% mIoU on SemanticKITTI and 31.8% mIoU on nuScenes), and further achieves promising results in the annotation-free point cloud semantic segmentation setting, showing its great potential for label-efficient learning. Shaofei Huang 0001, Yulu Gao, Zhen Wang 0003, Rui Wang 0032, Kehua Sheng, Bo Zhang 0069, Si Liu 0001 |
ACM Multimedia | 5 |
| 2023 | Consistency-aware Feature Learning for Hierarchical Fine-grained Visual ClassificationabstractHierarchical Fine-Grained Visual Classification (HFGVC) assigns a label sequence (e.g., ["Albatross'', "Laysan Albatross'']) with a coarse to fine hierarchy to each object. It remains challenging to achieve high accuracy and consistency due to the small inter-class difference, large intra-class variance, and difficulty in modeling relationships among classification tasks at different granularities. In this paper, we propose an effective Consistency-Aware Feature Learning (CAFL) method for HFGVC to improve prediction consistency and classification accuracy simultaneously. Our key idea is to encode the prediction consistency constraint into a weak supervision mechanism via forward deduction and backward induction over the label hierarchy. Furthermore, we develop a disentanglement and bidirectional reinforcement classification head to extract the features for the classifiers at different granularities. Together with the stop-gradient policy and attention mechanism, they enable each classifier to exploit the features from the ones at other granularities without suffering from their conflicting gradients in training. We evaluate our method on several commonly-used fine-grained public datasets, including CUB-200-2011, FGVC-Aircraft, and Stanford Cars. The results show that our method not only achieves state-of-the-art classification accuracy but also effectively reduces inconsistency errors by 50% under the hierarchical fine-grained classification setting. Rui Wang 0032, Cong Zou, Zixuan Zhu 0002, Lihua Jing |
ACM Multimedia | 1 |
| 2023 | Exploring the Robustness of Human Parsers Toward Common CorruptionsabstractHuman parsing aims to segment each pixel of the human image with fine-grained semantic categories. However, current human parsers trained with clean data are easily confused by numerous image corruptions such as blur and noise. To improve the robustness of human parsers, in this paper, we construct three corruption robustness benchmarks, termed LIP-C, ATR-C, and Pascal-Person-Part-C, to assist us in evaluating the risk tolerance of human parsing models. Inspired by the data augmentation strategy, we propose a novel heterogeneous augmentation-enhanced mechanism to bolster robustness under commonly corrupted conditions. Specifically, two types of data augmentations from different views, i.e., image-aware augmentation and model-aware image-to-image transformation, are integrated in a sequential manner for adapting to unforeseen image corruptions. The image-aware augmentation can enrich the high diversity of training images with the help of common image operations. The model-aware augmentation strategy that improves the diversity of input data by considering the model's randomness. The proposed method is model-agnostic, and it can plug and play into arbitrary state-of-the-art human parsing frameworks. The experimental results show that the proposed method demonstrates good universality which can improve the robustness of the human parsing models and even the semantic segmentation models when facing various image common corruptions. Meanwhile, it can still obtain approximate performance on clean data. Sanyi Zhang, Xiaochun Cao, Rui Wang 0032, Guo-Jun Qi, Jie Zhou 0001 |
IEEE Trans. Image Process. | 3 |
| 2022 | PMP-NET: Rethinking Visual Context for Scene Graph GenerationabstractScene graph generation aims to describe the contents in scenes by identifying the objects and their relationships. In previous works, visual context is widely utilized in message passing networks to generate the representations for classification. However, the noisy estimation of visual context limits model performance. In this paper, we revisit the concept of incorporating visual context via a randomly ordered bidirectional Long Short Temporal Memory (biLSTM) based baseline, and show that noisy estimation is worse than random. To alleviate the problem, we propose a new method, dubbed Progressive Message Passing Network (PMP-Net) that better estimates the visual context in a coarse to fine manner. Specifically, we first estimate the visual context with a random initiated scene graph, then refine it with multi-head attention. The experimental results on the benchmark dataset Visual Genome show that PMP-Net achieves better or comparable performance on all three tasks: scene graph generation (SGGen), scene graph classification (SGCls), and predicate classification (PredCls). Xuezhi Tong, Rui Wang 0032, Chuan Wang 0002, Sanyi Zhang, Xiaochun Cao |
ICASSP | 2 |
| 2021 | Updated Paired Regions for Shadow Detection from Single Image
Xiao Wang 0017, Siyuan Yao, Pengwen Dai, Rui Wang 0032, Xiaochun Cao |
BMVC | 4 |
| 2021 | Coarse-to-Fine Visual Place Recognition
Junkun Qi, Rui Wang 0032, Chuan Wang 0002, Xiaochun Cao |
ICONIP (4) | 2 |
| 2021 | Few-Shot Classification with Multi-task Self-supervised Learning
Rui Wang 0032, Sanyi Zhang, Xiaochun Cao |
ICONIP (4) | 2 |
| 2021 | Semantic Correspondence with Geometric Structure AnalysisabstractThis article studies the correspondence problem for semantically similar images, which is challenging due to the joint visual and geometric deformations. We introduce the Flip-aware Distance Ratio method (FDR) to solve this problem from the perspective of geometric structure analysis. First, a distance ratio constraint is introduced to enforce the geometric consistencies between images with large visual variations, whereas local geometric jitters are tolerated via a smoothness term. For challenging cases with symmetric structures, our proposed method exploits Curl to suppress the mismatches. Subsequently, image correspondence is formulated as a permutation problem, for which we propose a Gradient Guided Simulated Annealing (GGSA) algorithm to perform a robust discrete optimization. Experiments on simulated and real-world datasets, where both visual and geometric deformations are present, indicate that our method significantly improves the baselines for both visually and semantically similar images. Rui Wang 0032, Xiaochun Cao, Yuanfang Guo |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2020 | Fake Generated Painting Detection Via Frequency AnalysisabstractWith the development of deep neural networks, digital fake paintings can be generated by various style transfer algorithms. To detect the fake generated paintings, we analyze the fake generated and real paintings in Fourier frequency domain and observe statistical differences and artifacts. Based on our observations, we propose Fake Generated Painting Detection via Frequency Analysis (FGPD-FA) by extracting three types of features in frequency domain. Besides, we also propose a digital fake painting detection database for assessing the proposed method. Experimental results demonstrate the excellence of the proposed method in different testing conditions. Yuanfang Guo, Jinjie Wei, Rui Wang 0032, Yunhong Wang 0001 |
ICIP | 5 |
| 2019 | PCGAN: Partition-Controlled Human Image GenerationabstractHuman image generation is a very challenging task since it is affected by many factors. Many human image generation methods focus on generating human images conditioned on a given pose, while the generated backgrounds are often blurred. In this paper, we propose a novel Partition-Controlled GAN to generate human images according to target pose and background. Firstly, human poses in the given images are extracted, and foreground/background are partitioned for further use. Secondly, we extract and fuse appearance features, pose features and background features to generate the desired images. Experiments on Market-1501 and DeepFashion datasets show that our model not only generates realistic human images but also produce the human pose and background as we want. Extensive experiments on COCO and LIP datasets indicate the potential of our method. Rui Wang 0032, Xiaowei Tian, Cong Zou |
AAAI | 2 |
| 2019 | Semantic Correlations Loss: Improving Model Interpretability for Multi-class ClassificationabstractDespite that convolutional neural networks (CNNs) have recently demonstrated high-quality object classification, the trained models suffer from their extreme unexplainability. In this paper, we propose a general method, named as semantic correlation loss, for introducing common-sense knowledge to CNN architectures. In contrast to traditional cross-entropy loss which only considers the ground-truth class, we exploit to be aware of the accuracy of all classes. By adding this simple add-on, current multi-class classification models are able to improve on the ability of “making mistakes reasonably”. In addition, a slight performance gain is also achieved. Experimental results on CUB-200-2011, CIFAR -10 and 100 are provided to demonstrate the efficacy of our proposed method. Moreover, this novel loss is able to be applied in any setting as long as the labels of training data are included in the common sense knowledge base. Xuezhi Tong, Rui Wang 0032, Xiaochun Cao, Wenqi Ren |
IEEE BigData | 2 |
| 2019 | Weighted Focus-Attention Deep Network for Fine-grained Image ClassificationabstractFine-Grained Visual Classification (FGVC) is a challenging task, due to the small variation of visual representations from different categories. An effective solution is utilizing the bounding boxes centering the object parts to extract the discriminative representations. However, regular rectangles contains the background when the shape of the part is irregular, which may interfere with the classification. In this paper, we propose a weighted focus-attention deep network (FA-Net) to address the problem of background interference in fine-grained classification. In our FA-Net, a focus-attention module is proposed to identify the foreground region from the class activation map and remove the background. Two branches are employed to obtain the primary and secondary attention regions with focus-attention module, and a weighted layer is utilized to integrate the attention regions. Experiment results on three challenging fine-grained classification datasets (e.g., CUB-200-2011, Stanford Dogs and FGVC Aircraft) show that our FA-Net obtains state-of-the-art results and outperforms the other fine-grained algorithms. Cong Zou, Rui Wang 0032, Xiaochun Cao, Feixiao Lv |
IEEE BigData | 2 |
| 2018 | Audio Visual Attribute Discovery for Fine-Grained Object RecognitionabstractCurrent progresses on fine-grained recognition are mainly focus on learning the discriminative feature representation via introducing the visual supervisions e.g. part labels. However, it is time-consuming and needs the professional knowledge to obtain the accuracy annotations. Different from these existing methods based on the visual supervisions, in this paper, we introduce a novel feature named audio visual attributes via discovering the correlations between the visual and audio representations. Specifically, our unified framework is training with video-level category label, which consists of two important modules, the encoder module and the attribute discovery module, to encode the image and audio into vectors and learn the correlations between audio and images, respectively. On the encoder module, we present two types of feed forward convolutional neural network for the image and audio modalities. While an attention driven framework based on recurrent neural network is developed to generate the audio visual attribute representation. Thus, our proposed architecture can be implemented end-to-end in the step of inference. We exploit our models for the problem of fine-grained bird recognition on the CUB200-211 benchmark. The experimental results demonstrate that with the help of audio visual attribute, we achieve the superior or comparable performance to that of strongly supervised approaches on the bird recognition. Hua Zhang 0008, Xiaochun Cao, Rui Wang 0032 |
AAAI | 3 |
| 2018 | Running OS Kernel in Separate Domains: A New Architecture for Applications and OS Services QuarantineabstractContainer-based PaaS cloud is ease of use and cost-efficient, but vulnerable to attacks due to the weak isolation provided by the built-in containers. In this paper, we present a lightweight virtualization based kernel decomposition approach to securely isolate cloud tenants as well as the operating system (OS) services against various threats. Our design decouples existing OS kernels based on their functionality and isolates different kernel partitions in separate domains. The kernel partition that enables application execution is quarantined in an application domain, while other partitions that offer various services are isolated in separate service domains. The application owned by one tenant can run transparently in a dedicated application domain, with strong isolation to those owned by other tenants. Furthermore, the kernel partition approach effectively defeats the malware that requires support from different kernel services. We have implemented a prototype based on Linux kernel and Xen hypervisor. Our evaluation demonstrates that the proposed kernel decomposition approach can defeat various OS kernel-targeted attacks with minimal performance overhead. Weijuan Zhang, Xiaoqi Jia, Shengzhi Zhang, Rui Wang 0032, Peng Liu 0005 |
APSEC | 4 |
| 2018 | Focal Text: an Accurate Text Detection with Focal LossabstractText detection in natural scene images is an important and popular task in the computer vision community. Due to slanted characters and blurred images in natural environments, it is a challenging task under active research. In this paper, we propose a Focal Text Detection Network (FTDN), which could be trained well without abundant data. FTDN is able to segment text region and simultaneously regress text box at pixel-level. Specifically, combined with focal loss, our method can balance positive/negative and easy/hard samples to achieve better performance. Compared with previous methods, FTDN achieves better performance in terms of text detection accuracy in natural scene. It outperforms the state-of-the-art methods on the standard ICDAR 2015 dataset with 80.9% F-measure. Xiaowei Tian, Dao Wu, Rui Wang 0032, Xiaochun Cao |
ICIP | 3 |
| 2018 | Fake Colorized Image DetectionabstractImage forensics aims to detect the manipulation of digital images. Currently, splicing detection, copy-move detection, and image retouching detection are attracting significant attention from researchers. However, image editing techniques develop over time. An emerging image editing technique is colorization, in which grayscale images are colorized with realistic colors. Unfortunately, this technique may also be intentionally applied to certain images to confound object recognition algorithms. To the best of our knowledge, no forensic technique has yet been invented to identify whether an image is colorized. We observed that, compared with natural images, colorized images, which are generated by three state-of-the-art methods, possess statistical differences for the hue and saturation channels. Besides, we also observe statistical inconsistencies in the dark and bright channels, because the colorization process will inevitably affect the dark and bright channel values. Based on our observations, i.e., potential traces in the hue, saturation, dark, and bright channels, we propose two simple yet effective detection methods for fake colorized images: Histogram-based fake colorized image detection and feature encoding-based fake colorized image detection. Experimental results demonstrate that both proposed methods exhibit a decent performance against multiple state-of-the-art colorization approaches. Yuanfang Guo, Xiaochun Cao, Wei Zhang 0031, Rui Wang 0032 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2018 | Halftone Image Watermarking by Content Aware Double-Sided Embedding Error DiffusionabstractIn this paper, we carry out a performance analysis from a probabilistic perspective to introduce the error diffusion-based halftone visual watermarking (EDHVW) methods' expected performances and limitations. Then, we propose a new general EDHVW method, content aware double-sided embedding error diffusion (CaDEED), via considering the expected watermark decoding performance with specific content of the cover images and watermark, different noise tolerance abilities of various cover image content, and the different importance levels of every pixel (when being perceived) in the secret pattern (watermark). To demonstrate the effectiveness of CaDEED, we propose CaDEED with expectation constraint (CaDEED-EC) and CaDEED-noise visibility function (NVF) and importance factor (IF) (CaDEED-N&I). Specifically, we build CaDEED-EC by only considering the expected performances of specific cover images and watermark. By adopting the NVF and proposing the IF to assign weights to every embedding location and watermark pixel, respectively, we build the specific method CaDEED-N&I. In the experiments, we select the optimal parameters for NVF and IF via extensive experiments. In both the numerical and visual comparisons, the experimental results demonstrate the superiority of our proposed work. Yuanfang Guo, Oscar C. Au, Rui Wang 0032, Lu Fang 0001, Xiaochun Cao |
IEEE Trans. Image Process. | 3 |
| 2017 | Image Deblurring via Extreme Channels PriorabstractCamera motion introduces motion blur, affecting many computer vision tasks. Dark Channel Prior (DCP) helps the blind deblurring on scenes including natural, face, text, and low-illumination images. However, it has limitations and is less likely to support the kernel estimation while bright pixels dominate the input image. We observe that the bright pixels in the clear images are not likely to be bright after the blur process. Based on this observation, we first illustrate this phenomenon mathematically and define it as the Bright Channel Prior (BCP). Then, we propose a technique for deblurring such images which elevates the performance of existing motion deblurring algorithms. The proposed method takes advantage of both Bright and Dark Channel Prior. This joint prior is named as extreme channels prior and is crucial for achieving efficient restorations by leveraging both the bright and dark information. Extensive experimental results demonstrate that the proposed method is more robust and performs favorably against the state-of-the-art image deblurring methods on both synthesized and natural images. Yanyang Yan, Wenqi Ren, Yuanfang Guo, Rui Wang 0032, Xiaochun Cao |
CVPR | 4 |
| 2017 | Binarized Mode Seeking for Scalable Visual Pattern DiscoveryabstractThis paper studies visual pattern discovery in large-scale image collections via binarized mode seeking, where images can only be represented as binary codes for efficient storage and computation. We address this problem from the perspective of binary space mode seeking. First, a binary mean shift (bMS) is proposed to discover frequent patterns via mode seeking directly in binary space. The binomial-based kernel and binary constraint are introduced for binarized analysis. Second, we further extend bMS to a more general form, namely contrastive binary mean shift (cbMS), which maximizes the contrastive density in binary space, for finding informative patterns that are both frequent and discriminative for the dataset. With the binarized algorithm and optimization, our methods demonstrate significant computation (50×) and storage (32×) improvement compared to standard techniques operating in Euclidean space, while the performance does not largely degenerate. Furthermore, cbMS discovers more informative patterns by suppressing low discriminative modes. We evaluate our methods on both annotated ILSVRC (1M images) and un-annotated blind Flickr (10M images) datasets with million scale images, which demonstrates both the scalability and effectiveness of our algorithms for discovering frequent and informative patterns in large scale collection. Wei Zhang 0031, Xiaochun Cao, Rui Wang 0032, Yuanfang Guo, Zhineng Chen |
CVPR | 3 |
| 2017 | Deep Strip-Based Network with Cascade Learning for Scene Text LocalizationabstractScene text detection is currently a popular research topic in the computer vision community. However, it is a challenging task due to the variations of texts and clutter backgrounds. In this paper, we propose a novel framework for scene text localization. Based on the region proposal network, a Strip-based Text Detection Network (STDN) is developed with vertical anchor mechanism to predict the text/non-text strip-shaped proposals. Meanwhile, we incorporate the recurrent neural network layers in the proposed network to refine the predicted results. Specifically, hard example mining is performed to train the STDN with cascade learning, which has a remarkable improvement in precision. Besides, we exploit a clustering algorithm to generate anchor dimensions spontaneously without hand-picking, which is portable and time-saving. The text detection framework achieves the state-of-the-art performance on ICDAR2013 with 0.89 F-measure. Dao Wu, Rui Wang 0032, Pengwen Dai, Yueying Zhang, Xiaochun Cao |
ICDAR | 2 |
| 2017 | Contextual approach for identifying malicious Inter-Component privacy leaks in Android appsabstractInter-Component Communication (ICC) enables developers to create rich and innovative applications in Android platform. However, some privacy problems occur because of the interactions among multiple components. Since the flow of sensitive data across components may be legal or malicious, it is necessary to perform a precise ICC analysis to identify the malicious flow of sensitive data. In this paper, we propose a static taint analysis method, named IccChecker, to identify the malicious ICC-based privacy leaks in Android applications. IccChecker first tracks the potential flow of sensitive data across components and extracts the contextual factors which trigger the sensitive behavior. By leveraging the context information, our approach differentiates the malicious privacy leaks from the legal privacy information exchanges according to the proposed contextual policy. Moreover, we present a comprehensive assessment with benchmarks and real-world applications. Our evaluation results with benchmarks demonstrate that IccChecker improves the precision of ICC-based privacy leak detection. In the evaluation with real-world applications, our approach identifies 4 apps with ICC-based privacy leaks among 168 Google Play apps (2.3%) while 31 apps are identified from 49 malwares (63.3%). Daojuan Zhang, Yuanfang Guo, Dianjie Guo, Rui Wang 0032, Guangming Yu |
ISCC | 4 |
| 2017 | LEAF: Latent Extended Attribute Features Discovery for Visual ClassificationabstractTo improve the discrimination of attribute representation, in this paper, we propose to extend the traditional attribute representations via embedding the latent high-order structure between attributes. Specifically, our aim is to construct the Latent Extended Attribute Features (LEAF) for visual classification. Since there only exist weak label for each attribute, we firstly propose a feature selection method to explore the common feature structures across categories. After that, the attribute classifiers are trained based on the selected features. Then, the category specific graph is introduced, which is composed of single attributes and their co-occurrence attribute pairs. This attribute graph is used as the initialized representation of each image. Considering our aim, we should discover the discriminative latent structure between attributes and train the robust category classifiers. To that end, we develop a joint learning objective function which is composed of the high-order representation mining term and the classifier training term. The mining term can both preserve category-specific information and discover the common structure between categories. Based on the discovery representation, the robust visual classifiers could be trained by the classifier term. Finally, an alternating optimization method is designed to seek the optimal solution of our objective function. Experimental results on the challenging datasets demonstrate the advantages of our proposed model over existing work. Hua Zhang 0008, Rui Wang 0032, Changqing Zhang 0002, Xiaochun Cao |
ACM Multimedia | 2 |
| 2016 | SketchNet: Sketch Classification with Web ImagesabstractIn this study, we present a weakly supervised approach that discovers the discriminative structures of sketch images, given pairs of sketch images and web images. In contrast to traditional approaches that use global appearance features or relay on keypoint features, our aim is to automatically learn the shared latent structures that exist between sketch images and real images, even when there are significant appearance differences across its relevant real images. To accomplish this, we propose a deep convolutional neural network, named SketchNet. We firstly develop a triplet composed of sketch, positive and negative real image as the input of our neural network. To discover the coherent visual structures between the sketch and its positive pairs, we introduce the softmax as the loss function. Then a ranking mechanism is introduced to make the positive pairs obtain a higher score comparing over negative ones to achieve robust representation. Finally, we formalize above-mentioned constrains into the unified objective function, and create an ensemble feature representation to describe the sketch images. Experiments on the TUBerlin sketch benchmark demonstrate the effectiveness of our model and show that deep feature representation brings substantial improvements over other state-of-the-art methods on sketch classification. Hua Zhang 0008, Si Liu 0001, Changqing Zhang 0002, Wenqi Ren, Rui Wang 0032, Xiaochun Cao |
CVPR | 5 |
| 2016 | IacDroid: Preventing Inter-App Communication capability leaks in AndroidabstractInter-App Communication (IAC) plays an important role in Android platform to share data and services among applications. However, the existence of IAC capability leaks could lead to the unauthorized privileged operations. In this paper, we first investigate the usage of IAC in Android applications to show the prevalence of IAC in the Android development model. To mitigate the threat caused by IAC capability leaks in Android, we develop a real-time monitoring and control system, called IacDroid, which distinguishes and prevents the IAC capability leak accurately in both third-party and in-rom applications at runtime. IacDroid extends the Binder IPC mechanism and the system service to construct context-based component call chains between multiple applications. By leveraging the call chains, the permission system is extended to detect and prevent the IAC capability leaks. IacDroid also presents an intuitive client-side solution to help users control the IAC capability leaks. We implement the prototype in Android 4.3, and present a comprehensive assessment with 500 Google Play applications and 36 malicious applications. The experimental results demonstrate that IacDroid can effectively prevent the IAC capability leaks with a negligible performance overhead. Daojuan Zhang, Rui Wang 0032, Zimin Lin, Dianjie Guo, Xiaochun Cao |
ISCC | 2 |
| 2016 | An Adaptive Reversible Data Hiding Scheme for JPEG Images
Jiaxin Yin, Rui Wang 0032, Yuanfang Guo, Feng Liu 0001 |
IWDW | 2 |
| 2016 | MatchDR: Image Correspondence by Leveraging Distance Ratio ConstraintabstractImage correspondence is to establish the connections between coherent images, which can be quite challenging due to the visual and geometric deformations. This paper proposes a robust image correspondence technique from the perspective of spatial regularity. Specifically, the visual deformation is addressed by introducing the spatial information by enforcing the distance ratio constrain. At the same time, the geometric deformation is tolerated by adopting a smoothness term. Subsequently, image correspondence is formulated as permutation problem, for which, we propose a Gradient Guided Simulated Annealing method for robust optimization. Furthermore, our method is much more memory efficient, where the storage complexity is reduced from O(n4) to O(n2). The experiments on several datasets indicate that our proposed formulation and optimization significantly improve the baselines for both visually-similar and semantically-similar images, where both visual and geometric deformations are present. Rui Wang 0032, Wei Zhang 0031, Xiaochun Cao |
ACM Multimedia | 1 |
| 2015 | Multi-cue Augmented Face ClusteringabstractFace clustering is an important but challenging task since facial images always have huge variation due to change in facial expressions, head poses and partial occlusions, etc. Moreover, face clustering is actually an unsupervised problem which makes it more difficult to reach an accurate result. Fortunately, there are some cues that can be used to improve clustering performance. In this paper, two types of cues are employed. The first one is pairwise constraints: must-link and cannot-link constraints, which can be extracted from the temporal and spatial knowledge of data. The other is that each face is associated with a series of attributes (i.e, gender) which can contribute discrimination among faces. To take advantage of the above cues, we propose a new algorithm, Multi-cue Augmented Face Clustering (McAFC), which effectively incorporates the cues via graph-guided sparse subspace clustering technique. Specially, facial images from the same individual are encouraged to be connected while faces from different persons are restrained to be connected. Experiments on three face datasets from real-world videos show the improvements of our algorithm over the state-of-the-art methods. Chengju Zhou, Changqing Zhang 0002, Huazhu Fu, Rui Wang 0032, Xiaochun Cao |
ACM Multimedia | 4 |
| 2013 | Defending return-oriented programming based on virtualization techniquesabstractABSTRACT Over the past few years, return‐oriented programming (ROP) has drawn great attention of both academia and industry. Because of its Turing completeness, ROP reuses short instruction sequences already present in the victim program's address space to perform arbitrary computation. Hence, it can successfully bypass state‐of‐the‐art code integrity check mechanisms. In this paper, we look into using virtualization technologies to defeat return‐oriented programming. We design and implement HyperCropII, a virtualization‐based automatic runtime approach to defend such attacks. ROP attackers extract short instruction sequences ending in ret called “gadgets” and craft stack content to “chain” these gadgets together. We observe that a key characteristic of ROP is to fill the stack with plenty of addresses that are within the range of the program's libraries. Accordingly, we inspect the content of the stack to see if a potential ROP attack exists and quarantine the damages for further security purposes. We have implemented a proof‐of‐concept system based on the open source Xen hypervisor. The evaluation results exhibit that our solution is effective and efficient. Copyright © 2013 John Wiley & Sons, Ltd. Xiaoqi Jia, Rui Wang 0032, Shengzhi Zhang, Peng Liu 0005 |
Secur. Commun. Networks | 2 |
| 2010 | DepSim: A Dependency-Based Malware Similarity Comparison System
Yi Yang 0040, Lingyun Ying, Rui Wang 0032, Purui Su, Dengguo Feng |
Inscrypt | 3 |