VLDB 2026 Research / reviewers in the wild / expert
Jizhong Han
dblp:72/6837
· DBLP profile ↗
154ranked-venue papers
0as first author
80since 2021 · last 2026
0000-0003-1107-3873ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 62 · 34 since 2021Graphics, computer vision, multimedia, augmented reality and games · 55 · 38 since 2021Systems, architecture and hardware · 27 · 4 since 2021Databases, data management, data science and information retrieval · 12 · 3 since 2021Human-computer interaction and ubiquitous computing · 10 · 10 since 2021Computer networks · 8 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 since 2021Security and privacy · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Exploiting Synergistic Cognitive Biases to Bypass Safety in LLMsabstractLarge Language Models (LLMs) demonstrate impressive capabilities across diverse tasks, yet their safety mechanisms remain susceptible to adversarial exploitation of cognitive biases---systematic deviations from rational judgment. Unlike prior studies focusing on isolated biases, this work highlights the overlooked power of multi-bias interactions in undermining LLM safeguards. Specifically, we propose CognitiveAttack, a novel red-teaming framework that adaptively selects optimal ensembles from 154 human social psychology-defined cognitive biases, engineering them into adversarial prompts to effectively compromise LLM safety mechanisms. Experimental results reveal systemic vulnerabilities across 30 mainstream LLMs, particularly open-source variants. CognitiveAttack achieves a substantially higher attack success rate than the SOTA black-box method PAP (60.1% vs. 31.6%), exposing critical limitations in current defenses. Through quantitative analysis of successful jailbreaks, we further identify vulnerability patterns in safety-aligned LLMs under synergistic cognitive biases, validating multi-bias interactions as a potent yet underexplored attack vector. This work introduces a novel interdisciplinary perspective by bridging cognitive science and LLM safety, paving the way for more robust and human-aligned AI systems. Xikang Yang, Biyu Zhou, Xuehai Tang, Jizhong Han, Songlin Hu 0001 |
AAAI | 4 |
| 2026 | More Thinking, Less Talking: Internalizing Deliberative Safety into LLM ParametersabstractPrevailing safety alignment methods still leave Large Language Models (LLMs) vulnerable to sophisticated jailbreak attacks.To bolster defenses, explicit reasoning mechanisms like Safety-oriented Chain-of-Thought (SCoT) have emerged, significantly enhancing robustness.However, this transparency introduces a critical trade-off: the exposed reasoning process itself becomes a new attack surface, risking the leakage of harmful information and revealing the model's safety logic to adversaries.This paper directly confronts this dilemma, asking: Can we achieve the full benefits of deliberative safety without the costs of explicit reasoning generation?We propose Safety Reasoning Internalization to make the deliberative process in SCoT "available but not visible".This approach is grounded in a key theoretical insight: the corrective influence of an SCoT can be effectively approximated by a targeted, low-rank update to the model's Feed-Forward Network (FFN) layers.We operationalize this through Hierarchical Internalization of Adversarially-Guided Reasoning (HIAR), a layer-wise safety alignment framework that internalizes safety reasoning into an implicit computational pathway using Low-Rank Adaptation (LoRA).HIAR enables the model to reach a safe conclusion within a single forward pass, entirely eliminating the need to generate vulnerable SCoT text.Extensive experiments on various LLMs demonstrate that HIAR achieves a 43% lower Attack Success Rate (ASR) against distinct jailbreak attacks compared to strong baselines. Xuehai Tang, Biyu Zhou, Jizhong Han, Songlin Hu 0001 |
ACL (1) | 4 |
| 2026 | Resolving the Security-Auditability Dilemma with Auditable Latent Chain-of-Thought AlignmentabstractTo address the increasingly severe safety risk of large language models (LLMs), reasoningbased safety alignment methods have emerged.These methods overcome the limitations of 'shallow alignment' by exposing the model's Chain-of-Thought (CoT), enabling auditability of safety reasoning process through both training-phase supervision and post-generation verification.However, this transparency creates a critical vulnerability, a tension we define as the Security Auditability Dilemma: while explicit reasoning is a prerequisite for safety, its textual Auditable paradoxically transforms it into an optimization target for adaptive attackers and induces the model to unintentionally copy harmful content from its own reasoning context.To address this, we propose Auditable Latent CoT Alignment (ALCA), a framework that decouples internal reasoning from external output.ALCA shifts the safety deliberation process into a continuous latent space.This allows the safety reasoning process to guide the generation of harmless outputs, while eliminates the discrete textual surface that facilitates internal copying and adaptive attack.Yet, this process is not a black box.we introduce a restricted Self-Decoding mechanism that allows the model to reconstruct its latent reasoning into human-readable text for supervision under specific guidance.Extensive experiments show that ALCA achieves robustness alignment, reducing the success rate of adaptive jailbreak attacks by over 40% compared to strong baselines, while preserving performance.Our framework presents a path toward building LLMs that are both robustly secure and auditable. Biyu Zhou, Xuehai Tang, Jizhong Han, Songlin Hu 0001 |
ACL (1) | 4 |
| 2026 | A Cognitive Distribution and Behavior-Consistent Framework for Black-Box Attacks on Recommender Systems
Hongyue Zhang, Dongqin Liu, Honglei Lv, Jiao Dai, Jizhong Han |
DASFAA (1) | 9 |
| 2026 | FreeEdit: Mask-Free Reference-Based Image Editing With Multi-Modal InstructionabstractIntroducing user-specified visual concepts in image editing is highly practical as these concepts convey the user's intent more precisely than text-based descriptions. We propose FreeEdit, a novel approach for achieving such reference-based image editing, which can accurately reproduce the visual concept from the reference image based on user-friendly language instructions. Our approach leverages the multi-modal instruction encoder to encode language instructions to guide the editing process. This implicit way of locating the editing area eliminates the need for manual editing masks. To enhance the reconstruction of reference details, we introduce the Decoupled Residual Refer-Attention (DRRA) module. This module is designed to integrate fine-grained reference features extracted by a detail extractor into the image editing process in a residual way without interfering with the original self-attention. Given that existing datasets are unsuitable for reference-based image editing tasks, particularly due to the difficulty in constructing image triplets that include a reference image, we curate a high-quality dataset, FreeBench, using a newly developed twice-repainting scheme. FreeBench comprises the images before and after editing, detailed editing instructions, as well as a reference image that maintains the identity of the edited object, encompassing tasks such as object addition, replacement, and deletion. By conducting phased training on FreeBench followed by quality tuning, FreeEdit achieves high-quality zero-shot editing through convenient language instructions. We conduct extensive experiments to evaluate the effectiveness of FreeEdit across multiple task types, demonstrating its superiority over existing methods. Runze He, Linjiang Huang, Shaofei Huang 0001, Jialin Gao, Xiaoming Wei, Jiao Dai, Jizhong Han, Si Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2025 | Unleashing the Temporal-Spatial Reasoning Capacity of GPT for Training-Free Audio and Language Referenced Video Object SegmentationabstractIn this paper, we propose an Audio-Language-Referenced SAM 2 (AL-Ref-SAM 2) pipeline to explore the training-free paradigm for audio and language-referenced video object segmentation, namely AVS and RVOS tasks. The intuitive solution leverages GroundingDINO to identify the target object from a single frame and SAM 2 to segment the identified object throughout the video, which is less robust to spatiotemporal variations due to a lack of video context exploration. Thus, in our AL-Ref-SAM 2 pipeline, we propose a novel GPT-assisted Pivot Selection (GPT-PS) module to instruct GPT-4 to perform two-step temporal-spatial reasoning for sequentially selecting pivot frames and pivot boxes, thereby providing SAM 2 with a high-quality initial object prompt. Within GPT-PS, two task-specific Chain-of-Thought prompts are designed to unleash GPT’s temporal-spatial reasoning capacity by guiding GPT to make selections based on a comprehensive understanding of video and reference information. Furthermore, we propose a Language-Binded Reference Unification (LBRU) module to convert audio signals into language-formatted references, thereby unifying the formats of AVS and RVOS tasks in the same pipeline. Extensive experiments show that our training-free AL-Ref-SAM 2 pipeline achieves performances comparable to or even better than fully-supervised fine-tuning methods. Shaofei Huang 0001, Rui Ling, Tianrui Hui, Zongheng Tang, Xiaoming Wei, Jizhong Han, Si Liu 0001 |
AAAI | 7 |
| 2025 | DS-GCG: Enhancing LLM Jailbreaks with Token Suppression and Induction Dual-StrategyabstractIn intelligent collaborative systems, the role of Large Language Models (LLMs) is becoming increasingly significant, with their security and privacy being of paramount importance. Greedy Coordinate Gradient (GCG)-based adversarial approaches are a staple in red team testing for circumventing the security alignments of LLMs. Yet, these methods are challenged by issues such as convergence difficulties and pseudo-evasion, which can impede attack efficacy. Our research indicates that the high likelihood of rejection tokens appearing in the initial k positions of generated text is a major contributor to adversarial failures, and their suppression can significantly improve attack success rates. Building on these insights, we present DS-GCG, an innovative adversarial attack methodology that enhances GCG attack potency. It employs adjustable-position prefilling to quell refusal responses and incite harmful outputs, coupled with a bidirectional greedy gradient search to swiftly identify adversarial suffixes. DS-GCG's universal suffix approach not only mitigates refusals but also hastens convergence, offering an efficient and robust search strategy. Our experimental results on widely-used open-source LLMs, showcased on the AdvBench dataset, confirm the cutting-edge performance of DS-GCG. Xuehai Tang, Xikang Yang, Zhongjiang Yao, Jie Wen 0007, Jizhong Han, Songlin Hu 0001 |
CSCWD | 6 |
| 2025 | A Generative Approach for Alleviating Filter Bubbles in Collaborative FilteringabstractIn addressing the filter bubble phenomenon in recommender systems, existing research has integrated large language models (LLMs) but faced distribution discrepancies and challenges like hallucinations and over-reliance on text data. We propose a two-stage method called LEAD (LLM-Enhanced Augmentation of Data) to mitigate these issues. First, we leverage LLMs' to extract textual representations of user interests and item audiences as auxiliary information. This information acts as pseudo-labels to guide a Conditional Generative Adversarial Network (CGAN) in generating unexpected items that align with actual user data distributions. Our approach harnesses model knowledge while avoiding data distribution gaps. Second, we introduce a heterogeneous view alignment strategy to align the semantic space of LLMs with the representation space of collaborative signals, enhancing the quality of representations and reducing noise. Extensive experiments on three real-world datasets demonstrate that LEAD outperforms existing models in recommendation quality, particularly in diversity and distribution alignment, showcasing the potential of LLMs in evolving recommender systems. Our code is available at https://github.com/gusuccc/LEAD. Hongyue Zhang, Jizhong Han, Xiaodan Zhang 0004 |
CSCWD | 5 |
| 2025 | Evidence-Claim Relevance Decoupling for Multimodal Fact-CheckingabstractWith the proliferation of misinformation on social media, the need for fact-checking has increased, garnering the attention of researchers. The main task of fact-checking is to analyze the semantic relationship and logic between the evidence and the claim, so as to determine the veracity of the claim. The semantic relationship between the claim and the evidence can be strong or weak. However, previous studies ignores the disparities in the strength of semantic relationship between evidence and claims in the process of evidence modeling for fact-checking, neglects the role of weak semantically related evidence in verifying claims. Weakly semantically related evidence logically determines the veracity of the claim sometimes. This paper proposes an evidence-claim relevance decoupling (ECRD) framework for multimodal fact-checking, where the claim and the evidence could be separated by decoupling the strong and weak relevant semantic parts in multimodal dual-stream models. Specifically, we first project text and image modalities into two distinct spaces. One space focuses on tightly coupled semantics between the claim and evidence, reducing the gap. Another space explores the loosely coupled parts, analyzing the various meaning between claim and evidence to reveal potential underlying connections. Then, the fully connected visual and textual graphs are constructed for above two spaces. Finally, the responses obtained by interacting in above graphs are fed into the fusion layer. On FACTIFY dataset, which features diverse text-image relationships, and MultiFact dataset, experimental results show that the proposed ECRD framework could improve the performance of classification significantly compared to baseline models and produces a competitive performance. Jianliang Zeng, Fuqing Zhu, Jizhong Han, Songlin Hu 0001 |
CSCWD | 3 |
| 2025 | Gamma-Guard: Lightweight Residual Adapters for Robust Guardrails in Large Language ModelsabstractThis paper contains potentially offensive and harmful text.Large language models (LLMs) are widely deployed as zero-shot evaluators for answer grading, content moderation, and document ranking.Yet studies show that guard models (Guards)-LLMs fine-tuned for safety-remain vulnerable to "jailbreak" attacks, jeopardising downstream chatbots.We confirm this weakness on three public benchmarks (BeaverTails, XSTest, AdvBench) and trace it to representation shifts that arise in the embedding layer and cascade through the Transformer stack.To counteract the effect, we introduce Gamma-Guard: lightweight residual adapters inserted after the embeddings and at sparse intervals in the model.The adapters start with zero-scaled gates, so they retain the original behaviour; a brief adversarial finetuning phase then teaches them to denoise embeddings and refocus attention.With fewer than 0.1 % extra parameters and only a 2 % latency increase, Gamma-Guard lifts adversarial accuracy from ≤ 5 % to ≈ 95 % a 90 percentage-point gain while reducing cleandata accuracy by just 8 percentage points.Extensive ablations further show that robustness improvements persist across different layer placements and model sizes.To our knowledge, this is the first approach that directly augments large Guards with trainable adapters, providing a practical path toward safer large-scale LLM deployments. Lijia Lv, Yuanshu Zhao, Xuehai Tang, Jie Wen 0007, Jizhong Han, Songlin Hu 0001 |
EMNLP | 6 |
| 2025 | LyapLock: Bounded Knowledge Preservation in Sequential Large Language Model EditingabstractLarge Language Models often contain factually incorrect or outdated knowledge, giving rise to model editing methods for precise knowledge updates.However, current mainstream locate-then-edit approaches exhibit a progressive performance decline during sequential editing, due to inadequate mechanisms for long-term knowledge preservation.To tackle this, we model the sequential editing as a constrained stochastic programming.Given the challenges posed by the cumulative preservation error constraint and the gradually revealed editing tasks, LyapLock is proposed.It integrates queuing theory and Lyapunov optimization to decompose the long-term constrained programming into tractable stepwise subproblems for efficient solving.This is the first model editing framework with rigorous theoretical guarantees, achieving asymptotic optimal editing performance while meeting the constraints of long-term knowledge preservation.Experimental results show that our framework scales sequential editing capacity to over 10,000 edits while stabilizing general capabilities and boosting average editing efficacy by 11.89% over SOTA baselines.Furthermore, it can be leveraged to enhance the performance of baseline methods.Our code is released on https://github.com/caskcsg/LyapLock. Peng Wang 0028, Biyu Zhou, Xuehai Tang, Jizhong Han, Songlin Hu 0001 |
EMNLP | 4 |
| 2025 | AdaPPA: Adaptive Position Pre-Fill Jailbreak Attack Approach Targeting LLMsabstractJailbreak vulnerabilities in Large Language Models (LLMs) refer to methods that extract malicious content from the model by carefully crafting prompts or suffixes, which has garnered significant attention from the research community. However, traditional attack methods, which primarily focus on the semantic level, are easily detected by the model. These methods overlook the difference in the model’s alignment protection capabilities at different output stages. To address this issue, we propose an adaptive position pre-fill jailbreak attack approach for executing jailbreak attacks on LLMs. Our method leverages the model’s instruction-following capabilities to first output pre-filled safe content, then exploits its narrative-shifting abilities to generate harmful content. Extensive black-box experiments demonstrate our method can improve the attack success rate by 47% on the widely recognized secure model (Llama2) compared to existing approaches. Our code can be found at: https://github.com/Yummy416/AdaPPA. Lijia Lv, Weigang Zhang, Xuehai Tang, Jie Wen 0007, Feng Liu 0001, Jizhong Han, Songlin Hu 0001 |
ICASSP | 6 |
| 2025 | Resolution Attack: Exploiting Image Compression to Deceive Deep Neural NetworksabstractModel robustness is essential for ensuring the stability and reliability of machine learning systems. Despite extensive research on various aspects of model robustness, such as adversarial robustness and label noise robustness, the exploration of robustness towards different resolutions, remains less explored. To address this gap, we introduce a novel form of attack: the resolution attack. This attack aims to deceive both classifiers and human observers by generating images that exhibit different semantics across different resolutions. To implement the resolution attack, we propose an automated framework capable of generating dual-semantic images in a zero-shot manner. Specifically, we leverage large-scale diffusion models for their comprehensive ability to construct images and propose a staged denoising strategy to achieve a smoother transition across resolutions. Through the proposed framework, we conduct resolution attacks against various off-the-shelf classifiers. The experimental results exhibit high attack success rate, which not only validates the effectiveness of our proposed framework but also reveals the vulnerability of current classifiers towards different resolutions. Additionally, our framework, which incorporates features from two distinct objects, serves as a competitive tool for applications such as face swapping and facial camouflage. The code is available at https://github.com/ywj1/resolution-attack. Wangjia Yu, Xiaomeng Fu, Jizhong Han, Xiaodan Zhang 0004 |
ICLR | 4 |
| 2025 | ComSFI: A Community-Aware Approach to Early Rumor Propagation PredictionabstractInformation propagation prediction in social networks remains challenging due to the complex interactions between community structures. Existing models typically overlook multi-peaked cascade patterns that emerge from community-specific dynamics, especially during critical early propagation stages. To address these limitations, we propose ComSFI (Community-aware Susceptible-Forwarding-Immune), a dynamic prediction framework that captures both intra- and inter-community information diffusion. ComSFI makes three key contributions: (1) community-specific diffusion modeling through individualized SFI modules, (2) cross-community propagation modeling with a learnable threshold mechanism that identifies cascade initiation timing, and (3) integration of user interest profiles to estimate propagation willingness across community boundaries. Experiments on Weibo, Douban, and Memetracker datasets demonstrate that ComSFI outperforms state-of-the-art baselines by 3%-12% in cascade size prediction and temporal pattern accuracy, with particular effectiveness in early-stage rumor propagation prediction—a critical application for timely intervention. Our results establish ComSFI as a versatile framework for analyzing and predicting complex information diffusion patterns in real-world networks. Wei Zhou 0019, Ziang Hu, Jizhong Han, Tao Guo 0006 |
ICTAI | 4 |
| 2025 | Modality-Agnostic Deepfakes DetectionabstractAs AI-generated content (AIGC) thrives, deepfakes have expanded from single-modality falsification to cross-modal fake content creation, where either audio or visual components can be manipulated.While using two unimodal detectors can detect audio-visual deepfakes, cross-modal forgery clues could be overlooked.Existing multimodal deepfake detectors typically establish correspondence between the audio and visual modalities for binary real/fake classification and require the co-occurrence of both modalities.However, in real-world multi-modal applications, missing modality scenarios may occur where either modality is unavailable.In such cases, audio-visual detection methods are less practical than two independent unimodal methods.Consequently, the detector can not always obtain the number or type of manipulated modalities beforehand, necessitating a fake-modality-agnostic audio-visual detector.In this work, we introduce a comprehensive framework that is agnostic to fake modalities, which facilitates the identification of multimodal deepfakes and handles situations with missing modalities, regardless of the manipulations embedded in audio, video, or even cross-modal forms.To enhance the modeling of cross-modal forgery clues, we employ audio-visual speech recognition (AVSR) Jin Liu 0020, Jiao Dai, Xi Wang 0014, Shan Jia, Siwei Lyu, Jizhong Han |
IH&MMSec | 9 |
| 2025 | OMS: One More Step Noise Searching to Enhance Membership Inference Attacks for Diffusion ModelsabstractThe data-intensive nature of Diffusion models amplifies the risks of privacy infringements and copyright disputes, particularly when training on extensive unauthorized data scraped from the Internet. Membership Inference Attacks (MIA) aim to determine whether a data sample has been utilized by the target model during training, thereby serving as a pivotal tool for privacy preservation. Current MIA employs the prediction loss to distinguish between training member samples and non-members. These methods assume that, compared to non-members, members, having been encountered by the model during training result in a smaller prediction loss. However, this assumption proves ineffective in diffusion models due to the random noise sampled during the training process. Rather than estimating the loss, our approach examines this random noise and reformulate the MIA as a noise search problem, assuming that members are more feasible to find the noise used in the training process. We formulate this noise search process as an optimization problem and employ the fixed-point iteration to solve it. We analyze current MIA methods through the lens of the noise search framework and reveal that they rely on the first residual as the discriminative metric to differentiate members and non-members. Inspired by this observation, we introduce OMS, which augments existing MIA methods by iterating One More fixed-point Step to include a further residual, i.e., the second residual. We integrate our method into various MIA methods across different diffusion models. The experimental results validate the efficacy of our proposed approach. Xiaomeng Fu, Xi Wang 0014, Jin Liu 0020, Jiao Dai, Jizhong Han, Xingyu Gao 0001 |
IJCAI | 6 |
| 2025 | Visual Perception Uncertainty Learning for Hallucination Detection in Large Vision-Language ModelsabstractHallucination remains a significant challenge which constrains the development of large vision-language models (LVLMs). Therefore, reliable hallucination detection has become a critical step in LVLMs evaluation and real-world deployment. Many previous studies have explored hallucination detection in LVLMs, with uncertainty-based approaches being widely adopted due to the independence from external tools and relatively low resource consumption. However, we observe that uncertainty does not always completely correlate with hallucination. Therefore, uncertainty-based methods may fail in certain cases, such as instances exhibiting high uncertainty but non-hallucination. To address this issue, we propose a framework called Visual Perception Uncertainty Learning (VisPUL) for hallucination detection in LVLMs. Specifically, VisPUL integrates visual information into uncertainty learning directly, allowing to capture uncertainty and visual-text consistency simultaneously. VisPUL improves the insufficiency of uncertainty methods that rely only on text output, providing enhanced generalizability and reliability. Extensive experiments conducted on the M-HalDetect and POPE datasets, covering both open-ended and yes-or-no tasks. Experimental results demonstrate that VisPUL significantly outperforms several strong baseline methods across different LVLMs. Fuqing Zhu, Jizhong Han, Songlin Hu 0001 |
ACM Multimedia | 3 |
| 2025 | Multi-Directional Dynamical Framework for Information Propagation PredictionabstractAccurately predicting information propagation in social networks is crucial for applications such as social media analysis and rumor monitoring. Dynamical approaches, valued for their interpretability and ability to capture the underlying mechanisms of propagation, have become a leading research direction. However, most existing methods focus exclusively on forward prediction, overlooking valuable information from other temporal directions. Consequently, these approaches face two critical limitations: (1) inadequacy in handling missing values prevalent in real-world scenarios with sparse and irregular observations; and (2) susceptibility to error accumulation that significantly degrades long-term prediction accuracy. To address these limitations, we propose the Multi-Directional Dynamical Framework (MDDF), which integrates forward, backward, and interpolation dynamics to fully leverage observations from all temporal directions. MDDF formulates the propagation process as a multi-boundary value problem, enabling observations at any time point to constrain predictions across the entire timeline. We further introduce rigorous mathematical formulations for backward and interpolation dynamics in stochastic propagation processes, a unified framework to reconcile multi-directional predictions, and an adaptive ensemble mechanism that dynamically weights each component based on data characteristics. Extensive experiments on real-world datasets demonstrate that MDDF not only surpasses existing methods in prediction accuracy, but also exhibits exceptional robustness to missing data and long-term forecasting. Notably, MDDF achieves up to 27.4% lower MAE than the best baseline under high data sparsity, and reduces long-term prediction errors by up to 25% compared to forward-only approaches. Ziang Hu, Guang Wang 0001, Jizhong Han, Tao Guo 0006 |
TrustCom | 5 |
| 2025 | Anchor3DLane++: 3D Lane Detection via Sample-Adaptive Sparse 3D Anchor RegressionabstractIn this paper, we focus on the challenging task of monocular 3D lane detection. Previous methods typically adopt inverse perspective mapping (IPM) to transform the Front-Viewed (FV) images or features into the Bird-Eye-Viewed (BEV) space for lane detection. However, IPM's dependence on flat ground assumption and context information loss in BEV representations lead to inaccurate 3D information estimation. Though efforts have been made to bypass BEV and directly predict 3D lanes from FV representations, their performances still fall behind BEV-based methods due to a lack of structured modeling of 3D lanes. In this paper, we propose a novel BEV-free method named Anchor3DLane++ which defines 3D lane anchors as structural representations and makes predictions directly from FV features. We also design a Prototype-based Adaptive Anchor Generation (PAAG) module to generate sample-adaptive sparse 3D anchors dynamically. In addition, an Equal-Width (EW) loss is developed to leverage the parallel property of lanes for regularization. Furthermore, camera-LiDAR fusion is also explored based on Anchor3DLane++ to leverage complementary information. Extensive experiments on three popular 3D lane detection benchmarks show that our Anchor3DLane++ outperforms previous state-of-the-art methods. Code is available at: https://github.com/tusen-ai/Anchor3DLane. Shaofei Huang 0001, Zhenwei Shen, Zehao Huang, Yue Liao, Jizhong Han, Naiyan Wang, Si Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | Unlocking Generative Priors: A New Membership Inference Framework for Diffusion ModelsabstractDiffusion models pose risks of privacy breaches and copyright disputes, primarily stemming from the potential utilization of unauthorized data during the training phase. Membership inference is aimed to determine whether a specific sample has been used in the training process of a target model, representing a critical tool for privacy violation verification. However, the increased model complexity and stochasticity inherent in diffusion renders traditional shadow-model-based or metric-based methods ineffective when applied to diffusion models. Moreover, existing methods only yield binary classification labels which lack necessary comprehensibility in practical applications. In this paper, we explore a novel perspective for membership inference by leveraging the intrinsic generative priors within the diffusion model. Compared with unseen samples, training samples exhibit stronger generative priors within the diffusion model, enabling the successful reconstruction of substantially degraded training images. Consequently, we propose the Degrade Restore Compare (DRC) framework. In this framework, an image undergoes sequential degradation and restoration, and its membership is determined by comparing it with the restored counterpart. Experimental results verify that our approach not only significantly outperforms existing methods in terms of accuracy but also provides comprehensible decision criteria, offering evidence for potential privacy violations. Xiaomeng Fu, Xi Wang 0014, Jin Liu 0020, Jiao Dai, Jizhong Han, Xingyu Gao 0001 |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2024 | Uncertainty-Guided Modal Rebalance for Hateful Memes DetectionabstractHateful memes detection is a challenging multimodal understanding task that requires comprehensive learning of vision, language, and cross-modal interactions.Previous research has focused on developing effective fusion strategies for integrating hate information from different modalities.However, these methods excessively rely on cross-modal fusion features, ignoring the modality uncertainty caused by the contribution degree of each modality to hate sentiment and the modality imbalance caused by the dominant modality suppressing the optimization of another modality.To this end, this paper proposes an Uncertainty-guided Modal Rebalance (UMR) framework for hateful memes detection.The uncertainty of each meme is explicitly formulated by designing stochastic representation drawn from a Gaussian distribution for aggregating cross-modal features with unimodal features adaptively.The modality imbalance is alleviated by improving cosine loss from the perspectives of intermodal feature and weight vectors constraints.In this way, the suppressed unimodal representation ability in multimodal models would be unleashed, while the learning of modality contribution would be further promoted.Extensive experimental results demonstrate that the proposed UMR produces the state-of-the-art performance on four widely-used datasets. Chuanpeng Yang, Fuqing Zhu, Jizhong Han, Songlin Hu 0001 |
ACL (1) | 4 |
| 2024 | Uncertainty-Aware Cross-Modal Alignment for Hate Speech DetectionabstractHate speech detection has become an urgent task with the emergence of huge multimodal harmful content (, memes) on social media platforms. Previous studies mainly focus on complex feature extraction and fusion to learn discriminative information from memes. However, these methods ignore two key points: 1) the misalignment of image and text in memes caused by the modality gap, and 2) the uncertainty between modalities caused by the contribution degree of each modality to hate sentiment. To this end, this paper proposes an uncertainty-aware cross-modal alignment (UCA) framework for modeling the misalignment and uncertainty in multimodal hate speech detection. Specifically, we first utilize the cross-modal feature encoder to capture image and text feature representations in memes. Then, a cross-modal alignment module is applied to reduce semantic gaps between modalities by aligning the feature representations. Next, a cross-modal fusion module is designed to learn semantic interactions between modalities to capture cross-modal correlations, providing complementary features for memes. Finally, a cross-modal uncertainty learning module is proposed, which evaluates the divergence between unimodal feature distributions to to balance unimodal and cross-modal fusion features. Extensive experiments on five publicly available datasets show that the proposed UCA produces a competitive performance compared with the existing multimodal hate speech detection methods. Chuanpeng Yang, Fuqing Zhu, Jizhong Han, Songlin Hu 0001 |
LREC/COLING | 4 |
| 2024 | Generative Transferable Universal Adversarial Perturbation for Combating DeepfakesabstractRecently, Deepfake has posed a significant threat to our digital society. This technology allows for the modification of facial identity, expression, and attributes in facial images and videos. The misuse of Deepfake can invade personal privacy, damage individuals’ reputations, and have serious consequences. To counter this threat, researchers have proposed active defense methods using adversarial perturbation to distort Deepfake products which can hinder the dissemination of false information. However, the existing methods are primarily based on image-specific approaches, which are inefficient for large-scale data. To address these issues, we propose an end-to-end approach to generate universal perturbations for combating Deepfake. To further cope with diverse Deepfakes, we introduce an adaptive balancing strategy to combat multiple models simultaneously. Specifically, for different scenarios, we propose two types of universal perturbations. Disrupting Universal Perturbation (DUP) leads Deepfake models to generate distorted outputs. In contrast, Lapsing Universal Perturbation (LUP) tries to make the output consistent with the original image, allowing the correct information to continue propagating. Experiments demonstrate the effectiveness and better generalization of our proposed perturbation compared with state-of-the-art methods. Consequently, our proposed method offers a powerful and efficient solution for combating Deepfake, which can help preserve personal privacy and prevent reputational damage. Xi Wang 0014, Xiaomeng Fu, Jin Liu 0020, Zhaoxing Li, Yesheng Chai, Jizhong Han |
CSCWD | 7 |
| 2024 | Explainable Deepfake Detection with Human PromptsabstractFacial manipulation techniques pose a significant threat to society due to the prevalence of deepfake content on the internet. While previous efforts have focused on developing accurate deepfake detection models, these models may be limited in real-world scenarios due to the lack of confidence that human analysts have in their results. Therefore, this study presents a novel approach to improve the practicality of deepfake detection models by incorporating human understanding. We propose a human prompt based deepfake detection framework that overlays Human-enhanced artifacts attention onto image artifact attention, which utilizes vision prompts to improve the model’s responsiveness and feedback ability while preserving its precision and generalizability. The deepfake detection model achieves an AUC score of 0.99 on the FaceForensics++ dataset and exhibits graceful generalization when evaluated on the Celeb-DF dataset. Furthermore, the model generates "possible area of manipulation" that provides an intuitive signal to facilitate interpretation of the detection process, bridging the gap between machine and human perception of "fake". Our proposed approach can potentially mitigate the harm caused by deepfakes and provide a more reliable solution for real-world applications. Xiaorong Ma, Zhaoxing Li, Yesheng Chai, Liangjun Zang, Jizhong Han |
CSCWD | 6 |
| 2024 | Towards More Effective and Transferable Poisoning Attacks against Link Prediction on GraphsabstractWith the impressive performances achieved by graph representation learning models on tasks such as link prediction, their vulnerability to imperceptible adversarial perturbations has also come to light. However, adversarial attacks struggle to balance effectiveness and transferability, particularly when attempting to deliver effective attacks on various target models in one shot and achieving effective outcomes in both availability and integrity attack settings. This study explores a novel way to mitigate attack performance across various models under availability and integrity attack settings. To fulfill these objectives, we develop a Scoring & Update (SU) framework for performing adversarial attacks against link prediction on graphs. Specifically, we iteratively score perturbations through a parameter-frozen and transferable surrogate model and then update the perturbation with scoring feedback to learn a more effective and transferable adversarial perturbation. Extensive experiments on two real-world datasets show that our attack model is more effective than six attack baselines against six popular target models for graph representation under both availability and integrity attack settings. Code is available at https://github.com/anonymousaccept/STAA. Baojie Tian, Liangjun Zang, Haichao Fu, Jizhong Han, Songlin Hu 0001 |
CSCWD | 4 |
| 2024 | Capture Long-Range Dependency with Meta-Path Transformer for De-Anonymization of Q&A SitesabstractThe expeditious advancement of social question-and-answer (Q&A) platforms has led to the valuable yet challenging practice of anonymous knowledge sharing. Despite implementing various anonymity techniques, the persistent threat of potential privacy breaches remains a paramount concern. To tackle this issue, we introduce the task of de-anonymization within Q&A communities and provide a bilingual dataset (Chinese and English) for research. In this paper, we propose a novel de-anonymization framework called MPT, effectively improving the model’s ability to capture long-range dependencies between nodes by integrating graph neural networks(GNNs) and language models(LMs). Specifically, we use GNN to extract structural features, and then we encode and fuse node representations from multiple meta-paths using Transformer and attention mechanisms. Extensive experiments on Zhihu and Quora data sets show that our model significantly outperforms the baseline model. In addition, our model possesses a degree of interpretability, enabling a comprehensive comprehension of the underlying factors contributing to user privacy breaches and facilitating the implementation of appropriate safeguards. The dataset1and code2utilized in this study have been made publicly accessible. Baojie Tian, Liangjun Zang, Jizhong Han, Songlin Hu 0001 |
CSCWD | 3 |
| 2024 | Integrating Open-domain Knowledge via Large Language Model for Multimodal Fake News DetectionabstractMultimodal fake news propagation on social media has emerged as a primary concern for both the public and the government. Distinguishing fake news from genuine content is challenging due to their close resemblance. Moreover, fake news often presents information that contradicts real-world facts, underscoring the necessity of incorporating comprehensive open-domain knowledge. However, current knowledge-based detection methods struggle to improve two key areas in integrating open-domain knowledge for fake news detection: i) query acquisition for retrieval and ii) utilization of retrieved knowledge. The challenges stem from limitations in understanding the semantics and reasoning about the relationship between open-domain knowledge and the content of the news. With the emergence of the Large Language Model (LLM), Natural Language Processing (NLP) has witnessed a revolution where impressive semantic understanding and reasoning abilities are significantly improved. This paper proposes a novel Open-domain Knowledge Integrated (OKI) framework for multimodal fake news detection, featuring two LLM-based agents that collaboratively leverage open-domain knowledge. One agent generates appropriate queries to retrieve knowledge, while the other filters out irrelevant retrieved knowledge. Experimental results demonstrate a significant performance improvement of OKI over established baselines on the Weibo and Twitter datasets. Anbin Xie, Fuqing Zhu, Jizhong Han, Songlin Hu 0001 |
CSCWD | 3 |
| 2024 | Customize your NeRF: Adaptive Source Driven 3D Scene Editing via Local-Global Iterative TrainingabstractIn this paper, we target the adaptive source driven 3D scene editing task by proposing a CustomNeRF model that unifies a text description or a reference image as the editing prompt. However, obtaining desired editing results conformed with the editing prompt is nontrivial since there exist two significant challenges, including accurate editing of only foreground regions and multi-view consistency given a single-view reference image. To tackle the first challenge, we propose a Local-Global Iterative Editing (LGIE) training scheme that alternates between foreground region editing and full-image editing, aimed at foreground-only manipulation while preserving the background. For the second challenge, we also design a class-guided regularization that exploits class priors within the generation model to alleviate the inconsistency problem among different views in image-driven editing. Extensive experiments show that our CustomNeRF produces precise editing results under various real scenes for both text- and image-driven settings. The code is available at: https://github.com/hrz2000/CustomNeRF. Runze He, Shaofei Huang 0001, Xuecheng Nie, Tianrui Hui, Luoqi Liu, Jiao Dai, Jizhong Han, Guanbin Li, Si Liu 0001 |
CVPR | 7 |
| 2024 | Real Appearance Modeling for More General Deepfake Detection
Cai Yu, Xi Wang 0014, Zihao Xiao 0002, Jiao Dai, Jizhong Han, Yesheng Chai |
ECCV (52) | 7 |
| 2024 | Quartet: A Holistic Hybrid Parallel Framework for Training Large Language Models
Weigang Zhang, Biyu Zhou, Xing Wu 0002, Chaochen Gao, Xuehai Tang, Ruixuan Li 0001, Jizhong Han, Songlin Hu 0001 |
Euro-Par (2) | 8 |
| 2024 | Generative Universal Nullifying Perturbation for Countering Deepfakes Through Combined Unsupervised Feature Aggregation
Xi Wang 0014, Xiaomeng Fu, Jin Liu 0020, Zhaoxing Li, Jizhong Han |
ICANN (2) | 6 |
| 2024 | ConfR: Conflict Resolving for Generalizable Deepfake DetectionabstractDeepfake detectors often encounter performance degradation when tested on unseen forgery methods. Existing literature tries to capture common features among multiple source forgery domains. However, we show that conflict arises in the shared feature space when each domain expresses domain-specific bias. If left unresolved, this conflict might mislead the model to learn domain-specific features and lead to inferior generalization. In this paper, we propose a new learning approach, Conflict Resolving (ConfR), designed to minimize conflict and learn features that generalize across forgeries. ConfR incorporates two key elements: the Intra-Domain Consistency Preserving (ICP) loss ensures updating consistency within forgery types, and the Inter-Domain Conflict Resolving (ICR) Module resolves updating conflicts between different forgery types. Extensive experiments demonstrate that ConfR significantly improves upon the state-of-the-art method, highlighting its potential for more generalizable deepfake detection. Cai Yu, Xi Wang 0014, Zhaoxing Li, Yesheng Chai, Jiao Dai, Jizhong Han |
ICME | 8 |
| 2024 | Reference Prompted Model Adaptation for Referring Camouflaged Object DetectionabstractThe goal of referring camouflaged object detection is to identify and segment the specified object hidden in the surroundings given text or images as references. The previous method still faces limitations in learning discriminative object features and comprehensively exploiting reference information due to coarse reference-image fusion upon disunified network components. In this paper, we propose a novel Reference Prompted Model Adaptation (RPMA) pipeline that employs rich and fine-grained semantic knowledge in a generic segmentation network to enhance the Ref-COD model’s capability. Within RPMA, we design a Cross Reference Adapter (CRA) to integrate reference information into the generic segmentation network to prompt reference-relevant camouflaged image features, and also devise a Reference-guided Dynamic Convolution (RDC) for foreground-background segmentation via reference-generated kernels. Extensive experiments on the Ref-COD benchmark show that our method achieves new state-of-the-art performance. Xuewei Liu, Shaofei Huang 0001, Ruipu Wu, Hengyuan Zhao, Xiaoming Wei, Jizhong Han, Si Liu 0001 |
ICME | 7 |
| 2024 | HIDD: Human-perception-centric Incremental Deepfake DetectionabstractFacial manipulation techniques pose a significant societal threat due to the widespread dissemination of deepfake content on the internet. Existing efforts for deepfake detection exhibit inadequate generalization performance when encountering unseen or degraded samples. We attribute this limitation to the overfitting of minor forgery patterns and variations in data distribution among disparate datasets. To tackle this issue, we introduce an innovative human-perception-centric incremental deepfake detection framework to enhance the generalization capabilities of deepfake detection models through continuous learning from a limited set of new samples. Firstly, the model leverages human perceptual salience to discern and comprehend significant artifacts, thereby mitigating overfitting to minor features. Subsequently, in the incremental learning process, we utilize multi-perspective knowledge distillation and a replay strategy to maintain the performance of the old model and minimize the feature distance between old and new samples. This comprehensive approach mitigates feature-level overfitting and addresses distribution differences among various datasets in the incremental phase. We conducted thorough experiments on four benchmark datasets (FF++, DFDC-P, CDF2, and DFD), and the experimental results demonstrate the superior performance of our method. Xiaorong Ma, Yesheng Chai, Zhaoxing Li, Jiao Dai, Liangjun Zang, Jizhong Han |
ICME | 8 |
| 2024 | Explicit Correlation Learning for Generalizable Cross-Modal Deepfake DetectionabstractWith the rising prevalence of deepfakes, there is a growing interest in developing generalizable detection methods for various types of deepfakes. While effective in their specific modalities, traditional detection methods fall short in addressing the generalizability of detection across diverse cross-modal deepfakes. This paper aims to explicitly learn potential cross-modal correlation to enhance deepfake detection towards various generation scenarios. Our approach introduces a correlation distillation task, which models the inherent cross-modal correlation based on content information. This strategy helps to prevent the model from overfitting merely to audio-visual synchronization. Additionally, we present the Cross-Modal Deepfake Dataset (CMDFD), a comprehensive dataset with four generation methods to evaluate the detection of diverse cross-modal deepfakes. The experimental results on CMDFD and FakeAVCeleb datasets demonstrate the superior generalizability of our method over existing state-of-the-art methods. Our code and data can be found at https://github.com/ljj898/CMDFD-Dataset-and-Deepfake-Detection. Cai Yu, Shan Jia, Xiaomeng Fu, Jin Liu 0020, Jiao Dai, Xi Wang 0014, Siwei Lyu, Jizhong Han |
ICME | 9 |
| 2024 | CSFI for Social Media: Understanding and Predicting Cross-Community Information PropagationabstractSocial platform users are intricately interconnected, forming a complex social system. Personalized recommendations help users access information from diverse sources, fostering community communication. Micro-level dynamic analysis technology offers a scientific approach to measuring and predicting information transmission. While fundamental dissemination principles are studied, cross-community communication scenarios are under-researched. This study presents a community-centric communication dynamics model (CSFI) to explain information transmission within and across communities on social media. Tested on Sina Weibo data, the model improves retweet prediction accuracy by 11.3 % over baselines. Accurately predicting information propagation on social media is crucial for public opinion analysis. This research aids in predicting dissemination trajectories, reflecting public sentiment, and guiding targeted information strategies or interventions. Wei Zhou 0019, Ziang Hu, Jizhong Han, Tao Guo 0006 |
ICTAI | 4 |
| 2024 | HDDA: Human-perception-centric Deepfake Detection AdapterabstractFacial manipulation techniques pose a significant societal threat due to the prevalent presence of deepfake content online. Current deepfake detection methods demonstrate subpar generalization performance when applied to unseen samples. The cause of this limitation lies in the overfitting of minor forgery patterns and variations in data distribution across different datasets. To tackle this issue, we introduce an innovative Human-perception-centric Deepfake Detection Adapter, namely HDDA, to enhance the generalization ability of deepfake detection models. This adaptation primarily involves two stages. During the pre-training stage, the model utilizes human perception salience to spot significant artifacts, thus reducing overfitting to minor features. In the subsequent fine-tuning stage, we introduce an efficient parameter tuning module named Deepfake Detection Adapter. The Adapter introduces two types of lightweight yet specialized adapter modules to the pre-trained model while keeping the backbone network frozen. It fine-tunes the pre-trained model through the adapter to adapt new and unseen datasets, thereby enhancing generalization. We conducted comprehensive experiments on various standard deepfake detection benchmarks to validate the effectiveness of our approach, particularly in showcasing a compelling advantage under cross-dataset and cross-manipulation settings. Xiaorong Ma, Yesheng Chai, Jiao Dai, Zhaoxing Li, Liangjun Zang, Jizhong Han |
IJCNN | 7 |
| 2024 | EthGAN: Improving Ethereum Account Classification Accuracy via Data AugmentationabstractRecently, with the prevalent adoption of blockchain in the financial system, there has been an increasing of anomaly activities such as ponzi schemes, gambling and phishing fraud on Ethereum platforms, and an effective account classification method is urgently required. The existing account classification methods on Ethereum with high accuracy require a learning system to be trained with balanced datasets. However, the distribution of annotated labels for account identities published on third-party sites is relatively imbalanced. Therefore, in this paper, We propose a EthGAN framework which includes a high-dimensional node feature representation module and a few-shot account data augment module to improve the accuracy and robustness at imbalanced datasets. The high-dimensional node feature representation module captures features from statistical, temporal, and transaction structure, and the few-shot account data augmentation module based on generative adversarial network models generate few-shot samples to improve the diversity and representativeness of the training datasets. We conduct extensive experiments to evaluate the performance of our proposed EthGAN framework on real-world Ethereum transaction data. The average classification effect of our method is 10+% higher than that of existing methods. Experimental results demonstrate that our method outperforms state-of-the-art methods in Ethereum account classification. Xuehai Tang, Zhongjiang Yao, Huazhen Zhong, Yuanshu Zhao, Xiaodan Zhang 0004, Jizhong Han |
IJCNN | 7 |
| 2024 | Reinforcement Learning-powered Effectiveness and Efficiency Few-shot Jailbreaking Attack LLMsabstractThe widespread use of large language models (LLMs) has brought about security risks, including biases, discrimination, and ethical concerns. Reinforcement Learning from Human Feedback (RLHF), as a method to improve model security, still faces challenges such as objective management and misaligned generalization, leading to the emergence of jailbreak attacks. Existing methods implement jailbreak attacks by optimizing adversarial prompts or leveraging the in-context learning capabilities of LLMs, but they are limited in terms of efficiency and scalability. This paper proposes a reinforcement learning-based few-shot example selection method to enhance the effectiveness and efficiency of these attacks. The proposed method extends the GPT-2 architecture with an example selection module and employs strategies such as experience replay and entropy penalty to accelerate convergence and avoid local optima. Experimental results demonstrate that, compared to existing methods, this approach achieves a 100% increase in attack success rate on Vicuna-7B and a 2.4-second reduction in the time cost per harmful instruction generation on GPT-3.5. Xuehai Tang, Zhongjiang Yao, Jie Wen 0007, Yangchen Dong, Jizhong Han, Songlin Hu 0001 |
ISPA | 5 |
| 2024 | Pyramidal Cross-Modal Transformer with Sustained Visual Guidance for Multi-Label Image ClassificationabstractMulti-label image classification poses a formidable challenge due to the presence of multiple objects in each image, rendering it notably complex to decipher the visual content comprehensively. Discriminating between multiple objects necessitates the establishment of robust visual label dependencies. Previous methods attempt to formulate cross-modal interaction or one-shot co-occurrence relationship guidance. However, it not only exhibits limitations when handling occluded or blurry objects but also fails to fully leverage the diverse hierarchical properties for sustainably guiding the learning process of label dependencies. To sustainably establish hierarchical visual label dependencies, this paper introduces a Pyramidal Cross-modal Transformer framework for MLIC tasks. Specifically, the pyramidal visual guidance layer parses the visual features into a multi-resolution pyramid structure, allowing the updated visual-related information to provide sustained guidance for label semantics. This surpasses the conventional pre-processing of co-occurrence relationships. Besides, the hybrid modal interaction layer is proposed to effectively mitigate the semantic disparities between visual and label information with modal-blended indiscriminate attention, replacing vanilla self-attention. Several combination blocks consisting of these two layers are integrated and embedded within the encoder-decoder structure to facilitate the exploration of meticulous visual label dependencies. Extensive experiments on two widely-used benchmarks, including MS-COCO and PASCAL VOC 2007, consistently demonstrate that PCMT could provide state-of-the-art results. Ruyun Wang, Fuqing Zhu, Jizhong Han, Songlin Hu 0001 |
ICMR | 4 |
| 2024 | Unveiling Structural Memorization: Structural Membership Inference Attack for Text-to-Image Diffusion ModelsabstractWith the rapid advancements of large-scale text-to-image diffusion models, various practical applications have emerged, bringing significant convenience to society. However, model developers may misuse the unauthorized data to train diffusion models. These data are at risk of being memorized by the models, thus potentially violating citizens' privacy rights. Therefore, in order to judge whether a specific image is utilized as a member of a model's training set, Membership Inference Attack (MIA) is proposed to serve as a tool for privacy protection. Current MIA methods predominantly utilize pixel-wise comparisons as distinguishing clues, considering the pixel-level memorization characteristic of diffusion models. However, it is practically impossible for text-to-image models to memorize all the pixel-level information in massive training sets. Therefore, we move to the more advanced structure-level memorization. Observations on the diffusion process show that the structures of members are better preserved compared to those of nonmembers, indicating that diffusion models possess the capability to remember the structures of member images from training sets. Drawing on these insights, we propose a simple yet effective MIA method tailored for text-to-image diffusion models. Extensive experimental results validate the efficacy of our approach. Compared to current pixel-level baselines, our approach not only achieves state-of-the-art performance but also demonstrates remarkable robustness against various distortions. Xiaomeng Fu, Xi Wang 0014, Jin Liu 0020, Xingyu Gao 0001, Jiao Dai, Jizhong Han |
ACM Multimedia | 7 |
| 2024 | Dynamic Prompting of Frozen Text-to-Image Diffusion Models for Panoptic Narrative GroundingabstractPanoptic narrative grounding (PNG), whose core target is fine-grained image-text alignment, requires a panoptic segmentation of referred objects given a narrative caption. Previous discriminative methods achieve only weak or coarse-grained alignment by panoptic segmentation pretraining or CLIP model adaptation. Given the recent progress of text-to-image Diffusion models, several works have shown their capability to achieve fine-grained image-text alignment through cross-attention maps and improved general segmentation performance. However, the direct use of phrase features as static prompts to apply frozen Diffusion models to the PNG task still suffers from a large task gap and insufficient vision-language interaction, yielding inferior performance. Therefore, we propose an Extractive-Injective Phrase Adapter (EIPA) bypass within the Diffusion UNet to dynamically update phrase prompts with image features and inject the multimodal cues back, which leverages the fine-grained image-text alignment capability of Diffusion models more sufficiently. In addition, we also design a Multi-Level Mutual Aggregation (MLMA) module to reciprocally fuse multi-level image and phrase features for segmentation refinement. Extensive experiments on the PNG benchmark show that our method achieves new state-of-the-art performance. Tianrui Hui, Jing Zhang 0017, Bin Ma 0028, Xiaoming Wei, Jizhong Han, Si Liu 0001 |
ACM Multimedia | 7 |
| 2024 | Dynamic Mixed-Prototype Model for Incremental Deepfake DetectionabstractThe rapid advancement of deepfake technology poses significant threats to social trust. Although recent deepfake detectors have exhibited promising results on deepfakes of the same type as those present in training, their effectiveness degrades significantly on novel deepfakes crafted by unseen algorithms due to the gap in forgery patterns. Some studies have enhanced detectors by adapting to the continuously emerging deepfakes through incremental learning. Despite the progress, they overlooked the scarcity of novel samples that can easily lead to insufficient learning of forgery patterns. To mitigate this issue, we introduce the Dynamic Mixed-Prototype (DMP) model, which dynamically increases prototypes to adapt to novel deepfakes efficiently. Specifically, the DMP model adopts multiple prototypes to represent both real and fake classes, enabling learning novel patterns by expanding prototypes and jointly retaining knowledge learned in previous prototypes. Furthermore, we propose the Prototype-Guided Replay strategy and Prototype Representation Distillation loss, both of which effectively prevent forgetting learned knowledge based on the prototypical representation of samples. Our method surpasses existing incremental deepfake detectors across four datasets and can generalize to novel deepfakes by learning limited deepfake samples. Cai Yu, Xi Wang 0014, Zihao Xiao 0002, Jizhong Han, Yesheng Chai |
ACM Multimedia | 6 |
| 2024 | Parquet-Based CTR Model Training in Production EnvironmentabstractCTR(click through rate) model has played an important role in modern recommendation systems. Most of the recommendation models in industrial scenario are trained by TensorFlow. However, we observed that, TFRecord, the native k-v data format in TensorFlow, is not the best choice for CTR training. Those keys take up to 54% of the storage space in TFRecord formatted training data. To overcome this, we introduce Apache Parquet, a column-oriented data format, into CTR tasks to improve spatial efficiency. Besides, to use in production environment, we further give some high performance implementations of Parquet training scheme. Firstly, GPU data preprocessing method is adopted in replace of original Spark based solution to generate Parquet training data and accelerate data preprocessing. Secondly, we modify data loader in TensorFlow to consume the Parquet training data with high efficiency. Experimental results show that, the size of preprocessed Criteo dataset is 95.29% smaller in comparison to TFRecord and the data preprocessing time also reduces 99.6%. Without any model performance damage, we speed up the training process by 1.45x. Our scheme has applied to our internal business and has obtained similar performance benefits. Jinrong Guo, Biyu Zhou, Xiaokun Zhu, Yongjun Bao, Jizhong Han, Songlin Hu 0001 |
SMC | 6 |
| 2024 | Modality adaptation via feature difference learning for depth human parsing
Shaofei Huang 0001, Tianrui Hui, Fengguang Peng, Yuqiang Fang, Bin Ma 0028, Xiaoming Wei, Jizhong Han |
Comput. Vis. Image Underst. | 9 |
| 2024 | OSM-Net: One-to-Many One-Shot Talking Head Generation With Spontaneous Head MotionsabstractOne-shot talking head generation has no explicit head movement reference, thus it is difficult to generate talking heads with head motions. Some existing works only edit the mouth area and generate still talking heads, leading to unreal talking head performance. Other works construct one-to-one mapping between audio signal and head motion sequences, introducing ambiguity correspondences into the mapping since people can behave differently in head motions when speaking the same content. This unreasonable mapping form fails to model the diversity and produces either nearly static or even exaggerated head motions, which are unnatural and strange. Therefore, the one-shot talking head generation task is actually a one-to-many ill-posed problem and people present diverse head motions when speaking. Based on the above observation, we propose OSM-Net, aone-to-manyone-shot talking head generation network with natural head motions. OSM-Net constructs a motion space that contains rich and various clip-level head motion features. Each basis of the space represents a feature of meaningful head motion in a clip rather than just a frame, thus providing more coherent and natural motion changes in talking heads. The driving audio is mapped into the motion space, around which various motion features can be sampled within a reasonable range to achieve the one-to-many mapping. Besides, the landmark constraint and time window feature input improve the accurate expression feature extraction and video generation. Extensive experiments show that OSM-Net generates more natural realistic head motions under reasonable one-to-many mapping paradigm compared with other methods. Jin Liu 0020, Xi Wang 0014, Xiaomeng Fu, Yesheng Chai, Cai Yu, Jiao Dai, Jizhong Han |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2024 | Learning to Discover Forgery Cues for Face Forgery DetectionabstractLocating manipulation maps,i.e., pixel-level annotation of forgery cues, is crucial for providing interpretable detection results in face forgery detection. Related learning objects have also been widely adopted as auxiliary tasks to improve the classification performance of detectors whereas they require comparisons between paired real and forged faces to obtain manipulation maps as supervision. This requirement restricts their applicability to unpaired faces and contradicts real-world scenarios. Moreover, the used comparison methods annotate all changed pixels, including noise introduced by compression and upsampling. Using such maps as supervision hinders the learning of exploitable cues and makes models prone to overfitting. To address these issues, we introduce a weakly supervised model in this paper, named Forgery Cue Discovery (FoCus), to locate forgery cues in unpaired faces. Unlike some detectors that claim to locate forged regions in attention maps, FoCus is designed to sidestep their shortcomings of capturing partial and inaccurate forgery cues. Specifically, we propose a classification attentive regions proposal module to locate forgery cues during classification and a complementary learning module to facilitate the learning of richer cues. The produced manipulation maps can serve as better supervision to enhance face forgery detectors. Visualization of the manipulation maps of the proposed FoCus exhibits superior interpretability and robustness compared to existing methods. Experiments on five datasets and four multi-task models demonstrate the effectiveness of FoCus in both in-dataset and cross-dataset evaluations. Cai Yu, Xiaomeng Fu, Xi Wang 0014, Jiao Dai, Jizhong Han |
IEEE Trans. Inf. Forensics Secur. | 7 |
| 2023 | Anchor3DLane: Learning to Regress 3D Anchors for Monocular 3D Lane DetectionabstractMonocular 3D lane detection is a challenging task due to its lack of depth information. A popular solution is to first transform the front-viewed (FV) images or features into the bird-eye-view (BEV) space with inverse perspective mapping (IPM) and detect lanes from BEV features. However, the reliance of IPM on flat ground assumption and loss of context information make it inaccurate to restore 3D information from BEV representations. An attempt has been made to get rid of BEV and predict 3D lanes from FV representations directly, while it still underperforms other BEV-based methods given its lack of structured representation for 3D lanes. In this paper, we define 3D lane anchors in the 3D space and propose a BEV-free method named Anchor3DLane to predict 3D lanes directly from FV representations. 3D lane anchors are projected to the FV features to extract their features which contain both good structural and context information to make accurate predictions. In addition, we also develop a global optimization method that makes use of the equal-width property between lanes to reduce the lateral error of predictions. Extensive experiments on three popular 3D lane detection benchmarks show that our Anchor3DLane outperforms previous BEV-based methods and achieves state-of-the-art performances. The code is available at: https://github.com/tusenai/Anchor3DLane. Shaofei Huang 0001, Zhenwei Shen, Zehao Huang, Jiao Dai, Jizhong Han, Naiyan Wang, Si Liu 0001 |
CVPR | 6 |
| 2023 | Bridging Search Region Interaction with Template for RGB-T TrackingabstractRGB-T tracking aims to leverage the mutual enhancement and complement ability of RGB and TIR modalities for improving the tracking process in various scenarios, where cross-modal interaction is the key component. Some previous methods concatenate the RGB and TIR search region features directly to perform a coarse interaction process with redundant background noises introduced. Many other methods sample candidate boxes from search frames and conduct various fusion approaches on isolated pairs of RGB and TIR boxes, which limits the cross-modal interaction within local regions and brings about inadequate context modeling. To alleviate these limitations, we propose a novel Template-Bridged Search region Interaction (TBSI) module which exploits templates as the medium to bridge the cross-modal interaction between RGB and TIR search regions by gathering and distributing target-relevant object and environment contexts. Original templates are also updated with enriched multimodal contexts from the template medium. Our TBSI module is inserted into a ViT backbone for joint feature extraction, search-template matching, and cross-modal interaction. Extensive experiments on three popular RGB-T tracking benchmarks demonstrate our method achieves new state-of-the-art performances. Code is available at https://github.com/RyanHTR/TBSI. Tianrui Hui, Zizheng Xun, Fengguang Peng, Junshi Huang, Xiaoming Wei, Xiaolin Wei, Jiao Dai, Jizhong Han, Si Liu 0001 |
CVPR | 8 |
| 2023 | A Multi-source Domain Adaption Approach to Minority Disk Failure Prediction
Wang Wang, Xuehai Tang, Biyu Zhou, Yangchen Dong, Yuanhang Feng, Jizhong Han, Songlin Hu 0001 |
ICA3PP (2) | 6 |
| 2023 | OPT: One-shot Pose-Controllable Talking Head GenerationabstractOne-shot talking head generation produces lip-sync talking heads based on arbitrary audio and one source face. To guarantee the naturalness and realness, recent methods propose to achieve free pose control instead of simply editing mouth areas. However, existing methods do not preserve accurate identity of source face when generating head motions. To solve the identity mismatch problem and achieve high-quality free pose control, we present One-shot Pose-controllable Talking head generation network (OPT). Specifically, the Audio Feature Disentanglement Module separates content features from audios, eliminating the influence of speaker-specific information contained in arbitrary driving audios. Later, the mouth expression feature is extracted from the content feature and source face, during which the landmark loss is designed to enhance the accuracy of facial structure and identity preserving quality. Finally, to achieve free pose control, controllable head pose features from reference videos are fed into the Video Generator along with the expression feature and source face to generate new talking heads. Extensive quantitative and qualitative experimental results verify that OPT generates high-quality pose-controllable talking heads with no identity mismatch problem, outperforming previous SOTA methods. Jin Liu 0020, Xi Wang 0014, Xiaomeng Fu, Yesheng Chai, Cai Yu, Jiao Dai, Jizhong Han |
ICASSP | 7 |
| 2023 | Large Pose Friendly Face Reenactment using subtle motionsabstractFace reenactment aims to synthesis a photo-realistic video of the source face by imitating the motion and expression of the driving video while keeping the source appearance (i.e. identity). Although good results have achieved recently, most state-of-the-art methods remain vulnerable to extreme conditions, which greatly restricts the application in the real world. Among various extreme conditions, the large pose problem is the most common one. We clarify that the large pose problem is mainly caused by the severe motion change between the source image and the current driving frame. An intuitive solution is to divide the severe motion change into a sequence of subtle motions. Therefore, we propose a new scheme that exploring the temporal coherence between previous neighbor frame and current frame. The smaller motion change between consecutive frames help to solve the large pose problem. Furthermore, a calibration net is designed to eliminate the error accumulation of the previous step. Extensive experiments demonstrate that our method performs better on large pose face reenactment than the state-of-the-art in terms of large pose cases and visual quality. Xiaomeng Fu, Xi Wang 0014, Jin Liu 0020, Jiao Dai, Jizhong Han |
ICME | 5 |
| 2023 | Semantic Stage-Wise Learning for Knowledge DistillationabstractKnowledge distillation enhances the performance of the student model by transferring knowledge from the teacher model. Moreover, the attention mechanism has been introduced recently to enable each layer of the student to learn knowledge from all teacher layers, which brings about considerable optimization. However, noted that features from different layers, such as shallow and deep layers, might have a big semantic gap, and compulsively aligning one student layer to all teacher layers would mislead the learning process. To tackle this problem, an effective framework called Semantic Stage-Wise learning for Knowledge Distillation (SSWKD) is presented in this paper. We divide all layers into shallow and deep stages, and only allow feature alignment within the same stage to alleviate semantic mismatch. In addition, with the observation that the performance of deep networks relies more on some key features rather than evenly on all of them, a crucial feature enhancement method based on KL divergence is then proposed for SSWKD, forcing the student to pay more attention to critical features of the teacher. Extensive experiments and visualizations show that our SSWKD outperforms other distillation methods on CIFAR-100 and COCO2017 datasets for image classification, object detection, and instance segmentation tasks. Dongqin Liu, Wei Zhou 0019, Zhaoxing Li, Jiao Dai, Jizhong Han, Ruixuan Li 0001, Songlin Hu 0001 |
ICME | 6 |
| 2023 | FONT: Flow-guided One-shot Talking Head Generation with Natural Head MotionsabstractOne-shot talking head generation has received growing attention in recent years, with various creative and practical applications. An ideal natural and vivid generated talking head video should contain natural head pose changes. However, it is challenging to map head pose sequences from driving audio since there exists a natural gap between audio-visual modalities. In this work, we propose a Flow-guided One-shot model that achieves NaTural head motions(FONT) over generated talking heads. Specifically, we design a probabilistic CVAE-based model to predict head pose sequences from driving audio and source face. Then we develop a keypoint predictor that produces unsupervised keypoints describing the facial structure information from the source face, driving audio and pose sequences. Finally, a flow- guided occlusion-aware generator is employed to produce photo-realistic talking head videos from the estimated keypoints and source face. Extensive experimental results prove that FONT generates talking heads with natural head poses and synchronized mouth shapes, outperforming other compared methods. Jin Liu 0020, Xi Wang 0014, Xiaomeng Fu, Yesheng Chai, Cai Yu, Jiao Dai, Jizhong Han |
ICME | 7 |
| 2023 | Discovering Sounding Objects by Audio Queries for Audio Visual SegmentationabstractAudio visual segmentation (AVS) aims to segment the sounding objects for each frame of a given video. To distinguish the sounding objects from silent ones, both audio-visual semantic correspondence and temporal interaction are required. The previous method applies multi-frame cross-modal attention to conduct pixel-level interactions between audio features and visual features of multiple frames simultaneously, which is both redundant and implicit. In this paper, we propose an Audio-Queried Transformer architecture, AQFormer, where we define a set of object queries conditioned on audio information and associate each of them to particular sounding objects. Explicit object-level semantic correspondence between audio and visual modalities is established by gathering object information from visual features with predefined audio queries. Besides, an Audio-Bridged Temporal Interaction module is proposed to exchange sounding object-relevant information among multiple frames with the bridge of audio features. Extensive experiments are conducted on two AVS benchmarks to show that our method achieves state-of-the-art performances, especially 7.1% M_J and 7.6% M_F gains on the MS3 setting. Shaofei Huang 0001, Hongji Zhu, Jiao Dai, Jizhong Han, Wenge Rong, Si Liu 0001 |
IJCAI | 6 |
| 2023 | Enriching Phrases with Coupled Pixel and Object Contexts for Panoptic Narrative GroundingabstractPanoptic narrative grounding (PNG) aims to segment things and stuff objects in an image described by noun phrases of a narrative caption. As a multimodal task, an essential aspect of PNG is the visual-linguistic interaction between image and caption. The previous two-stage method aggregates visual contexts from offline-generated mask proposals to phrase features, which tend to be noisy and fragmentary. The recent one-stage method aggregates only pixel contexts from image features to phrase features, which may incur semantic misalignment due to lacking object priors. To realize more comprehensive visual-linguistic interaction, we propose to enrich phrases with coupled pixel and object contexts by designing a Phrase-Pixel-Object Transformer Decoder (PPO-TD), where both fine-grained part details and coarse-grained entity clues are aggregated to phrase features. In addition, we also propose a Phrase-Object Contrastive Loss (POCL) to pull closer the matched phrase-object pairs and push away unmatched ones for aggregating more precise object contexts from more phrase-relevant object tokens. Extensive experiments on the PNG benchmark show our method achieves new state-of-the-art performance with large margins. Tianrui Hui, Junshi Huang, Xiaoming Wei, Xiaolin Wei, Jiao Dai, Jizhong Han, Si Liu 0001 |
IJCAI | 7 |
| 2023 | MFR-Net: Multi-faceted Responsive Listening Head Generation via Denoising Diffusion ModelabstractFace-to-face communication is a common scenario including roles of speakers and listeners. Most existing research methods focus on producing speaker videos, while the generation of listener heads remains largely overlooked. Responsive listening head generation is an important task that aims to model face-to-face communication scenarios by generating a listener head video given a speaker video and a listener head image. An ideal generated responsive listening video should respond to the speaker with attitude or viewpoint expressing while maintaining diversity in interaction patterns and accuracy in listener identity information. To achieve this goal, we propose the Multi-Faceted Responsive Listening Head Generation Network (MFR-Net). Specifically, MFR-Net employs the probabilistic denoising diffusion model to predict diverse head pose and expression features. In order to perform multi-faceted response to the speaker video, while maintaining accurate listener identity preservation, we design the Feature Aggregation Module to boost listener identity features and fuse them with other speaker-related features. Finally, a renderer finetuned with identity consistency loss produces the final listening head videos. Our extensive experiments demonstrate that MFR-Net not only achieves multi-faceted responses in diversity and speaker identity information but also in attitude and viewpoint expression. Jin Liu 0020, Xi Wang 0014, Xiaomeng Fu, Yesheng Chai, Cai Yu, Jiao Dai, Jizhong Han |
ACM Multimedia | 7 |
| 2023 | CoP: Chain-of-Pose for Image Animation in Large Pose ChangesabstractImage animation involves generating a video of a source image imitating the pose of a driving video. Despite recent advancements in the image animation task, most state-of-the-art methods remain vulnerable to large pose changes. In cases of large pose changes, existing methods struggle to model the complex nonlinear motion and yield distorted results, which greatly restricts their application in the real world. To tackle this problem, we present a novel approach called Chain-of-Pose (CoP) that decomposes large pose changes into a sequence of intermediate pose changes. This enables us to handle simplified pose changes and improves the accuracy of pose estimation. Furthermore, to better preserve the appearance of the source object, we introduce the Appearance Refinement Module (ARM) that effectively integrates the appearance texture feature of the source image with the structural pose feature from the pose chain. Our experimental results demonstrate that our method qualitatively and quantitatively outperforms state-of-the-art approaches on four diverse datasets, comprising talking faces, human bodies, and pixel animals. Notably, our approach significantly improves video quality in the case of large object pose changes. Our code is attached to the supplementary material. Xiaomeng Fu, Xi Wang 0014, Jin Liu 0020, Shuhui Wang, Jiao Dai, Jizhong Han |
ACM Multimedia | 6 |
| 2023 | Invariant Meets Specific: A Scalable Harmful Memes Detection FrameworkabstractHarmful memes detection is a challenging task in the field of multimodal information processing due to the semantic gap between different modalities. Current research on this task mainly focuses on multimodal dual-stream models. However, the existing works ignore the misalignment of the memes caused by the modality gap. Moreover, the cross-modal interaction in the dual-stream models is insufficient to identify harmful memes. To this end, this paper proposes a scalable invariant and specific modality (ISM) representations framework via graph neural networks. The proposed ISM framework provides a comprehensive and disentangled view for memes and promotes inter-modal interaction. Specifically, ISM projects each modality to two distinct spaces. The first space is modality-invariant, learning the corresponding commonalities and reducing the modality gap. The second space is modality-specific, holding the distinctive characteristics of each modality and complementing the common latent features captured in invariant spaces. Then, we construct fully connected visual and textual graphs for each space. The unimodal graphs are fused to dynamically balance inter-modal and intra-modal relationships, which are complementary to the dual-stream models. Finally, an adaptive module is designed to weigh the proportion of each fusion graph for memes. Moreover, the mainstream multimodal dual-stream models could be employed as the backbone flexibly. Extensive experiments on five publicly available datasets show that the proposed ISM provides a stable improvement over baselines and produces a competitive performance compared with the existing harmful memes detection methods. Chuanpeng Yang, Fuqing Zhu, Jizhong Han, Songlin Hu 0001 |
ACM Multimedia | 3 |
| 2023 | Language-Aware Spatial-Temporal Collaboration for Referring Video SegmentationabstractGiven a natural language referring expression, the goal of referring video segmentation task is to predict the segmentation mask of the referred object in the video. Previous methods only adopt 3D CNNs upon the video clip as a single encoder to extract a mixed spatio-temporal feature for the target frame. Though 3D convolutions are able to recognize which object is performing the described actions, they still introduce misaligned spatial information from adjacent frames, which inevitably confuses features of the target frame and leads to inaccurate segmentation. To tackle this issue, we propose a language-aware spatial-temporal collaboration framework that contains a 3D temporal encoder upon the video clip to recognize the described actions, and a 2D spatial encoder upon the target frame to provide undisturbed spatial features of the referred object. For multimodal features extraction, we propose a Cross-Modal Adaptive Modulation (CMAM) module and its improved version CMAM+ to conduct adaptive cross-modal interaction in the encoders with spatial- or temporal-relevant language features which are also updated progressively to enrich linguistic global context. In addition, we also propose a Language-Aware Semantic Propagation (LASP) module in the decoder to propagate semantic information from deep stages to the shallow stages with language-aware sampling and assignment, which is able to highlight language-compatible foreground visual features and suppress language-incompatible background visual features for better facilitating the spatial-temporal collaboration. Extensive experiments on four popular referring video segmentation benchmarks demonstrate the superiority of our method over the previous state-of-the-art methods. Tianrui Hui, Si Liu 0001, Shaofei Huang 0001, Guanbin Li, Wenguan Wang, Luoqi Liu, Jizhong Han |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2022 | ESimCSE: Enhanced Sample Building Method for Contrastive Learning of Unsupervised Sentence EmbeddingabstractContrastive learning has been attracting much attention for learning unsupervised sentence embeddings. The current state-of-the-art unsupervised method is the unsupervised SimCSE (unsup-SimCSE). Unsup-SimCSE takes dropout as a minimal data augmentation method, and passes the same input sentence to a pre-trained Transformer encoder (with dropout turned on) twice to obtain the two corresponding embeddings to build a positive pair. As the length information of a sentence will generally be encoded into the sentence embeddings due to the usage of position embedding in Transformer, each positive pair in unsup-SimCSE actually contains the same length information. And thus unsup-SimCSE trained with these positive pairs is probably biased, which would tend to consider that sentences of the same or similar length are more similar in semantics. Through statistical observations, we find that unsup-SimCSE does have such a problem. To alleviate it, we apply a simple repetition operation to modify the input sentence, and then pass the input sentence and its modified counterpart to the pre-trained Transformer encoder, respectively, to get the positive pair. Additionally, we draw inspiration from the community of computer vision and introduce a momentum contrast, enlarging the number of negative pairs without additional calculations. The proposed two modifications are applied on positive and negative pairs separately, and build a new sentence embedding method, termed Enhanced Unsup-SimCSE (ESimCSE). We evaluate the proposed ESimCSE on several benchmark datasets w.r.t the semantic text similarity (STS) task. Experimental results show that ESimCSE outperforms the state-of-the-art unsup-SimCSE by an average Spearman correlation of 2.02% on BERT-base. Xing Wu 0002, Chaochen Gao, Liangjun Zang, Jizhong Han, Zhongyuan Wang 0006, Songlin Hu 0001 |
COLING | 4 |
| 2022 | Smoothed Contrastive Learning for Unsupervised Sentence EmbeddingabstractUnsupervised contrastive sentence embedding models, e.g., unsupervised SimCSE, use the InfoNCE loss function in training. Theoretically, we expect to use larger batches to get more adequate comparisons among samples and avoid overfitting. However, increasing batch size leads to performance degradation when it exceeds a threshold, which is probably due to the introduction of false-negative pairs through statistical observation. To alleviate this problem, we introduce a simple smoothing strategy upon the InfoNCE loss function, termed Gaussian Smoothed InfoNCE (GS-InfoNCE). In other words, we add random Gaussian noise as an extension to the negative pairs without increasing the batch size. Through experiments on the semantic text similarity tasks, though simple, the proposed smoothing strategy brings improvements to unsupervised SimCSE. Xing Wu 0002, Chaochen Gao, Yipeng Su, Jizhong Han, Zhongyuan Wang 0006, Songlin Hu 0001 |
COLING | 4 |
| 2022 | PG-Pass: Targeted Online Password Guessing Model based on Pointer Generator NetworkabstractExisting targeted online password guessing models were based on Probabilistic Context-Free Grammars (PCFG), having inherent disadvantages that guessing structures were always the same for different users. This problem would lead to poor guessing efficiency. In order to solve this problem, we propose a targeted online password guessing model (PG-Pass) composed of the pointer generator network. It could automatically learn the impact of personal information on passwords and guess the target user’s password more accurately. Through extensive experiments, we obtain the optimal parameters. The results show that with only Personally Identifiable Information (PII), the guessing success rate of the PG-Pass model could reach 19.49% in guessing once, which is ten times higher than TarGuess-I. When guessing 100 times, the guessing success rate could be 41.07%, proving the effectiveness of the proposed model in this paper. Yong Li 0007, Ruixin Shi, Jizhong Han |
CSCWD | 5 |
| 2022 | Language-Bridged Spatial-Temporal Interaction for Referring Video Object SegmentationabstractReferring video object segmentation aims to predict foreground labels for objects referred by natural language expressions in videos. Previous methods either depend on 3D ConvNets or incorporate additional 2D ConvNets as encoders to extract mixed spatial-temporal features. However, these methods suffer from spatial misalignment or false distractors due to delayed and implicit spatial-temporal interaction occurring in the decoding phase. To tackle these limitations, we propose a Language-Bridged Duplex Transfer (LBDT) module which utilizes language as an intermediary bridge to accomplish explicit and adaptive spatial-temporal interaction earlier in the encoding phase. Concretely, cross-modal attention is performed among the temporal encoder, referring words and the spatial encoder to aggregate and transfer language-relevant motion and appearance information. In addition, we also propose a Bilateral Channel Activation (BCA) module in the decoding phase for further denoising and highlighting the spatial-temporal consistent features via channel-wise activation. Extensive experiments show our method achieves new state-of-the-art performances on four popular benchmarks with 6.8% and 6.9% absolute AP gains on A2D Sentences and J-HMDB Sentences respectively, while consuming around 7× less computational overhead11https://github.com/dzh19990407/LBDT. Tianrui Hui, Junshi Huang, Xiaoming Wei, Jizhong Han, Si Liu 0001 |
CVPR | 5 |
| 2022 | Contrastive Learning for Session-Based Recommendation
Wanhui Qian, Dongqin Liu, Yipeng Su, Jizhong Han, Ruixuan Li 0001 |
ICANN (4) | 6 |
| 2022 | Cross-Layer Aggregation with Transformers for Multi-Label Image ClassificationabstractMulti-label image classification task aims to predict multiple object labels in a given image and faces the challenge of variable-sized objects. Limited by the size of CNN convolution kernels, existing CNN-based methods have difficulty capturing global dependencies and effectively fusing multiple layers features, which is critical for this task. Recently, transformers have utilized multi-head attention to extract feature with long range dependencies. Inspired by this, this paper proposes a Cross-layer Aggregation with Transformers (CAT) framework, which leverages transformers to capture the long range dependencies of CNN-based features with Long Range Dependencies module and aggregate the features layer by layer with Cross-Layer Fusion module. To make the framework efficient, a multi-head pre-max attention is designed to reduce the computation cost when fusing the high-resolution features of lower-layers. On two widely-used benchmarks (i.e., VOC2007 and MS-COCO), CAT provides a stable improvement over the baseline and produces a competitive performance. Weibo Zhang, Fuqing Zhu, Jizhong Han, Tao Guo 0006, Songlin Hu 0001 |
ICASSP | 3 |
| 2022 | UFI: A Unified Feature Interaction Framework for Multi-Label Image ClassificationabstractMulti-label image classification (MLIC) is a more challenging task compared with single-label image classification due to multiple concepts targets, and complex visual relationships should be formulated. Convolutional Neural Network (CNN) and Visual Transformer (ViT) have shown superior performance in local and global feature representations, respectively. However, the interactions between local and global features are neglected in current works. To further formulate the critical interactions, this paper designs a Unified Feature Interaction (UFI) framework, aiming to integrate the selected local features with global features based on CNN and ViT, simultaneously. The proposed UFI includes two key modules: Class-Related Feature Selection (CRFS) and Feature Interaction Attention (FIA) modules. Specifically, according to the activation map, CRFS selects target regions by the preliminary calculation of predicted scores. FIA enables the significant local-global feature interaction based on the selected target regions and whole image. We initially attempted to interact with local and global features for multi-label image classification. UFI provides a stable improvement over the baseline and produces a new state-of-the-art result on MS-COCO and VOC2007. Weibo Zhang, Ziang Yang, Fuqing Zhu, Jizhong Han, Songlin Hu 0001 |
ICME | 5 |
| 2022 | Focus by Prior: Deepfake Detection Based on Prior-AttentionabstractNowadays advanced facial manipulation techniques produce deepfake videos more realistically, which makes deepfake detection more difficult. To capture subtle and intricate artifacts, recent works attempt to enhance low-level textural information by attention-based framework. However, these methods require complex simulated data or extra supervision. Highly dependent on training settings, these methods not only have high training costs but also are prone to overfitting. To address this issue, we propose a novel perspective of deepfake detection via so-called prior-attention. Specifically, we introduce prior textural information, such as edge and noise, to model the attention maps explicitly. Benefiting from these natural “attention maps”, our model significantly enhances discriminative information without additional supervision. Furthermore, we design a Feature Abstraction Block (FAB) to facilitate cross-layer features interaction and insert it into distinct layers of CNN to detect the inconsistencies at multiple spatial levels. Extensive experiments demonstrate that our method achieves performance comparable to state-of-the-art methods. Cai Yu, Jiao Dai, Xi Wang 0014, Weibo Zhang, Jin Liu 0020, Jizhong Han |
ICME | 7 |
| 2022 | MakeItSmile: Detail-Enhanced Smiling Face ReenactmentabstractGiven a target face and a driving face, face reenactment aims to transfer attributes from the driving face to the target face. In the last decade, a great number of methods have been proposed to generate realistic reenacted faces. However, when these methods are applied to generate a smiling face, most of them can only get a mouth with blurry teeth, making the reenacted face unrealistic. This problem is mainly caused by incomplete tooth structure in the target face image under the setting of one-shot reenactment. In order to obtain smiling reenacted faces with detailed tooth structure, our method uses the tooth information from the driving face rather than the target face. Furthermore, to better represent the tooth structure and expressions of the driving face, we extract the texture with a carefully designed geometry-aware encoder. By training the encoder with tooth segmentation task and non-identity classification task, we acquire refined tooth representations and meanwhile derive the non-identity part of the driving face. We also design a specific generator to fuse the tooth texture features into the target face. Moreover, we add a mouth loss function to further ensure the high definition of the smiling reenacted face. We compare our method to existing state-of-the-art approaches. The experiments show that our method gets comparable results on non-smiling face reenactment and has superior performance on smiling face reenactment. Xiaomeng Fu, Xi Wang 0014, Jin Liu 0020, Wantao Liu, Jiao Dai, Jizhong Han |
IJCNN | 6 |
| 2022 | Improving disk failure detection accuracy via data augmentationabstractFrequently happening of disk failures seriously affects the dependability and service quality of cloud data centers. Recently, machine learning (ML) based methods are popularly adopted to proactively predict forthcoming disk failures via supervised learning. However, the high imbalance of failure samples and healthy samples is a huge obstacle for existing detection methods to establish high performance detection model. This paper presents a data augmentation method MSGMD, which can efficiently generate high quality failure samples to alleviate the data imbalance of the training set, so as to effectively improve the performance of any supervised failure detection models. First, MSGMD converts failure samples (multivariate time series) into multiple univariate time series via decomposing the spatial relations among features. Then it learns the temporal correlation of each feature via a policy-based reinforcement learning model trained in an adversarial way. After that, it generates failure samples by combining feature series sampled from learned distribution. Finally, it filters out low quality generated samples with a confidence-based method. Experimental results on real-world datasets show that, through data augmentation, MSGMD can improve the FDR and F1-Score of the state-of-the-art disk failure detection model by 31.59% and 30.74% respectively on average. Wang Wang, Xuehai Tang, Biyu Zhou, Wenjie Xiao, Jizhong Han, Songlin Hu 0001 |
IWQoS | 5 |
| 2022 | Multimodal Hate Speech Detection via Cross-Domain Knowledge TransferabstractNowadays, the hate speech diffusion of texts and images in social network has become the mainstream compared with the diffusion of texts-only, raising the pressing needs of multimodal hate speech detection task. Current research on this task mainly focuses on the construction of multimodal models without considering the influence of the unbalanced and widely distributed samples for various attacks in hate speech. In this situation, introducing enhanced knowledge is necessary for understanding the attack category of hate speech comprehensively. Due to the high correlation between hate speech detection and sarcasm detection tasks, this paper makes an initial attempt of common knowledge transfer based on the above two tasks, where hate speech detection and sarcasm detection are defined as primary and auxiliary tasks, respectively. A scalable cross-domain knowledge transfer (CDKT) framework is proposed, where the mainstream vision-language transformer could be employed as backbone flexibly. Three modules are included, bridging the semantic, definition and domain gaps simultaneously between primary and auxiliary tasks. Specifically, semantic adaptation module formulates the irrelevant parts between image and text in primary and auxiliary tasks, and disentangles with the text representation to align the visual and word tokens. Definition adaptation module assigns different weights to the training samples of auxiliary task by measuring the correlation between samples of the auxiliary and primary task. Domain adaptation module minimizes the feature distribution gap of samples in two tasks. Extensive experiments show that the proposed CDKT provides a stable improvement compared with baselines and produces a competitive performance compared with some existing multimodal hate speech detection methods. Chuanpeng Yang, Fuqing Zhu, Guihua Liu, Jizhong Han, Songlin Hu 0001 |
ACM Multimedia | 4 |
| 2021 | An Adaptive Hybrid Framework for Cross-domain Aspect-based Sentiment AnalysisabstractCross-domain aspect-based sentiment analysis aims to utilize the useful knowledge in a source domain to extract aspect terms and predict their sentiment polarities in a target domain. Recently, methods based on adversarial training have been applied to this task and achieved promising results. In such methods, both the source and target data are utilized to learn domain-invariant features through deceiving a domain discriminator. However, the task classifier is only trained on the source data, which causes the aspect and sentiment information lying in the target data can not be exploited by the task classifier. In this paper, we propose an Adaptive Hybrid Framework (AHF) for cross-domain aspect-based sentiment analysis. We integrate pseudo-label based semi-supervised learning and adversarial training in a unified network. Thus the target data can be used not only to align the features via the training of domain discriminator, but also to refine the task classifier. Furthermore, we design an adaptive mean teacher as the semi-supervised part of our network, which can mitigate the effects of noisy pseudo labels generated on the target data. We conduct experiments on four public datasets and the experimental results show that our framework significantly outperforms the state-of-the-art methods. Fuqing Zhu, Pu Song, Jizhong Han, Tao Guo 0006, Songlin Hu 0001 |
AAAI | 4 |
| 2021 | Entity and Relation Matching Consensus for Entity AlignmentabstractEntity alignment aims to match synonymous entities across different knowledge graphs, which is a fundamental task for knowledge integration. Recently, researchers have devoted to leveraging rich information within relations to enhance entity alignment. They explicitly incorporate relations in entity representation and alignment, demonstrating remarkable results. However, affected by the semantic assumptions from early works, these works represent a relation by combining all the entities it connects, ignoring the semantic independence between entity and relation. Moreover, since these works perform alignment by comparing embedding similarity, they fail to consider a graph level alignment and tend to find local false correspondences. Jinzhu Yang, Wei Zhou 0019, Wanhui Qian, Xin Wang 0086, Jizhong Han, Songlin Hu 0001 |
CIKM | 6 |
| 2021 | Collaborative Spatial-Temporal Modeling for Language-Queried Video Actor SegmentationabstractLanguage-queried video actor segmentation aims to predict the pixel-level mask of the actor which performs the actions described by a natural language query in the target frames. Existing methods adopt 3D CNNs over the video clip as a general encoder to extract a mixed spatio-temporal feature for the target frame. Though 3D convolutions are amenable to recognizing which actor is performing the queried actions, it also inevitably introduces misaligned spatial information from adjacent frames, which confuses features of the target frame and yields inaccurate segmentation. Therefore, we propose a collaborative spatial-temporal encoder-decoder framework which contains a 3D temporal encoder over the video clip to recognize the queried actions, and a 2D spatial encoder over the target frame to accurately segment the queried actors. In the decoder, a Language-Guided Feature Selection (LGFS) module is proposed to flexibly integrate spatial and temporal features from the two encoders. We also propose a Cross-Modal Adaptive Modulation (CMAM) module to dynamically recombine spatial- and temporal-relevant linguistic features for multimodal feature interaction in each stage of the two encoders. Our method achieves new state-of-the-art performance on two popular benchmarks with less computational overhead than previous approaches. Tianrui Hui, Shaofei Huang 0001, Si Liu 0001, Guanbin Li, Wenguan Wang, Jizhong Han, Fei Wang 0032 |
CVPR | 7 |
| 2021 | Fed-Tra: Improving Accuracy of Deep Learning Model on Non-iid in Federated Learning
Wenjie Xiao, Xuehai Tang, Biyu Zhou, Wang Wang, Yangchen Dong, Liangjun Zang, Jizhong Han, Songlin Hu 0001 |
ICA3PP (1) | 7 |
| 2021 | Aligning the training and evaluation of unsupervised text style TransferabstractIn the text style transfer task, models modify the attribute style of given texts while keeping the style-irrelevant content unchanged. Previous work has proposed many approaches on the non-parallel corpus (without style-to-style training pairs). These approaches are mostly motivated by heuristic intuition and fail to precisely control texts’ attributes, such as the amount of preserved semantics, which leaves discrepancies between training and evaluation. This paper proposes a novel training method based on the evaluation metrics to address the discrepancy issue. Specifically, the model first evaluates different aspects of the transferred texts and provides the differentiable quality approximations by employing extra supervising modules. Then the model is optimized by bridging the gap between approximations and expectations. Extensive experiments conducted on two sentiment style datasets demonstrate the effectiveness of our proposal compared with some competitive baselines. Wanhui Qian, Fuqing Zhu, Jinzhu Yang, Jizhong Han, Songlin Hu 0001 |
ICASSP | 4 |
| 2021 | Topic Sequence Embedding for User Identity Linkage from Heterogeneous Behavior DataabstractIn social media, user identity linkage is a vital information security issue of identifying users’ private information across multiple online social networks. With the popularity of behavior-rich social services, existing methods attempt to align users through encoding behaviors. However, most of the efforts suffer from the high variety and heterogeneity of behavior data across social networks, resulting in a limitation of modeling user intrinsic characteristics. To address the above issues, we focus on keyword-based topics to formulate user’s variety behaviors for user identity linkage. In this paper, a novel Topic Sequence Embedding (TSeqE) method is proposed to embed contextual information of topics to represent users’ intrinsic characteristics for identity linkage. Furthermore, we introduce a domain-adversarial training strategy to tackle the behavior heterogeneity problem. Our experiments on two real-world datasets demonstrate that TSeqE produces a significant improvement compared with several strong baselines. Jinzhu Yang, Wei Zhou 0019, Wanhui Qian, Jizhong Han, Songlin Hu 0001 |
ICASSP | 4 |
| 2021 | Li-Net: Large-Pose Identity-Preserving Face Reenactment NetworkabstractFace reenactment is a challenging task, as it is difficult to maintain accurate expression, pose and identity simultaneously. Most existing methods directly apply driving facial landmarks to reenact source faces and ignore the intrinsic gap between two identities, resulting in the identity mismatch issue. Besides, they neglect the entanglement of expression and pose features when encoding driving faces, leading to inaccurate expressions and visual artifacts on large-pose reenacted faces. To address these problems, we propose a Large-pose Identity-preserving face reenactment network, LI-Net. Specifically, the Landmark Transformer is adopted to adjust driving landmark images, which aims to narrow the identity gap between driving and source landmark images. Then the Face Rotation Module and the Expression Enhancing Generator decouple the transformed landmark image into pose and expression features, and reenact those attributes separately to generate identity-preserving faces with accurate expressions and poses. Both qualitative and quantitative experimental results demonstrate the superiority of our method. Jin Liu 0020, Zhaoxing Li, Cai Yu, Shuqiao Zou, Jiao Dai, Jizhong Han |
ICME | 8 |
| 2021 | DLFMNet: End-to-End Detection and Localization of Face Manipulation Using Multi-Domain FeaturesabstractRecently, more and more realistic facial manipulation images and videos, known as DeepFakes, have been created and rapidly circulated in social media. Therefore, it is crucial to develop effective and efficient methods to detect the malicious DeepFakes. Previous approaches all adopt a two-step pipeline with multiple separate models, i.e., first face detection and then face forensics, and lacks robustness against compressed data. In this paper, we propose an end-to-end framework for detection and localization of face manipulation, named DLFMNet, which effectively integrates face detection and face forensics into one model, avoiding intermediate processes like image cropping and feature re-extraction. In addition, to capture richer and more robust manipulated clues, we exploit multi-domain features that takes advantages of two different but complementary domains (i.e., RGB and noise). The evaluations on FaceForensics++ dataset demonstrate the effectiveness of our proposed DLFMNet. https://github.com/LightningChan/DLFMNet. Jin Liu 0020, Cai Yu, Shuqiao Zou, Jiao Dai, Jizhong Han |
ICME | 7 |
| 2021 | Scene Graph Generation With Hierarchical ContextabstractScene graph generation has received increasing attention in recent years. Enhancing the predicate representations is an important entry point to this task. There are various methods to fully investigate the context of representation enhancement. In this brief, we analyze the decisive factors that can significantly affect the relation detection results. Our analysis shows that spatial correlations between objects, focused regions of objects, and global hints related to the relations have strong influences in relation prediction and contradiction elimination. Based on our analysis, we propose a hierarchical context network (HCNet) to generate a scene graph. HCNet consists of three contexts, including interaction context, depression context, and global context, which integrates information from pair, object, and graph levels. The experiments show that our method outperforms the state-of-the-art methods on the Visual Genome (VG) data set. Guanghui Ren, Lejian Ren, Yue Liao, Si Liu 0001, Bo Li 0006, Jizhong Han, Shuicheng Yan |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2020 | Symmetric Metric Learning with Adaptive Margin for RecommendationabstractMetric learning based methods have attracted extensive interests in recommender systems. Current methods take the user-centric way in metric space to ensure the distance between user and negative item to be larger than that between the current user and positive item by a fixed margin. While they ignore the relations among positive item and negative item. As a result, these two items might be positioned closely, leading to incorrect results. Meanwhile, different users usually have different preferences, the fixed margin used in those methods can not be adaptive to various user biases, and thus decreases the performance as well. To address these two problems, a novel Symmetic Metric Learning with adaptive margin (SML) is proposed. In addition to the current user-centric metric, it symmetically introduces a positive item-centric metric which maintains closer distance from positive items to user, and push the negative items away from the positive items at the same time. Moreover, the dynamically adaptive margins are well trained to mitigate the impact of bias. Experimental results on three public recommendation datasets demonstrate that SML produces a competitive performance compared with several state-of-the-art methods. Fuqing Zhu, Wanhui Qian, Liangjun Zang, Jizhong Han, Songlin Hu 0001 |
AAAI | 6 |
| 2020 | An Event-Oriented Neural Ranking Model for News RetrievalabstractEvent-oriented news retrieval (ENR) is the task of retrieving news articles related to the specific event in response to the event-oriented query. Previous approaches usually focus on optimizing traditional retrieval models through hand-crafted features from the perspective of new articles. However, these approaches often fail to work well in reality, as they do not consider the essential natures of the event, i.e., dynamics, coupling. In this paper, we propose a novel and effective event-oriented neural ranking model for news retrieval (ENRMNR). Our model exploits a deep attention mechanism to tackle the dynamics and coupling derived from event evolution. Specifically, the word-level bidirectional attention allows the model to identify which query words about the subevent are related to the news article words, and vice-versa, in order to tackle the dynamics. Moreover, the hierarchical attention at passage-level and document-level allows it to capture fine-grained event representations for the coupling between different events within a news article. Experimental results on real-world datasets demonstrate that ENRMNR model significantly outperforms competitive models. Wanhui Qian, Liangjun Zang, Fuqing Zhu, Ruixuan Li 0001, Jizhong Han, Songlin Hu 0001 |
CIKM | 7 |
| 2020 | Early Detection of Fake News by Utilizing the Credibility of News, Publishers, and Users based on Weakly Supervised LearningabstractThe dissemination of fake news significantly affects personal reputation and public trust.Recently, fake news detection has attracted tremendous attention, and previous studies mainly focused on finding clues from news content or diffusion path.However, the required features of previous models are often unavailable or insufficient in early detection scenarios, resulting in poor performance.Thus, early fake news detection remains a tough challenge.Intuitively, the news from trusted and authoritative sources or shared by many users with a good reputation is more reliable than other news.Using the credibility of publishers and users as prior weakly supervised information, we can quickly locate fake news in massive news and detect them in the early stages of dissemination.In this paper, we propose a novel Structure-aware Multi-head Attention Network (SMAN), which combines the news content, publishing, and reposting relations of publishers and users, to jointly optimize the fake news detection and credibility prediction tasks.In this way, we can explicitly exploit the credibility of publishers and users for early fake news detection.We conducted experiments on three real-world datasets, and the results show that SMAN can detect fake news in 4 hours with an accuracy of over 91%, which is much faster than the state-of-the-art models. Chunyuan Yuan, Qianwen Ma, Wei Zhou 0019, Jizhong Han, Songlin Hu 0001 |
COLING | 4 |
| 2020 | Referring Image Segmentation via Cross-Modal Progressive ComprehensionabstractReferring image segmentation aims at segmenting the foreground masks of the entities that can well match the description given in the natural language expression. Previous approaches tackle this problem using implicit feature interaction and fusion between visual and linguistic modalities, but usually fail to explore informative words of the expression to well align features from the two modalities for accurately identifying the referred entity. In this paper, we propose a Cross-Modal Progressive Comprehension (CMPC) module and a Text-Guided Feature Exchange (TGFE) module to effectively address the challenging task. Concretely, the CMPC module first employs entity and attribute words to perceive all the related entities that might be considered by the expression. Then, the relational words are adopted to highlight the correct entity as well as suppress other irrelevant ones by multimodal graph reasoning. In addition to the CMPC module, we further leverage a simple yet effective TGFE module to integrate the reasoned multimodal features from different levels with the guidance of textual information. In this way, features from multi-levels could communicate with each other and be refined based on the textual context. We conduct extensive experiments on four popular referring segmentation benchmarks and achieve new state-of-the-art performances. Code is available at https://github.com/spyflying/CMPC-Refseg. Shaofei Huang 0001, Tianrui Hui, Si Liu 0001, Guanbin Li, Yunchao Wei, Jizhong Han, Luoqi Liu, Bo Li 0006 |
CVPR | 6 |
| 2020 | Tail: An Automated and Lightweight Gradient Compression Framework for Distributed Deep LearningabstractExisting gradient compression schemes fail to automatically determine the compression ratio or are accompanied by high compression overhead. To address this, we present Tail, an automated and lightweight gradient compression framework stacked by three modules, quantization, sparsification, and encoding. Without any hand-tuned effort, quantization module automatically adjusts the compression ratio along training iterations to retain accuracy first. Then, sparsification and encoding modules are successively applied to the quantized gradient to further improve compression ratio. Moreover, Tail reduces the compression overhead by approximate computing in the automated decision-making process. Experiments validate that Tail can reduce communication traffic by an order of magnitude while retaining or even improving model accuracy. Jinrong Guo, Songlin Hu 0001, Wang Wang, Chunrong Yao, Jizhong Han, Ruixuan Li 0001 |
DAC | 5 |
| 2020 | RE-GCN: Relation Enhanced Graph Convolutional Network for Entity Alignment in Heterogeneous Knowledge Graphs
Jinzhu Yang, Wei Zhou 0019, Lingwei Wei, Junyu Lin 0002, Jizhong Han, Songlin Hu 0001 |
DASFAA (2) | 5 |
| 2020 | Linguistic Structure Guided Context Modeling for Referring Image Segmentation
Tianrui Hui, Si Liu 0001, Shaofei Huang 0001, Guanbin Li, Sansi Yu, Faxi Zhang, Jizhong Han |
ECCV (10) | 7 |
| 2020 | Structural Position Network for Aspect-Based Sentiment Classification
Pu Song, Wei Jiang 0028, Fuqing Zhu, Jizhong Han, Songlin Hu 0001 |
ICANN (2) | 5 |
| 2020 | Accelerating Distributed Deep Learning By Adaptive Gradient QuantizationabstractTo accelerate distributed deep learning, gradient quantization technique is widely used to reduce the communication cost. However, the existing quantization schemes suffer from either model accuracy degradation or low compression ratio (arisen from a redundant setting of quantization level or high overhead in determining the level). In this work, we propose a novel adaptive quantization scheme (AdaQS) to explore the balance between model accuracy and quantization level. AdaQS determines the quantization level automatically according to gradient's mean to standard deviation ratio (MSDR). Then, to reduce the quantization overhead, we employ a computationally-friendly way of moment estimation to calculate the MSDR. Finally, theoretical analysis of AdaQS's convergence is conducted for non-convex objectives. Experiments demonstrate that AdaQS performs excellently on very deep model GoogleNet with 2.55% accuracy improvement relative to vanilla SGD and achieves 1.8x end-to-end speedup on AlexNet in a distributed cluster with 4*4 GPUs. Jinrong Guo, Wantao Liu, Wang Wang, Jizhong Han, Ruixuan Li 0001, Songlin Hu 0001 |
ICASSP | 4 |
| 2020 | FSSPOTTER: Spotting Face-Swapped Video by Spatial and Temporal CluesabstractRecent advances in face generation and manipulation have enabled the creation of sophisticated face-swapped videos, also known as DeepFakes, which brings great potential threats to our society. Hence, it is crucial to develop effective approaches to distinguish them. Currently, face-swapped videos produced by existing methods are prone to exhibit some subtle spatial and temporal manipulated traces, which can be utilized as distinctive clues for face-swapped video detection. In this paper, we propose a unified framework, named FSSpotter, to explore rich spatial and temporal information in the video simultaneously. It consists of a Spatial Feature Extractor (SFE), which aims to discover spatial evidences within a single frame, and a Temporal Feature Aggregator (TFA), which is responsible for capturing temporal inconsistencies between frames. Moreover, a novel data processing strategy is adopted to highlight the inconsistencies of forged face with its surrounding regions. The evaluations on Deepfakes of FaceForensics++, DeepfakeTIMIT, UADFV and Celeb-DF datasets demonstrate that the proposed approach achieves better or comparable performance on AUC scores. Jin Liu 0020, Guangzhi Zhou, Hongchao Gao, Jiao Dai, Jizhong Han |
ICME | 7 |
| 2020 | A Multi-head Self-relation Network for Scene Text RecognitionabstractThe text embedded in scene images can be seen everywhere in our lives. However, recognizing text from natural scene images is still a challenge because of its diverse shapes and distorted patterns. Recently, advanced recognition networks generally treat scene text recognition as a sequence prediction task. Although achieving excellent performance, these recognition networks consider the feature map cells as independent individuals and update cells state without utilizing the information of their related cells. And the local receptive field of traditional convolutional neural network (CNN) makes a single cell that cannot cover the whole text region in an image. Due to these issues, the existing recognition networks cannot extract the global context information in a visual scene. To deal with the above problems, we propose a Multi-head Self-relation Network(MSRN) for scene text recognition in this paper. The MSRN consists of several multihead self-relation layers, which are designed for extracting the global context information of a visual scene. Then the information of the related cells can be fused by multi-head self-relation layer. Furthermore, experiments over several public datasets demonstrate that our proposed recognition network achieves superior performance on several benchmark datasets including IC03, IC13, IC15, SVT-Perspective. Junwei Zhou 0005, Hongchao Gao, Jiao Dai, Dongqin Liu, Jizhong Han |
ICPR | 5 |
| 2020 | A Rating Bias Formulation based on Fuzzy Set for RecommendationabstractIn recommender systems, the user uncertain preference results in unexpected ratings. Previous approaches (e.g., BiasMF) only adjust the rating value based on the bias vector, ignoring the uncertainty of rating. This paper makes an initial attempt in integrating the influence of user uncertain degree and user rating bias into the matrix factorization framework, simultaneously. An approach based on fuzzy set, called fuZzy Matrix Factorization (ZMF), is proposed. Specifically, a fuzzy set of like is defined for each user, and the membership function is utilized to measure the degree of an item belonging to the fuzzy set. Then, the user uncertain preference matrix is obtained, which could explain and represent the user bias and uncertainty effectively. Furthermore, to enhance the computational impact on sparse matrix, the uncertain preference is formulated as a side-information for fusion. Besides, the proposed approach could be extended to others due to independency on additional data sources. Experimental results on three datasets show that ZMF produces an effective improvement. Fuqing Zhu, Jiao Dai, Liangjun Zang, Yipeng Su, Jizhong Han, Songlin Hu 0001 |
IJCNN | 6 |
| 2020 | Graph Convolutional Networks for Target-oriented Opinion Words Extraction with Adversarial TrainingabstractThe task of Target-oriented Opinion Words Extraction aims to extract the corresponding opinion words for a given opinion target from the sentence. Recently, the methods based on recurrent neural networks have shown promising results for this task. However, these approaches only considered the sequential information of the sentences and ignored the syntactic structure. In this paper, we propose a novel graph convolutional network with adversarial training to extract the opinion words. We present a graph convolutional network based on dependency tree to learn the syntactic representation of the input. Besides, we train our model with the mixture of original examples and adversarial examples, which can improve the robustness of the model. We conduct experiments on four benchmarking datasets and the results illustrate that our proposed model consistently outperforms the state-of-the-art methods. Wei Jiang 0028, Po Song, Yipeng Su, Tao Guo 0006, Jizhong Han, Songlin Hu 0001 |
IJCNN | 6 |
| 2020 | Exploiting Heterogeneous Artist and Listener Preference Graph for Music Genre ClassificationabstractMusic genres are useful for indexing, organizing, searching, and recommending songs and albums. Therefore, the automatic classification of music genres is an essential part of almost all kinds of music applications. Recent works focus on exploiting text, audio, or multi-modal information for genre classification, without considering the influence of the artists' and listeners' preference. However, intuitively, artists have their composing preferences, and listeners also have their music tastes. Both of them provide helpful hints to the music genre from different views, which are crucial to improve classification performance. Chunyuan Yuan, Qianwen Ma, Junyang Chen 0001, Wei Zhou 0019, Xiaodan Zhang 0004, Xuehai Tang, Jizhong Han, Songlin Hu 0001 |
ACM Multimedia | 7 |
| 2020 | Hierarchical Interaction Networks with Rethinking Mechanism for Document-Level Sentiment Analysis
Lingwei Wei, Dou Hu 0001, Wei Zhou 0019, Xuehai Tang, Xiaodan Zhang 0004, Xin Wang 0086, Jizhong Han, Songlin Hu 0001 |
ECML/PKDD (3) | 7 |
| 2020 | Beyond Statistical Relations: Integrating Knowledge Relations into Style Correlations for Multi-Label Music Style ClassificationabstractAutomatically labeling multiple styles for every song is a comprehensive application in all kinds of music websites. Recently, some researches explore review-driven multi-label music style classification and exploit style correlations for this task. However, their methods focus on mining the statistical relations between different music styles and only consider shallow style relations. Moreover, these statistical relations suffer from the underfitting problem because some music styles have little training data. To tackle these problems, we propose a novel knowledge relations integrated framework (KRF) to capture the complete style correlations, which jointly exploits the inherent relations between music styles according to external knowledge and their statistical relations. Based on the two types of relations, we use graph convolutional network to learn the deep correlations between styles automatically. Experimental results show that our framework significantly outperforms the state-of-the-art methods. Further studies demonstrate that our framework can effectively alleviate the underfitting problem and learn meaningful style correlations. Qianwen Ma, Chunyuan Yuan, Wei Zhou 0019, Jizhong Han, Songlin Hu 0001 |
WSDM | 4 |
| 2020 | ORDNet: Capturing Omni-Range Dependencies for Scene ParsingabstractLearning to capture dependencies between spatial positions is essential to many visual tasks, especially the dense labeling problems like scene parsing. Existing methods can effectively capture long-range dependencies with self-attention mechanism while short ones by local convolution. However, there is still much gap between long-range and short-range dependencies, which largely reduces the models' flexibility in application to diverse spatial scales and relationships in complicated natural scene images. To fill such a gap, we develop a Middle-Range (MR) branch to capture middle-range dependencies by restricting self-attention into local patches. Also, we observe that the spatial regions which have large correlations with others can be emphasized to exploit long-range dependencies more accurately, and thus propose a Reweighed Long-Range (RLR) branch. Based on the proposed MR and RLR branches, we build an Omni-Range Dependencies Network (ORDNet) which can effectively capture short-, middle- and long-range dependencies. Our ORDNet is able to extract more comprehensive context information and well adapt to complex spatial variance in scene images. Extensive experiments show that our proposed ORDNet outperforms previous state-of-the-art methods on three scene parsing benchmarks including PASCAL Context, COCO Stuff and ADE20K, demonstrating the superiority of capturing omni-range dependencies in deep models for scene parsing task. Shaofei Huang 0001, Si Liu 0001, Tianrui Hui, Jizhong Han, Bo Li 0006, Jiashi Feng, Shuicheng Yan |
IEEE Trans. Image Process. | 4 |
| 2020 | Yet another approach to understanding news event evolution
Shangwen Lv, Longtao Huang, Liangjun Zang, Wei Zhou 0019, Jizhong Han, Songlin Hu 0001 |
World Wide Web | 5 |
| 2019 | A Fuzzy Set Based Approach for Rating BiasabstractIn recommender systems, the user uncertain preference results in unexpected ratings. This paper makes an initial attempt in integrating the influence of user uncertain degree into the matrix factorization framework. Specifically, a fuzzy set of like for each user is defined, and the membership function is utilized to measure the degree of an item belonging to the fuzzy set. Furthermore, to enhance the computational effect on sparse matrix, the uncertain preference is formulated as a side-information for fusion. Experimental results on three real-world datasets show that the proposed approach produces stable improvements compared with others. Jiao Dai, Fuqing Zhu, Liangjun Zang, Songlin Hu 0001, Jizhong Han |
AAAI | 6 |
| 2019 | SAM-Net: Integrating Event-Level and Chain-Level Attentions to Predict What Happens NextabstractScripts represent knowledge of event sequences that can help text understanding. Script event prediction requires to measure the relation between an existing chain and the subsequent event. The dominant approaches either focus on the effects of individual events, or the influence of the chain sequence. However, only considering individual events will lose much semantic relations within the event chain, and only considering the sequence of the chain will introduce much noise. With our observations, both the individual events and the event segments within the chain can facilitate the prediction of the subsequent event. This paper develops self attention mechanism to focus on diverse event segments within the chain and the event chain is represented as a set of event segments. We utilize the event-level attention to model the relations between subsequent events and individual events. Then, we propose the chain-level attention to model the relations between subsequent events and event segments within the chain. Finally, we integrate event-level and chain-level attentions to interact with the chain to predict what happens next. Comprehensive experiment results on the widely used New York Times corpus demonstrate that our model achieves better results than other state-of-the-art baselines by adopting the evaluation of Multi-Choice Narrative Cloze task. Shangwen Lv, Wanhui Qian, Longtao Huang, Jizhong Han, Songlin Hu 0001 |
AAAI | 4 |
| 2019 | Text Recognition using local correlation
Hongchao Gao, Xi Wang 0014, Jizhong Han, Ruixuan Li 0001 |
BMVC | 4 |
| 2019 | A Time-Series Sockpuppet Detection Method for Dynamic Social Relationships
Wei Zhou 0019, Jingli Wang, Junyu Lin 0002, Jizhong Han, Songlin Hu 0001 |
DASFAA (1) | 5 |
| 2019 | Multi-hop Selector Network for Multi-turn Response Selection in Retrieval-based ChatbotsabstractChunyuan Yuan, Wei Zhou, Mingming Li, Shangwen Lv, Fuqing Zhu, Jizhong Han, Songlin Hu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Chunyuan Yuan, Wei Zhou 0019, Shangwen Lv, Fuqing Zhu, Jizhong Han, Songlin Hu 0001 |
EMNLP/IJCNLP (1) | 6 |
| 2019 | Imbalanced Sentiment Classification Enhanced with Discourse Marker
Tao Zhang 0101, Xing Wu 0002, Jizhong Han, Songlin Hu 0001 |
ICANN (4) | 4 |
| 2019 | AccUDNN: A GPU Memory Efficient Accelerator for Training Ultra-Deep Neural NetworksabstractWith the implementation of mainstream DL frameworks, scarce GPU memory resource is the primary bottleneck that hinders the trainability and training efficiency of ultra-deep neural networks (UDNN). Prior memory optimization works focus on removing the trainability restriction but leave the training efficiency out of consideration. To fill the gap, we present "AccUDNN", an accelerator that aims to make full use of finite GPU memory resource to speed up the training process of UDNN in this paper. AccUDNN mainly includes two modules: memory optimizer and hyperparameter tuner. Memory optimizer develops a novel performance-model guided dynamic swap out/in strategy to meet trainability first and further remedy the efficiency degradation in other swapping strategies. Then, a hyperparameter tuner is designed to explore the efficiency-optimal minibatch size and the matched learning rate after applying the dynamic swapping strategy. Evaluations demonstrate that AccUDNN cuts down the GPU memory requirement of ResNet-152 from more than 24GB to 8GB. In turn, given 12GB GPU memory budget, the efficiency-optimal minibatch size can reach 4.2x larger than Caffe and finally improve the scaling efficiency (speedup) of 8 GPUs' cluster by 1.9x. Jinrong Guo, Wantao Liu, Wang Wang, Chunrong Yao, Jizhong Han, Ruixuan Li 0001, Songlin Hu 0001 |
ICCD | 5 |
| 2019 | Jointly Embedding the Local and Global Relations of Heterogeneous Graph for Rumor DetectionabstractThe development of social media has revolutionized the way people communicate, share information and make decisions, but it also provides an ideal platform for publishing and spreading rumors. Existing rumor detection methods focus on finding clues from text content, user profiles, and propagation patterns. However, the local semantic relation and global structural information in the message propagation graph have not been well utilized by previous works. In this paper, we present a novel global-local attention network (GLAN) for rumor detection, which jointly encodes the local semantic and global structural information. We first generate a better integrated representation for each source tweet by fusing the semantic information of related retweets with the attention mechanism. Then, we model the global relationships among all source tweets, retweets, and users as a heterogeneous graph to capture the rich structural information for rumor detection. We conduct experiments on three real-world datasets, and the results demonstrate that GLAN significantly outperforms the state-of-the-art models in both rumor detection and early detection scenarios. Chunyuan Yuan, Qianwen Ma, Wei Zhou 0019, Jizhong Han, Songlin Hu 0001 |
ICDM | 4 |
| 2019 | Learning Review Representations from user and Product Level Information for Spam DetectionabstractOpinion spam has become a widespread problem in social media, where hired spammers write deceptive reviews to promote or demote products to mislead the consumers for profit or fame. Existing works mainly focus on manually designing discrete textual or behavior features, which cannot capture complex global semantics of reviews. Although recent works apply deep learning methods to learn review-level semantic features, their models ignore the impact of the user-level and product-level information on learning review semantics and the inherent user-review-product relationship information. In this paper, we propose a Hierarchical Fusion Attention Network (HFAN) to automatically learn the semantics of reviews from user and product level. Specifically, we design a multiattention unit to extract user(product)-related review information. Then, we use orthogonal decomposition and fusion attention to learn a user, review, and product representation from the review information. Finally, we take the review as a relation between user and product entity and apply TransH to jointly encode this relationship into review representation. Experimental results obtained more than 10% absolute precision improvement over the state-of-the-art performances on four real-world datasets, which show the effectiveness and versatility of the model. Chunyuan Yuan, Wei Zhou 0019, Qianwen Ma, Shangwen Lv, Jizhong Han, Songlin Hu 0001 |
ICDM | 5 |
| 2019 | Self-Representation Convolutional Neural NetworksabstractThe traditional convolutional neural networks (CNNs) learn numbers of fixed kernels (filters), which are used to obtain the representations of fixed patterns. Therefore, the knowledge representations of CNNs are limited to the number of kernels. In this paper, we present a Self-Representation Convolutional (SRC) layer to obtain richer knowledge representations of images by fully considering the self-similarity between adjacent pixels. SRC layer comprises a learnable local correlation measurement which measures the importance of adjacent pixels to the current pixel and two learnable linear parameters that perform linear projection on adjacent pixels and the weighted sum vectors, respectively. Compared with regular convolutional layers, the SRC layers can not only obtain comparable knowledge representations, but also reduce by a factor of 3× to 56× in the number of learnable parameters. Empirically, CNNs with SRC layers, called Self-Representation Convolutional Neural Networks (SRCNN), achieve strong performances on a range of visual datasets (SVHN, CIFAR-10 and CIFAR-100) while enjoying significant parameters and FLOPs savings. Hongchao Gao, Xi Wang 0014, Jizhong Han, Songlin Hu 0001, Ruixuan Li 0001 |
ICME | 4 |
| 2019 | SPL: Exploiting Unlabeled Data for Multi-label Image ClassificationabstractThe utilization of the unlabeled data provides a beneficial attempt for improving the generalization ability of the convolutional neural network (CNN) model, just as what is applied in person re-identification task. Different from that, multi-label image classification aims to predict multiple labels for each given image. The unlabeled data should be properly assigned multiple labels for regularizing the training process of CNN model. To make full use of the unlabeled data, this paper proposes a soft pseudo labeling (SPL) method for multi-label image classification. Specifically, the unlabeled samples are first generated by DCGAN and WGAN-GP. Then, the virtual multiple labels of the generated unlabeled samples are assigned based on an initial confidence value by SoftMax function. Finally, both the generated samples and original training samples are fed into the network as input, in order to learn a CNN model with stronger generalization ability. On three public multi-label image classification datasets (i.e., WIDER-Attribute, NUS-WIDE and MS-COCO), SPL provides a stable improvement over the baseline and produces a competitive performance compared with some existing multi-label image classification methods. Weibo Zhang, Fuqing Zhu, Jiao Dai, Songlin Hu 0001, Jizhong Han, Tao Guo 0006 |
ICME | 5 |
| 2019 | Fusion Convolutional Attention Network for Opinion Spam Detection
Qianwen Ma, Chunyuan Yuan, Wei Zhou 0019, Jizhong Han, Songlin Hu 0001 |
ICONIP (1) | 5 |
| 2019 | Mask and Infill: Applying Masked Language Model for Sentiment TransferabstractThis paper focuses on the task of sentiment transfer on non-parallel text, which modifies sentiment attributes (e.g., positive or negative) of sentences while preserving their attribute-independent contents. Existing methods adopt RNN encoder-decoder structure to generate a new sentence of a target sentiment word by word, which is trained on a particular dataset from scratch and have limited ability to produce satisfactory sentences. When people convert the sentiment attribute of a given sentence, a simple but effective approach is to only replace the sentiment tokens of the sentence with other expressions indicative of the target sentiment, instead of building a new sentence from scratch. Such a process is very similar to the task of Text Infilling or Cloze. With this intuition, we propose a two steps approach: Mask and Infill. In the \emph{mask} step, we identify and mask the sentiment tokens of a given sentence. In the \emph{infill} step, we utilize a pre-trained Masked Language Model (MLM) to infill the masked positions by predicting words or phrases conditioned on the context\footnote{In this paper, \emph{content} and \emph{context} are equivalent, \emph{style}, \emph{attribute} and \emph{label} are equivalent.}and target sentiment. We evaluate our model on two review datasets \emph{Yelp} and \emph{Amazon} by quantitative, qualitative, and human evaluations. Experimental results demonstrate that our model achieve state-of-the-art performance on both accuracy and BLEU scores. Xing Wu 0002, Tao Zhang 0101, Liangjun Zang, Jizhong Han, Songlin Hu 0001 |
IJCAI | 4 |
| 2019 | A Span-based Joint Model for Opinion Target Extraction and Target Sentiment ClassificationabstractTarget-Based Sentiment Analysis aims at extracting opinion targets and classifying the sentiment polarities expressed on each target. Recently, token based sequence tagging methods have been successfully applied to jointly solve the two tasks, which aims to predict a tag for each token. Since they do not treat a target containing several words as a whole, it might be difficult to make use of the global information to identify that opinion target, leading to incorrect extraction. Independently predicting the sentiment for each token may also lead to sentiment inconsistency for different words in an opinion target. In this paper, inspired by span-based methods in NLP, we propose a simple and effective joint model to conduct extraction and classification at span level rather than token level. Our model first emulates spans with one or more tokens and learns their representation based on the tokens inside. And then, a span-aware attention mechanism is designed to compute the sentiment information towards each span. Extensive experiments on three benchmark datasets show that our model consistently outperforms the state-of-the-art methods. Longtao Huang, Tao Guo 0006, Jizhong Han, Songlin Hu 0001 |
IJCAI | 4 |
| 2019 | Ensemble Attention For Text Recognition In Natural ImagesabstractRecognizing text from natural images is a challenging and hot research topic in computer vision, yet not completely solved. The recent methods regard this task as a sequence labeling problem. In this task, there is a strong correspondence between the position of the input image patches sequence and the output character sequence. However, most of the recent recognition systems rarely consider this local information of the input sequence when recognizing the current character. In contrast to this, we present a Local Restricted Attention (LRA) mechanism to encode the current vector by considering adjacent vectors of the input sequence. We propose an ensemble decoder block which combines LRA mechanism with a regular decoder mechanism. This block not only brings significant improvement of recognition results under shorter training time but also can be easily embedded in other recognition frameworks. In addition, we propose a scene text recognition network based on the ensemble decoder. The experimental performances show that the proposed model achieves the state-of-the-art on several benchmark datasets including IIIT-5K, SVT, CUTE80, SVT-Perspective and ICDARs. Hongchao Gao, Xi Wang 0014, Jizhong Han, Ruixuan Li 0001 |
IJCNN | 4 |
| 2019 | LMLSTM: Extract Event-Oriented Keyphrase From News Stream
Longtao Huang, Liangjun Zang, Jizhong Han, Songlin Hu 0001 |
IJCNN | 4 |
| 2019 | Jily: Cost-Aware AutoScaling of Heterogeneous GPU for DNN Inference in Public CloudabstractRecently, a large number of DNN inference services have emerged in public clouds, making the low-cost deployment of DNN inference services a hot research topic. Previous studies have failed to take into account GPU heterogeneity and batch processing, both of which will seriously affect the financial cost as well as the latency. In this paper, we study the problem of DNN inference service deployment in public cloud, considering both GPU heterogeneity and batch processing. The goal is to minimize the financial costs under the constraint of latency. We propose Jily, an autoscaling scheduler for DNN inference services to minimize the cost while satisfying the given latency SLO. Jily finds the optimal heterogeneous GPU instance provisioning through a DNN inference model profiler, a latency estimator, a workload predictor and a cost-aware scaler. Simulation results demonstrate that Jily can reduce average cost by up to 28% compared to a state-of-the-art autoscaling approach. Further, Jily has been proved to have good versatility and robustness under different batching mechanisms and latency SLO constrains. Zhaoxing Wang, Xuehai Tang, Qiuyang Liu, Jizhong Han |
IPCCC | 4 |
| 2019 | Finding Images by Dialoguing with ImageabstractImage retrieval in complicated scene is a challenging task that requires the comprehensive understanding of an image. In this paper, we propose a scene graph based image retrieval framework that combines the scene graph generation with image retrieval and fine tuning the searching results via a dialogue mechanism. Specifically, we proposed an image retrieval oriented scene graph generation model that takes an image and a text describing the image as inputs. The additional text input is used to control the generated scene graph. It provides information for a newly introduced attributes head to better predict the attributes and helps constructing an adjacency matrix at the same time. Graph Convolutional Network is further used to gather information among nodes for precise relation estimation. Moreover, modification on the scene graph can be done by changing the text. Our proposed approach achieves the state-of-the-art performances in both scene graph based image retrieval and scene graph generation in the Visual Genome dataset. Lejian Ren, Si Liu 0001, Han Huang 0002, Jizhong Han, Shuicheng Yan, Bo Li 0006 |
ACM Multimedia | 4 |
| 2019 | DFPE: Explaining Predictive Models for Disk Failure PredictionabstractRecent research works on disk failure prediction achieve a high detection rate and a low false alarm rate with complex models at the cost of explainability. The lack of explainability is likely to hide bias or overfitting in the models, resulting in bad performance in real-world applications. To address the problem, we propose a new explanation method DFPE designed for disk failure prediction to explain failure predictions made by a model and infer prediction rules learned by a model. DFPE explains failure predictions by performing a series of replacement tests to find out the failure causes while it explains models by aggregating explanations for the failure predictions. A presented use case on a real-world dataset shows that compared to current explanation methods, DFPE can explain more about failure predictions and models with more accuracy. Thus it helps to target and handle the hidden bias and overfitting, measures feature importances from a new perspective and enables intelligent failure handling. Yanwen Xie, Dan Feng 0001, Fang Wang 0001, Xuehai Tang, Jizhong Han |
MSST | 5 |
| 2019 | A GPU memory efficient speed-up scheme for training ultra-deep neural networks: posterabstractUltra-deep neural network(UDNN) tends to yield higher-quality model but its training process is often difficult to handle. Scarce GPU DRAM capacity is the primary bottleneck that limits the depth of neural network and the range of trainable minibatch size. In this paper, we present a scheme that dedicates to make the utmost use of finite GPU memory resource to speed up the training process for UDNN. Firstly, a performance-model guided dynamic swap out/in strategy between GPU and host memory is carefully orchestrated to tackle the out-of-memory problem without introducing performance penalty. Then, a hyperparameter (minibatch size, learning rate) tuning policy is designed to explore the optimal configuration after applying the swap strategy from the perspectives of training time and final accuracy simultaneously. Finally, we verify the effectiveness of our scheme in both single and distributed GPU mode. Jinrong Guo, Wantao Liu, Wang Wang, Qu Lu, Songlin Hu 0001, Jizhong Han, Ruixuan Li 0001 |
PPoPP | 6 |
| 2019 | A Multimodal Text Matching Model for Obfuscated Language Identification in Adversarial Communication?abstractObfuscated language is created to avoid censorship in adversarial communication such as sensitive information conveying, strong sentiment expression, secret actions plan, and illegal trading. The obfuscated sentences are usually generated by replacing one word with another to conceal the textual content. Intelligence and security agencies identify such adversarial messages by scanning with a watch-list of red-flagged terms. Though semantic expansion techniques are adopted, the precision and recall of the identification is limited due to the ambiguity and the unbounded creation way. To this end, this paper frames the obfuscated language identification problem as a text matching task, where each message is checked whether matches a red-flagged term. We propose a multimodal text matching model which combining textual and visual features. The proposed model extends a Bi-directional Long Short Term Memory network with a visual-level representation component to achieve the given task. Comparative experiments on real-world dataset demonstrate that the proposed method could achieve a better performance than the previous methods. Longtao Huang, Junyu Lin 0002, Jizhong Han, Songlin Hu 0001 |
WWW | 4 |
| 2018 | Cross-Domain Human Parsing via Adversarial Feature and Label AdaptationabstractHuman parsing has been extensively studied recently due to its wide applications in many important scenarios. Mainstream fashion parsing models (i.e., parsers) focus on parsing the high-resolution and clean images. However, directly applying the parsers trained on benchmarks of high-quality samples to a particular application scenario in the wild, e.g., a canteen, airport or workplace, often gives non-satisfactory performance due to domain shift. In this paper, we explore a new and challenging cross-domain human parsing problem: taking the benchmark dataset with extensive pixel-wise labeling as the source domain, how to obtain a satisfactory parser on a new target domain without requiring any additional manual labeling? To this end, we propose a novel and efficient cross-domain human parsing model to bridge the cross-domain differences in terms of visual appearance and environment conditions and fully exploit commonalities across domains. Our proposed model explicitly learns a feature compensation network, which is specialized for mitigating the cross-domain differences. A discriminative feature adversarial network is introduced to supervise the feature compensation to effectively reduces the discrepancy between feature distributions of two domains. Besides, our proposed model also introduces a structured label adversarial network to guide the parsing results of the target domain to follow the high-order relationships of the structured labels shared across domains. The proposed framework is end-to-end trainable, practical and scalable in real applications. Extensive experiments are conducted where LIP dataset is the source domain and 4 different datasets including surveillance videos, movies and runway shows without any annotations, are evaluated as target domains. The results consistently confirm data efficiency and performance advantages of the proposed method for the challenging cross-domain human parsing problem. Si Liu 0001, Yao Sun 0004, Defa Zhu, Guanghui Ren, Jiashi Feng, Jizhong Han |
AAAI | 7 |
| 2018 | Sibyl: Host Load Prediction with an Efficient Deep Learning Model in Cloud Computing
Xuehai Tang, Jizhong Han, Peng Wang 0028 |
ICA3PP (2) | 3 |
| 2018 | OME: An Optimized Modeling Engine for Disk Failure Prediction in Heterogeneous DatacenterabstractNowadays, there are lots of disks from various disk models in datacenter. It is a challenge to make failure prediction for all disk models with high precision and high coverage. One-for-one modeling, transfer learning modeling and one-for-all modeling are proposed to address the challenge. However, none of them works well for all disk models and the automation problem for method selection and parameter tuning still persists. In this paper, we propose OME, an optimized modeling engine for disk failure prediction in heterogeneous datacenter. It builds a basis predictive model with one-for-all modeling and searches for the optimized with one-for-one and transfer learning modeling for every disk model. To achieve automation, OME employs a simple but effective transfer learning method, does cross-validation for comparison, prunes the tuning space, and constructs a directed acyclic graph for parallelism. Evaluation on a dataset from a real-world datacenter shows that OME outperforms a one-for-all predictive model from previous work by 18.5% overall, and the improvement for 43.3% disk models reaches over 30%. Yanwen Xie, Dan Feng 0001, Fang Wang 0001, Jizhong Han, Xuehai Tang |
ICCD | 5 |
| 2018 | A Hybrid Model Based on the Rating Bias and Textual Bias for Recommender Systems
Jiao Dai, Songlin Hu 0001, Jizhong Han |
ICONIP (2) | 4 |
| 2018 | Multi-stage Gradient Compression: Overcoming the Communication Bottleneck in Distributed Deep Learning
Qu Lu, Wantao Liu, Jizhong Han, Jinrong Guo |
ICONIP (1) | 3 |
| 2018 | An Interactivity-Based Personalized Mutual Reinforcement Model for Microblog Topic Summarization
Lu Zhang 0038, Liangjun Zang, Longtao Huang, Jizhong Han, Songlin Hu 0001 |
PRICAI (1) | 4 |
| 2017 | Dynamic Forest Model for Sentiment Classification
Jiao Dai, Jizhong Han |
ICONIP (5) | 4 |
| 2017 | Performance analysis and optimization for chunked network coding based wireless cooperative downloading systemsabstractDense network coding (NC) is widely used in wireless cooperative downloading systems. Wireless devices have limited computing resources. Researchers have recently found that dense NC is not suitable because of its high coding complexity, and it is necessary to use chunked NC in wireless environments. However, chunked NC can cause more communications, and the amount of communications is affected by the chunk size. Therefore, setting a suitable chunk size to improve the overall perfor-mance of chunked NC is a prerequisite for applying it in wireless cooperative downloading systems. Most of the existing studies on chunked NC focus on centralized wireless broadcasting systems, which are different from wireless cooperative downloading systems with distributed features. Accordingly, we study the performance of chunked NC based wireless cooperative downloading systems. First, an analysis model is established using a Markov process taking the distributed features into consideration, and then the block collection completion time of encoded blocks for cooperative downloading is optimized based on the analysis model. Furthermore, queuing theory is used to model the decoding process of the chunked NC. Combining queuing theory with the analysis model, the decoding completion time for cooperative downloading is optimized, and the optimal chunk size is derived. Numerical simulation shows that the block collection completion time and the decode completion time can be largely reduced after optimization. Xiuxiu Wen, Junyu Lin 0002, Guangsheng Feng, Hongwu Lv, Jizhong Han |
Frontiers Inf. Technol. Electron. Eng. | 6 |
| 2016 | An efficient graph data processing system for large-scale social network service applicationsabstractSummary Trust in social network draws more and more attentions from both the academia and industry fields. Public opinion analysis is a direct way to increase the trust in social network. Because the public opinion analysis can be expressed naturally by the graph algorithm and graph data are the default data organization mechanism used in large‐scale social network service applications, more and more research works apply the graph processing system to deal with the public opinion analysis. As the data volume is growing rapidly, the distributed graph systems are introduced to process the large‐scale public opinion analysis. Most of graph algorithms introduce a large number of data iterations, so the synchronization requirements between successive iterations can severely jeopardize the effectiveness of parallel operations, which makes the data aggregation and analysis operations become slower. In this paper, we propose a large‐scale graph data processing system to address these issues, which includes a graph data processing model, Arbor. Arbor develops a new graph data organization format to represent the social relationship, and the format can not only save storage space but also accelerate graph data processing operations. Furthermore, Arbor substitutes time‐constrained synchronization operations with non‐time‐constrained control message transmissions to increase the degree of parallelism. Based on the system, we put forward two most frequently used graph applications on Arbor: shortest path and PageRank. In order to evaluate the system, we compare Arbor with the other graph processing systems using large‐scale experimental graph data, and the results show that it outperforms the state‐of‐the‐art systems. Copyright © 2014 John Wiley & Sons, Ltd. Wei Zhou 0019, Jizhong Han, Zhiyong Xu 0003 |
Concurr. Comput. Pract. Exp. | 2 |
| 2014 | PADM: Page Rank-Based Anomaly Detection Method of Log Sequences by Graph ComputingabstractWith the popularity of various software applications in cloud computing, software exception becomes an important issue. How to detect the exceptions more quickly seems to be crucial for the software service company. To solve the above problem, this paper presents an efficient log anomaly detection method named PADM (Page Rank-based Anomaly Detection Method) based on the graph computing algorithm. In this method, the logs are transformed into a graph to represent the complex relationship between the log records, then we design an extended Page Rank algorithm based on the graph to get the importance score for each log. After that, we compare the scores to that of the training logs to determine whether they are abnormal or not. Finally, we compare PADM with other anomaly detection methods on the real logs, and the results show that it outperforms the currently widely used mechanisms with higher accuracy, lower time complexity and better scalability. Xiaoben Yan, Wei Zhou 0019, Jizhong Han, Ge Fu |
CloudCom | 5 |
| 2014 | Performance Evaluation of Light-Weighted Virtualization for PaaS in Clouds
Xuehai Tang, Yifang Wang 0003, Qingqing Feng, Jizhong Han |
ICA3PP (1) | 6 |
| 2014 | A hybrid erasure-coded ECC scheme to improve performance and reliability of solid state drivesabstractThe high performance and ever-increasing capacity of flash memory has led to the rapid adoption of Solid-State Disks (SSDs) in mass storage systems. In order to increase disk capacity, multi-level cells (MLC) are used in the design of SSDs, but the use of such SSDs in persistent storage systems raise concerns for users due to the low reliability of such disks. In this paper, we present a hybrid erasure-coded (EECC) architecture that incorporates ECC schemes and erasure codes to improve both performance and reliability. As weak error-correction codes have faster decoding speed than complex error correction codes (ECC), we propose the use of weak-ECC at the segment level rather than complex ECC. To compensate the reduced correction ability of weak-ECC, we use an erasure code that is striped across segments rather than pages or blocks. We use a small sized HDD to store parities so that we can leverage parallelism across multiple devices and remove the parity updates from the critical write path. We carry out simulation experiments based on Disksim to demonstrate that our proposed scheme is able reduce the SSD average read-latency by up to 31.23% and along with tolerance from double chip failures, it dramatically reduces the uncorrectable page error rate. Pradeep Subedi, Ping Huang 0001, Xubin He, Ming Zhang 0026, Jizhong Han |
IPCCC | 5 |
| 2014 | Marbor: A novel large-scale graph data storage and processing frameworkabstractIn this paper, we propose Marbor, a novel graph data processing framework to analyze the large-scale data in social network services. It develops an efficient graph organization model to minimize the costs of graph data accesses and reduce the memory consumption. In addition, we present a novel control message method in Marbor to improve the synchronization iterations performance. During the graph data processing, in each iteration, it analyzes the relationships among tasks and forwards the tasks to the next iteration with control messages, so no synchronization operations are used. We compare Marbor with other graph processing methods on several large-scale real world SNS datasets with two widely used applications, and the results show that Marbor outperforms the current mechanisms. Wei Zhou 0019, Jizhong Han, Zhiyong Xu 0003 |
IPCCC | 3 |
| 2014 | Online Anomaly Detection by Improved Grammar Compression of Log SequencesabstractNowadays, log sequences mining techniques are widely used in detecting anomalies for Internet services. The state-of-the-art anomaly detection methods either need significant computational costs, or require specific assumptions that the test logs are holding certain data distribution patterns in order to be effective. Therefore, it is very difficult to achieve real time responses and it greatly reduces the effectiveness of these mechanisms in reality. To address these issues, we propose an innovative anomaly detection strategy called CADM. In CADM, the relative entropy between test logs and normal logs is exploited to discover the anomalous levels. Instead of calculating the relative entropy based on certain predefined data distribution models, our solution inspects the relationship between relative entropy and compression size with an improved grammar-based compression method. No assumptions are needed. In addition, our mechanism has excellent scalability with only O(n) computational complexity. It can generate the detection results on the fly. Experimental analysis with both synthetic and real world logs proves that CADM is superior to the other methods. It can achieve very high anomaly detection accuracy with the minimal computational overhead. It is suitable for log mining tasks and can be applied on a broad variety of application fields. Wei Zhou 0019, Jizhong Han, Dan Meng 0002, Zhiyong Xu 0003 |
SDM | 4 |
| 2013 | An adaptive hierarchical caching scheme for remotely sensed databaseabstractWith the rapid growth of users and data volume, the state-of-the-art remotely sensed databases have been posed on grand challenges by the online geospatial applications. They cannot provide low-latency retrieval of big remotely sensed data and fail to adapt for various access patterns from concurrent users. We propose an adaptive hierarchical caching scheme called RSCache for remotely sensed database to overcome the deficiencies. The big data are split into large amounts of small tiles and stored as key-value objects in NoSQL database with Hilbert order. Moreover, a distributed object placement model is designed for different access patterns with consideration of spatio-temporal proximity. Besides, RSCache forms a large memory cache by collecting the exclusive memory pools allocated by cluster nodes and provides global unified view of data manipulation from application's perspective. We implement RSCache middleware prototype on scale-out HBase cluster and confirm its effectiveness with comprehensive experiments in real applications. Yunqin Zhong, Jizhong Han, Jinyun Fang |
IGARSS | 2 |
| 2013 | HDKV: supporting efficient high-dimensional similarity search in key-value storesabstractSUMMARY Key‐value stores are widely used on large‐scale data management in the cloud environment. However, they can only naturally support key‐based queries, and do not have efficient solutions for value‐based queries. Thus, dealing with high‐dimensional data in key‐value stores is still a big challenge. State‐of‐the‐art solutions apply value‐based tree‐structure indexes to solve this issue. These methods suffer from the curse of dimensionality and cannot achieve satisfactory performance. They also bring serious load unbalancing problem among servers, and result in dramatic system scalability degradation. Meanwhile, similarity search in high‐dimensional data space becomes more and more popular in today's cloud applications. Due to the lack of efficient algorithms for value‐based queries, users have to wait for a long time before the results are returned. To address this issue, we propose a novel approach called high‐dimensional similarity query in key‐value stores (HDKV), which can generate similarity results in a short time and maintain good database scalability. In HDKV, a strict order‐preserving hash function is designed to map nearby objects in the high‐dimensional space onto adjacent keys of a continuous linear space in key‐value stores. With this strategy, many expensive random accesses are replaced with more efficient scan accesses. The experimental evaluation on real world data set shows that compared to the state‐of‐the‐art methods, HDKV can dramatically reduce the search time with little impact on the accuracy. Copyright © 2012 John Wiley & Sons, Ltd. Wei Zhou 0019, Jizhong Han, Jiao Dai, Zhiyong Xu 0003 |
Concurr. Comput. Pract. Exp. | 2 |
| 2012 | SDM: A Stripe-Based Data Migration Scheme to Improve the Scalability of RAID-6abstractIn large scale data storage systems, RAID-6 has received more attention due to its capability to tolerate concurrent failures of any two disks, providing a higher level of reliability. However, a challenging issue is its scalability, or how to efficiently expand the disks. The main reason causing this problem is the typical fault tolerant scheme of most RAID-6 systems known as Maximum Distance Separable (MDS) codes, which offer data protection against disk failures with optimal storage efficiency but they are difficult to scale. To address this issue, we propose a novel Stripe-based Data Migration (SDM) scheme for large scale storage systems based on RAID-6 to achieve higher scalability. SDM is a stripe-level scheme, and the basic idea of SDM is optimizing data movements according to the future parity layout, which minimizes the overhead of data migration and parity modification. SDM scheme also provides uniform data distribution, fast data addressing and migration. We have conducted extensive mathematical analysis of applying SDM to various popular RAID-6 coding methods such as RDP, P-Code, H-Code, HDP, X-Code, and EVENODD. The results show that, compared to existing scaling approaches, SDM decreases more than 72.7% migration I/O operations and saves the migration time by up to 96.9%, which speeds up the scaling process by a factor of up to 32. Chentao Wu, Xubin He, Jizhong Han, Huailiang Tan, Changsheng Xie 0001 |
CLUSTER | 3 |
| 2012 | Magicube: High Reliability and Low Redundancy Storage Architecture for Cloud ComputingabstractHigh reliability, high performance and low (space) cost are three important priorities for storage systems. However, it's hard for a cloud storage system to fit them all because they are conflict with each other. Currently, widely used cloud storage systems, such as Google File System (GFS), Hadoop Distributed File System (HDFS) and Amazon's Simple Storage Service (S3), can well meet with high reliability and high performance, but suffer from a huge extra space overhead because of their multi-replication policy. In this paper, we introduce Magicube - a high reliable and low redundancy storage architecture for cloud computing. With only one replica in HDFS, and an (n, k) algorithm for fault-tolerant, it satisfies both low space overhead and high reliability simultaneously. By executing the fault-tolerant process in the background, the performance of Magicube is also good. According to our experiments' result, Magicube can work well for batch processing jobs. Qingqing Feng, Jizhong Han, Dan Meng 0002 |
NAS | 2 |
| 2012 | An Anomaly Detection Algorithm Based on Lossless CompressionabstractAnomaly detection is essential in network security. It has been researched for decades. Many anomaly detection methods have been proposed. Because of the simplicity of principles, statistical and Markovian methods dominate these approaches. However, their effectiveness is constrained by specific preconditions, which make them work for only appropriate data sets which satisfy their premises. Other than statistical and Markovian model, information theory provides a different perspective about anomaly detection. However, the computation of information theoretic measures is still based on statistics. In this paper, we present a novel, information theoretic anomaly detection framework. Instead of statistics, it employs lossless compression for measuring the information quantity, and detects outliers according to compression result. We also discuss the selection of underlying compression algorithm, and choose a grammar compression for utilizing the structure of data. With grammar compression, our method overcomes the shortcomings of statistical and Markovian methods. In addition, the implementation and operation of our method is even simpler than traditional approaches. We test our method on four data sets about text analyzing, host intrusion detection and bug detection. Experimental results show that, even traditional methods fail in some situations, our simple method works well in all cases. Jizhong Han, Jinyun Fang |
NAS | 2 |
| 2012 | A Transparent Control-Flow Based Approach to Record-Replay Non-deterministic BugsabstractRecord-replay is effective to reproduce non-deterministic bugs, and has gained attentions in research community. However, current approaches fall short of handling nondeterministic bugs in multi-processor platforms and distributed systems due to several reasons. First, multi-thread programs on multi-processor platforms, which are common in today's distributed systems, are difficult to be recorded and replayed because of data-races. Second, increasing systems scale makes production environment more sensitive to perturbation from recording. Even hacking control scripts has been unacceptable because of the boosting complexity comes from variety of programs and large number of computing cores. Third, when deployed in distributed systems, large scale will also multiply recording traces, which overwhelms developers, and also slows down the whole system dramatically. To address the above issues, we propose following mechanisms to efficiently record-reply in multi-processor distributed systems: control-flow based record-replay, low-perturbation loading and proportion sampling. We have implemented these mechanisms in ReBranch -- a practical record-replay system for debugging multi-thread programs in multi-processor platforms and distributed systems. ReBranch has already shown its power on dealing with real bugs. We also present our debugging experiences using ReBranch with a case study on handling a bug in memcached -- an important component in many commercial systems. Jizhong Han, Jinyun Fang |
NAS | 2 |
| 2010 | Accelerating Spatial Data Processing with MapReduceabstractMap Reduce is a key-value based programming model and an associated implementation for processing large data sets. It has been adopted in various scenarios and seems promising. However, when spatial computation is expressed straightforward by this key-value based model, difficulties arise due to unfit features and performance degradation. In this paper, we present methods as follows: 1) a splitting method for balancing workload, 2) pending file structure and redundant data partition dealing with relation between spatial objects, 3) a strip-based two-direction plane sweeping algorithm for computation accelerating. Based on these methods, ANN(All nearest neighbors) query and astronomical cross-certification are developed. Performance evaluation shows that the Map Reduce-based spatial applications outperform the traditional one on DBMS. Jizhong Han, Bibo Tu, Jiao Dai, Wei Zhou 0019 |
ICPADS | 2 |
| 2010 | Reproducing non-deterministic bugs with lightweight recording in production environmentsabstractReproducing non-deterministic bugs is challenging. Recording program execution in production environments and reproducing bugs is an effective way to re-enable cyclic debugging. Unfortunately, most current record-replay approaches introduce large perturbations to either environments and/or execution flow, in addition to performance penalty and high storage overhead, which make them impracticable to be deployed in production environments. This paper presents Snitchaser - a fully user-space record-replay tool which can faithfully reproduce bugs by replaying system calls which are recorded with negligible perturbation and recording overhead. This is achieved by 1) a novel, lightweight system call interception mechanism without patching the binary instructions to reduce the perturbation to execution flow; 2) system call latch to save signal semantic; 3) periodic checkpointing to reduce the storage overhead. Snitchaser focuses on bugs caused by asynchronous events on heavily loaded, high throughput servers. Experimental results show that Snitchaser is capable of reproducing non-deterministic bugs efficiently at nearly no performance penalty. We also present two case studies on dealing with existing bugs in Lighttpd - a popular software used in many large scale systems. Jizhong Han, Haiping Fu, Xubin He, Jinyun Fang |
IPCCC | 2 |
| 2010 | Multi-dimensional Index on Hadoop Distributed File SystemabstractIn this paper, we present an approach to construct a built-in block-based hierarchical index structures, like R-tree, to organize data sets in one, two, or higher dimensional space and improve the query performance towards the common query types (e.g., point query, range query) on Hadoop distributed file system (HDFS). The query response time for data sets that are stored in HDFS can be significantly reduced by avoiding exhaustive search on the corresponding data sets in the presence of index structures. The basic idea is to adopt the conventional hierarchical structure to HDFS, and several issues, including index organization, index node size, buffer management, and data transfer protocol, are considered to reduce the query response time and data transfer overhead through network. Experimental evaluation demonstrates that the built-in index structure can efficiently improve query performance, and serve as cornerstones for structured or semi-structured data management. Haojun Liao, Jizhong Han, Jinyun Fang |
NAS | 2 |
| 2009 | Co-match: fast and efficient packet inspection for multiple flowsabstractPacket inspection is widely employed in application-layer protocol analyzing systems to enable accurate protocol identification. Many existing systems, however, fail to meet the requirement of keeping up with wire speed in networking. There are two limitations: (1) software-based matching schemes are usually in a sequential manner which is slow and inefficient; (2) fast hardware-based matching schemes are inapplicable to network packet processing for lacking of intrinsic support for multiple flows. Yingke Xie, Mingshu Wang, Jizhong Han, Chengde Han |
ANCS | 4 |
| 2009 | Implementing WebGIS on Hadoop: A case study of improving small file I/O performance on HDFSabstractHadoop framework has been widely used in various clusters to build large scale, high performance systems. However, Hadoop distributed file system (HDFS) is designed to manage large files and suffers performance penalty while managing a large amount of small files. As a consequence, many web applications, like WebGIS, may not take benefits from Hadoop. In this paper, we propose an approach to optimize I/O performance of small files on HDFS. The basic idea is to combine small files into large ones to reduce the file number and build index for each file. Furthermore, some novel features such as grouping neighboring files and reserving several latest version of data are considered to meet the characteristics of WebGIS access patterns. Preliminary experiment results show that our approach achieves better performance. Xuhui Liu, Jizhong Han, Yunqin Zhong, Chengde Han, Xubin He |
CLUSTER | 2 |
| 2009 | SJMR: Parallelizing spatial join with MapReduce on clustersabstractMapReduce is a widely used parallel programming model and computing platform. With MapReduce, it is very easy to develop scalable parallel programs to process data-intensive applications on clusters of commodity machines. However, it does not directly support heterogeneous related data sets processing, which is common in operations like spatial joins. This paper presents SJMR (Spatial Join with MapReduce), a novel parallel algorithm to relieve the problem. The strategies include strip-based plane sweeping algorithm, tile-based spatial partitioning function and duplication avoidance technology. We evalauted the performance of SJMR algorithm in various situations with the real world data sets. It demonstrates the applicability of computing-intensive spatial applications with MapReduce on small scale clusters. Jizhong Han, Zhiyong Xu 0003 |
CLUSTER | 2 |
| 2009 | Efficient Java Communication Libraries over InfiniBandabstractThis paper presents our current research efforts on efficient Java communication libraries over InfiniBand. The use of Java for network communications still delivers insufficient performance and does not exploit the performance and other special capabilities (RDMA and QoS) of high-speed networks, especially for this interconnect. In order to increase its Java communication performance, InfiniBand has been supported in our high performance sockets implementation, Java Fast Sockets (JFS), and it has been greatly improved the efficiency of Java Direct InfiniBand (Jdib), our low-level communication layer, enabling zero-copy RDMA capability in Java. According to our experimental results, Java communication performance has been improved significantly, reducing start-up latencies from 34 mus down to 12 and 7 mus for JFS and Jdib, respectively, whereas peak bandwidth has been increased from 0.78 Gbps sending serialized data up to 6.7 and 11.2 Gbps for JFS and Jdib, respectively. Finally, it has been analyzed the impact of these communication improvements on parallel Java applications, obtaining significant speedup increases of up to one order of magnitude on 128 cores. Guillermo L. Taboada, Juan Touriño, Ramón Doallo, Jizhong Han |
HPCC | 5 |
| 2009 | uStream: A User-Level Stream Protocol over InfinibandabstractAs one of the most popular high speed networks, InfiniBand demonstrates several enhanced features, such as RDMA and zero-copy mechanisms, which offer high bandwidth and low latency. Communication stacks IPoIB and SDP (Sockets Direct Protocol) have been proposed on InfiniBand for sockets based applications to take advantage of these features. However, these protocols are inefficient to utilize the performance capabilities provided by the physical network. In order to fully exploit the high performance of InfiniBand, we present uStream, a relatively simple and efficient protocol on the user-level socket layer with a stream interface to enable RDMA capability and zero-copy mechanism. Experiment results have shown that the performance of uStream is comparable to raw InfiniBand Verbs/RDMA interface, and uStream outperforms SDP with 7.9 ¿s minimum latency and 10.4 Gbps peak bandwidth in our testbed. In addition, a Java communication library, jStream, is built upon uStream to enable Java to use RDMA directly. Preliminary results have shown Java clusters communication performance has been improved to a level similar to C programs over InfiniBand through combined uStream and jStream. Jizhong Han, Jinjun Gao, Xubin He |
ICPADS | 2 |
| 2009 | Accelerating MapReduce with Distributed Memory CacheabstractMapReduce is a partition-based parallel programming model and framework enabling easy development of scalable parallel programs on clusters of commodity machines. In order to make time-intensive applications benefit from MapReduce on small scale clusters, this paper proposes a new method to improve the performance of MapReduce by using distributed memory cache as a high speed access between map tasks and reduce tasks. Map outputs sent to the distributed memory cache can be gotten by reduce tasks as soon as possible. Experiment results show that our prototype’s performance is much better than that of the original on small scale clusters. To our knowledge, this is the first effort to accelerate MapReduce with the help of distributed memory cache. Jizhong Han, Shengzhong Feng |
ICPADS | 2 |
| 2009 | An efficient design for fast memory registration in RDMA
Li Ou, Xubin He, Jizhong Han |
J. Netw. Comput. Appl. | 3 |
| 2008 | PROD: Relayed file retrieving in overlay networksabstractTo share and exchange the files among Internet users, peer-to-peer (P2P) applications build another layer of overlay networks on top of the Internet infrastructure. In P2P file sharing systems, a file request takes two steps. First, a routing message is generated by the client (request initiator) and spread to the overlay network. After the process finishes, the location information of the requested file is returned to the client. In the second step, the client establishes direct connection(s) with the peer(s) who store a copy of that file to start the retrieving process. While numerous research projects have been conducted to design efficient, high-performance routing algorithms, few work concentrated on file retrieving performance. In this paper, we propose a novel and efficient algorithm - PROD to improve the file retrieving performance in DHT based overlay networks. In PROD, when a file or a portion of a file is transferred from a source peer to the client, instead of creating just one direct link between these two peers, we build an application level connection chain. Along the chain, multiple network links are established. Each intermediate peer on this chain uses a store-and-forward mechanism for the data transfer. PROD also introduces a novel topological based strategy to choose these peers and guarantees the transmission delay of each intermediate link is much lower than the direct link. We conducted extensive simulation experiments and the results shown that PROD can greatly reduce the transfer time per file in DHT base P2P systems. Zhiyong Xu 0003, Dan Stefanescu, Laxmi N. Bhuyan, Jizhong Han |
IPDPS | 5 |
| 2007 | Collaborative Memory Pool in Cluster SystemabstractWith the developments of network technologies, many mechanisms have been introduced to improve system performance in cluster systems by exploiting remote idle memory. However, none of them can satisfy the requirements from different applications. Most methods can only improve the performance of a particular type of applications but not for others. One important reason is they failed to provide unified interfaces. In this paper, we propose collaborative memory pool (CMP) to solve this problems. CMP brings scalability and high performance. It has five features: (1) Providing malloc-like interfaces, block device interfaces and kernel API for different applications, which benefit both user-level and kernel-level applications; (2) Retaining traditional VM mechanism, programmers and uses have the freedom to select CMP or not; (3) Improving kernel applications performance by eliminating remote swapping; (4) Avoiding loan while in debt problem with dynamic workload; (5) Providing optional memory servers to further improve performance. In our testbed with CMP-based swap devices, Qsort gets 83.28% improvement comparing with the case using disk-based swap devices. Xuhui Liu, Jizhong Han, Lisheng Zhang, Zhiyong Xu 0003 |
ICPP | 4 |
| 2007 | Scalable and Decentralized Content-Aware Dispatching in Web ClustersabstractIn this paper, we propose a novel and efficient content-aware dispatching algorithm. Our approach eliminates the potential bottleneck and the single point of failure problems completely by using totally decentralized P2P architecture. It is scalable, the system throughput increases nearly linearly with the increased number of servers. Meanwhile, it does not introduce heavy communication overhead among back-end servers which appeared in the previous decentralized mechanisms. Our simulation results show that our approach is superior to the previous solutions. Zhiyong Xu 0003, Jizhong Han, Laxmi N. Bhuyan |
IPCCC | 2 |
| 2007 | An SRP Target Mode to Improve Read Performance of SRP-Based IB-SANs
Zhiying Jiang, Jizhong Han, Xigui Wang, Yonghao Zhou, Xubin He |
ISPA | 3 |
| 2006 | Block-Level Storage Security Architectures
Jizhong Han, Zhensong Wang |
ICCSA (1) | 2 |