EDBT 2026 Demo / reviewers in the wild / expert
Fan Yin
dblp:24/8079
· DBLP profile ↗
27ranked-venue papers
10as first author
24since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 9 first-author · 18 since 2021Security and privacy · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Magnet: Multi-turn Tool-use Data Synthesis and Distillation via Graph TranslationabstractFan Yin, Zifeng Wang, I-Hung Hsu, Jun Yan, Ke Jiang, Yanfei Chen, Jindong Gu, Long Le, Kai-Wei Chang, Chen-Yu Lee, Hamid Palangi, Tomas Pfister. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Fan Yin, Zifeng Wang 0002, I-Hung Hsu, Jun Yan 0001, Yanfei Chen, Jindong Gu, Long T. Le, Kai-Wei Chang 0001, Chen-Yu Lee, Hamid Palangi, Tomas Pfister |
ACL (1) | 1 |
| 2025 | DUNE: Sim2Real Transfer for Depth-based Navigation in Unstructured Dynamic Indoor EnvironmentsabstractCollision-free navigation in dynamic environments, especially with moving pedestrians, is crucial for mobile robots. This paper introduces DUNE, a depth-based policy trained in simulation for collision-free navigation of Ackermann mobile robots in unstructured indoor environments. DUNE uses a CNN-LSTM to encode depth vision and past actions, while an actor-critic network controls velocity and steering. Additionally, a depth filter bridges the sim-to-real gap. Unlike prior works relying on LIDAR and complex algorithms, DUNE achieves state-of-the-art performance with only egocentric depth perception and lightweight neural networks, both in simulation and real-world tasks. Chaoyi Xu, Jilong Wang 0016, Fan Yin |
ICASSP | 5 |
| 2025 | BingoGuard: LLM Content Moderation Tools with Risk LevelsabstractMalicious content generated by large language models (LLMs) can pose varying degrees of harm.
Although existing LLM-based moderators can detect harmful content, they struggle to assess risk levels and may miss lower-risk outputs.
Accurate risk assessment allows platforms with different safety thresholds to tailor content filtering and rejection. In this paper, we introduce per-topic severity rubrics for 11 harmful topics and build BingoGuard, an LLM-based moderation system designed to predict both binary safety labels and severity levels.
To address the lack of annotations on levels of severity, we propose a scalable generate-then-filter framework that first generates responses across different severity levels and then filters out low-quality responses. Using this framework, we create BingoGuardTrain, a training dataset with 54,897 examples covering a variety of topics, response severity, styles, and BingoGuardTest, a test set with 988 examples explicitly labeled based on our severity rubrics that enables fine-grained analysis on model behaviors on different severity levels. Our BingoGuard-8B, trained on BingoGuardTrain, achieves the state-of-the-art performance on several moderation benchmarks, including WildGuardTest and HarmBench, as well as BingoGuardTest, outperforming best public models, WildGuard, by 4.3\%. Our analysis demonstrates that incorporating severity levels into training significantly enhances detection performance and enables the model to effectively gauge the severity of harmful responses. Warning: this paper includes red-teaming examples that may be harmful in nature. Fan Yin, Philippe Laban, Yilun Zhou, Yixin Mao, Vaibhav Vats, Linnea Ross, Divyansh Agarwal, Caiming Xiong, Chien-Sheng Wu |
ICLR | 1 |
| 2025 | Short-Term Prediction for Waste Heat Recovery Power Generation Based on Crossformer Network in Coke Dry Quenching ProcessabstractUtilizing the waste heat recovered from the coke dry quenching process for power generation is an important energy-saving measure for steel enterprises. However, the discontinuous nature of the coke discharging operation leads to instability in waste heat supply and constant load fluctuations in the process of recovering the sensible heat from red hot coke, thereby affecting the power generation. Consequently, short-term prediction of power generation in coke dry quenching process is essential to provide guidance for the power scheduling of steel companies. In this paper, a prediction model is developed based on the Crossformer network. Experimental results based on real production data show that the proposed prediction model can relatively accurately predict power generation in the coke dry quenching process, which provides guidance for the short-term power scheduling of the steel company. Xiaochong Chen, Fan Yin, Yiheng Chen, Jie Hu 0013, Jundong Wu, Min Wu 0002 |
IECON | 2 |
| 2025 | OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL CyclesabstractWe introduce *OpenVLThinker*, one of the first open-source large vision–language models (LVLMs) to exhibit sophisticated chain-of-thought reasoning, achieving notable performance gains on challenging visual reasoning tasks. While text-based reasoning models (e.g., Deepseek R1) show promising results in text-only tasks, distilling their reasoning into LVLMs via supervised fine-tuning (SFT) often results in performance degradation due to imprecise visual grounding. Conversely, purely reinforcement learning (RL)-based methods face a large search space, hindering the emergence of reflective behaviors in smaller models (e.g., 7B LVLMs). Surprisingly, alternating between SFT and RL ultimately results in significant performance improvements after a few iterations. Our analysis reveals that the base model rarely exhibits reasoning behaviors initially, but SFT effectively surfaces these latent actions and narrows the RL search space, accelerating the development of reasoning capabilities. Each subsequent RL stage further refines the model's reasoning skills, producing higher-quality SFT data for continued self-improvement. OpenVLThinker-7B consistently advances performance across six benchmarks demanding mathematical and general reasoning, notably improving MathVista by 3.2\%, EMMA by 1.4\%, and HallusionBench by 2.7\%. Beyond demonstrating the synergy between SFT and RL for complex reasoning tasks, our findings provide early evidence towards achieving R1-style reasoning in multimodal contexts. Yihe Deng, Hritik Bansal, Fan Yin, Nanyun Peng 0001, Wei Wang 0010, Kai-Wei Chang 0001 |
NeurIPS | 3 |
| 2024 | Synchronous Faithfulness Monitoring for Trustworthy Retrieval-Augmented GenerationabstractRetrieval-augmented language models (RALMs) have shown strong performance and wide applicability in knowledge-intensive tasks.However, there are significant trustworthiness concerns as RALMs are prone to generating unfaithful outputs, including baseless information or contradictions with the retrieved context.This paper proposes SYNCHECK, a lightweight monitor that leverages fine-grained decoding dynamics including sequence likelihood, uncertainty quantification, context influence, and semantic alignment to synchronously detect unfaithful sentences.By integrating efficiently measurable and complementary signals, SYNCHECK enables accurate and immediate feedback and intervention, achieving 0.85 AUROC in detecting faithfulness errors across six long-form retrieval-augmented generation tasks, improving prior best method by 4%.Leveraging SYNCHECK, we further introduce FOD, a faithfulness-oriented decoding algorithm guided by beam search for long-form retrieval-augmented generation.Empirical results demonstrate that FOD outperforms traditional strategies such as abstention, reranking, or contrastive decoding significantly in terms of faithfulness, achieving over 10% improvement across six datasets. Di Wu 0054, Jia-Chen Gu, Fan Yin, Nanyun Peng 0001, Kai-Wei Chang 0001 |
EMNLP | 3 |
| 2024 | Characterizing Truthfulness in Large Language Model Generations with Local Intrinsic DimensionabstractWe study how to characterize and predict the truthfulness of texts generated from large language models (LLMs), which serves as a crucial step in building trust between humans and LLMs. Although several approaches based on entropy or verbalized uncertainty have been proposed to calibrate model predictions, these methods are often intractable, sensitive to hyperparameters, and less reliable when applied in generative tasks with LLMs. In this paper, we suggest investigating internal activations and quantifying LLM’s truthfulness using the local intrinsic dimension (LID) of model activations. Through experiments on four question answering (QA) datasets, we demonstrate the effectiveness of our proposed method. Additionally, we study intrinsic dimensions in LLMs and their relations with model layers, autoregressive language modeling, and the training of LLMs, revealing that intrinsic dimensions can be a powerful approach to understanding LLMs. Fan Yin, Jayanth Srinivasa, Kai-Wei Chang 0001 |
ICML | 1 |
| 2024 | On Prompt-Driven Safeguarding for Large Language ModelsabstractPrepending model inputs with safety prompts is a common practice for safeguarding large language models (LLMs) against queries with harmful intents. However, the underlying working mechanisms of safety prompts have not been unraveled yet, restricting the possibility of automatically optimizing them to improve LLM safety. In this work, we investigate how LLMs’ behavior (i.e., complying with or refusing user queries) is affected by safety prompts from the perspective of model representation. We find that in the representation space, the input queries are typically moved by safety prompts in a "higher-refusal" direction, in which models become more prone to refusing to provide assistance, even when the queries are harmless. On the other hand, LLMs are naturally capable of distinguishing harmful and harmless queries without safety prompts. Inspired by these findings, we propose a method for safety prompt optimization, namely DRO (Directed Representation Optimization). Treating a safety prompt as continuous, trainable embeddings, DRO learns to move the queries’ representations along or opposite the refusal direction, depending on their harmfulness. Experiments with eight LLMs on out-of-domain and jailbreak benchmarks demonstrate that DRO remarkably improves the safeguarding performance of human-crafted safety prompts, without compromising the models’ general performance. Chujie Zheng, Fan Yin, Hao Zhou 0012, Fandong Meng, Jie Zhou 0016, Kai-Wei Chang 0001, Minlie Huang, Nanyun Peng 0001 |
ICML | 2 |
| 2024 | SAMS : One-Shot Learning for the Segment Anything Model Using Similar ImagesabstractThe field of computer vision is currently transitioning from closed-set to open-set tasks. Vision foundation models have already demonstrated success in open-set scenarios. Building on these models, the utilization of a feature supervision framework can further enhance results. Our paper introduces a new method called SAMS (Segment Anything Model using Similar Images), which is a type of feature supervision framework. It is designed to segment specific masks from visual supervision features. Our framework comprises a pre-trained segmentation model and an efficient, novel prompt generation model capable of generating new prompts based on pre-extracted image features. This innovation eliminates the need for manually crafted prompts in the mask generation phase by integrating the principles of one-shot or few-shot learning with visual instructions from similar images. The effectiveness of the SAMS method is evident in its performance across various tasks, particularly in open-set tasks where traditional models tend to struggle. The pretrained model not only achieves impressive mean Intersection over Union (mIOU) scores without incurring additional time loss, but also demonstrates potential for further improvement through targeted module training. Fan Yin, Yifei Wei, Chaoyi Xu |
IJCNN | 1 |
| 2024 | Enhancing Large Vision Language Models with Self-Training on Image ComprehensionabstractLarge vision language models (LVLMs) integrate large language models (LLMs) with pre-trained vision encoders, thereby activating the perception capability of the model to understand image inputs for different queries and conduct subsequent reasoning. Improving this capability requires high-quality vision-language data, which is costly and labor-intensive to acquire. Self-training approaches have been effective in single-modal settings to alleviate the need for labeled data by leveraging model's own generation. However, effective self-training remains a challenge regarding the unique visual perception and reasoning capability of LVLMs. To address this, we introduce **S**elf-**T**raining on **I**mage **C**omprehension (**STIC**), which emphasizes a self-training approach specifically for image comprehension. First, the model self-constructs a preference dataset for image descriptions using unlabeled images. Preferred responses are generated through a step-by-step prompt, while dis-preferred responses are generated from either corrupted images or misleading prompts. To further self-improve reasoning on the extracted visual information, we let the model reuse a small portion of existing instruction-tuning data and append its self-generated image descriptions to the prompts. We validate the effectiveness of STIC across seven different benchmarks, demonstrating substantial performance gains of 4.0% on average while using 70% less supervised fine-tuning data than the current method. Further studies dive into various components of STIC and highlight its potential to leverage vast quantities of unlabeled images for self-training. Yihe Deng, Pan Lu, Fan Yin, Ziniu Hu, Sheng Shen 0001, Quanquan Gu, James Zou 0001, Kai-Wei Chang 0001, Wei Wang 0010 |
NeurIPS | 3 |
| 2024 | Red Teaming Language Model Detectors with Language ModelsabstractAbstract The prevalence and strong capability of large language models (LLMs) present significant safety and ethical risks if exploited by malicious users. To prevent the potentially deceptive usage of LLMs, recent work has proposed algorithms to detect LLM-generated text and protect LLMs. In this paper, we investigate the robustness and reliability of these LLM detectors under adversarial attacks. We study two types of attack strategies: 1) replacing certain words in an LLM’s output with their synonyms given the context; 2) automatically searching for an instructional prompt to alter the writing style of the generation. In both strategies, we leverage an auxiliary LLM to generate the word replacements or the instructional prompt. Different from previous works, we consider a challenging setting where the auxiliary LLM can also be protected by a detector. Experiments reveal that our attacks effectively compromise the performance of all detectors in the study with plausible generations, underscoring the urgent need to improve the robustness of LLM-generated text detection systems. Code is available at https://github.com/shizhouxing/LLM-Detector-Robustness. Zhouxing Shi, Fan Yin, Xiangning Chen, Kai-Wei Chang 0001, Cho-Jui Hsieh |
Trans. Assoc. Comput. Linguistics | 3 |
| 2023 | Efficient Shapley Values Estimation by Amortization for Text ClassificationabstractDespite the popularity of Shapley Values in explaining neural text classification models, computing them is prohibitive for large pretrained models due to a large number of model evaluations.In practice, Shapley Values are often estimated with a small number of stochastic model evaluations.However, we show that the estimated Shapley Values are sensitive to random seed choices -the top-ranked features often have little overlap across different seeds, especially on examples with longer input texts.This can only be mitigated by aggregating thousands of model evaluations, which on the other hand, induces substantial computational overheads.To mitigate the trade-off between stability and efficiency, we develop an amortized model that directly predicts each input feature's Shapley Value without additional model evaluations.It is trained on a set of examples whose Shapley Values are estimated from a large number of model evaluations to ensure stability.Experimental results on two text classification datasets demonstrate that our amortized model estimates Shapley Values accurately with up to 60 times speedup compared to traditional methods.Furthermore, the estimated values are stable as the inference is deterministic.We release our code at https://github.com/yangalan123/ Amortized-Interpretability. Chenghao Yang 0001, Fan Yin, He He 0001, Kai-Wei Chang 0001, Xiaofei Ma 0001, Bing Xiang |
ACL (1) | 2 |
| 2023 | Did You Read the Instructions? Rethinking the Effectiveness of Task Definitions in Instruction LearningabstractLarge language models (LLMs) have shown impressive performance in following natural language instructions to solve unseen tasks.However, it remains unclear whether models truly understand task definitions and whether the human-written definitions are optimal.In this paper, we systematically study the role of task definitions in instruction learning.We first conduct an ablation analysis informed by human annotations to understand which parts of a task definition are most important, and find that model performance only drops substantially when removing contents describing the task output, in particular label information.Next, we propose an automatic algorithm to compress task definitions to a minimal supporting set of tokens, and find that 60% of tokens can be removed while maintaining or even improving model performance.Based on these results, we propose two strategies to help models better leverage task instructions: (1) providing only key information for tasks in a common structured format, and (2) adding a metatuning stage to help the model better understand the definitions.With these two strategies, we achieve a 4.2 Rouge-L improvement over 119 unseen test tasks. Fan Yin, Jesse Vig, Philippe Laban, Shafiq R. Joty, Caiming Xiong, Chien-Sheng Wu |
ACL (1) | 1 |
| 2023 | Active Instruction Tuning: Improving Cross-Task Generalization by Training on Prompt Sensitive TasksabstractInstruction tuning (IT) achieves impressive zero-shot generalization results by training large language models (LLMs) on a massive amount of diverse tasks with instructions.However, how to select new tasks to improve the performance and generalizability of IT models remains an open question.Training on all existing tasks is impractical due to prohibiting computation requirements, and randomly selecting tasks can lead to suboptimal performance.In this work, we propose active instruction tuning based on prompt uncertainty, a novel framework to identify informative tasks, and then actively tune the models on the selected tasks.We represent the informativeness of new tasks with the disagreement of the current model outputs over perturbed prompts.Our experiments on NIV2 and Self-Instruct datasets demonstrate that our method consistently outperforms other baseline strategies for task selection, achieving better out-of-distribution generalization with fewer training tasks.Additionally, we introduce a task map that categorizes and diagnoses tasks based on prompt uncertainty and prediction probability.We discover that training on ambiguous (prompt-uncertain) tasks improves generalization while training on difficult (prompt-certain and low-probability) tasks offers no benefit, underscoring the importance of task selection for instruction tuning. 1 Po-Nien Kung, Fan Yin, Di Wu 0054, Kai-Wei Chang 0001, Nanyun Peng 0001 |
EMNLP | 2 |
| 2023 | Dynosaur: A Dynamic Growth Paradigm for Instruction-Tuning Data CurationabstractInstruction tuning has emerged to enhance the capabilities of large language models (LLMs) to comprehend instructions and generate appropriate responses.Existing methods either manually annotate or employ LLM (e.g., GPTseries) to generate data for instruction tuning.However, they often overlook associating instructions with existing annotated datasets.In this paper, we propose DYNOSAUR, a dynamic growth paradigm for the automatic curation of instruction-tuning data.Based on the metadata of existing datasets, we use LLMs to automatically construct instruction-tuning data by identifying relevant data fields and generating appropriate instructions.By leveraging the existing annotated datasets, DYNOSAUR offers several advantages: 1) it reduces the API cost for generating instructions (e.g., it costs less than $12 USD by calling GPT-3.5-turbo for generating 800K instruction tuning samples; 2) it provides high-quality data for instruction tuning (e.g., it performs better than ALPACA and FLAN on SUPER-NI and LONGFORM with comparable data sizes); and 3) it supports the continuous improvement of models by generating instruction-tuning data when a new annotated dataset becomes available.We further investigate a continual learning scheme for learning with the ever-growing instruction-tuning dataset, and demonstrate that replaying tasks with diverse instruction embeddings not only helps mitigate forgetting issues but generalizes to unseen tasks better. Da Yin, Xiao Liu 0032, Fan Yin, Ming Zhong 0005, Hritik Bansal, Jiawei Han 0001, Kai-Wei Chang 0001 |
EMNLP | 3 |
| 2023 | CleanCLIP: Mitigating Data Poisoning Attacks in Multimodal Contrastive LearningabstractMultimodal contrastive pretraining has been used to train multimodal representation models, such as CLIP, on large amounts of paired image-text data. However, previous studies have revealed that such models are vulnerable to backdoor attacks. Specifically, when trained on backdoored examples, CLIP learns spurious correlations between the embedded backdoor trigger and the target label, aligning their representations in the joint embedding space. Injecting even a small number of poisoned examples, such as 75 examples in 3 million pretraining data, can significantly manipulate the model’s behavior, making it difficult to detect or unlearn such correlations. To address this issue, we propose CleanCLIP, a finetuning framework that weakens the learned spurious associations introduced by backdoor attacks by independently re-aligning the representations for individual modalities. We demonstrate that unsupervised finetuning using a combination of multimodal contrastive and unimodal self-supervised objectives for individual modalities can significantly reduce the impact of the backdoor attack. Additionally, we show that supervised finetuning on task-specific labeled image data removes the backdoor trigger from the CLIP vision encoder. We show empirically that CleanCLIP maintains model performance on benign examples while erasing a range of backdoor attacks on multimodal contrastive learning. Code and pretrained checkpoints are available at https://github.com/nishadsinghi/CleanCLIP. Hritik Bansal, Fan Yin, Nishad Singhi, Aditya Grover, Kai-Wei Chang 0001 |
ICCV | 2 |
| 2023 | Should I Stop or Should I Go: Early Stopping with Heterogeneous PopulationsabstractRandomized experiments often need to be stopped prematurely due to the treatment having an unintended harmful effect. Existing methods that determine when to stop an experiment early are typically applied to the data in aggregate and do not account for treatment effect heterogeneity. In this paper, we study the early stopping of experiments for harm on heterogeneous populations. We first establish that current methods often fail to stop experiments when the treatment harms a minority group of participants. We then use causal machine learning to develop CLASH, the first broadly-applicable method for heterogeneous early stopping. We demonstrate CLASH's performance on simulated and real data and show that it yields effective early stopping for both clinical trials and A/B tests. Hammaad Adam, Fan Yin, Huibin Hu, Neil A. Tenenholtz, Lorin Crawford, Lester Mackey, Allison Koenecke |
NeurIPS | 2 |
| 2022 | On the Sensitivity and Stability of Model Interpretations in NLPabstractRecent years have witnessed the emergence of a variety of post-hoc interpretations that aim to uncover how natural language processing (NLP) models make predictions.Despite the surge of new interpretation methods, it remains an open problem how to define and quantitatively measure the faithfulness of interpretations, i.e., to what extent interpretations reflect the reasoning process by a model.We propose two new criteria, sensitivity and stability, that provide complementary notions of faithfulness to the existed removal-based criteria.Our results show that the conclusion for how faithful interpretations are could vary substantially based on different notions.Motivated by the desiderata of sensitivity and stability, we introduce a new class of interpretation methods that adopt techniques from adversarial robustness.Empirical results show that our proposed methods are effective under the new criteria and overcome limitations of gradient-based methods on removal-based criteria.Besides text classification, we also apply interpretation methods and metrics to dependency parsing.Our results shed light on understanding the diverse set of interpretations. Fan Yin, Zhouxing Shi, Cho-Jui Hsieh, Kai-Wei Chang 0001 |
ACL (1) | 1 |
| 2022 | ADDMU: Detection of Far-Boundary Adversarial Examples with Data and Model Uncertainty EstimationabstractAdversarial Examples Detection (AED) is a crucial defense technique against adversarial attacks and has drawn increasing attention from the Natural Language Processing (NLP) community.Despite the surge of new AED methods, our studies show that existing methods heavily rely on a shortcut to achieve good performance.In other words, current search-based adversarial attacks in NLP stop once model predictions change, and thus most adversarial examples generated by those attacks are located near model decision boundaries.To surpass this shortcut and fairly evaluate AED methods, we propose to test AED methods with Far Boundary (FB) adversarial examples.Existing methods show worse than random guess performance under this scenario.To overcome this limitation, we propose a new technique, ADDMU, adversary detection with data and model uncertainty, which combines two types of uncertainty estimation for both regular and FB adversarial example detection.Our new method outperforms previous methods by 3.6 and 6.0 AUC points under each scenario.Finally, our analysis shows that the two types of uncertainty provided by ADDMU can be leveraged to characterize adversarial examples and identify the ones that contribute most to model's robustness in adversarial training. Fan Yin, Yao Li 0015, Cho-Jui Hsieh, Kai-Wei Chang 0001 |
EMNLP | 1 |
| 2022 | Bayesian Nonparametric Learning for Point Processes with Spatial Homogeneity: A Spatial Analysis of NBA Shot LocationsabstractBasketball shot location data provide valuable summary information regarding players to coaches, sports analysts, fans, statisticians, as well as players themselves. Represented by spatial points, such data are naturally analyzed with spatial point process models. We present a novel nonparametric Bayesian method for learning the underlying intensity surface built upon a combination of Dirichlet process and Markov random field. Our method has the advantage of effectively encouraging local spatial homogeneity when estimating a globally heterogeneous intensity surface. Posterior inferences are performed with an efficient Markov chain Monte Carlo (MCMC) algorithm. Simulation studies show that the inferences are accurate and the method is superior compared to a wide range of competing methods. Application to the shot location data of $20$ representative NBA players in the 2017-2018 regular season offers interesting insights about the shooting patterns of these players. A comparison against the competing method shows that the proposed method can effectively incorporate spatial contiguity into the estimation of intensity surfaces. Fan Yin, Jieying Jiao, Guanyu Hu 0002 |
ICML | 1 |
| 2022 | Differentially private hierarchical tree with high efficiency
Hui Zhu 0006, Fan Yin, Shuangrong Peng, Xiaohu Tang 0004 |
Comput. Secur. | 2 |
| 2022 | Achieving Efficient and Privacy-Preserving Cross-Domain Big Data Deduplication in CloudabstractSecure data deduplication can significantly reduce the communication and storage overheads in cloud storage services, and has potential applications in our big data-driven society. Existing data deduplication schemes are generally designed to either resist brute-force attacks or ensure the efficiency and data availability, but not both conditions. We are also not aware of any existing scheme that achieves accountability, in the sense of reducing duplicate information disclosure (e.g., to determine whether plaintexts of two encrypted messages are identical). In this paper, we investigate a three-tier cross-domain architecture, and propose an efficient and privacy-preserving big data deduplication in cloud storage (hereafter referred to as EPCDD). EPCDD achieves both privacy-preserving and data availability, and resists brute-force attacks. In addition, we take accountability into consideration to offer better privacy assurances than existing schemes. We then demonstrate that EPCDD outperforms existing competing schemes, in terms of computation, communication and storage overheads. In addition, the time complexity of duplicate search in EPCDD is logarithmic. Xue Yang 0003, Rongxing Lu, Kim-Kwang Raymond Choo, Fan Yin, Xiaohu Tang 0004 |
IEEE Trans. Big Data | 4 |
| 2022 | Achieving Practical Symmetric Searchable Encryption With Search Pattern Privacy Over CloudabstractDynamic symmetric searchable encryption (SSE), which enables a data user to securely search and dynamically update the encrypted documents stored in a semi-trusted cloud server, has received considerable attention in recent years. However, the search and update operations in many previously reported SSE schemes will bring some additional privacy leakages, e.g., search pattern privacy, forward privacy and backward privacy. To the best of our knowledge, none of the existing dynamic SSE schemes preserves the search pattern privacy, and many backward private SSE schemes still leak some critical information, e.g., the identifiers containing a specific keyword currently in the database. Therefore, aiming at the above challenges, in this article, we design a practical SSE scheme, which not only supports the search pattern privacy but also enhances the backward privacy. Specifically, we first leverage the$k$-anonymity and encryption to design an obfuscating technique. Then, based on the obfuscating technique, pseudorandom function and pseudorandom generator, we design a basic dynamic SSE scheme to support single keyword queries and simultaneously achieve search pattern privacy and enhanced backward privacy. Furthermore, we also extend our proposed scheme to support more efficient boolean queries. Security analysis demonstrates that our proposed scheme can achieve the desired privacy properties, and the extensive performance evaluations also show that our proposed scheme is indeed efficient in terms of communication overhead and computational cost. Yandong Zheng, Rongxing Lu, Jun Shao 0001, Fan Yin, Hui Zhu 0001 |
IEEE Trans. Serv. Comput. | 4 |
| 2021 | Achieve efficient position-heap-based privacy-preserving substring-of-keyword query over cloud
Fan Yin, Rongxing Lu, Yandong Zheng, Jun Shao 0001, Xue Yang 0003, Xiaohu Tang 0004 |
Comput. Secur. | 1 |
| 2020 | On the Robustness of Language Encoders against Grammatical ErrorsabstractWe conduct a thorough study to diagnose the behaviors of pre-trained language encoders (ELMo, BERT, and RoBERTa) when confronted with natural grammatical errors.Specifically, we collect real grammatical errors from non-native speakers and conduct adversarial attacks to simulate these errors on clean text data.We use this approach to facilitate debugging models on downstream applications.Results confirm that the performance of all tested models is affected but the degree of impact varies.To interpret model behaviors, we further design a linguistic acceptability task to reveal their abilities in identifying ungrammatical sentences and the position of errors.We find that fixed contextual encoders with a simple classifier trained on the prediction of sentence correctness are able to locate error positions.We also design a cloze test for BERT and discover that BERT captures the interaction between errors and specific tokens in context.Our results shed light on understanding the robustness and behaviors of language encoders against grammatical errors. Fan Yin, Quanyu Long, Kai-Wei Chang 0001 |
ACL | 1 |
| 2019 | Entity-Relation Extraction as Multi-Turn Question AnsweringabstractIn this paper, we propose a new paradigm for the task of entity-relation extraction.We cast the task as a multi-turn question answering problem, i.e., the extraction of entities and relations is transformed to the task of identifying answer spans from the context.This multi-turn QA formalization comes with several key advantages: firstly, the question query encodes important information for the entity/relation class we want to identify; secondly, QA provides a natural way of jointly modeling entity and relation; and thirdly, it allows us to exploit the well developed machine reading comprehension (MRC) models.Experiments on the ACE and the CoNLL04 corpora demonstrate that the proposed paradigm significantly outperforms previous best models.We are able to obtain the stateof-the-art results on all of the ACE04, ACE05 and CoNLL04 datasets, increasing the SOTA results on the three datasets to 49.4 (+1.0), 60.2 (+0.6) and 68.9 (+2.1), respectively.Additionally, we construct a newly developed dataset RESUME in Chinese, which requires multi-step reasoning to construct entity dependencies, as opposed to the single-step dependency extraction in the triplet exaction in previous datasets.The proposed multi-turn QA model also achieves the best performance on the RESUME dataset. 1 Xiaoya Li 0001, Fan Yin, Zijun Sun, Xiayu Li, Arianna Yuan, Duo Chai, Mingxin Zhou, Jiwei Li 0001 |
ACL (1) | 2 |
| 2019 | Glyce: Glyph-vectors for Chinese Character RepresentationsabstractIt is intuitive that NLP tasks for logographic languages like Chinese should benefit from the use of the glyph information in those languages. However, due to the lack of rich pictographic evidence in glyphs and the weak generalization ability of standard computer vision models on character data, an effective way to utilize the glyph information remains to be found. In this paper, we address this gap by presenting Glyce, the glyph-vectors for Chinese character representations. We make three major innovations: (1) We use historical Chinese scripts (e.g., bronzeware script, seal script, traditional Chinese, etc) to enrich the pictographic evidence in characters; (2) We design CNN structures (called tianzege-CNN) tailored to Chinese character image processing; and (3) We use image-classification as an auxiliary task in a multi-task learning setup to increase the model's ability to generalize. We show that glyph-based models are able to consistently outperform word/char ID-based models in a wide range of Chinese NLP tasks. When combing with BERT, we are able to set new state-of-the-art results for a variety of Chinese NLP tasks, including language modeling, tagging (NER, CWS, POS), sentence pair classification (BQ, LCQMC, XNLI, NLPCC-DBQA), single sentence classification tasks (ChnSentiCorp, the Fudan corpus, iFeng), dependency parsing, and semantic role labeling. For example, the proposed model achieves an F1 score of 81.6 on the OntoNotes dataset of NER, +1.5 over BERT; it achieves an almost perfect accuracy of 99.8\% on the the Fudan corpus for text classification. Yuxian Meng, Wei Wu 0044, Fei Wang 0060, Xiaoya Li 0001, Ping Nie, Fan Yin, Muyu Li, Qinghong Han, Xiaofei Sun 0001, Jiwei Li 0001 |
NeurIPS | 6 |