Yukun Yan

dblp:206/7211 · DBLP profile ↗
← Back
29ranked-venue papers
4as first author
28since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 1 first-author · 19 since 2021Databases, data management, data science and information retrieval · 8 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Chunks as Arms: Multi-Armed Bandit-Guided Sampling for Long-Context LLM Preference Optimization
abstract
Shaohua Duan, Pengcheng Huang, Xinze Li, Zhenghao Liu, Xiaoyuan Yi, Yukun Yan, Shuo Wang, Yu Gu, Ge Yu, Maosong Sun. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Shaohua Duan, Pengcheng Huang 0004, Zhenghao Liu 0001, Xiaoyuan Yi, Yukun Yan, Shuo Wang 0013, Yu Gu 0002, Ge Yu 0001, Maosong Sun 0001
ACL (1)6
2026 Empirical Analysis of Decoding Biases in Masked Diffusion Models
abstract
Pengcheng Huang, Tianming Liu, Zhenghao Liu, Yukun Yan, Shuo Wang, Tong Xiao, Zulong Chen, Maosong Sun. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Pengcheng Huang 0004, Zhenghao Liu 0001, Yukun Yan, Shuo Wang 0013, Tong Xiao 0001, Zulong Chen, Maosong Sun 0001
ACL (1)4
2026 Long-Chain Reasoning Distillation via Adaptive Prefix Alignment
abstract
Zhenghao Liu, Zhuoyang Wu, Xinze Li, Yukun Yan, Shuo Wang, Zulong Chen, Yu Gu, Ge Yu, Maosong Sun. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zhenghao Liu 0001, Zhuoyang Wu, Yukun Yan, Shuo Wang 0013, Zulong Chen, Yu Gu 0002, Ge Yu 0001, Maosong Sun 0001
ACL (1)4
2026 CheckRLM: Effective Knowledge-Thought Coherence Checking in Retrieval-Augmented Reasoning
abstract
Dingling Xu, Ruobing Wang, Qingfei Zhao, Yukun Yan, Zhichun Wang, Daren Zha, Shi Yu, Zhenghao Liu, Shuo Wang, Xu Han, Maosong Sun. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Dingling Xu, Qingfei Zhao, Yukun Yan, Zhichun Wang, Daren Zha, Shi Yu 0001, Zhenghao Liu 0001, Shuo Wang 0013, Maosong Sun 0001
ACL (1)4
2026 HIPPO: Enhancing the Table Understanding Capability of LLMs Through Hybrid-Modal Preference Optimization
Haolan Wang, Zhenghao Liu 0001, Xiaocui Yang, Yu Gu 0002, Yukun Yan, Qi Shi 0002, Fangfang Li 0002, Ge Yu 0001
DASFAA (4)6
2026 LISRec: Modeling User Preferences with Learned Item Shortcuts for Sequential Recommendation
abstract
User-item interaction histories are pivotal for sequential recommendation systems but often include noise, such as unintended clicks or actions that fail to reflect genuine user preferences. To address this, we propose Learned Item Shortcuts for Sequential Recommendation (LISRec), a novel framework that explicitly captures stable preferences by extracting personalized semantic shortcuts from historical interactions. LISRec first learns task-agnostic semantic representations to assess item similarities, then constructs a personalized semantic graph over all user-interacted items. By identifying the maximal semantic connectivity subset within this graph, LISRec selects the most representative items as semantic shortcuts to guide user preference modeling. This focused representation filters out irrelevant actions while preserving the diversity of genuine interests. Experimental results on the Yelp and Amazon Product datasets illustrate that LISRec achieves a 13% improvement over baseline recommendation models, showing its effectiveness in capturing stable user interests. Further analysis indicates that shortcut-based histories better capture user preferences, making more accurate and relevant recommendations. All codes and datasets are available at https://github.com/NEUIR/LISRec.
Haidong Xin, Zhenghao Liu 0001, Sen Mei, Yukun Yan, Shi Yu 0001, Shuo Wang 0013, Zulong Chen, Yu Gu 0002, Ge Yu 0001, Chenyan Xiong
KDD (1)4
2026 Mixture-of-Retrieval Experts for Reasoning-Guided Multimodal Knowledge Exploitation
abstract
Multimodal Retrieval-Augmented Generation (MRAG) has shown promise in mitigating hallucinations in Multimodal Large Language Models (MLLMs) by incorporating external knowledge. However, existing methods typically adhere to rigid retrieval paradigms by mimicking fixed retrieval trajectories and thus fail to fully exploit the knowledge of different retrieval experts through dynamic interaction based on the model's knowledge needs or evolving reasoning states. To overcome this limitation, we introduce Mixture-of-Retrieval Experts (MoRE), a novel framework that enables MLLMs to collaboratively interact with diverse retrieval experts for more effective knowledge exploitation. Specifically, MoRE learns to dynamically determine which expert to engage with, conditioned on the evolving reasoning state. To effectively train this capability, we propose Stepwise Group Relative Policy Optimization (Step-GRPO), which goes beyond sparse outcome-based supervision by encouraging MLLMs to interact with multiple retrieval experts and synthesize fine-grained rewards, thereby teaching the MLLM to fully coordinate all experts when answering a given query. Experimental results on diverse open-domain QA benchmarks demonstrate the effectiveness of MoRE, achieving average performance gains of over 7% compared to competitive baselines. Notably, MoRE exhibits strong adaptability by dynamically coordinating heterogeneous experts to precisely locate relevant information, validating its capability for robust, reasoning-driven expert collaboration. All codes and data are released on https://github.com/OpenBMB/MoRE.
Zhenghao Liu 0001, Yishan Li, Yukun Yan, Shuo Wang 0013, Yu Gu 0002, Minghe Yu 0001, Ge Yu 0001, Maosong Sun 0001
SIGIR5
2026 ReAlign: Optimizing the Visual Document Retriever with Reasoning-Guided Fine-Grained Alignment
Yifan Ji, Zhenghao Liu 0001, Yukun Yan, Zulong Chen, Shuo Wang 0013, Yu Gu 0002, Ge Yu 0001
SIGIR5
2026 Ensuring consistency with benign predictions: Differential privacy-guided certified defense against poisoning-based backdoor attacks
Yukun Yan, Jie Zhang 0073, Peng Tang 0002, Rui Chen 0012, Qilong Han, Haibo Hu 0001, Qing Guo 0005
Inf. Sci.1
2025 Enhancing Open-Domain Task-Solving Capability of LLMs via Autonomous Tool Integration from GitHub
abstract
Bohan Lyu, Xin Cong, Heyang Yu, Pan Yang, Cheng Qian, Zihe Wang, Yujia Qin, Yining Ye, Yaxi Lu, Chen Qian, Zhong Zhang, Yukun Yan, Yankai Lin, Zhiyuan Liu, Maosong Sun. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Bohan Lyu 0001, Xin Cong, Heyang Yu, Pan Yang 0022, Cheng Qian 0008, Yujia Qin, Yining Ye, Yaxi Lu, Zhong Zhang 0004, Yukun Yan, Yankai Lin 0001, Zhiyuan Liu 0001, Maosong Sun 0001
ACL (1)12
2025 RankCoT: Refining Knowledge for Retrieval-Augmented Generation through Ranking Chain-of-Thoughts
abstract
Retrieval-Augmented Generation (RAG) enhances the performance of Large Language Models (LLMs) by incorporating external knowledge.However, LLMs still encounter challenges in effectively utilizing the knowledge from retrieved documents, often being misled by irrelevant or noisy information.To address this issue, we introduce RankCoT, a knowledge refinement method that incorporates reranking signals in generating CoT-based summarization for knowledge refinement based on given query and all retrieval documents.During training, RankCoT prompts the LLM to generate Chain-of-Thought (CoT) candidates based on the query and individual documents.It then fine-tunes the LLM to directly reproduce the best CoT from these candidate outputs based on all retrieved documents, which requires LLM to filter out irrelevant documents during generating CoT-style summarization.Additionally, RankCoT incorporates a self-reflection mechanism that further refines the CoT outputs, resulting in higher-quality training data.Our experiments demonstrate the effectiveness of RankCoT, showing its superior performance over other knowledge refinement models.Further analysis reveals that RankCoT can provide shorter but effective refinement results, enabling the generator to produce more accurate answers.All code and data are available at https://github.com/NEUIR/RankCoT.
Mingyan Wu, Zhenghao Liu 0001, Yukun Yan, Shi Yu 0001, Zheni Zeng, Yu Gu 0002, Ge Yu 0001
ACL (1)3
2025 RAGEval: Scenario Specific RAG Evaluation Dataset Generation Framework
abstract
Kunlun Zhu, Yifan Luo, Dingling Xu, Yukun Yan, Zhenghao Liu, Shi Yu, Ruobing Wang, Shuo Wang, Yishan Li, Nan Zhang, Xu Han, Zhiyuan Liu, Maosong Sun. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Kunlun Zhu, Dingling Xu, Yukun Yan, Zhenghao Liu 0001, Shi Yu 0001, Shuo Wang 0013, Yishan Li, Xu Han 0007, Zhiyuan Liu 0001, Maosong Sun 0001
ACL (1)4
2025 LegalDuet: Learning Fine-Grained Representations for Legal Judgment Prediction via a Dual-View Contrastive Learning
Buqiang Xu, Zhenghao Liu 0001, Huiyuan Xie, Xiaoyuan Yi, Shuo Wang 0013, Yukun Yan, Liner Yang, Yu Gu 0002, Ge Yu 0001
ADMA (1)7
2025 Why Stop at One Error? Benchmarking LLMs as Data Science Code Debuggers for Multi-Hop and Multi-Bug Errors
abstract
LLMs are transforming software development, yet most code benchmarks still emphasize syntactic or functional correctness in simple, single-error cases.These settings miss the core difficulty of real-world data science debugging, where runtime errors propagate across multiple lines (multi-hop) and often appear in sets (multi-bug).We introduce DSDBench: Data Science Debugging Benchmark, the first benchmark to systematically evaluate LLMs on this challenge.Unlike general debugging benchmark suites such as SWE-bench, DSD-Bench targets non-expert, data-centric scripting, where practitioners rely heavily on blackbox libraries and write exploratory code that is error-prone and difficult to debug.Evaluations of state-of-the-art LLMs reveal large performance gaps: even frontier models that excel at code generation fail to reliably trace and resolve these errors, exposing a critical "generation versus understanding" gap.DSDBench provides a resource to drive progress toward more robust and trustworthy AI-assisted data science.
Zhiyu Yang 0001, Shuo Wang 0013, Yukun Yan, Yang Deng 0002
EMNLP3
2025 ExpandR: Teaching Dense Retrievers Beyond Queries with LLM Guidance
abstract
Large language models (LLMs) have demonstrated significant potential in enhancing dense retrieval through query augmentation.However, most existing methods treat the LLM and the retriever as separate modules, overlooking the alignment between generation and ranking objectives.In this work, we propose Ex-pandR, a unified LLM-augmented dense retrieval framework that jointly optimizes both the LLM and the retriever.ExpandR employs the LLM to generate semantically rich query expansions, which are leveraged to enhance the retriever's training.Simultaneously, the LLM is trained using Direct Preference Optimization (DPO), guided by a carefully designed reward function that balances retrieval effectiveness and generation consistency.This joint optimization paradigm enables mutual adaptation between the LLM and the retriever, resulting in query expansions that are both informative and well-suited for retrieval.Experimental results on multiple benchmarks show that Ex-pandR consistently outperforms strong baselines, achieving more than a 5% improvement in retrieval performance.
Sijia Yao, Pengcheng Huang 0004, Zhenghao Liu 0001, Yu Gu 0002, Yukun Yan, Shi Yu 0001, Ge Yu 0001
EMNLP5
2025 RAG-DDR: Optimizing Retrieval-Augmented Generation Using Differentiable Data Rewards
abstract
Retrieval-Augmented Generation (RAG) has proven its effectiveness in mitigating hallucinations in Large Language Models (LLMs) by retrieving knowledge from external resources. To adapt LLMs for the RAG systems, current approaches use instruction tuning to optimize LLMs, improving their ability to utilize retrieved knowledge. This supervised fine-tuning (SFT) approach focuses on equipping LLMs to handle diverse RAG tasks using different instructions. However, it trains RAG modules to overfit training signals and overlooks the varying data preferences among agents within the RAG system. In this paper, we propose a Differentiable Data Rewards (DDR) method, which end-to-end trains RAG systems by aligning data preferences between different RAG modules. DDR works by collecting the rewards to optimize each agent in the RAG system with the rollout method, which prompts agents to sample some potential responses as perturbations, evaluates the impact of these perturbations on the whole RAG system, and subsequently optimizes the agent to produce outputs that improve the performance of the RAG system. Our experiments on various knowledge-intensive tasks demonstrate that DDR significantly outperforms the SFT method, particularly for LLMs with smaller-scale parameters that depend more on the retrieved knowledge. Additionally, DDR exhibits a stronger capability to align the data preference between RAG modules. The DDR method makes the generation module more effective in extracting key information from documents and mitigating conflicts between parametric memory and external knowledge. All codes are available at https://github.com/OpenMatch/RAG-DDR.
Sen Mei, Zhenghao Liu 0001, Yukun Yan, Shuo Wang 0013, Shi Yu 0001, Zheni Zeng, Ge Yu 0001, Zhiyuan Liu 0001, Maosong Sun 0001, Chenyan Xiong
ICLR4
2025 VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents
abstract
Retrieval-augmented generation (RAG) is an effective technique that enables large language models (LLMs) to utilize external knowledge sources for generation. However, current RAG systems are solely based on text, rendering it impossible to utilize vision information like layout and images that play crucial roles in real-world multi-modality documents. In this paper, we introduce VisRAG, which tackles this issue by establishing a vision-language model (VLM)-based RAG pipeline. In this pipeline, instead of first parsing the document to obtain text, the document is directly embedded using a VLM as an image and then retrieved to enhance the generation of a VLM. Compared to traditional text-based RAG, VisRAG maximizes the retention and utilization of the data information in the original documents, eliminating the information loss introduced during the parsing process. We collect both open-source and synthetic data to train the retriever in VisRAG and explore a variety of generation methods. Experiments demonstrate that VisRAG outperforms traditional RAG in both the retrieval and generation stages, achieving a 20–40% end-to-end performance gain over traditional text-based RAG pipeline. Further analysis reveals that VisRAG is efficient in utilizing training data and demonstrates strong generalization capability, positioning it as a promising solution for RAG on multi-modality documents. Our code and data are available at https://github.com/openbmb/visrag.
Shi Yu 0001, Chaoyue Tang, Bokai Xu, Junbo Cui, Junhao Ran, Yukun Yan, Zhenghao Liu 0001, Shuo Wang 0013, Xu Han 0007, Zhiyuan Liu 0001, Maosong Sun 0001
ICLR6
2025 Benchmarking Retrieval-Augmented Generation in Multi-Modal Contexts
abstract
With the rapid advancement of Multi-modal Large Language Models (MLLMs), their capability in understanding both images and text has greatly improved. However, their potential for leveraging multi-modal contextual information in Retrieval-Augmented Generation (RAG) remains largely underexplored. To address this gap, this paper introduces Multi-Modal Retrieval-Augmented Generation (M2RAG), a benchmark designed to evaluate the effectiveness of Multi-modal Large Language Models in leveraging knowledge from multi-modal retrieval documents. The benchmark comprises four tasks: image captioning, multi-modal question answering, multi-modal fact verification, and image reranking. All tasks are set in an open-domain setting, requiring RAG models to retrieve query-relevant information from a multi-modal document collection and use it as contextual input for RAG modeling. To enhance the context utilization capabilities of MLLMs, we also introduce Multi-Modal Retrieval-Augmented Instruction Tuning (MM-RAIT), an instruction tuning method that optimizes MLLMs within multi-modal contexts. Our experiments demonstrate the effectiveness of MM-RAIT by significantly improving the quality of responses generated by different RAG models, outperforming MiniCPM-V 2.6 and Qwen2-VL with 34% and 33% gains, respectively. All data and code are available at https://github.com/NEUIR/M2RAG.
Zhenghao Liu 0001, Xingsheng Zhu, Tianshuo Zhou, Xiaoyuan Yi, Yukun Yan, Ge Yu 0001, Maosong Sun 0001
ACM Multimedia6
2025 ParamMute: Suppressing Knowledge-Critical FFNs for Faithful Retrieval-Augmented Generation
abstract
Large language models (LLMs) integrated with retrieval-augmented generation (RAG) have improved factuality by grounding outputs in external evidence. However, they remain susceptible to unfaithful generation, where outputs contradict retrieved context despite its relevance and accuracy. Existing approaches aiming to improve faithfulness primarily focus on enhancing the utilization of external context, but often overlook the persistent influence of internal parametric knowledge during generation. In this work, we investigate the internal mechanisms behind unfaithful generation and identify a subset of mid-to-deep feed-forward networks (FFNs) that are disproportionately activated in such cases. Building on this insight, we propose Parametric Knowledge Muting through FFN Suppression (ParamMute), a framework that improves contextual faithfulness by suppressing the activation of unfaithfulness-associated FFNs and calibrating the model toward retrieved knowledge. To evaluate our approach, we introduce CoFaithfulQA, a benchmark specifically designed to evaluate faithfulness in scenarios where internal knowledge conflicts with accurate external evidence. Experimental results show that ParamMute significantly enhances faithfulness across both CoFaithfulQA and the established ConFiQA benchmark, achieving substantial reductions in reliance on parametric memory. These findings underscore the importance of mitigating internal knowledge dominance and provide a new direction for improving LLM trustworthiness in RAG. All codes are available at https://github.com/OpenBMB/ParamMute.
Pengcheng Huang 0004, Zhenghao Liu 0001, Yukun Yan, Xiaoyuan Yi, Zhiyuan Liu 0001, Maosong Sun 0001, Tong Xiao 0001, Ge Yu 0001, Chenyan Xiong
NeurIPS3
2025 Building a Coding Assistant via the Retrieval-Augmented Language Model
abstract
Pretrained language models have shown strong effectiveness in code-related tasks, such as code retrieval, code generation, code summarization, and code completion tasks. In this article, we propose COde assistaNt viA retrieval-augmeNted language model (CONAN), which aims to build a code assistant by mimicking the knowledge-seeking behaviors of humans during coding. Specifically, it consists of a code structure-aware retriever (CONAN-R) and a dual-view code representation-based retrieval-augmented generation model (CONAN-G). CONAN-R pretrains CodeT5 using Code-Documentation Alignment and Masked Entity Prediction tasks to make language models code structure-aware and learn effective representations for code snippets and documentation. Then CONAN-G designs a dual-view code representation mechanism for implementing a retrieval-augmented code generation model. CONAN-G regards the code documentation descriptions as prompts, which help language models better understand the code semantics. Our experiments show that CONAN achieves convincing performance on different code generation tasks and significantly outperforms previous retrieval augmented code generation models. Our further analyses show that CONAN learns tailored representations for both code snippets and documentation by aligning code-documentation data pairs and capturing structural semantics by masking and predicting entities in the code data. Additionally, the retrieved code snippets and documentation provide necessary information from both program language and natural language to assist the code generation process. CONAN can also be used as an assistant for Large Language Models (LLMs), providing LLMs with external knowledge in shorter code document lengths to improve their effectiveness on various code tasks. It shows the ability of CONAN to extract necessary information and help filter out the noise from retrieved code documents.
Hanbin Wang, Zhenghao Liu 0001, Shi Yu 0001, Shuo Wang 0013, Yukun Yan, Yu Gu 0002, Ge Yu 0001
ACM Trans. Inf. Syst.6
2024 UltraLink: An Open-Source Knowledge-Enhanced Multilingual Supervised Fine-tuning Dataset
abstract
Haoyu Wang, Shuo Wang, Yukun Yan, Xujia Wang, Zhiyu Yang, Yuzhuang Xu, Zhenghao Liu, Liner Yang, Ning Ding, Xu Han, Zhiyuan Liu, Maosong Sun. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Shuo Wang 0013, Yukun Yan, Xujia Wang, Zhiyu Yang 0001, Yuzhuang Xu, Zhenghao Liu 0001, Liner Yang, Ning Ding 0002, Xu Han 0007, Zhiyuan Liu 0001, Maosong Sun 0001
ACL (1)3
2024 DPC: Filtering Out Patch-Based Poisoned Samples with Differential Privacy
Yukun Yan, Peng Tang 0002, Rui Chen 0012, Qilong Han, Ruochen Du
ESORICS (2)1
2024 Delta-CoMe: Training-Free Delta-Compression with Mixed-Precision for Large Language Models
abstract
Fine-tuning is a crucial process for adapting large language models (LLMs) to diverse applications. In certain scenarios, such as multi-tenant serving, deploying multiple LLMs becomes necessary to meet complex demands. Recent studies suggest decomposing a fine-tuned LLM into a base model and corresponding delta weights, which are then compressed using low-rank or low-bit approaches to reduce costs. In this work, we observe that existing low-rank and low-bit compression methods can significantly harm the model performance for task-specific fine-tuned LLMs (e.g., WizardMath for math problems). Motivated by the long-tail distribution of singular values in the delta weights, we propose a delta quantization approach using mixed-precision. This method employs higher-bit representation for singular vectors corresponding to larger singular values. We evaluate our approach on various fine-tuned LLMs, including math LLMs, code LLMs, chat LLMs, and even VLMs. Experimental results demonstrate that our approach performs comparably to full fine-tuned LLMs, surpassing both low-rank and low-bit baselines by a considerable margin. Additionally, we show that our method is compatible with various backbone LLMs, such as Llama-2, Llama-3, and Mistral, highlighting its generalizability.
Bowen Ping, Shuo Wang 0013, Hanqing Wang 0003, Xu Han 0007, Yuzhuang Xu, Yukun Yan, Yun Chen 0007, Baobao Chang, Zhiyuan Liu 0001, Maosong Sun 0001
NeurIPS6
2023 Nested Named Entity Recognition as Building Local Hypergraphs
abstract
Named entity recognition is a fundamental task in natural language processing. Based on the sequence labeling paradigm for flat named entity recognition, multiple methods have been developed to handle the nested structures. However, they either require fixed recognition order or introduce complex hypergraphs. To tackle this problem, we propose a novel model named Local Hypergraph Builder Network (LHBN) that builds multiple simpler local hypergraphs to capture named entities instead of a single complex full-size hypergraph. The proposed model has three main properties: (1) The named entities that share boundaries are captured in the same local hypergraph. (2) The boundary information is enhanced by building local hypergraphs. (3) The hypergraphs can be built bidirectionally to take advantage of the identification direction preference of different named entities. Experiments illustrate that our model outperforms previous state-of-the-art methods on four widely used nested named entity recognition datasets: ACE04, ACE05, GENIA, and KBP17. The code is available at https://github.com/yanyk13/local-hypergraph-building-network.git.
Yukun Yan, Bingling Cai, Sen Song
AAAI1
2023 Towards Defending Against Byzantine LDP Amplified Gain Attacks
Yukun Yan, Qingqing Ye 0001, Haibo Hu 0001, Rui Chen 0012, Qilong Han, Leixia Wang
DASFAA (1)1
2023 Fair_FM: An Improved Functional Mechanism to Provide Better Privacy Protection and Fairness Guarantee
abstract
Machine learning algorithms are currently used in various fields. The success of these algorithms is following the access of large amounts of sensitive personal data, which raises sociological concerns about privacy and fairness. At present, existing works only focus on providing fairness guarantee for the specific attribute that “can affect fairness” without considering the impact of other attributes on that attribute. Therefore, this paper takes the classical logistic regression model as the research object, uses Decision Tree Analysis and Bayesian Network to complete the analysis, and uses functional mechanism to achieve fair privacy protection. For the functional mechanism, we analyze that the Chebyshev polynomial expansion is a less error approach. Theoretical analysis and experimental results show that proposed algorithms can effectively achieve differential privacy and fairness guarantee while maintaining good utility.
Zuotian Han, Yukun Yan, Qilong Han
IJCNN2
2023 PFED-AGG: A Personalized Private Federated Learning Aggregation Algorithm
abstract
Federated learning is a special kind of distributed machine learning, in which multiple clients work together to solve a machine learning problem with the collaboration of a central server, and the clients only need to upload parameters for server aggregation instead of uploading raw data, so the privacy of the clients can be protected. However, existing research shows that an attacker who obtains the parameters uploaded by the client can reverse the privacy information of the client, and Federated Learning applies a local differential privacy approach to protect the information of the parameters uploaded by the client from being leaked. However, this privacy protection approach assumes the same level of privacy protection for all clients. To the best of our knowledge, there needs to be work that satisfies the personalized privacy needs of clients. To address this problem, we propose a personalized local differential privacy-based federation framework that satisfies the personalized privacy needs of clients and better protects the privacy of clients by making the specific privacy needs of clients inaccessible to the server. We have conducted extensive experiments on six benchmark datasets, and our approach works better and achieves personalized privacy protection compared to the same privacy-preserving method.
Yongjie Zhu, Yukun Yan, Qilong Han
IJCNN2
2023 Personalized sampling graph collection with local differential privacy for link prediction
Linyu Jiang, Yukun Yan, Zhihong Tian 0001, Zuobin Xiong, Qilong Han
World Wide Web (WWW)2
2018 Object-oriented Neural Programming (OONP) for Document Understanding
abstract
We propose Object-oriented Neural Programming (OONP), a framework for semantically parsing documents in specific domains.Basically, OONP reads a document and parses it into a predesigned object-oriented data structure that reflects the domain-specific semantics of the document.An OONP parser models semantic parsing as a decision process: a neural netbased Reader sequentially goes through the document, and builds and updates an intermediate ontology during the process to summarize its partial understanding of the text.OONP supports a big variety of forms (both symbolic and differentiable) for representing the state and the document, and a rich family of operations to compose the representation.An OONP parser can be trained with supervision of different forms and strength, including supervised learning (SL) , reinforcement learning (RL) and hybrid of the two.Our experiments on both synthetic and real-world document parsing tasks have shown that OONP can learn to handle fairly complicated ontology with training data of modest sizes.* The work was done when these authors worked as interns at DeeplyCurious.ai.
Zhengdong Lu, Xianggen Liu, Haotian Cui, Yukun Yan, Daqi Zheng
ACL (1)4