EDBT 2026 Demo / reviewers in the wild / expert
Yubing Ren
dblp:331/1171
· DBLP profile ↗
18ranked-venue papers
3as first author
18since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 3 first-author · 14 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Exons-Detect: Identifying and Amplifying Exonic Tokens via Hidden-State Discrepancy for Robust AI-Generated Text DetectionabstractThe rapid advancement of large language models has increasingly blurred the boundary between human-written and AI-generated text, raising societal risks such as misinformation dissemination, authorship ambiguity, and threats to intellectual property rights.These concerns highlight the urgent need for effective and reliable detection methods.While existing training-free approaches often achieve strong performance by aggregating token-level signals into a global score, they typically assume uniform token contributions, making them less robust under short sequences or localized token modifications.To address these limitations, we propose Exons-Detect, a training-free method for AI-generated text detection based on an exon-aware token reweighting perspective.Exons-Detect identifies and amplifies informative exonic tokens by measuring hiddenstate discrepancy under a dual-model setting, and computes an interpretable translation score from the resulting importance-weighted token sequence.Empirical evaluations demonstrate that Exons-Detect achieves state-of-the-art detection performance and exhibits strong robustness to adversarial attacks and varying input lengths.In particular, it attains a 2.2% relative improvement in average AUROC over the strongest prior baseline on DetectRL.Code and data are available at https://github.com/ Xiaoweizhu57/Exons-Detect. Yubing Ren, Fang Fang 0009, Shi Wang 0002, Yanan Cao 0001, Li Guo 0001 |
ACL (1) | 2 |
| 2025 | PIG: Privacy Jailbreak Attack on LLMs via Gradient-based Iterative In-Context OptimizationabstractLarge Language Models (LLMs) excel in various domains but pose inherent privacy risks.Existing methods to evaluate privacy leakage in LLMs often use memorized prefixes or simple instructions to extract data, both of which well-alignment models can easily block.Meanwhile, Jailbreak attacks bypass LLM safety mechanisms to generate harmful content, but their role in privacy scenarios remains underexplored.In this paper, we examine the effectiveness of jailbreak attacks in extracting sensitive information, bridging privacy leakage and jailbreak attacks in LLMs.Moreover, we propose PIG, a novel framework targeting Personally Identifiable Information (PII) and addressing the limitations of current jailbreak methods.Specifically, PIG identifies PII entities and their types in privacy queries, uses in-context learning to build a privacy context, and iteratively updates it with three gradient-based strategies to elicit target PII.We evaluate PIG and existing jailbreak methods using two privacy-related datasets.Experiments on four white-box and two blackbox LLMs show that PIG outperforms baseline methods and achieves state-of-the-art (SoTA) results.The results underscore significant privacy risks in LLMs, emphasizing the need for stronger safeguards. Yidan Wang 0001, Yanan Cao 0001, Yubing Ren, Fang Fang 0009, Zheng Lin 0001, Binxing Fang |
ACL (1) | 3 |
| 2025 | From Trade-off to Synergy: A Versatile Symbiotic Watermarking Framework for Large Language ModelsabstractThe rise of Large Language Models (LLMs) has heightened concerns about the misuse of AI-generated text, making watermarking a promising solution.Mainstream watermarking schemes for LLMs fall into two categories: logits-based and sampling-based.However, current schemes entail trade-offs among robustness, text quality, and security.To mitigate this, we integrate logits-based and sampling-based schemes, harnessing their respective strengths to achieve synergy.In this paper, we propose a versatile symbiotic watermarking framework with three strategies: serial, parallel, and hybrid.The hybrid framework adaptively embeds watermarks using token entropy and semantic entropy, optimizing the balance between detectability, robustness, text quality, and security.Furthermore, we validate our approach through comprehensive experiments on various datasets and models.Experimental results indicate that our method outperforms existing baselines and achieves state-of-the-art (SOTA) performance.We believe this framework provides novel insights into diverse watermarking paradigms.Our code is available at https://github.com/redwyd/SymMark. Yidan Wang 0001, Yubing Ren, Yanan Cao 0001, Binxing Fang |
ACL (1) | 2 |
| 2025 | Reliably Bounding False Positives: A Zero-Shot Machine-Generated Text Detection Framework via Multiscaled Conformal PredictionabstractThe rapid advancement of large language models has raised significant concerns regarding their potential misuse by malicious actors.As a result, developing effective detectors to mitigate these risks has become a critical priority.However, most existing detection methods focus excessively on detection accuracy, often neglecting the societal risks posed by high false positive rates (FPRs).This paper addresses this issue by leveraging Conformal Prediction (CP), which effectively constrains the upper bound of FPRs.While directly applying CP constraints FPRs, it also leads to a significant reduction in detection performance.To overcome this trade-off, this paper proposes a Zero-Shot Machine-Generated Text Detection Framework via Multiscaled Conformal Prediction (MCP), which both enforces the FPR constraint and improves detection performance.This paper also introduces RealDet, a high-quality dataset that spans a wide range of domains, ensuring realistic calibration and enabling superior detection performance when combined with MCP.Empirical evaluations demonstrate that MCP effectively constrains FPRs, significantly enhances detection performance, and increases robustness against adversarial attacks across multiple detectors and datasets. Yubing Ren, Yanan Cao 0001, Xixun Lin, Fang Fang 0009, Yangxi Li |
ACL (1) | 2 |
| 2025 | ToP: a Structured Pathway for Document-level Information Extraction with Large Language ModelsabstractIn Natural Language Processing (NLP), Document-level Information Extraction (DocIE) poses a significant challenge, requiring analyzing and synthesizing information across extensive contexts. This complexity extends beyond the scale of data, delving into the nuanced interplay of context, semantics, and text elements. Recent advancements in Large Language Models (LLMs) have opened new prospects for DocIE. However, their application is limited by the models’ struggles with extended contexts and complex narrative structures, a gap highlighted when comparing the efficacy of direct LLM prompting against Chain-of-Thought (CoT) prompting. Addressing these challenges, we introduce the Tree of Plans (Top) framework, a novel approach that mirrors human decision-making by engaging in slow, deliberate reasoning. Top decomposes the DocIE task into three structured stages: (1) Task-specific Seed Plan Decomposition, (2) LLM-driven Plan Sampling and Voting, and (3) Instruction-prompted Plan Execution. Each stage leverages zero-shot prompting of an LLM, ensuring coherent progression and comprehensive document-level analysis. We evaluate Top on four challenging DocIE datasets, which cover document-level relation extraction and document-level event extraction tasks, both requiring complex reasoning over the long context. Pengfei Yin, Yubing Ren, Fang Fang 0009, Boxiang Hu |
IJCNN | 2 |
| 2025 | Exploring Polyglot Harmony: On Multilingual Data Allocation for Large Language Models PretrainingabstractLarge language models (LLMs) have become integral to a wide range of applications worldwide, driving an unprecedented global demand for effective multilingual capabilities. Central to achieving robust multilingual performance is the strategic allocation of language proportions within training corpora. However, determining optimal language ratios is highly challenging due to intricate cross-lingual interactions and sensitivity to dataset scale. This paper introduces CLIMB (Cross-Lingual Interaction-aware Multilingual Balancing), a novel framework designed to systematically optimize multilingual data allocation. At its core, CLIMB introduces a cross-lingual interaction-aware language ratio, explicitly quantifying each language’s effective allocation by capturing inter-language dependencies. Leveraging this ratio, CLIMB proposes a principled two-step optimization procedure—first equalizing marginal benefits across languages, then maximizing the magnitude of the resulting language allocation vectors—significantly simplifying the inherently complex multilingual optimization problem. Extensive experiments confirm that CLIMB can accurately measure cross-lingual interactions across various multilingual settings. LLMs trained with CLIMB-derived proportions consistently achieve state-of-the-art multilingual performance, even achieve competitive performance with open-sourced LLMs trained with more tokens. Yubing Ren, Fengze Liu, Haobin Lin, Bingni Zhang, Taifeng Wang |
NeurIPS | 2 |
| 2025 | DNA-DetectLLM: Unveiling AI-Generated Text via a DNA-Inspired Mutation-Repair ParadigmabstractThe rapid advancement of large language models (LLMs) has blurred the line between AI-generated and human-written text. This progress brings societal risks such as misinformation, authorship ambiguity, and intellectual property concerns, highlighting the urgent need for reliable AI-generated text detection methods. However, recent advances in generative language modeling have resulted in significant overlap between the feature distributions of human-written and AI-generated text, blurring classification boundaries and making accurate detection increasingly challenging. To address the above challenges, we propose a DNA-inspired perspective, leveraging a repair-based process to directly and interpretably capture the intrinsic differences between human-written and AI-generated text. Building on this perspective, we introduce **DNA-DetectLLM**, a zero-shot detection method for distinguishing AI-generated and human-written text. The method constructs an ideal AI-generated sequence for each input, iteratively repairs non-optimal tokens, and quantifies the cumulative repair effort as an interpretable detection signal. Empirical evaluations demonstrate that our method achieves state-of-the-art detection performance and exhibits strong robustness against various adversarial attacks and input lengths. Specifically, DNA-DetectLLM achieves relative improvements of **5.55\%** in AUROC and **2.08\%** in F1 score across multiple public benchmark datasets. Code and data are available at https://github.com/Xiaoweizhu57/DNA-DetectLLM. Yubing Ren, Fang Fang 0009, Qingfeng Tan, Shi Wang 0002, Yanan Cao 0001 |
NeurIPS | 2 |
| 2025 | EnsemJudge: Enhancing Reliability in Chinese LLM-Generated Text Detection Through Diverse Model Ensembles
Zhuoshang Wang, Yubing Ren, Guoyu Zhao, Hao Li 0156, Yanan Cao 0001 |
NLPCC (4) | 2 |
| 2025 | Bridging the Gap: Aligning Language Model Generation with Structured Information Extraction via Controllable State TransitionabstractLarge language models (LLMs) achieve superior performance in generative tasks. However, due to the natural gap between language model generation and structured information extraction in three dimensions: task type, output format, and modeling granularity, they often fall short in structured information extraction, a crucial capability for effective data utilization on the web. In this paper, we define the generation process of the language model as the controllable state transition, aligning the generation and extraction processes to ensure the integrity of the output structure and adapt to the goals of the information extraction task. Furthermore, we propose the Structure2Text decider to help the language model understand the fine-grained extraction information, which converts the structured output into natural language and makes state decisions, thereby focusing on the task-specific information kernels, and alleviating language model hallucinations and incorrect content generation. We conduct extensive experiments and detailed analyses on myriad information extraction tasks, including named entity recognition, relation extraction, and event argument extraction. Our method not only achieves significant performance improvements but also considerably enhances the model's capability to generate precise and relevant content, making the extracted content easy to parse. Hao Li 0156, Yubing Ren, Yanan Cao 0001, Fang Fang 0009, Zheng Lin 0001, Shi Wang 0002 |
WWW | 2 |
| 2024 | Teaching Large Language Models to Translate on Low-resource Languages with Textbook PromptingabstractLarge Language Models (LLMs) have achieved impressive results in Machine Translation by simply following instructions, even without training on parallel data. However, LLMs still face challenges on low-resource languages due to the lack of pre-training data. In real-world situations, humans can become proficient in their native languages through abundant and meaningful social interactions and can also learn foreign languages effectively using well-organized textbooks. Drawing inspiration from human learning patterns, we introduce the Translate After LEarNing Textbook (TALENT) approach, which aims to enhance LLMs’ ability to translate low-resource languages by learning from a textbook. TALENT follows a step-by-step process: (1) Creating a Textbook for low-resource languages. (2) Guiding LLMs to absorb the Textbook’s content for Syntax Patterns. (3) Enhancing translation by utilizing the Textbook and Syntax Patterns. We thoroughly assess TALENT’s performance using 112 low-resource languages from FLORES-200 with two LLMs: ChatGPT and BLOOMZ. Evaluation across three different metrics reveals that TALENT consistently enhances translation performance by 14.8% compared to zero-shot baselines. Further analysis demonstrates that TALENT not only improves LLMs’ comprehension of low-resource languages but also equips them with the knowledge needed to generate accurate and fluent sentences in these languages. Ping Guo 0002, Yubing Ren, Yue Hu 0002, Yunpeng Li 0006, Jiarui Zhang 0003, Xingsheng Zhang, Heyan Huang |
LREC/COLING | 2 |
| 2024 | DEIE: Benchmarking Document-level Event Information Extraction with a Large-scale Chinese News DatasetabstractA text corpus centered on events is foundational to research concerning the detection, representation, reasoning, and harnessing of online events. The majority of current event-based datasets mainly target sentence-level tasks, thus to advance event-related research spanning from sentence to document level, this paper introduces DEIE, a unified large-scale document-level event information extraction dataset with over 56,000+ events and 242,000+ arguments. Three key features stand out: large-scale manual annotation (20,000 documents), comprehensive unified annotation (encompassing event trigger/argument, summary, and relation at once), and emergency events annotation (covering 19 emergency types). Notably, our experiments reveal that current event-related models struggle with DEIE, signaling a pressing need for more advanced event-related research in the future. Yubing Ren, Yanan Cao 0001, Hao Li 0156, Zixuan ZM Ma, Fang Fang 0009, Ping Guo 0002 |
LREC/COLING | 1 |
| 2024 | Sorting, Reasoning, and Extraction: An Easy-to-Hard Reasoning Framework for Document-Level Event Argument ExtractionabstractDocument-level event argument extraction is a crucial task to help understand event information. Existing methods mostly ignore the different extraction difficulties of arguments, and the lack of task planning significantly affects the extraction and reasoning abilities of the model. In this paper, we innovatively analyze the difficulty of arguments and propose a novel framework for reasoning from easy to hard, aiming to use the information of simple arguments to help the extraction of difficult arguments in a human-like way. Specifically, our framework consists of three core modules: sorting, reasoning, and extraction. The sorting module first sorts the argument roles according to the current context and plans the reasoning path from easy to hard. Then, the reasoning module performs information reasoning based on the reasoning path to help capture the information of difficult arguments. Finally, the extraction module utilizes the reasoning information to complete argument extraction. Experimental results on the RAMS and WikiEvents datasets show the great advantages of our proposed approach. In particular, we obtain new state-of-the-art (SOTA) performance in multiple scenarios. Hao Li 0156, Yanan Cao 0001, Yubing Ren, Fang Fang 0009, Lanxue Zhang, Shi Wang 0002 |
ICASSP | 3 |
| 2024 | Steering Large Language Models for Cross-lingual Information RetrievalabstractIn today's digital age, accessing information across language barriers poses a significant challenge, with conventional search systems often struggling to interpret and retrieve multilingual content accurately. Addressing this issue, our study introduces a novel integration of applying Large Language Models (LLMs) as Cross-lingual Readers in information retrieval systems, specifically targeting the complexities of cross-lingual information retrieval (CLIR). We present an innovative approach: Activation Steered Multilingual Retrieval (ASMR) that employs "steering activations''-a method to adjust and direct the LLM's focus-enhancing its ability to understand user queries and generate accurate, language-coherent responses. ASMR adeptly combines a Multilingual Dense Passage Retrieval (mDPR) system with an LLM, overcoming the limitations of traditional search engines in handling diverse linguistic inputs. This approach is particularly effective in managing the nuances and intricacies inherent in various languages. Rigorous testing on established benchmarks such as XOR-TyDi QA, and MKQA demonstrates that ASMR not only meets but surpasses existing standards in CLIR, achieving state-of-the-art performance. The results of our research hold significant implications for understanding the inherent features of how LLMs understand and generate natural languages, offering an attempt towards more inclusive, effective, and linguistically diverse information access on a global scale. Ping Guo 0002, Yubing Ren, Yue Hu 0002, Yanan Cao 0001, Yunpeng Li 0006, Heyan Huang |
SIGIR | 2 |
| 2024 | Query in Your Tongue: Reinforce Large Language Models with Retrievers for Cross-lingual Search Generative ExperienceabstractIn the contemporary digital landscape, search engines play an invaluable role in information access, yet they often face challenges in Cross-Lingual Information Retrieval (CLIR). Though attempts are made to improve CLIR, current methods still leave users grappling with issues such as misplaced named entities and lost cultural context when querying in non-native languages. While some advances have been made using Neural Machine Translation models and cross-lingual representation, these are not without limitations. Enter the paradigm shift brought about by Large Language Models (LLMs), which have transformed search engines from simple retrievers to generators of contextually relevant information. This paper introduces the Multilingual Information Model for Intelligent Retrieval (MIMIR). Built on the power of LLMs, MIMIR directly responds in the language of the user's query, reducing the need for post-search translations. Our model's architecture encompasses a dual-module system: a retriever for searching multilingual documents and a responder for crafting answers in the user's desired language. Through a unique unified training framework, with the retriever serving as a reward model supervising the responder, and in turn, the responder producing synthetic data to refine the retriever's proficiency, MIMIR's retriever and responder iteratively enhance each other. Performance evaluations via CLEF and MKQA benchmarks reveal MIMIR's superiority over existing models, effectively addressing traditional CLIR challenges. Ping Guo 0002, Yue Hu 0002, Yanan Cao 0001, Yubing Ren, Yunpeng Li 0006, Heyan Huang |
WWW | 4 |
| 2023 | Retrieve-and-Sample: Document-level Event Argument Extraction via Hybrid Retrieval AugmentationabstractRecent studies have shown the effectiveness of retrieval augmentation in many generative NLP tasks.These retrieval-augmented methods allow models to explicitly acquire prior external knowledge in a non-parametric manner and regard the retrieved reference instances as cues to augment text generation.These methods use similarity-based retrieval, which is based on a simple hypothesis: the more the retrieved demonstration resembles the original input, the more likely the demonstration label resembles the input label.However, due to the complexity of event labels and sparsity of event arguments, this hypothesis does not always hold in document-level EAE.This raises an interesting question: How do we design the retrieval strategy for document-level EAE?We investigate various retrieval settings from the input and label distribution views in this paper.We further augment document-level EAE with pseudo demonstrations sampled from event semantic regions that can cover adequate alternatives in the same context and event schema.Through extensive experiments on RAMS and WikiEvents, we demonstrate the validity of our newly introduced retrieval-augmented methods and analyze why they work. Yubing Ren, Yanan Cao 0001, Ping Guo 0002, Fang Fang 0009, Zheng Lin 0001 |
ACL (1) | 1 |
| 2023 | Mitigating Long-Tail Language Representation Collapsing via Cross-Lingual Bootstrapped Unsupervised Fine-TuningabstractLarge Language Models have shown great capability to comprehend natural language and provide reasonable responses. However, previous researches have shown weak performance of these models on low-resource (long-tail) languages. It remains to be a problem to mitigate the performance gap between long-tail languages and rich-resource ones, which is referred to as long-tail language representation collapsing. Though some previous works can generate pseudo-parallel corpora with the auto-regressive generation, this generation progress is time-consuming and remains low quality, particularly for long-tail languages. In this paper, we propose a (X) Cross-lingual Bootstrapped Unsupervised Fine-tuning Framework (X-BUFF) to mitigate long-tail language representation collapsing. X-BUFF iteratively updates cross-lingual PLMs in a curriculum way. In each iteration of X-BUFF, we (1) select sentences with complementary semantics from monolingual corpora in long-tail languages. (2) match these selected sentences with semantic equivalent sentences in many other languages to create parallel sentence pairs, which we then merge with previous sentence pairs to build a larger and more difficult bootstrapped parallel queue. (3) fine-tune the PLMs with the bootstrapped parallel queue. Extensive experiments show that X-BUFF can mitigate the long-tail language representation collapsing problem in cross-lingual PLMs and achieve significant improvements over the previous baselines on several cross-lingual evaluation benchmarks. Ping Guo 0002, Yue Hu 0002, Yubing Ren, Yunpeng Li 0006, Jiarui Zhang 0003, Xingsheng Zhang |
ECAI | 3 |
| 2023 | A Multi-granularity Similarity Enhanced Model for Implicit Event Argument Extraction
Yanhe Fu, Yi Liu 0067, Yanan Cao 0001, Yubing Ren, Qingyue Wang, Fang Fang 0009, Cong Cao 0001 |
NLPCC (2) | 4 |
| 2022 | CLIO: Role-interactive Multi-event Head Attention Network for Document-level Event ExtractionabstractTransforming the large amounts of unstructured text on the Internet into structured event knowledge is a critical, yet unsolved goal of NLP, especially when addressing document-level text. Existing methods struggle in Document-level Event Extraction (DEE) due to its two intrinsic challenges: (a) Nested arguments, which means one argument is the sub-string of another one. (b) Multiple events, which indicates we should identify multiple events and assemble the arguments for them. In this paper, we propose a role-interactive multi-event head attention network (CLIO) to solve these two challenges jointly. The key idea is to map different events to multiple subspaces (i.e. multi-event head). In each event subspace, we draw the semantic representation of each role closer to its corresponding arguments, then we determine whether the current event exists. To further optimize event representation, we propose an event representation enhancing strategy to regularize pre-trained embedding space to be more isotropic. Our experiments on two widely used DEE datasets show that CLIO achieves consistent improvements over previous methods. Yubing Ren, Yanan Cao 0001, Fang Fang 0009, Ping Guo 0002, Zheng Lin 0001, Yi Liu 0067 |
COLING | 1 |