VLDB 2026 Research / reviewers in the wild / expert
Ziyi Ni
dblp:242/9922
· DBLP profile ↗
11ranked-venue papers
3as first author
10since 2021 · last 2026
0000-0001-6459-4715ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Language models and text generation · 53% Efficient and distributed learning · 18% Motion planning and robot control · 10% | |
| Software engineering, system software, and programming languages
3 papers |
Program synthesis and code generation · 62% Software maintenance and evolution · 38% | |
| Network and information security
2 papers |
Security and privacy of machine learning · 72% Privacy and data protection · 28% |
Topics — the 18 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Program synthesis and code generation
code agent |
1.9 | 2 | 2026 | GitTaskBench: A Benchmark for Code Agents Solving Real-World Tasks Through Code Repository Leveraging · AAAI 2026 RepoMaster: Autonomous Exploration and Understanding of GitHub Repositories for Complex Task Solving · NeurIPS 2025 |
Natural language and speech › Language models and text generation
LLM agents |
1.2 | 2 | 2026 | SE-Agent: Self-Evolution Trajectory Optimization in Multi-Step Reasoning with LLM-Based Agents · NeurIPS 2025 GitTaskBench: A Benchmark for Code Agents Solving Real-World Tasks Through Code Repository Leveraging · AAAI 2026 |
Natural language and speech › Language models and text generation
large language model safety |
1.0 | 1 | 2026 | SafetyMem: Adaptive Jailbreak Defense via Dual-Component Safety Memory · ACL (1) 2026 |
Security and privacy of machine learning
adversarial robustness |
1.0 | 1 | 2026 | SafetyMem: Adaptive Jailbreak Defense via Dual-Component Safety Memory · ACL (1) 2026 |
Security and privacy of machine learning › large language model safety
jailbreak defense |
1.0 | 1 | 2026 | SafetyMem: Adaptive Jailbreak Defense via Dual-Component Safety Memory · ACL (1) 2026 |
Natural language and speech › Language models and text generation › large language model reasoning
multi-step reasoning |
0.9 | 1 | 2025 | SE-Agent: Self-Evolution Trajectory Optimization in Multi-Step Reasoning with LLM-Based Agents · NeurIPS 2025 |
Machine learning › Reinforcement learning
self-evolution |
0.9 | 1 | 2025 | SE-Agent: Self-Evolution Trajectory Optimization in Multi-Step Reasoning with LLM-Based Agents · NeurIPS 2025 |
Robotics › Motion planning and robot control
trajectory optimization |
0.9 | 1 | 2025 | SE-Agent: Self-Evolution Trajectory Optimization in Multi-Step Reasoning with LLM-Based Agents · NeurIPS 2025 |
Software maintenance and evolution
code reuse |
0.9 | 1 | 2025 | RepoMaster: Autonomous Exploration and Understanding of GitHub Repositories for Complex Task Solving · NeurIPS 2025 |
Machine learning › Efficient and distributed learning
energy-efficient learning |
0.8 | 1 | 2024 | SpikeLM: Towards General Spike-Driven Language Modeling via Elastic Bi-Spiking Mechanisms · ICML 2024 |
Natural language and speech › Language models and text generation
language modeling |
0.8 | 1 | 2024 | SpikeLM: Towards General Spike-Driven Language Modeling via Elastic Bi-Spiking Mechanisms · ICML 2024 |
Machine learning › Efficient and distributed learning
model merging |
0.8 | 1 | 2024 | Mitigating Training Imbalance in LLM Fine-Tuning via Selective Parameter Merging · EMNLP 2024 |
Machine learning › Deep learning architectures and training
spiking neural network |
0.8 | 1 | 2024 | SpikeLM: Towards General Spike-Driven Language Modeling via Elastic Bi-Spiking Mechanisms · ICML 2024 |
Natural language and speech › Language models and text generation › large language model › large language model adaptation
supervised fine-tuning |
0.8 | 1 | 2024 | Mitigating Training Imbalance in LLM Fine-Tuning via Selective Parameter Merging · EMNLP 2024 |
Privacy and data protection › consumer privacy
demand privacy |
0.4 | 1 | 2019 | UMBRELLA: user demand privacy preserving framework based on association rules and differential privacy in social networks · Sci. China Inf. Sci. 2019 |
Privacy and data protection
differential privacy |
0.4 | 1 | 2019 | UMBRELLA: user demand privacy preserving framework based on association rules and differential privacy in social networks · Sci. China Inf. Sci. 2019 |
Software maintenance and evolution › issue management
issue resolution |
0.3 | 1 | 2025 | SE-Agent: Self-Evolution Trajectory Optimization in Multi-Step Reasoning with LLM-Based Agents · NeurIPS 2025 |
Data mining › pattern mining
association rule mining |
0.1 | 1 | 2019 | UMBRELLA: user demand privacy preserving framework based on association rules and differential privacy in social networks · Sci. China Inf. Sci. 2019 |
Methods — techniques the papers use, named apart from their topics
safety memory · 2.0benchmark evaluation · 2.0adversarial memory expansion · 2.0revision · 1.7refinement · 1.7recombination · 1.7monte carlo tree search · 1.7module-dependency graphs · 0.9large language model · 0.9function call graph · 0.9model merging · 0.8elastic bi-spiking mechanism · 0.8differential privacy · 0.8association rules · 0.8ablation study · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GitTaskBench: A Benchmark for Code Agents Solving Real-World Tasks Through Code Repository LeveragingabstractBeyond scratch coding, exploiting large-scale code repositories (e.g., GitHub) for practical tasks is vital in real-world software development, yet current benchmarks rarely evaluate code agents in such authentic, workflow-driven scenarios. To bridge this gap, we introduce GitTaskBench, a benchmark designed to systematically assess this capability via 54 realistic tasks across 7 modalities and 7 domains. Each task pairs a relevant repository with an automated, human-curated evaluation harness specifying practical success criteria. Beyond measuring execution and task success, we also propose the alpha-value metric to quantify the economic benefit of agent performance, which integrates task success rates, token cost, and average developer salaries. Experiments across three state-of-the-art agent frameworks with multiple advanced LLMs show that leveraging code repositories for complex task solving remains challenging: even the best-performing system, OpenHands+Claude 3.7, solves only 48.15% of tasks. Error analysis attributes over half of failures to seemingly mundane yet critical steps like environment setup and dependency resolution, highlighting the need for more robust workflow management and increased timeout preparedness. By releasing GitTaskBench, we aim to drive progress and attention toward repository-aware code reasoning, execution, and deployment---moving agents closer to solving complex, end-to-end real-world tasks. Ziyi Ni, Huacan Wang, Shuo Lu, Wang You, Zhenheng Tang, Sen Hu 0005, Bo Li 0117, Binxing Jiao, Daxin Jiang, Yuntao Du 0001 |
AAAI | 1 |
| 2026 | SafetyMem: Adaptive Jailbreak Defense via Dual-Component Safety MemoryabstractCurrent defenses for Large Language Models (LLMs) often suffer from a "memory gap": parameter-modifying methods are computationally rigid, while inference-time filters cannot retain or reuse defense knowledge across interactions.To address this, we propose Safet-yMem, a novel framework that secures LLMs through a dual-component safety memory system.SafetyMem consists of Semantic Safety Memory (SSM), which consolidates diverse jailbreak attempts into a structured knowledge base of attack patterns, and Episodic Safety Memory (ESM), which maintains an evolving set of procedural rules refined from historical detection failures.Unlike static defenses, Safe-tyMem allows the model to "remember" and adapt to emerging adversarial strategies without parameter retraining.To further enhance robustness, we introduce an adversarial memory expansion mechanism that proactively generates challenging variants to solidify these memories.Experiments on standard and stealthy jailbreak benchmarks show that SafetyMem substantially reduces attack success rates while preserving efficiency and interpretability, consistently outperforming state-of-the-art baselines across multiple LLMs. Ziyi Ni, Huacan Wang, Lei Sha |
ACL (1) | 2 |
| 2025 | Dynamic Cognitive Bias: Hallucination and Forgetting in the Cognitive Dynamics of LLMsabstractLarge Language Models (LLMs) have achieved remarkable success across a wide range of tasks, showcasing their transformative potential in natural language processing. However, these models face two persistent challenges: hallucination and forgetting. Hallucination occurs when models generate fabricated or inaccurate information in response to unfamiliar or out-of-distribution inputs, while forgetting refers to the inability to retain or recall previously learned knowledge during continued training. Although these phenomena have been extensively studied in isolation, their intrinsic connection remains largely un-explored. In this paper, we investigate the dialectical relationship between hallucination and forgetting through gradient dynamics during training. We propose a novel theory, Dynamic Cognitive Bias Entropy (HDCB), to jointly quantify the trade-offs between these two phenomena. To validate our findings, we conduct experiments based on two benchmark datasets, BIG-Bench Hard (BBH) and SQuAD, using LLaMA 3.2 1B and GPT 3.5 turbo models. Our results demonstrate that the order of training data significantly impacts the balance between hallucination and forgetting, with "known-first" training minimizing forgetting but increasing hallucination, and "unknown-first" training reducing hallucination but exacerbating forgetting. These findings, supported by entropy-based proofs and gradient-driven insights, provide a unified theoretical foundation and practical guidance for improving the reliability and robustness of LLMs. Ziyi Ni, Sen Tang, Xuhao Guo, Liguo Sun |
IJCNN | 1 |
| 2025 | SE-Agent: Self-Evolution Trajectory Optimization in Multi-Step Reasoning with LLM-Based AgentsabstractLarge Language Model (LLM)-based agents have recently shown impressive capabilities in complex reasoning and tool use via multi-step interactions with their environments. While these agents have the potential to tackle complicated tasks, their problem-solving process—agents' interaction trajectory leading to task completion—remains underexploited. These trajectories contain rich feedback that can navigate agents toward the right directions for solving problems correctly. Although prevailing approaches, such as Monte Carlo Tree Search (MCTS), can effectively balance exploration and exploitation, they ignore the interdependence among various trajectories and lack the diversity of search spaces, which leads to redundant reasoning and suboptimal outcomes. To address these challenges, we propose SE-Agent, a Self-Evolution framework that enables Agents to optimize their reasoning processes iteratively. Our approach revisits and enhances former pilot trajectories through three key operations: revision, recombination, and refinement. This evolutionary mechanism enables two critical advantages: (1) it expands the search space beyond local optima by intelligently exploring diverse solution paths guided by previous trajectories, and (2) it leverages cross-trajectory inspiration to efficiently enhance performance while mitigating the impact of suboptimal reasoning paths. Through these mechanisms, SE-Agent achieves continuous self-evolution that incrementally improves reasoning quality. We evaluate SE-Agent on SWE-bench Verified to resolve real-world GitHub issues. Experimental results across five strong LLMs show that integrating SE-Agent delivers up to 55% relative improvement, achieving state-of-the-art performance among all open-source agents on SWE-bench Verified. Yifu Guo, Jiaye Lin, Huacan Wang, Yuzhen Han, Sen Hu 0005, Ziyi Ni, Mingguang Chen |
NeurIPS | 6 |
| 2025 | RepoMaster: Autonomous Exploration and Understanding of GitHub Repositories for Complex Task SolvingabstractThe ultimate goal of code agents is to solve complex tasks autonomously.
Although large language models (LLMs) have made substantial progress in code generation, real-world tasks typically demand full-fledged code repositories rather than simple scripts. Building such repositories from scratch remains a major challenge. Fortunately, GitHub hosts a vast, evolving collection of open-source repositories, which developers frequently reuse as modular components for complex tasks. Yet, existing frameworks like OpenHands and SWE-Agent still struggle to effectively leverage these valuable resources.
Relying solely on README files provides insufficient guidance, and deeper exploration reveals two core obstacles: overwhelming information and tangled dependencies of repositories, both constrained by the limited context windows of current LLMs.
To tackle these issues, we propose RepoMaster, an autonomous agent framework designed to explore and reuse GitHub repositories for solving complex tasks.
For efficient understanding, RepoMaster constructs function-call graphs, module-dependency graphs, and hierarchical code trees to identify essential components, providing only identified core elements to the LLMs rather than the entire repository.
During autonomous execution, it progressively explores related components using our exploration tools and prunes information to optimize context usage.
Evaluated on the adjusted MLE-bench, RepoMaster achieves a 110\% relative boost in valid submissions over the strongest baseline OpenHands.
On our newly released GitTaskBench, RepoMaster lifts the task-pass rate from 40.7% to 62.9% while reducing token usage by 95%.
Our code and demonstration materials are publicly available at https://github.com/QuantaAlpha/RepoMaster. Huacan Wang, Ziyi Ni, Shuo Lu, Sen Hu 0005, Jiaye Lin, Yifu Guo, Yuntao Du 0001 |
NeurIPS | 2 |
| 2024 | Mitigating Training Imbalance in LLM Fine-Tuning via Selective Parameter MergingabstractSupervised fine-tuning (SFT) is crucial for adapting Large Language Models (LLMs) to specific tasks.In this work, we demonstrate that the order of training data can lead to significant training imbalances, potentially resulting in performance degradation.Consequently, we propose to mitigate this imbalance by merging SFT models fine-tuned with different data orders, thereby enhancing the overall effectiveness of SFT.Additionally, we introduce a novel technique, "parameter-selection merging," which outperforms traditional weightedaverage methods on five datasets.Further, through analysis and ablation studies, we validate the effectiveness of our method and identify the sources of performance improvements. Yiming Ju, Ziyi Ni, Xingrun Xing, Zhixiong Zeng, Siqi Fan 0001, Zheng Zhang 0006 |
EMNLP | 2 |
| 2024 | ViLaS: Exploring the Effects of Vision and Language Context in Automatic Speech RecognitionabstractEnhancing automatic speech recognition (ASR) performance by leveraging additional multimodal information has shown promising results in previous studies. However, most of these works have primarily focused on utilizing visual cues derived from human lip motions. In fact, context-dependent visual and linguistic cues can also benefit in many scenarios. In this paper, we first propose ViLaS (Vision and Language into Automatic Speech Recognition), a novel multimodal ASR model based on the continuous integrate-and-fire (CIF) mechanism, which can integrate visual and textual context simultaneously or separately, to facilitate speech recognition. Next, we introduce an effective training strategy that improves performance in modal-incomplete test scenarios. Then, to explore the effects of integrating vision and language, we create VSDial, a multimodal ASR dataset with multimodal context cues in both Chinese and English versions. Finally, empirical results are reported on the public Flickr8K and self-constructed VSDial datasets. We explore various cross-modal fusion schemes, analyze fine-grained cross-modal alignment on VSDial, and provide insights into the effects of integrating multimodal information on speech recognition. Ziyi Ni, Minglun Han, Linghui Meng 0001, Jing Shi 0003, Bo Xu 0002 |
ICASSP | 1 |
| 2024 | SpikeLM: Towards General Spike-Driven Language Modeling via Elastic Bi-Spiking MechanismsabstractTowards energy-efficient artificial intelligence similar to the human brain, the bio-inspired spiking neural networks (SNNs) have advantages of biological plausibility, event-driven sparsity, and binary activation. Recently, large-scale language models exhibit promising generalization capability, making it a valuable issue to explore more general spike-driven models. However, the binary spikes in existing SNNs fail to encode adequate semantic information, placing technological challenges for generalization. This work proposes the first fully spiking mechanism for general language tasks, including both discriminative and generative ones. Different from previous spikes with 0,1 levels, we propose a more general spike formulation with bi-directional, elastic amplitude, and elastic frequency encoding, while still maintaining the addition nature of SNNs. In a single time step, the spike is enhanced by direction and amplitude information; in spike frequency, a strategy to control spike firing rate is well designed. We plug this elastic bi-spiking mechanism in language modeling, named SpikeLM. It is the first time to handle general language tasks with fully spike-driven models, which achieve much higher accuracy than previously possible. SpikeLM also greatly bridges the performance gap between SNNs and ANNs in language modeling. Our code is available at https://github.com/Xingrun-Xing/SpikeLM. Xingrun Xing, Zheng Zhang 0006, Ziyi Ni, Shitao Xiao, Yiming Ju, Siqi Fan 0001, Yequan Wang |
ICML | 3 |
| 2024 | Bridge the Query and Document: Contrastive Learning for Generative Document RetrievalabstractGenerative retrieval has garnered significant attention for its end-to-end optimization and exceptional performance. Compared with the dense retrieval paradigm, the generative retrieval paradigm maps a query to a relevant document ID only relying on its model parameters, greatly simplifying the retrieval process. However, generative retrieval faces two challenges: it does not explicitly model the semantic relevance between query and document, and there exists a gap between the representation of query and document. To this end, we propose the Contrastive Search Index (ConSI), a simple but effective contrastive learning framework for generative document retrieval, to address the above challenges. Experiments show that the proposed ConSI consistently surpasses the previous generative retrieval baselines, NCI. Further analysis of different factors and indicators verifies the performance enhancement brought by our method. Besides, our ConSI also achieves excellent performance in the dense retrieval paradigm, demonstrating that the designed framework boosts representation learning ability and can be directly used as a dense retriever. Ziyi Ni, Zefa Hu, Bo Xu 0002 |
IJCNN | 2 |
| 2023 | Matching-Based Term Semantics Pre-Training for Spoken Patient Query UnderstandingabstractMedical Slot Filling (MSF) task aims to convert medical queries into structured information, playing an essential role in diagnosis dialogue systems. However, the lack of sufficient term semantics learning makes existing approaches hard to capture semantically identical but colloquial expressions of terms in medical conversations. In this work, we formalize MSF into a matching problem and propose a Term Semantics Pre-trained Matching Network (TSPMN) that takes both terms and queries as input to model their semantic inter-action. To learn term semantics better, we further design two self-supervised objectives, including Contrastive Term Discrimination (CTD) and Matching-based Mask Term Modeling (MMTM). CTD determines whether it is the masked term in the dialogue for each given term, while MMTM directly predicts the masked ones. Experimental results on two Chinese benchmarks show that TSPMN outperforms strong baselines, especially in few-shot settings1. Zefa Hu, Xiuyi Chen, Minglun Han, Ziyi Ni, Jing Shi 0003, Bo Xu 0002 |
ICASSP | 5 |
| 2019 | UMBRELLA: user demand privacy preserving framework based on association rules and differential privacy in social networks
Chunliu Yan, Ziyi Ni, Bin Cao 0003, Rongxing Lu, Shaohua Wu 0002, Qinyu Zhang 0001 |
Sci. China Inf. Sci. | 2 |