Xinhao Song

dblp:388/1833 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2026
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Theoretical computer science
2 papers
Automated reasoning and model checking · 100%
Artificial intelligence
2 papers
Reinforcement learning · 82% Graph learning · 18%
Network and information security
2 papers
Cryptographic primitives and cryptanalysis · 100%

Topics — the 8 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Cryptographic primitives and cryptanalysis › cryptanalysis
SAT-based cryptanalysis
1.622025
Bridging Crypto with ML-based Solvers: the SAT Formulation and Benchmarks · NeurIPS 2025
Learning Plaintext-Ciphertext Cryptographic Problems via ANF-based SAT Instance Representation · NeurIPS 2024
Machine learning › Reinforcement learning › reward design
reward shaping
1.012026
SEARL: Joint Optimization of Policy and Tool Graph Memory for Self-Evolving Agents · ACL (1) 2026
Machine learning › Reinforcement learning
self-improving agent
1.012026
SEARL: Joint Optimization of Policy and Tool Graph Memory for Self-Evolving Agents · ACL (1) 2026
Automated reasoning and model checking › satisfiability › SAT solving
conflict-driven clause learning
0.912025
Bridging Crypto with ML-based Solvers: the SAT Formulation and Benchmarks · NeurIPS 2025
Automated reasoning and model checking
satisfiability
0.912025
Bridging Crypto with ML-based Solvers: the SAT Formulation and Benchmarks · NeurIPS 2025
Automated reasoning and model checking › satisfiability
SAT solving
0.812024
Learning Plaintext-Ciphertext Cryptographic Problems via ANF-based SAT Instance Representation · NeurIPS 2024
Machine learning › Graph learning
graph neural network
0.212024
Learning Plaintext-Ciphertext Cryptographic Problems via ANF-based SAT Instance Representation · NeurIPS 2024
Machine learning › Graph learning › graph neural network
message passing
0.212024
Learning Plaintext-Ciphertext Cryptographic Problems via ANF-based SAT Instance Representation · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

message passing · 2.3graph neural network · 2.3ANF-based SAT representation · 2.3neural network · 1.7machine learning · 1.7hyperparameter optimization · 1.7reinforcement learning with verifiable rewards · 1.0experience memory · 1.0
YearPublicationVenuePosition
2026 SEARL: Joint Optimization of Policy and Tool Graph Memory for Self-Evolving Agents
abstract
Recent advances in Reinforcement Learning with Verifiable Rewards (RLVR) have demonstrated significant potential in single-turn reasoning tasks.With the paradigm shift toward self-evolving agentic learning, models are increasingly expected to learn from trajectories by synthesizing tools or accumulating explicit experiences.However, prevailing methods typically rely on large-scale LLMs or multi-agent frameworks, which hinder their deployment in resource-constrained environments.The inherent sparsity of outcome-based rewards also poses a substantial challenge, as agents typically receive feedback only upon completion of tasks.To address these limitations, we introduce a Tool-Memory based self-evolving agentic framework SEARL.Unlike approaches that directly utilize interaction experiences, our method constructs a structured experience memory that integrates planning with execution.This provides a novel state abstraction that facilitates generalization across analogous contexts, such as tool reuse.Consequently, agents extract explicit knowledge from historical data while leveraging inter-trajectory correlations to densify reward signals.We evaluate our framework on knowledge reasoning and mathematics tasks, demonstrating its effectiveness in achieving more practical and efficient learning 1 .
Xinshun Feng, Xinhao Song, Gongshen Liu
ACL (1)2
2025 ALIS: Aligned LLM Instruction Security Strategy for Unsafe Input Prompt
abstract
In large language models, existing instruction tuning methods may fail to balance the performance with robustness against attacks from user input like prompt injection and jailbreaking. Inspired by computer hardware and operating systems, we propose an instruction tuning paradigm named Aligned LLM Instruction Security Strategy (ALIS) to enhance model performance by decomposing user inputs into irreducible atomic instructions and organizing them into instruction streams which will guide the response generation of model. ALIS is a hierarchical structure, in which user inputs and system prompts are treated as user and kernel mode instructions respectively. Based on ALIS, the model can maintain security constraints by ignoring or rejecting the input instructions when user mode instructions attempt to conflict with kernel mode instructions. To build ALIS, we also develop an automatic instruction generation method for training ALIS, and give one instruction decomposition task and respective datasets. Notably, the ALIS framework with a small model to generate instruction streams still improve the resilience of LLM to attacks substantially without any lose on general capabilities.
Xinhao Song, Sufeng Duan, Gongshen Liu
COLING1
2025 Bridging Crypto with ML-based Solvers: the SAT Formulation and Benchmarks
abstract
The Boolean Satisfiability Problem (SAT) plays a crucial role in cryptanalysis, enabling tasks like key recovery and distinguisher construction. Conflict-Driven Clause Learning (CDCL) has emerged as the dominant paradigm in modern SAT solving, and machine learning has been increasingly integrated with CDCL-based SAT solvers to tackle complex cryptographic problems. However, the lack of a unified evaluation framework, inconsistent input formats, and varying modeling approaches hinder fair comparison. Besides, cryptographic SAT instances also differ structurally from standard SAT problems, and the absence of standardized datasets further complicates evaluation. To address these issues, we introduce SAT4CryptoBench, the first comprehensive benchmark for assessing machine learning–based solvers in cryptanalysis. SAT4CryptoBench provides diverse SAT datasets in both Arithmetic Normal Form (ANF) and Conjunctive Normal Form (CNF), spanning various algorithms, rounds, and key sizes. Our framework evaluates three levels of machine learning integration: standalone distinguishers for instance classification, heuristic enhancement for guiding solving strategies, and hyperparameter optimization for adapting to specific problem distributions. Experiments demonstrate that ANF-based networks consistently achieve superior performance over CNF-based networks in learning cryptographic features. Nonetheless, current ML techniques struggle to generalize across algorithms and instance sizes, with computational overhead potentially offsetting benefits on simpler cases. Despite this, ML-driven optimization strategies notably improve solver efficiency on cryptographic SAT instances. Besides, we propose BASIN, a bitwise solver taking plaintext-ciphertext bitstrings as inputs. Crucially, its superior performance on high-round problems highlights the importance of input modeling and the advantage of direct input representations for complex cryptographic structures.
Xinhao Zheng, Xinhao Song, Bolin Qiu, Yang Li 0197, Zhongteng Gui, Junchi Yan
NeurIPS2
2024 Learning Plaintext-Ciphertext Cryptographic Problems via ANF-based SAT Instance Representation
abstract
Cryptographic problems, operating within binary variable spaces, can be routinely transformed into Boolean Satisfiability (SAT) problems regarding specific cryptographic conditions like plaintext-ciphertext matching. With the fast development of learning for discrete data, this SAT representation also facilitates the utilization of machine-learning approaches with the hope of automatically capturing patterns and strategies inherent in cryptographic structures in a data-driven manner. Existing neural SAT solvers consistently adopt conjunctive normal form (CNF) for instance representation, which in the cryptographic context can lead to scale explosion and a loss of high-level semantics. In particular, extensively used XOR operations in cryptographic problems can incur an exponential number of clauses. In this paper, we propose a graph structure based on Arithmetic Normal Form (ANF) to efficiently handle the XOR operation bottleneck. Additionally, we design an encoding method for AND operations in these ANF-based graphs, demonstrating improved efficiency over alternative general graph forms for SAT. We then propose CryptoANFNet, a graph learning approach that trains a classifier based on a message-passing scheme to predict plaintext-ciphertext satisfiability. Using ANF-based SAT instances, CryptoANFNet demonstrates superior scalability and can naturally capture higher-order operational information. Empirically, CryptoANFNet achieves a 50x speedup over heuristic solvers and outperforms SOTA learning-based SAT solver NeuroSAT, with 96\% vs. 91\% accuracy on small-scale and 72\% vs. 55\% on large-scale datasets from real encryption algorithms. We also introduce a key-solving algorithm that simplifies ANF-based SAT instances from plaintext and ciphertext, enhancing key decryption accuracy from 76.5\% to 82\% and from 72\% to 75\% for datasets generated from two real encryption algorithms.
Xinhao Zheng, Yang Li 0197, Cunxin Fan, Huaijin Wu, Xinhao Song, Junchi Yan
NeurIPS5