Zhenyun Yin

dblp:437/2255 · DBLP profile ↗
← Back
1ranked-venue papers
1as first author
1since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Language models and text generation · 81% Trustworthy machine learning · 19%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
chain-of-thought reasoning
1.012026
Deliberative Searcher: Improving LLM Reliability via Reinforcement Learning with Constraints · ACL (1) 2026
Machine learning › Trustworthy machine learning › calibration
confidence calibration
1.012026
Deliberative Searcher: Improving LLM Reliability via Reinforcement Learning with Constraints · ACL (1) 2026
Natural language and speech › Language models and text generation
large language model reasoning
1.012026
Deliberative Searcher: Improving LLM Reliability via Reinforcement Learning with Constraints · ACL (1) 2026
Natural language and speech › Language models and text generation › trustworthy language model
large language model reliability
1.012026
Deliberative Searcher: Improving LLM Reliability via Reinforcement Learning with Constraints · ACL (1) 2026
Natural language and speech › Language models and text generation
retrieval-augmented generation
1.012026
Deliberative Searcher: Improving LLM Reliability via Reinforcement Learning with Constraints · ACL (1) 2026
Natural language and speech › Language models and text generation › large language model inference
test-time compute
0.312026
Deliberative Searcher: Improving LLM Reliability via Reinforcement Learning with Constraints · ACL (1) 2026

Methods — techniques the papers use, named apart from their topics

lagrangian multipliers · 1.0constrained reinforcement learning · 1.0confidence-weighted aggregation · 1.0
YearPublicationVenuePosition
2026 Deliberative Searcher: Improving LLM Reliability via Reinforcement Learning with Constraints
abstract
Large language models with search capabilities frequently exhibit miscalibrated confidence, producing incorrect answers with high certainty.We present Deliberative Searcher, a reasoning-primary framework that integrates search operations into chain-of-thought generation while maintaining explicit confidence calibration.Our method employs constrained reinforcement learning with adaptive Lagrangian multipliers to jointly optimize correctness and reliability.Experiments across five benchmarks demonstrate substantial improvements: our 7B model reduces average false-certain rates from 54% in baselines to 2%, while our 72B variant achieves competitive accuracy with closedsource models and reduces false-certain rates to 9%.The well-calibrated confidence scores also enable more efficient test-time compute: instead of standard majority voting, we use confidence-weighted aggregation and match the performance of 16-sample majority voting with only 4 samples, a 4× reduction in inference compute.These results establish calibrated confidence as a foundation for both trustworthy outputs and adaptive test-time compute, demonstrating the value of the proposed constrained RL framework in search-augmented language models.
Zhenyun Yin, Xuhong Wang, Xingjun Ma, Yingchun Wang 0004
ACL (1)1