Jianyuan Zhong

dblp:239/5133 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
7since 2021 · last 2026
0009-0006-9954-9563ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Language models and text generation · 57% Knowledge representation and reasoning · 18% Reinforcement learning · 9%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Electronic design automation · 100%
Network and information security
1 paper
Security and privacy of machine learning · 100%

Topics — the 18 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
reasoning verification
1.922026
Solve-Detect-Verify: Inference-Time Scaling with Flexible Generative Verifier · ACL (1) 2026
Dyve: Thinking Fast and Slow for Dynamic Process Verification · EMNLP 2025
Natural language and speech › Language models and text generation
test-time scaling
1.322026
Solve-Detect-Verify: Inference-Time Scaling with Flexible Generative Verifier · ACL (1) 2026
Dyve: Thinking Fast and Slow for Dynamic Process Verification · EMNLP 2025
Natural language and speech › Language models and text generation
chain-of-thought reasoning
1.012026
Solve-Detect-Verify: Inference-Time Scaling with Flexible Generative Verifier · ACL (1) 2026
Knowledge, reasoning and agents › Knowledge representation and reasoning
commonsense reasoning
0.912025
Dependency Matters: Enhancing LLM Reasoning with Explicit Knowledge Grounding · NeurIPS 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge-based systems
knowledge-grounded reasoning
0.912025
Dependency Matters: Enhancing LLM Reasoning with Explicit Knowledge Grounding · NeurIPS 2025
Natural language and speech › Language models and text generation
large language model reasoning
0.912025
Dependency Matters: Enhancing LLM Reasoning with Explicit Knowledge Grounding · NeurIPS 2025
Machine learning › Reinforcement learning › reinforcement learning from human feedback
process reward model
0.912025
Dyve: Thinking Fast and Slow for Dynamic Process Verification · EMNLP 2025
Electronic design automation › machine learning for EDA
circuit representation learning
0.912025
DeepGate4: Efficient and Effective Representation Learning for Circuit Design at Scale · ICLR 2025
Electronic design automation
hardware verification and test
0.912025
DeepGate4: Efficient and Effective Representation Learning for Circuit Design at Scale · ICLR 2025
Electronic design automation
logic synthesis
0.912025
DeepGate4: Efficient and Effective Representation Learning for Circuit Design at Scale · ICLR 2025
Electronic design automation › hardware verification and test
testability analysis
0.912025
DeepGate4: Efficient and Effective Representation Learning for Circuit Design at Scale · ICLR 2025
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.812024
GuardT2I: Defending Text-to-Image Models from Adversarial Prompts · NeurIPS 2024
Security and privacy of machine learning › adversarial attack › large language model attack
adversarial prompts
0.812024
GuardT2I: Defending Text-to-Image Models from Adversarial Prompts · NeurIPS 2024
Natural language and speech › Information extraction and text analysis › text classification
humor detection
0.412019
UR-FUNNY: A Multimodal Language Dataset for Understanding Humor · EMNLP/IJCNLP (1) 2019
Natural language and speech › Language models and text generation › natural language understanding
multimodal language understanding
0.412019
UR-FUNNY: A Multimodal Language Dataset for Understanding Humor · EMNLP/IJCNLP (1) 2019
Natural language and speech › Language models and text generation › decoding
best-of-n selection
0.312025
Dyve: Thinking Fast and Slow for Dynamic Process Verification · EMNLP 2025
Machine learning › Graph learning › graph neural network
graph transformer
0.312025
DeepGate4: Efficient and Effective Representation Learning for Circuit Design at Scale · ICLR 2025
Natural language and speech › Question answering and dialogue systems › reasoning-based question answering
multi-hop question answering
0.312025
Dependency Matters: Enhancing LLM Reasoning with Explicit Knowledge Grounding · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

large language model · 3.4graph transformer · 1.7GAT · 1.7CUDA kernel · 1.7adversarial prompt detection · 1.5process reward model · 1.0step-wise verification · 0.9reasoning models · 0.9monte carlo estimation · 0.9LLM-as-a-judge · 0.9text guidance embedding transformation · 0.8
YearPublicationVenuePosition
2026 Solve-Detect-Verify: Inference-Time Scaling with Flexible Generative Verifier
abstract
Figure 1: The Solve-Detect-Verify (SDV) pipeline transforms linguistic signals into efficiency.Left: On AIME 2024, SDV achieves 83.3% accuracy (vs.63.3% for GenPRM) while using 6x fewer verification tokens by pruning redundant reasoning.Right: The pipeline is powered by FlexiVe , a unified verifier.Unlike process-based verifiers that incur per-step overhead, FlexiVe analyzes traces holistically.It employs a "pragmatic" consensus strategy: parallel "Fast Thinking" checks (∼0.1k tokens) provide an initial semantic intuition, escalating to deliberative "Slow Thinking" (∼4k tokens) only when the model exhibits verbalized uncertainty.
Jianyuan Zhong, Zeju Li, Xiangyu Wen 0001, Kezhi Li, Qiang Xu 0001
ACL (1)1
2025 Dyve: Thinking Fast and Slow for Dynamic Process Verification
abstract
Large Language Models (LLMs) have advanced significantly in complex reasoning, often leveraging external verifiers to improve multi-step process reliability.However, existing process verification methods face critical limitations: discriminative Process Reward Models (PRMs) often provide overly simplistic binary feedback and struggle with incomplete reasoning traces, while sophisticated Generative Reward Models (GenRMs) can be computationally expensive.Furthermore, curating quality supervision data for process verifier is of challenging.Therefore, we present Dyve, a dynamic process verifier that enhances reasoning error detection in LLMs by integrating fast (System 1) and slow (System 2) thinking, inspired by Kahneman's Systems Theory.Dyve adaptively applies immediate token-level confirmation for straightforward steps and comprehensive analysis for complex ones.To address data challenges and enable its adaptive fast and slow thinking, Dyve employs a novel step-wise consensus-filtered supervision strategy.This strategy leverages Monte Carlo estimation, LLM-as-a-Judge, and specialized reasoning models to extract the high-quality training signals from noisy rollouts.Experimental results on ProcessBench and the MATH dataset confirm that Dyve significantly outperforms existing process-based verifiers and boosts performance in Best-of-N settings, while maintaining computational efficiency through strategic resource allocation.Our code, data and model are released at: https://github.com/ staymylove/
Jianyuan Zhong, Zeju Li, Xiangyu Wen 0001, Qiang Xu 0001
EMNLP1
2025 DeepGate4: Efficient and Effective Representation Learning for Circuit Design at Scale
abstract
Circuit representation learning has become pivotal in electronic design automation, enabling critical tasks such as testability analysis, logic reasoning, power estimation, and SAT solving. However, existing models face significant challenges in scaling to large circuits due to limitations like over-squashing in graph neural networks and the quadratic complexity of transformer-based models. To address these issues, we introduce \textbf{DeepGate4}, a scalable and efficient graph transformer specifically designed for large-scale circuits. DeepGate4 incorporates several key innovations: (1) an update strategy tailored for circuit graphs, which reduce memory complexity to sub-linear and is adaptable to any graph transformer; (2) a GAT-based sparse transformer with global and local structural encodings for AIGs; and (3) an inference acceleration CUDA kernel that fully exploit the unique sparsity patterns of AIGs. Our extensive experiments on the ITC99 and EPFL benchmarks show that DeepGate4 significantly surpasses state-of-the-art methods, achieving 15.5\% and 31.1\% performance improvements over the next-best models. Furthermore, the Fused-DeepGate4 variant reduces runtime by 35.1\% and memory usage by 46.8\%, making it highly efficient for large-scale circuit analysis. These results demonstrate the potential of DeepGate4 to handle complex EDA tasks while offering superior scalability and efficiency.
Shan Huang 0010, Jianyuan Zhong, Zhengyuan Shi, Guohao Dai 0001, Ningyi Xu, Qiang Xu 0001
ICLR3
2025 Dependency Matters: Enhancing LLM Reasoning with Explicit Knowledge Grounding
abstract
Large language models (LLMs) often produce reasoning steps that are superficially coherent yet internally inconsistent, leading to unreliable outputs. Since such failures typically arise from implicit or poorly-grounded knowledge, we introduce \emph{Grounded Reasoning in Dependency (GRiD)}, a novel dependency-aware reasoning framework that explicitly grounds reasoning steps in structured knowledge. GRiD represents reasoning as a graph consisting of interconnected knowledge extraction nodes and reasoning nodes, enforcing logical consistency through explicit dependencies. Each reasoning step is validated via a lightweight, step-wise verifier that ensures logical correctness relative to its premises. Extensive experiments across diverse reasoning benchmarks—including StrategyQA, CommonsenseQA, GPQA, and TruthfulQA—demonstrate that GRiD substantially improves reasoning accuracy, consistency, and faithfulness compared to recent state-of-the-art structured reasoning methods. Notably, GRiD enhances performance even when applied purely as a lightweight verification module at inference time, underscoring its generalizability and practical utility. Code is available at: https://github.com/cure-lab/GRiD.
Xiangyu Wen 0001, Min Li 0019, Junhua Huang, Jianyuan Zhong, Zeju Li, Yongxiang Huang, Mingxuan Yuan, Qiang Xu 0001
NeurIPS4
2024 DeepGate3: Towards Scalable Circuit Representation Learning
abstract
Circuit representation learning has shown promising results in advancing the field of Electronic Design Automation (EDA). Existing models, such as DeepGate Family, primarily utilize Graph Neural Networks (GNNs) to encode circuit netlists into gate-level embeddings. However, the scalability of GNN-based models is fundamentally constrained by architectural limitations, impacting their ability to generalize across diverse and complex circuit designs. To address these challenges, we introduce DeepGate3, an enhanced architecture that integrates Transformer modules following the initial GNN processing. This novel architecture not only retains the robust gate-level representation capabilities of its predecessor, DeepGate2, but also enhances them with the ability to model subcircuits through a novel pooling transformer mechanism. DeepGate3 is further refined with multiple innovative supervision tasks, significantly enhancing its learning process and enabling superior representation of both gate-level and subcircuit structures. Our experiments demonstrate marked improvements in scalability and generalizability over traditional GNN-based approaches, establishing a significant step forward in circuit representation learning technology.
Zhengyuan Shi, Sadaf Khan, Jianyuan Zhong, Min Li 0019, Qiang Xu 0001
ICCAD4
2024 GuardT2I: Defending Text-to-Image Models from Adversarial Prompts
abstract
Recent advancements in Text-to-Image models have raised significant safety concerns about their potential misuse for generating inappropriate or Not-Safe-For-Work contents, despite existing countermeasures such as Not-Safe-For-Work classifiers or model fine-tuning for inappropriate concept removal. Addressing this challenge, our study unveils GuardT2I a novel moderation framework that adopts a generative approach to enhance Text-to-Image models’ robustness against adversarial prompts. Instead of making a binary classification, GuardT2I utilizes a large language model to conditionally transform text guidance embeddings within the Text-to-Image models into natural language for effective adversarial prompt detection, without compromising the models’ inherent performance. Our extensive experiments reveal that GuardT2I outperforms leading commercial solutions like OpenAI-Moderation and Microsoft Azure Moderator by a significant margin across diverse adversarial scenarios. Our framework is available at https://github.com/cure-lab/GuardT2I.
Ruiyuan Gao 0001, Jianyuan Zhong, Qiang Xu 0001
NeurIPS4
2021 Attention Is All You Need In Speech Separation
abstract
Recurrent Neural Networks (RNNs) have long been the dominant architecture in sequence-to-sequence learning. RNNs, however, are inherently sequential models that do not allow parallelization of their computations. Transformers are emerging as a natural alternative to standard RNNs, replacing recurrent computations with a multi-head attention mechanism.In this paper, we propose the SepFormer, a novel RNN-free Transformer-based neural network for speech separation. The Sep-Former learns short and long-term dependencies with a multi-scale approach that employs transformers. The proposed model achieves state-of-the-art (SOTA) performance on the standard WSJ0-2/3mix datasets. It reaches an SI-SNRi of 22.3 dB on WSJ0-2mix and an SI-SNRi of 19.5 dB on WSJ0-3mix. The SepFormer inherits the parallelization advantages of Transformers and achieves a competitive performance even when downsampling the encoded representation by a factor of 8. It is thus significantly faster and it is less memory-demanding than the latest speech separation systems with comparable performance.
Cem Subakan, Mirco Ravanelli, Samuele Cornell, Mirko Bronzi, Jianyuan Zhong
ICASSP5
2020 Multi-Task Self-Supervised Learning for Robust Speech Recognition
abstract
Despite the growing interest in unsupervised learning, extracting meaningful knowledge from unlabelled audio remains an open challenge. To take a step in this direction, we recently proposed a problem-agnostic speech encoder (PASE), that combines a convolutional encoder followed by multiple neural networks, called workers, tasked to solve self-supervised problems (i.e., ones that do not require manual annotations as ground truth). PASE was shown to capture relevant speech information, including speaker voice-print and phonemes. This paper proposes PASE+, an improved version of PASE for robust speech recognition in noisy and reverberant environments. To this end, we employ an online speech distortion module, that contaminates the input signals with a variety of random disturbances. We then propose a revised encoder that better learns short- and long-term speech dynamics with an efficient combination of recurrent and convolutional networks. Finally, we refine the set of workers used in self-supervision to encourage better cooperation. Results on TIMIT, DIRHA and CHiME-5 show that PASE+ significantly outperforms both the previous version of PASE as well as common acoustic features. Interestingly, PASE+ learns transferable representations suitable for highly mismatched acoustic conditions.
Mirco Ravanelli, Jianyuan Zhong, Santiago Pascual, Pawel Swietojanski, João Monteiro 0002, Jan Trmal, Yoshua Bengio
ICASSP2
2019 UR-FUNNY: A Multimodal Language Dataset for Understanding Humor
abstract
Md Kamrul Hasan, Wasifur Rahman, AmirAli Bagher Zadeh, Jianyuan Zhong, Md Iftekhar Tanveer, Louis-Philippe Morency, Mohammed (Ehsan) Hoque. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Md. Kamrul Hasan 0003, Wasifur Rahman, Amir Zadeh 0001, Jianyuan Zhong, Md. Iftekhar Tanveer, Louis-Philippe Morency, Mohammed E. Hoque 0001
EMNLP/IJCNLP (1)4