Yukang Yang

dblp:274/6411 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
9since 2021 · last 2025
0000-0001-6807-3052ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 4 first-author · 9 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Generative modeling · 31% Trustworthy machine learning · 20% Knowledge representation and reasoning · 12%
Computer graphics and multimedia
1 paper
Visual content generation and editing · 77% Image and video coding · 23%

Topics — the 19 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
2.332025
Training-Free Guidance Beyond Differentiability: Scalable Path Steering with Tree Search in Diffusion and Flow Models · NeurIPS 2025
Gradient Guidance for Diffusion Models: An Optimization Perspective · NeurIPS 2024
GlyphControl: Glyph Conditional Control for Visual Text Generation · NeurIPS 2023
Machine learning › Trustworthy machine learning › interpretability
mechanistic interpretability
1.122025
Causal Head Gating: A Framework for Interpreting Roles of Attention Heads in Transformers · NeurIPS 2025
Emergent Symbolic Mechanisms Support Abstract Reasoning in Large Language Models · ICML 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning
abstract reasoning
0.912025
Emergent Symbolic Mechanisms Support Abstract Reasoning in Large Language Models · ICML 2025
Machine learning › Trustworthy machine learning
interpretability
0.912025
Causal Head Gating: A Framework for Interpreting Roles of Attention Heads in Transformers · NeurIPS 2025
Natural language and speech › Language models and text generation
large language model
0.912025
Emergent Symbolic Mechanisms Support Abstract Reasoning in Large Language Models · ICML 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning › logic-based reasoning
symbolic reasoning
0.912025
Emergent Symbolic Mechanisms Support Abstract Reasoning in Large Language Models · ICML 2025
Machine learning › Generative modeling › diffusion model › guided diffusion
training-free guidance
0.912025
Training-Free Guidance Beyond Differentiability: Scalable Path Steering with Tree Search in Diffusion and Flow Models · NeurIPS 2025
Machine learning › Trustworthy machine learning › interpretability › neural network interpretation
transformer interpretability
0.912025
Causal Head Gating: A Framework for Interpreting Roles of Attention Heads in Transformers · NeurIPS 2025
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
tree search
0.912025
Training-Free Guidance Beyond Differentiability: Scalable Path Steering with Tree Search in Diffusion and Flow Models · NeurIPS 2025
Machine learning › Optimization for machine learning
convergence analysis
0.812024
Gradient Guidance for Diffusion Models: An Optimization Perspective · NeurIPS 2024
Machine learning › Generative modeling › diffusion model
guided diffusion
0.812024
Gradient Guidance for Diffusion Models: An Optimization Perspective · NeurIPS 2024
Machine learning › Optimization for machine learning › optimization
optimization theory
0.812024
Gradient Guidance for Diffusion Models: An Optimization Perspective · NeurIPS 2024
Computer vision › Image recognition and object detection › object detection
detection transformer
0.712023
Rank-DETR for High Quality Object Detection · NeurIPS 2023
Computer vision › Image recognition and object detection
object detection
0.712023
Rank-DETR for High Quality Object Detection · NeurIPS 2023
Machine learning › Learning theory
ranking
0.712023
Rank-DETR for High Quality Object Detection · NeurIPS 2023
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.712023
GlyphControl: Glyph Conditional Control for Visual Text Generation · NeurIPS 2023
Visual content generation and editing
visual text generation
0.712023
GlyphControl: Glyph Conditional Control for Visual Text Generation · NeurIPS 2023
Machine learning › Deep learning architectures and training
attention mechanism
0.312025
Causal Head Gating: A Framework for Interpreting Roles of Attention Heads in Transformers · NeurIPS 2025
Image and video coding
quality assessment
0.212023
GlyphControl: Glyph Conditional Control for Visual Text Generation · NeurIPS 2023

Methods — techniques the papers use, named apart from their topics

diffusion model · 2.2tree search · 0.9flow model · 0.9circuit analysis · 0.9causal mediation analysis · 0.9causal head gating · 0.9attention head analysis · 0.9ablation · 0.9gradient guidance · 0.8first-order optimization · 0.8glyph conditioning · 0.7
YearPublicationVenuePosition
2025 Emergent Symbolic Mechanisms Support Abstract Reasoning in Large Language Models
abstract
Many recent studies have found evidence for emergent reasoning capabilities in large language models (LLMs), but debate persists concerning the robustness of these capabilities, and the extent to which they depend on structured reasoning mechanisms. To shed light on these issues, we study the internal mechanisms that support abstract reasoning in LLMs. We identify an emergent symbolic architecture that implements abstract reasoning via a series of three computations. In early layers, symbol abstraction heads convert input tokens to abstract variables based on the relations between those tokens. In intermediate layers, symbolic induction heads perform sequence induction over these abstract variables. Finally, in later layers, retrieval heads predict the next token by retrieving the value associated with the predicted abstract variable. These results point toward a resolution of the longstanding debate between symbolic and neural network approaches, suggesting that emergent reasoning in neural networks depends on the emergence of symbolic mechanisms.
Yukang Yang, Declan Campbell, Kaixuan Huang, Mengdi Wang 0001, Jonathan D. Cohen 0003, Taylor W. Webb
ICML1
2025 Training-Free Guidance Beyond Differentiability: Scalable Path Steering with Tree Search in Diffusion and Flow Models
abstract
Training-free guidance enables controlled generation in diffusion and flow models, but most methods rely on gradients and assume differentiable objectives. This work focuses on training-free guidance addressing challenges from non-differentiable objectives and discrete data distributions. We propose TreeG: Tree Search-Based Path Steering Guidance, applicable to both continuous and discrete settings in diffusion and flow models. TreeG offers a unified framework for training-free guidance by proposing, evaluating, and selecting candidates at each step, enhanced with tree search over active paths and parallel exploration. We comprehensively investigate the design space of TreeG over the candidate proposal module and the evaluation function, instantiating TreeG into three novel algorithms. Our experiments show that TreeG consistently outperforms top guidance baselines in symbolic music generation, small molecule design, and enhancer DNA design with improvements of 29.01%, 26.38%, and 18.43%. Additionally, we identify an inference-time scaling law showing TreeG's scalability in inference-time computation.
Yukang Yang, Hui Yuan 0002, Mengdi Wang 0001
NeurIPS2
2025 Causal Head Gating: A Framework for Interpreting Roles of Attention Heads in Transformers
abstract
We present causal head gating (CHG), a scalable method for interpreting the functional roles of attention heads in transformer models. CHG learns soft gates over heads and assigns them a causal taxonomy—facilitating, interfering, or irrelevant—based on their impact on task performance. Unlike prior approaches in mechanistic interpretability, which are hypothesis-driven and require prompt templates or target labels, CHG applies directly to any dataset using standard next-token prediction. We evaluate CHG across multiple large language models (LLMs) in the Llama 3 model family and diverse tasks, including syntax, commonsense, and mathematical reasoning, and show that CHG scores yield causal, not merely correlational, insight validated via ablation and causal mediation analyses. We also introduce contrastive CHG, a variant that isolates sub-circuits for specific task components. Our findings reveal that LLMs contain multiple sparse task-sufficient sub-circuits, that individual head roles depend on interactions with others (low modularity), and that instruction following and in-context learning rely on separable mechanisms.
Andrew Nam, Henry Conklin, Yukang Yang, Thomas L. Griffiths 0001, Jonathan D. Cohen 0003, Sarah-Jane Leslie
NeurIPS3
2025 Anatomical prior-based vertebral landmark detection for spinal disorder diagnosis
Yukang Yang, Ming Sun 0019, Shiji Song, Gao Huang 0001
Artif. Intell. Medicine1
2024 Gradient Guidance for Diffusion Models: An Optimization Perspective
abstract
Diffusion models have demonstrated empirical successes in various applications and can be adapted to task-specific needs via guidance. This paper studies a form of gradient guidance for adapting a pre-trained diffusion model towards optimizing user-specified objectives. We establish a mathematical framework for guided diffusion to systematically study its optimization theory and algorithmic design. Our theoretical analysis spots a strong link between guided diffusion models and optimization: gradient-guided diffusion models are essentially sampling solutions to a regularized optimization problem, where the regularization is imposed by the pre-training data. As for guidance design, directly bringing in the gradient of an external objective function as guidance would jeopardize the structure in generated samples. We investigate a modified form of gradient guidance based on a forward prediction loss, which leverages the information in pre-trained score functions and provably preserves the latent structure. We further consider an iteratively fine-tuned version of gradient-guided diffusion where guidance and score network are both updated with newly generated samples. This process mimics a first-order optimization iteration in expectation, for which we proved $\tilde{\mathcal{O}}(1/K)$ convergence rate to the global optimum when the objective function is concave. Our code is released at https://github.com/yukang123/GGDMOptim.git.
Hui Yuan 0002, Yukang Yang, Minshuo Chen, Mengdi Wang 0001
NeurIPS3
2023 Rank-DETR for High Quality Object Detection
abstract
Modern detection transformers (DETRs) use a set of object queries to predict a list of bounding boxes, sort them by their classification confidence scores, and select the top-ranked predictions as the final detection results for the given input image. A highly performant object detector requires accurate ranking for the bounding box predictions. For DETR-based detectors, the top-ranked bounding boxes suffer from less accurate localization quality due to the misalignment between classification scores and localization accuracy, thus impeding the construction of high-quality detectors. In this work, we introduce a simple and highly performant DETR-based object detector by proposing a series of rank-oriented designs, combinedly called Rank-DETR. Our key contributions include: (i) a rank-oriented architecture design that can prompt positive predictions and suppress the negative ones to ensure lower false positive rates, as well as (ii) a rank-oriented loss function and matching cost design that prioritizes predictions of more accurate localization accuracy during ranking to boost the AP under high IoU thresholds. We apply our method to improve the recent SOTA methods (e.g., H-DETR and DINO-DETR) and report strong COCO object detection results when using different backbones such as ResNet-$50$, Swin-T, and Swin-L, demonstrating the effectiveness of our approach. Code is available at \url{https://github.com/LeapLabTHU/Rank-DETR}.
Yifan Pu, Weicong Liang, Yiduo Hao, Yuhui Yuan, Yukang Yang, Chao Zhang 0001, Han Hu 0001, Gao Huang 0001
NeurIPS5
2023 GlyphControl: Glyph Conditional Control for Visual Text Generation
abstract
Recently, there has been an increasing interest in developing diffusion-based text-to-image generative models capable of generating coherent and well-formed visual text. In this paper, we propose a novel and efficient approach called GlyphControl to address this task. Unlike existing methods that rely on character-aware text encoders like ByT5 and require retraining of text-to-image models, our approach leverages additional glyph conditional information to enhance the performance of the off-the-shelf Stable-Diffusion model in generating accurate visual text. By incorporating glyph instructions, users can customize the content, location, and size of the generated text according to their specific requirements. To facilitate further research in visual text generation, we construct a training benchmark dataset called LAION-Glyph. We evaluate the effectiveness of our approach by measuring OCR-based metrics, CLIP score, and FID of the generated visual text. Our empirical evaluations demonstrate that GlyphControl outperforms the recent DeepFloyd IF approach in terms of OCR accuracy, CLIP score, and FID, highlighting the efficacy of our method.
Yukang Yang, Dongnan Gui, Yuhui Yuan, Weicong Liang, Haisong Ding, Han Hu 0001, Kai Chen 0001
NeurIPS1
2022 ASD-ISSPA: Adaptive Stochastic Droppath and Interactive Slow Semi-Polarized Attention based on FPN for Object Detection
abstract
Feature Pyramid Network (FPN) has been widely used to combine features at different scales for extracting semantically strong representations in object detection. However, most of the previous approaches only focus on the original image resolution to allocate pyramidal layers without considering the image size change in the processing network. In this work, an algorithm is proposed to determine the feature layer based on the ratio of the target box area to the original image area, which improves the network adaptability to the scale change. Further-more, a novel Adaptive Stochastic Droppath and Interactive Slow Semi-Polarized Attention (ASD-ISSPA) is proposed to retain the valuable information in the feature maps of FPN. ASD-ISSPA can be split into two parts, Adaptive Stochastic Droppath Attention (ASDA) and Interactive Slow Semi-Polarized Attention (ISSPA). The ASDA module acts on the bottom-up pathway in the pyramid to collect the loss information after convolution, while the ISSPA module deals with the top-down path in the pyramid to capture the semantics before upsampling and the interaction details after upsampling. The integrated two-path attention mechanism effectively regains the loss information of the convolution and removes the redundancy of the deconvolution so that our model can improve the category prediction accuracy and reduce the discrepancy between the predicted and ground-truth boxes. Compared with existing methods, the proposed attention extract network, called ASD-ISSPA, achieves competitive results on the PASCAL VOC dataset.
Yukang Yang, Lu Lu 0011
IJCNN1
2022 A multi-scale keypoint estimation network with self-supervision for spinal curvature assessment of idiopathic scoliosis from the imperfect dataset
Yukang Yang, Ming Sun 0019, Cody Bunger, Cheng Wu 0002
Artif. Intell. Medicine3