Zhen Huang 0007

dblp:22/3870-7 · DBLP profile ↗
← Back
23ranked-venue papers
4as first author
22since 2021 · last 2026
0000-0003-0902-2321ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 2 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 3 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Bridging the Language Gap: Uncovering and Aligning Shared Circuits for Multi-Hop Reasoning in Multilingual LLMs
abstract
Large language models (LLMs) present a paradox: they can correctly answer a multi-hop factual query in a high-resource language like English, yet fail on the identical query in another language. This raises a fundamental question about the nature of multilingual knowledge: are facts missing, or merely inaccessible? The underlying mechanisms for this knowledge gap have remained largely unexplored. In this work, we resolve this question by introducing a mechanistic interpretability framework that traces the causal pathways of multi-hop knowledge reasoning. Our analysis reveals a core, non-obvious finding: cross-lingual inconsistencies do not stem from a knowledge deficit. Instead, factual knowledge is robustly stored in a set of **shared, language-agnostic semantic neurons**. The failure originates from **misaligned attention pathways**, where a common set of critical attention heads fails to correctly route information along the reasoning chain to the appropriate knowledge neurons in lower-resource languages. This mechanistic diagnosis motivates a targeted alignment strategy: a surgical fine-tuning of only these critical heads. Experiments demonstrate that our method achieves significant improvements in multilingual multi-hop factuality—with positive cross-lingual transfer—while uniquely preserving general model capabilities, offering a scalable and mechanistically-grounded approach to building more reliable multilingual models.
Zhen Huang 0007, Yonggang Zhang 0003, Xinmei Tian 0001, Xu Shen 0001, Jieping Ye
AAAI2
2026 Mixture-of-Trees: Learning to Select and Weigh Reasoning Paths for Efficient LLM Inference
abstract
We introduce Mixture-of-Trees (MoT), a novel framework that integrates sparse expert activation with structured tree-based reasoning for efficient LLM inference. MoT employs a learned gating mechanism to selectively activate only the most relevant expert reasoning trees for each problem, where experts use models of varying capacities based on task complexity. The framework features three key innovations: (1) sparse expert activation through unified gating networks, (2) specialized expert trees that leverage domain-specific expertise while optimizing the quality-efficiency trade-off, and (3) collaborative debate mechanisms for conflicting solutions. Additionally, MoT includes a shared baseline tree with early stopping—activated experts perform lightweight validation and terminate early when confidence is high. Experiments across five benchmarks (GSM8K, MATH, AIME 2024, MMLU, HotpotQA) show that MoT achieves 2-7 percentage point accuracy improvements while reducing LLM calls by 37-40% compared to existing multi-path methods.
Yangbo Wei, Zhen Huang 0007, Shaoqiang Lu, Junhong Qian, Dongge Qin, Ting-Jung Lin, Wei W. Xing, Lei He 0001
AAAI2
2026 MedAtlas: Evaluating LLMs for Multi-Round, Multi-Task Medical Reasoning Across Diverse Imaging Modalities and Clinical Text
abstract
Artificial intelligence has demonstrated significant potential in clinical decision-making; however, developing models capable of adapting to diverse real-world scenarios and performing complex diagnostic reasoning remains a major challenge. Existing medical multi-modal benchmarks are typically limited to single-image, single-turn tasks, lacking multi-modal medical image integration and failing to capture the longitudinal and multi-modal interactive nature inherent to clinical practice. To address this gap, we introduce MedAtlas, a novel benchmark framework designed to evaluate large language models on realistic medical reasoning tasks. MedAtlas is characterized by four key features: multi-round visual question answering (VQA), Joint reasoning of multiple modalities of medical images, multi-task integration, and high clinical fidelity. It supports four core tasks: open-ended multi-round VQA, closed-ended multi-round VQA, multi-image joint reasoning, and comprehensive disease diagnosis. Each case is derived from real diagnostic workflows and incorporates temporal interactions between textual medical histories and multiple imaging modalities, including CT, MRI, PET, ultrasound, X-ray, etc., requiring models to perform deep integrative reasoning across images and clinical texts. MedAtlas provides expert-annotated gold standards for all tasks. Furthermore, we propose two novel evaluation metrics: Stage Chain Accuracy (SCA) and Error Propagation Suppression Coefficient (EPSC). Benchmark results with existing multi-modal models reveal substantial performance gaps in multi-stage clinical reasoning. MedAtlas establishes a challenging evaluation platform to advance the development of robust and trustworthy medical AI.
Ronghao Xu, Zhen Huang 0007, Yangbo Wei, Xiaoqian Zhou, Zihang Jiang, Shaohua Kevin Zhou
AAAI2
2026 Cross-Modal Masked Compositional Concept Modeling for Enhancing Visio-Linguistic Compositionality
abstract
Contrastively trained vision-language models like CLIP, have made remarkable progress in learning joint image-text representations, but still face challenges in compositional understanding.They often exhibit a "bag-of-words" behavior-struggling to capture the object relations, attribute-object bindings, and word order dependencies.This limitation arises not only from the reliance on global, single-vector representations for optimization, but also from the insufficient exploitation and modeling of the rich compositional information inherently present in paired image text data.In this work, we propose MACCO (MAsked Compositional Concept MOdeling), a framework that masks compositional concepts in one modality and reconstructs them conditioned on the full contextual information from the other, enabling the model to capture and align cross-modal compositional structures more effectively.To facilitate this process, we introduce two auxiliary objectives that jointly align and regularize masked features both inter-modally and intramodally.Extensive experiments on five compositional benchmarks, along with in-depth analyses, demonstrate that our approach not only significantly enhances compositionality in VLMs but also improves their ability to capture syntactic structure and linguistic information.Additionally, the improved compositionality also benefits text-to-image generation and multimodal large language model.
Wei Li 0317, Zhen Huang 0007, Xinmei Tian 0001
ACL (1)2
2026 VFlow: Discovering Optimal Agentic Workflows for Verilog Generation
Yangbo Wei, Zhen Huang 0007, Lei He 0001, Ting-Jung Lin, Wei W. Xing
ASP-DAC2
2026 PASS-Tr: PAtch-wise swin slice attention to leverage generalization of 2D large vision model to universal lesion detection
Jingsong Liu, Zhen Huang 0007, Xun Ma, Peter J. Schüffler, Nassir Navab, Shaohua Kevin Zhou
Medical Image Anal.3
2025 Visual Evidence Prompting Mitigates Hallucinations in Large Vision-Language Models
abstract
Wei Li, Zhen Huang, Houqiang Li, Le Lu, Yang Lu, Xinmei Tian, Xu Shen, Jieping Ye. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Wei Li 0317, Zhen Huang 0007, Houqiang Li, Le Lu 0001, Xinmei Tian 0001, Xu Shen 0001, Jieping Ye
ACL (1)2
2025 Interpret and Improve In-Context Learning via the Lens of Input-Label Mappings
abstract
Chenghao Sun, Zhen Huang, Yonggang Zhang, Le Lu, Houqiang Li, Xinmei Tian, Xu Shen, Jieping Ye. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Zhen Huang 0007, Yonggang Zhang 0003, Le Lu 0001, Houqiang Li, Xinmei Tian 0001, Xu Shen 0001, Jieping Ye
ACL (1)2
2025 Self-Attention to Operator Learning-based 3D-IC Thermal Simulation
abstract
Thermal management in 3D ICs is increasingly challenging due to higher power densities. Traditional PDESolving based methods, while accurate, are too slow for iterative design. Machine learning approaches like FNO provide faster alternatives but suffer from high-frequency information loss and high-fidelity data dependency. We introduce Self-Attention UNet Fourier Neural Operator (SAU-FNO), a novel framework combining self-attention and U-Net with FNO to capture longrange dependencies and model local high-frequency features effectively. Transfer learning is employed to fine-tune low-fidelity data, minimizing the need for extensive high-fidelity datasets and speeding up training. Experiments demonstrate that SAUFNO achieves state-of-the-art thermal prediction accuracy and provides an $842 \times$ speedup over traditional FEM methods, making it an efficient tool for advanced 3D IC thermal simulations.
Zhen Huang 0007, Wenkai Yang, Muxi Tang, Depeng Xie, Ting-Jung Lin, Yu Zhang 0086, Wei W. Xing, Lei He 0001
DAC1
2025 A Unified Framework for Few-Shot Medical Image Classification via Multi-agent Description Generation and Refined Contrastive Learning
Shenghao Chen, Zhen Huang 0007, Xiaoqian Zhou, Shaohua Kevin Zhou
ICIC (28)2
2025 Eliminating Ambiguities in One-Shot Medical Landmark Detection via Mask Drawing
Zhen Huang 0007, Xiaoqian Zhou, Shaohua Kevin Zhou
ICIC (5)2
2025 PATE: Enhancing Few-Shot Pathological Image Classification via Prompt-Based Text-Image Embedding Adaptation
Shenghao Chen, Zhen Huang 0007, Xiaoqian Zhou, Chunjiang Wang, Shaohua Kevin Zhou
MICCAI (6)2
2025 SimCroP: Radiograph Representation Learning with Similarity-Driven Cross-Granularity Pre-training
Rongsheng Wang 0003, Fenghe Tang, Qingsong Yao, Rui Yan 0009, Zhen Huang 0007, Haoran Lai, Zhiyang He, Xiaodong Tao, Zihang Jiang, Shaohua Kevin Zhou
MICCAI (5)6
2025 U-RWKV: Lightweight Medical Image Segmentation with Direction-Adaptive RWKV
Hongbo Ye, Fenghe Tang, Peiang Zhao, Zhen Huang 0007, Dexin Zhao, Minghao Bian, Shaohua Kevin Zhou
MICCAI (11)4
2025 ModelGen: Automating Semiconductor Parameter Extraction with Large Language Model Agents
abstract
Device models require large numbers of parameters to characterize complex physical effects. Although the latest advancements in machine learning and automated tools have drastically improved efficiency over the classic methods, they still demand a considerable amount of human intervention in the loop to gain accuracy. This drastically limits further automation. Inspired by the success of Multimodal Large Language Models (MLLMs) in addressing tasks across diverse fields, we propose ModelGen, the first in-depth study to leverage MLLMs with RAG (Retrieval-Augmented Generation) to significantly reduce human effort in parameter extraction for compact model. Our contributions include (1) Automated Agentic Workflow Construction that learns to build and refine extraction workflows through iterative optimization, (2) MLLM Judge, a visual scoring mechanism that evaluates fitting quality using actual device characteristic plots rather than simple numerical metrics, and (3) Model-specific RAG for providing relevant domain knowledge during the extraction process. Experimental results demonstrate that ModelGen achieves a 26.8%–33.1% improvement in pass@1,3,5 compared to base LLM methods. The system completes complex model extractions for BSIMs and ASM-HEMT in hours (up to 168× faster) rather than days or weeks, making parameter extraction more accessible to non-experts while maintaining professional engineer-level accuracy.
Yangbo Wei, Zhanfei Chen, Jinlong Yan, Ting-Jung Lin, Zhen Huang 0007, Wei W. Xing, Lei He 0001
ACM Trans. Design Autom. Electr. Syst.7
2024 Interpretable Composition Attribution Enhancement for Visio-linguistic Compositional Understanding
abstract
Contrastively trained vision-language models such as CLIP have achieved remarkable progress in vision and language representation learning.Despite the promising progress, their proficiency in compositional reasoning over attributes and relations (e.g., distinguishing between "the car is underneath the person" and "the person is underneath the car") remains notably inadequate.We investigate the cause for this deficient behavior is the composition attribution issue, where the attribution scores (e.g., attention scores or GradCAM scores) for relations (e.g., underneath) or attributes (e.g., red) in the text are substantially lower than those for object terms.In this work, we show such issue is mitigated via a novel framework called CAE (Composition Attribution Enhancement).This generic framework incorporates various interpretable attribution methods to encourage the model to pay greater attention to composition words denoting relationships and attributes within the text.Detailed analysis shows that our approach enables the models to adjust and rectify the attribution of the texts.Extensive experiments across seven benchmarks reveal that our framework significantly enhances the ability to discern intricate details and construct more sophisticated interpretations of combined visual and linguistic elements.
Wei Li 0317, Zhen Huang 0007, Xinmei Tian 0001, Le Lu 0001, Houqiang Li, Xu Shen 0001, Jieping Ye
EMNLP2
2024 From Yes-Men to Truth-Tellers: Addressing Sycophancy in Large Language Models with Pinpoint Tuning
abstract
Large Language Models (LLMs) tend to prioritize adherence to user prompts over providing veracious responses, leading to the sycophancy issue. When challenged by users, LLMs tend to admit mistakes and provide inaccurate responses even if they initially provided the correct answer. Recent works propose to employ supervised fine-tuning (SFT) to mitigate the sycophancy issue, while it typically leads to the degeneration of LLMs' general capability. To address the challenge, we propose a novel supervised pinpoint tuning (SPT), where the region-of-interest modules are tuned for a given objective. Specifically, SPT first reveals and verifies a small percentage (<5%) of the basic modules, which significantly affect a particular behavior of LLMs. i.e., sycophancy. Subsequently, SPT merely fine-tunes these identified modules while freezing the rest. To verify the effectiveness of the proposed SPT, we conduct comprehensive experiments, demonstrating that SPT significantly mitigates the sycophancy issue of LLMs (even better than SFT). Moreover, SPT introduces limited or even no side effects on the general capability of LLMs. Our results shed light on how to precisely, effectively, and efficiently explain and improve the targeted ability of LLMs.
Wei Chen 0005, Zhen Huang 0007, Liang Xie 0003, Binbin Lin 0001, Houqiang Li, Le Lu 0001, Xinmei Tian 0001, Deng Cai 0001, Yonggang Zhang 0003, Wenxiao Wang 0001, Xu Shen 0001, Jieping Ye
ICML2
2023 Teaching What You Should Teach: A Data-Based Distillation Method
abstract
In real teaching scenarios, an excellent teacher always teaches what he (or she) is good at but the student is not. This gives the student the best assistance in making up for his (or her) weaknesses and becoming a good one overall. Enlightened by this, we introduce the "Teaching what you Should Teach" strategy into a knowledge distillation framework, and propose a data-based distillation method named "TST" that searches for desirable augmented samples to assist in distilling more efficiently and rationally. To be specific, we design a neural network-based data augmentation module with priori bias to find out what meets the teacher's strengths but the student's weaknesses, by learning magnitudes and probabilities to generate suitable data samples. By training the data augmentation module and the generalized distillation paradigm alternately, a student model is learned with excellent generalization ability. To verify the effectiveness of our method, we conducted extensive comparative experiments on object recognition, detection, and segmentation tasks. The results on the CIFAR-100, ImageNet-1k, MS-COCO, and Cityscapes datasets demonstrate that our method achieves state-of-the-art performance on almost all teacher-student pairs. Furthermore, we conduct visualization studies to explore what magnitudes and probabilities are needed for the distillation process.
Shitong Shao, Huanran Chen, Zhen Huang 0007, Linrui Gong, Shuai Wang 0048, Xinxiao Wu
IJCAI3
2022 Cloth-Changing Person Re-identification from A Single Image with Gait Prediction and Regularization
abstract
Cloth-Changing person re-identification (CC-ReID) aims at matching the same person across different locations over a long-duration, e.g., over days, and therefore inevitably has cases of changing clothing. In this paper, we focus on handling well the CC-ReID problem under a more challenging setting, i.e., just from a single image, which enables an efficient and latency-free person identity matching for surveillance. Specifically, we introduce Gait recognition as an auxiliary task to drive the Image ReID model to learn cloth-agnostic representations by leveraging personal unique and cloth-independent gait information, we name this framework as GI-ReID. GI-ReID adopts a two-stream architecture that consists of an image ReID-Stream and an auxiliary gait recognition stream (Gait-Stream). The Gait-Stream, that is discarded in the inference for high efficiency, acts as a regulator to encourage the ReID-Stream to capture cloth-invariant biometric motion features during the training. To get temporal continuous motion cues from a single image, we design a Gait Sequence Prediction (GSP) module for Gait-Stream to enrich gait information. Finally, a semantics consistency constraint over two streams is enforced for effective knowledge regularization. Extensive experiments on multiple image-based Cloth-Changing ReID benchmarks, e.g., LTCC, PRCC, Real28, and VC-Clothes, demonstrate that GI-ReID performs favorably against the state-of-the-art methods.
Xin Jin 0014, Tianyu He, Kecheng Zheng, Zhiheng Yin, Xu Shen 0001, Zhen Huang 0007, Ruoyu Feng 0001, Jianqiang Huang 0001, Zhibo Chen 0001, Xian-Sheng Hua 0001
CVPR6
2022 Exploring Fourier Prior for Single Image Rain Removal
abstract
Deep convolutional neural networks (CNNs) have become dominant in the task of single image rain removal. Most of current CNN methods, however, suffer from the problem of overfitting on one single synthetic dataset as they neglect the intrinsic prior of the physical properties of rain streaks. To address this issue, we propose a simple but effective prior - Fourier prior to improve the generalization ability of an image rain removal model. The Fourier prior is a kind of property of rainy images. It is based on a key observation of us - replacing the Fourier amplitude of rainy images with that of clean images greatly suppresses the synthetic and real-world rain streaks. This means the amplitude contains most of the rain streak information and the phase keeps the similar structures of the background. So it is natural for single image rain removal to process the amplitude and phase information of the rainy images separately. In this paper, we develop a two-stage model where the first stage restores the amplitude of rainy images to clean rain streaks, and the second stage restores the phase information to refine fine-grained background structures. Extensive experiments on synthetic rainy data demonstrate the power of Fourier prior. Moreover, when trained on synthetic data, a robust generalization ability to real-world images can also be obtained. The code will be publicly available at https://github.com/willinglucky/ExploringFourier-Prior-for-Single-Image-Rain-Removal.
Xin Guo 0018, Xueyang Fu, Man Zhou 0003, Zhen Huang 0007, Jialun Peng, Zhengjun Zha
IJCAI4
2021 Revisiting Knowledge Distillation: An Inheritance and Exploration Framework
abstract
Knowledge Distillation (KD) is a popular technique to transfer knowledge from a teacher model or ensemble to a student model. Its success is generally attributed to the privileged information on similarities/consistency between the class distributions or intermediate feature representations of the teacher model and the student model. However, directly pushing the student model to mimic the probabilities/features of the teacher model to a large extent limits the student model in learning undiscovered knowledge/features. In this paper, we propose a novel inheritance and exploration knowledge distillation framework (IE-KD), in which a student model is split into two parts - inheritance and exploration. The inheritance part is learned with a similarity loss to transfer the existing learned knowledge from the teacher model to the student model, while the exploration part is encouraged to learn representations different from the inherited ones with a dis-similarity loss. Our IE-KD framework is generic and can be easily combined with existing distillation or mutual learning methods for training deep neural networks. Extensive experiments demonstrate that these two parts can jointly push the student model to learn more diversified and effective representations, and our IE-KD can be a general technique to improve the student network to achieve SOTA performance. Furthermore, by applying our IE-KD to the training of two networks, the performance of both can be improved w.r.t. deep mutual learning.
Zhen Huang 0007, Xu Shen 0001, Jun Xing, Tongliang Liu, Xinmei Tian 0001, Houqiang Li, Bing Deng, Jianqiang Huang 0001, Xian-Sheng Hua 0001
CVPR1
2021 3D Local Convolutional Neural Networks for Gait Recognition
abstract
The goal of gait recognition is to learn the unique spatiotemporal pattern about the human body shape from its temporal changing characteristics. As different body parts behave differently during walking, it is intuitive to model the spatio-temporal patterns of each part separately. However, existing part-based methods equally divide the feature maps of each frame into fixed horizontal stripes to get local parts. It is obvious that these stripe partition-based methods cannot accurately locate the body parts. First, different body parts can appear at the same stripe (e.g., arms and the torso), and one part can appear at different stripes in different frames (e.g., hands). Second, different body parts possess different scales, and even the same part in different frames can appear at different locations and scales. Third, different parts also exhibit distinct movement patterns (e.g., at which frame the movement starts, the position change frequency, how long it lasts). To overcome these issues, we propose novel 3D local operations as a generic family of building blocks for 3D gait recognition backbones. The proposed 3D local operations support the extraction of local 3D volumes of body parts in a sequence with adaptive spatial and temporal scales, locations and lengths. In this way, the spatio-temporal patterns of the body parts are well learned from the 3D local neighborhood in partspecific scales, locations, frequencies and lengths. Experiments demonstrate that our 3D local convolutional neural networks achieve state-of-the-art performance on popular gait datasets. Code is available at: https://github.com/yellowtownhz/3DLocalCNN.
Zhen Huang 0007, Dixiu Xue, Xu Shen 0001, Xinmei Tian 0001, Houqiang Li, Jianqiang Huang 0001, Xian-Sheng Hua 0001
ICCV1
2020 Spatio-Temporal Inception Graph Convolutional Networks for Skeleton-Based Action Recognition
abstract
Skeleton-based human action recognition has attracted much attention with the prevalence of accessible depth sensors. Recently, graph convolutional networks (GCNs) have been widely used for this task due to their powerful capability to model graph data. The topology of the adjacency graph is a key factor for modeling the correlations of the input skeletons. Thus, previous methods mainly focus on the design/learning of the graph topology. But once the topology is learned, only a single-scale feature and one transformation exist in each layer of the networks. Many insights, such as multi-scale information and multiple sets of transformations, that have been proven to be very effective in convolutional neural networks (CNNs), have not been investigated in GCNs. The reason is that, due to the gap between graph-structured skeleton data and conventional image/video data, it is very challenging to embed these insights into GCNs. To overcome this gap, we reinvent the split-transform-merge strategy in GCNs for skeleton sequence processing. Specifically, we design a simple and highly modularized graph convolutional network architecture for skeleton-based action recognition. Our network is constructed by repeating a building block that aggregates multi-granularity information from both the spatial and temporal paths. Extensive experiments demonstrate that our network outperforms state-of-the-art methods by a significant margin with only 1/5 of the parameters and 1/10 of the FLOPs.
Zhen Huang 0007, Xu Shen 0001, Xinmei Tian 0001, Houqiang Li, Jianqiang Huang 0001, Xian-Sheng Hua 0001
ACM Multimedia1