Huijia Wu

dblp:188/6224 · DBLP profile ↗
← Back
17ranked-venue papers
3as first author
14since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 HEV Generative Sandbox: A Framework for Assessing Domain-Specific Social Risks Through Human-LLM Simulation
abstract
Deploying Large Language Models (LLMs) in specialized domains introduces significant societal and compliance risks, including bias amplification, misinformation propagation, and privacy violations. These risks predominantly emerge from the dynamic interactions between LLMs and humans in specific contexts. Different domains face unique distribution of hazards, and varying interaction modalities introduce distinct levels of exposure and vulnerability. However, current risk assessment frameworks lack a systematic methodology to capture this dynamic interplay. In this work, we introduce the HEV Generative Sandbox, a novel risk evaluation framework that simulates human-LLM behavior to quantify domain-contextual risks across three interdependent dimensions: 1) Hazard (H): Domain-specific threats inherent to a given context; 2) Exposure (E): The extent to which the LLM and its users are subjected to hazardous scenarios; 3) Vulnerability (V): The susceptibility of the system to risk due to human interaction or model weaknesses. Our approach pioneers "domain-rooted scenario generation", wherein we sample contextual distributions from domain-specific corpora and simulate diverse inputs. By unifying dynamic scenario simulation, causal risk decomposition, and closed-loop evaluation, the HEV Generative Sandbox provides a scalable, domain-sensitive methodology for responsible LLM deployment. This work contributes to advancing the safe deployment of LLMs by providing a comprehensive and automated risk evaluation framework.
Zhiyi Hou, Xiaoang Xu, Shuo Wang 0013, Huijia Wu, Kaicheng Yu, Yang Yu 0011, ChengXiang Zhai
AAAI5
2026 From Chaos to Cure: A Prefix Heuristics Guided Model-Agnostic Adaptive Detoxification Framework
abstract
The impressive performance of large language models (LLMs) also brings inherent toxicity risks, prompting the need for effective detoxification to support responsible deployment. Prevailing methods generally follow an inflexible model-specific fashion, addressing only individual models or model families. Moreover, overlooking the underlying toxic risks involved in the input prefix can lead to toxic accumulation during autoregressive generation. Existing methods rely on external strong attribute interventions to address this issue, which further exacerbates contextual semantic inconsistencies and makes it difficult to balance toxicity efficacy and generation quality. To address these concerns, we propose a novel Model-Agnostic Adaptive Detoxification (MAAD) framework. To address accumulating toxicity, we present prefix heuristics that serve as contextual signals, guiding the base LLM toward safer generation. Along this line, we construct an antidote dataset to support a lightweight model, Detoxifier, which steers the base LLM to make in-scope and reliable detoxifying distribution adjustments while preserving fluency and contextual understanding. Designed as an easy-to-deploy module, Detoxifier requires a small amount of data and can be seamlessly applied to various base LLMs with one-off training. Since over-purifying often reduces diversity, we also propose a dynamic truncation method called CW-cutoff sampling to trade off language model quality and diversity. Extensive experiments demonstrate that MAAD strikes a better balance between detoxification effectiveness and generation quality, while also maintaining model utility.
Yuhu Shang, Xiang Cheng 0003, Yimeng Ren 0001, Huijia Wu, Xuexiong Luo, Kangkang Lu 0002, Jian Zhao 0018, Zhaofeng He 0001
AAAI4
2026 Interpretable Safety Alignment via SAE-Constructed Low-Rank Subspace Adaptation
abstract
Dianyun Wang, Qingsen Ma, Yuhu Shang, Zhifeng Lu, Zhenbo Xu, Lechen Ning, Huijia Wu, Zhaofeng He. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Dianyun Wang, Qingsen Ma, Yuhu Shang, Zhifeng Lu, Zhenbo Xu, Lechen Ning, Huijia Wu, Zhaofeng He 0001
ACL (1)7
2026 Rethinking Class-Incremental Learning From a Dynamic Imbalanced Learning Perspective
abstract
Deep neural networks suffer from catastrophic forgetting when continually learning new concepts. In this paper, we analyze this problem from a data imbalance point of view. We argue that the imbalance between old task and new task data contributes to forgetting of the old tasks. Moreover, the increasing imbalance ratio during incremental learning further aggravates the problem. To address the dynamic imbalance issue, we propose Uniform Prototype Contrastive Learning (UPCL), where uniform and compact features are learned. Specifically, we generate a set of non-learnable uniform prototypes before each task starts. Then we assign these uniform prototypes to each class and guide the feature learning through prototype contrastive learning. We also dynamically adjust the relative margin between old and new classes so that the feature distribution will be maintained balanced and compact. Finally, we demonstrate through extensive experiments that the proposed method achieves state-of-the-art performance on several benchmark including CIFAR-100, ImageNet-100, TinyImageNet, Food-101, and CUB-200. Experimental results show that our approach not only effectively addresses the issue of imbalanced old data in memory but also tackles the problem of imbalanced new data distributions.
Leyuan Wang, Liuyu Xiang, Yunlong Wang 0003, Huijia Wu, Huafeng Yang, Jingqian Liu, Zhaofeng He 0001
IEEE Trans. Multim.4
2025 Psyche-Wave: Fusing Vector-Quantized Morphology and LLM-Inferred Semantics from Millimeter-Wave SCG for Psychological State Decoding
abstract
This paper introduces Psyche-Wave, a novel paradigm for non-contact psychological state assessment, addressing the challenge that existing methods struggle to reconcile signal representation robustness with deep physiological semantic understanding. The proposed framework is built upon high-fidelity Seismocardiogram (SCG) and respiratory signals, captured by a proprietary high-sampling-rate millimeter-wave (mmWave) radar system. Psyche-Wave features a parallel dual-branch architecture for complementary feature extraction. The first, a Data-Driven Morphological Branch, employs Vector Quantization (VQ) to encode the Mel spectrogram of the SCG signal into a codebook-based representation, yielding a noise-resilient morphological embedding. The second, a Knowledge-Driven Semantic Branch, leverages a Large Language Model (LLM) to infer deep contextual relationships from medically significant physiological parameters—including heart rate variability, cardiac time intervals, and cardiopulmonary coupling—outputting a rich semantic embedding. These complementary embeddings are then integrated through a dedicated fusion module and passed to a downstream classifier for precise emotion and personality trait evaluation. Comprehensive evaluations on a newly collected high-fidelity dataset, referred to as mmHeart-Pro, and the public ReMAP dataset demonstrate state-of-the-art performance. This work pioneers a new path that fuses data-driven morphological analysis with knowledge-driven semantic reasoning, significantly advancing the accuracy and interpretability of non-contact psychological sensing.
Yiwei Ru, Zhenbo Xu, Yanlin Xu, Huijia Wu, Zhaofeng He 0001, Zhenan Sun
BIBM5
2025 Select-Then-Decompose: From Empirical Analysis to Adaptive Selection Strategy for Task Decomposition in Large Language Models
abstract
Large language models (LLMs) have demonstrated remarkable reasoning and planning capabilities, driving extensive research into task decomposition.Existing task decomposition methods focus primarily on memory, tool usage, and feedback mechanisms, achieving notable success in specific domains, but they often overlook the trade-off between performance and cost.In this study, we first conduct a comprehensive investigation on task decomposition, identifying six categorization schemes.Then, we perform an empirical analysis of three factors that influence the performance and cost of task decomposition: categories of approaches, characteristics of tasks, and configuration of decomposition and execution models, uncovering three critical insights and summarizing a set of practical principles.Building on this analysis, we propose the Select-Then-Decompose strategy, which establishes a closed-loop problemsolving process composed of three stages: selection, execution, and verification.This strategy dynamically selects the most suitable decomposition approach based on task characteristics and enhances the reliability of the results through a verification module.Comprehensive evaluations across multiple benchmarks show that the Select-Then-Decompose consistently lies on the Pareto frontier, demonstrating an optimal balance between performance and cost.Our code is publicly available at https://github.com/summervvind/ Select-Then-Decompose.
Shuodi Liu, Yingzhuo Liu, Zi Wang 0014, Huijia Wu, Liuyu Xiang, Zhaofeng He 0001
EMNLP5
2025 Improving Food Recognition with Retrieval-Augmented and Domain-Adaptive LVLMs
abstract
Food recognition is pivotal in enhancing intelligent food recommendation systems and nutritional management, contributing to balanced diets and overall health. Although Large Vision-Language Models (LVLMs) have demonstrated impressive performances across various domains, their performance on the food recognition task still lags behind traditional vision models. To bridge this gap, this paper proposes two methods to improve the food recognition capabilities of LVLMs: Retrieval-Augmented Recognition (RAR) and Domain-Adaptive Recognition (DAR). On the one hand, the training-free RAR utilizes a vision model to retrieve relevant image-category pairs from an image-category memory pre-built from the training set, thus incorporating the categorical information into the input of LVLMs to enhance food recognition performance. On the other hand, DAR employs a two-stage training process by first pre-training LVLMs on diverse food analysis tasks and then fine-tuning LVLMs using food recognition data. Extensive evaluations on two large-scale food recognition datasets demonstrate that both RAR and DAR improve the food recognition performance of LVLMs and, compred to RAR, DAR achieves a higher precision that outperforms traditional vision models.
Dehua Ma, Zhenbo Xu, Tianshun Xing, Huijia Wu, Zhaofeng He 0001
ICASSP6
2025 Eye Movements as Images: A Multimodal Framework for Eye Movements Representation
abstract
Eye movements are increasingly popular for enhancing natural language processing and modeling individual states. Although specialized methods have been developed to represent eye movements for various tasks, effectively modeling the complex dynamics of eye movements and the heterogeneity with stimulus text remains challenging. This paper proposes a text-guided eye movement representation framework that introduces a novel perspective by converting raw eye movement sequences into line graph images and encoding them with a powerful pre-trained vision transformer. To address the disparities between eye movements and text, we guide their temporal alignment using human reading order and combine Canonical Correlation Analysis with Optimal Transport to fuse the two modalities. This approach not only significantly simplifies the design of specialized models but also has the potential to become a universal representation for eye movements. Experimental results on six different domain tasks show that the proposed method achieves state-of-the-art performance. We release the source code at https://github.com/wulalahalala/VLEM.
Dongsen Zhang, Peipei Li 0002, Zekun Li 0001, Yiwei Ru, Huijia Wu, Zhaofeng He 0001
ICASSP5
2025 FoodWeight1.4M: A Large-scale Multi-modal Dataset for Weight Estimation
abstract
Large vision language models (VLMs) excel in visual tasks but struggle with weight estimation, hindering 3D perception and embodied intelligence. To address the lack of large-scale weight datasets, we present FoodWeight1.4M, derived from real-world supermarket scenarios. It contains 1.4 million high-quality images across 1,550 food categories, with weights precisely measured and rigorously filtered, making it the first large-scale weight estimation dataset. The weight estimation performance of current VLMs were tested and found to be unsatisfactory, which can be significantly improved by instruction tuning using Food-Weight1.4M. Moreover, we propose two strategies, Category-Guided and Reference Calibration, to enhance weight estimation without fine-tuning. Experiments confirm their effectiveness in improving multi-modal weight perception. Furthermore, experimental results show that pre-training on FoodWeight1.4M can benefit other food analysis tasks. Our dataset will be publicly available soon.
Zhenbo Xu, Dehua Ma, Liuyu Xiang, Huijia Wu, Zhaofeng He 0001
ICME6
2025 A*-Thought: Efficient Reasoning via Bidirectional Compression for Low-Resource Settings
abstract
Large Reasoning Models (LRMs) achieve superior performance by extending the thought length. However, a lengthy thinking trajectory leads to reduced efficiency. Most of the existing methods are stuck in the assumption of overthinking and attempt to reason efficiently by compressing the Chain-of-Thought, but this often leads to performance degradation. To address this problem, we introduce A*-Thought, an efficient tree search-based unified framework designed to identify and isolate the most essential thoughts from the extensive reasoning chains produced by these models. It formulates the reasoning process of LRMs as a search tree, where each node represents a reasoning span in the giant reasoning space. By combining the A* search algorithm with a cost function specific to the reasoning path, it can efficiently compress the chain of thought and determine a reasoning path with high information density and low cost. In addition, we also propose a bidirectional importance estimation mechanism, which further refines this search process and enhances its efficiency beyond uniform sampling. Extensive experiments on several advanced math tasks show that A*-Thought effectively balances performance and efficiency over a huge search space. Specifically, A*-Thought can improve the performance of QwQ-32B by 2.39$\times$ with low-budget and reduce the length of the output token by nearly 50\% with high-budget. The proposed method is also compatible with several other LRMs, demonstrating its generalization capability. The code can be accessed at: https://github.com/AI9Stars/AStar-Thought.
Xiaoang Xu, Shuo Wang 0013, Zhenghao Liu 0001, Huijia Wu, Peipei Li 0002, Zhiyuan Liu 0001, Maosong Sun 0001, Zhaofeng He 0001
NeurIPS5
2025 Memory-Constrained DiskANN: Efficient Approximate Nearest Neighbor Search Under Resource Constraints
Yuhang Lou, Linyun Ma, Yan Ruan, Huijia Wu, Minhua Lu
SISAP5
2025 Can large language models independently complete tasks? A dynamic evaluation framework for multi-turn task planning and completion
Junlin Cui, Huijia Wu, Liuyu Xiang, Xiangang Li, Yaodong Yang 0001, Zhaofeng He 0001
Neurocomputing3
2024 HyperMoE: Towards Better Mixture of Experts via Transferring Among Experts
abstract
The Mixture of Experts (MoE) for language models has been proven effective in augmenting the capacity of models by dynamically routing each input token to a specific subset of experts for processing.Despite the success, most existing methods face a challenge for balance between sparsity and the availability of expert knowledge: enhancing performance through increased use of expert knowledge often results in diminishing sparsity during expert selection.To mitigate this contradiction, we propose HyperMoE, a novel MoE framework built upon Hypernetworks.This framework integrates the computational processes of MoE with the concept of knowledge transferring in multi-task learning.Specific modules generated based on the information of unselected experts serve as supplementary information, which allows the knowledge of experts not selected to be used while maintaining selection sparsity.Our comprehensive empirical evaluations across multiple datasets and backbones establish that HyperMoE significantly outperforms existing MoE methods under identical conditions concerning the number of experts.
Zihan Qiu, Huijia Wu, Zhaofeng He 0001, Jie Fu 0001
ACL (1)3
2024 SafeLLMs: A Benchmark for Secure Bilingual Evaluation of Large Language Models
Wenhan Liang, Huijia Wu, Yuhu Shang, Zhaofeng He 0001
NLPCC (2)2
2017 A Dynamic Window Neural Network for CCG Supertagging
abstract
Combinatory Category Grammar (CCG) supertagging is a task to assign lexical categories to each word in a sentence. Almost all previous methods use fixed context window sizes to encode input tokens. However, it is obvious that different tags usually rely on different context window sizes. This motivates us to build a supertagger with a dynamic window approach, which can be treated as an attention mechanism on the local contexts. We find that applying dropout on the dynamic filters is superior to the regular dropout on word embeddings. We use this approach to demonstrate the state-of-the-art CCG supertagging performance on the standard test set.
Huijia Wu, Jiajun Zhang 0001, Chengqing Zong
AAAI1
2017 Shortcut Sequence Tagging
Huijia Wu, Jiajun Zhang 0001, Chengqing Zong
NLPCC1
2016 An Empirical Exploration of Skip Connections for Sequential Tagging
abstract
In this paper, we empirically explore the effects of various kinds of skip connections in stacked bidirectional LSTMs for sequential tagging. We investigate three kinds of skip connections connecting to LSTM cells: (a) skip connections to the gates, (b) skip connections to the internal states and (c) skip connections to the cell outputs. We present comprehensive experiments showing that skip connections to cell outputs outperform the remaining two. Furthermore, we observe that using gated identity functions as skip mappings works pretty well. Based on this novel skip connections, we successfully train deep stacked bidirectional LSTM models and obtain state-of-the-art results on CCG supertagging and comparable results on POS tagging.
Huijia Wu, Jiajun Zhang 0001, Chengqing Zong
COLING1