Jiacong Hu

dblp:136/3061 · DBLP profile ↗
← Back
16ranked-venue papers
6as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 4 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 2 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Enhancing Attention Patterns in Vision Transformers for Robustness
Haofei Zhang, Hanyang Yuan, Haoze Jiang, Jiacong Hu, Shengxuming Zhang, Mingli Song
ICIC (7)7
2025 Beyond the Label: Unveiling Fairness through Dynamic Attribute Projections in Classification
abstract
Image classification has been widely adopted in critical applications such as face recognition and medical imaging, but its prediction fairness has raised significant concerns. Existing fairness evaluation specifications and metrics have inherent limitations, which overlook certain correlations between target features and sensitive attributes. In this work, we introduce a novel evaluation specification for image classification models based on dynamic perturbations to address this challenge. Specifically, we propose an Attribute Projection Perturbation Strategy (APPS) and a projection-based fairness metric system to quantify the upper and lower bounds of fairness perturbations. By employing projection factors, sensitive attributes that may influence task-specific properties are mapped onto a unified dimension, enabling a multi-perspective examination and evaluation of the impact of these attributes on the fairness of prediction outcomes. Compared to existing metrics, the proposed evaluation specification demonstrates superior objectivity and interpretability across 24 image classification models, including CNN and ViT architectures.
Haoze Jiang, Zunlei Feng, Jiacong Hu, Binde Hu, Mingli Song, Yuanyu Wan
ICME3
2025 Assessing Safety Risks and Quantization-aware Safety Patching for Quantized Large Language Models
abstract
Quantized large language models (LLMs) have gained increasing attention and significance for enabling deployment in resource-constrained environments. However, emerging studies on a few calibration dataset-free quantization methods suggest that quantization may compromise the safety capabilities of LLMs, underscoring the urgent need for systematic safety evaluations and effective mitigation strategies. In this paper, we present comprehensive safety evaluations across various mainstream quantization techniques and diverse calibration datasets, utilizing widely accepted safety benchmarks. To address the identified safety vulnerabilities, we propose a quantization-aware safety patching framework, Q-resafe, to efficiently restore the safety capabilities of quantized LLMs while minimizing any adverse impact on utility. Extensive experiment results demonstrate that Q-resafe successfully re-aligns the safety of quantized LLMs with their pre-quantization counterparts, even under challenging evaluation scenarios. Project page: https://github.com/Thecommonirin/Qresafe.
Kejia Chen 0007, Jiawen Zhang 0005, Jiacong Hu, Yu Wang 0176, Jian Lou 0001, Zunlei Feng, Mingli Song
ICML3
2025 DenseSAM: Semantic Enhance SAM for Efficient Dense Object Segmentation
abstract
Dense object segmentation is essential for various applications, particularly in pathology image and remote sensing image analysis. However, distinguishing numerous similar and densely packed objects in this task presents significant challenges. Several methods, including CNN- and ViT-based approaches, have been proposed to tackle these issues. Yet, models trained on limited datasets exhibit limited generalization ability. The Segment Anything Model (SAM) has recently achieved significant progress in zero-shot segmentation but relies heavily on precise positional guidance. However, providing numerous accurate location prompts in dense scenarios is time-consuming. To overcome this limitation, we conducted an in-depth exploration of the SAM mechanism and found that its strong generalization ability stems from the encoder’s edge detection capability, which is semantically independent, making location prompts essential for segmentation. This insight inspired the development of DenseSAM, which replaces location prompts with semantic guidance for automatic segmentation in dense scenarios. Specifically, it uses local details to weaken the edges of background objects, leverages global context to enhance intra-class feature similarity, while further increasing contrast with the background, and integrates a dual-head decoding process to enable lightweight automatic semantic segmentation. Extensive experiments on pathology images demonstrate that DenseSAM delivers remarkable performance with minimal training parameters, providing a cost-effective and efficient solution. Moreover, experiments on remote sensing images further validate its excellent scalability, making DenseSAM suitable for various dense object segmentation domains. The code is available at https://github.com/imAzhou/DenseSAM.
Linyun Zhou, Jiacong Hu, Shengxuming Zhang, Xiangtong Du, Mingli Song, Xiuming Zhang, Zunlei Feng
IJCAI2
2025 Tree of Preferences for Diversified Recommendation
abstract
Diversified recommendation has attracted increasing attention from both researchers and practitioners, which can effectively address the homogeneity of recommended items. Existing approaches predominantly aim to infer the diversity of user preferences from observed user feedback. Nonetheless, due to inherent data biases, the observed data may not fully reflect user interests, where underexplored preferences can be overwhelmed or remain unmanifested. Failing to capture these preferences can lead to suboptimal diversity in recommendations. To fill this gap, this work aims to study diversified recommendation from a data-bias perspective. Inspired by the outstanding performance of large language models (LLMs) in zero-shot inference leveraging world knowledge, we propose a novel approach that utilizes LLMs' expertise to uncover underexplored user preferences from observed behavior, ultimately providing diverse and relevant recommendations. To achieve this, we first introduce Tree of Preferences (ToP), an innovative structure constructed to model user preferences from coarse to fine. ToP enables LLMs to systematically reason over the user's rationale behind their behavior, thereby uncovering their underexplored preferences. To guide diversified recommendations using uncovered preferences, we adopt a data-centric approach, identifying candidate items that match user preferences and generating synthetic interactions that reflect underexplored preferences. These interactions are integrated to train a general recommender for diversification. Moreover, we scale up overall efficiency by dynamically selecting influential users during optimization. Extensive evaluations of both diversity and relevance show that our approach outperforms existing methods in most cases and achieves near-optimal performance in others, with reasonable inference latency.
Hanyang Yuan, Tongya Zheng, Jiarong Xu, Xintong Hu, Renhong Huang, Shunyu Liu 0001, Jiacong Hu, Jiawei Chen 0007, Mingli Song
NeurIPS8
2024 Human Identification at a Distance: Challenges, Methods and Results on the Competition HID 2024
abstract
Human identification at a distance (HID) faces challenges due to the difficulty of acquiring traditional biometric modalities like face and fingerprints. Gait recognition offers a viable solution since it can be captured at a distance. To advance the algorithm development and provide fair evaluations, the International Competition on Human Identification at a Distance (HID) has been held annually since 2020, with HID 2024 marking the fifth edition. Despite increased difficulty, participants demonstrated remarkable capabilities, surpassing previous accuracy levels. This paper, co-authored by competition organizers and top participants, provides a comprehensive summary of HID 2024, including an overview of the competition, and insights into the methods employed by the top teams. Specifically, inspired by the achievements of the 5 competitions of HID, we also provide the insights for the future directions on gait recognition.
Shiqi Yu 0001, Weiming Wu, Jiacong Hu, Zepeng Wang 0002, Runsheng Wang, Yunfei Ni, Yongzhen Huang, Liang Wang 0001, Md. Atiqur Rahman Ahad
IJCB3
2024 Improving Adversarial Robustness via Feature Pattern Consistency Constraint
Jiacong Hu, Jingwen Ye, Zunlei Feng, Jiazhen Yang, Shunyu Liu 0001, Xiaotian Yu, Lingxiang Jia, Mingli Song
IJCAI1
2024 Hundredfold Accelerating for Pathological Images Diagnosis and Prognosis through Self-reform Critical Region Focusing
Xiaotian Yu, Haoming Luo, Jiacong Hu, Xiuming Zhang, Yijun Bei, Mingli Song, Zunlei Feng
IJCAI3
2024 Transformer Doctor: Diagnosing and Treating Vision Transformers
abstract
Due to its powerful representational capabilities, Transformers have gradually become the mainstream model in the field of machine vision. However, the vast and complex parameters of Transformers impede researchers from gaining a deep understanding of their internal mechanisms, especially error mechanisms. Existing methods for interpreting Transformers mainly focus on understanding them from the perspectives of the importance of input tokens or internal modules, as well as the formation and meaning of features. In contrast, inspired by research on information integration mechanisms and conjunctive errors in the biological visual system, this paper conducts an in-depth exploration of the internal error mechanisms of Transformers. We first propose an information integration hypothesis for Transformers in the machine vision domain and provide substantial experimental evidence to support this hypothesis. This includes the dynamic integration of information among tokens and the static integration of information within tokens in Transformers, as well as the presence of conjunctive errors therein. Addressing these errors, we further propose heuristic dynamic integration constraint methods and rule-based static integration constraint methods to rectify errors and ultimately improve model performance. The entire methodology framework is termed as Transformer Doctor, designed for diagnosing and treating internal errors within transformers. Through a plethora of quantitative and qualitative experiments, it has been demonstrated that Transformer Doctor can effectively address internal errors in transformers, thereby enhancing model performance.
Jiacong Hu, Hao Chen 0041, Kejia Chen 0007, Yang Gao 0001, Jingwen Ye, Xingen Wang, Mingli Song, Zunlei Feng
NeurIPS1
2024 Vision Mamba Mender
abstract
Mamba, a state-space model with selective mechanisms and hardware-aware architecture, has demonstrated outstanding performance in long sequence modeling tasks, particularly garnering widespread exploration and application in the field of computer vision. While existing works have mixed opinions of its application in visual tasks, the exploration of its internal workings and the optimization of its performance remain urgent and worthy research questions given its status as a novel model. Existing optimizations of the Mamba model, especially when applied in the visual domain, have primarily relied on predefined methods such as improving scanning mechanisms or integrating other architectures, often requiring strong priors and extensive trial and error. In contrast to these approaches, this paper proposes the Vision Mamba Mender, a systematic approach for understanding the workings of Mamba, identifying flaws within, and subsequently optimizing model performance. Specifically, we present methods for predictive correlation analysis of Mamba's hidden states from both internal and external perspectives, along with corresponding definitions of correlation scores, aimed at understanding the workings of Mamba in visual recognition tasks and identifying flaws therein. Additionally, tailored repair methods are proposed for identified external and internal state flaws to eliminate them and optimize model performance. Extensive experiments validate the efficacy of the proposed methods on prevalent Mamba architectures, significantly enhancing Mamba's performance.
Jiacong Hu, Anda Cao, Zunlei Feng, Shengxuming Zhang, Lingxiang Jia, Mingli Song
NeurIPS1
2024 Model LEGO: Creating Models Like Disassembling and Assembling Building Blocks
abstract
With the rapid development of deep learning, the increasing complexity and scale of parameters make training a new model increasingly resource-intensive. In this paper, we start from the classic convolutional neural network (CNN) and explore a paradigm that does not require training to obtain new models. Similar to the birth of CNN inspired by receptive fields in the biological visual system, we draw inspiration from the information subsystem pathways in the biological visual system and propose Model Disassembling and Assembling (MDA). During model disassembling, we introduce the concept of relative contribution and propose a component locating technique to extract task-aware components from trained CNN classifiers. For model assembling, we present the alignment padding strategy and parameter scaling strategy to construct a new model tailored for a specific task, utilizing the disassembled task-aware components. The entire process is akin to playing with LEGO bricks, enabling arbitrary assembly of new models, and providing a novel perspective for model creation and reuse. Extensive experiments showcase that task-aware components disassembled from CNN classifiers or new models assembled using these components closely match or even surpass the performance of the baseline, demonstrating its promising results for model reuse. Furthermore, MDA exhibits diverse potential applications, with comprehensive experiments exploring model decision route analysis, model compression, knowledge distillation, and more.
Jiacong Hu, Jingwen Ye, Yang Gao 0001, Xingen Wang, Zunlei Feng, Mingli Song
NeurIPS1
2024 Association Pattern-aware Fusion for Biological Entity Relationship Prediction
abstract
Deep learning-based methods significantly advance the exploration of associations among triple-wise biological entities (e.g., drug-target protein-adverse reaction), thereby facilitating drug discovery and safeguarding human health. However, existing researches only focus on entity-centric information mapping and aggregation, neglecting the crucial role of potential association patterns among different entities. To address the above limitation, we propose a novel association pattern-aware fusion method for biological entity relationship prediction, which effectively integrates the related association pattern information into entity representation learning. Additionally, to enhance the missing information of the low-order message passing, we devise a bind-relation module that considers the strong bind of low-order entity associations. Extensive experiments conducted on three biological datasets quantitatively demonstrate that the proposed method achieves about 4%-23% hit@1 improvements compared with state-of-the-art baselines. Furthermore, the interpretability of association patterns is elucidated in detail, thus revealing the intrinsic biological mechanisms and promoting it to be deployed in real-world scenarios. Our data and code are available at https://github.com/hry98kki/PatternBERP.
Lingxiang Jia, Yuchen Ying, Zunlei Feng, Zipeng Zhong, Shaolun Yao, Jiacong Hu, Mingjiang Duan, Xingen Wang, Jie Song 0011, Mingli Song
NeurIPS6
2022 Model Doctor: A Simple Gradient Aggregation Strategy for Diagnosing and Treating CNN Classifiers
abstract
Recently, Convolutional Neural Network (CNN) has achieved excellent performance in the classification task. It is widely known that CNN is deemed as a 'blackbox', which is hard for understanding the prediction mechanism and debugging the wrong prediction. Some model debugging and explanation works are developed for solving the above drawbacks. However, those methods focus on explanation and diagnosing possible causes for model prediction, based on which the researchers handle the following optimization of models manually. In this paper, we propose the first completely automatic model diagnosing and treating tool, termed as Model Doctor. Based on two discoveries that 1) each category is only correlated with sparse and specific convolution kernels, and 2) adversarial samples are isolated while normal samples are successive in the feature space, a simple aggregate gradient constraint is devised for effectively diagnosing and optimizing CNN classifiers. The aggregate gradient strategy is a versatile module for mainstream CNN classifiers. Extensive experiments demonstrate that the proposed Model Doctor applies to all existing CNN classifiers, and improves the accuracy of 16 mainstream CNN classifiers by 1%~5%.
Zunlei Feng, Jiacong Hu, Sai Wu, Xiaotian Yu, Jie Song 0011, Mingli Song
AAAI2
2021 A Location Constrained Dual-Branch Network for Reliable Diagnosis of Jaw Tumors and Cysts
Jiacong Hu, Zunlei Feng, Yining Mao, Jie Lei 0002, Mingli Song
MICCAI (7)1
2004 Technique of the second satellite of CBERS-1 image multi-stage information extraction on land use and land cover
abstract
Since Oct. 21st 2003, on which the second satellite of CBERS-1 was successfully launched, professionals in the field of remote sensing pay more attention if information about the Earth's surface can be effectively extracted from its collection data. A method on extracting LULC information is discussed using data imaging on Feb. 2nd 2003 in This work. Remote sensing has inherent correlations with geo-knowledge. There are only few men interpreting remote sensing image with spectral knowledge. In this paper, a technique of multi-stage information extraction on LULC is present considering geo-knowledge by analyzing single band valve of sample end members, profile spectrum and constructing spectral band index. Different separable information of LULC is extracted at different stages, during which certain geo-knowledge is fused. In the end, this method is applied to an experiment area-Yun Nan DaLi, which obtains a better classified result compared with other supervised classified and unsupervised classified methods at the same area. Hence, it is useful to extract land use and land cover information from CBERS-1 image and extent application fields of CBERS-1.
Jiacong Hu, Yanmin Shuai, Suhong Liu, Jiangtao Li 0002, Min Li 0002, Qijiang Zhu
IGARSS1
2003 Study on the quality of hyperspectral vegetation data observed in the field
abstract
A measurement model on spectra quality is presented through a bigram composed of a spectra quality grade and a metadata integrality grade. Quantitative describing datasets and qualitative describing datasets of spectra quality are extracted with spectra enveloping line analysis, spectra line profile analysis, principles of relative parameters matching and spectra prior knowledge. The quality grade is converted from subordinative degree of eigenpoints and eigenvalues from quantitative datasets. The metadata integrality grade is obtained by visiting each node in a multicross tree by which metadata about vegetation spectra is organized. The two grades make up a bigram by which one can evaluate vegetation spectra quality.
Yanmin Shuai, Qijiang Zhu, Shihao Tang, Shuhong Liu, Jiacong Hu
IGARSS5