Yi-Jie Huang

dblp:184/6951 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
7since 2021 · last 2025
0000-0002-8320-9277ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 CLIP Brings Better Features to Visual Aesthetics Learners
abstract
Image Aesthetics Assessment (IAA) is a challenging task due to its subjective nature and expensive manual annotations. Recent large-scale vision-language models, such as Contrastive Language-Image Pre-training (CLIP), have shown their promising representation capability for various downstream tasks. However, the application of CLIP to resource-constrained and low-data IAA tasks remains limited. While few attempts to leverage CLIP in IAA have mainly focused on carefully designed prompts, we extend beyond this by allowing models from different domains and with different model sizes to acquire knowledge from CLIP. To achieve this, we propose a unified and flexible two-phase CLIP-based Semi-supervised Knowledge Distillation (CSKD) paradigm, aiming to learn a lightweight IAA model while leveraging CLIP’s strong generalization capability. Specifically, CSKD employs a feature alignment strategy to facilitate the distillation of heterogeneous CLIP teacher and IAA student models, effectively transferring valuable features from pre-trained visual representations to two lightweight IAA models, respectively. To efficiently adapt to downstream IAA tasks in a low-data regime, the two strong visual aesthetics learners then conduct distillation with unlabeled examples for refining and transferring the task-specific knowledge collaboratively. Extensive experiments demonstrate that the proposed CSKD achieves state-of-the-art performance on multiple widely used IAA benchmarks. Furthermore, analysis of attention distance and entropy before and after feature alignment shows the effective transfer of CLIP’s feature representation to IAA models, which not only provides valuable guidance for the model initialization of IAA but also enhances the aesthetic feature representation of IAA models. Code will be made publicly available.
Liwu Xu, Jinjin Xu, Yuzhe Yang 0001, Xilu Wang 0001, Yi-Jie Huang
ICME5
2025 Reproducibility Companion Paper: u-LLaVA: Unifying Multi-Modal Tasks via Large Language Model
Jinjin Xu, Xilu Wang 0001, Liwu Xu, Yuzhe Yang 0001, Xiang Li 0179, Fanyi Wang, Yanchun Xie, Yi-Jie Huang, Yunfan Hu
ICMR8
2025 Open-Set Image Tagging with Multi-Grained Text Supervision
abstract
This paper introduces the Recognize Anything Plus Model (RAM++), an open-set image tagging model effectively leveraging multi-grained text supervision. Previous approaches (e.g., CLIP) primarily utilize global text supervision paired with images, leading to sub-optimal performance in recognizing multiple individual semantic tags. In contrast, RAM++ seamlessly integrates individual tag supervision with global text supervision, all within a unified alignment framework. This integration not only ensures efficient recognition of predefined tag categories, but also enhances generalization capabilities for diverse open-set categories. Furthermore, RAM++ employs large language models (LLMs) to convert semantically constrained tag supervision into more expansive tag description supervision, thereby enriching the scope of open-set visual description concepts. Comprehensive evaluations on various image recognition benchmarks demonstrate RAM++ exceeds existing state-of-the-art (SOTA) open-set image tagging models on most aspects. Specifically, for predefined commonly used tag categories, RAM++ showcases 10.2 mAP and 15.4 mAP enhancements over CLIP on OpenImages and ImageNet. For open-set categories beyond predefined, RAM++ records improvements of 5.0 mAP and 6.4 mAP over CLIP and RAM respectively on OpenImages. For diverse human-object interaction phrases, RAM++ achieves 7.8 mAP and 4.7 mAP improvements on the HICO benchmark.
Yi-Jie Huang, Youcai Zhang, Rui Feng 0001, Yuejie Zhang, Yanchun Xie, Lei Zhang 0001
ACM Multimedia2
2025 Toward Realistic Hierarchical Object Detection: Problem, Benchmark, and Solution
abstract
With the continuous advancement of deep learning, object detection has made remarkable progress in accurately identifying a wide range of object categories, even within increasingly complex scenes. However, as the number of categories grows, visual concepts naturally organize into a label hierarchy. We contend that existing hierarchical classification and detection methods predominantly prioritize fine-grained prediction, potentially leading to inconsistencies with realistic human perception. From this perspective, we investigate the Hierarchical Object Detection (HOD) problem to better align with real-world perception. To address the lack of benchmarks in the field, we build a large-scale HOD benchmark termed RHOD with open-source datasets, comprising 740 categories. To better align the hierarchical object detectors towards realistic perception, we propose a new evaluation metric named Hierarchical Average Precision (HAP). Furthermore, we present a novel hierarchical object detection method that includes two components, Tree Soft Labeling (TSL) and Hierarchical Extension and Suppression (HES). Our method mitigates the issue of overconfidence in fine-grained predictions, which has been prevalent in previous approaches. We evaluate a range of existing methods on the RHOD benchmark, including plain, hierarchical, and open-vocabulary models. Additionally, we perform comprehensive experiments to assess the performance of our proposed method. The experimental results show that our method achieves state-of-the-art performance on the RHOD benchmark.
Juexiao Feng, Yuhong Yang 0008, Mengyao Lyu, Tianxiang Hao 0001, Yi-Jie Huang, Yanchun Xie, Jungong Han, Liuyu Xiang, Guiguang Ding
IEEE Trans. Circuits Syst. Video Technol.5
2024 u-LLaVA: Unifying Multi-Modal Tasks via Large Language Model
abstract
Recent advancements in multi-modal large language models (MLLMs) have led to substantial improvements in visual understanding, primarily driven by sophisticated modality alignment strategies. However, predominant approaches prioritize global or regional comprehension, with less focus on fine-grained, pixel-level tasks. To address this gap, we introduce u-LLaVA, an innovative unifying multi-task framework that integrates pixel, regional, and global features to refine the perceptual faculties of MLLMs. We commence by leveraging an efficient modality alignment approach, harnessing both image and video datasets to bolster the model’s foundational understanding across diverse visual contexts. Subsequently, a joint instruction tuning method with task-specific projectors and decoders for end-to-end downstream training is presented. Furthermore, this work contributes a novel mask-based multi-task dataset comprising 277K samples, crafted to challenge and assess the fine-grained perception capabilities of MLLMs. The overall framework is simple, effective, and achieves state-of-the-art performance across multiple benchmarks. We make model, data, and code publicly accessible at https://github.com/OPPOMKLab/u-LLaVA.
Jinjin Xu, Liwu Xu, Yuzhe Yang 0001, Xiang Li 0179, Fanyi Wang, Yanchun Xie, Yi-Jie Huang
ECAI7
2022 3D Graph-Connectivity Constrained Network for Hepatic Vessel Segmentation
abstract
Segmentation of hepatic vessels from 3D CT images is necessary for accurate diagnosis and preoperative planning for liver cancer. However, due to the low contrast and high noises of CT images, automatic hepatic vessel segmentation is a challenging task. Hepatic vessels are connected branches containing thick and thin blood vessels, showing an important structural characteristic or a prior: the connectivity of blood vessels. However, this is rarely applied in existing methods. In this paper, we segment hepatic vessels from 3D CT images by utilizing the connectivity prior. To this end, a graph neural network (GNN) used to describe the connectivity prior of hepatic vessels is integrated into a general convolutional neural network (CNN). Specifically, a graph attention network (GAT) is first used to model the graphical connectivity information of hepatic vessels, which can be trained with the vascular connectivity graph constructed directly from the ground truths. Second, the GAT is integrated with a lightweight 3D U-Net by an efficient mechanism called the plug-in mode, in which the GAT is incorporated into the U-Net as a multi-task branch and is only used to supervise the training procedure of the U-Net with the connectivity prior. The GAT will not be used in the inference stage, and thus will not increase the hardware and time costs of the inference stage compared with the U-Net. Therefore, hepatic vessel segmentation can be well improved in an efficient mode. Extensive experiments on two public datasets show that the proposed method is superior to related works in accuracy and connectivity of hepatic vessel segmentation.
Ruikun Li 0004, Yi-Jie Huang, Huai Chen, Yizhou Yu, Dahong Qian, Lisheng Wang
IEEE J. Biomed. Health Informatics2
2021 3-D RoI-Aware U-Net for Accurate and Efficient Colorectal Tumor Segmentation
abstract
Segmentation of colorectal cancerous regions from 3-D magnetic resonance (MR) images is a crucial procedure for radiotherapy. Automatic delineation from 3-D whole volumes is in urgent demand yet very challenging. Drawbacks of existing deep-learning-based methods for this task are two-fold: 1) extensive graphics processing unit (GPU) memory footprint of 3-D tensor limits the trainable volume size, shrinks effective receptive field, and therefore, degrades speed and segmentation performance and 2) in-region segmentation methods supported by region-of-interest (RoI) detection are either blind to global contexts, detail richness compromising, or too expensive for 3-D tasks. To tackle these drawbacks, we propose a novel encoder-decoder-based framework for 3-D whole volume segmentation, referred to as 3-D RoI-aware U-Net (3-D RU-Net). 3-D RU-Net fully utilizes the global contexts covering large effective receptive fields. Specifically, the proposed model consists of a global image encoder for global understanding-based RoI localization, and a local region decoder that operates on pyramid-shaped in-region global features, which is GPU memory efficient and thereby enables training and prediction with large 3-D whole volumes. To facilitate the global-to-local learning procedure and enhance contour detail richness, we designed a dice-based multitask hybrid loss function. The efficiency of the proposed framework enables an extensive model ensemble for further performance gain at acceptable extra computational costs. Over a dataset of 64 T2-weighted MR images, the experimental results of four-fold cross-validation show that our method achieved 75.5% dice similarity coefficient (DSC) in 0.61 s per volume on a GPU, which significantly outperforms competing methods in terms of accuracy and efficiency. The code is publicly available.
Yi-Jie Huang, Qi Dou 0001, Zi-Xian Wang, Li-Zhi Liu, Chao-Feng Li, Lisheng Wang, Hao Chen 0011, Rui-Hua Xu
IEEE Trans. Cybern.1
2020 Rectifying Supporting Regions With Mixed and Active Supervision for Rib Fracture Recognition
abstract
Automatic rib fracture recognition from chest X-ray images is clinically important yet challenging due to weak saliency of fractures. Weakly Supervised Learning (WSL) models recognize fractures by learning from large-scale image-level labels. In WSL, Class Activation Maps (CAMs) are considered to provide spatial interpretations on classification decisions. However, the high-responding regions, namely Supporting Regions of CAMs may erroneously lock to regions irrelevant to fractures, which thereby raises concerns on the reliability of WSL models for clinical applications. Currently available Mixed Supervised Learning (MSL) models utilize object-level labels to assist fitting WSL-derived CAMs. However, as a prerequisite of MSL, the large quantity of precisely delineated labels is rarely available for rib fracture tasks. To address these problems, this paper proposes a novel MSL framework. Firstly, by embedding the adversarial classification learning into WSL frameworks, the proposed Biased Correlation Decoupling and Instance Separation Enhancing strategies guide CAMs to true fractures indirectly. The CAM guidance is insensitive to shape and size variations of object descriptions, thereby enables robust learning from bounding boxes. Secondly, to further minimize annotation cost in MSL, a CAM-based Active Learning strategy is proposed to recognize and annotate samples whose Supporting Regions cannot be confidently localized. Consequently, the quantity demand of object-level labels can be reduced without compromising the performance. Over a chest X-ray rib-fracture dataset of 10966 images, the experimental results show that our method produces rational Supporting Regions to interpret its classification decisions and outperforms competing methods at an expense of annotating 20% of the positive samples with bounding boxes.
Yi-Jie Huang, Xiuying Wang 0001, Qu Fang, Renzhen Wang, Huai Chen, Hao Chen 0011, Deyu Meng, Lisheng Wang
IEEE Trans. Medical Imaging1
2017 Genetic algorithm-based interval type-2 fuzzy model identification for people with type-1 diabetes
abstract
In this paper, the glucose regulation system is identified by interval type-2 fuzzy neural network (IT2FNN) based on genetic algorithm (GA) used to adapt the model parameters. The IT2FNN is constructed to identify the glucose regulation system of the people with diabetes type-1. The centers and widths of the memberships and output weights of the IT2FNN can be tuned by optimizing GA. The simulation result shows that the glucose-insulin behavior can be well identified by the advocated identification scheme.
Tsung-Chih Lin, Yi-Jie Huang, Josephine I-Ju Lin, Valentina Emilia Balas, Seshadhri Srinivasan
FUZZ-IEEE2
2016 Frequency domain digital watermark recognition using image code sequences with a back-propagation neural network
Chih-Ta Yen, Yi-Jie Huang
Multim. Tools Appl.2