EDBT 2026 Demo / reviewers in the wild / expert
Junjia Huang
dblp:302/4194
· DBLP profile ↗
11ranked-venue papers
9as first author
11since 2021 · last 2026
0009-0006-2914-6232ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 7 first-author · 8 since 2021Artificial intelligence and machine learning · 6 · 5 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DreamFuse: Toward Realistic and Seamless Image Fusion Across Diverse ScenariosabstractImage fusion seeks to seamlessly integrate foreground objects with background scenes, producing realistic and harmonious fused images. While existing methods often insert objects directly, adaptive and interactive fusion-requiring contextual adaptation and foreground-background interplay-remains a challenging yet critical task. To address this, we first propose a pipeline for generating high-quality fusion data. By combining iterative in-context learning with existing tools, we curate a diverse cross-scene dataset supporting three core tasks: object integration, replacement, and attribute-referenced editing. Leveraging this, we introduce DreamFuse, a unified diffusion-based approach that jointly optimizes these capabilities. DreamFuse exploits the Diffusion Transformer (DiT) architecture, using its attention mechanism to extract and align foreground-background features for coherent fusion. For flexible control, we incorporate a Positional Affine mechanism, enabling precise spatial and scale adjustments while supporting diverse text-driven fusion. Furthermore, we employ Localized Direct Preference Optimization (L-DPO), refining the model via human feedback to enhance harmony and consistency. Extensive experimental results demonstrate DreamFuse's superiority over state-of-the-art approaches across multiple metrics. Junjia Huang, Pengxiang Yan, Jiyang Liu, Jie Wu 0030, Liang Lin 0004, Guanbin Li |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | Universal Scale Transformer for Histology Image SegmentationabstractHistology image segmentation is a critical prerequisite to pathological diagnosis. Accurate segmentation of these images can significantly aid physicians by facilitating quicker and more precise diagnostic decisions. A notable challenge in this area arises from the fact that different objects within histology images require segmentation at varying magnifications. However, most existing models are typically limited to performing segmentation at one pre-determined magnification. In this paper, we propose a novel universal scale transformer model (UniScaleFormer), which employs a scale-aware approach and integrates textual input to uniformly segment objects in histology images across various magnifications. Our method adopts an end-to-end architecture and utilizes a candidate mask query mechanism, specifically designed to identify and distinguish objects at different image scales. Moreover, we develop a Scale-Aware Module that enhances our network's ability to recognize the magnification level of input histology images, by using a scale query with extracted visual features. Then the scale query is integrated with mask queries, facilitating the incorporation of scale information. Experimental results demonstrate that the proposed method effectively achieves competitive results on various segmentation benchmarks at different magnifications. The code will be released at https://github.com/lhaof/UniScale. Junjia Huang, Haofeng Li, Yuanhuan Xiong, Guanbin Li |
IEEE Trans. Medical Imaging | 1 |
| 2025 | DreamLayer: Simultaneous Multi-Layer Generation via Diffusion Model
Junjia Huang, Pengxiang Yan, Jinhang Cai, Jiyang Liu, Guanbin Li |
ICCV | 1 |
| 2025 | DreamFuse: Adaptive Image Fusion with Diffusion TransformerabstractImage fusion seeks to seamlessly integrate foreground objects with background scenes, producing realistic and harmonious fused images. Unlike existing methods that directly insert objects into the background, adaptive and interactive fusion remains a challenging yet appealing task. It requires the foreground to adjust or interact with the background context, enabling more coherent integration. To address this, we propose an iterative human-in-the-loop data generation pipeline, which leverages limited initial data with diverse textual prompts to generate fusion datasets across various scenarios and interactions, including placement, holding, wearing, and style transfer. Building on this, we introduce DreamFuse, a novel approach based on the Diffusion Transformer (DiT) model, to generate consistent and harmonious fused images with both foreground and background information. DreamFuse employs a Positional Affine mechanism to inject the size and position of the foreground into the background, enabling effective foreground-background interaction through shared attention. Furthermore, we apply Localized Direct Preference Optimization guided by human feedback to refine DreamFuse, enhancing background consistency and foreground harmony. DreamFuse achieves harmonious fusion while generalizing to text-driven attribute editing of the fused results. Experimental results demonstrate that our method outperforms state-of-the-art approaches across multiple metrics. Junjia Huang, Pengxiang Yan, Jiyang Liu, Jie Wu 0030, Liang Lin 0004, Guanbin Li |
ICCV | 1 |
| 2024 | UniCell: Universal Cell Nucleus Classification via Prompt LearningabstractThe recognition of multi-class cell nuclei can significantly facilitate the process of histopathological diagnosis. Numerous pathological datasets are currently available, but their annotations are inconsistent. Most existing methods require individual training on each dataset to deduce the relevant labels and lack the use of common knowledge across datasets, consequently restricting the quality of recognition. In this paper, we propose a universal cell nucleus classification framework (UniCell), which employs a novel prompt learning mechanism to uniformly predict the corresponding categories of pathological images from different dataset domains. In particular, our framework adopts an end-to-end architecture for nuclei detection and classification, and utilizes flexible prediction heads for adapting various datasets. Moreover, we develop a Dynamic Prompt Module (DPM) that exploits the properties of multiple datasets to enhance features. The DPM first integrates the embeddings of datasets and semantic categories, and then employs the integrated prompts to refine image representations, efficiently harvesting the shared knowledge among the related cell types and data sources. Experimental results demonstrate that the proposed method effectively achieves the state-of-the-art results on four nucleus detection and classification benchmarks. Code and models are available at https://github.com/lhaof/UniCell Junjia Huang, Haofeng Li, Guanbin Li |
AAAI | 1 |
| 2023 | Affine-Consistent Transformer for Multi-Class Cell Nuclei DetectionabstractMulti-class cell nuclei detection is a fundamental prerequisite in the diagnosis of histopathology. It is critical to efficiently locate and identify cells with diverse morphology and distributions in digital pathological images. Most existing methods take complex intermediate representations as learning targets and rely on inflexible post-refinements while paying less attention to various cell density and fields of view. In this paper, we propose a novel Affine-Consistent Transformer (AC-Former), which directly yields a sequence of nucleus positions and is trained collaboratively through two sub-networks, a global and a local network. The local branch learns to infer distorted input images of smaller scales while the global network outputs the large-scale predictions as extra supervision signals. We further introduce an Adaptive Affine Transformer (AAT) module, which can automatically learn the key spatial transformations to warp original images for local network training. The AAT module works by learning to capture the transformed image regions that are more valuable for training the model. Experimental results demonstrate that the proposed method significantly outperforms existing state-of-the-art algorithms on various benchmarks. Junjia Huang, Haofeng Li, Guanbin Li |
ICCV | 1 |
| 2023 | Prompt-Based Grouping Transformer for Nucleus Detection and Classification
Junjia Huang, Haofeng Li, Weijun Sun, Guanbin Li |
MICCAI (8) | 1 |
| 2023 | Estimating Human Weight From a Single ImageabstractBody weight, as one of the biometric traits, has been studied in both the forensic and medical domains. However, estimating weight directly from 2-D images is particularly challenging since visual inspection is rather sensitive to the distance between the subject and camera, even for frontal view images. In this case, the widely used body mass index (BMI), which is associated with body height and weight, can be employed as a measure of weight to indicate health conditions. Previous works on the estimation of BMI have predominantly focused on using multiple 2-D images, 3-D images, or facial images; however, these cues are not always available. To address this issue, we explore the feasibility of obtaining BMI from a single 2-D body image with the dual-branch regression framework proposed in this work. More specifically, the framework comprises an anthropometric feature computation branch and a deep learning-based feature extraction branch. One aggregation layer maps all the features to an estimated BMI value. In addition, a new public 2-D image-to-BMI dataset, which contains 4189 images (1477 males and 2712 females) from approximately 3000 subjects with attributes including gender, age, height, and weight, was collected and released to facilitate the study. Extensive experiments confirm that the proposed framework combining anthropometric features and deep features outperforms the single-type feature approaches to BMI estimation in most cases. Zhi Jin 0002, Junjia Huang, Wenjin Wang 0002, Aolin Xiong, Xiaojun Tan |
IEEE Trans. Multim. | 2 |
| 2022 | Attentive Symmetric Autoencoder for Brain MRI Segmentation
Junjia Huang, Haofeng Li, Guanbin Li |
MICCAI (5) | 1 |
| 2022 | Attention guided deep features for accurate body mass index estimation
Zhi Jin 0002, Junjia Huang, Aolin Xiong, Yuxian Pang, Wenjin Wang 0002, Beichen Ding |
Pattern Recognit. Lett. | 2 |
| 2021 | Seeing Health with Eyes: Feature Combination for Image-Based Human BMI EstimationabstractBody Mass Index (BMI) is an important measurement of human obesity and health, which can provide useful information for plenty of practical purposes, such as monitoring, re-identification, and health care. Recently, some data-driven advances have been proposed to estimate BMI by 2D or 3D features from face images, frontal-body images and RGB-D images. However, due to the privacy issue or limitations of 3D cameras, the required data is hard to be obtained. More importantly, each of the previous works has only studied for a single type of features, hence it is worth investigating whether combinations of different features are more effective. To address this issue, we analyze the correlation of various features extracted from 2D body images with the estimated BMI, and then propose an accurate BMI estimation method with the optimal feature combination. Extensive experiments demonstrate that the proposed method outperforms these image-based BMI estimation methods which only utilize a single type feature in most cases. Code has been made available at : https://github.com/FVL2020/Features_for_BMI_estimation. Junjia Huang, Chenming Shang, Aolin Xiong, Yuxian Pang, Zhi Jin 0002 |
ICME | 1 |