EDBT 2026 Demo / reviewers in the wild / expert
Hengjia Li
dblp:255/5148
· DBLP profile ↗
12ranked-venue papers
3as first author
12since 2021 · last 2026
0009-0007-6184-9437ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 9 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RoadBench: A Vision-Language Foundation Model and Benchmark for Road Damage UnderstandingabstractAccurate road damage detection is crucial for timely infrastructure maintenance and public safety, but existing vision-only datasets and models lack the rich contextual understanding that textual information can provide. To address this limitation, we introduce RoadBench, the first multimodal benchmark for comprehensive road damage understanding. This dataset pairs high-resolution images of road damages with detailed textual descriptions, providing a richer context for model training. We also present RoadCLIP, a novel vision-language model that builds upon CLIP by integrating domain-specific enhancements. It includes a disease-aware positional encoding that captures spatial patterns of road defects and a mechanism for injecting road-condition priors to refine the model’s understanding of road damages. We further employ a GPT-driven data generation pipeline to expand the image–text pairs in Road-Bench, greatly increasing data diversity without exhaustive manual annotation. Experiments demonstrate that Road-CLIP achieves state-of-the-art performance on road damage recognition tasks, significantly outperforming existing vision-only models by 19.2%. These results highlight the advantages of integrating visual and textual information for enhanced road condition analysis, setting new benchmarks for the field and paving the way for more effective infrastructure monitoring through multimodal learning. Xi Xiao 0003, Yunbei Zhang, Janet Wang, Yuxiang Wei 0004, Hengjia Li, Yanshu Li, Xiao Wang 0004, Swalpa Kumar Roy, Tianyang Wang 0004 |
WACV | 6 |
| 2025 | Local Conditional Controlling for Text-to-Image Diffusion ModelsabstractDiffusion models have exhibited impressive prowess in the text-to-image task. Recent methods add image-level structure controls, e.g., edge and depth maps, to manipulate the generation process together with text prompts to obtain desired images. This controlling process is globally operated on the entire image, which limits the flexibility of control regions. In this paper, we explore a novel and practical task setting: local control. It focuses on controlling specific local region according to user-defined image conditions, while the remaining regions are only conditioned by the original text prompt. However, it is non-trivial to achieve it. The naive manner of directly adding local conditions may lead to the local control dominance problem, which forces the model to focus on the controlled region and neglect object generation in other regions. To mitigate this problem, we propose Regional Discriminate Loss to update the noised latents, aiming at enhanced object generation in non-control regions. Furthermore, the proposed Focused Token Response suppresses weaker attention scores which lack the strongest response to enhance object distinction and reduce duplication. Lastly, we adopt Feature Mask Constraint to reduce quality degradation in images caused by information differences across the local control region. All proposed strategies are operated at the inference stage. Extensive experiments demonstrate that our method can synthesize high-quality images aligned with the text prompt under local control conditions. Yang Yang 0002, Zekai Luo, Hengjia Li, Zheng Yang 0008, Xiaofei He 0001, Wei Zhao 0019, Qinglin Lu, Wei Liu 0005, Boxi Wu 0001 |
AAAI | 5 |
| 2025 | Magicid: Hybrid Preference Optimization for Id-Consistent and Dynamic-Preserved Video Customization
Hengjia Li, Lifan Jiang, Hongwei Yi, Boxi Wu 0001, Deng Cai 0001 |
ICCV | 1 |
| 2025 | PersonalVideo: High ID-Fidelity Video Customization without Dynamic and Semantic DegradationabstractThe current text-to-video (T2V) generation has made significant progress in synthesizing realistic general videos, but it is still under-explored in identity-specific human video generation with customized ID images. The key challenge lies in maintaining high ID fidelity consistently while preserving the original motion dynamic and semantic following after the identity injection. Current video identity customization methods mainly rely on reconstructing given identity images on text-to-image models, which have a divergent distribution with the T2V model. This process introduces a tuning-inference gap, leading to dynamic and semantic degradation. To tackle this problem, we propose a novel framework, dubbed $\textbf{PersonalVideo}$, that applies a mixture of reward supervision on synthesized videos instead of the simple reconstruction objective on images. Specifically, we first incorporate identity consistency reward to effectively inject the reference's identity without the tuning-inference gap. Then we propose a novel semantic consistency reward to align the semantic distribution of the generated videos with the original T2V model, which preserves its dynamic and semantic following capability during the identity injection. With the non-reconstructive reward training, we further employ simulated prompt augmentation to reduce overfitting by supervising generated results in more semantic scenarios, gaining good robustness even with only a single reference image. Extensive experiments demonstrate our method's superiority in delivering high identity faithfulness while preserving the inherent video generation qualities of the original T2V model, outshining prior methods. Hengjia Li, Haonan Qiu, Shiwei Zhang 0001, Xiang Wang 0012, Yujie Wei 0001, Yingya Zhang, Boxi Wu 0001, Deng Cai 0001 |
ICCV | 1 |
| 2025 | DreamRelation: Relation-Centric Video Customization
Yujie Wei 0001, Shiwei Zhang 0001, Hangjie Yuan, Biao Gong, Longxiang Tang, Xiang Wang 0012, Haonan Qiu, Hengjia Li, Yingya Zhang, Hongming Shan |
ICCV | 8 |
| 2025 | SAMRefiner: Taming Segment Anything Model for Universal Mask RefinementabstractIn this paper, we explore a principal way to enhance the quality of widely pre-existing coarse masks, enabling them to serve as reliable training data for segmentation models to reduce the annotation cost. In contrast to prior refinement techniques that are tailored to specific models or tasks in a close-world manner, we propose SAMRefiner, a universal and efficient approach by adapting SAM to the mask refinement task. The core technique of our model is the noise-tolerant prompting scheme. Specifically, we introduce a multi-prompt excavation strategy to mine diverse input prompts for SAM (\ie, distance-guided points, context-aware elastic bounding boxes, and Gaussian-style masks) from initial coarse masks. These prompts can collaborate with each other to mitigate the effect of defects in coarse masks. In particular, considering the difficulty of SAM to handle the multi-object case in semantic segmentation, we introduce a split-then-merge (STM) pipeline. Additionally, we extend our method to SAMRefiner++ by introducing an additional IoU adaption step to further boost the performance of the generic SAMRefiner on the target dataset. This step is self-boosted and requires no additional annotation. The proposed framework is versatile and can flexibly cooperate with existing segmentation methods. We evaluate our mask framework on a wide range of benchmarks under different settings, demonstrating better accuracy and efficiency. SAMRefiner holds significant potential to expedite the evolution of refinement tools. Our code is available at https://github.com/linyq2117/SAMRefiner. Yuqi Lin, Hengjia Li, Wenqi Shao, Zheng Yang 0008, Jun Zhao 0009, Xiaofei He 0001, Ping Luo 0002, Kaipeng Zhang |
ICLR | 2 |
| 2025 | Singing Timbre Popularity Assessment Based on Multimodal Large Foundation Model
Ruibin Yuan, Ziqi Geng, Hengjia Li, Xingwei Qu, Songye Chen, Haoying Fu, Roger B. Dannenberg |
ACM Multimedia | 4 |
| 2025 | LoRA-Composer: Leveraging Low-Rank Adaptation for Multi-Concept Customization in Training-Free Diffusion ModelsabstractCustomization generation techniques have significantly advanced the synthesis of specific concepts across varied contexts. Multi-concept customization emerges as the challenging task within this domain. Existing approaches often rely on training a fusion matrix of multiple Low-Rank Adaptations (LoRAs) to merge various concepts into a single image. However, we identify this straightforward method faces two major challenges: 1) concept confusion, where the model struggles to preserve distinct individual characteristics, and 2) concept vanishing, where the model fails to generate the intended subjects. To address these issues, we introduce LoRA-Composer, a training-free framework designed for seamlessly integrating multiple LoRAs, thereby enhancing the harmony among different concepts within generated images. LoRA-Composer addresses concept vanishing through concept injection constraints, enhancing visibility via an expanded cross-attention mechanism. To combat concept confusion, concept isolation constraints are introduced, refining the self-attention computation. Furthermore, we propose two inference techniques to accelerate inference speed without performance degradation and enhance the accuracy of the generated region, respectively. Extensive experiments demonstrate that LoRA-Composer significantly outperforms standard baselines, especially in scenarios without image-based conditions such as canny edge or pose estimation. Yang Yang 0002, Chaotian Song, Hengjia Li, Qinglin Lu, Deng Cai 0001, Xiaofei He 0001, Boxi Wu 0001, Wei Liu 0005 |
IEEE Trans. Image Process. | 6 |
| 2024 | TagCLIP: A Local-to-Global Framework to Enhance Open-Vocabulary Multi-Label Classification of CLIP without TrainingabstractContrastive Language-Image Pre-training (CLIP) has demonstrated impressive capabilities in open-vocabulary classification. The class token in the image encoder is trained to capture the global features to distinguish different text descriptions supervised by contrastive loss, making it highly effective for single-label classification. However, it shows poor performance on multi-label datasets because the global feature tends to be dominated by the most prominent class and the contrastive nature of softmax operation aggravates it. In this study, we observe that the multi-label classification results heavily rely on discriminative local features but are overlooked by CLIP. As a result, we dissect the preservation of patch-wise spatial information in CLIP and proposed a local-to-global framework to obtain image tags. It comprises three steps: (1) patch-level classification to obtain coarse scores; (2) dual-masking attention refinement (DMAR) module to refine the coarse scores; (3) class-wise reidentification (CWR) module to remedy predictions from a global perspective. This framework is solely based on frozen CLIP and significantly enhances its multi-label classification performance on various benchmarks without dataset-specific training. Besides, to comprehensively assess the quality and practicality of generated tags, we extend their application to the downstream task, i.e., weakly supervised semantic segmentation (WSSS) with generated tags as image-level pseudo labels. Experiments demonstrate that this classify-then-segment paradigm dramatically outperforms other annotation-free segmentation methods and validates the effectiveness of generated tags. Our code is available at https://github.com/linyq2117/TagCLIP. Yuqi Lin, Minghao Chen 0001, Kaipeng Zhang, Hengjia Li, Zheng Yang 0008, Dongqin Lv, Binbin Lin 0001, Haifeng Liu 0001, Deng Cai 0001 |
AAAI | 4 |
| 2024 | Pseudo Label Refinery for Unsupervised Domain Adaptation on Cross-Dataset 3D Object DetectionabstractRecent self-training techniques have shown notable improvements in unsupervised domain adaptation for 3D object detection (3D UDA). These techniques typically select pseudo labels, i.e., 3D boxes, to supervise models for the target domain. However, this selection process inevitably introduces unreliable 3D boxes, in which 3D points cannot be definitively assigned as foreground or background. Previous techniques mitigate this by reweighting these boxes as pseudo labels, but these boxes can still poison the training process. To resolve this problem, in this paper, we propose a novel pseudo label refinery framework. Specifically, in the selection process, to improve the reliability of pseudo boxes, we propose a complementary augmentation strategy. This strategy involves either removing all points within an unre-liable box or replacing it with a high-confidence box. More-over, the point numbers of instances in high-beam datasets are considerably higher than those in low-beam datasets, also degrading the quality of pseudo labels during the training process. We alleviate this issue by generating additional proposals and aligning RoI features across different domains. Experimental results demonstrate that our method effectively enhances the quality of pseudo labels and consistently surpasses the state-of-the-art methods on six autonomous driving benchmarks. Code will be available at https://github.com/Zhanwei-Z/PERE. Zhanwei Zhang, Minghao Chen 0001, Hengjia Li, Binbin Lin 0001, Ping Li 0006, Wenxiao Wang 0001, Boxi Wu 0001, Deng Cai 0001 |
CVPR | 5 |
| 2024 | Few-shot Hybrid Domain Adaptation of Image GeneratorabstractCan a pre-trained generator be adapted to the hybrid of multiple target domains and generate images with integrated attributes of them? In this work, we introduce a new task -- Few-shot $\textit{Hybrid Domain Adaptation}$ (HDA). Given a source generator and several target domains, HDA aims to acquire an adapted generator that preserves the integrated attributes of all target domains, without overriding the source domain's characteristics. Compared with $\textit{Domain Adaptation}$ (DA), HDA offers greater flexibility and versatility to adapt generators to more composite and expansive domains. Simultaneously, HDA also presents more challenges than DA as we have access only to images from individual target domains and lack authentic images from the hybrid domain. To address this issue, we introduce a discriminator-free framework that directly encodes different domains' images into well-separable subspaces. To achieve HDA, we propose a novel directional subspace loss comprised of a distance loss and a direction loss. Concretely, the distance loss blends the attributes of all target domains by reducing the distances from generated images to all target subspaces. The direction loss preserves the characteristics from the source domain by guiding the adaptation along the perpendicular to subspaces. Experiments show that our method can obtain numerous domain-specific attributes in a single adapted generator, which surpasses the baseline methods in semantic similarity, image fidelity, and cross-domain consistency. Hengjia Li, Yang Liu 0212, Linxuan Xia, Yuqi Lin, Wenxiao Wang 0001, Tu Zheng, Zheng Yang 0008, Xiaohui Zhong, Xiaobo Ren, Xiaofei He 0001 |
ICLR | 1 |
| 2021 | Fixed-Lens camera setup and calibrated image registration for multifocus multiview 3D reconstructionabstractImage-based 3D reconstruction or 3D photogrammetry of small-scale objects including insects and biological specimens is challenging due to the use of a high magnification lens with inherently limited depth of field, and the object’s fine structures. Therefore, the traditional 3D reconstruction techniques cannot be applied without additional image preprocessing. One such preprocessing technique is multifocus stacking/fusion that combines a set of partially focused images captured at different distances from the same viewing angle to create a single in-focus image. We found that the image formation is not properly considered by the traditional multifocus image capture and stacking techniques. The resulting in-focus images contain artifacts that violate the perspective projection. A 3D reconstruction using such images often fails to produce accurate 3D models of the captured objects. This paper shows how this problem can be solved effectively by a new multifocus multiview 3D reconstruction procedure which includes a new Fixed-Lens multifocus image capture and a calibrated image registration technique using analytic homography transformation. The experimental results using the real and synthetic images demonstrate the effectiveness of the proposed solutions by showing that both the fixed-lens image capture and multifocus stacking with calibrated image alignment significantly reduce the errors in the camera poses and produce more complete 3D reconstructed models as compared with those by the conventional moving lens image capture and multifocus stacking. Shah Ariful Hoque Chowdhury, Chuong V. Nguyen, Hengjia Li, Richard I. Hartley |
Neural Comput. Appl. | 3 |