Enming Zhang

dblp:200/4357 · DBLP profile ↗
← Back
19ranked-venue papers
6as first author
16since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 4 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Learning Optimal Prompt Ensemble for Multi-source Visual Prompt Transfer
abstract
Prompt tuning has emerged as a lightweight strategy for adapting foundation models to downstream tasks, particularly for resource-constrained systems. As pre-trained prompts become valuable assets, combining multiple source prompts offers a promising approach to enhance generalization for new tasks by leveraging complementary knowledge. However, naive aggregation often overlooks different source prompts have different contribution potential to the target task. To address this, we propose HGPrompt, a dynamic framework that learns optimal ensemble weights. These weights are optimized by jointly maximizing an information-theoretic metric for transferability and minimizing gradient conflicts via a novel regularization strategy. Specifically, we propose a differentiable prompt transferability metric to captures the discriminability of prompt-induced features on the target task. Meanwhile, HGPrompt match the gradient variances with respect to different source prompts based on Hessian and Fisher Information, ensuring stable and coherent knowledge transfer while suppressing gradient conflicts among them. Extensive experiments on the large-scale VTAB benchmark demonstrate the state-of-the-art performance of HGPrompt, validating its effectiveness in learning an optimal ensemble for effective multi-source prompt transfer.
Enming Zhang, Liwen Cao, Yanru Wu, Yang Li 0104
AAAI1
2026 Evaluating the Perceptual Robustness of Vision-Language Models for Autonomous Driving in Corner Cases
Peizhe Gong, Enming Zhang, Ruixi Qiao, Xingyuan Dai, Xiaoyan Gong, Qinghai Miao
IV2
2026 DyStaFusion: Dynamic State-Space fusion network for multimodal tourist emotion dynamics prediction in social media
Enming Zhang
Appl. Intell.1
2026 DAMGO: Dynamic adaptive multi-graph optimization and fusion for multimodal recommendation
Enming Zhang, Xianying Huang
Expert Syst. Appl.1
2025 Towards Comprehensive Lecture Slides Understanding: Large-Scale Dataset and Effective Method
Enming Zhang, Yingying Zhu 0005, Xiang Bai
ICCV1
2025 MiniDrive: More Efficient Vision-Language Models with Multi-level 2D Features as Text Tokens for Autonomous Driving
Enming Zhang, Xingyuan Dai, Min Huang 0009, Qinghai Miao
PRCV (11)1
2025 A New Numerical Method for Fast Prediction of Wheel Tread Wear for Stacker Cranes
Minggong Yu, Enming Zhang, Johannes Fottner
SIMULTECH2
2025 SKDF: A Simple Knowledge Distillation Framework for Distilling Open-Vocabulary Knowledge to Open-World Object Detector
abstract
Open World Object Detection (OWOD) is a novel computer vision task with a considerable challenge, bridging the gap between classic object detection (OD) and real-world object detection. In addition to detecting and classifying seen/known objects, OWOD algorithms are expected to localize all potential unseen/unknown objects and incrementally learn them. The large pre-trained vision-language grounding models (VLM, e.g., GLIP) have rich knowledge about the open world, but are limited by text prompts and cannot localize indescribable objects. However, there are many detection scenarios in which pre-defined language descriptions are unavailable during inference. In this paper, we attempt to specialize the VLM model for OWOD tasks by distilling its open-world knowledge into a language-agnostic detector. Surprisingly, we observe that the simple knowledge distillation approach leads to unexpected performance for unknown object detection, even with a small amount of data. Unfortunately, knowledge distillation for unknown objects severely affects the learning of detectors with conventional structures, leading to catastrophic damage to the model's ability to learn about known objects. To alleviate these problems, we propose the down-weight training strategy for knowledge distillation from vision-language model to single visual modality one. Meanwhile, we propose the cascade decoupled decoders that decouple the learning of localization and recognition to reduce the impact of category interactions of known and unknown objects on the localization learning process. Ablation experiments demonstrate that both of them are effective in mitigating the impact of open-world knowledge distillation on the learning of known objects. Additionally, to alleviate the current lack of comprehensive benchmarks for evaluating the ability of the open-world detector to detect unknown objects in the open world, we refine the benchmark for evaluating the performance of unknown object detection by augmenting annotations for unknown objects which we name"IntensiveSet$\scriptstyle\spadesuit$♠". Comprehensive experiments performed on OWOD, MS-COCO, and our proposed benchmarks demonstrate the effectiveness of our methods.
Shuailei Ma, Ying Wei 0007, Enming Zhang, Peihao Chen
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 Optimization of Prompt Learning via Multi-Knowledge Representation for Vision-Language Models
abstract
Vision-language models (VLMs), such as CLIP, play a foundational role in various cross-modal applications. To fully leverage the potential of VLMs in adapting to downstream tasks, context optimization methods such as prompt tuning are essential. However, one key limitation is the lack of diversity in prompt templates, whether they are hand-crafted or learned through additional modules. This limitation restricts the capabilities of pretrained VLMs and can result in incorrect predictions in downstream tasks. To address this challenge, we propose context optimization with multi-knowledge representation (CoKnow), a framework that enhances prompt learning for VLMs with rich contextual knowledge. To facilitate CoKnow during inference, we train lightweight semantic knowledge mappers, which are capable of generating multi-knowledge representations for an input image without requiring additional priors. Experimentally, we conduct extensive experiments on 11 publicly available datasets, demonstrating that CoKnow outperforms a series of previous methods.
Enming Zhang, Bingke Zhu, Yingying Chen 0003, Qinghai Miao, Ming Tang 0001, Jinqiao Wang
IEEE Trans. Multim.1
2025 Transferability-Guided Cross-Domain Cross-Task Transfer Learning
abstract
We propose two novel transferability metrics fast optimal transport-based conditional entropy (F-OTCE) and joint correspondence OTCE (JC-OTCE) to evaluate how much the source model (task) can benefit the learning of the target task and to learn more generalizable representations for cross-domain cross-task transfer learning. Unlike the original OTCE metric that requires evaluating the empirical transferability on auxiliary tasks, our metrics are auxiliary-free such that they can be computed much more efficiently. Specifically, F-OTCE estimates transferability by first solving an optimal transport (OT) problem between source and target distributions and then uses the optimal coupling to compute the negative conditional entropy (NCE) between the source and target labels. It can also serve as an objective function to enhance downstream transfer learning tasks including model finetuning and domain generalization (DG). Meanwhile, JC-OTCE improves the transferability accuracy of F-OTCE by including label distances in the OT problem, though it incurs additional computation costs. Extensive experiments demonstrate that F-OTCE and JC-OTCE outperform state-of-the-art auxiliary-free metrics by 21.1% and 25.8%, respectively, in correlation coefficient with the ground-truth transfer accuracy. By eliminating the training cost of auxiliary tasks, the two metrics reduce the total computation time of the previous method from 43 min to 9.32 and 10.78 s, respectively, for a pair of tasks. When applied in the model finetuning and DG tasks, F-OTCE shows significant improvements in the transfer accuracy in few-shot classification experiments, with up to 4.41% and 2.34% accuracy gains, respectively.
Yang Tan 0004, Enming Zhang, Yang Li 0104, Shao-Lun Huang, Xiao-Ping Zhang 0002
IEEE Trans. Neural Networks Learn. Syst.2
2024 Never Lost in the Middle: Mastering Long-Context Question Answering with Position-Agnostic Decompositional Training
abstract
Junqing He, Kunhao Pan, Xiaoqun Dong, Zhuoyang Song, LiuYiBo LiuYiBo, Qianguosun Qianguosun, Yuxin Liang, Hao Wang, Enming Zhang, Jiaxing Zhang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Junqing He, Kunhao Pan, Xiaoqun Dong, Zhuoyang Song, LiuYiBo LiuYiBo, Qianguosun Qianguosun, Enming Zhang
ACL (1)9
2024 PSALM: Pixelwise SegmentAtion with Large Multi-modal Model
Zheng Zhang 0022, Yeyao Ma, Enming Zhang, Xiang Bai
ECCV (34)3
2024 DDOWOD: DiffusionDet for open-world object detection
Enming Zhang, Ying Wei 0007, Jiakun Xia, Xinghong Liu, Shuailei Ma
Pattern Recognit. Lett.2
2023 ICDAR 2023 Competition on Detecting Tampered Text in Images
Dongliang Luo, Rui Yang 0041, Xianjin Liu, Jishen Zeng, Enming Zhang, Ziming Huang, Xiang Bai
ICDAR (2)7
2023 A knowledge service framework for fault diagnosis of low-earth orbit satellite constellation
abstract
The rapid expansion of low-earth orbit satellite constellations poses a significant challenge for the operation and maintenance of thousands of satellites. In-orbit fault diagnosis helps minimize the cost of repairs, and ensure the overall reliability of the satellite constellation by identifying faults in their early stages. Compared to traditional fault detection methods, fault diagnosis requires knowledge to enhance very limited data and to infer cascading failures. In this paper, we propose a novel knowledge service framework that combines data-driven and knowledge-driven models to improve the efficiency and effectiveness of fault diagnosis. We implement a fault diagnosis service based on a constructed knowledge graph, which affords an assistant decision support function. The in-orbit deployment and simulation experiment verify the feasibility of the proposed framework, which shows great potential for fault diagnosis of low-earth orbit satellite constellation.
Fei Teng 0001, Enming Zhang, Qibo Sun
ICWS3
2023 Parallel matters: Efficient polyp segmentation with parallel structured feature augmentation modules
abstract
Abstract The large variations of polyp sizes and shapes and the close resemblances of polyps to their surroundings call for features with long‐range information in rich scales and strong discrimination. This article proposes two parallel structured modules for building those features. One is the Transformer Inception module (TI) which applies Transformers with different reception fields in parallel to input features and thus enriches them with more long‐range information in more scales. The other is the Local‐Detail Augmentation module (LDA) which applies the spatial and channel attentions in parallel to each block and thus locally augments the features from two complementary dimensions for more object details. Integrating TI and LDA, a new Transformer encoder based framework, Parallel‐Enhanced Network (PENet), is proposed, where LDA is specifically adopted twice in a coarse‐to‐fine way for accurate prediction. PENet is efficient in segmenting polyps with different sizes and shapes without the interference from the background tissues. Experimental comparisons with state‐of‐the‐arts methods show its merits.
Xianyong Fang, Kaibing Wang, Yuqing Shi, Linbo Wang 0001, Enming Zhang, Zhengyi Liu
IET Image Process.6
2020 Extracting Features from Online Forums to Meet Social Needs of Breast Cancer Patients
abstract
Breast cancer patients go through many ordeals when they undergo treatments. Many of these issues are personal, social, or professional. As many of them are not directly medical in nature, these issues are not discussed with their healthcare providers and hence, not included in their treatment plan. However, these issues are vital for the patients' complete recovery. We present a novel approach that acts as the first step in including such personal and social issues resulting from breast cancer treatment into a patient's treatment plan. There are numerous online forums where patients share their experiences and post questions about their treatments and subsequent side effects. We collected data from one such forum called "Online Breast Cancer Forum". On this forum, users (patients) have created threads across many related topics and shared their experiences and questions. We use these message threads to identify critical issues faced by the patient and how they are related to their treatment. We convert the forum data into a bipartite network and turn the network nodes into a high-dimensional feature space. In this feature space, we perform community detection to unearth latent connections between patients and topics. We claim that these latent connections, along with the known ones, will help to create a new knowledge base that will eventually help physicians to estimate non-medical issues for a prescribed treatment. This new knowledge will help the physicians plan a more adaptive and personalized treatment and be better prepared by anticipating potential problems beforehand. We evaluated our method on two baseline methods and show that our method outperforms the baseline methods by 25% on a manually labeled reference dataset.
Maitreyi Mokashi, Enming Zhang, Josette F. Jones, Sunandan Chakraborty
COMPASS2
2018 Finding the Best EHR Training Dataset: A Cross-over Trial
Josette F. Jones, Rukhaiya Fatima, Saptarshi Purkayastha, Anand Kulanthaivel, Enming Zhang
AMIA5
2016 Towards the creation of a novel career-based Health Informatics (HI) curriculum assessment instrument: Mapping HI job competencies to HI curriculum competencies
Anand Kulanthaivel, Enming Zhang, Shilpa Katta, Josette F. Jones
AMIA2