EDBT 2026 Demo / reviewers in the wild / expert
Longfei Han
dblp:150/2367
· DBLP profile ↗
26ranked-venue papers
5as first author
21since 2021 · last 2026
0000-0003-2135-6228ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 2 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CLIP-based knowledge projector for image-text matching
Dingwen Zhang, Longfei Han, Huaxiang Zhang 0001, Li Liu 0031, Junwei Han 0001 |
Inf. Process. Manag. | 3 |
| 2026 | DecoupleNet: Domain-specific task decoupling network for low-light image enhancement
Peiliang Huang, Xianmin Chen, Xiaoxu Feng, Qiangqiang Wang, Dingwen Zhang, Longfei Han, Junwei Han 0001 |
Pattern Recognit. | 6 |
| 2026 | Retinex-RAWMamba: Bridging Demosaicing and Denoising for Low-Light RAW Image EnhancementabstractLow-light image enhancement, particularly in cross-domain tasks such as mapping from the raw domain to the sRGB domain, remains a significant challenge. Many deep learning-based methods have been developed to address this issue and have shown promising results in recent years. However, single-stage methods, which attempt to unify the complex mapping across both domains, leading to limited denoising performance. In contrast, existing two-stage approaches typically overlook the characteristic of demosaicing within the Image Signal Processing (ISP) pipeline, leading to color distortions under varying lighting conditions, especially in low-light scenarios. To address these issues, we propose a novel Mamba-based method customized for low light RAW images, called RAWMamba, to effectively handle raw images with different CFAs. Furthermore, we introduce a Retinex Decomposition Module (RDM) grounded in Retinex prior, which decouples illumination from reflectance to facilitate more effective denoising and automatic non-linear exposure correction, reducing the effect of manual linear illumination enhancement. By bridging demosaicing and denoising, better enhancement for low light RAW images is achieved. Experimental evaluations conducted on public datasets SID and MCR demonstrate that our proposed RAWMamba achieves state-of-the-art performance on cross-domain mapping. The code is available at https://github.com/Cynicarlos/RetinexRawMamba. Xianmin Chen, Longfei Han, Peiliang Huang, Xiaoxu Feng, Dingwen Zhang, Junwei Han 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | 3D Extended Object Tracking Based on Extruded B-Spline Side View ProfilesabstractObject tracking is an essential task for autonomous systems. With the advancement of 3D sensors, these systems can better perceive their surroundings using effective 3D Extended Object Tracking (EOT) methods. Based on the observation that common road users are symmetrical on the right and left sides in the traveling direction, we focus on the side view profile of the object. In order to leverage of the development in 2D EOT and balance the number of parameters of a shape model in the tracking algorithms, we propose a method for 3D extended object tracking (EOT) by describing the side view profile of the object with B-spline curves and forming an extrusion to obtain a 3D extent. The use of B-spline curves exploits their flexible representation power by allowing the control points to move freely. The algorithm is developed into an Extended Kalman Filter (EKF). For a through evaluation of this method, we use simulated traffic scenario of different vehicle models and real-world open dataset containing both radar and lidar data. Longfei Han, Klaus Kefferpütz, Jürgen Beyerer |
FUSION | 1 |
| 2025 | EEG2Gaussian: Decoding and Visualizing Visual-Evoked EEG for VR Scenes Using 3D Gaussian SplattingabstractDecoding and visualizing brain activity evoked by visual stimuli is critical for both understanding neural mechanisms and advancing brain-computer interfaces (BCIs). However, non-invasive signals such as Electroencephalogram (EEG) present significant challenges due to their inherently low signal-to-noise ratios. Although recent deep learning methods have resolved this task, most approaches are confined to 2D visualizations that fail to capture the complexities of real-world 3D perception. In this research, we investigate the relationship between EEG signals and 3D visual stimuli presented in virtual reality (VR) scenes, aiming to extract taskrelevant semantics from the EEG responses elicited by these stimuli. We introduce EEG2Gaussian, a novel framework for decoding and visualizing visual-evoked EEG signals by reconstructing immersive VR scenes using 3D Gaussian Splatting. The framework consists of three stages. The preprocessing stage removes noise and artifacts from raw EEG signals to provide cleaner input for subsequent processing. In the encoding stage, we propose a Neural Temporal-Frequency Encoder (NTF-Encoder) to extract temporal and frequency features using fused channel and band attention mechanisms, and disentangles them into high-level and low-level semantic representations. In the decoding stage, a 3D EEG Decoder takes these multi-level features through separate pathways as conditional inputs to guide the reconstruction of semantically consistent VR scenes. Furthermore, we construct a VR-EEG dataset that pairs real-time EEG recordings with VR scenes, and analyze how different types of scenes affect EEG responses across frequency bands. Our experimental results show that EEG2Gaussian can reconstruct VR scenes that are semantically aligned with the visual stimuli. Ablation studies verify the effectiveness of channel and band attention in EEG feature encoding, and demonstrate that combining high-level and low-level semantic features enhances the consistency and interpretability of the reconstructed scenes. Qiang Qu 0004, Xiaoming Chen 0006, Longfei Han, Yiran Shen 0001 |
ISMAR | 4 |
| 2025 | Hierarchical candidate recursive network for highlight restoration in endoscopic videos
Chenchu Xu, Jiangnan Wu, Dong Zhang 0009, Longfei Han, Dingwen Zhang, Junwei Han 0001 |
Expert Syst. Appl. | 4 |
| 2025 | LLaVA-Endo: a large language-and-vision assistant for gastrointestinal endoscopy
Jieru Yao, Xueran Li, Longfei Han, Yiwen Jia, Nian Liu 0002, Dingwen Zhang, Junwei Han 0001 |
Frontiers Comput. Sci. | 4 |
| 2025 | MHKD: Multi-Step Hybrid Knowledge Distillation for Low-Resolution Whole Slide Images Glomerulus DetectionabstractGlomerulus detection is a critical component of renal histopathology assessment, essential for diagnosing glomerulonephritis. To mitigate the increasing workload on pathologists, AI-assisted diagnostic methods based on high-resolution digital pathology whole slide images have been developed. However, these current AI-assisted approaches are limited to high-resolution whole slide images, necessitating expensive digital scanner equipment, high image storage costs, and significant computational complexity. To address this limitation, this paper pioneers a method for facilitating glomerulus detection in low-resolution human kidney pathology images. Specifically, we propose a novel multi-step hybrid knowledge distillation method. Our method distills both the global features and the semantic information through a hybrid knowledge distillation strategy that integrates offline and online knowledge distillation, where the information from high-resolution pathological images is successively transferred to student model from the global features in the shallow network layers to the semantic information of the back-end through a multi-step training strategy. Experimental results on two datasets show that the proposed method achieves effective detection outcomes for low-resolution kidney pathology images. Compared to other state-of-the-art detection techniques, our method achieves an ${AP}_{0.5:0.95}$ improvement of 23.1% on the private LN dataset and 15.9% on the public HUBMAP dataset. Xiangsen Zhang, Longfei Han, Chenchu Xu, Zhaohui Zheng 0004, Jin Ding, Xianghui Fu, Dingwen Zhang, Junwei Han 0001 |
IEEE J. Biomed. Health Informatics | 2 |
| 2025 | Prompting Vision-Language Model for Nuclei Instance Segmentation and ClassificationabstractNuclei instance segmentation and classification are a fundamental and challenging task in whole slide Imaging (WSI) analysis. Most dense nuclei prediction studies rely heavily on crowd labelled data on high-resolution digital images, leading to a time-consuming and expertise-required paradigm. Recently, Vision-Language Models (VLMs) have been intensively investigated, which learn rich cross-modal correlation from large-scale image-text pairs without tedious annotations. Inspired by this, we build a novel framework, called PromptNu, aiming at infusing abundant nuclei knowledge into the training of the nuclei instance recognition model through vision-language contrastive learning and prompt engineering techniques. Specifically, our approach starts with the creation of multifaceted prompts that integrate comprehensive nuclear knowledge, including visual insights from the GPT-4V model, statistical analyses, and expert insights from the pathology field. Then, we propose a novel prompting methodology that consists of two pivotal vision-language contrastive learning components: the Prompting Nuclei Representation Learning (PNuRL) and the Prompting Nuclei Dense Prediction (PNuDP), which adeptly integrates the expertise embedded in pre-trained VLMs and multifaceted prompts into the feature extraction and prediction process, respectively. Comprehensive experiments on six datasets with extensive WSI scenarios demonstrate the effectiveness of our method for both nuclei instance segmentation and classification tasks. The code is available at https://github.com/NucleiDet/PromptNu. Jieru Yao, Guangyu Guo 0001, Zhaohui Zheng 0004, Longfei Han, Dingwen Zhang, Junwei Han 0001 |
IEEE Trans. Medical Imaging | 5 |
| 2024 | RL-SeqISP: Reinforcement Learning-Based Sequential Optimization for Image Signal ProcessingabstractHardware image signal processing (ISP), aiming at converting RAW inputs to RGB images, consists of a series of processing blocks, each with multiple parameters. Traditionally, ISP parameters are manually tuned in isolation by imaging experts according to application-specific quality and performance metrics, which is time-consuming and biased towards human perception due to complex interaction with the output image. Since the relationship between any single parameter’s variation and the output performance metric is a complex, non-linear function, optimizing such a large number of ISP parameters is challenging. To address this challenge, we propose a novel Sequential ISP parameter optimization model, called the RL-SeqISP model, which utilizes deep reinforcement learning to jointly optimize all ISP parameters for a variety of imaging applications. Concretely, inspired by the sequential tuning process of human experts, the proposed model can progressively enhance image quality by seamlessly integrating information from both the image feature space and the parameter space. Furthermore, a dynamic parameter optimization module is introduced to avoid ISP parameters getting stuck into local optima, which is able to more effectively guarantee the optimal parameters resulting from the sequential learning strategy. These merits of the RL-SeqISP model as well as its high efficiency are substantiated by comprehensive experiments on a wide range of downstream tasks, including two visual analysis tasks (instance segmentation and object detection), and image quality assessment (IQA), as compared with representative methods both quantitatively and qualitatively. In particular, even using only 10% of the training data, our model outperforms other SOTA methods by an average of 7% mAP on two visual analysis tasks. Zhikun Zhao, Congyan Lang, Mingxuan Cai, Longfei Han, Juan Wang 0012, Bing Li 0001 |
AAAI | 6 |
| 2024 | Scalable Radar-based Roadside Perception: Self-localization and Occupancy Heat Map for Traffic Analysisabstract4D mmWave radar sensors are suitable for roadside perception in city-scale Intelligent Transportation Systems (ITS) due to their long sensing range, weatherproof functionality, simple mechanical design, and low manufacturing cost. In this work, we investigate radar-based ITS for scalable traffic analysis. Localization of these radar sensors at city scale is a fundamental task in ITS. For flexible sensor setups, it requires even more effort. To address this task, we propose a self-localization approach that matches two descriptions of the "road": the one from the geometry of the motion trajectories of cumulatively observed vehicles, and the other one from the aerial laser scan. An Iterative Closest Point (ICP) algorithm is used to register the motion trajectory in the road section of the laser scan. The resulting estimate of the transformation matrix represents the sensor pose in a global reference frame. We evaluate the results and show that the method outperforms other map-based radar localization methods, especially for the orientation estimation. Beyond the localization result, we project radar sensor data onto a city-scale laser scan and generate a scalable occupancy heat map as a traffic analysis tool. This is demonstrated using two radar sensors monitoring an urban area in the real world. Longfei Han, Qiuyu Xu, Klaus Kefferpütz, Gordon Elger, Jürgen Beyerer |
IV | 1 |
| 2024 | ASPS: Augmented Segment Anything Model for Polyp Segmentation
Huiqian Li, Dingwen Zhang, Jieru Yao, Longfei Han, Zhongyu Li 0006, Junwei Han 0001 |
MICCAI (9) | 4 |
| 2024 | M-RRFS: A Memory-Based Robust Region Feature Synthesizer for Zero-Shot Object Detection
Peiliang Huang, Dingwen Zhang, De Cheng, Longfei Han, Pengfei Zhu 0001, Junwei Han 0001 |
Int. J. Comput. Vis. | 4 |
| 2024 | Contextual Dependency Vision Transformer for spectrogram-based multivariate time series analysis
Jieru Yao, Longfei Han, Kaihui Yang, Guangyu Guo 0001, Nian Liu 0002, Xiankai Huang, Zhaohui Zheng 0004, Dingwen Zhang, Junwei Han 0001 |
Neurocomputing | 2 |
| 2024 | Position-based anchor optimization for point supervised dense nuclei detection
Jieru Yao, Longfei Han, Guangyu Guo 0001, Zhaohui Zheng 0004, Runmin Cong, Xiankai Huang, Jin Ding, Kaihui Yang, Dingwen Zhang, Junwei Han 0001 |
Neural Networks | 2 |
| 2024 | Pixel Distillation: Cost-Flexible Distillation Across Image Sizes and Heterogeneous NetworksabstractPrevious knowledge distillation (KD) methods mostly focus on compressing network architectures, which is not thorough enough in deployment as some costs like transmission bandwidth and imaging equipment are related to the image size. Therefore, we propose Pixel Distillation that extends knowledge distillation into the input level while simultaneously breaking architecture constraints. Such a scheme can achieve flexible cost control for deployment, as it allows the system to adjust both network architecture and image quality according to the overall requirement of resources. Specifically, we first propose an input spatial representation distillation (ISRD) mechanism to transfer spatial knowledge from large images to student's input module, which can facilitate stable knowledge transfer between CNN and ViT. Then, a Teacher-Assistant-Student (TAS) framework is further established to disentangle pixel distillation into the model compression stage and input compression stage, which significantly reduces the overall complexity of pixel distillation and the difficulty of distilling intermediate knowledge. Finally, we adapt pixel distillation to object detection via an aligned feature for preservation (AFP) strategy for TAS, which aligns output dimensions of detectors at each stage by manipulating features and anchors of the assistant. Comprehensive experiments on image classification and object detection demonstrate the effectiveness of our method. Guangyu Guo 0001, Dingwen Zhang, Longfei Han, Nian Liu 0002, Ming-Ming Cheng, Junwei Han 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Learning to Exploit the Sequence-Specific Prior Knowledge for Image Processing Pipelines OptimizationabstractThe hardware image signal processing (ISP) pipeline is the intermediate layer between the imaging sensor and the downstream application, processing the sensor signal into an RGB image. The ISP is less programmable and consists of a series of processing modules. Each processing module handles a subtask and contains a set of tunable hyperparameters. A large number of hyperparameters form a complex mapping with the ISP output. The industry typically relies on manual and time-consuming hyperparameter tuning by image experts, biased towards human perception. Recently, several automatic ISP hyperparameter optimization methods using downstream evaluation metrics come into sight. However, existing methods for ISP tuning treat the high-dimensional parameter space as a global space for optimization and prediction all at once without inducing the structure knowledge of ISP. To this end, we propose a sequential ISP hyperparameter prediction framework that utilizes the sequential relationship within ISP modules and the similarity among parameters to guide the model sequence process. We validate the proposed method on object detection, image segmentation, and image quality tasks. Haina Qin, Longfei Han, Weihua Xiong, Juan Wang 0012, Bing Li 0001, Weiming Hu 0004 |
CVPR | 2 |
| 2023 | Giving Text More Imagination Space for Image-text MatchingabstractImage-text matching is a hot topic in multi-modal analysis. The existing image-text matching algorithms focus on bridging the heterogeneity gap and mapping the feature into a common space under strong alignment assumption. However, these methods have unsatisfactory performance under the weak alignment scenario, which assumes that the text contains more abstract information, and the number of entities in the text is always fewer than objects in image. This is the first time, from our knowledge, to solve the image-text matching problem from the perspective of information difference with weak alignment. In order to both narrow the cross-modal heterogeneity gap and balance the information discrepancy, we proposed an imagination network to enrich the text modality based on pre-trained framework, which is helpful for image-text matching. The imagination network utilizes reinforcement learning to enhance the semantic information for text modality, and an action refinement strategy is designed to constrain the freedom and divergence of imagination. The experiment results show the superiority and generality of the proposed framework based on two pre-trained models, CLIP and BLIP on two most frequently-used datasets MSCOCO and Flickr30K. Longfei Han, Dingwen Zhang, Li Liu 0031, Junwei Han 0001, Huaxiang Zhang 0001 |
ACM Multimedia | 2 |
| 2023 | A Mutual Enhancement Framework for Specular Highlight Detection and Removal
Ge Huang, Jieru Yao, Peiliang Huang, Longfei Han |
PRCV (11) | 4 |
| 2023 | BMAnet: Boundary Mining With Adversarial Learning for Semi-Supervised 2D Myocardial Infarction SegmentationabstractAutomatic segmentation of myocardial infarction (MI) regions in late gadolinium-enhanced cardiac magnetic resonance images is an essential step in the computed diagnosis of myocardial infarction. Most of the current myocardial infarction region segmentation methods are based on fully supervised deep learning. However, cardiologists' annotation of myocardial infarction regions in cardiac magnetic resonance images during the diagnosis process is time-consuming and expensive. This paper proposes a semi-supervised myocardial infarction segmentation. It consists of two models: 1) a boundary mining model and 2) an adversarial learning model. The boundary mining model can solve the boundary ambiguity problem by enlarging the gap between the foreground and background features, thus segmenting the myocardial infarction region accurately. The adversarial learning model can make the boundary mining model learn from additional unlabeled data by evaluating the segmentation performance and providing pseudo supervision, which significantly increases the robustness of the boundary mining model. We conduct extensive experiments on an in-house myocardial magnetic resonance dataset. The experimental results on six evaluation metrics demonstrate that our method achieves excellent results in myocardial infarction segmentation and outperforms the state-of-the-art semi-supervised methods. Chenchu Xu, Dong Zhang 0009, Longfei Han, Yanping Zhang 0001, Jie Chen 0025, Shuo Li 0001 |
IEEE J. Biomed. Health Informatics | 4 |
| 2022 | Attention-Aware Learning for Hyperparameter Prediction in Image Processing Pipelines
Haina Qin, Longfei Han, Juan Wang 0012, Congxuan Zhang, Bing Li 0001, Weiming Hu 0004 |
ECCV (19) | 2 |
| 2020 | Deep supervised learning with mixture of neural networks
Yaxian Hu, Senlin Luo, Longfei Han, Limin Pan, Tiemei Zhang |
Artif. Intell. Medicine | 3 |
| 2017 | Soft-Margin Mixture of Regressions
Dong Huang 0007, Longfei Han, Fernando De la Torre |
CVPR | 2 |
| 2017 | Self-paced Mixture of RegressionsabstractMixture of regressions (MoR) is the well-established and effective approach to model discontinuous and heterogeneous data in regression problems. Existing MoR approaches assume smooth joint distribution for its good anlaytic properties. However, such assumption makes existing MoR very sensitive to intra-component outliers (the noisy training data residing in certain components) and the inter-component imbalance (the different amounts of training data in different components). In this paper, we make the earliest effort on Self-paced Learning (SPL) in MoR, i.e., Self-paced mixture of regressions (SPMoR) model. We propose a novel self-paced regularizer based on the Exclusive LASSO, which improves inter-component balance of training data. As a robust learning regime, SPL pursues confidence sample reasoning. To demonstrate the effectiveness of SPMoR, we conducted experiments on both the sythetic examples and real-world applications to age estimation and glucose estimation. The results show that SPMoR outperforms the state-of-the-arts methods. Longfei Han, Dingwen Zhang, Dong Huang 0007, Xiaojun Chang, Senlin Luo, Junwei Han 0001 |
IJCAI | 1 |
| 2017 | An Intelligible Risk Stratification Model Based on Pairwise and Size Constrained KmeansabstractHaving a system to stratify individuals according to risk is key to clinical disease prevention. This allows individuals identified at different risk tiers to benefit from further investigation and intervention. But the same risk score estimated for two different persons does not mean they need the same further investigation or represent the similarity health condition between two persons. Meanwhile, users still do not know a prior what most of the risk tiers are, and how many tiers should be found in risk stratification. In this paper, the proposed pairwise and size constrained Kmeans (PSCKmeans) method simultaneously integrates the limited supervised information and the size constraints to screen the high-risk population based on similarity measurement, and gets a feasible and balanced stratification solution to avoid cluster with few points. Results on China Health and Nutrition Survey public dataset and follow-up dataset show that the proposed PSCKmeans method can naturally grade the risk of diabetes into four tiers, and achieve 73.8%, 85.1%, and 0.95% sensitivity, specificity, and ratio of minimum to expected on testing data. The proposed method compares favorably with eight previous semisupervised clustering methods; it demonstrates that semisupervised clustering by unifying multiple forms of constraints can guide a good partition that is more relevant for the domain and find new categories through prior knowledge. Finally, this risk stratification model can provide a tool for risk stratification of clinical disease and be used for further intervention for people with similar health condition. Longfei Han, Senlin Luo, Huaiqing Wang, Limin Pan, Xincheng Ma, Tiemei Zhang |
IEEE J. Biomed. Health Informatics | 1 |
| 2015 | Rule Extraction From Support Vector Machines Using Ensemble Learning Approach: An Application for Diagnosis of DiabetesabstractDiabetes mellitus is a chronic disease and a worldwide public health challenge. It has been shown that 50-80% proportion of T2DM is undiagnosed. In this paper, support vector machines are utilized to screen diabetes, and an ensemble learning module is added, which turns the "black box" of SVM decisions into comprehensible and transparent rules, and it is also useful for solving imbalance problem. Results on China Health and Nutrition Survey data show that the proposed ensemble learning method generates rule sets with weighted average precision 94.2% and weighted average recall 93.9% for all classes. Furthermore, the hybrid system can provide a tool for diagnosis of diabetes, and it supports a second opinion for lay users. Longfei Han, Senlin Luo, Jianmin Yu, Limin Pan, Songjing Chen |
IEEE J. Biomed. Health Informatics | 1 |