EDBT 2026 Demo / reviewers in the wild / expert
Kanghao Chen
dblp:302/4949
· DBLP profile ↗
17ranked-venue papers
7as first author
17since 2021 · last 2026
0000-0002-8447-4400ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 4 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EvDiff3D: Event-Aware Diffusion Repair for High-Fidelity Event-Based 3D ReconstructionabstractEvent cameras are bio-inspired sensors that capture visual information through asynchronous brightness changes, offering distinct advantages including high temporal resolution and wide dynamic range. While prior research has investigated event-based 3D reconstruction for extreme scenarios, existing methods face inherent limitations and fail to fully exploit the unique characteristics of event data. In this paper, we present EvDiff3D, a novel two-stage 3D reconstruction framework that integrates event-based geometric constraints with an event-aware diffusion prior for appearance refinement. Our key insight lies in bridging the gap between physically grounded event-based reconstruction and data-driven appearance repair through a unified cyclical pipeline. In the first stage, we reconstruct a coarse 3D scene under supervision from event loss and event-based monocular depth constraints to preserve structural fidelity. The second stage fine-tunes an event-aware diffusion model based on a pretrained video diffusion model as a repair prior to enhance the appearance in under-constrained regions. Based on the diffusion model, our pipeline operates within a reconstruction-generation cycle that progressively refines both geometry and appearance using only event data. Extensive experiments on synthetic and real-world datasets demonstrate that EvDiff3D significantly outperforms existing methods in perceptual quality and structural consistency. Kanghao Chen, Lin Wang 0025, Zeyu Wang 0003 |
AAAI | 1 |
| 2026 | T-Rex-Omni: Integrating Negative Visual Prompt in Generic Object DetectionabstractObject detection methods have evolved from closed-set to open-set paradigms over the years. Current open-set object detectors, however, remain constrained by their exclusive reliance on positive indicators based on given prompts like text descriptions or visual exemplars. This positive-only paradigm experiences consistent vulnerability to visually similar but semantically different distractors. We propose T-Rex-Omni, a novel framework that addresses this limitation by incorporating negative visual prompts to negate hard negative distractors. Specifically, we first introduce a unified visual prompt encoder that jointly processes positive and negative visual prompts. Next, a training-free Negating Negative Computing (NNC) module is proposed to dynamically suppress negative responses during the probability computing stage. To further boost performance through fine-tuning, our Negating Negative Hinge (NNH) loss enforces discriminative margins between positive and negative embeddings. T-Rex-Omni supports flexible deployment in both positive-only and joint positive-negative inference modes, accommodating either user-specified or automatically generated negative examples. Extensive experiments demonstrate remarkable zero-shot detection performance, significantly narrowing the performance gap between visual-prompted and text-prompted methods while showing particular strength in long-tailed scenarios (51.2 AP_r on LVIS-minival). This work establishes negative prompts as a crucial new dimension for advancing open-set visual recognition systems. Jiazhou Zhou, Kanghao Chen, Lutao Jiang, Yuanhuiyi Lyu, Ying-Cong Chen, Lei Zhang 0001 |
AAAI | 3 |
| 2026 | Dual-modality adaptation in vision-language models for continual learning
Jiayang Zeng, Wentao Zhang 0005, Kanghao Chen, Jiantao Tan, Wei-Shi Zheng 0001 |
Neural Networks | 3 |
| 2026 | EvLight++: Low-Light Video Enhancement With an Event Camera: A Large-Scale Real-World Dataset, Novel Method, and MoreabstractEvent cameras offer significant advantages for low-light video enhancement, primarily due to their high dynamic range. Current research, however, is severely limited by the absence of large-scale, real-world, and spatio-temporally aligned event-video datasets. To address this, we introduce a large-scale dataset with over 30,000 pairs of frames and events captured under varying illumination. This dataset was curated using a robotic arm that traces a consistent non-linear trajectory, achieving spatial alignment precision under 0.03 mm and temporal alignment with errors under 0.01 s for 90% of the dataset. Based on the dataset, we propose EvLight++, a novel event-guided low-light video enhancement approach designed for robust performance in real-world scenarios. First, we design a multi-scale holistic fusion branch to integrate structural and textural information from both images and events. To counteract variations in regional illumination and noise, we introduce Signal-to-Noise Ratio (SNR)-guided regional feature selection, enhancing features from high SNR regions and augmenting those from low SNR regions by extracting structural information from events. To incorporate temporal information and ensure temporal coherence, we further introduce a recurrent module and temporal loss in the whole pipeline. Extensive experiments on ours and the synthetic SDSD dataset demonstrate that EvLight++ significantly outperforms both single image- and video-based methods by 1.37 dB and 3.71 dB, respectively. To further explore its potential in downstream tasks like semantic segmentation and monocular depth estimation, we extend our datasets by adding pseudo segmentation and depth labels via meticulous annotation efforts with foundation models. Experiments under diverse low-light scenes show that the enhanced results achieve a 15.97% improvement in mIoU for semantic segmentation. Kanghao Chen, Guoqiang Liang 0003, Yunfan Lu, Lin Wang 0025 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | PLNK: Prompt Learning With Neutral Knowledge for Few-Shot Out-of-Distribution DetectionabstractRecent developments in few-shot out-of-distribution (OOD) detection have yielded remarkable performance, benefiting from large pre-trained vision-language models (VLMs). Our prior work focuses on using in-distribution (ID) knowledge as references to learn richer knowledge beyond the textual semantics of class labels, which is prone to cause the model overconfidence and results in a limited score gap between ID and OOD data. In this paper, rather than treating ID knowledge as references, we propose Prompt Learning with Neutral Knowledge (PLNK) to better differentiate ID from OOD data. Our key insight lies in leveraging diverse neutral knowledge to improve ID discrimination while alleviating the inherent model overconfidence on OOD data induced by ID knowledge, thereby capturing the notable discrepancy between ID and OOD data. By introducing neutral knowledge with a balanced degree of similarity to both ID and OOD data, we amplify the discrepancy between the learnable prompt and the references (i.e., diverse neutral knowledge) for ID data, while reducing it for OOD data. In this way, the simple yet effective PLNK framework brings a notable score gap between ID and OOD data, thereby improving OOD detection. Moreover, we incorporate the visual neutral prompt with richer semantics alongside the original text-only reference. Comprehensive experiments show that our method consistently surpasses current state-of-the-art methods. The codes will be released publicly. Xinhua Lu, Runhe Lai, Yanqi Wu, Kanghao Chen, Zhiming Dai, Wei-Shi Zheng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Augmenting Continual Learning of Diseases with LLM-Generated Visual ConceptsabstractContinual learning is essential for medical image classification systems to adapt to dynamically evolving clinical environments. The integration of multimodal information can significantly enhance continual learning of image classes. However, while existing approaches do utilize textual modality information, they solely rely on simplistic templates with a class name, thereby neglecting richer semantic information. To address these limitations, we propose a novel framework that harnesses visual concepts generated by large language models (LLMs) as discriminative semantic guidance. Our method dynamically constructs a visual concept pool with a similarity-based filtering mechanism to prevent redundancy. Then, to integrate the concepts into the continual learning process, we employ a crossmodal image-concept attention module, coupled with an attention loss. Through attention, the module can leverage the semantic knowledge from relevant visual concepts and produce classrepresentative fused features for classification. Experiments on medical and natural image datasets show our method achieves state-of-the-art performance. Jiantao Tan, Peixian Ma, Zhiming Dai, Kanghao Chen |
BIBM | 4 |
| 2025 | FA: Forced Prompt Learning of Vision-Language Models for Out-of-Distribution DetectionabstractPre-trained vision-language models (VLMs) have advanced out-of-distribution (OOD) detection recently. However, existing CLIP-based methods often focus on learning OOD-related knowledge to improve OOD detection, showing limited generalization or reliance on external large-scale auxiliary datasets. In this study, instead of delving into the intricate OOD-related knowledge, we propose an innovative CLIP-based framework based on Forced prompt leArning (FA), designed to make full use of the In-Distribution (ID) knowledge and ultimately boost the effectiveness of OOD detection. Our key insight is to learn a prompt (i.e., forced prompt) that contains more diversified and richer descriptions of the ID classes beyond the textual semantics of class labels. Specifically, it promotes better discernment for ID images, by forcing more notable semantic similarity between ID images and the learnable forced prompt. Moreover, we introduce a forced coefficient, encouraging the forced prompt to learn more comprehensive and nuanced descriptions of the ID classes. In this way, FA is capable of achieving notable improvements in OOD detection, even when trained without any external auxiliary datasets, while maintaining an identical number of trainable parameters as CoOp. Extensive empirical evaluations confirm our method consistently outperforms current state-of-the-art methods. Code is available at https://github.com/0xFAFA/FA. Xinhua Lu, Runhe Lai, Yanqi Wu, Kanghao Chen, Wei-Shi Zheng 0001 |
ICCV | 4 |
| 2025 | Elite-EvGS: Learning Event-based 3D Gaussian Splatting by Distilling Event-to-Video PriorsabstractEvent cameras are bio-inspired sensors that output asynchronous and sparse event streams, instead of fixed frames. Benefiting from their distinct advantages, such as high dynamic range and high temporal resolution, event cameras have been applied to address 3D reconstruction, important for robotic mapping. Recently, neural rendering techniques, such as 3D Gaussian splatting (3DGS), have been shown successful in 3D reconstruction. However, it still remains under-explored how to develop an effective event-based 3DGS pipeline. In particular, as 3DGS typically depends on high-quality initialization and dense multiview constraints, a potential problem appears for the 3DGS optimization with events given its inherent sparse property. To this end, we propose a novel event-based 3DGS framework, named Elite-EvGS. Our key idea is to distill the prior knowledge from the off-the-shelf event-to-video (E2V) models to effectively reconstruct 3D scenes from events in a coarse-to-fine optimization manner. Specifically, to address the complexity of 3DGS initialization from events, we introduce a novel warm-up initialization strategy that optimizes a coarse 3DGS from the frames generated by E2V models and then incorporates events to refine the details. Then, we propose a progressive event supervision strategy that employs the window-slicing operation to progressively reduce the number of events used for supervision. This subtly relives the temporal randomness of the event frames, benefiting the optimization of local textural and global structural details. Experiments on the benchmark datasets demonstrate that Elite-EvGS can reconstruct 3D scenes with better textural and structural details. Meanwhile, our method yields plausible performance on the captured real-world data, including diverse challenging conditions, such as fast motion and low light scenes. For demo and more results, please check our project page. Kanghao Chen, Lin Wang 0025 |
ICRA | 2 |
| 2025 | Hierarchical Vision-Language Learning for Medical Out-of-Distribution Detection
Runhe Lai, Xinhua Lu, Kanghao Chen, Qichao Chen, Wei-Shi Zheng 0001 |
MICCAI (5) | 3 |
| 2025 | Event-Guided Consistent Video Enhancement with Modality-Adaptive Diffusion PipelineabstractRecent advancements in low-light video enhancement (LLVE) have increasingly leveraged both RGB and event cameras to improve video quality under challenging conditions. However, existing approaches share two key drawbacks. First, they are tuned for steady low-light scenes, so their performance drops when illumination varies. Second, they assume every sensing modality is always available, while real systems may lose or corrupt one of them. These limitations make the methods brittle in dynamic, real-world settings. In this paper, we propose EVDiffuser, a novel framework for consistent LLVE that integrates RGB and event data through a modality-adaptive diffusion pipeline. By harnessing the powerful priors of video diffusion models, EVDiffuser enables consistent video enhancement and generalization to diverse scenarios under varying illumination, where RGB or events may even be absent. Specifically, we first design a modality-agnostic conditioning mechanism based on a diffusion pipeline by treating the two modalities as optional conditions, which is fine-tuned using augmented and integrated datasets. Furthermore, we introduce a modality-adaptive guidance rescaling that dynamically adjusts the contribution of each modality according to sensor-specific characteristics. Additionally, we establish a benchmark that accounts for varying illumination and diverse real-world scenarios, facilitating future research on consistent event-guided LLVE. Our experiments demonstrate state-of-the-art performance across challenging scenarios (i.e., varying illumination) and sensor-based settings (e.g., event-only, RGB-only), highlighting the generalization of our framework. Kanghao Chen, Guoqiang Liang 0003, Lutao Jiang, Zeyu Wang 0003, Ying-Cong Chen |
NeurIPS | 1 |
| 2025 | PASS: Path-selective State Space Model for Event-based RecognitionabstractEvent cameras are bio-inspired sensors that capture intensity changes asynchronously with distinct advantages, such as high temporal resolution. Existing methods for event-based object/action recognition predominantly sample and convert event representation at every fixed temporal interval (or frequency). However, they are constrained to processing a limited number of event lengths and show poor frequency generalization, thus not fully leveraging the event's high temporal resolution. In this paper, we present our PASS framework, exhibiting superior capacity for spatiotemporal event modeling towards a larger number of event lengths and generalization across varying inference temporal frequencies. Our key insight is to learn adaptively encoded event features via the state space models (SSMs), whose linear complexity and generalization on input frequency make them ideal for processing high temporal resolution events. Specifically, we propose a Path-selective Event Aggregation and Scan (PEAS) module to encode events into features with fixed dimensions by adaptively scanning and selecting aggregated event presentation. On top of it, we introduce a novel Multi-faceted Selection Guiding (MSG) loss to minimize the randomness and redundancy of the encoded features during the PEAS selection process. Our method outperforms prior methods on five public datasets and shows strong generalization across varying inference frequencies with less accuracy drop (ours -8.62% v.s. -20.69% for the baseline). Moreover, our model exhibits strong long spatiotemporal modeling for a broader distribution of event length (1-10^9), precise temporal perception, and effective generalization for real-world scenarios. Code and checkpoints will be released upon acceptance. Jiazhou Zhou, Kanghao Chen, Lin Wang 0025 |
NeurIPS | 2 |
| 2024 | Towards Robust Event-guided Low-Light Image Enhancement: A Large-Scale Real-World Event-Image Dataset and Novel ApproachabstractEvent camera has recently received much attention for low-light image enhancement (LIE) thanks to their distinct advantages, such as high dynamic range. However, current research is prohibitively restricted by the lack of large-scale, real-world, and spatial-temporally aligned event-image datasets. To this end, we propose a real-world (indoor and outdoor) dataset comprising over 30K pairs of images and events under both low and normal illumination conditions. To achieve this, we utilize a robotic arm that traces a consistent non-linear trajectory to curate the dataset with spatial alignment precision under 0.03mm. We then introduce a matching alignment strategy, rendering 90% of our dataset with errors less than 0.01s. Based on the dataset, we propose a novel event-guided LIE approach, called EvLight, towards robust performance in real-world low-light scenes. Specifically, we first design the multiscale holistic fusion branch to extract holistic structural and textural information from both events and images. To ensure robustness against variations in the regional illumination and noise, we then introduce a Signal-to-Noise-Ratio (SNR)-guided regional feature selection to selectively fuse features of images from regions with high SNR and enhance those with low SNR by extracting regional structure information from events. Extensive experiments on our dataset and the synthetic SDSD dataset demonstrate our EvLight significantly surpasses the frame-based methods, e.g., [4] by 1.14 dB and 2.62 dB, respectively. Guoqiang Liang 0003, Kanghao Chen, Yunfan Lu, Lin Wang 0025 |
CVPR | 2 |
| 2024 | LaSe-E2V: Towards Language-guided Semantic-aware Event-to-Video ReconstructionabstractEvent cameras harness advantages such as low latency, high temporal resolution, and high dynamic range (HDR), compared to standard cameras. Due to the distinct imaging paradigm shift, a dominant line of research focuses on event-to-video (E2V) reconstruction to bridge event-based and standard computer vision. However, this task remains challenging due to its inherently ill-posed nature: event cameras only detect the edge and motion information locally. Consequently, the reconstructed videos are often plagued by artifacts and regional blur, primarily caused by the ambiguous semantics of event data. In this paper, we find language naturally conveys abundant semantic information, rendering it stunningly superior in ensuring semantic consistency for E2V reconstruction. Accordingly, we propose a novel framework, called LaSe-E2V, that can achieve semantic-aware high-quality E2V reconstruction from a language-guided perspective, buttressed by the text-conditional diffusion models. However, due to diffusion models' inherent diversity and randomness, it is hardly possible to directly apply them to achieve spatial and temporal consistency for E2V reconstruction. Thus, we first propose an Event-guided Spatiotemporal Attention (ESA) module to condition the event data to the denoising pipeline effectively. We then introduce an event-aware mask loss to ensure temporal coherence and a noise initialization strategy to enhance spatial consistency. Given the absence of event-text-video paired data, we aggregate existing E2V datasets and generate textual descriptions using the tagging models for training and evaluation. Extensive experiments on three datasets covering diverse challenging scenarios (e.g., fast motion, low light) demonstrate the superiority of our method. Demo videos for the results are attached to the project page. Kanghao Chen, Jiazhou Zhou, Zeyu Wang 0003, Lin Wang 0025 |
NeurIPS | 1 |
| 2023 | PCCT: Progressive Class-Center Triplet Loss for Imbalanced Medical Image ClassificationabstractImbalanced training data in medical image diagnosis is a significant challenge for diagnosing rare diseases. For this purpose, we propose a novel two-stage Progressive Class-Center Triplet (PCCT) framework to overcome the class imbalance issue. In the first stage, PCCT designs a class-balanced triplet loss to coarsely separate distributions of different classes. Triplets are sampled equally for each class at each training iteration, which alleviates the imbalanced data issue and lays solid foundation for the successive stage. In the second stage, PCCT further designs a class-center involved triplet strategy to enable a more compact distribution for each class. The positive and negative samples in each triplet are replaced by their corresponding class centers, which prompts compact class representations and benefits training stability. The idea of class-center involved loss can be extended to the pair-wise ranking loss and the quadruplet loss, which demonstrates the generalization of the proposed framework. Extensive experiments support that the PCCT framework works effectively for medical image classification with imbalanced training images. On four challenging class-imbalanced datasets (two skin datasets Skin7 and Skin 198, one chest X-ray dataset ChestXray-COVID, and one eye dataset Kaggle EyePACs), the proposed approach respectively obtains the mean F1 score 86.20, 65.20, 91.32, and 87.18 over all classes and 81.40, 63.87, 82.62, and 79.09 for rare classes, achieving state-of-the-art performance and outperforming the widely used methods for the class imbalance issue. Kanghao Chen, Weixian Lei, Wei-Shi Zheng 0001 |
IEEE J. Biomed. Health Informatics | 1 |
| 2022 | Expert with Outlier Exposure for Continual Learning of New DiseasesabstractCurrent intelligent diagnosis systems struggle to continually learn to diagnose more and more diseases due to catastrophic forgetting of old knowledge when learning new knowledge. Although storing small old data for subsequent continual learning can effectively help alleviate the forgetting issue, the heavy data imbalance between old classes and to-be-learned new classes in classifier training often causes biased prediction towards the new classes just learned by the updated classifier. In this study, an outlier detection technique is novelly applied to train an additional expert classifier for new classes to help alleviate the class imbalance issue and discriminate the learned new classes from old classes during inference (instead of the training phase). Specially, the stored small data of old classes are considered as outliers during training the expert classifier, such that the output probability distributions from the expert classifier are expected to be obviously different between test data of the old classes and those of the new classes. Such difference between old classes and new classes can be used to fine-tune the original output from the updated classifier which is responsible for prediction of all learned (old and new) classes. During inference, a novel ensemble strategy is proposed to combine the predictions from the updated classifier, the expert classifier, and the previously learned old classifier. The proposed learning and inference framework can be easily combined with existing continual learning strategies. Empirical evaluations on three medical image datasets and one natural image dataset show that the proposed framework can effectively improve continual learning performance. Zhengjing Xu, Kanghao Chen, Wei-Shi Zheng 0001, Zhijun Tan |
BIBM | 2 |
| 2022 | Improving Class Balancing at Both Feature Extractor and Classifier HeadabstractTraining data are often imbalanced across classes in practice, and such class imbalance issue often causes model predictions biased toward majority classes during inference. Different from existing solutions which employ various training strategies to alleviate the class imbalance issue, this study proposes a novel two-head model architecture to help alleviate the issue. One auxiliary classifier head helps the feature extractor of the classifier more fairly learn to extract features for each class, and the main classifier head learns in a more class-balanced manner by dividing each majority class into multiple clusters in advance and considering each cluster as a new class. Extensive empirical evaluations on four class-imbalanced image datasets showed that the proposed approach achieves state-of-the-art classification performance. Kanghao Chen, Huijuan Lu, Wei-Shi Zheng 0001 |
ICME | 1 |
| 2021 | Alleviating Data Imbalance Issue with Perturbed Input During Inference
Kanghao Chen, Huijuan Lu, Chenghua Zeng, Wei-Shi Zheng 0001 |
MICCAI (5) | 1 |