EDBT 2026 Demo / reviewers in the wild / expert
Bin Chen 0022
dblp:22/5523-22
· DBLP profile ↗
14ranked-venue papers
0as first author
13since 2021 · last 2026
0000-0002-3979-021XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 10 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SAM-IAD: Injecting specific knowledge into SAM for industrial anomaly detection
Yichi Chen 0002, Bin Chen 0022, Weizhi Xian, Xinyi Gong, Jianwen Han, Xian Tao |
Knowl. Based Syst. | 2 |
| 2026 | MPFR: Memory prompt feature reconstruction for continual anomaly detection and segmentation
Yichi Chen 0002, Xian Tao, Bin Chen 0022, Pang-jo Chun, Xinmiao Zhou |
Pattern Recognit. | 3 |
| 2026 | Neighborhood Attention-based Feature Reconstruction for Image Anomaly Detection and LocalizationabstractWith the advancement of machine vision technology, automated vision inspection systems are needed in broad quality control scenarios. This article proposes a neighborhood attention-based feature reconstruction method for image anomaly detection and localization (NAFRAD). To address the challenges of data scarcity, low visibility, and irregular defect shapes in unsupervised anomaly detection, we introduce a feature reconstruction framework that preserves high-level abstract features rather than focusing on pixel-level reconstruction. This approach enhances model robustness and generalizability by leveraging neighborhood attention (NA) mechanisms, which simultaneously capture local details and the global context through a sliding window strategy. The NA-based autoencoder reconstructs normal features by aggregating local inductive biases with translational equivariance, enabling precise anomaly localization. Extensive experiments on the MVTec Anomaly Detection (MVTec AD) dataset—comprising 15 categories with 5,354 images—demonstrate the superiority of NAFRAD. It achieves state-of-the-art performance with AUROC \({}_{I}\) = 99.02, AUROC \({}_{P}\) = 98.99, and AP = 79.40, outperforming existing methods by 3.6% in AP and 0.89% in AUROC \({}_{P}\) . The framework’s effectiveness is validated through ablation studies, visualization of feature reconstruction, and comparisons with eight leading unsupervised methods. The code is made public at https://github.com/Math-Computer/NAFRAD . Weizhi Xian, Yichi Chen 0002, Bin Chen 0022, Leong Hou U, Shiyou Liu, Yong Feng 0002, Mingliang Zhou 0001, Sam Kwong |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2025 | OV-DQUO: Open-Vocabulary DETR with Denoising Text Query Training and Open-World Unknown Objects SupervisionabstractOpen-vocabulary detection aims to detect objects from novel categories beyond the base categories on which the detector is trained. However, existing open-vocabulary detectors trained on base category data tend to assign higher confidence to trained categories and confuse novel categories with the background. To resolve this, we propose OV-DQUO, an Open-Vocabulary DETR with Denoising text Query training and open-world Unknown Objects supervision. Specifically, we introduce a wildcard matching method. This method enables the detector to learn from pairs of unknown objects recognized by the open-world detector and text embeddings with general semantics, mitigating the confidence bias between base and novel categories. Additionally, we propose a denoising text query training strategy. It synthesizes foreground and background query-box pairs from open-world unknown objects to train the detector through contrastive learning, enhancing its ability to distinguish novel objects from the background. We conducted extensive experiments on the OV-COCO and OV-LVIS benchmarks, achieving new state-of-the-art results of 45.6 AP50 and 39.3 mAP on novel categories, respectively. Bin Chen 0022, Bin Kang, Yulin Li 0003, Weizhi Xian, Yichi Chen 0002 |
AAAI | 2 |
| 2025 | DeCLIP: Decoupled Learning for Open-Vocabulary Dense PerceptionabstractDense visual prediction tasks have been constrained by their reliance on predefined categories, limiting their applicability in real-world scenarios where visual concepts are unbounded. While Vision-Language Models (VLMs) like CLIP have shown promise in open-vocabulary tasks, their direct application to dense prediction often leads to suboptimal performance due to limitations in local feature representation. In this work, we present our observation that CLIP’s image tokens struggle to effectively aggregate information from spatially or semantically related regions, resulting in features that lack local discriminability and spatial consistency. To address this issue, we propose DeCLIP, a novel framework that enhances CLIP by decoupling the self-attention module to obtain "content" and "context" features respectively. The "content" features are aligned with image crop representations to improve local discriminability, while "context" features learn to retain the spatial correlations under the guidance of vision foundation models, such as DINO. Extensive experiments demonstrate that DeCLIP significantly outperforms existing methods across multiple open-vocabulary dense prediction tasks, including object detection and semantic segmentation. Code is available at https://github.com/xiaomoguhz/DeCLIP. Bin Chen 0022, Yulin Li 0003, Bin Kang, Yichi Chen 0002, Zhuotao Tian |
CVPR | 2 |
| 2025 | CalibCLIP: Contextual Calibration of Dominant Semantics for Text-Driven Image RetrievalabstractExisting Visual Language Models (VLMs) suffer structural limitations where a few low contribution tokens may excessively capture global semantics, dominating the information aggregation process and suppressing the discriminative features in text-driven image retrieval tasks. To address this, we introduce \textbf{CalibCLIP}, a training-free method designed to calibrate the suppressive effect of dominant tokens. Specifically, in the visual space, we propose the Contrastive Visual Enhancer (CVE), which decouples visual features into target and low information regions. Subsequently, it identifies dominant tokens and dynamically suppresses their representations.In the textual space, we introduce the Discriminative Concept Calibrator (DCC), which aims to differentiate between general and discriminative concepts within the text query. By mitigating the challenges posed by generic concepts and improving the representations of discriminative concepts, DCC strengthens the differentiation among similar samples. Finally, extensive experiments demonstrate consistent improvements across seven benchmarks spanning three image retrieval tasks, underscoring the effectiveness of CalibCLIP. Code is available at: https://github.com/kangbin98/CalibCLIP Bin Kang, Bin Chen 0022, Yulin Li 0003, Junzhi Zhao, Junle Wang, Zhuotao Tian |
ACM Multimedia | 2 |
| 2025 | Less Is More, but Where? Dynamic Token Compression via LLM-Guided Keyframe PriorabstractRecent advances in Video Large Language Models (VLLMs) have achieved remarkable video understanding capabilities, yet face critical efficiency bottlenecks due to quadratic computational growth with lengthy visual token sequences of long videos. While existing keyframe sampling methods can improve temporal modeling efficiency, additional computational cost is introduced before feature encoding, and the binary frame selection paradigm is found suboptimal. Therefore, in this work, we propose **Dy**namic **To**ken compression via LLM-guided **K**eyframe prior (**DyToK**), a training-free paradigm that enables dynamic token compression by harnessing VLLMs' inherent attention mechanisms. Our analysis reveals that VLLM attention layers naturally encoding query-conditioned keyframe priors, by which DyToK dynamically adjusts per-frame token retention ratios, prioritizing semantically rich frames while suppressing redundancies. Extensive experiments demonstrate that DyToK achieves state-of-the-art efficiency-accuracy tradeoffs. DyToK shows plug-and-play compatibility with existing compression methods, such as VisionZip and FastV, attaining 2.5x faster inference while preserving accuracy across multiple VLLMs, such as LLaVA-OneVision and Qwen2.5-VL. Code and models will be made publicly available. Yulin Li 0003, Haokun Gui, Ziyang Fan, Bin Kang, Bin Chen 0022, Zhuotao Tian |
NeurIPS | 6 |
| 2025 | Enhancing Multimodal Anomaly Detection via Asymmetric Dual-Branch Reverse Distillation
Zihe Chen, Bin Chen 0022, Yichi Chen 0002 |
Vis. Comput. | 2 |
| 2024 | Co-Salient Object Detection via Discriminative Prototypes ContrastabstractCo-salient object detection aims to detect co-salient objects in a group of images, combining collaborative segmentation and saliency detection, which is more challenging. Recent deep learning approaches identify co-salient objects by capturing the attention of consistent patterns within a group of images. However, due to limited semantic discriminability, these approaches often generate redundant attention unrelated to the co-salient object, resulting in inaccuracies. To address this, we refer to prototypical contrastive learning, and propose a prototype generation module to create discriminative prototypes representing both intra-group consistency and inter-group variance. These prototypes guide our proposed collaborative attention generation module, effectively enhancing co-salient object detection by highlighting relevant deep features. To ensure prototype’s discriminability, we add contrastive supervision for multi-task learning. Additionally, we design a position-independent contrast loss function to enhance intra-group consistency representation. Experiments demonstrate the superiority of our approach over existing state-of-the-art approaches on three challenging benchmarks, i.e., CoCA,CoSOD3k,and CoSal2015. Bin Chen 0022, Wenrui Fan, Yongjiang Liu |
ICASSP | 2 |
| 2024 | Unsupervised Multi-Modal Medical Image Registration via query-selected attention and decoupled Contrastive LearningabstractThe translation-based unsupervised deformable image registration method has become one of the classic methods for multi-modal image registration. In this paper, we propose a novel translation-based unsupervised deformable image registration approach. We use entropy to assess the significance of features in different image regions, and select those with minimal entropy to aid in network training. Additionally, we utilize an attention matrix to maintain feature relations in the source domain. By replacing the the widely used InfoNCE loss with the decoupled contrastive learning (DCL) loss, we eliminate the negative-positive-coupling (NPC) effect. Evaluation on two commonly-used datasets demonstrates the superior performance of our approach. The code is available at https://github.com/hzrhit/QSDCL-DFMIR. Zhenrong Huang, Bin Chen 0022 |
ICME | 2 |
| 2024 | Multi-Attribute Consistency Driven Visual Language Framework for Surface Defect DetectionabstractVisual Language Pre-training models encounter significant challenges stemming from the scarcity of data and the presence of ambiguous cues in industrial defect detection tasks. In this work, we propose a multi-attribute consistency-driven defect detection (MACD) framework to optimize text prompts in a coarse-to-fine trajectory. To bridge differences in domain knowledge, we build a structured attribute repository that contains descriptions of various defects’ inherent attributes. Based on this, we propose a multi-attribute consistency (MAC) module that can adequately model the global alignment between sentences with multiple attributes and defect images. Furthermore, we design a refined cross-alignment (RCA) module to determine the fine-grained correspondence between each attribute and the region within the image. Finally, the proposed method is experimentally validated on two benchmarks, resulting in significant performance improvements in a wide range of defective scenarios. Bin Kang, Bin Chen 0022, Weizhi Xian, Huifeng Chang |
ICME | 2 |
| 2024 | Perceptual Quality Analysis in Deep Domains Using Structure Separation and High-Order MomentsabstractImages are composed of “things” (i.e., structured objects) and “stuff” (i.e., textured surfaces), which have completely different effects on the human visual system (HVS). A good image quality assessment (IQA) method should fully consider the visual salience effects of image structures and the masking effects of image textures. In this article, we propose a perceptual quality analysis model using structure separation and high-order moments (SSHMPQA) in the deep domain. First, we use a total variation (TV) model to separate the perceptual structures in images from their deep feature maps, thereby maintaining meaningful object shapes with texture suppression and defining perceptual structure-aware distances in the deep domain. Then, we use the first- to fourth-order moments to calculate the mean, skewness and kurtosis of the probability distributions of the deep features. On this basis, we define a perceptual texture-aware distance in the deep domain. We then formulate the final model by solving a well-defined perceptual optimization problem. The proposed SSHMPQA model has good interpretability and is data-driven; moreover, the model does not require a complex and long training process because the optimization problem is convex and has an exact analytical solution. To verify the effectiveness of our model, comprehensive experiments are conducted. The experimental results show that the proposed model is superior to other state-of-the-art traditional and deep learning-based full-reference (FR) IQA methods. Weizhi Xian, Mingliang Zhou 0001, Bin Fang 0001, Tao Xiang 0001, Weijia Jia 0001, Bin Chen 0022 |
IEEE Trans. Multim. | 6 |
| 2024 | LGFDR: local and global feature denoising reconstruction for unsupervised anomaly detection
Yichi Chen 0002, Bin Chen 0022, Weizhi Xian |
Vis. Comput. | 2 |
| 2005 | Perspective image modeling for segmentationabstractThis paper proposes a novel thresholding method for real time perspective image segmentation based on an image gray level probability distribution model. A perspective image can be modeled as a mixture of the background and the object in the image, which are farther modeled as two independent Gaussian distributions. Therefore, given an input perspective image, our method firstly derives the gray level distributions of the background and the object, respectively, based on the image histogram analysis. Then the optimum threshold can be obtained from the intersection of the derived distributions of the background and the object. Experimental results on currency watermark image segmentation for real time printing quality inspection are rather encouraging. Lei He 0007, Bin Chen 0022 |
SMC | 2 |