Jin Wan

dblp:52/8781 · DBLP profile ↗
← Back
38ranked-venue papers
6as first author
34since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 21 · 4 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 3 first-author · 17 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Graph of Verification: Structured Verification of LLM Reasoning with Directed Acyclic Graphs
abstract
Verifying the complex and multi-step reasoning of Large Language Models (LLMs) is a critical challenge, as holistic methods often overlook localized flaws. Step-by-step validation is a promising alternative, yet existing methods are often rigid. They struggle to adapt to diverse reasoning structures, from formal proofs to informal natural language narratives. To address this adaptability gap, we propose the Graph of Verification (GoV), a novel framework for adaptable and multi-granular verification. GoV's core innovation is its flexible node block architecture. This mechanism allows GoV to adaptively adjust its verification granularity—from atomic steps for formal tasks to entire paragraphs for natural language—to match the native structure of the reasoning process. This flexibility allows GoV to resolve the fundamental trade-off between verification precision and robustness. Experiments on both well-structured and loosely-structured benchmarks demonstrate GoV's versatility. The results show that GoV's adaptive approach significantly outperforms both holistic baselines and other state-of-the-art decomposition-based methods, establishing a new standard for training-free reasoning verification.
Jiwei Fang, Bin Zhang 0052, Changwei Wang 0001, Jin Wan, Zhiwei Xu 0005
AAAI4
2026 Normality-enhanced knowledge distillation network for unsupervised industrial anomaly detection
Gang Li 0005, Tianjiao Chen, Jin Wan, Mingle Zhou, Delong Han, Min Li 0033
Eng. Appl. Artif. Intell.3
2026 Enhancing mixture-of-experts model with prior knowledge for infrared and visible image fusion in complex degraded environments
Gang Li 0005, Chengrun Jiang, Jin Wan, Mingle Zhou, Delong Han
Expert Syst. Appl.4
2026 A Novel Dataset and Lightweight Distillation Baseline for Highlight Transparent Object Detection
Gang Li 0005, Qinghui Chen, Qunshu Zhang, Jin Wan, Maomao Xiong, Cong Bai, Dagang Li 0001, Wenyin Zhang, Jinglin Zhang 0004, Shengyong Chen
Int. J. Comput. Vis.6
2026 Shadow vanishing point detection via combined human/shadow adaptive modulation
Jin Wan, Hui Yin 0002, Zhenyao Wu, Xinyi Wu 0002, Song Wang 0002
Signal Process. Image Commun.1
2026 Boosting Small Object Detection via High-Frequency Feature Oriented Network
abstract
Small Object Detection (SOD) aims to accurately identify and locate small objects in images. However, existing methods usually focus on exploring spatial domain features, neglecting high-frequency features that preserve fine-grained details such as texture and edge information. To overcome this limitation, we propose a High-Frequency Feature-Oriented Network (HFFO-Net). First, we introduce the Channel- wise Frequency Modulation Module (CFMM), which leverages the 2D Discrete Cosine Transform (DCT) to accentuate salient frequency components while mitigating noise interference. Second, we design a High-Frequency Oriented Module (HFOM), which utilizes the Channel Selection Branch (CSB) and Spatial Selection Branch (SSB) to highlight small objects in the channel and spatial region. Third, we introduce a Dual-Query Attention Fusion Mechanism (DQAFM), which reduces the semantic gap between spatial and frequency features and achieves better feature fusion through bidirectional cross-attention. Extensive experiments are implemented, and the corresponding results demonstrate that HFFO-Net excels at detecting small objects.
Min Li 0033, Zhaofei Hao, Gang Li 0005, Jin Wan, Delong Han, Mingle Zhou
IEEE Signal Process. Lett.4
2026 Multimodal Industrial Anomaly Detection via Geometric Prior
abstract
The purpose of multimodal industrial anomaly detection is to detect complex geometric shape defects such as subtle surface deformations and irregular contours that are difficult to detect in 2D-based methods. However, current multimodal industrial anomaly detection lacks the effective use of crucial geometric information like surface normal vectors and 3D shape topology, resulting in low detection accuracy. In this paper, we propose a novel Geometric Prior-based Anomaly Detection network (GPAD). Firstly, we propose a point cloud expert model to perform fine-grained geometric feature extraction, employing differential normal vector computation to enhance the geometric details of the extracted features and generate geometric prior. Secondly, we propose a two-stage fusion strategy to efficiently leverage the complementarity of multimodal data as well as the geometric prior inherent in 3D points. We further propose attention fusion and anomaly regions segmentation based on geometric prior, which enhance the model’s ability to perceive geometric defects. Extensive experiments show that our multimodal industrial anomaly detection model outperforms the State-of-the-art (SOTA) methods in detection accuracy on both MVTec-3D AD and Eyecandies datasets.
Min Li 0033, Gang Li 0005, Jin Wan, Delong Han
IEEE Trans. Circuits Syst. Video Technol.5
2025 Few-Shot Relation Extraction via Semantically Related Negative Samples
abstract
Few-shot relation extraction aims to identify and classify specific semantic relations between entities from text with a small number of annotated examples. Recent studies have shown that innovative model designs and learning strategies can significantly enhance the model’s generalization ability and performance, even with limited training samples. However, the samples used for training face the dual challenges of scarce annotated data and semantic ambiguity. Given the limitations of the existing SaCon framework in negative sample generation and discrimination efficiency, we propose a dynamic adversarial negative sample enhancement strategy. This strategy introduces adversarial perturbations before encoding the pre-trained language model by constructing a multi-granularity semantic space. Specifically, we first design an entity permutation mechanism to randomly exchange the subject/object entities of sentences in a small batch to generate negative sample clusters with similar semantics but misplaced relations, then, we integrate a multi-view contrastive learning framework to embed adversarial samples into the feature space topology optimization process.To strengthen boundary-sensitive features, an adaptive margin ranking loss function is proposed to dynamically adjust the representation distance constraints of positive and negative samples, forcing the model to capture the deep semantic invariance of relational predicates under limited samples. This method aims to optimize the traditional negative sample random sampling paradigm, actively explore the semantic space through an adversarial generation mechanism, and construct “difficult samples” with minimal semantic deviation through gradient back propagation, thereby improving the model’s ability to parse implicit relational patterns. The experiment result shows that this framework effectively alleviates the risk of overfitting in small sample scenarios by decoupling relational semantics and surface syntactic features, and its dynamic loss design provides mathematical guarantees for orthogonal separation of feature space. Our code is available online at https://github.com/hhy-test/ESCR.
Delong Han, Hongyu Hao, Jin Wan, Gang Li 0005, Min Li 0033, Mingle Zhou
ECAI3
2025 STAD: Joint Spatial-Temporal Dimension and Channel Correlation for Time Series Anomaly Detection
abstract
Accurately identifying real anomalies and pseudo-anomalies in complex multi-dimensional time series data has been a difficult problem in time series anomaly detection. To solve this problem, this paper proposes a new framework, STAD, that joint temporal and spatial dimensions. This framework guides the model to capture the correlation information between channels It also aims to learn the deep feature representation of sequences by mining potential information in the spatialtemporal dimension. It can effectively distinguish between true and false anomalies by comparing information from spatial and temporal dimensions. STAD identifies and integrates correlated channels by using a correlation aggregation mechanism to join multiple channels and detect anomalies.In addition, the KAN mixer designed in this paper can effectively extract features from different spatial locations in the spatial-temporal dimension. Through extensive experiments on several public datasets, STAD demonstrates its superiority in terms of accuracy and robustness.
Mingle Zhou, Xingli Wang, Delong Han, Jin Wan
ICASSP4
2025 HGCF: Hierarchical Geometry-Color Fusion for Multimodal Industrial Anomaly Detection
abstract
While current multimodal anomaly detection methods predominantly employ intermediate fusion strategies, they often suffer from inadequate cross-modal interaction and irreversible information loss during feature alignment processes. To overcome these limitations, we propose Hierarchical Geometry-Color Fusion (HGCF), a novel framework that establishes deep synergistic relationships between RGB texture features and point cloud geometric representations. Firstly, we propose a bidirectional cross-modal early fusion mechanism that enables complementary information exchange between point cloud and RGB modalities at the input level. Secondly, we introduce a local self-supervised geometric color reconstruction network with group-wise feature alignment, enhancing fine-grained feature extraction through joint color-geometry reconstruction tasks. Finally, we propose a local window spatial-consistent attention fusion, which achieves semantic consistency and spatial consistency by emphasizing local mutation features to improve the detection of subtle anomalies. Extensive experiments show our model achieves 99.1% I-AUROC on MVTec 3D-AD and 91.7% on Eyecandies, both surpassing state-of-the-art methods.
Min Li 0033, Delong Han, Jin Wan, Gang Li 0005
ACM Multimedia5
2025 Exploring Multimodal Prompts For Unsupervised Continuous Anomaly Detection
abstract
Unsupervised Continuous Anomaly Detection (UCAD) is gaining attention for effectively addressing the catastrophic forgetting and heavy computational burden issues in traditional Unsupervised Anomaly Detection (UAD). However, existing UCAD approaches that rely solely on visual information are insufficient to capture the manifold of normality in complex scenes, thereby impeding further gains in anomaly detection accuracy. To overcome this limitation, we propose an unsupervised continual anomaly detection framework grounded in multimodal prompting. Specifically, we introduce a Continual Multimodal Prompt Memory Bank (CMPMB) that progressively distills and retains prototypical normal patterns from both visual and textual domains across consecutive tasks, yielding a richer representation of normality. Furthermore, we devise a Defect-Semantic-Guided Adaptive Fusion Mechanism (DSG-AFM) that integrates an Adaptive Normalization Module (ANM) with a Dynamic Fusion Strategy (DFS) to jointly enhance detection accuracy and adversarial robustness. Benchmark experiments on MVTec AD and VisA datasets show that our approach achieves state-of-the-art (SOTA) performance on image-level AUROC and pixel-level AUPR metrics.
Mingle Zhou, Jin Wan, Gang Li 0005, Min Li 0033
ACM Multimedia3
2025 Intra-domain self generalization network for intelligent fault diagnosis of bearings under unseen working conditions
Zhijun Ren, Linbo Zhu, Tantao Lin, Yongsheng Zhu, Jin Wan
Adv. Eng. Informatics7
2025 Multimodal feature cooperative refinement for few-shot anomaly detection
Delong Han, Gang Li 0005, Mingle Zhou, Jin Wan, Min Li 0033
Adv. Eng. Informatics5
2025 DualPhys-GS: Dual physically-guided 3D Gaussian splatting for underwater scene reconstruction
Guangzhi Han, Jin Wan, Yuan Gao 0033, Delong Han
Comput. Graph.3
2025 MemMambaAD: Memory-augmented state space model for multivariate time series anomaly detection
Gang Li 0005, Mingchao Ge, Jin Wan, Delong Han, Min Li 0033, Mingle Zhou
Eng. Appl. Artif. Intell.3
2025 Scene text image super-resolution with semantic-aware interaction
Mingle Zhou, Jin Wan, Delong Han, Min Li 0033, Gang Li 0005
Eng. Appl. Artif. Intell.3
2025 Dual-domain divide-and-conquer for scene text image super-resolution
Zhengqian Feng, Jin Wan, Mingle Zhou
Knowl. Based Syst.4
2025 Reconsidering learnable fine-grained text prompts for few-shot anomaly detection in visual-language models
Delong Han, Mingle Zhou, Jin Wan, Min Li 0033, Gang Li 0005
Neural Networks4
2025 EnIter: Enhancing Iterative Multi-View Depth Estimation with Universal Contextual Hints
abstract
Iterative inference approaches have shown promising success in the task of multi-view depth estimation. However, these methods put excessive emphasis on the universal inter-view correspondences while neglecting the correspondence ambiguity in regions of low texture and depth discontinuous areas. Thus, they are prone to produce inaccurate or even erroneous depth estimations, which is further exacerbated due to cumulative errors especially in the iterative pipeline, providing unreliable information in many real-world scenarios. In this article, we revisit this issue from the intra-view contextual hints and introduce a novel enhancing iterative approach, named EnIter. Concretely, at the beginning of each iteration, we present a Depth Intercept (DI) modulator to provide more accurate depth by aggregating neighbor uncertainty, correlation volume of reference and normal. This plug and play modulator is effective at intercepting the erroneous depth estimations with implicit guidance from the universal correlation contextual hints, especially for the challenging regions. Furthermore, at the end of each iteration, we refine the depth map with another plug and play modulator termed as Depth Refine (DR). It mines the latent structure knowledge of reference contextual hints and establishes one-way dependency using local attention from reference features to depth, yielding delicate depth in detail. Extensive experiment demonstrates that our method not only achieves state-of-the-art performance over existing models but also exhibits remarkable universality in popular iterative pipelines, e.g., CasMVS, UCSNet, TransMVS, and UniMVS.
Qianqian Du 0002, Hui Yin 0002, Lang Nie, Jin Wan
ACM Trans. Multim. Comput. Commun. Appl.5
2024 Time Series Anomaly Detection via Temporal Dependencies and Multivariate Correlations Integrating
Gang Li 0005, Mingchao Ge, Mingle Zhou, Jin Wan, Delong Han
ICONIP (3)4
2024 Neural architecture search for multi-sensor information fusion-based intelligent fault diagnosis
Tantao Lin, Zhijun Ren, Linbo Zhu, Yongsheng Zhu, Jin Wan
Adv. Eng. Informatics7
2024 Estimating intrinsic characteristics of images for shadow removal
Hui Yin 0002, Jin Wan, Zhenyao Wu, Xinyi Wu 0002, Song Wang 0002
Comput. Graph.4
2024 CRFormer: A cross-region transformer for shadow removal
Jin Wan, Hui Yin 0002, Zhenyao Wu, Xinyi Wu 0002, Song Wang 0002
Image Vis. Comput.1
2024 A novel neural network architecture utilizing parametric-logarithmic-modulus-based activation function: Theory, algorithm, and applications
Zidong Wang 0001, Jin Wan, Guoping Lu, Weibo Liu 0001
Knowl. Based Syst.3
2024 Reference-Based Image Dehazing With Internal and External Contrastive Learning
abstract
Collecting paired pixel-aligned hazy/haze-free image pairs in real-world is arduous for full-supervised image dehazing. Alternatively, methods employing unpaired hazy/clear images have been developed, yet their learning ability about content information of the hazy images is easily disturbed by content-independent clear images, causing artifact problems, particularly for thick hazy images. To address the above issues, we propose a new reference-based image dehazing paradigm with hazy/reference images, where the reference image is clear and taken at the same scene as the hazy image. Therefore, how to maximize the reference value from the hazy/reference images with similar content but unaligned pixels becomes a key issue. Here, we construct a reference-based contrastive learning framework to realize the effective utilization of hazy/reference image pairs. Specifically, internal contrastive learning is designed to preserve the local content invariance between the dehazed images and hazy images in a patch-wise contrastive manner, while the other external contrastive learning learns the global content consistency between the dehazed images and reference images in an overall contrastive manner. Additionally, we design a style consistency loss committee consisting of a regular adversarial loss and a style loss. The former aims to ensure each dehazed image consistent with the overall style distribution of the entire reference set, while the latter is intended to make each dehazed image have an exclusive style with the corresponding reference image. Extensive experiments corroborate that the reference-based dehazing paradigm is recommendable and reliable, and the proposed method performs admirably against other state-of-the-art methods.
Hui Yin 0002, Ai-Xin Chong, Jin Wan
IEEE Trans. Circuits Syst. Video Technol.4
2023 Joint Memory Propagation and Rectification for Video Object Segmentation
Hui Yin 0002, Jin Wan, Jianhuan Chen
ICIG (4)4
2023 SA-Net: Scene-Aware Network for Cross-domain Stereo Matching
Ai-Xin Chong, Hui Yin 0002, Jin Wan, Qianqian Du 0002
Appl. Intell.3
2023 Generalized predictive control using improved recurrent fuzzy neural network for a boiler-turbine unit
Jin Wan
Eng. Appl. Artif. Intell.2
2022 Style-Guided Shadow Removal
Jin Wan, Hui Yin 0002, Zhenyao Wu, Xinyi Wu 0002, Song Wang 0002
ECCV (19)1
2022 Is It Necessary to Transfer Temporal Knowledge for Domain Adaptive Video Semantic Segmentation?
Xinyi Wu 0002, Zhenyao Wu, Jin Wan, Lili Ju, Song Wang 0002
ECCV (27)3
2022 Multi-hierarchy feature extraction and multi-step cost aggregation for stereo matching
Ai-Xin Chong, Hui Yin 0002, Jin Wan
Neurocomputing4
2022 Edge Aware Network for Image Dehazing
abstract
The learning-based methods have recently shown their advantages in the image dehazing task. However, most existing learning-based methods do not pay much attention to the restoration in the edges of the hazy image, resulting in the edge blur of the dehazing results. To mitigate this issue, in this letter, we propose a novel Edge Aware Network (EA-Net) for image dehazing, which can simultaneously model edge features and contextual features into a single network for restoring haze-free image with sharp edges. Firstly, we extract the multi-scales contextual features of hazy image by a progressive fusion way. Furthmore, the abundant edge features are inferred by a edge subnetwork with Edge Feature Extraction Module(EFEM). Finally, we present an Edge Attention (EA) mechanism to couple the edge features with contextual features at various resolutions for sufficiently leveraging these complementary features. Due to the rich edge information and feature fusion strategy, the fused features can make the haze-free image to be clearer, especially at the edges, which is very important for the high-level vision tasks. Extensive experiments demonstrate that the proposed method achieves significant improvements over the state-of-the-art methods.
Hui Yin 0002, Jin Wan, Ai-Xin Chong
IEEE Signal Process. Lett.3
2021 Pop-net: A self-growth network for popping out the salient object in videos
abstract
Abstract It is a big challenge for unsupervised video segmentation without any object annotation or prior knowledge. In this article, we formulate a completely unsupervised video object segmentation network which can pop out the most salient object in an input video by self‐growth, called Pop‐Net. Specifically, in this article, a novel self‐growth strategy which helps a base segmentation network to gradually grow to stick out the salient object as the video goes on, is introduced. To solve the sample generation problem for the unsupervised method, the sample generation module which fuses the appearance and motion saliency is proposed. Furthermore, the proposed sample optimization module improves the samples by using contour constrains for each self‐growth step. Experimental results on several datasets (DAVIS, DAVSOD, VideoSD, Segtrack‐v2) show the effectiveness of the proposed method. In particular, the state‐of‐the‐art methods on completely unfamiliar datasets (no fine‐tuned datasets) are performed.
Hui Yin 0002, Jin Wan
IET Comput. Vis.4
2021 ADSCN: Adaptive dense skip connection network for railway infrastructure displacement monitoring images super-resolution
Hui Yin 0002, Jin Wan, Shi-Jie Zhang
Multim. Tools Appl.2
2020 Pilot Misoperation Control Based on L1 Adaptation and INDI Methods
abstract
The pilot's error is one of the important factors affecting flight safety, which may cause serious consequences. This article mainly discusses the continuous operating deviation and sudden large-scale misoperation during the flight process. For the two types of error models, the controllers are designed to reduce the impact after the misoperation occurs, and simulation verification is performed. Simulation results show that the L1 adaptive controller can effectively suppress the impact of continuous operating deviation on control performance and improve flight comfort; the incremental nonlinear dynamic inverse controller can effectively reduce the impact of large misoperation and restore the original flight state.
Jin Wan, Dan Huang 0002, Lei Song 0005, Shan Fu
ICARCV1
2020 Progressive residual networks for image super-resolution
Jin Wan, Hui Yin 0002, Ai-Xin Chong
Appl. Intell.1
2019 Adaptive convolutional neural network for large change in video object segmentation
abstract
This study tackles the semi‐supervised segmentation task for the objects that have large motion or appearance change in a video sequence, which is very challenging to the existing methods of video object segmentation (VOS). In this study, a novel adaptive approach is presented, named adaptive convolutional neural network for large change VOS, which determines when and how to fine‐tune the convolutional neural network through the motion metric and the appearance metric among consecutive video frames. Additionally, a lightweight optimisation algorithm for the predictive binary mask is introduced which is effective for pixel prediction by eliminating the discrete points cluster. To illustrate the advantages of this approach, experiments have been performed on four VOS datasets, which demonstrate that the proposed method is highly effective and could achieve the state‐of‐the‐art on these datasets.
Hui Yin 0002, Jin Wan
IET Comput. Vis.4
2009 Implicit Interaction: A Modality for Ambient Exercise Monitoring
Jin Wan, Michael J. O'Grady, Gregory M. P. O'Hare
INTERACT (2)1