EDBT 2026 Demo / reviewers in the wild / expert
Fan Zhang 0045
dblp:21/3626-45
· DBLP profile ↗
80ranked-venue papers
12as first author
75since 2021 · last 2026
0000-0002-0343-3499ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 45 · 6 first-author · 45 since 2021Graphics, computer vision, multimedia, augmented reality and games · 22 · 3 first-author · 19 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 2 first-author · 8 since 2021Databases, data management, data science and information retrieval · 9 · 2 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TimeSAF: Towards LLM-Guided Semantic Asynchronous Fusion for Time Series ForecastingabstractDespite the recent success of large language models (LLMs) in time-series forecasting, most existing methods still adopt a Deep Synchronous Fusion strategy, where dense interactions between textual and temporal features are enforced at every layer of the network.This design overlooks the inherent granularity mismatch between modalities and leads to what we term semantic perceptual dissonance: highlevel abstract semantics provided by the LLM become inappropriately entangled with the lowlevel, fine-grained numerical dynamics of time series, making it difficult for semantic priors to effectively guide forecasting.To address this issue, we propose TimeSAF, a new framework based on hierarchical asynchronous fusion.Unlike synchronous approaches, TimeSAF explicitly decouples unimodal feature learning from cross-modal interaction.It introduces an independent cross-modal semantic fusion trunk, which uses learnable queries to aggregate global semantics from the temporal and prompt backbones in a bottom-up manner, and a stage-wise semantic refinement decoder that asynchronously injects these high-level signals back into the temporal backbone.This mechanism provides stable and efficient semantic guidance while avoiding interference with lowlevel temporal dynamics.Extensive experiments on standard long-term forecasting benchmarks show that TimeSAF significantly outperforms state-of-the-art baselines, and further exhibits strong generalization in both few-shot and zero-shot transfer settings. Fan Zhang 0045, Shiming Fan, Hua Wang 0012 |
ACL (1) | 1 |
| 2026 | EEO-TFV: Escape-Explore Optimizer for Web-Scale Time-Series Forecasting and Vision AnalysisabstractTransformer-based foundation models have achieved remarkable progress in tasks such as time-series forecasting and image segmentation. However, they frequently suffer from error accumulation in multivariate long-sequence prediction and exhibit vulnerability to out-of-distribution samples in image-related tasks. Furthermore, these challenges become particularly pronounced in large-scale Web data analysis tasks, which typically involve complex temporal patterns and multimodal features. This complexity substantially increases optimization difficulty, rendering models prone to stagnation at saddle points within high-dimensional parameter spaces. To address these issues, we propose a lightweight Transformer architecture in conjunction with a novel Escape-Explore Optimizer (EEO). The optimizer enhances both exploration and generalization while effectively avoiding sharp minima and saddle-point traps. Experimental results show that, in representative Web data scenarios, our method achieves performance on par with state-of-the-art models across 11 time-series benchmark datasets and the Synapse medical image segmentation task. Moreover, it demonstrates superior generalization and stability, thereby validating its potential as a versatile cross-task foundation model for Web-scale data mining and analysis. Hua Wang 0012, Jinghao Lu, Fan Zhang 0045 |
WWW | 3 |
| 2026 | Time-TK: A Multi-Offset Temporal Interaction Framework Combining Transformer and Kolmogorov-Arnold Networks for Time Series Forecasting
Fan Zhang 0045, Shiming Fan, Hua Wang 0012 |
WWW | 1 |
| 2026 | Dual-channel transformer: Integrating independence and dependence for time series forecasting
Zhigen Huang, Fan Zhang 0045, Yepeng Liu 0003 |
Expert Syst. Appl. | 2 |
| 2026 | Anatomy-Aware Text-Visual Fusion with Dual-Perspective Prompts for Fine-Grained Lumbar Spine Segmentation
Sheng Lian, Jianlong Cai, Dengfeng Pan, Guang-Yong Chen, Fan Zhang 0045, Jialun Pei, Shuo Li 0001 |
Int. J. Comput. Vis. | 6 |
| 2026 | DDformer: Transformer with dynamic variable fusion and dynamic difference attention for multivariate time series long-term forecasting
Hua Wang 0012, Fan Zhang 0045 |
Neurocomputing | 3 |
| 2026 | AlignTime: Interperiodic phase alignment sampling for time-series forecasting
Min Wang 0051, Hua Wang 0012, Fan Zhang 0045 |
Inf. Process. Manag. | 3 |
| 2026 | FSMamba: A dual-expert architecture with fast global attention and local-enhanced state-space mamba for time series forecasting
Shiming Fan, Hua Wang 0012, Fan Zhang 0045 |
Knowl. Based Syst. | 3 |
| 2026 | DTFNet: A dual-modal time-frequency fusion network for non-stationary time series modeling
Fan Zhang 0045, Xiaofeng Zhang 0003, Hua Wang 0012 |
Knowl. Based Syst. | 2 |
| 2026 | Correctformer: A transformer architecture for correcting periodic drift in time-series forecasting
Min Wang 0051, Hua Wang 0012, Fan Zhang 0045 |
Neural Networks | 3 |
| 2026 | VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative ModelsabstractVideo generation has witnessed significant advancements, yet evaluating these models remains a challenge. A comprehensive evaluation benchmark for video generation is indispensable for two reasons: 1) Existing metrics do not fully align with human perceptions; 2) An ideal evaluation system should provide insights to inform future developments of video generation. To this end, we present VBench++, a comprehensive benchmark suite that dissects "video generation quality" into specific, hierarchical, and disentangled dimensions, each with tailored prompts and evaluation methods. VBench++ has several appealing properties: 1) Comprehensive Dimensions: VBench++ comprises 16 dimensions in text-to-video generation (e.g., subject identity inconsistency, motion smoothness, temporal flickering, and spatial relationship, etc). The evaluation metrics with fine-grained levels reveal individual models' strengths and weaknesses. 2) Human Alignment: We also provide a dataset of human preference annotations to validate our benchmarks' alignment with human perception, for each evaluation dimension respectively. 3) Valuable Insights: We look into current models' ability across various evaluation dimensions, and various content types. We also investigate the gaps between video and image generation models. 4) Versatile Benchmarking: VBench++ is designed to evaluate a wide range of video generation tasks, including text-to-video and image-to-video. We introduce a high-quality Image Suite with an adaptive aspect ratio to enable fair evaluations across different image-to-video generation settings. Beyond assessing technical quality, VBench++ evaluates the trustworthiness of video generative models, providing a more holistic view of model performance. 5) Full Open-Sourcing: We fully open-source VBench++, including all prompts, the Image Suite, evaluation methods, generated videos, and human preference annotations. Fan Zhang 0045, Yinan He, Jiashuo Yu, Ziyue Dong, Qianli Ma 0008, Nattapol Chanpaisit, Chenyang Si, Yuming Jiang 0003, Yaohui Wang 0001, Ying-Cong Chen, Limin Wang 0002, Dahua Lin, Yu Qiao 0001, Ziwei Liu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | Dual-gated transformer with local context aggregation for weakly-supervised medical image anomaly detection
Fan Zhang 0045, Demin Liu, Yakun Ju |
Pattern Recognit. | 1 |
| 2026 | Multi-scale temporal correlation multi-dimensional decomposition network for time series analysis
Fan Zhang 0045, Lele Yuan, Hua Wang 0012 |
Pattern Recognit. | 1 |
| 2026 | SAFAformer: Scale-Aware Frequency-Adaptive Guidance for Nighttime Flare RemovalabstractNighttime flare removal is challenging due to the difficulty of acquiring real-world paired data. Existing methods, trained on synthetic pipelines, often struggle to generalize to real-world scenarios. A key limitation of these pipelines is their focus on single-flare scenes, whereas real-world conditions frequently involve more complex cases, such as multi-flare and composite flare scenarios, which are difficult to simulate effectively. This discrepancy significantly hampers model performance in practical applications. Through detailed analysis, we uncover a fundamental characteristic of flare degradation: regardless of whether the scene is synthetic single-flare, real-world single-flare, or multi-flare, the degradation information exhibits a similar distribution across frequency subbands—predominantly concentrated in the low-frequency region, with a minor presence in the high-frequency region. Notably, the severity of the glare effect correlates with an even stronger concentration in the low-frequency domain. This finding suggests that targeted frequency modeling can bridge the gap between synthetic and real-world domains, forming a principled approach to improving generalization. Building on this insight, we propose the Scale-Aware Frequency-Adaptive Guidance Network for Nighttime Flare Removal (SAFAformer), which integrates a Frequency-Adaptive Guidance Module (FAGM) and a Scale-Aware Transformer Block (SATB) to leverage frequency-domain properties during training. Extensive experiments demonstrate that SAFAformer achieves state-of-the-art performance in flare removal compared to existing methods. Our code and pre-trained models are available on GitHub for validation. Fan Zhang 0045, Min Gan, Guang-Yong Chen, C. L. Philip Chen |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Evaluation Agent: Efficient and Promptable Evaluation Framework for Visual Generative ModelsabstractRecent advancements in visual generative models have enabled high-quality image and video generation, opening diverse applications. However, evaluating these models often demands sampling hundreds or thousands of images or videos, making the process computationally expensive, especially for diffusion-based models with inherently slow sampling. Moreover, existing evaluation methods rely on rigid pipelines that overlook specific user needs and provide numerical results without clear explanations. In contrast, humans can quickly form impressions of a model’s capabilities by observing only a few samples. To mimic this, we propose the Evaluation Agent framework, which employs human-like strategies for efficient, dynamic, multi-round evaluations using only a few samples per round, while offering detailed, user-tailored analyses. It offers four key advantages: 1) efficiency, 2) promptable evaluation tailored to diverse user needs, 3) explainability beyond single numerical scores, and 4) scalability across various models and tools. Experiments show that Evaluation Agent reduces evaluation time to 10% of traditional methods while delivering comparable results. The Evaluation Agent framework is fully open-sourced to advance research in visual generative models and their efficient evaluation. Fan Zhang 0045, Shulin Tian, Yu Qiao 0001, Ziwei Liu 0002 |
ACL (1) | 1 |
| 2025 | A Multiscale Edge-Guided Polynomial Approximation Network for Medical Image Segmentation
Fuxian Sui, Hua Wang 0012, Fan Zhang 0045 |
CVM (1) | 3 |
| 2025 | HIFNet: Medical Image Segmentation Network Utilizing Hierarchical Attention Feature Fusion
Hua Wang 0012, Fan Zhang 0045 |
CVM (1) | 3 |
| 2025 | GauUpdate: New Object Insertion in 3D Gaussian Fields with Consistent Global Illuminationabstract3D Gaussian Splatting (3DGS) is a prevailing technique to reconstruct large-scale 3D scenes from multiview images for novel view synthesis, like a room, a block, and even a city. Such large-scale scenes are not static with changes constantly happening in these scenes, like a new building being built or a new decoration being set up. To keep the reconstructed 3D Gaussian fields up-to-date, a naive way is to reconstruct the whole scene after changing, which is extremely costly and inefficient. In this paper, we propose a new method called GauUpdate that allows partially updating an old 3D Gaussian field with new objects from a new 3D Gaussian field. However, simply inserting the new objects leads to inconsistent appearances because the old and new Gaussian fields may have different lighting environments from each other. GauUpdate addresses this problem by applying inverse rendering techniques in the 3DGS to recover both the materials and environmental lights. Based on the materials and lighting, we relight the new objects in the old 3D Gaussian field for consistent global illumination. For an accurate estimation of the materials and lighting, we put additional constraints on the materials and lighting conditions, that these two fields share the same materials but different environment lights, to improve their qualities. We conduct experiments on both synthetic scenes and real-world scenes to evaluate GauUpdate, which demonstrate that GauUpdate achieves realistic object insertion in 3D Gaussian fields with consistent appearances. Chengwei Ren, Fan Zhang 0045, Liangchao Xu, Liang Pan, Ziwei Liu 0002, Wenping Wang 0001, Yuan Liu 0025 |
ICCV | 2 |
| 2025 | ShotBench: Expert-Level Cinematic Understanding in Vision-Language ModelsabstractRecent Vision-Language Models (VLMs) have shown strong performance in general-purpose visual understanding and reasoning, but their ability to comprehend the visual grammar of movie shots remains underexplored and insufficiently evaluated. To bridge this gap, we present \textbf{ShotBench}, a dedicated benchmark for assessing VLMs’ understanding of cinematic language. ShotBench includes 3,049 still images and 500 video clips drawn from more than 200 films, with each sample annotated by trained annotators or curated from professional cinematography resources, resulting in 3,608 high-quality question-answer pairs. We conduct a comprehensive evaluation of over 20 state-of-the-art VLMs across eight core cinematography dimensions. Our analysis reveals clear limitations in fine-grained perception and cinematic reasoning of current VLMs. To improve VLMs capability in cinematography understanding, we construct a large-scale multimodal dataset, named ShotQA, which contains about 70k Question-Answer pairs derived from movie shots.
Besides, we propose ShotVL and train this VLM model with a two-stage training strategy, integrating both supervised fine-tuning and Group Relative Policy Optimization (GRPO). Experimental results demonstrate that our model achieves substantial improvements, surpassing all existing strongest open-source and proprietary models evaluated on ShotBench, establishing a new state-of-the-art performance. Jingwen He, Dian Zheng, Yuhao Dong, Fan Zhang 0045, Yinan He, Weichao Chen 0001, Yu Qiao 0001, Wanli Ouyang, Shengjie Zhao 0001, Ziwei Liu 0002 |
NeurIPS | 6 |
| 2025 | MESA-Net: Multi-Scale Enhanced Spatial Attention Network for medical image segmentation
Demin Liu, Hua Wang 0012, Fan Zhang 0045 |
Comput. Graph. | 5 |
| 2025 | A channel-independent network based on wavelet enhancement for long-term time series forecasting
Zhigen Huang, Fan Zhang 0045, Yepeng Liu 0003 |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | Robust memory-based graph neural networks for noisy and sparse graphs
Linling Jiang, Hua Wang 0012, Fan Zhang 0045 |
Eng. Appl. Artif. Intell. | 4 |
| 2025 | Periodic decomposition and feature enhancement fusion for traffic forecasting
Xiaofei Kong, Hua Wang 0012, Fan Zhang 0045 |
Eng. Appl. Artif. Intell. | 4 |
| 2025 | Probabilistic intervals prediction based on adaptive regression with attention residual connections and covariance constraints
Fan Zhang 0045, Min Wang 0051, Lin Li 0078, Yepeng Liu 0003, Hua Wang 0012 |
Eng. Appl. Artif. Intell. | 1 |
| 2025 | A decoupled network with variable graph convolution and temporal external attention for long-term multivariate time series forecasting
Yepeng Liu 0003, Zhigen Huang, Fan Zhang 0045, Xiaofeng Zhang 0003 |
Expert Syst. Appl. | 3 |
| 2025 | MCNR: Multiscale feature-based latent data component extraction linear regression model
Jinghao Lu, Fan Zhang 0045, Xiaofeng Zhang 0003, Yujuan Sun, Hua Wang 0012 |
Expert Syst. Appl. | 2 |
| 2025 | ADMNet: An adaptive downsampling multi-frequency multi-channel network for long-term time series forecasting
Lele Yuan, Hua Wang 0012, Fan Zhang 0045 |
Expert Syst. Appl. | 3 |
| 2025 | A Lightweight Channel Correlation Invertible Network for Image DenoisingabstractABSTRACT In recent years, deep learning has made significant progress in image denoising. However, the complexity of advanced methods' systems is also increasing, which will increase the calculation cost and hinder the convenient analysis and comparison of methods. Therefore, a lightweight model based on invertible networks is proposed. The invertible network has great advantages in image denoising. It is lightweight, memory‐saving, and information‐lossless in backpropagation. To effectively remove the noise and restore a clean image, the high‐frequency part of the image is resampled and modeled to remove the impact of noise better. The channel context block is proposed to better focus on useful channels and improve the network's perception of useful information in images while ensuring the complexity and computing cost. At the same time, the residual structure with channel correlation modeling is used to extract the features in the convolutional flow, to effectively retain the details and texture of the image, and learn more details of the spatial features of the image, so as to prevent the blur and distortion of the image in the denoising process. The proposed method allows the model to enjoy lower computational complexity on the premise of ensuring performance. Fuxian Sui, Hua Wang 0012, Fan Zhang 0045 |
IET Image Process. | 3 |
| 2025 | SCA-Net: Seasonal Cycle-Aware Model Emphasizing Global and Local Features for Time Series ForecastingabstractRecent advances in transformer architectures have significantly improved performance in time‐series forecasting. Despite the excellent performance of attention mechanisms in global modeling, they often overlook local correlations between seasonal cycles. Drawing on the idea of trend‐seasonality decomposition, we design a seasonal cycle‐aware time‐series forecasting model (SCA‐Net). This model uses a dual‐branch extraction architecture to decompose time series into seasonal and trend components, modeling them based on their intrinsic features, thereby improving prediction accuracy and model interpretability. We propose a method combining global modeling and local feature extraction within seasonal cycles to capture the global view and explore latent features. Specifically, we introduce a frequency‐domain attention mechanism for global modeling and use multiscale dilated convolution to capture local correlations within each cycle, ensuring more comprehensive and accurate feature extraction. For simpler trend components, we apply a regression method and merge the output with the seasonal components via residual connections. To improve seasonal cycle identification, we design an adaptive decomposition method that extracts trend components layer by layer, enabling better decomposition and more useful information extraction. Extensive experiments on eight classic datasets show that SCA‐Net achieves a performance improvement of 12.1% in multivariate forecasting and 15.6% in univariate forecasting compared to the baseline. Min Wang 0051, Hua Wang 0012, Zhen Hua, Fan Zhang 0045 |
Int. J. Intell. Syst. | 4 |
| 2025 | Adaptive decoupled strategy for robust and efficient low-rank matrix decomposition
Min Gan, Fan Zhang 0045, Xiang-Xiang Su, Guang-Yong Chen |
Neurocomputing | 3 |
| 2025 | Traffic prediction based on spatio-temporal feature embedding fusion and gate operation optimization
Xiaotong Geng, Fan Zhang 0045, Hua Wang 0012 |
Neurocomputing | 2 |
| 2025 | Unsupervised bidirectional generative smoothing framework with frequency decomposition and attention enhancement
Jiafu Zeng, Yepeng Liu 0003, Fan Zhang 0045 |
Neurocomputing | 4 |
| 2025 | An Interactive Attention Mechanism Network Integrating the C¹ Activation Function for Time Series ForecastingabstractDecomposing time series into odd and even component sequences is an effective method in time series analysis. However, this data partitioning sometimes leads to the weakening or even disappearance of local features in the original sequence within the odd and even component sequences, thereby reducing the accuracy of the model. To address this issue, we propose a novel neural network with an interactive attention mechanism in this paper. In order to allow the odd and even component sequences obtained after decomposition to capture more global information from the time series and compensate for the lost local features, we introduce odd-even fusion components. Through an interactive attention mechanism, the information of the odd component sequence, even component sequence, and odd-even fusion component sequence complement each other, yielding feature sub-sequences with different temporal relationship weights. Furthermore, we introduce an improved spatial attention submodule with C1 activation functions to better preserve local feature mappings. The segmented polynomial curve C1 activation function PP(x) not only incurs low computational overhead but also effectively alleviates the vanishing gradient problem, resulting in improved feature recognition capabilities. The C1 functional characteristics ensure continuity during backpropagation, guaranteeing stability during the model training process. Experimental results on multiple real-world datasets demonstrate the superior predictive and generalization capabilities of our model for time series forecasting tasks. Lele Yuan, Hua Wang 0012, Fan Zhang 0045 |
IEEE Internet Things J. | 3 |
| 2025 | THATSN: Temporal hierarchical aggregation tree structure network for long-term time-series forecasting
Fan Zhang 0045, Min Wang 0051, Hua Wang 0012 |
Inf. Sci. | 1 |
| 2025 | Minimum variance weighted broad cascade network structure for imbalanced classification
Wuxing Chen, Zhiwen Yu 0002, Kaixiang Yang 0001, Jun Jiang 0003, Fan Zhang 0045, C. L. Philip Chen |
Knowl. Based Syst. | 5 |
| 2025 | CAWformer: A cross variable attention with discrete wavelet denoising for multivariate time series forecasting
Shiming Fan, Hua Wang 0012, Fan Zhang 0045 |
Knowl. Based Syst. | 3 |
| 2025 | Driven by textual knowledge: A Text-View Enhanced Knowledge Transfer Network for lung infection region segmentation
Lexin Fang, Xuemei Li 0001, Yunyang Xu, Fan Zhang 0045, Caiming Zhang 0001 |
Medical Image Anal. | 4 |
| 2025 | Transforming time and space: efficient video super-resolution with hybrid attention and deformable transformers
Linling Jiang, Fan Zhang 0045, Caiming Zhang 0001 |
Vis. Comput. | 3 |
| 2025 | Mask autoencoder for enhanced image reconstruction with position coding offset and combined masking
Yuenan Wang, Hua Wang 0012, Fan Zhang 0045 |
Vis. Comput. | 3 |
| 2024 | VBench: Comprehensive Benchmark Suite for Video Generative ModelsabstractVideo generation has witnessed significant advance-ments, yet evaluating these models remains a challenge. A comprehensive evaluation benchmark for video generation is indispensable for two reasons: 1) Existing metrics do not fully align with human perceptions; 2) An ideal eval-uation system should provide insights to inform future de-velopments of video generation. To this end, we present VBench, a comprehensive benchmark suite that dissects “video generation quality” into specific, hierarchical, and disentangled dimensions, each with tailored prompts and evaluation methods. VBench has three appealing proper-ties: 1) Comprehensive Dimensions: VBench comprises 16 dimensions in video generation (e.g., subject identity in-consistency, motion smoothness, temporal flickering, and spatial relationship, etc.). The evaluation metrics with fine-grained levels reveal individual models' strengths and weaknesses. 2) Human Alignment: We also provide a dataset of human preference annotations to validate our benchmarks' alignment with human perception, for each evaluation dimension respectively. 3) Valuable Insights: We look into current models' ability across various evaluation dimensions, and various content types. We also investi-gate the gaps between video and image generation models. We will open-source VBench, including all prompts, evaluation methods, generated videos, and human preference an-notations, and also include more video generation models in VBench to drive forward the field of video generation. Yinan He, Jiashuo Yu, Fan Zhang 0045, Chenyang Si, Yuming Jiang 0003, Yuanhan Zhang, Tianxing Wu 0002, Qingyang Jin, Nattapol Chanpaisit, Yaohui Wang 0001, Limin Wang 0002, Dahua Lin, Yu Qiao 0001, Ziwei Liu 0002 |
CVPR | 4 |
| 2024 | Skip-Timeformer: Skip-Time Interaction Transformer for Long Sequence Time-Series Forecasting
Hua Wang 0012, Fan Zhang 0045 |
IJCAI | 3 |
| 2024 | Using piecewise polynomial activation functions and relevance attention for long-term time-series predictionabstractBecause of the introduction of self-attention mechanisms, various variants of Transformers have demonstrated significant potential for applications in time-series forecasting in recent years. This mechanism enhances the ability of the model to effectively capture correlations at different positions in a sequence, thereby endowing it with powerful global perception capabilities. However, traditional attention computation methods face challenges in accurately calculating attention scores and lack the flexibility required to model complex relationships. To address these issues, we propose a relevance attention mechanism. This mechanism emphasizes the model’s perception of input variations, enabling it to focus flexibly on different features and effectively perform global feature extraction. Unlike convolutional neural network (CNN) structures, the self-attention mechanism cannot selectively model local features. Therefore, we introduce an adaptive multiscale feature fusion module based on CNNs specifically designed to model local features. This module combines local features with global correlations to capture the overall characteristics of the time series more effectively. Finally, considering the limitations of existing activation functions such as Rectified Linear Unit with a leak (LeakyReLU), which exhibit issues such as C0continuity and neuron death, we adopt a piecewise polynomial function, called C2piecewise polynomial activation (PPA), to provide a more robust and effective activation mechanism for neural network training. Our proposed model, which uses piecewise polynomial activation functions and relevance attention for long-term time-series prediction, outperforms existing models on six benchmark datasets, providing robust support for the further advancement and practical applications in the field of long-term time-series forecasting. Linling Jiang, Fan Zhang 0045 |
IJCNN | 2 |
| 2024 | Computing nodes for plane data points by constructing cubic polynomial with constraints
Hua Wang 0012, Fan Zhang 0045 |
Comput. Aided Geom. Des. | 2 |
| 2024 | CF-DAN: Facial-expression recognition based on cross-fusion dual-attention networkabstractRecently, facial-expression recognition (FER) has primarily focused on images in the wild, including factors such as face occlusion and image blurring, rather than laboratory images. Complex field environments have introduced new challenges to FER. To address these challenges, this study proposes a cross-fusion dual-attention network. The network comprises three parts: (1) a cross-fusion grouped dual-attention mechanism to refine local features and obtain global information; (2) a proposed C2 activation function construction method, which is a piecewise cubic polynomial with three degrees of freedom, requiring less computation with improved flexibility and recognition abilities, which can better address slow running speeds and neuron inactivation problems; and (3) a closed-loop operation between the self-attention distillation process and residual connections to suppress redundant information and improve the generalization ability of the model. The recognition accuracies on the RAF-DB, FERPlus, and AffectNet datasets were 92.78%, 92.02%, and 63.58%, respectively. Experiments show that this model can provide more effective solutions for FER tasks. Fan Zhang 0045, Gongguan Chen, Hua Wang 0012, Caiming Zhang 0001 |
Comput. Vis. Media | 1 |
| 2024 | Frequency-aware robust multidimensional information fusion framework for remote sensing image segmentation
Junyu Fan, Jinjiang Li 0001, Yepeng Liu 0003, Fan Zhang 0045 |
Eng. Appl. Artif. Intell. | 4 |
| 2024 | Combining optical flow and Swin Transformer for Space-Time video super-resolution
Hua Wang 0012, Fan Zhang 0045 |
Eng. Appl. Artif. Intell. | 4 |
| 2024 | Probabilistic interval prediction method based on shape-adaptive quantile regressionabstractAbstract This article introduces customized screening ensemble with shape‐adaptive quantile regression (CseAQR), a novel probabilistic interval forecasting method built upon the quantile regression model. CseAQR utilizes ensemble learning to perform adaptive quantile regression prediction, which can handle the heteroscedasticity feature in time series data by using a weighted adaptive allocation loss function to enhance the adaptability of the basic quantile regression model on the dataset. The model performance predictor is used to select the optimal ensemble learner combination, assign reasonable adaptive weights to it, and obtain a preliminary prediction interval through weighted aggregation. Combining ensemble learners not only improves the accuracy and robustness of prediction intervals but also ensures the commutativity required for conformal prediction. Finally, the conformal prediction method is applied to locally adjust the prediction interval, constructing a more consistently aligned prediction interval with the actual data on a narrower basis. Lin Li 0078, Hua Wang 0012, Yepeng Liu 0003, Fan Zhang 0045 |
Expert Syst. J. Knowl. Eng. | 4 |
| 2024 | CrossWaveNet: A dual-channel network with deep cross-decomposition for Long-term Time Series Forecasting
Siyuan Huang 0006, Yepeng Liu 0003, Fan Zhang 0045, Jinjiang Li 0001, Caiming Zhang 0001 |
Expert Syst. Appl. | 3 |
| 2024 | A stock series prediction model based on variational mode decomposition and dual-channel attention network
Yepeng Liu 0003, Siyuan Huang 0006, Xiaoyi Tian 0002, Fan Zhang 0045, Feng Zhao 0006, Caiming Zhang 0001 |
Expert Syst. Appl. | 4 |
| 2024 | Spatio-temporal Fourier enhanced heterogeneous graph learning for traffic forecasting
Hua Wang 0012, Fan Zhang 0045 |
Expert Syst. Appl. | 3 |
| 2024 | MEAformer: An all-MLP transformer with temporal external attention for long-term time series forecasting
Siyuan Huang 0006, Yepeng Liu 0003, Haoyi Cui, Fan Zhang 0045, Jinjiang Li 0001, Xiaofeng Zhang 0003, Caiming Zhang 0001 |
Inf. Sci. | 4 |
| 2024 | Elevation Information-Guided Multimodal Fusion Robust Framework for Remote Sensing Image SegmentationabstractCurrently, the task of remote sensing image segmentation still faces some challenges, such as variations in illumination, shadows, and occlusions present in remote sensing images. Additionally, there may be similarities and confusions between different types of terrain features. In this paper, we aim to explore how to utilize information exchange between multiple modalities to reduce the impact of interfering factors. To fully exploit the complementary information between different modalities, we establish an information exchange mechanism between optical images (visible light + infrared) features and Digital Surface Model (DSM) features. This allows them to interact and express themselves in a shared feature space, facilitating the acquisition of complementary information from different modalities. Furthermore, through a multimodal fusion encoder and decoder based on Transformer design, the optical features and DSM features are integrated, enabling the learning of high-level semantic representations in different dimensions. Extensive subjective, objective comparative experiments, and ablation experiments are conducted on the ISPRS Vaihingen and Potsdam datasets to evaluate the proposed method. The mIoU on the Vaihingen and Potsdam datasets reached 85.06% and 87.6% respectively, while the OA reached 92.01% and 91.92% respectively. The source code will be available at https://github.com/JunyuFan/MIEFNet. Junyu Fan, Jinjiang Li 0001, Zhen Hua, Fan Zhang 0045, Caiming Zhang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2024 | Fast and highly coupled model for time series forecasting
Hua Wang 0012, Yepeng Liu 0003, Fan Zhang 0045 |
Multim. Tools Appl. | 5 |
| 2024 | Attention Filtering Network Based on Branch Transformer for Change Detection in Remote Sensing ImagesabstractThe emergence of high-resolution (HR) remote sensing imagery showcases the continual advancements in remote sensing technology but also sets higher demands for related tasks in the field, including remote sensing image change detection. Due to their outstanding performance in extracting salient features, convolutional neural networks (CNNs) have played a significant role and become widely utilized in many computer vision tasks. The encoder–decoder structure has confirmed the effectiveness of integrating multilevel feature information, as it allows for the synthesis of both local and global information of features. The exploration of the potential relationships between multilevel features and their efficient integration remains of significant importance. Furthermore, thanks to the advent of the transformer, many modern approaches have seen great improvements in high-level semantic understanding of images. In this article, we propose an attention-filtering network based on a branch transformer for effective change detection in remote sensing images. A hybrid attention fusion module (HAFM) is used to efficiently fuse features of different granularities and perform progressive information filtering on the extracted multilevel features to obtain an effective change feature. We also propose a branch transformer block (BTB) to efficiently aggregate global long-range dependencies and spatial details from the change feature. Extensive comparative experiments conducted on three different HR remote sensing datasets have verified the effectiveness of our method. Yu Shangguan, Jinjiang Li 0001, Yepeng Liu 0003, Fan Zhang 0045, Caiming Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | DUDB: Deep Unfolding-Based Dual-Branch Feature Fusion Network for Pan-Sharpening Remote Sensing ImagesabstractThe proposed method aims to enhance the fusion of high-resolution multispectral (MS) images (HRMS) by extracting spatial and spectral features from panchromatic (PAN) images and MS images. However, existing pan-sharpening methods often suffer from the problem of missing spatial and spectral detail information. To better preserve these details, we introduce a dual-branch feature fusion pan-sharpening network based on deep unfolding. In this network, we utilize the algorithm unfolding iterative module (AUIF-Block) to continuously acquire detailed information from both MS and PAN images for image reconstruction. By leveraging the adaptive channel and spatial feature enhancement module (DEM-Block), the network can adjust spatial and channel features adaptively, leading to more accurate feature extraction and more complete image reconstruction. Finally, the detail-based fusion module (DBFM-Block) is employed to integrate and enrich the content of detailed information extracted from different channels, resulting in improved fusion performance. Experiments were conducted on QuickBird (QB) and WorldView-2 (WV2) datasets. Through qualitative analysis and quantitative comparisons, we demonstrate that this method outperforms existing approaches. Hailin Tao, Jinjiang Li 0001, Zhen Hua, Fan Zhang 0045 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Force-Directed Graph Layouts Revisited: A New Force Based on the T-DistributionabstractIn this article, we propose the t-FDP model, a force-directed placement method based on a novel bounded short-range force (t-force) defined by Student's t-distribution. Our formulation is flexible, exerts limited repulsive forces for nearby nodes and can be adapted separately in its short- and long-range effects. Using such forces in force-directed graph layouts yields better neighborhood preservation than current methods, while maintaining low stress errors. Our efficient implementation using a Fast Fourier Transform is one order of magnitude faster than state-of-the-art methods and two orders faster on the GPU, enabling us to perform parameter tuning by globally and locally adjusting the t-force in real-time for complex graphs. We demonstrate the quality of our approach by numerical evaluation against state-of-the-art approaches and extensions for interactive exploration. Fahai Zhong, Mingliang Xue, Jian Zhang 0070, Fan Zhang 0045, Rui Ban, Oliver Deussen, Yunhai Wang |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2024 | Deep recurrent residual channel attention network for single image super-resolution
Yepeng Liu 0003, Dezhi Yang, Fan Zhang 0045, Qingsong Xie, Caiming Zhang 0001 |
Vis. Comput. | 3 |
| 2023 | FAMC-Net: Frequency Domain Parity Correction Attention and Multi-Scale Dilated Convolution for Time Series ForecastingabstractIn recent years, time series forecasting models based on the Transformer framework have shown great potential, but they suffer from the inherent drawback of high computational complexity and only focus on global modeling. Inspired by trend-seasonality decomposition, we propose a method that combines global modeling with local feature extraction within the seasonal cycle. It aims at capturing the global view while fully exploring the potential features within each seasonal cycle and better expressing the long-term and periodic characteristics of time series. We introduce a frequency domain parity correction block to compute global attention and utilize multi-scale dilated convolution to extract local correlations within each cycle. Additionally, we adopt a dual-branch structure to separately model the seasonality and trend based on their intrinsic features, improving prediction performance and enhancing model interpretability. This model is implemented on a completely single-layer decoder architecture, breaking through the traditional encoder-decoder architecture paradigm and reducing computational complexity to a certain extent. We conducted sufficient experimental validation on eight benchmark datasets, and the results demonstrate its superior performance compared to existing methods in both univariate and multivariate forecasting. Min Wang 0051, Hua Wang 0012, Fan Zhang 0045 |
CIKM | 3 |
| 2023 | An Improved Lightweight YOLOv5 for Remote Sensing Images
Shihao Hou, Linwei Fan, Fan Zhang 0045 |
ICANN (2) | 3 |
| 2023 | Truncated Weighted Nuclear Norm Regularization and Sparsity for Image DenoisingabstractThe attribute of signal sparsity is widely used to sparse representaion. The existing nuclear norm minimization and weighted nuclear norm minimization may achieve a suboptimal in real application with the inaccurate approximation of rank function. This paper presents a novel denoising method that preserves fine structures in the image by imposing L1norm constraints on the wavelet transform coefficients and low rank on high-frequency components of group similar patches. An efficient proximal operator of Truncated Weighted Nuclear Norm (TWNN) is proposed to accurately recover the underlying high-frequency components of low rank patches. By combining a wavelet domain sparse preservation prior with TWNN, the proposed method significantly improves the reconstruction accuracy, leading to a higher PSNR/SSIM and visual quality than state of the art approaches. MingYan Zhang, Feng Zhao 0006, Fan Zhang 0045, Yepeng Liu 0003, Alan C. Evans |
ICIP | 4 |
| 2023 | Time-Series Forecasting Through Contrastive Learning with a Two-Dimensional Self-attention Mechanism
Linling Jiang, Fan Zhang 0045, Caiming Zhang 0001 |
ICONIP (2) | 2 |
| 2023 | Multi-scale Multi-step Dependency Graph Neural Network for Multivariate Time-Series Forecasting
Kaiqiang Zhang, Linling Jiang, Fan Zhang 0045 |
ICONIP (8) | 4 |
| 2023 | Resformer: Combine quadratic linear transformation with efficient sparse Transformer for long-term series forecastingabstractWith the continuous development of deep learning, long sequence time-series forecasting (LSTF) has attracted more and more attention in power consumption prediction, traffic prediction and stock prediction. In recent studies, various improved models of Transformer are favored. While these models have made breakthroughs in reducing the time and space complexity of Transformer, there are still some problems, such as the predictive power of the improved model being slightly lower than that of Transformer. And these models ignore the importance of special values in the time series. To solve these problems, we designed a more concise network named Resformer, which has four significant characteristics: (1) The fully sparse self-attention mechanism achieves O(𝐿𝑙𝑜𝑔𝐿) time complexity. (2) The AMS module is used to process the special values of time series and has comparable performance on sequences dependency alignment. (3) Using quadratic linear transformation, a simple LT module is designed to replace the self-attention mechanism. It effectively reduces redundant information. (4) The DistPooling method based on data distribution is proposed to suppress redundant information and noise. A large number of experiments on real data sets show that the Resformer method is superior to the existing improved model and standard Transformer method. Gongguan Chen, Hua Wang 0012, Yepeng Liu 0003, Fan Zhang 0045 |
Intell. Data Anal. | 5 |
| 2023 | DFNet: Decomposition fusion model for long sequence time-series forecasting
Fan Zhang 0045, Hua Wang 0012 |
Knowl. Based Syst. | 1 |
| 2023 | OM3: An Ordered Multi-level Min-Max Representation for Interactive Progressive Visualization of Time SeriesabstractWe present a novel multi-level representation of time series called OM3 that facilitates efficient interactive progressive visualization of large data stored in a database and supports various interactions such as resizing, panning, zooming, and visual query. Based on our proposed line-segment aggregation, this representation can produce error-free line visualizations that preserve the shape of a time series in windows of arbitrary sizes. To reduce the interaction latency, we develop an incremental tree-based query strategy to support progressive visualizations, allowing a finer control on the accuracy-time tradeoff. We quantitatively compare OM3 with state-of-the-art methods, including a method implemented on a leading time-series database InfluxDB, in two settings with databases residing either in the local area network or on the cloud. Results show that OM^3 maintains a low latency within 300~ms on the web browser and a high data reduction ratio regardless of the data size (ranging from millions to billions of records), achieving around 1,000 times faster than the state-of-the-art methods on the largest dataset experimented with. Yunhai Wang, Xin Chen 0075, Yue Zhao 0033, Fan Zhang 0045, Eugene Wu 0002, Chi-Wing Fu, Xiaohui Yu 0001 |
Proc. ACM Manag. Data | 5 |
| 2023 | Multi-Scale Video Super-Resolution Transformer With Polynomial ApproximationabstractVideo super-resolution techniques aim to obtain high-resolution equivalents of existing low-resolution videos through a series of operations. In recent research, transformers have been increasingly popular because of their remarkable abilities in parallel computing and efficient extraction of space-time sequence features from videos. Moreover, combining self-attention and multi-scale methods has yielded excellent results. However, the combination of the two methods has limitations, current up-sampling methods struggle to match the global modeling capacity of self-attention mechanisms. Therefore, this paper proposes three strategies to combine the two methods. Based on the approximation strategy, we first construct a new bilinear up-sampling method for multi-scale acquisition. Convolution and cross-attention techniques are then used to correct and align features at different scales to prevent large deviations in feature extraction at a specific scale, which can affect subsequent feature extraction. Finally, to effectively solve the common computational complexity,$ C^{0} $continuity, and neuron death problems of existing activation functions, a new method to construct the activation function is proposed. The cubic spline function is used to construct a new activation function approximating tanh. The new activation function is$ C^{2} $continuous, which is piecewise defined by cubic polynomial curves. In this study, better results were achieved on three public video super-resolution test sets: REDS4, Vid4, and Vimeo-90K-T. Experiments demonstrated that the proposed method could provide a new solution for video super-resolution tasks. Fan Zhang 0045, Gongguan Chen, Hua Wang 0012, Jinjiang Li 0001, Caiming Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | CADUI: Cross-Attention-Based Depth Unfolding Iteration Network for Pansharpening Remote Sensing ImagesabstractPansharpening is an important technology for remote sensing imaging systems to obtain high-resolution multispectral (HRMS) images. It mainly obtains high-resolution multi-spectral (HRMS) images with uniform spectral distribution and rich spatial details by fusing low-resolution multi-spectral (LRMS) images and high-spatial-resolution panchromatic (PAN) images. Therefore, how to extract features completely and reconstruct images with high quality is critical to obtain ideal fusion images. In this paper, we propose a new pansharpening method, called the Cross Attention-based Depth Unfolding Iteration Network for Pan-sharpening remote sensing images (CADUI), which achieves the desired fusion effect by iteratively optimizing the deep prior regularization and combining it with a cross-attention mechanism. The network consists of two parts: optimized iterations of deep prior regularization (DEIN-Block) and cross-attention mechanism (CAFM-Block). Among them, DEIN-Block introduces the depth prior as an implicit regularization and improves the adaptability and representation ability of the relevant data of the reconstructed image through iteration. CAFM-Block realizes dual-branch fusion through cross-attention fusion and channel-attention fusion to achieve better fusion results. Simulation experiments and real experiments are carried out on the standard datasets QuikBird (QB) and WorldView-2 (WV2). Through quantitative comparison and qualitative analysis, it is proved that the method is superior to the existing methods. Jinjiang Li 0001, Fan Zhang 0045, Linwei Fan |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Neural-Singular-Hessian: Implicit Neural Representation of Unoriented Point Clouds by Enforcing Singular HessianabstractNeural implicit representation is a promising approach for reconstructing surfaces from point clouds. Existing methods combine various regularization terms, such as the Eikonal and Laplacian energy terms, to enforce the learned neural function to possess the properties of a Signed Distance Function (SDF). However, inferring the actual topology and geometry of the underlying surface from poor-quality unoriented point clouds remains challenging. In accordance with Differential Geometry, the Hessian of the SDF is singular for points within the differential thin-shell space surrounding the surface. Our approach enforces the Hessian of the neural implicit function to have a zero determinant for points near the surface. This technique aligns the gradients for a near-surface point and its on-surface projection point, producing a rough but faithful shape within just a few iterations. By annealing the weight of the singular-Hessian term, our approach ultimately produces a high-fidelity reconstruction result. Extensive experimental results demonstrate that our approach effectively suppresses ghost geometry and recovers details from unoriented point clouds with better expressiveness than existing fitting-based methods. Zixiong Wang, Rui Xu 0016, Fan Zhang 0045, Peng-Shuai Wang, Shuang-Min Chen, Shi-Qing Xin, Wenping Wang 0001, Changhe Tu |
ACM Trans. Graph. | 4 |
| 2022 | Prediction of stock market index based on ISSA-BP neural network
Junhong Guo, Hua Wang 0012, Fan Zhang 0045 |
Expert Syst. Appl. | 4 |
| 2021 | AutoEncoder for Neuroimage
Fan Zhang 0045, Jianxin Zhang 0001, Ahmad Chaddad, Fenghua Guo, Wenbin Zhang 0002, Ji Zhang 0001, Alan C. Evans |
DEXA (2) | 2 |
| 2021 | Estimation of Human Sensitivity for Curvature Gain of Redirected Walking TechnologyabstractThe curvature gain of redirected walking method enables users to explore virtual spaces that are larger than real spaces. The estimation of human sensitivity for curvature gain is important for redirected walking. Herein, we conduct two experiments under different path conditions. By adopting the psychophysical “method of limits” to be a new approach, the first experiment re-estimates the sensitivity for a curved path and further proves that people are more sensitive for right-curved paths than left-curved paths. The second experiment investigates the characteristics of preorder paths on the sensitivity to postorder paths, and finds that the existing of preorder paths significantly increases the sensitivity for the postorder path. Moreover, we find an interesting moderation effect on the direction consistency of preorder and postorder paths: the larger curved preorder path can significantly decrease the sensitivity for postorder path when they have the same direction. If their directions diverge, the effect of preorder path is not significant. Yulong Bian, Chenglei Yang, Fan Zhang 0045, Yanshuai Zhao, Juan Liu 0008, Xiangxu Meng, Linwei Fan |
MobileHCI | 4 |
| 2021 | Low-light image enhancement based on multi-illumination estimation
Xiaomei Feng, Jinjiang Li 0001, Zhen Hua, Fan Zhang 0045 |
Appl. Intell. | 4 |
| 2021 | Detail preserving image denoising with patch-based structure similarity via sparse representation and SVD
Miaowen Shi, Fan Zhang 0045, Suwei Wang, Caiming Zhang 0001, Xuemei Li 0001 |
Comput. Vis. Image Underst. | 2 |
| 2021 | Exploring the effect of virtual reality relaxation environment on white coat hypertension in blood pressure measurement
Haokai Ma, Yulong Bian, Yingbin Wang, Chao Zhou 0012, Wenxiu Geng, Fan Zhang 0045, Juan Liu 0008, Chenglei Yang |
J. Biomed. Informatics | 6 |
| 2021 | Image smoothing based on histogram equalized content-aware patches and direction-constrained sparse gradients
Yepeng Liu 0003, Fan Zhang 0045, Yongxia Zhang, Xuemei Li 0001, Caiming Zhang 0001 |
Signal Process. | 2 |
| 2020 | MR Environments Constructed for a Large Indoor Physical Space
Huan Xing, Chenglei Yang, Xiyu Bao, Sheng Li 0008, Wei Gai, Juan Liu 0008, Yuliang Shi, Gerard de Melo, Fan Zhang 0045, Xiangxu Meng |
CGI | 10 |
| 2020 | Computing knots by quadratic and cubic polynomial curvesabstractA new method is presented to determine parameter values (knot) for data points for curve and surface generation. With four adjacent data points, a quadratic polynomial curve can be determined uniquely if the four points form a convex polygon. When the four data points do not form a convex polygon, a cubic polynomial curve with one degree of freedom is used to interpolate the four points, so that the interpolant has better shape, approximating the polygon formed by the four data points. The degree of freedom is determined by minimizing the cubic coefficient of the cubic polynomial curve. The advantages of the new method are, firstly, the knots computed have quadratic polynomial precision, i.e., if the data points are sampled from a quadratic polynomial curve, and the knots are used to construct a quadratic polynomial, it reproduces the original quadratic curve. Secondly, the new method is affine invariant, which is significant, as most parameterization methods do not have this property. Thirdly, it computes knots using a local method. Experiments show that curves constructed using knots computed by the new method have better interpolation precision than for existing methods. Fan Zhang 0045, Jinjiang Li 0001, Peiqiang Liu |
Comput. Vis. Media | 1 |
| 2019 | Rotbav: A Toolkit for Constructing Mixed Reality Apps with Real-Time Roaming in Large Indoor Physical SpacesabstractThis paper presents a toolkit called Rotbav for easily constructing mixed reality (MR) apps that can be experienced in real time in large indoor physical space via HoloLens. It resolves the problem that existing MR devices, e.g. HoloLens, are unable to scan and model an entire large scene with several rooms at once. We introduce a custom data structure called VorPa, based on the Voronoi diagram, to implement path editing, accelerated rendering and location effectively. Our experiments and applications show that the toolkit is convenient and easy to use for constructing MR apps targeting large indoor physical spaces, in which users can roam in real time. Huan Xing, Xiyu Bao, Fan Zhang 0045, Wei Gai, Juan Liu 0008, Yuliang Shi, Gerard de Melo, Chenglei Yang, Xiangxu Meng |
VR | 3 |
| 2018 | Formula for computing knots with minimum stress and stretching energies
Xuemei Li 0001, Fan Zhang 0045, Guoning Chen, Caiming Zhang 0001 |
Sci. China Inf. Sci. | 2 |
| 2015 | Enlarging Image by Constrained Least Square Approach with Shape Preserving
Fan Zhang 0045, Xin Zhang 0079, Xueying Qin, Caiming Zhang 0001 |
J. Comput. Sci. Technol. | 1 |