Xinfeng Zhang 0001

dblp:12/7627-1 · also Xin-Feng Zhang 0001 · DBLP profile ↗
← Back
16ranked-venue papers in the field
0as first author
10since 2021 · last 2025
0000-0002-7517-3868ORCID · conflict

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 9Information Retrieval & Web Search · 4Knowledge Engineering, Semantic Web & Information Systems · 2Other / Interdisciplinary · 1
YearPublicationVenuePosition
2025 Mixture-of-KAN for Multivariate Time Series Forecasting
abstract
Multivariate time series forecasting is a crucial task that predicts the future states based on historical inputs. Although current deep learning-based methods have made significant advancements, they still face the criticism of lacking interpretability. The rise of the Kolmogorov-Arnold Network (KAN) provides a new perspective to implement an efficient and interpretable deep learning-based method for forecasting time series. However, we find there are two main challenges in the application of KAN in time series forecasting: how to select the appropriate one from various KAN variants and how to train the deep KAN-based network. To this end, we propose the multi-layer mixture-of-KAN network, which achieves excellent performance while retaining KAN's ability to be transformed into a combination of symbolic functions. The core module is the mixture-of-KAN layer, which uses a mixture-of-experts structure to assign variables to best-matched KAN experts. Then, we analyze the shortcomings of parameter initialization in the original KAN and provide an effective initialization method to alleviate training instability. Extensive experimental results demonstrate that our proposed method is effective in multivariate time series forecasting. Codes are released in https://github.com/2448845600/EasyTSF.
Zhenduo Zhang, Xinfeng Zhang 0001, Yiling Wu, Zhe Wu 0006
CIKM3
2025 MoRLACS: A Monocular RGBD-based Locomotion Approach for CAVE Systems
abstract
Navigation within Cave Automatic Virtual Environment (CAVE) systems often faces challenges due to limited physical space and the necessity for seamless user interaction. Traditional solutions typically rely on multi-view tracking systems or constrained locomotion techniques, which can interrupt immersion and hinder usability. In this paper, we introduce MoRLACS, a novel locomotion approach for CAVE systems that leverages a single RGBD camera. This hybrid framework integrates small-scale physical walking with controller-based large-scale exploration through a tailored guidance method. By accurately tracking the user's head position in the real world and synchronizing it with the virtual camera, MoRLACS enables natural walking within confined CAVE spaces and supports extended interaction in larger virtual environments. Preliminary user experiments demonstrate the approach's effectiveness, revealing improvements in usability and a heightened sense of presence. These findings underscore the potential of MoRLACS to enrich user experiences in immersive CAVE settings and offer valuable design insights for integrating 3D sensor data into multimedia interaction frameworks.
Haopeng Lu, Qian Yin 0002, Li Song 0001, Xinfeng Zhang 0001, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001
ICMR5
2024 Dynamic point cloud compression with spatio-temporal transformer-style modeling
abstract
The essence of dynamic point cloud compression lies in the effective modeling of temporal context information, which poses significant challenges owing to the unstructured and sparse characteristics of point clouds. Existing dynamic compression methods exhibit a limited capacity to capture and leverage inter-frame information. Consequently, in this paper, we propose a Dynamic Point Cloud Compression framework with Spatio-Temporal Transformer-style Modeling (DPCC-STTM) to compress point cloud sequences within a latent space. To effectively extract and fully utilize temporal context, we introduce a spatio-temporal transformer-style modeling module, which performs effective modeling of the rich temporal information based on the correlation of temporal content. Furthermore, we introduce a multi-scale temporal processing module that captures temporal correlations across short and long ranges of multi-frame point clouds. This module also fuses modeled temporal information to enhance the prediction accuracy of potential features for the current frame. Extensive experiments demonstrate the superiority of our proposed framework, validated through both objective evaluation and subjective perception.
Xinfeng Zhang 0001, Xiaoqi Ma, Yingzhan Xu, Kai Zhang 0007, Li Zhang 0006
DCC2
2024 Compact Visual Data Representation for Multimedia Search and Analytics
abstract
With the exponential growth of multimedia in various forms, the volume of acquired visual data has dramatically increased while their value intensity remains relatively low. This presents significant challenges in multimedia search and analytics. In this tutorial, we aim to introduce recent advances of compact visual data representation techniques that enable efficient, flexible, and reliable multimedia search and analytics. We will explore the shift from traditional visual information representation techniques, such as video coding, to biologically inspired information processing paradigms, like digital retina based coding and representation. We will also discuss the representation of point cloud data and Artificial Intelligence Generated Content (AIGC) data, which are becoming increasingly popular in modern machine vision technologies. Additionally, we will discuss the recent advances in quality assessment technologies for multimedia signals under various novel and challenging scenarios. Finally, we will introduce the recent standardization activities in media coding including Video Coding for Machine (VCM). This tutorial aims to stimulate fruitful discussions, encourage innovative research, and drive advancements in the field of semantic and visual communication, multimedia search, analytics, computing as well as generative AI.
Shiqi Wang 0001, Xinfeng Zhang 0001
ICMR2
2024 FATO: Frequency Attention Transformer for Omnidirectional Image Super-Resolution
abstract
Benefiting from the 360 • field of view (FoV) of the omnidirectional images (ODIs), users could enjoy an immersive experience with head-mounted devices or computers.High-resolution ODIs can provide pleasing visual experience and boost the performance of related visual tasks.Therefore, Super-resolution (SR) is an essential technique during the application of ODIs.However, traditional SR methods fail to enhance the most widely utilized equirectangular projection (ERP) format ODIs due to projection distortions.Existing ODI-SR methods take the latitude-related position information as a prior, but lack the adaptation to the ERP content distribution characteristics.To address this issue, we propose a novel Frequency Attention Transformer ODI-SR (FATO) network focusing on highfrequency details of ODIs.In particular, we transform an ODI into fine-grained patches in the frequency domain through Discrete Cosine Transform (DCT).After that, we design a frequency selfattention mechanism to capture the relationship between different frequency patches.Subsequently, we introduce a frequency loss function to further constrain the network.Extensive experimental results demonstrate that the proposed FATO achieves superior performance over state-of-the-art methods on ODIs.
Hongyu An, Xinfeng Zhang 0001, Shijie Zhao 0001, Li Zhang 0006
MMAsia2
2023 Adaptive Graph Neural Diffusion for Traffic Demand Forecasting
abstract
This paper studies the problem of spatial-temporal modeling for traffic demand forecasting. In practice, the temporal-spatial dependencies are complex. Conventional methods using graph convolutional networks and gated recurrent units cannot fully explore the patterns of demand evolution. Therefore, we propose Adaptive Graph Neural Diffusion (AGND) for spatial-temporal graph modeling. Specifically, complex spatial relations are modeled with a diffusion process by the graph neural diffusion. The spatial attention mechanism and a data-driven semantic adjacency matrix are used to describe the diffusivity function in the graph neural diffusion, which provides both local and global spatial information. Long-term temporal dependencies are modeled by the temporal attention mechanism. The proposed method is applied to two real-world datasets, and the results show that the proposed method outperforms state-of-the-art methods.
Yiling Wu, Xinfeng Zhang 0001, Yaowei Wang 0001
CIKM2
2022 Parametric Non-local In-loop Filter for Future Video Coding
abstract
In-loop filter has been comprehensively explored during the development of video coding standards to suppress compression artifacts. However, the existing in-loop filters in Versatile Video Coding (VVC) mainly take advantage of the image local similarity. Although some non-local based in-loop filters can make up for this short-coming, the unsupervised parameter selection scheme, which is widely used by non-local filters, limits the content adaptability. Given this, we propose a parametric non-local in-loop filter (PNLF) that fully considers the non-local characteristics and trains the filter coefficients based on the video content. In the filtering process, the reference samples based on the non-local similarity are first derived for each to-be-filtered sample. Then to-be-filtered samples are grouped into specific classes based on multiple features. For each class, filter coefficients are online trained in the encoder and transmitted to the decoder. Finally, the filtering process is conducted using the online-selected coefficients. Simulation results reveal that the proposed approach achieves 0.70%, 1.43%, and 2.09% bit-rate savings on average compared to VTM-11.0 under All Intra (AI), Random Access (RA), and Low-Delay B (LDB) configurations, respectively. The sequences used in the experiment include Class AI, A2, B, C, D, E, F, and SCC. Compared to the non-local structure-based filter (NLSF) [1], our proposed PNLF with fast block matching scheme [2] applied on B-frames and P-frames can achieve better performance gain with lower software and hardware complexity under RA and LDB configurations.
Xuewei Meng, Chuanmin Jia, Xinfeng Zhang 0001, Meng Lei, Shanshe Wang, Lin Li 0062, Siwei Ma 0001
DCC3
2021 Optimized Adaptive Loop Filter in Versatile Video Coding
abstract
In the Versatile Video Coding (VVC) standard, adaptive loop filter (ALF), including Geometry transformation-based Adaptive Loop Filter (GALF) and Cross Component Adaptive Loop Filter (CCALF), plays an essential role in reducing compression artifacts. However, it also has high coding complexity and requires many picture buffer accesses in the encoder that will increase external memory access and is unfriendly to the software and hardware design. Therefore, we propose an optimized ALF framework, including the parallel design of GALF and CCALF, the adaptive parameter decision of GALF, and one-pass CCALF scheme by effectively estimating the CCALF filtering distortion without conducting filter operation. Compared to VTM-8.0, the proposed method can reduce the picture buffer access from 152 to 1 and achieve roughly 25% time-savings of the ALF module with negligible coding performance change under RA configuration. Some of the proposed methods have been adopted in the VVC reference software.
Xuewei Meng, Jiaqi Zhang 0007, Chuanmin Jia, Xinfeng Zhang 0001, Shanshe Wang, Siwei Ma 0001
DCC4
2021 Flow-Grounded Dynamic Texture Synthesis for Video Compression
abstract
The basic ingredients of modern video coding standards are block-based prediction and transforms. However, when dealing with video contents containing dynamic textures (DT), the existing prediction schemes usually failed due to temporal variability and randomness of DT, which results in more bit cost on residual coding compared with other contents. In view of this point, a novel video compression scheme for DT is proposed in this work. In particular, wavelet-based analysis on motion characteristics of DT is firstly presented and based on the analysis, we introduce a flow-grounded texture synthesis method for video compression. Instead of conventional inter prediction, synthesized DT contents are used for reconstruction at the decoder. The proposed scheme has been fully integrated into the test model of Versatile Video Coding standard, VTM-10.0, for validation and a subjective test has also been carried out. Experimental results show that bitrate savings can be achieved by 40% on average at comparable visual quality.
Suhong Wang, Xinfeng Zhang 0001, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001
DCC2
2021 No-reference image quality assessment for contrast-changed images via a semi-supervised robust PCA model
Jingchao Cao, Ran Wang 0001, Yuheng Jia, Xinfeng Zhang 0001, Shiqi Wang 0001, Sam Kwong
Inf. Sci.4
2020 Gradient-Based Early Termination of CU Partition in VVC Intra Coding
abstract
The quaternary tree with nested binary and ternary tree structure is an efficient coding unit (CU) partitioning method adopted in Versatile Video Coding (VVC). Compared with quaternary tree only in HEVC, its flexible block sizes improve the coding performance significantly at the cost of computation load increase due to the recursive and nested searching for the best CU structure. In this paper, an early termination algorithm is proposed to skip unnecessary searches in CU size decision. Based on directional gradients, pre-determine the likelihood of binary partition or ternary partition in horizontal or vertical direction for current block, thus the impertinent partition pattern is skipped. The experimental results show that the proposed method is able to save the encoding time up to 51% with about 1.2% BD-rate degradation compared with VVC software reference VTM5.0.
Tao Zhang 0013, Chenchen Gu, Xinfeng Zhang 0001, Siwei Ma 0001
DCC4
2019 Perceptual Video Coding Based on Visual Saliency Modulated Just Noticeable Distortion
abstract
To reduce the perceptual redundancy in the video coding process, human visual system (HVS)-based visual attention and visual sensitivity can be utilized due to their intrinsic natures. Just Noticeable Distortion (JND) is one of widely used models to simulate human visual sensitivity, while visual saliency map has been popular for years in image processing to describe the visual attention feature, which has been proved by the ability to enhance the visual sensitivity effect. In this paper, we proposed a perceptual video coding (PVC) scheme with visual saliency modulated JND model to suppress the DCT coefficient without resulting in noteworthy subjective quality degradation. The experimental results show that the PVC scheme with the proposed VS-JND model can save bit rates up to 35.58% in high bit rates case with the similar subjective quality compared with that of HEVC software reference code HM 16.12.
Ruiqin Xiong, Xinfeng Zhang 0001, Shanshe Wang, Siwei Ma 0001
DCC3
2018 Locally Refined Motion Compensation for Future Video Coding
abstract
Motion compensation plays a key role in high efficiency video coding. The popular video compression standards, such as H.264/AVC and HEVC, adopt block based motion compensation technique due to its high compression efficiency and relatively low computational complexity. However, block based motion compensation may not be in accordance with the actual object boundary, potentially leading to low prediction accuracy especially in the high-texture areas. In this paper, we propose a locally refined motion compensation method to address this issue. In particular, the image segmentation is applied on the prediction block indicated by a motion vector rather than the original block to avoid explicit signaling. Furthermore, the local content is analyzed to select one segmented region and subsequently the prediction of this region is generated based on the local motion filed. Experimental results show that the proposed algorithm can achieve 0.8%, 1.1% and 1.7% bitrate savings for Random Access, Lowdelay-B and Lowdelay-P configurations respectively without introducing noticeable computational complexity.
Zhao Wang 0004, Shiqi Wang 0001, Xinfeng Zhang 0001, Shanshe Wang, Siwei Ma 0001
DCC3
2017 Band-Wise Adaptive Sparsity Regularization for Quantized Compressed Sensing Exploiting Nonlocal Similarity
abstract
The theory of compressive sensing (CS) has attracted considerable research interests from signal and image processing communities. And in practice, because of the considerations of data storage and transmission, scalar quantization is necessary to be implemented on the CS measurements. In this paper, we propose an adaptive bandwise sparsity regularization to handle the recovery problem of quantized compressive sensing. The sparsity regularization constraints every patch by using bandwise distribution model in transform domain. In addition, we bring in the quantization cost function to quantify the influence of measurement quantization. Experimental results demonstrate that our CS recovery strategy achieves significant performance improvements over the current state-of-the-art schemes with both unquantized measurements and quantized measurements.
Ruiqin Xiong, Xinfeng Zhang 0001, Siwei Ma 0001
DCC3
2017 Visual attention analysis and prediction on human faces
Xiongkuo Min, Guangtao Zhai, Ke Gu 0001, Jing Liu 0002, Shiqi Wang 0001, Xinfeng Zhang 0001, Xiaokang Yang 0001
Inf. Sci.6
2016 From Visual Search to Video Compression: A Compact Representation Framework for Video Feature Descriptors
abstract
Visual feature descriptors have been successfully deployed in a wide range of applications, e.g. visual retrieval and analysis. To transmit these descriptors over bandwidth-limited networks, a high efficiency feature coding technique is highly desired to maximize compression capability and achieve compact feature representations. In this paper, a hybrid visual feature descriptor compression framework is presented and implemented in the encoding and decoding loops of texture videos. In particular, the multiple-hypothesis prediction is employed to effectively remove redundancies originated not only from spatial and temporal similarities, but also from reconstructed video frames. As the ultimate purpose of the transmitted descriptors is retrieval, the rate-accuracy optimization (RAO) technique is proposed to obtain the best tradeoff between the rate and retrieval performance. Such paradigm enables the conventional video stream to achieve high efficient retrieval/analysis with very low bitrate consumption. Moreover, we also demonstrate that texture video compression can also benefit from the additional information provided by the transmitted descriptors, leading to significantly improvement of coding efficiency on top of the high efficiency video coding (HEVC) standard. Extensive simulations have shown that the proposed method can offer significant bitrate reduction in representing both the descriptors and texture video frames, and meanwhile providing desirable retrieval performance.
Xiang Zhang 0004, Siwei Ma 0001, Shiqi Wang 0001, Shanshe Wang, Xinfeng Zhang 0001, Wen Gao 0001
DCC5