VLDB 2026 Research / reviewers in the wild / expert
Yiming Wang 0008
dblp:71/3182-8
· DBLP profile ↗
21ranked-venue papers
9as first author
21since 2021 · last 2026
0000-0001-5683-909XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 5 first-author · 16 since 2021Artificial intelligence and machine learning · 5 · 4 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Structural enhanced robust discriminative least squares regression for image classificationabstractDiscriminative Least Squares Regression (DLSR) improves classification performance by enhancing inter-class separability through the ϵ-dragging method. However, it overlooks the impact of intra-class structure on the model’s discriminative ability and is sensitive to noise interference. While many DLSR variants have been proposed to preserve intra-class structure modeling and noise robustness, they often fail to achieve a balanced trade-off between these goals and model complexity. To solve these challenges, this paper develops a new model, called Structural Enhanced Robust Discriminative Least Squares Regression (SERDLSR). Specifically, SERDLSR introduces a projection distance regularization to promote the compactness of intra-class samples in the projected space. Additionally, this model integrates local sparsity and local low-rank constraints, the former improves intra-class consistency, while the latter preserves the underlying low-dimensional structure of intra-class samples. By integrating the above three types of constraints, SERDLSR effectively strengthens the intra-class structure of projected feature representations while enhancing robustness against the inherent high-dimensional noise in the dataset. Extensive experiments on face recognition (AR, ORL, CMU PIE, FERET, and Georgia Tech datasets), biometric recognition (PolyU Palmprint dataset), and object recognition (COIL-20 dataset) demonstrate that SERDLSR achieves superior classification performance. Zhangjing Yang, Yiming Wang 0008, Pu Huang 0004, Fanlong Zhang |
Expert Syst. Appl. | 3 |
| 2026 | Dual-Scale Transformer with Variable Bitrate Synchronization for Neural Video CompressionabstractNeural video compression (NVC) has emerged as a promising paradigm for improving rate-distortion performance. However, existing neural video codecs predominantly rely on convolutional neural networks (CNNs) with limited local receptive fields to generate the latent representations, often neglecting global–local spatial correlations. This leads to suboptimal feature modeling and redundancy in the latent space. To address this limitation, we propose a novel Dual-Scale Transformer (DST) block specifically tailored for NVC, which effectively enhances coding efficiency. The DST block incorporates a Global–Local (Shifted) Window-based Self-Attention (GL(S)WSA) mechanism to jointly capture global structure information and local texture details. Moreover, we design a Cross-Gated Feed-Forward Network (CGFFN) to adaptively modulate complementary components, producing more compact and expressive latent representations. Furthermore, to overcome the drawbacks of traditional asynchronous training and further boost rate-distortion performance, we introduce a Variable Bitrate Synchronization (VBRS) strategy that leverages multi-GPU parallel training, with each GPU dedicated to a specific bitrate and synchronized via gradient backpropagation for joint optimization. Experimental results demonstrate that our proposed method achieves the higher coding performance compared to the previous state-of-the-art (SOTA) methods and significantly outperforms H.266/VVC (VTM-13.2) under various low delay B (LDB) coding configurations. Yiming Wang 0008, Yaojun Wu 0001, Zhaobin Zhang, Qian Huang 0008, Bin Tang 0002, Zhangjing Yang, Kai Zhang 0007, Li Zhang 0006 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2025 | Learned Video Compression With Refined Adaptive Flow Pyramid And Coordinate-Aware AttentionabstractIn video compression, motion estimation and motion compensation are critical for achieving efficient encoding. Although the commonly used SpyNet and bilinear interpolation have contributed in improving the compression efficiency, they still have limitations. SpyNet often loses details and fails to fully utilize the feature extraction capabilities of deep networks. Furthermore, bilinear interpolation inherently attenuates high-frequency information, leading to frame blurring and distortion. In this paper, we propose a novel video compression algorithm. To overcome the limitations of SpyNet, we propose a refined adaptive flow pyramid network. This network uses a multi-scale feature pyramid to capture more details. Moreover, we use an iterative cost volume refinement engine that improves the feature representation of the network and iteratively improve the accuracy of motion estimation. In addition, to overcome the limitation of bilinear interpolation, we propose a coordinate-aware attention module, which captures more high-frequency information to improve the accuracy of motion compensation. Experimental results show that our method outperforms VTM-13.2 (LDP) in terms of PSNR. Qian Huang 0008, Xin Li 0090, Yiming Wang 0008 |
ICASSP | 4 |
| 2025 | Learned Video Compression with Spatial Correlation Priors and Hierarchical Temporal AttentionabstractAccurately predicting the probability distribution of quantized latent representations is a critical challenge for entropy models in learned video compression (LVC). Existing mainstream LVC methods typically adopt ready-made entropy models based on image compression, which fail to fully exploit the information of spatial-temporal correlation. To address this issue, we propose a spatial correlation priors and hierarchical temporal attention (SCP-HTA) model, which exploits the spatial correlation information from the current video frames and refine the temporal information from the context. First, we extract the spatial correlation of the current frame to guide the generation of masks, enabling the frame to leverage more information during encoding and decoding process. Additionally, to obtain more accurate temporal information, we introduce a hierarchical temporal attention module at channel level when we generate the context. Experimental results demonstrate that the proposed SCP-HTA model achieve 15.76% bitrate saving in PSNR and 61.82% in MS-SSIM on average across all test datasets when compared with VTM-13.2 (LDP). Qian Huang 0008, Wenchao Shan, Zaipeng Xie, Yiming Wang 0008 |
ICIP | 5 |
| 2025 | Neighbor-Aware Feature-Driven Motion Compensation for Learned Video CompressionabstractLearned video compression (LVC) methods typically align spatial-temporal transformation features with optical flow to perform motion compensation. However, existing LVC methods typically rely on a single reference feature to provide local detail features. This ignores global structural information, resulting in limited capabilities when dealing with fast motion or occlusion scenarios. In addition, with the encoding of P-frames, errors accumulate, leading to a degradation in reconstruction quality. To address these issues, we propose the Neighbor-Aware Feature-Driven Motion Compensation (NAFD-MC) that utilizes the spatial-temporal correlations of the neighboring features to explore the global structural information and the local detail information. Furthermore, we introduce the Synergy Filtering Module (SFM) to enhance inter-frame consistency and alleviate the error accumulation. Experimental results demonstrate that our method outperforms the H.266/VVC reference software VTM-13.2 in public benchmark datasets. Hao Lu 0013, Qian Huang 0008, Ziyang Yin, Zaipeng Xie, Yiming Wang 0008 |
ICIP | 5 |
| 2025 | Efficient Local-Global Collaboration Transcoding for JPEG AIabstractIn the past decade, learning-based image compression has made significant advancements, with the Joint Photographic Experts Group (JPEG) working towards the launch of the first neural network-based image coding standard, JPEG AI. However, traditional codecs like JPEG and HEVC intra coding remain the most widely used. A key challenge arises when attempting to compress images pre-encoded with these traditional codecs using JPEG AI, as such images often carry artifacts that JPEG AI is not fully optimized to handle, leading to significant loss in performance. This paper addresses this issue by proposing a novel transcoding framework designed to optimize pre-encoded images for JPEG AI, bridging the gap between traditional and learning-based compression methods. The framework employs two branches for processing luminance and chrominance components, with a local-global modulation module (LGMM) for luminance and a dynamic fusion module (DFM) for chrominance. Experimental results demonstrate that the proposed scheme achieves up to 22.36% rate savings on images pre-encoded with HEVC intra and up to 22.29% rate savings on images pre-encoded with JPEG using JPEG AI reference software, highlighting its effectiveness in enhancing the performance of JPEG AI for real-world applications. Yiming Wang 0008, Zhaobin Zhang, Yaojun Wu 0001, Qian Huang 0008, Bin Tang 0002, Kai Zhang 0007, Li Zhang 0136 |
ICME | 1 |
| 2025 | Neural Video Compression with In-Loop Contextual Filtering and Out-of-Loop Reconstruction EnhancementabstractThis paper explores the application of enhancement filtering techniques in neural video compression. Specifically, we categorize these techniques into in-loop contextual filtering and out-of-loop reconstruction enhancement based on whether the enhanced representation affects the subsequent coding loop. In-loop contextual filtering refines the temporal context by mitigating error propagation during frame-by-frame encoding. However, its influence on both the current and subsequent frames poses challenges in adaptively applying filtering throughout the sequence. To address this, we introduce an adaptive coding decision strategy that dynamically determines filtering application during encoding. Additionally, out-of-loop reconstruction enhancement is employed to refine the quality of reconstructed frames, providing a simple yet effective improvement in coding efficiency. To the best of our knowledge, this work presents the first systematic study of enhancement filtering in the context of conditional-based neural video compression. Extensive experiments demonstrate a 7.71% reduction in bit rate compared to state-of-the-art neural video codecs, validating the effectiveness of the proposed approach. Yaojun Wu 0001, Chaoyi Lin, Yiming Wang 0008, Semih Esenlik, Zhaobin Zhang, Kai Zhang 0007, Li Zhang 0006 |
ACM Multimedia | 3 |
| 2025 | STFE-VC: Spatio-temporal feature enhancement for learned video compression
Yiming Wang 0008, Qian Huang 0008, Bin Tang 0002, Xin Li 0090, Xing Li 0005 |
Expert Syst. Appl. | 1 |
| 2025 | Multiscale motion-aware and spatial-temporal-channel contextual coding network for learned video compression
Yiming Wang 0008, Qian Huang 0008, Bin Tang 0002, Xin Li 0090, Xing Li 0005 |
Knowl. Based Syst. | 1 |
| 2024 | Learned Video Compression with Spatial-Temporal OptimizationabstractPrevious optical flow based video compression is gradually replaced by unsupervised deformable convolution (DCN) based method. This is mainly due to the fact that the motion vector (MV) estimated by the existing optical flow network is not accurate and may introduce extra artifacts. However, DCN based method is difficult for training owing to the lack of explicit guidance in the feature space. In this work, we propose a learned video compression with spatial-temporal optimization. Specifically, we first propose the spatial-temporal motion refinement module to improve the accuracy of MV estimated by the optical flow network for prediction. Then, we propose the In-loop filter module to remove compression artifacts and improve the reconstructed frame quality. Finally, comprehensive experimental results demonstrate our proposed method outperforms the recent learned methods on three benchmark datasets. Moreover, our method also beats the H.266/VVC in terms of MS-SSIM metrics. Yiming Wang 0008, Qian Huang 0008, Bin Tang 0002, Wenchao Shan |
ICASSP | 1 |
| 2024 | Temporal context video compression with flow-guided feature prediction
Yiming Wang 0008, Qian Huang 0008, Bin Tang 0002, Huashan Sun, Zhuang Miao |
Expert Syst. Appl. | 1 |
| 2024 | A survey of feature matching methodsabstractAbstract Feature matching plays a crucial role in computer vision, with applications in visual localization, simultaneous localization and mapping (SLAM), image stitching, and more. It establishes correspondences between sets of feature points from multiple images, enabling various tasks. Over the years, feature matching has witnessed significant development, with an increasing number of methods being applied. However, different methods exhibit different degrees of applicability in different scenarios and requirements due to their different rationales. To cope with these issues, a comprehensive analysis and comparison of matching methods are essential. Existing reviews often lack coverage of deep learning models and focus more on feature detection and description, neglecting the matching process. This survey investigates feature detection, description, and matching techniques within the feature‐based image‐matching pipeline. Representative methods, their mechanisms, and application scenarios are also briefly introduced. In addition, comprehensive evaluations of classical and state‐of‐the‐art methods are conducted through extensive experiments on representative datasets. Particularly, matching‐based applications are compared to fully demonstrate the advantages of the methods. Lastly, this survey highlights current problems and development directions in matching methods, serving as a reference for researchers in the field. Qian Huang 0008, Yiming Wang 0008, Huashan Sun |
IET Image Process. | 3 |
| 2024 | DMCVS: Decomposed motion compensation-based video stabilizationabstractAbstract With the popularity of handheld devices, video stabilization is becoming increasingly important. In previous studies, many methods have been proposed to stabilize shaky videos. However, these methods fail to balance between image content integrity and stability. Some methods sacrifice image content for better stability. Other methods ignore the subtle jitters, which leads to poor stability. This work innovatively proposes a video stabilization method based on decomposed motion compensation. First, a grid‐based motion statistics method is adopted for motion estimation, which obtains more accurate motion vectors according to matched likelihood estimates. Then, the motion compensation is inherently decomposed into two parts: linear motion compensation and auxiliary motion compensation. Linear motion compensation removes complex jitter by constructing linear path constraints to obtain a more stable camera path. Auxiliary motion compensation uses a moving average filter to remove the high‐frequency jitter as a supplement and preserve more image content. The two components are combined with individual weights to derive the final transform matrix and warp the original frames. Experimental results show that our method outperforms the previous methods on NUS and DeepStab datasets qualitatively and quantitatively. Qian Huang 0008, Jiwen Liu, Chuanxu Jiang, Yiming Wang 0008 |
IET Image Process. | 4 |
| 2024 | Fusing angular features for skeleton-based action recognition using multi-stream graph convolution networkabstractAbstract Distinguishing similar actions has been a challenging challenge in skeleton‐based action recognition. Since the joint coordinates in these actions are similar, it is difficult to accomplish the recognition task using traditional joint features. To address this issue, the use of angle features to capture subtle nuances in various body parts, along with a critical angle enhancement module that assigns weights to different angle feature representations for a given action are proposed, highlighting the critical angle feature representation. The approach is evaluated using a three‐stream ensemble method on three large action recognition datasets, NTU‐RGB+D, NTU‐RGB+D 120, and Kinetics‐400. The experimental results demonstrate that incorporating angular information can effectively complement joint and skeletal features, leading to improved recognition of similar actions and enhanced model performance and robustness. Qian Huang 0008, Mingzhou Shang, Yiming Wang 0008 |
IET Image Process. | 4 |
| 2024 | Ship detection based on YOLO algorithm for visible imagesabstractAbstract Ship detection is a crucial task for waterway surveillance and channel optimization, especially in close proximity to the shore. However, detecting ship in visible image‐based detection remains a challenge due to the limited nature of visible image datasets. To address this issue, the Inland Ships Data Set (ISDS) is constructed to facilitate research on ship identification. On the other hand, most detection methods struggle to accurately identify ships that are small in size. Therefore, a visible image‐based ship detection model is proposed that employs a multi‐scale weighted feature fusion structure with the YOLOv4 detection model to improve the efficacy of small ship detection. Specifically, the YOLOv4 model is improved through fusing multi‐scale feature, redesigning priori frame, and enhancing loss function. The model, named YOLOv4‐MSW (i.e. YOLOv4 based on Multi‐Scale Weighted feature fusion), exhibits improved performance on ship detection in experiments conducted on the ISDS dataset, outperforming the original YOLOv4 model by improving the average precision (AP) by 4.87% and the recall rate by 10.03%. Meanwhile, the model achieve better detection accuracy and improve the average precision rate by at least 0.86% compared to existing learned object detection methods. The code related to this work are released at https://github.com/Sunhuashan/YOLOv4‐MSW . The whole dataset is available at https://drive.google.com/drive/folders/1fzJ2fcqiko6lFwqIEGghMceoQgv‐8jBy . Qian Huang 0008, Huashan Sun, Yiming Wang 0008 |
IET Image Process. | 3 |
| 2024 | Adaptive video stabilization based on feature point detection and full-reference stability assessment
Yiming Wang 0008, Qian Huang 0008, Jiwen Liu, Chuanxu Jiang, Mingzhou Shang |
Multim. Tools Appl. | 1 |
| 2023 | FGC-VC: Flow-Guided Context Video CompressionabstractDeep video compression has attracted more and more attention in recent years. Previous works rely on feature space operations, which may cause the offset maps overflow degrading reconstructed frame quality. In this work, we propose a flow- guided module to guide the offset maps learning explicitly and alleviate offset maps overflow. Moreover, we introduce a context scheme to explore the temporal prior and fuse the hyper prior model to improve the compression ratio. For coding speed, we drop the time-consuming auto regressive module. Experimental results demonstrate that our method out-performs the previous learning-based schemes and traditional codecs. Compared to x265 with medium preset, our approach brings average 38.53% and 54.67% bit rate savings in PSNR and MS-SSIM metrics, respectively. Yiming Wang 0008, Qian Huang 0008, Bin Tang 0002, Huashan Sun |
ICIP | 1 |
| 2023 | End-to-End Variable-Rate Image Compression with Bi-Resolution Spatial-Channel Context AggregationabstractRecently, neural network-based image compression techniques have demonstrated remarkable compression performance. The use of context-adaptive entropy models greatly enhances the rate-distortion (R-D) performance by effectively capturing spatial redundancy in latent representations. However, latent representations still contain some spatial correlations(e.g. same spatial structure), it needs to be eliminated by further processing. And many compression models are single-rate model, which is difficult to cover a big range of bitrate. In order to address this issue, we propose a novel variable-rate image compression algorithm that efficiently leverages bi-resolution spatial-channel information through learned mechanisms. In this paper, we first proposed a BRP network to divide our latent representations and side information into HR and LR components, eliminating the spatial redundancy in same location. Combining the spatial-channel context, we proposed a BSC context model, including a decreasing-granularity checkerboard pattern and channel grouping based on cosine slicing strategy. To cover a wide range of bitrate, we take a weight map as input to control bit allocation, achieving multiple compression rates. Our experimental results show that our method provides a better rate-distortion trade-off than BPG, JPEG and other recent image compression methods based on deep learning. Qian Huang 0008, Yiming Wang 0008, Huashan Sun |
MMAsia | 3 |
| 2023 | Optical Flow based Feature Prediction and Decomposed Context for Video CompressionabstractIn recent years, there have been a growing interest in developing end-to-end neural video codecs. Previous works generally use a past decoded frame as reference directly, utilizing the motion information between it and the input frame to reduce temporal redundancy. However, this approach may lead to high bit rate consumption of the motion and fails to take advantage of the prior information in other reconstructed frames. In this work, We propose a learned video coding framework with optical flow based feature prediction module and decomposed context module. Specifically, we employ the previous optical flow to generate a warped frame, and along with other reconstructions, they are used for a more accurate reference forecasting, thereby reducing the bit rate required for motion compression. Moreover, based on the conditional coding framework, our decomposed context module explores conditional context in past decoded frames and further reduces additional spatiotemporal correlations. Experimental results demonstrate that our approach yields better performance than previous learned video compression methods and traditional standard codecs. For example, our neural codec achieves 28.94% coding gain over HEVC in PSNR metric and about 2.00% coding gain over VVC in MS-SSIM metric. Huashan Sun, Qian Huang 0008, Yiming Wang 0008, Ruoyu Hao |
MMAsia | 3 |
| 2023 | Video stabilization: A comprehensive survey
Yiming Wang 0008, Qian Huang 0008, Chuanxu Jiang, Jiwen Liu, Mingzhou Shang, Zhuang Miao |
Neurocomputing | 1 |
| 2022 | Intelligent Video Surveillance Platform Based on FFmpeg and Yolov5abstractWith the development of multimedia, video surveillance systems are becoming more popular. However, the current video surveillance systems have a general function and are unable to provide Intelligent perception. Chuanxu Jiang, Yanfang Wang 0005, Qian Huang 0008, Yiming Wang 0008, Yuhan Dai |
MMAsia | 4 |