VLDB 2026 Research / reviewers in the wild / expert
Jinxiang Liu
dblp:70/6217
· DBLP profile ↗
8ranked-venue papers
5as first author
7since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Photovoltaic power interval prediction method based on spatio-temporal fusion diffusion modelabstractProbability prediction of photovoltaic power plays an important role in improving photovoltaic absorption capacity. Conditional diffusion model takes numerical weather prediction (NWP) data as input, improving the performance of photovoltaic power day-ahead probability prediction. However, with the improvement of NWP system, there is some redundancy in NWP data. In addition, the previous research on diffusion model did not consider the temporal and spatial correlation between photovoltaic power stations, so there is room for further improvement in performance. Therefore, a day-to-day forecasting method of photovoltaic power interval based on spatio-temporal fusion diffusion model is proposed. Firstly, Pearson and mutual information are used to analyze NWP data to reduce the input dimension of the model. Secondly, a spatio-temporal fusion diffusion model based on graph attention networks and Transformer is constructed to make full use of the spatio-temporal correlation between photovoltaic power stations. Then, the inverse diffusion method of denoising diffusion implicit model is constructed, and multiple groups of photovoltaic power data are generated based on the same NWP input, and the kernel density is estimated and analyzed to obtain the prediction interval. Finally, experiments are carried out on public data sets and compared with other algorithms. The results show that the proposed method can achieve higher comprehensive probability prediction performance. Yishi Chen, Jinxiang Liu, Jingxin He, Taozhan Zhang, Senlin Zhang |
IECON | 2 |
| 2024 | Audio-Visual Segmentation via Unlabeled Frame ExploitationabstractAudio-visual segmentation (AVS) aims to segment the sounding objects in video frames. Although great progress has been witnessed, we experimentally reveal that current methods reach marginal performance gain within the use of the unlabeled frames, leading to the underutilization issue. To fully explore the potential of the unlabeled frames for AVS, we explicitly divide them into two categories based on their temporal characteristics, i.e., neighboring frame (NF) and distantframe (DF). NFs, temporally adjacent to the labeled frame, often contain rich motion information that assists in the accurate localization of sounding objects. Contrary to NFs, DFs have long temporal distaaces from the labeled frame, which share semantic-similar objects with appearance variations. Considering their unique characteristics, we propose a versatile framework that effectively leverages them to tackle AVS. Specifically, for NFs, we exploit the motion cues as the dynamic guidance to improve the objectness localization. Besides, we exploit the semantic cues in DFs by treating them as valid augmentations to the labeled frames, which are then used to enrich data diversity in a self-training manner. Extensive experimental results demonstrate the versatility and superiority of our method, unleashing the power of the abundant unlabeled frames. Jinxiang Liu, Fei Zhang 0016, Chen Ju, Ya Zhang 0002, Yanfeng Wang 0001 |
CVPR | 1 |
| 2024 | Annotation-free Audio-Visual SegmentationabstractThe objective of Audio-Visual Segmentation (AVS) is to localise the sounding objects within visual scenes by accurately predicting pixel-wise segmentation masks. To tackle the task, it involves a comprehensive consideration of both the data and model aspects. In this paper, first, we initiate a novel pipeline for generating artificial data for the AVS task without extra manual annotations. We leverage existing image segmentation and audio datasets and match the image-mask pairs with its corresponding audio samples using category labels in segmentation datasets, that allows us to effortlessly compose (image, audio, mask) triplets for training AVS models. The pipeline is annotation-free and scalable to cover a large number of categories. Additionally, we introduce a lightweight model SAMA-AVS which adapts the pre-trained segment anything model (SAM) to the AVS task. By introducing only a small number of trainable parameters with adapters, the proposed model can effectively achieve adequate audio-visual fusion and interaction in the encoding stage with vast majority of parameters fixed. We conduct extensive experiments, and the results show our proposed model remarkably surpasses other competing methods. Moreover, by using the proposed model pretrained with our synthetic data, the performance on real AVSBench data is further improved, achieving 83.17 mIoU on S4 subset and 66.95 mIoU on MS3 set. The project page is https://jinxiang-liu.github.io/anno-free-AVS/. Jinxiang Liu, Yu Wang 0027, Chen Ju, Chaofan Ma, Ya Zhang 0002, Weidi Xie |
WACV | 1 |
| 2023 | Distilling Vision-Language Pre-Training to Collaborate with Weakly-Supervised Temporal Action LocalizationabstractWeakly-supervised temporal action localization (WTAL) learns to detect and classify action instances with only category labels. Most methods widely adopt the off-the-shelf Classification-Based Pre-training (CBP) to generate video features for action localization. However, the different optimization objectives between classification and localization, make temporally localized results suffer from the serious incomplete issue. To tackle this issue without additional annotations, this paper considers to distill free action knowledge from Vision-Language Pre-training (VLP), as we surprisingly observe that the localization results of vanilla VLP have an over-complete issue, which is just complementary to the CBP results. To fuse such complementarity, we propose a novel distillation-collaboration framework with two branches acting as CBP and VLP respectively. The framework is optimized through a dual-branch alternate training strategy. Specifically, during the B step, we distill the confident background pseudo-labels from the CBP branch; while during the F step, the confident foreground pseudo-labels are distilled from the VLP branch. As a result, the dualbranch complementarity is effectively fused to promote one strong alliance. Extensive experiments and ablation studies on THUMOS14 and ActivityNet1.2 reveal that our method significantly outperforms state-of-the-art methods. Chen Ju, Kunhao Zheng, Jinxiang Liu, Peisen Zhao, Ya Zhang 0002, Jianlong Chang, Qi Tian 0001, Yanfeng Wang 0001 |
CVPR | 3 |
| 2022 | Exploiting Transformation Invariance and Equivariance for Self-supervised Sound LocalisationabstractWe present a simple yet effective self-supervised framework for audio-visual representation learning, to localize the sound source in videos. To understand what enables to learn useful representations, we systematically investigate the effects of data augmentations, and reveal that (1) composition of data augmentations plays a critical role, i.e. explicitly encouraging the audio-visual representations to be invariant to various transformations (transformation invariance); (2) enforcing geometric consistency substantially improves the quality of learned representations, i.e. the detected sound source should follow the same transformation applied on input video frames (transformation equivariance). Extensive experiments demonstrate that our model significantly outperforms previous methods on two sound localization benchmarks, namely, Flickr-SoundNet and VGG-Sound. Additionally, we also evaluate audio retrieval and cross-modal retrieval tasks. In both cases, our self-supervised models demonstrate superior retrieval performances, even competitive with the supervised approach in audio retrieval. This reveals the proposed framework learns strong multi-modal representations that are beneficial to sound localisation and generalization to further applications. The project page is https://jinxiang-liu.github.io/SSL-TIE. Jinxiang Liu, Chen Ju, Weidi Xie, Ya Zhang 0002 |
ACM Multimedia | 1 |
| 2022 | Sea Surface Green Algae Density Estimation Using Ship-Borne GEO-Satellite Reflection ObservationsabstractIn recent years, global navigation satellite systems-reflectometry (GNSS-R) technology has been increasingly considered for applications in sea surface monitoring. This paper presents a new method to retrieve the density of sea surface green algae by using the reflected signals of geostationary Earth orbit (GEO) satellites collected by shipborne receiver. Because GEO satellites are stationary relative to a fixed receiver on the earth’s surface, the reflected GEO satellite (GEO-R) signals are not affected by Doppler frequency or elevation angle, which can greatly simplify the modeling of the reflected power and realize continuous green algae monitoring in the same area. Specifically, the influence of green algae on GEO-R power through varying reflection coefficient and roughness was analyzed. Then, an empirical model was established to retrieve the green algae density by using the GEO-R power. Finally, the experimental data collected in the Qingdao Jiaozhou bay were used to verify the developed models, and the results show that the inversion accuracy of the green algae density model is better than 4%. Wei Ban, Nanshan Zheng, Kegen Yu, Kefei Zhang 0003, Jinxiang Liu |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2022 | A 3D Mesh-Based Lifting-and-Projection Network for Human Pose TransferabstractHuman pose transfer has typically been modeled as a 2D image-to-image translation problem. This formulation ignores the human body shape prior in 3D space and inevitably causes implausible artifacts, especially when facing occlusion. To address this issue, we propose alifting-and-projectionframework to perform pose transfer in the 3D mesh space. The core of our framework is a foreground generation module, that consists of two novel networks: a lifting-and-projection network (LPNet) and an appearance detail compensating network (ADCNet). To leverage the human body shape prior, LPNet exploits the topological information of the body mesh to learn an expressive visual representation for the target person in the 3D mesh space. To preserve texture details, ADCNet is further introduced to enhance the feature produced by LPNet with the source foreground image. Such design of the foreground generation module enables the model to better handle difficult cases such as those with occlusions. Experiments on the iPER and Fashion datasets empirically demonstrate that the proposed lifting-and-projection framework is effective and outperforms the existing image-to-image-based and mesh-based methods on human pose transfer task in both self-transfer and cross-transfer settings. Jinxiang Liu, Yangheng Zhao, Siheng Chen, Ya Zhang 0002 |
IEEE Trans. Multim. | 1 |
| 2000 | Direct Minutiae Extraction from Gray-Level Fingerprint Image by Relationship ExaminationabstractA new method to extract the minutiae from a gray-level fingerprint image is presented in this paper. The proposed method determines the minutiae by examining the relationship between the ridges and furrows. The minutiae change the relationship between the ridges and furrows, so the minutiae can be detected by finding the relationship change. By following the ridges and furrows along the local direction, the relationship change can be extracted. Experimental results demonstrate that the method developed based on the new approach is insensitive to noise and hence extracts minutiae better. Jinxiang Liu, ZhongYang Huang, Kap Luk Chan |
ICIP | 1 |