VLDB 2026 Research / reviewers in the wild / expert
Xiaofeng Jin
dblp:127/0912
· DBLP profile ↗
18ranked-venue papers
3as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 6 since 2021Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021Computer networks · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FePS-Net: A Frequency Enhanced Prosody-Style Fusion Network for Deepfake Speech Detection
Yankai Zhao, Xiaofeng Jin, Guirong Wang |
ICIC (14) | 3 |
| 2025 | Zero-Shot Character Recognition Method of Korean Ancient Documents Based on the Chinese and Korean Characters Unified IDS Encoding
Mengling Zhao, Xiaofeng Jin, Guirong Wang, Yankai Zhao |
ADMA (1) | 2 |
| 2025 | U-SAM: Upgrade Segment Anything Model With Semantic-Aware and Memory-EfficientabstractSegment Anything Model (SAM) has achieved remarkable success in the field of class-agnostic image segmentation by utilizing points or boxes as prompts. However, we identify two significant limitations when compared to traditional image segmentation models: (1) Trained in a category-agnostic interactive segmentation manner, SAM lacks the ability to discern object granularity and semantics, rendering it ineffective for traditional instance, semantic, and panoptic segmentation tasks. (2) SAM’s inefficient use of instance-independent visual features and tokens necessitates maintaining unique features and tokens for each instance, leading to excessive GPU memory consumption and diminished segmentation efficiency. To address these issues, we propose the Universal Segment Anything Model (U-SAM), a semantic-aware and memory-efficient segmentation model designed to perform both promptable and traditional segmentation tasks within a compact and unified framework. Specifically, U-SAM enhances SAM by integrating the Multi-Scale Semantic-Aware Image Encoder (S2IE), thus providing multi-scale semantic features for achieving traditional image segmentation tasks. Additionally, U-SAM is equipped with a Twin Token Mask Decoder (T2MD) which reduces GPU memory overhead by substituting replicated visual features with replicated tokens. Extensive experiments across interactive, instance, semantic, and panoptic segmentation demonstrate U-SAM’s promising results. Notably, U-SAM is 9× smaller and 10× faster than SAM, showing strong performance in zero-shot segmentation. Moreover, U-SAM surpasses the SOTA object-prompter-based model, RSPrompter, by achieving a 6.2% increase in PQ, operating 14× faster, and cutting training memory usage by 61%. Xiaofeng Jin, Jie Hu 0018, Jianghang Lin, Shengchuan Zhang, Liujuan Cao |
ICASSP | 1 |
| 2025 | Rendering Anywhere You See: Renderability Field-guided Gaussian SplattingabstractScene view synthesis, which generates novel views from limited perspectives, is increasingly vital for applications like virtual reality, augmented reality, and robotics. Unlike object-based tasks, such as generating 360° views of a car, scene view synthesis handles entire environments where non-uniform observations pose unique challenges for stable rendering quality. To address this issue, we propose a novel approach: renderability field-guided gaussian splatting (RF-GS). This method quantifies input inhomogeneity through a renderability field, guiding pseudo-view sampling to enhanced visual consistency. To ensure the quality of wide-baseline pseudo-views, we train an image restoration model to map point projections to visiblelight styles. Additionally, our validated hybrid data optimization strategy effectively fuses information of pseudo-view angles and source view textures. Comparative experiments on simulated and real-world data show that our method outperforms existing approaches in rendering stability. Xiaofeng Jin, Matteo Frosi, Jianfei Ge, Jiangjian Xiao, Matteo Matteucci |
IROS | 1 |
| 2025 | OpenFusion++: An Open-vocabulary Real-time Scene Understanding SystemabstractReal-time open-vocabulary scene understanding is essential for efficient 3D perception in applications such as vision-language navigation, embodied intelligence, and augmented reality. However, existing methods suffer from imprecise instance segmentation, static semantic updates, and limited handling of complex queries. To address these issues, we present OpenFusion++, a TSDF-based real-time 3D semantic-geometric reconstruction system. Our approach refines 3D point clouds by fusing confidence maps from foundational models, dynamically updates global semantic labels via an adaptive cache based on instance area, and employs a dual-path encoding framework that integrates object attributes with environmental context for precise query responses. Experiments on the ICL, Replica, ScanNet, and ScanNet++ datasets demonstrate that OpenFusion++ significantly outperforms the baseline in both semantic accuracy and query responsiveness. Xiaofeng Jin, Matteo Frosi, Matteo Matteucci |
IROS | 1 |
| 2025 | Universal Image Segmentation With EfficiencyabstractIn this paper, we present UISE, a unified image segmentation framework that achieves efficient performance across various segmentation tasks, eliminating the need for multiple specialized pipelines. UISE employs dynamic convolutions between universal segmentation kernels and image feature maps, enabling a single pipeline for different tasks such as panoptic, instance, semantic, and video instance segmentation. To address computational requirements, we introduce a feature pyramid aggregator for image feature extraction and a separable dynamic decoder for generating segmentation kernels. The aggregator re-parameterizes interpolation-first modules in a convolution-first manner, resulting in a significant acceleration of the pipeline without incurring additional costs. The decoder incorporates multi-head cross-attention through separable dynamic convolution, enhancing both efficiency and accuracy. Extensive experiments are conducted to validate UISE's performance across different segmentation tasks. To the best of our knowledge, UISE is the first universal segmentation framework that delivers competitive performance in terms of both speed and accuracy when compared to current state-of-the-art models. Jie Hu 0018, Liujuan Cao, Xiaofeng Jin, Shengchuan Zhang, Rongrong Ji |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Balanced Active Sampling for Person Re-identificationabstractActive learning is attracting more and more attention in person re-identification (Re-ID), as it is promising in the scalability of Re-ID models to satisfy performance with reduced labeling cost. Active sampling of pair-wise images in Re-ID is a highly imbalanced problem, where negative pairs are the vast majority. To avoid sampled pairs being dominated by the negative relationship, previous works tend to sample pairs with confident positive relationships in various ways. However, it is a waste of the labeling budget as most sampled pairs will be positive and already have a very close distance. Thus, there is no significant improvement in the model performance. In this paper, we first argue that balanced sampling is the key to active learning for Re-ID. Along this line, we propose a naïve balanced sampling method based on the global estimation of the most confusing distance. It is further improved by the label-wise estimation and diversity measurement. We also formulate the training of Re-ID models as a constrained clustering problem, where labeled positive and negative pairs are as must-link and cannot-link. Then the model training is based on the pseudo labels. Extensive experiments on benchmarks evaluate the effectiveness and superiority of the proposed methods. Specifically, it achieves comparable performance with supervised counterparts with less than 0.1% pair-wise annotation, which significantly surpasses the state-of-the-art. Leqi Shen, Guiguang Ding, Zhiheng Zhou 0001, Tianshi Xu, Xiaofeng Jin, Yuheng Huang 0005 |
ICME | 6 |
| 2024 | Camera Bias Regularization for Person Re-identificationabstractPerson re-identification (Re-ID) is to match persons captured by non-overlapping cameras. Due to the discrepancies between cameras caused by illumination, background, or viewpoint, the underlying difficulty for Re-ID is the camera bias problem, which leads to the large gap of within-identity features from different cameras. With limited cross-camera annotation, Re-ID models tend to learn camera-related features, instead of identity-related features. Consequently, Re-ID models suffer from poor transfer ability from seen to unseen domains. In this paper, we investigate the camera bias problem in both supervised and unsupervised learning. In particular, we propose a novel Camera Bias Regularization (CBR) term to reduce the feature distribution gap between cameras. The CBR works by simultaneously enlarging the distance of intra-camera distributions between positive and negative pairs, and reducing the distance of positive pairs’ distributions between intra-camera and cross-camera. In addition, a Cross-Camera (CC) clustering method is also designed for unsupervised learning, which puts more emphasis on cross-camera pairs than intra-camera ones during the clustering process. Extensive experiments are conducted to validate the effectiveness of the proposed CBR and CC. Specifically, with only a plain ResNet-50, it achieves 56.7% mAP and 40.7% mAP on the challenging MSMT17 dataset in supervised and unsupervised settings respectively, which surpasses most state-of-the-arts. Leqi Shen, Guiguang Ding, Zhiheng Zhou 0001, Tianshi Xu, Xiaofeng Jin, Yuheng Huang 0005 |
ICME | 6 |
| 2024 | SPformer: Hybrid Sequential-Parallel Architectures for Automatic Speech RecognitionabstractIn recent years, the capability to interact with multi-scale information has been regarded as a crucial aspect of the Automatic Speech Recognition (ASR) encoder’s abilities. Conformer and Branchformer, representing sequential and parallel architectural designs, respectively, facilitate the interaction between global and local information, achieving state-of-the-art performance. However, sequential architectures struggle with explicability in the interaction process and rigid model design, while parallel architectures face challenges in integration difficulties and limited interaction. To address these issues, we propose the SPformer, effectively combining sequential connection and parallel branch architectures. It allows dynamic interaction between convolution and self-attention while utilizing branch structures. The SPformer’s performance, both in-domain and out-of-domain, surpasses that of Conformer and E-Branchformer, as demonstrated by our experiments on public ASR datasets. Mingdong Yu, Xiaofeng Jin, Guirong Wang, Bo Wang 0105 |
ICME | 2 |
| 2024 | Pre-training Encoder-Decoder for Minority Language Speech RecognitionabstractAlthough significant progress has been made in the field of speech recognition for widely used languages, minority languages remain constrained by a shortage of high-quality data and labels. This issue not only limits the development of speech recognition technology for minority language communities but also impacts the protection of linguistic and cultural diversity. To address this issue, we propose a pre-training encoder-decoder model Predformer. Inspired by contrastive learning methods, our research in the pre-training encoder phase involves comparing the differences in speech across various data representations to deeply mine the implicit representational information in the data. At the same time, in the pre-training decoder phase, autoregressive reconstruction of codes is used to capture the contextual relationships between semantics. Building on this, to alleviate the complexities and uncertainties encountered in processing long sequences, we propose the introduction of a self-attention with a memory module during the pre-training phase of the decoder, aimed at enhancing learning efficiency and effectively overcoming these challenges. Experimental results show that the Predformer model significantly improves the performance of nine minority language speech recognition model baselines and can transfer learning between similar speeches. Bo Wang 0105, Xiaofeng Jin, Mingdong Yu, Guirong Wang |
IJCNN | 2 |
| 2024 | DINO-ViT Enhanced Diffusion for Multi-exemplar-based Image TranslationabstractWe have developed a framework for multi-exemplar-based image translation. Most exemplar-based image translation methods allow only one target image for appearance transfer. In these existing methods, GANs are often used as generators, and features extracted from pre-trained CNNs are used as visual descriptors. By comparison, our framework allows users to provide one or more images as exemplars to realize the appearance transfer of different objects while preserving the structure of the source image. Methods based on GANs are typically limited to a specific domain. To overcome this, we choose a diffusion model as the generator in our framework, allowing image translation across arbitrary domains. DINO-ViT, as a visual transformer trained by self-supervision, its deep features have rich semantic properties and visual information compared to CNNs’. Therefore, we use the deep features extracted from pretrained DINO-ViT as visual descriptors to achieve more accurate image translation. The naive combination of translation results from different exemplars may lead to visual disharmony due to rough edges and differences in object visual characteristics. To address this, we introduce small amounts of noise on the combined image to reduce inconsistencies and use the reverse process of diffusion to denoise, then we can obtain an image that is almost identical to the combined image but without artifacts. Our framework offers higher-quality image translation and more flexible image editing, and we have demonstrated the effectiveness and superiority of our approach in several image translation tasks. Jiangjian Xiao, Xiaofeng Jin, Xiaojing Gu, Gen Xu |
IJCNN | 4 |
| 2024 | Separable Spatial-Temporal Residual Graph for Cloth-Changing Group Re-IdentificationabstractGroup re-identification (GReID) aims to correctly associate group images belonging to the same group identity, which is a crucial task for video surveillance. Existing methods only model the member feature representations inside each image (regarded as spatial members), which leads to potential failures in long-term video surveillance due to cloth-changing behaviors. Therefore, we focus on a new task called cloth-changing group re-identification (CCGReID), which needs to consider group relationship modeling in GReID and robust group representation against cloth-changing members. In this paper, we propose the separable spatial-temporal residual graph (SSRG) for CCGReID. Unlike existing GReID methods, SSRG considers both spatial members inside each group image and temporal members among multiple group images with the same identity. Specifically, SSRG constructs full graphs for each group identity within the batched data, which will be completely and non-redundantly separated into the spatial member graph (SMG) and temporal member graph (TMG). SMG aims to extract group features from spatial members, and TMG improves the robustness of the cloth-changing members by feature propagation. The separability enables SSRG to be available in the inference rather than only assisting supervised training. The residual guarantees efficient SSRG learning for SMG and TMG. To expedite research in CCGReID, we develop two datasets, including GroupPRCC and GroupVC, based on the existing CCReID datasets. The experimental results show that SSRG achieves state-of-the-art performance, including the best accuracy and low degradation (only 2.15% on GroupVC). Moreover, SSRG can be well generalized to the GReID task. As a weakly supervised method, SSRG surpasses the performance of some supervised methods and even approaches the best performance on the CSG dataset. Jian-Huang Lai, Xiaohua Xie, Xiaofeng Jin, Sien Huang |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | OAM Spatial Field Digital Modulation System for Physical-Level Secure CommunicationabstractIn this paper, we investigate a novel physical layer secure (PLS) potential technique which is summarized into a category of index modulation (IM) scheme, namely the spatial field digital modulation (SFDM). In an SFDM communication system, bit streams are modulated in the spatial distribution of electromagnetic (EM) field. Orbital angular momentum (OAM) mode groups can be utilized for beamforming and thus realize an OAM based SFDM (OAM-SFDM) system. This kind of IM systems that utilize the distribution of EM field to transmit information can radiate different signals to different directions and possess inherent anti-eavesdropping capability. Since an eavesdropper lacks prior knowledge, it is difficult to construct appropriate and effective estimators for signal demodulation. However, such system is vulnerable when potential eavesdroppers adopt clustering algorithm, a type of non-realtime demodulation. We propose two PLS strategies with different computational complexity, the power hopping (PH) method and the symbol variation (SV) method, to confront single-antenna eavesdroppers and multi-antenna eavesdroppers, respectively. Both approaches are key-free, and do not require any prior channel state information (CSI) about eavesdroppers. The mentioned methods could constantly disrupt the statistical properties of eavesdroppers’ channels and even their received signals. Theoretical analysis and numerical simulations have been conducted to validate the feasibility of these two schemes. Furthermore, a prototype of OAM-SFDM based PLS communication system is built in a realistic scenario. The experimental results demonstrate that both PH method and SV method can effectively resist clustering algorithm in their corresponding application scenarios. Yuqi Chen 0006, Xiaowen Xiong, Shilie Zheng, Zhaohui Yang 0001, Zelin Zhu, Bingchen Pan, Bincai Wu, Xiaonan Hui, Xiaofeng Jin, Xianbin Yu, Xianmin Zhang 0001 |
IEEE Trans. Wirel. Commun. | 9 |
| 2021 | Adaptively Fusing Complete Multi-resolution Features for Human Pose Estimation
Yuezhen Huang, Xiaofeng Jin, Yuheng Huang 0005, Tianshi Xu |
ICIG (2) | 4 |
| 2021 | Dual Gated Learning for Visible-Infrared Person Re-identification
Yuheng Huang 0005, Jincai Xian, Xiaofeng Jin, Tianshi Xu |
ICIG (2) | 4 |
| 2019 | Variational Representation Learning for Vehicle Re-IdentificationabstractVehicle Re-identification is attracting more and more attention in recent years. One of the most challenging problems is to learn an efficient representation for a vehicle from its multi-viewpoint images. Existing methods tend to derive features of dimensions ranging from thousands to tens of thousands. In this work we proposed a deep learning based framework that can lead to an efficient representation of vehicles. While the dimension of the learned features can be as low as 256, experiments on different datasets show that the Top-1 and Top-5 retrieval accuracies exceed multiple state-of-the-art methods. The key to our framework is two-fold. Firstly, variational feature learning is employed to generate variational features which are more discriminating. Secondly, long short-term memory (LSTM) is used to learn the relationship among different viewpoints of a vehicle. The LSTM also plays as an encoder to downsize the features. Saghir Ahmed Saghir Alfasly, Yongjian Hu, Tiancai Liang, Xiaofeng Jin, Qingli Zhao |
ICIP | 4 |
| 2017 | Mode Division Multiplexing Communication Using Microwave Orbital Angular Momentum: An Experimental StudyabstractMode division multiplexing (MDM) using orbital angular momentum (OAM) is a recently developed physical layer transmission technique, which has obtained intensive interest among optics, millimeter-wave, and radio frequency due to its capability to enhance communication capacity while retaining an ultra-low receiver complexity. In this paper, the system model based on OAM-MDM is mathematically analyzed and it is theoretically concluded that such system architecture can bring a vast reduction in receiver complexity without capacity penalty compared with conventional line-of-sight multiple-in-multiple-out systems under the same physical constraint. Furthermore, a$4\times 4$OAM-MDM communication experiment adopting a pair of easily realized Cassegrain reflector antennas capable of multiplexing/demultiplexing four orthogonal OAM modes of$l = {-3}$, −2, +2, and +3 is carried out at a microwave frequency of 10 GHz. The experimental results show high spectral efficiency as well as low receiver complexity. Weite Zhang, Shilie Zheng, Xiaonan Hui, Ruofan Dong, Xiaofeng Jin, Hao Chi, Xianmin Zhang 0001 |
IEEE Trans. Wirel. Commun. | 5 |
| 2006 | An Entropy-based Evaluation Model of Business Strategic PerformanceabstractWith the rapid development of science and technology, product life cycle is shortening, and market competition is fiercer than before. To achieve continuous development, a company has to develop and implement right strategies to obtain and maintain its sustainable competitiveness. In the strategic management process, the company needs to evaluate the performance of its strategies so that corrective actions could be taken and new strategies be generated. This paper first reviews the development of researches on strategic performance evaluation. It then constructs a conceptual model of business strategic performance evaluation. A performance evaluation parameter system is established, which includes financial and non-financial factors, to construct an entropy-based model for evaluating strategic performance of companies. Finally, the paper does an empirical research on using the model to comprehensively evaluate strategic performance of 30 corporations listed in Shanghai Stock Exchange and Shenzhen Stock Exchange in appliance and electronics industry. The implications generated from the research are concluded that non-financial performances in market, technology, and governance, play an important role for a corporation to gain long-term competitive advantage. Xiaofeng Jin |
SMC | 2 |