Xingyu Gao 0001

dblp:32/4831-1 · DBLP profile ↗
← Back
78ranked-venue papers
10as first author
53since 2021 · last 2026
0000-0002-4660-8092ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 42 · 8 first-author · 26 since 2021Artificial intelligence and machine learning · 26 · 2 first-author · 18 since 2021Computer networks · 9 · 1 first-author · 8 since 2021Databases, data management, data science and information retrieval · 9 · 1 first-author · 6 since 2021Security and privacy · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
YearPublicationVenuePosition
2026 Biologically-Inspired Evolutionary Domain Symbiosis for Few-shot and Zero-shot Point Cloud Semantic Segmentation
abstract
Few-shot and zero-shot point cloud semantic segmentation aim to accurately segment novel categories using limited or no labeled samples, respectively. However, existing methods face significant challenges including domain shifts between support and query sets and the inability to handle both few-shot and zero-shot scenarios within a unified framework. To address these issues, we propose a biologically-inspired Evolutionary Domain Symbiosis Network EDS-Net for unified few-shot and zero-shot point cloud semantic segmentation. Specifically, inspired by natural symbiotic evolution, we propose a Symbiotic Evolution Module (SEM) that models co-adaptation between support and query features through self-correlation and cross-correlation mechanisms. Second, motivated by genetic crossover mechanisms, we introduce a Vision-Semantic Bridging Module (VSBM) that treats visual prototypes and semantic prototypes as two “parent” individuals, creating fused offspring prototypes through adaptive crossover operations and mutation strategies for zero-shot scenarios. Third, we develop a multi-generational evolutionary optimization framework employing an adaptive gating network to learn optimal fusion weights across different evolutionary stages. Extensive experiments demonstrate that EDS-Net with biological interpretability achieves state-of-the-art performance on both few-shot and zero-shot settings.
Changshuo Wang 0001, Zhijian Hu, Zaiyang Yu, Yibin Wu, Mingkun Xu, Yusong Wang 0003, Xingyu Gao 0001, Prayag Tiwari
AAAI8
2026 Sharpness-aware Federated Graph Learning
abstract
One of many impediments to applying graph neural networks (GNNs) in processing large-volume real-world graph-structured data is that it disapproves of a centralized training scheme which involves gathering data belonging to different organizations due to privacy concerns. As a distributed data processing scheme, federated graph learning (FGL) enables learning GNN models collaboratively without sharing participants' private data. Though theoretically feasible, a core challenge in FGL systems is the variation of local training data distributions among clients, also known as the data heterogeneity problem. Most existing solutions suffer from two problems: (1) The typical optimizer based on empirical risk minimization tends to cause local models to fall into sharp valleys and weakens their generalization to out-of-distribution graph data. (2) The prevalent dimensional collapse in the learned representations of local graph data has an adverse impact on the classification capacity of the GNN model. To this end, we formulate a novel optimization objective that is aware of the sharpness (i.e., the curvature of the loss surface) of local GNN models. By minimizing the loss function and its sharpness simultaneously, we seek out model parameters in a flat region with uniformly low loss values, thus improving the generalization over heterogeneous data. By introducing a regularizer based on the correlation matrix of local representations, we relax the correlations of representations generated by individual local graph samples, so as to alleviate the dimensional collapse of the learned model. The proposed Sharpness-aware fEderated grAph Learning (SEAL) algorithm can enhance the classification accuracy and generalization ability of local GNN models in federated graph learning. Experimental studies on several graph classification benchmarks show that SEAL consistently outperforms SOTA FGL baselines and provides gains for more participants.
Ruiyu Li, Peige Zhao, Guangxia Li, Xingyu Gao 0001, Zhiqiang Xu 0003
WSDM5
2026 DrugBLIP: exploring the protein-molecule interaction mechanisms with a multi-task learning graph transformer
abstract
MOTIVATION: Traditional drug discovery methods are costly and inefficient, while existing deep learning approaches remain limited by task specificity and practical applicability. Accurately modeling protein-molecule interactions is critical for advancing virtual screening, docking, and drug design. RESULTS: We propose DrugBLIP, a multi-task graph transformer model based on SE(3)-equivariant architectures, to unify protein-molecule interaction learning. By integrating contrastive learning, matching tasks, and docking optimization, DrugBLIP captures 3D spatial relationships through a hybrid graph transformer framework. Evaluations demonstrate state-of-the-art performance: DrugBLIP achieves an AUROC of 0.8217 and BEDROC of 0.5743 on virtual screening, outperforming traditional and deep learning baselines by 10%-127% across metrics. It also attains 91.2% top-1 docking success on CASF-2016 and 41.8% target fishing accuracy, showcasing robustness in diverse scenarios. Additionally, DrugBLIP reduces computational time by 700× compared to traditional docking tools. AVAILABILITY AND IMPLEMENTATION: Code is available at https://github.com/Wolkenwandler/DrugBLIP and archived at Zenodo with DOI: 10.5281/zenodo.16990700.
Rubo Wang, Xingyu Gao 0001, Peilin Zhao
Bioinform.2
2026 Counterfactual distribution intervention for few-shot class-incremental learning
Jicheng Yuan, Wenfa Li, Lusi Li, Liping Zhang 0014, Enhao Ning, Xingyu Gao 0001, Xin Ning 0001
Knowl. Based Syst.6
2026 PHANet: Contrastive hypergraph structures and prototype memory for discriminative skeleton-based action recognition
Chen Pang 0001, Guangqi Wen, Chunmeng Kang, Xingyu Gao 0001, Lei Lyu 0001
Knowl. Based Syst.5
2026 MSDformer: Multi-Scale Discrete Transformer for Time Series Generation
abstract
Discrete Token Modeling (DTM), which employs vector quantization techniques, has demonstrated remarkable success in modeling non-natural language modalities, particularly in time series generation. While our prior work SDformer established the first DTM-based framework to achieve state-of-the-art performance in this domain, two critical limitations persist in existing DTM approaches: 1) their inability to capture multi-scale temporal patterns inherent to complex time series data, and 2) the absence of theoretical foundations to guide model optimization. To address these challenges, we proposes a novel multi-scale DTM-based time series generation method, called Multi-Scale Discrete Transformer (MSDformer). MSDformer employs a multi-scale time series tokenizer to learn discrete token representations at multiple scales, which jointly characterize the complex nature of time series data. Subsequently, MSDformer applies a multi-scale autoregressive token modeling technique to capture the multi-scale patterns of time series within the discrete latent space. Theoretically, we validate the effectiveness of the DTM method and the rationality of MSDformer through the rate-distortion theorem. Comprehensive experiments demonstrate that MSDformer significantly outperforms state-of-the-art methods. Both theoretical analysis and experimental results demonstrate that incorporating multi-scale information and modeling multi-scale patterns can substantially enhance the quality of generated time series in DTM-based approaches.
Shibo Feng, Xi Xiao 0001, Zhong Zhang 0014, Qing Li 0006, Xingyu Gao 0001, Peilin Zhao
IEEE Trans. Pattern Anal. Mach. Intell.6
2026 Prompt-Free Efficient Adaptation of Segment Anything Model for Remote Sensing Landslide Detection
abstract
Landslides pose serious risks to human safety and infrastructure, making accurate and timely detection essential for disaster assessment. Although deep learning has advanced landslide recognition from remote sensing imagery, existing methods typically rely on task-specific designs, which limit their generalization ability across diverse geographic environments. The Segment Anything Model (SAM) offers strong generalization and zero-shot segmentation ability, yet its dependence on manually provided prompts and limited suitability for remote sensing imagery constrain its effectiveness in landslide detection. To address these limitations, we propose PF-SAM, a prompt-free and parameter-efficient adaptation framework tailored for landslide segmentation. PF-SAM eliminates manual prompts by autonomously constructing the prompt embeddings that drive SAM’s segmentation process and adapting SAM to landslide remote sensing data. Specifically, we first propose the CNN-Token Adapter, which extracts multi-scale CNN features and transforms them into prompt tokens, providing SAM with localized geometric and textural cues essential for detecting landslides of diverse shapes and sizes. Then, we propose the Global Prompt Generation Mechanism (GPGM) that further enhances segmentation: the Dense Prompt Generation Module (DPGM) generates the coarse mask as the dense prompt to guide SAM toward the correct landslide regions, while the Sparse Prompt Generation Module (SPGM) selects high-confidence embeddings to generate sparse prompts that enhance the precise localization of landslide regions. Extensive experiments on multiple landslide datasets show that PF-SAM achieves state-of-the-art performance, substantially surpassing existing methods in detection accuracy and robustness.
Zuolei Li, Xingyu Gao 0001, Zhenyu Chen 0003, Yawen Duan
IEEE Trans. Circuits Syst. Video Technol.2
2026 STKPS-Net: Spatio-Temporal Key Patch Selection Network for Few Shot Anomalous Action Recognition
abstract
For providing timely warnings and preventing potential damages, it is crucial to detect anomalous actions that threaten public safety through surveillance cameras. Compared to normal actions, anomalous actions often occupy only a small portion of surveillance videos and exhibit more complex manifestations in terms of time and space. Considering that normal action recognition methods fail to highlight crucial information from small-sized patches, we propose the Spatio-temporal Key Patch Selection Network (STKPS-Net). It includes a spatially adaptive key patch selection module to select small but informative patches, and a long-short feature map spatio-temporal relation module to capture dynamic changes in anomalous actions. Additionally, a spatio-temporal refined loss is introduced to enhance fine-grained feature learning. Experimental results on the HMDB51, Kinetics, and UCF-Crime v2 datasets show that our STKPS-Net achieves state-of-the-art performance in few-shot anomalous action recognition, outperforming the most competitive methods by 1.2% on the anomalous action dataset UCF-Crime v2.
Jinsheng Xiao, Ruidi Chen, Xingyu Gao 0001, Hailong Shi, Zhongyuan Wang 0001
IEEE Trans. Inf. Forensics Secur.4
2026 Multi-Perspective Visual Contrastive Decoding for Reliable Assistance
abstract
Multimodal Large Language Models (MLLMs) offer promising capabilities for assisting individuals with blindness and low vision (BLV), but their effectiveness is compromised when processing BLV-captured images, which typically suffer from three fundamental challenges: quality degradation, object incompleteness, and spatial misalignment. This article presents MPVCD (Multi-Perspective Visual Contrastive Decoding), a novel framework that addresses these challenges through visual contrastive decoding techniques. MPVCD implements three specialized perspectives: Noise Contrastive Decoding addresses quality issues by comparing predictions between original and noise-injected images; Retrieval Contrastive Decoding tackles object incompleteness by retrieving semantically similar images from a memory bank; and Focus Contrastive Decoding resolves spatial misalignment by focusing on detected object regions. These perspectives are dynamically balanced through an Adaptive Perspective Integration that optimizes token selection based on prediction confidence. Our comprehensive experiments across diverse datasets demonstrate MPVCD’s effectiveness in reducing hallucinations under varied scenarios. By generating more accurate and reliable visual descriptions, MPVCD represents a significant advancement toward assistive technologies that BLV users can confidently rely on for environmental understanding and decision-making.
Bocheng Pan, Hailong Shi, Xingyu Gao 0001
ACM Trans. Internet Things3
2026 Personalized Subgraph Federated Learning With Decoupled Data Heterogeneity in Mobile-Edge Computing
abstract
Graph Federated Learning (FL) has attracted extensive attention in recent years due to its ability to train global graph models in a distributed manner without exposing original local data. However, fine-grained data heterogeneity remains largely overlooked in collaborative graph model training. Existing graph FL methods that address heterogeneity are mostly adapted from traditional FL and fail to account for the unique complexity of graph-specific heterogeneity. Specifically, graph heterogeneity can be further decomposed into feature heterogeneity and structural heterogeneity, which are tightly coupled during local training. To address this issue, we propose a novel local graph module, Feature and Structure Decoupling Convolution (FSD-Conv), designed to disentangle the interplay between feature bias and structural bias. With FSD-Conv, clients can learn feature-related yet structure-unbiased representations, thereby alleviating the adverse impact of graph heterogeneity in federated training. Furthermore, we introduce FedFSD, a personalized graph FL framework that achieves effective personalized model aggregation through an explainable neural network operating in a low-dimensional space. Extensive experiments on six graph datasets under both disjoint and overlapping client partitioning schemes demonstrate the effectiveness of FedFSD in handling complex graph data heterogeneity.
Bisheng Tang, Xiaojun Chen 0004, Shaopu Wang, Yuexin Xuan, Zhendong Zhao, Xingyu Gao 0001
IEEE Trans. Mob. Comput.6
2026 CWRNN-INVR: A Coupled WarpRNN Based Implicit Neural Video Representation
abstract
Implicit Neural Video Representation (INVR) has emerged as a novel approach for video representation and compression, using learnable grids and neural networks. Existing methods focus on developing new grid structures efficient for latent representation and neural network architectures with large representation capability, lacking the study on their roles in video representation. In this paper, the difference between INVR based on neural network and INVR based on grid is first investigated from the perspective of video information composition to specify their own advantages, i.e., neural network for general structure while grid for specific detail. Accordingly, an INVR based on mixed neural network and residual grid framework is proposed, where the neural network is used to represent the regular and structured information and the residual grid is used to represent the remaining irregular information in a video. A Coupled WarpRNN-based multi-scale motion representation and compensation module is specifically designed to explicitly represent the regular and structured information, thus terming our method as CWRNN-INVR. For the irregular information, a mixed residual grid is learned where the irregular appearance and motion information are represented together. The mixed residual grid can be combined with the coupled WarpRNN in a way that allows for network reuse. Experiments show that our method achieves the best reconstruction results compared with the existing methods, with an average PSNR of 33.73 dB on the UVG dataset under the 3M model and outperforms existing INVR methods in other downstream tasks. The code can be found athttps://github.com/yiyang-sdu/CWRNN-INVR.git.
Yanbo Gao, Shuai Li 0005, Jinglin Zhang 0001, Hui Yuan 0001, Mao Ye 0001, Xingyu Gao 0001
IEEE Trans. Multim.8
2026 PosFormer: Generalizable Indoor Positioning via Global-Local Feature Fusion Network
abstract
Ultra-wideband (UWB), a short-pulse radio technology with high time resolution, has shown great potential for high-accuracy indoor positioning. However, UWB-based localization still faces significant challenges in multipath-rich nonline-of-sight (NLOS) environments. Traditional geometry-based positioning methods tend to overestimate distances, while existing deep learning approaches often fail to capture the complex spatial and temporal characteristics of UWB signal propagation, leading to limited positioning accuracy. To address these limitations, we propose a Transformer and convolutional neural network (CNN) dual-fusion network inspired by multipath physical information for indoor positioning, termed PosFormer. The Transformer and CNN modules of PosFormer capture the long-distance dependencies and spatial characteristics of channel impulse responses (CIRs), effectively extracting both global and local multiscale features. In addition, we design a nonadjacent anchor-subsets (NAASs) scheme to enrich the diversity of CIRs and introduce a lightweight transfer learning (TL) framework to improve deployment robustness across environments with limited labeled data. Extensive experiments on public industrial datasets demonstrate that PosFormer achieves a mean absolute error (MAE) of 17.64 cm, significantly outperforming classical models such as CNN, long short-term memory (LSTM), and Transformer. In industrial hall environments with numerous metal obstacles, TL enables the pretrained model to achieve a positioning accuracy of 35.92 cm with only 20% of the fingerprint samples, highlighting its practical value for cross-environment deployment.
Xing Xiao, Xingyu Gao 0001, Menggang Sheng, Deyi Peng
IEEE Trans. Neural Networks Learn. Syst.3
2026 mmWave Radar-based Personalized Multi-object Vital Signs Monitoring
abstract
Frequency Modulated Continuous Wave (FMCW)-based mmWave radar has attracted widespread attention because of its non-contact and high spatial resolution for vital signs monitoring. Meanwhile, current studies focus mainly on how to improve the detection performance of steady multiple objects or unsteady single objects. In this work, we propose an innovative method for identity-based multi-object vital signs monitoring under unsteady scenarios. The method automatically distinguishes between steady and motion states, and conducts a best-effort vital signs monitoring during unsteady scenarios. To this end, we design a weight vector enhancement method combined with object spatial positioning for differentiating multiple objects, and identify each object according to the gait-based EfficientNet model. We also design a steady-state detector based on the MobileNet-V2 network to find the slots of object keeping steady for vital signs monitoring and then apply the variational mode decomposition (VMD) algorithm to extract the respiratory and heart rates of a single object. The experimental results showed that the mean absolute error of respiratory rate and heart rate decreased to 1.37 bpm and 2.56 bpm respectively in the case of multiple objects. In addition, the steady-state detector achieves close to 98.1% accuracy in recognizing motion types, and the average recognition rate of identity recognition based on gait features reaches about 93.26%.
Jiefan Qiu, Xingyu Gao 0001, Dongfu Zhu, Mengqi Jiang, Jiahan Song, Hailong Shi
ACM Trans. Multim. Comput. Commun. Appl.2
2025 TS-LIF: A Temporal Segment Spiking Neuron Network for Time Series Forecasting
abstract
Spiking Neural Networks (SNNs) offer a promising, biologically inspired approach for processing spatiotemporal data, particularly for time series forecasting. However, conventional neuron models like the Leaky Integrate-and-Fire (LIF) struggle to capture long-term dependencies and effectively process multi-scale temporal dynamics. To overcome these limitations, we introduce the Temporal Segment Leaky Integrate-and-Fire (TS-LIF) model, featuring a novel dual-compartment architecture. The dendritic and somatic compartments specialize in capturing distinct frequency components, providing functional heterogeneity that enhances the neuron's ability to process both low- and high-frequency information. Furthermore, the newly introduced direct somatic current injection reduces information loss during intra-neuronal transmission, while dendritic spike generation improves multi-scale information extraction. We provide a theoretical stability analysis of the TS-LIF model and explain how each compartment contributes to distinct frequency response characteristics. Experimental results show that TS-LIF outperforms traditional SNNs in time series forecasting, demonstrating better accuracy and robustness, even with missing data. TS-LIF advances the application of SNNs in time-series forecasting, providing a biologically inspired approach that captures complex temporal dynamics and offers potential for practical implementation in diverse forecasting scenarios.
Shibo Feng, Wanjin Feng, Xingyu Gao 0001, Peilin Zhao, Zhiqi Shen 0001
ICLR3
2025 IgGM: A Generative Model for Functional Antibody and Nanobody Design
abstract
Immunoglobulins are crucial proteins produced by the immune system to identify and bind to foreign substances, playing an essential role in shielding organisms from infections and diseases. Designing specific antibodies opens new pathways for disease treatment. With the rise of deep learning, AI-driven drug design has become possible, leading to several methods for antibody design. However, many of these approaches require additional conditions that differ from real-world scenarios, making it challenging to incorporate them into existing antibody design processes. Here, we introduce IgGM, a generative model for the de novo design of immunoglobulins with functional specificity. IgGM simultaneously generates antibody sequences and structures for a given antigen, consisting of three core components: a pre-trained language model for extracting sequence features, a feature learning module for identifying pertinent features, and a prediction module that outputs designed antibody sequences and the predicted complete antibody-antigen complex structure. IgGM effectively predicts structures and designs novel antibodies and nanobodies. This makes it highly applicable in a wide range of practical situations related to antibody and nanobody design. Code is available at: https://github.com/TencentAI4S/IgGM.
Rubo Wang, Fandi Wu, Xingyu Gao 0001, Jiaxiang Wu 0001, Peilin Zhao, Jianhua Yao 0001
ICLR3
2025 Efficient Parallel Training Methods for Spiking Neural Networks with Constant Time Complexity
abstract
Spiking Neural Networks (SNNs) often suffer from high time complexity $O(T)$ due to the sequential processing of $T$ spikes, making training computationally expensive. In this paper, we propose a novel Fixed-point Parallel Training (FPT) method to accelerate SNN training without modifying the network architecture or introducing additional assumptions. FPT reduces the time complexity to $O(K)$, where $K$ is a small constant (usually $K=3$), by using a fixed-point iteration form of Leaky Integrate-and-Fire (LIF) neurons for all $T$ timesteps. We provide a theoretical convergence analysis of FPT and demonstrate that existing parallel spiking neurons can be viewed as special cases of our approach. Experimental results show that FPT effectively simulates the dynamics of original LIF neurons, significantly reducing computational time without sacrificing accuracy. This makes FPT a scalable and efficient solution for real-world applications, particularly for long-duration simulations.
Wanjin Feng, Xingyu Gao 0001, Wenqian Du 0005, Hailong Shi, Peilin Zhao, Chunyan Miao
ICML2
2025 OMS: One More Step Noise Searching to Enhance Membership Inference Attacks for Diffusion Models
abstract
The data-intensive nature of Diffusion models amplifies the risks of privacy infringements and copyright disputes, particularly when training on extensive unauthorized data scraped from the Internet. Membership Inference Attacks (MIA) aim to determine whether a data sample has been utilized by the target model during training, thereby serving as a pivotal tool for privacy preservation. Current MIA employs the prediction loss to distinguish between training member samples and non-members. These methods assume that, compared to non-members, members, having been encountered by the model during training result in a smaller prediction loss. However, this assumption proves ineffective in diffusion models due to the random noise sampled during the training process. Rather than estimating the loss, our approach examines this random noise and reformulate the MIA as a noise search problem, assuming that members are more feasible to find the noise used in the training process. We formulate this noise search process as an optimization problem and employ the fixed-point iteration to solve it. We analyze current MIA methods through the lens of the noise search framework and reveal that they rely on the first residual as the discriminative metric to differentiate members and non-members. Inspired by this observation, we introduce OMS, which augments existing MIA methods by iterating One More fixed-point Step to include a further residual, i.e., the second residual. We integrate our method into various MIA methods across different diffusion models. The experimental results validate the efficacy of our proposed approach.
Xiaomeng Fu, Xi Wang 0014, Jin Liu 0020, Jiao Dai, Jizhong Han, Xingyu Gao 0001
IJCAI7
2025 DR-VQA: Decompose-then-Reconstruct for Visual Question Answering in BLV Assistance
abstract
Visual impairment affects over 200 million individuals globally, creating significant challenges in daily visual tasks. While vision-language models offer transformative assistive potential, existing systems based on Multimodal Large Language Models (MLLMs) face a serious cross-contamination problem when processing real-world images captured by blind and low-vision (BLV) users: when jointly processing imperfect images and specific questions, current models are often misled by question assumptions rather than adhering to visual facts, generating hallucinations about objects not present in the image. We introduce DR-VQA (Decompose-then-Reconstruct Visual Question Answering), a novel framework that balances user intent with visual facts. Our approach prevents cross-contamination through structured reasoning. Our approach deliberately separates image processing from question analysis, ensuring model-generated descriptions are strictly based on image facts without being influenced by questions. Subsequently, through a structured decomposition mechanism, the system generates targeted sub-questions relevant to user intent, gradually aligning visual descriptions with user needs while minimizing question bias. During final synthesis, a memory-reset LLM reconstructs the reasoning chain with detailed information to generate responses that either provide evidence-supported conclusions or transparently acknowledge information limitations. Experimental evaluations demonstrate our framework's effectiveness in reducing hallucination risks while improving answer accuracy. By systematically balancing user intent with factual visual evidence, this work advances BLV-assistive technologies from probabilistic outputs to reliable visual assistance services.
Bocheng Pan, Hailong Shi, Xingyu Gao 0001
ACM Multimedia3
2025 Geometric Algebra-Enhanced Bayesian Flow Network for RNA Inverse Design
abstract
With the development of biotechnology, RNA therapies have shown great potential. However, different from proteins, the sequences corresponding to a single RNA three-dimensional structure are more abundant. Most of the existing RNA design methods merely take into account the secondary structure of RNA, or are only capable of generating a limited number of candidate sequences. To address these limitations, we propose a geometric-algebra-enhanced $\textbf{B}$ayesian $\textbf{F}$low $\textbf{N}$etwork for the inverse design of $\textbf{R}$NA, called $\textbf{RBFN}$. RBFN uses a Bayesian Flow Network to model the distribution of nucleotide sequences in RNA, enabling the generation of more reasonable RNA sequences. Meanwhile, considering the more flexible characteristics of RNA conformations, we utilize geometric algebra to enhance the modeling ability of the RNA three-dimensional structure, facilitating a better understanding of RNA structural properties. In addition, due to the scarcity of RNA structures and the limitation that there are only four types of nucleic acids, we propose a new time-step distribution sampling to address the scarcity of RNA structure data and the relatively small number of nucleic acid types. Evaluation on the single-state fixed-backbone re-design benchmark and multi-state fixed-backbone benchmark indicates that RBFN can outperform existing RNA design methods in various RNA design tasks, enabling effective RNA sequence design.
Rubo Wang, Xingyu Gao 0001, Peilin Zhao
NeurIPS2
2025 Adaptive Depth-Converted-Scale Convolution for Self-Supervised Monocular Depth Estimation
abstract
Self-supervised monocular depth estimation (MDE) has received increasing interests in the last few years. The objects in the scene, including the object size and relationship among different objects, are the main clues to extract the scene structure. However, previous works lack the explicit handling of the changing sizes of the object due to the change of its depth. Especially in a monocular video, the size of the same object is continuously changed, resulting in size and depth ambiguity. To address this problem, we propose a Depth-converted-Scale Convolution (DcSConv) enhanced monocular depth estimation framework, by incorporating the prior relationship between the object depth and object scale to extract features from appropriate scales of the convolution receptive field. The proposed DcSConv focuses on the adaptive scale of the convolution filter instead of the local deformation of its shape. It establishes that the scale of the convolution filter matters no less (or even more in the evaluated task) than its local deformation. Moreover, a Depth-converted-Scale aware Fusion (DcS-F) is developed to adaptively fuse the DcSConv features and the conventional convolution features. Our DcSConv enhanced monocular depth estimation framework can be applied on top of existing CNN based methods as a plug-and-play module to enhance the conventional convolution block. Extensive experiments with different baselines have been conducted on the KITTI benchmark and our method achieves the best results with an improvement up to 11.6% in terms of SqRel reduction. Ablation study also validates the effectiveness of each proposed module.
Yanbo Gao, Huibin Bai, Huasong Zhou, Xingyu Gao 0001, Shuai Li 0005, Hui Yuan 0001, Wei Hua 0002, Tian Xie 0011
IEEE Trans. Circuits Syst. Video Technol.4
2025 Deep Learning to Hash With Application to Cross-View Nearest Neighbor Search
abstract
Learning hash functions for approximate nearest neighbor search of high-dimensional data has received a surge of interests in recent years. Most existing methods are often concerned with learning hash functions for nearest neighbor search on high-dimensional data from a single source. In many real-world applications, data can be collected from diverse sources or represented using different feature descriptors. This raises an open challenge, i.e., the Cross-View Nearest Neighbor Search (CVNNS), where the representation of a query instance can be different from that of target instances to be retrieved in database. The key challenge of cross-view search is to learn an effective shared representation which can effectively connect the query instance and the target instances to be retrieved. In this paper, we present a new cross-view nearest neighbor search scheme by applying the emerging deep learning to hash techniques. In particular, we investigate two different architectures of deep Restricted Boltzmann Machines (RBMs) for learning to hash toward cross-view nearest neighbor search, and conduct extensive experiments to examine their empirical performance on diverse settings of cross-view image retrieval tasks. The encouraging results show that our technique outperforms the state-of-the-art approaches.
Xingyu Gao 0001, Zhenyu Chen 0003, Boshen Zhang, Jianze Wei
IEEE Trans. Circuits Syst. Video Technol.1
2025 Scribble-Supervised Video Object Segmentation via Scribble Enhancement
abstract
Current video object segmentation methods heavily rely on pixel-level mask annotations when training, which are expensive and time-consuming to acquire. To address this problem, some approaches try to train with sparse scribble annotations and take sparse target scribble as initial information for inference. However, due to the sparsity of scribble annotations, the performance is often limited, and the corresponding loss function needs to be designed. Inspired by the powerful ability of Segment Anything Model (SAM) to leverage prompt for segmentation, we argue that this problem can be alleviated by improving the quality of scribble. Therefore, we propose SEVOS, a framework for scribble-supervised video object segmentation, which contains a scribble enhancement algorithm and an semi-supervised video object segmentation network. Specifically, the scribble enhancement algorithm first samples corresponding positive sample points and negative sample points from target scribbles, and then feeds them into the SAM in turn, achieving high-quality scribble enhancement without human intervention. This algorithm augments the scribble-annotated video dataset, which is used for additional training of the model. Furthermore, we design a post-processing enhancement algorithm to further improve the prediction results. The obtained model outperforms state-of-the-art methods with a considerable performance gap, indicating the generalization and effectiveness of the proposed model.
Xingyu Gao 0001, Zuolei Li, Hailong Shi, Zhenyu Chen 0003, Peilin Zhao
IEEE Trans. Circuits Syst. Video Technol.1
2025 Unsupervised Feature Enrichment and Fidelity Preservation Learning Framework for Skeleton-Based Action Recognition
abstract
Unsupervised skeleton-based action recognition has achieved remarkable progress recently. Existing unsupervised learning methods suffer from severe overfitting problem, and thus small networks are used, significantly reducing the representation capability. To address this problem, the overfitting mechanism behind the unsupervised learning for skeleton-based action recognition is first investigated. It is observed that skeleton is already a relatively high-level and low-dimension feature, but not in the same manifold as the features for action recognition. Simply applying the existing unsupervised learning method tends to produce features that discriminate the different samples rather than action classes, resulting in the overfitting problem. To address this problem, this paper proposes an Unsupervised spatial-temporal Feature Enrichment and Fidelity Preservation (U-FEFP) learning framework to generate rich distributed features that contain all the information of a skeleton sample. A spatial-temporal feature transformation subnetwork is developed using channel-wise topology refinement graph convolutional block and graph convolutional gated recurrent unit block as the basic feature extraction network. The unsupervised Bootstrap Your Own Latent-based learning is utilized to generate rich distributed features, and the unsupervised pretext task-based learning is employed to preserve the information contained in the skeleton. The two unsupervised learning ways are collaborated as U-FEFP to produce robust and discriminative representations. Experimental results on four widely used benchmarks, namely NTU-RGB+D-60, PKU-MMD, NTU-RGB+D-120 and AAV-Human dataset, demonstrate that the proposed U-FEFP obtains the best result compared with the state-of-the-art unsupervised learning methods.
Chuankun Li, Shuai Li 0005, Yanbo Gao, Xingyu Gao 0001, Ping Chen 0004, Wanqing Li 0001
IEEE Trans. Circuits Syst. Video Technol.4
2025 Unlocking Generative Priors: A New Membership Inference Framework for Diffusion Models
abstract
Diffusion models pose risks of privacy breaches and copyright disputes, primarily stemming from the potential utilization of unauthorized data during the training phase. Membership inference is aimed to determine whether a specific sample has been used in the training process of a target model, representing a critical tool for privacy violation verification. However, the increased model complexity and stochasticity inherent in diffusion renders traditional shadow-model-based or metric-based methods ineffective when applied to diffusion models. Moreover, existing methods only yield binary classification labels which lack necessary comprehensibility in practical applications. In this paper, we explore a novel perspective for membership inference by leveraging the intrinsic generative priors within the diffusion model. Compared with unseen samples, training samples exhibit stronger generative priors within the diffusion model, enabling the successful reconstruction of substantially degraded training images. Consequently, we propose the Degrade Restore Compare (DRC) framework. In this framework, an image undergoes sequential degradation and restoration, and its membership is determined by comparing it with the restored counterpart. Experimental results verify that our approach not only significantly outperforms existing methods in terms of accuracy but also provides comprehensible decision criteria, offering evidence for potential privacy violations.
Xiaomeng Fu, Xi Wang 0014, Jin Liu 0020, Jiao Dai, Jizhong Han, Xingyu Gao 0001
IEEE Trans. Inf. Forensics Secur.7
2025 Looking Clearer With Text: A Hierarchical Context Blending Network for Occluded Person Re-Identification
abstract
Existing occluded person re-identification (re-ID) methods mainly learn limited visual information for occluded pedestrians from images. However, textual information, which can describe various human appearance attributes, is rarely fully utilized in the task. To address this issue, we propose a Text-guided Hierarchical Context Blending Network ( THCB-Net) for occluded person re-ID. Specifically, at the data level, informative multi-modal inputs are first generated to make full use of the auxiliary role of textual information and make image data have a strong inductive bias for occluded environments. At the feature expression level, we design a novel Hierarchical Context Blending (HCB) module that can adaptively integrate shallow appearance features obtained by CNNs and multi-scale semantic features from visual transformer encoder. At the model optimization level, a Multi-modal Feature Interaction (MFI) module is proposed to learn the multi-modal information of pedestrians from texts and images, then guide the visual transformer encoder and HCB module to further learn discriminative identity information for occluded pedestrians through Image-Multimodal Contrastive (IMC) learning. Extensive experiments on standard occluded person re-ID benchmarks demonstrate that the proposed THCB-Net outperforms state-of-the-art methods.
Changshuo Wang 0001, Xingyu Gao 0001, Meiqing Wu, Siew-Kei Lam, Shuting He, Prayag Tiwari
IEEE Trans. Inf. Forensics Secur.2
2025 Uncertainty-Aware Bilateral Transformer for Accurate and Reliable Iris Segmentation
abstract
Iris segmentation is a deterministic and critical part of the iris recognition system. However, its performance is usually degraded by data uncertainty in acquisition and annotation, impeding more accurate recognition of the iris recognition system. In the paper, we propose a bilateral self-attention by exploring spatial and visual relationships to effectively distinguish between iris and non-iris regions, then design a bilateral Transformer by enhancing spatial perception and hierarchical feature fusion to mitigate the impact of acquisition uncertainty. Besides, iris segmentation uncertainty learning is developed to estimate the uncertainty map according to prediction discrepancy. With the estimated uncertainty, a weighting scheme and a regularization term are designed to minimize the effect of annotation uncertainty. To investigate data uncertainty, the paper presents a challenging near-infrared iris dataset named UTIris. It comprises 3,690 images with high acquisition uncertainty and provides rich segmentation masks to explore annotation uncertainty. Furthermore, we manually label a large-scale iris dataset, ND-0405 [1], with additional binary maps of iris masks to evaluate segmentation performance. Experimental results on UTIris and four other databases demonstrate the effectiveness of the proposed method in iris segmentation, and its segmentation improvement consequently promotes recognition accuracy.
Jianze Wei, Xingyu Gao 0001, Yunlong Wang 0003, Ran He 0001, Zhenan Sun
IEEE Trans. Inf. Forensics Secur.2
2025 A Unified Framework for Bandit Online Multiclass Prediction
abstract
Bandit online multiclass prediction plays an important role in many real-world applications. In this paper, we propose a unified Bandit Online Multiclass Prediction (BOMP) framework. This framework is based on our proposed margin-based gradient descent approach. Its update step provides an unbiased estimate of the surrogate loss gradient and has a lower variance than existing methods. It also enables our algorithms to update even for incorrect predictions by penalizing the wrong classes. The link function of the framework can evolve over time, gradually incorporating online data information including second-order information into the potential functions. Based on the proposed framework, we investigate first-order and second-order bandit online multiclass prediction algorithms. Theoretical analysis demonstrates the superiority of our proposed update rule and bandit online multiclass prediction framework. Finally, we compare our proposed first-order and second-order bandit online multiclass prediction algorithms with several state-of-the-art methods on two synthetic and four real-world datasets. The encouraging results show that our proposed algorithms significantly outperform state-of-the-art techniques.
Wanjin Feng, Xingyu Gao 0001, Peilin Zhao, Steven C. H. Hoi
IEEE Trans. Knowl. Data Eng.2
2025 Deep Mutual Distillation for Unsupervised Domain Adaptation Person Re-Identification
abstract
Unsupervised domain adaptation person re-identification (UDA person re-ID) aims at transferring the knowledge on the source domain with expensive manual annotation to the unlabeled target domain. Most of the recent papers leverage pseudo-labels for the target images to accomplish this task. However, the noise in the generated labels hinders the identification system from learning discriminative features. To address this problem, we propose a deep mutual distillation (DMD) to generate reliable pseudo-labels for UDA person re-ID. The proposed DMD applies two parallel branches for feature extraction, and each branch serves as the teacher of the other to generate pseudo-labels for its training. This mutually reinforcing optimization framework enhances the reliability of pseudo-labels, improving the identification performance. In addition, we present a bilateral graph representation (BGR) to describe the pedestrian images. BGR mimics the person re-identification of the human to aggregate the identity features according to the visual similarity and attribute consistency. Experimental results on Market-1501 and Duke demonstrate the effectiveness and generalization of the proposed method.
Xingyu Gao 0001, Zhenyu Chen 0003, Jianze Wei, Rubo Wang, Zhijun Zhao
IEEE Trans. Multim.1
2025 Brain-Inspired Fast- and Slow-Update Prompt Tuning for Few-Shot Class-Incremental Learning
abstract
Few-shot class-incremental learning (FSCIL) aims to learn new classes incrementally with a limited number of samples per class. Foundation models combined with prompt tuning showcase robust generalization and zero-shot learning (ZSL) capabilities, endowing them with potential advantages in transfer capabilities for FSCIL. However, existing prompt tuning methods excel in optimizing for stationary datasets, diverging from the inherent sequential nature in the FSCIL paradigm. To address this issue, taking inspiration from the "fast and slow mechanism" of the complementary learning systems (CLSs) in the brain, we present fast- and slow-update prompt tuning FSCIL (FSPT-FSCIL), a brain-inspired prompt tuning method for transferring foundation models to the FSCIL task. We categorize the prompts into two groups: fast-update prompts and slow-update prompts, which are interactively trained through meta-learning. Fast-update prompts aim to learn new knowledge within a limited number of iterations, while slow-update prompts serve as meta-knowledge and aim to strike a balance between rapid learning and avoiding catastrophic forgetting. Through experiments on multiple benchmark tests, we demonstrate the effectiveness and superiority of FSPT-FSCIL. The code is available at https://github.com/qihangran/FSPT-FSCIL.
Hang Ran, Xingyu Gao 0001, Lusi Li, Weijun Li 0002, Songsong Tian, Gang Wang 0023, Hailong Shi, Xin Ning 0001
IEEE Trans. Neural Networks Learn. Syst.2
2025 Multimodal Consistency Suppression Factor for Fake News Detection
abstract
Recent multimodal fake news detection methods often use the consistency between textual and visual contents to determine the truth or fake of news information. Higher levels of textual-visual consistency typically lead to a greater likelihood of classifying a news item as real. However, a critical observation reveals that creators of most fake news intentionally select images that align with the textual content, thereby enhancing the credibility of the news. Consequently, high consistency between textual and visual contents alone cannot guarantee the authenticity of the information. To address this problem, we introduce a novel approach termed Multimodal Consistency-based Suppression Factor to modulate the significance of textual-visual consistency in information assessment. When the textual-visual matching is high, this suppression factor reduces the influence of consistency during the judgment process. Moreover, we use contrastive language-image pre-training (CLIP) model to extract features and measure the consistency level between modalities to guide multimodal fusion. In addition, we also use a method of compressing and fusing modal information based on variational autoencoder (VAE) to reconstruct CLIP features, learning the shared representation of different modal information of CLIP. Finally, extensive experiments were conducted on three publicly datasets, Weibo, Twitter, and Weibo21, and the results confirmed that our method outperformed the state-of-the-art methods in the field and had 0.8%, 2.6%, and 4.1% effect improvement on the accuracy rate.
Zhulin Tao, Xingyu Gao 0001, Xi Wang 0014, Xianglin Huang
ACM Trans. Multim. Comput. Commun. Appl.4
2025 Interventional Feature Generation for Few-shot Learning
abstract
Few-shot learning (FSL) aims to classify a novel object into a specific category under limited training samples. This is a challenging task since (1) the features expressed by pre-trained knowledge introduce perceived bias and then constrain the classification space, and (2) the use of general hallucination techniques based on global features fails to escape the limited classification space, resulting in sub-optimal improvements. To solve these issues, this article proposes an interventional feature generation (IFG) method. Specifically, we first use the relations of the categories or instances as interventional operations to implicitly constrain the feature representations (pre-trained knowledge) into different classification subsets. Then, we employ a parameter-free feature generation strategy to enrich each subset’s training samples of the support category. In other words, IFG provides a multi-subsets learning strategy to reduce the influence of perceived bias, enrich the diversity of generated features, and improve the robustness of the few-shot classifier. We apply our method to four benchmark datasets and observe state-of-the-art performance across all experiments. Specifically, compared to the baseline on the Mini-ImageNet dataset, our approach yields accuracy improvements of 6.03% and 3.46% for 1 and 5 support training samples, respectively. Furthermore, the proposed interventional feature generation technique can improve classifier performance in other FSL methods, demonstrating its versatility and potential for broader applications. The code is available at https://github.com/ShuoWangCS/IFG-FSL/ .
Shuo Wang 0008, Jinda Lu, Huixia Ben, Yanbin Hao, Xingyu Gao 0001, Meng Wang 0001
ACM Trans. Multim. Comput. Commun. Appl.5
2025 Maximizing Long-Term Task Completion Ratio of UAV-Enabled Wirelessly Powered MEC Systems
abstract
Unmanned Aerial Vehicle (UAV)-enabled wirelessly powered Mobile Edge Computing (MEC) is emerging as a powerful technology for boosting computational capability and energy supplementation in Internet of Things (IoT). This work addresses the long-term task completion ratio maximization problem in UAV-enabled wirelessly powered MEC systems. Besides the large number of optimization parameters, the environment can only be partially observed as the UAVs cannot cover the whole network area. Then, it is very challenging to obtain good solutions due to the lack of global information. We introduce a novel distributed Multi-Agent Deep Reinforcement Learning (MADRL) framework for optimizing UAVs’ actions and resource allocation, considering the constraints of tasks that vary in size, arrival times, and required computation completion time. To decouple the complicated parameters, we divide the problem into two manageable subproblems—UAVs’ action decision and resource allocation under a given UAV’s action. We employ a distributed Deep Reinforcement Learning (DRL) scheme for the former subproblem to cope with the partially observable nature. By revealing some important properties of the later subproblem, we design an efficient two-stage optimal algorithm to minimize the total consumed energy of nodes while maximizing the task-completing number. Extensive simulations validate the effectiveness of the proposed framework, achieving over a 50% improvement in task completion ratio compared to baseline schemes in some scenarios.
Shaojun Zhu, Bingcheng Zhu, Kaikai Chi, Jiefan Qiu, Hailong Shi, Xingyu Gao 0001
ACM Trans. Multim. Comput. Commun. Appl.6
2024 REDIR: Refocus-Free Event-Based De-occlusion Image Reconstruction
Hailong Shi, Jinsheng Xiao, Xingyu Gao 0001
ECCV (80)5
2024 A Coarse-to-Fine Fusion Network for Event-Based Image Deblurring
Hailong Shi, Xingyu Gao 0001
IJCAI3
2024 Unveiling Structural Memorization: Structural Membership Inference Attack for Text-to-Image Diffusion Models
abstract
With the rapid advancements of large-scale text-to-image diffusion models, various practical applications have emerged, bringing significant convenience to society. However, model developers may misuse the unauthorized data to train diffusion models. These data are at risk of being memorized by the models, thus potentially violating citizens' privacy rights. Therefore, in order to judge whether a specific image is utilized as a member of a model's training set, Membership Inference Attack (MIA) is proposed to serve as a tool for privacy protection. Current MIA methods predominantly utilize pixel-wise comparisons as distinguishing clues, considering the pixel-level memorization characteristic of diffusion models. However, it is practically impossible for text-to-image models to memorize all the pixel-level information in massive training sets. Therefore, we move to the more advanced structure-level memorization. Observations on the diffusion process show that the structures of members are better preserved compared to those of nonmembers, indicating that diffusion models possess the capability to remember the structures of member images from training sets. Drawing on these insights, we propose a simple yet effective MIA method tailored for text-to-image diffusion models. Extensive experimental results validate the efficacy of our approach. Compared to current pixel-level baselines, our approach not only achieves state-of-the-art performance but also demonstrates remarkable robustness against various distortions.
Xiaomeng Fu, Xi Wang 0014, Jin Liu 0020, Xingyu Gao 0001, Jiao Dai, Jizhong Han
ACM Multimedia5
2024 SDformer: Similarity-driven Discrete Transformer For Time Series Generation
abstract
The superior generation capabilities of Denoised Diffusion Probabilistic Models (DDPMs) have been effectively showcased across a multitude of domains. Recently, the application of DDPMs has extended to time series generation tasks, where they have significantly outperformed other deep generative models, often by a substantial margin. However, we have discovered two main challenges with these methods: 1) the inference time is excessively long; 2) there is potential for improvement in the quality of the generated time series. In this paper, we propose a method based on discrete token modeling technique called Similarity-driven Discrete Transformer (SDformer). Specifically, SDformer utilizes a similarity-driven vector quantization method for learning high-quality discrete token representations of time series, followed by a discrete Transformer for data distribution modeling at the token level. Comprehensive experiments show that our method significantly outperforms competing approaches in terms of the generated time series quality while also ensuring a short inference time. Furthermore, without requiring retraining, SDformer can be directly applied to predictive tasks and still achieve commendable results.
Shibo Feng, Zhong Zhang 0014, Xi Xiao 0001, Xingyu Gao 0001, Peilin Zhao
NeurIPS5
2024 3D Person Re-Identification Based on Global Semantic Guidance and Local Feature Aggregation
abstract
Person re-identification (Re-ID) has played an extremely crucial role in ensuring social safety and has attracted considerable research attention. 3D shape information is an important clue to understand the posture and shape of pedestrians. However, most existing person Re-ID methods learn pedestrian feature representations from images, ignoring the real 3D human body structure and the spatial relationship between the pedestrians and interferents. To address this problem, our devise a new point cloud Re-ID network (PointReIDNet), designed to obtain 3D shape representations of pedestrians from point clouds of 3D scenes. The model consists of modules, namely global semantic guidance module and local feature extraction module. The global semantic guidance module is designed by enhancing the point cloud feature representation in similar feature neighborhoods and to reduce the interference caused by 3D shape reconstruction or noise. Further, to provide an efficient representation of point clouds, we propose space cover convolution (SC-Conv), which efficiently encodes information on human shapes in local point clouds by constructing anisotropic geometries in the coordinate neighborhoods. Extensive experiments are conducted on four holistic person Re-ID datasets, one occlusion person Re-ID dataset and one point cloud classification dataset. The results exhibit significant improvements over point-cloud-based person Re-ID methods. In particular, the proposed efficient PointReIDNet decreases the number of parameters from 2.30M to 0.35M with an insignificant drop in performance. The source code is available at: https://github.com/changshuowang/PointReIDNet.
Changshuo Wang 0001, Xin Ning 0001, Weijun Li 0002, Xiao Bai 0001, Xingyu Gao 0001
IEEE Trans. Circuits Syst. Video Technol.5
2024 Multi-Faceted Knowledge-Driven Graph Neural Network for Iris Segmentation
abstract
Accurate iris segmentation, especially around the iris inner and outer boundaries, is still a formidable challenge. Pixels within these areas are difficult to semantically distinguish since they have similar visual characteristics and close spatial positions. To tackle this problem, the paper proposes an iris segmentation graph neural network (ISeGraph) for accurate segmentation. ISeGraph regards individual pixels as nodes within the graph and constructs self-adaptive edges according to multi-faceted knowledge, including visual similarity, positional correlation, and semantic consistency for feature aggregation. Specifically, visual similarity strengthens the connections between nodes sharing similar visual characteristics, while positional correlation assigns weights according to the spatial distance between nodes. In contrast to the above knowledge, semantic consistency maps nodes into a semantic space and learns pseudo-labels to define relationships based on label consistency. ISeGraph leverages multi-faceted knowledge to generate self-adaptive relationships for accurate iris segmentation. Furthermore, a pixel-wise adaptive normalization module is developed to increase the feature discriminability. It takes informative features in the shallow layer as a reference to improve the segmentation features from a statistical perspective. Experimental results on three iris datasets illustrate that the proposed method achieves superior performance in iris segmentation, increasing the segmentation accuracy in areas near the iris boundaries.
Jianze Wei, Yunlong Wang 0003, Xingyu Gao 0001, Ran He 0001, Zhenan Sun
IEEE Trans. Inf. Forensics Secur.3
2024 Knowledge Enhanced Vision and Language Model for Multi-Modal Fake News Detection
abstract
The rapid dissemination of fake news and rumors through the Internet and social media platforms poses significant challenges and raises concerns in the public sphere. Automatic detection of fake news plays a crucial role in mitigating the spread of misinformation. While recent approaches have focused on leveraging neural networks to improve textual and visual representations in multi-modal fake news analysis, they often overlook the potential of incorporating knowledge information to verify facts within news articles. In this paper, we propose a knowledge enhanced vision and language model for multi-modal fake news detection. Our proposed model integrates information from large scale open knowledge graphs to augment its ability to discern the veracity of news content. Unlike previous methods that utilize separate models to extract textual and visual features, we synthesize a unified model capable of extracting both types of features simultaneously. To represent news articles, we introduce a graph structure where nodes encompass entities, relationships extracted from the textual content, and objects depicted in associated images. By utilizing the knowledge graph, we establish meaningful relationships between nodes within the news articles. Experimental evaluations on a real-world multi-modal dataset from Twitter demonstrate significant performance improvement by incorporating knowledge information.
Xingyu Gao 0001, Xi Wang 0014, Zhenyu Chen 0003, Wei Zhou 0019, Steven C. H. Hoi
IEEE Trans. Multim.1
2024 Self-Adaptive Graph With Nonlocal Attention Network for Skeleton-Based Action Recognition
abstract
Graph convolutional networks (GCNs) have achieved encouraging progress in modeling human body skeletons as spatial-temporal graphs. However, existing methods still suffer from two inherent drawbacks. Firstly, these models process the input data based on the physical structure of the human body, which leads to some latent correlations among joints being ignored. Furthermore, the key temporal relationships between nonadjacent frames are overlooked, preventing to fully learn the changes of the body joints along the temporal dimension. To address these issues, we propose an innovative spatial-temporal model by introducing a self-adaptive GCN (SAGCN) with global attention network, collectively termed SAGGAN. Specifically, the SAGCN module is proposed to construct two additional dynamic topological graphs to learn the common characteristics of all data and represent a unique pattern for each sample, respectively. Meanwhile, the global attention module (spatial attention (SA) and temporal attention (TA) modules) is designed to extract the global connections between different joints in a single frame and model temporal relationships between adjacent and nonadjacent frames in temporal sequences. In this manner, our network can capture richer features of actions for accurate action recognition and overcome the defect of the standard graph convolution. Extensive experiments on three benchmark datasets (NTU-60, NTU-120, and Kinetics) have demonstrated the superiority of our proposed method.
Chen Pang 0001, Xingyu Gao 0001, Zhenyu Chen 0003, Lei Lyu 0001
IEEE Trans. Neural Networks Learn. Syst.2
2024 On the Effectiveness of Sampled Softmax Loss for Item Recommendation
abstract
The learning objective plays a fundamental role to build a recommender system. Most methods routinely adopt either pointwise (e.g., binary cross-entropy) or pairwise (e.g., BPR) loss to train the model parameters, while rarely pay attention to softmax loss, which assumes the probabilities of all classes sum up to 1, due to its computational complexity when scaling up to large datasets or intractability for streaming data where the complete item space is not always available. The sampled softmax (SSM) loss emerges as an efficient substitute for softmax loss. Its special case, InfoNCE loss, has been widely used in self-supervised learning and exhibited remarkable performance for contrastive learning. Nonetheless, limited recommendation work uses the SSM loss as the learning objective. Worse still, none of them explores its properties thoroughly and answers “Does SSM loss suit for item recommendation?” and “What are the conceptual advantages of SSM loss, as compared with the prevalent losses?”, to the best of our knowledge. In this work, we aim at offering a better understanding of SSM for item recommendation. Specifically, we first theoretically reveal three model-agnostic advantages: (1) mitigating popularity bias, which is beneficial to long-tail recommendation; (2) mining hard negative samples, which offers informative gradients to optimize model parameters; and (3) maximizing the ranking metric, which facilitates top- K performance. However, based on our empirical studies, we recognize that the default choice of cosine similarity function in SSM limits its ability in learning the magnitudes of representation vectors. As such, the combinations of SSM with the models that also fall short in adjusting magnitudes (e.g., matrix factorization) may result in poor representations. One step further, we provide mathematical proof that message passing schemes in graph convolution networks can adjust representation magnitude according to node degree, which naturally compensates for the shortcoming of SSM. Extensive experiments on four benchmark datasets justify our analyses, demonstrating the superiority of SSM for item recommendation. Our implementations are available in both TensorFlow 1 and PyTorch. 2
Jiancan Wu, Xiang Wang 0010, Xingyu Gao 0001, Jiawei Chen 0007, Hongcheng Fu
ACM Trans. Inf. Syst.3
2023 Mixtron: Bandit Online Multiclass Prediction with Implicit Feedback
abstract
The exploitation and exploration dilemma is a crucial issue in bandit online multiclass prediction. Conventional algorithms typically resort to either random sample or estimate uncertainty for exploration. In contrast, we propose a novel scheme that focuses solely on exploitation with implicit feedback. To ensure efficient information feedback even when predictions are incorrect, we introduce mixed losses into our proposed scheme. We derive two mixed versions of the multiclass hinge loss and the logistic loss, along with their corresponding algorithms. Context-free and context-aware experiments are conducted to evaluate the performance of our proposed algorithms. Experiment results demonstrate that random sampling are unnecessary if a reasonable loss function is employed. By compared with several state-of-the-art baselines on both synthetic and real-world datasets, our proposed algorithms show their superior performance. Remarkably, they even outperform the Perceptron algorithm with full-information feedback in some cases.
Wanjin Feng, Hailong Shi, Peilin Zhao, Xingyu Gao 0001
ICDM4
2023 Contextual Measures for Iris Recognition
abstract
The iris patterns of the human contain a large amount of randomly distributed and irregularly shaped microstructures. These microstructures make the human iris informative biometric traits. To learn identity representation from them, this paper regards each iris region as a potential microstructure and proposes contextual measures (CM) to model the correlations between them. CM adopts two parallel branches to learn global and local contexts in iris image. The first one is the globally contextual measure branch. It measures the global context involving the relationships between all regions for feature aggregation and is robust to local occlusions. Besides, we improve its spatial perception considering the positional randomness of the microstructures. The other one is the locally contextual measure branch. This branch considers the role of local details in the phenotypic distinctiveness of iris patterns and learns a series of relationship atoms to capture contextual information from a local perspective. In addition, we develop the perturbation bottleneck to make sure that the two branches learn divergent contexts. It introduces perturbation to limit the information flow from input images to identity features, forcing CM to learn discriminative contextual information for iris recognition. Experimental results suggest that global and local contexts are two different clues critical for accurate iris recognition. The superior performance on four benchmark iris datasets demonstrates the effectiveness of the proposed approach in within-database and cross-database scenarios.
Jianze Wei, Yunlong Wang 0003, Huaibo Huang, Ran He 0001, Zhenan Sun, Xingyu Gao 0001
IEEE Trans. Inf. Forensics Secur.6
2023 Discriminative Feature Mining Based on Frequency Information and Metric Learning for Face Forgery Detection
abstract
Face forgery detection has received considerable attention due to security concerns about abnormal faces generated by face forgery technology. While recent researches have made prominent progress, they still suffer from two limitations: a) the learned features supervised by softmax loss are insufficiently discriminative, since the softmax loss fails to explicitly boost inter-class separability and intra-class compactness; b) hand-crafted features are unable to effectively mine forgery patterns from frequency domain. To address the two problems, this paper proposes a novel frequency-aware discriminative feature learning framework. Specifically, we design an innovative single-center loss which compresses mere intra-class variations of natural faces while encouraging inter-class differences between natural and manipulated faces in the embedding space. Supervised by such a loss, more discriminative features can be learned with less optimization difficulty. As for frequency-related features, a frequency feature adaptively generated module is developed to capture frequency clues in a data-driven manner. Besides, to better fuse the features of both RGB domain and frequency domain, this paper devises a fusion module based on positional correlation of features. The effectiveness and superiority of our framework have been proved by extensive experiments and our approach achieves state-of-the-art performance in both in-dataset and cross-dataset evaluation.
Jiaming Li 0017, Hongtao Xie 0001, Lingyun Yu 0002, Xingyu Gao 0001, Yongdong Zhang 0001
IEEE Trans. Knowl. Data Eng.4
2023 Contrastive Multi-Level Graph Neural Networks for Session-Based Recommendation
abstract
Session-based recommendation (SBR) aims to predict the next item at a certain time point based on anonymous user behavior sequences. Existing methods typically model session representation based on simple item transition information. However, since session-based data consists of limited users' short-term interactions, modeling session representation by capturing fixed item transition information from a single dimension suffers from data sparsity. In this paper, we propose a novel contrastive multi-level graph neural networks (CM-GNN) to better exploit complex and high-order item transition information. Specifically, CM-GNN applies local-level graph convolutional network (L-GCN) and global-level graph convolutional network (G-GCN) on the current session and all the sessions respectively, to effectively capture pairwise relations over all the sessions by aggregation strategy. Meanwhile, CM-GNN applies hyper-level graph convolutional network (H-GCN) to capture high-order information among all the item transitions. CM-GNN further introduces an attention-based fusion module to learn pairwise relation-based session representation by fusing the item representations generated by L-GCN and G-GCN. CM-GNN averages the item representations obtained by H-GCN to obtain high-order relation-based session representation. Moreover, to convert the high-order item transition information into the pairwise relation-based session representation, CM-GNN maximizes the mutual information between the representations derived from the fusion module and the average pool layer by contrastive learning paradigm. We conduct extensive experiments on several widely used benchmark datasets to validate the efficacy of the proposed method. The encouraging results demonstrate that our proposed method outperforms the state-of-the-art SBR techniques.
Fuyun Wang, Xingyu Gao 0001, Zhenyu Chen 0003, Lei Lyu 0001
IEEE Trans. Multim.2
2023 Dilated Convolution-based Feature Refinement Network for Crowd Localization
abstract
As an emerging computer vision task, crowd localization has received increasing attention due to its ability to produce more accurate spatially predictions. However, continuous scale variations in complex crowd scenes lead to tiny individuals at the edges, so that existing methods cannot achieve precise crowd localization. Aiming at alleviating the above problems, we propose a novel Dilated Convolution-based Feature Refinement Network (DFRNet) to enhance the representation learning capability. Specifically, the DFRNet is built with three branches that can capture the information of each individual in crowd scenes more precisely. More specifically, we introduce a Feature Perception Module to model long-range contextual information at different scales by adopting multiple dilated convolutions, thus providing sufficient feature information to perceive tiny individuals at the edge of images. Afterwards, a Feature Refinement Module is deployed at multiple stages of the three branches to facilitate the mutual refinement of feature information at different scales, thus further improving the expression capability of multi-scale contextual information. By incorporating the above modules, DFRNet can locate individuals in complex scenes more precisely. Extensive experiments on multiple datasets demonstrate that the proposed method has more advanced performance compared to existing methods and can be more accurately adapted to complex crowd scenes.
Xingyu Gao 0001, Jinyang Xie, Zhenyu Chen 0003, Anan Liu, Zhenan Sun, Lei Lyu 0001
ACM Trans. Multim. Comput. Commun. Appl.1
2023 Learning Semantic Representation on Visual Attribute Graph for Person Re-identification and Beyond
abstract
Person re-identification (re-ID) aims to match pedestrian pairs captured from different cameras. Recently, various attribute-based models have been proposed to combine the pedestrian attribute as an auxiliary semantic information to learn a more discriminative pedestrian representation. However, these methods usually directly concatenate the visual branch and attribute branch embeddings as the final pedestrian representation, which ignores the semantic relation between the pedestrian revealed by attribute similarity. To capture and explore such semantic relation, we propose a unified pedestrian representation framework, called Visual Attribute Graph Embedding Network (VAGEN), to simultaneously learn attribute and visual representation. We unify the visual embedding and attribute similarity into a Visual Attribute Graph, where pedestrian is considered as a node and attribute similarity as an edge. Then, we learn graph node embedding to generate pedestrian representation through Graph Neural Network. Except for this unified representation for visual and attribute embeddings, VAGEN also conducts implicitly hard example mining for visual similar false-positive results, which has not been explored yet among existing attribute-based methods. We conduct extensive empirical studies on several person re-ID datasets to evaluate our proposed algorithm from different aspects. The results show that our proposed method outperforms state-of-the-art techniques with considerable margins.
Geyu Tang, Xingyu Gao 0001, Zhenyu Chen 0003
ACM Trans. Multim. Comput. Commun. Appl.2
2022 Parameterization of Cross-token Relations with Relative Positional Encoding for Vision MLP
abstract
Vision multi-layer perceptrons (MLPs) have shown promising performance in computer vision tasks, and become the main competitor of CNNs and vision Transformers. They use token-mixing layers to capture cross-token interactions, as opposed to the multi-head self-attention mechanism used by Transformers. However, the heavily parameterized token-mixing layers naturally lack mechanisms to capture local information and multi-granular non-local relations, thus their discriminative power is restrained. To tackle this issue, we propose a new positional spacial gating unit (PoSGU). It exploits the attention formulations used in the classical relative positional encoding (RPE), to efficiently encode the cross-token relations for token mixing. It can successfully reduce the current quadratic parameter complexity O(N2) of vision MLPs to $O(N)$ and O(1). We experiment with two RPE mechanisms, and further propose a group-wise extension to improve their expressive power with the accomplishment of multi-granular contexts. These then serve as the key building blocks of a new type of vision MLP, referred to as PosMLP. We evaluate the effectiveness of the proposed approach by conducting thorough experiments, demonstrating an improved or comparable performance with reduced parameter complexity. For instance, for a model trained on ImageNet1K, we achieve a performance improvement from 72.14% to 74.02% and a learnable parameter reduction from 19.4M to 18.2M. Code could be found at https://github.com/Zhicaiwww/PosMLP https://github.com/Zhicaiwww/PosMLP.
Zhicai Wang, Yanbin Hao, Xingyu Gao 0001, Hao Zhang 0047, Shuo Wang 0008, Tingting Mu, Xiangnan He 0001
ACM Multimedia3
2022 Self-Supervised Auxiliary Domain Alignment for Unsupervised 2D Image-Based 3D Shape Retrieval
abstract
Unsupervised 2D image-based 3D shape retrieval aims to match the similar 3D unlabeled shapes when given a 2D labeled sample. Although a lot of methods have made a certain degree of progress, the performance of this task is still restricted due to the lack of target labels resulting in tremendous domain gap. In this paper, we aim to explore the discriminative representation of the unlabeled target 3D shapes and facilitate the procedure of domain adaptation by taking full advantage of multi-view information. To achieve the above goals, we propose an effective self-supervised auxiliary domain alignment (SADA) for unsupervised 2D image-based 3D shape retrieval. SADA mainly contains multi-view guided self-supervised feature learning and two auxiliary domain alignments, including intermediate domain alignment and multi-domain alignment. Firstly, we group multiple views of each 3D shape into two sub-target domains based on the view similarities and regard each other as the constraint to optimize the feature learning in an unsupervised manner. To reduce the difficulty of directly aligning the domain discrepancy, we combine the source labeled samples and target samples (pseudo labels) with the same category to generate an intermediate domain, which translates the source-target alignment into source-intermediate and intermediate-target alignments. Moreover, to explore the inner characteristics of target 3D shapes and provide more clues for better adaptation, multi-domain alignment is proposed to convert the source and single target domain alignment to the source and multiple target domain (one target domain and two sub-target domains) alignments. The adversarial training and semantic alignment are employed to fully excavate the relations between source domain and multiple target domains. Experiments on two challenging datasets show that the proposed method achieves competing performance in the unsupervised 2D image-based 3D shape retrieval task.
Anan Liu, Chenyu Zhang 0003, Wenhui Li 0001, Xingyu Gao 0001, Zhengya Sun, Xuanya Li
IEEE Trans. Circuits Syst. Video Technol.4
2022 Task-Adaptive Attention for Image Captioning
abstract
Attention mechanisms are now widely used in image captioning models. However, most attention models only focus on visual features. When generating syntax related words, little visual information is needed. In this case, these attention models could mislead the word generation. In this paper, we propose Task-Adaptive Attention module for image captioning, which can alleviate this misleading problem and learn implicit non-visual clues which can be helpful for the generation of non-visual words. We further introduce a diversity regularization to enhance the expression ability of the Task-Adaptive Attention module. Extensive experiments on the MSCOCO captioning dataset demonstrate that by plugging our Task-Adaptive Attention module into a vanilla Transformer-based image captioning model, performance improvement can be achieved.
Chenggang Yan 0001, Yiming Hao, Liang Li 0003, Jian Yin 0003, Anan Liu, Zhendong Mao 0001, Zhenyu Chen 0003, Xingyu Gao 0001
IEEE Trans. Circuits Syst. Video Technol.8
2022 Long Short-Term Relation Transformer With Global Gating for Video Captioning
abstract
Video captioning aims to generate a natural language sentence to describe the main content of a video. Since there are multiple objects in videos, taking full exploration of the spatial and temporal relationships among them is crucial for this task. The previous methods wrap the detected objects as input sequences, and leverage vanilla self-attention or graph neural network to reason about visual relations. This cannot make full use of the spatial and temporal nature of a video, and suffers from the problems of redundant connections, over-smoothing, and relation ambiguity. In order to address the above problems, in this paper we construct a long short-term graph (LSTG) that simultaneously captures short-term spatial semantic relations and long-term transformation dependencies. Further, to perform relational reasoning over the LSTG, we design a global gated graph reasoning module (G3RM), which introduces a global gating based on global context to control information propagation between objects and alleviate relation ambiguity. Finally, by introducing G3RM into Transformer instead of self-attention, we propose the long short-term relation transformer (LSRT) to fully mine objects' relations for caption generation. Experiments on MSVD and MSR-VTT datasets show that the LSRT achieves superior performance compared with state-of-the-art methods. The visualization results indicate that our method alleviates problem of over-smoothing and strengthens the ability of relational reasoning.
Liang Li 0003, Xingyu Gao 0001, Jincan Deng, Yunbin Tu, Zhengjun Zha, Qingming Huang
IEEE Trans. Image Process.2
2022 Dynamic-Aware Federated Learning for Face Forgery Video Detection
abstract
The spread of face forgery videos is a serious threat to information credibility, calling for effective detection algorithms to identify them. Most existing methods have assumed a shared or centralized training set. However, in practice, data may be distributed on devices of different enterprises that cannot be centralized to share due to security and privacy restrictions. In this article, we propose a Federated Learning face forgery detection framework to train a global model collaboratively while keeping data on local devices. In order to make the detection model more robust, we propose a novel Inconsistency-Capture module (ICM) to capture the dynamic inconsistencies between adjacent frames of face forgery videos. The ICM contains two parallel branches. The first branch takes the whole face of adjacent frames as input to calculate a global inconsistency representation. The second branch focuses only on the inter-frame variation of critical regions to capture the local inconsistency. To the best of our knowledge, this is the first work to apply federated learning to face forgery video detection, which is trained with decentralized data. Extensive experiments show that the proposed framework achieves competitive performance compared with existing methods that are trained with centralized data, with higher-level security and privacy guarantee.
Ziheng Hu, Hongtao Xie 0001, Lingyun Yu 0002, Xingyu Gao 0001, Zhihua Shang, Yongdong Zhang 0001
ACM Trans. Intell. Syst. Technol.4
2021 Unsupervised adversarial domain adaptation with similarity diffusion for person re-identification
Geyu Tang, Xingyu Gao 0001, Zhenyu Chen 0003, Huicai Zhong
Neurocomputing2
2020 Parsing-Based View-Aware Embedding Network for Vehicle Re-Identification
abstract
Vehicle Re-Identification is to find images of the same vehicle from various views in the cross-camera scenario. The main challenges of this task are the large intra-instance distance caused by different views and the subtle inter-instance discrepancy caused by similar vehicles. In this paper, we propose a parsing-based view-aware embedding network (PVEN) to achieve the view-aware feature alignment and enhancement for vehicle ReID. First, we introduce a parsing network to parse a vehicle into four different views and then align the features by mask average pooling. Such alignment provides a fine-grained representation of the vehicle. Second, in order to enhance the view-aware features, we design a common-visible attention to focus on the common visible views, which not only shortens the distance among intra-instances, but also enlarges the discrepancy of inter-instances. The PVEN helps capture the stable discriminative information of vehicle under different views. The experiments conducted on three datasets show that our model outperforms state-of-the-art methods by a large margin.
Dechao Meng, Liang Li 0003, Xuejing Liu, Zhengjun Zha, Xingyu Gao 0001, Shuhui Wang, Qingming Huang
CVPR7
2020 Fine-grained Feature Alignment with Part Perspective Transformation for Vehicle ReID
abstract
Given a query image, vehicle Re-Identification is to search the same vehicle in multi-camera scenarios, which are attracting much attention in recent years. However, vehicle ReID severely suffers from the perspective variation problem. For different vehicles with similar color and type which are taken from different perspectives, all visual patterns are misaligned and warped, which is hard for the model to find out the exact discriminative regions. In this paper, we propose part perspective transformation module (PPT) to map the different parts of vehicle into a unified perspective respectively. The PPT disentangles the vehicle features of different perspectives and then aligns them in a fine-grained level. Further, we propose a dynamically batch hard triplet loss to select the common visible regions of the compared vehicles. Our approach helps the model to generate the perspective invariant features and find out the exact distinguishable regions for vehicle ReID. Extensive experiments on three standard vehicle ReID datasets show the effectiveness of our method.
Dechao Meng, Liang Li 0003, Shuhui Wang, Xingyu Gao 0001, Zhengjun Zha, Qingming Huang
ACM Multimedia4
2020 Learning salient features to prevent model drift for correlation tracking
Yu Zhang 0102, Xingyu Gao 0001, Zhenyu Chen 0003, Huicai Zhong, Liang Li 0003, Chenggang Yan 0001, Tao Shen 0004
Neurocomputing2
2020 Advancing Image Understanding in Poor Visibility Environments: A Collective Benchmark Study
abstract
Existing enhancement methods are empirically expected to help the high-level end computer vision task: however, that is observed to not always be the case in practice. We focus on object or face detection in poor visibility enhancements caused by bad weathers (haze, rain) and low light conditions. To provide a more thorough examination and fair comparison, we introduce three benchmark sets collected in real-world hazy, rainy, and low-light conditions, respectively, with annotated objects/faces. We launched the UG2+ challenge Track 2 competition in IEEE CVPR 2019, aiming to evoke a comprehensive discussion and exploration about whether and how low-level vision techniques can benefit the high-level automatic visual recognition in various scenarios. To our best knowledge, this is the first and currently largest effort of its kind. Baseline results by cascading existing enhancement and detection models are reported, indicating the highly challenging nature of our new data as well as the large room for further technical innovations. Thanks to a large participation from the research community, we are able to analyze representative team solutions, striving to better identify the strengths and limitations of existing mindsets as well as the future directions.
Wenhan Yang, Ye Yuan 0012, Wenqi Ren, Jiaying Liu 0001, Walter J. Scheirer, Zhangyang Wang, Taiheng Zhang, Qiaoyong Zhong, Di Xie, Shiliang Pu, Yuqiang Zheng, Yanyun Qu, Yuhong Xie, Hao Jiang 0014, Siyuan Yang 0001, Yan Liu 0041, Xiaochao Qu, Pengfei Wan 0001, Shuai Zheng 0005, Minhui Zhong, Taiyi Su, Lingzhi He, Yandong Guo, Yao Zhao 0001, Zhenfeng Zhu, Jinxiu Liang, Jingwen Wang 0003, Yuhui Quan, Yong Xu 0007, Bo Liu 0112, Xin Liu 0012, Tingyu Lin 0003, Xiaochuan Li 0001, Feng Lu 0005, Lin Gu 0003, Shengdi Zhou, Cong Cao 0005, Cheng Chi 0003, Chubin Zhuang, Zhen Lei 0001, Stan Z. Li, Shizheng Wang, Ruizhe Liu, Dong Yi, Zheming Zuo, Jianning Chi, Huan Wang 0014, Kai Wang 0036, Yixiu Liu, Xingyu Gao 0001, Zhenyu Chen 0003, Yongzhou Li, Huicai Zhong, Jing Huang 0017, Heng Guo 0003, Jianfei Yang 0001, Wenjuan Liao, Jiangang Yang, Liguo Zhou, Mingyue Feng, Likun Qin
IEEE Trans. Image Process.56
2020 Mining Spatial-Temporal Similarity for Visual Tracking
abstract
Correlation filter (CF) is a critical technique to improve accuracy and speed in the field of visual object tracking. Despite being studied extensively, most existing CF methods suffer from failing to make the most of the inherent spatial-temporal prior of videos. To address this limitation, as consecutive frames are eminently resemble in most videos, we investigate a novel scheme to predict targets' future state by exploiting previous observations. Specifically, in this paper, we propose a prediction based CF tracking framework by learning the spatial-temporal similarity of consecutive frames for sample managing, template regularization, and training response pre-weighting. We model the learning problem theoretically as a novel objective and provide effective optimization algorithms to solve the learning task. In addition, we implement two CF trackers with different features. Extensive experiments are conducted on three popular benchmarks to validate our scheme. The encouraging results demonstrate that the proposed scheme can significantly boost the accuracy of CF tracking, and the two trackers achieve competitive performances against state-of-the-art trackers. We finally present a comprehensive analysis on the efficacy of our proposed method and the efficiency of our trackers to facilitate real-world visual tracking applications.
Yu Zhang 0102, Xingyu Gao 0001, Zhenyu Chen 0003, Huicai Zhong, Hongtao Xie 0001, Chenggang Yan 0001
IEEE Trans. Image Process.2
2020 Spatiotemporal Recurrent Convolutional Networks for Recognizing Spontaneous Micro-Expressions
abstract
Recently, the recognition task of spontaneous facial micro-expressions has attracted much attention with its various real-world applications. Plenty of handcrafted or learned features have been employed for a variety of classifiers and achieved promising performances for recognizing micro-expressions. However, the micro-expression recognition is still challenging due to the subtle spatiotemporal changes of micro-expressions. To exploit the merits of deep learning, we propose a novel deep recurrent convolutional networks based micro-expression recognition approach, capturing the spatiotemporal deformations of micro-expression sequence. Specifically, the proposed deep model is constituted of several recurrent convolutional layers for extracting visual features and a classificatory layer for recognition. It is optimized by an end-to-end manner and obviates manual feature design. To handle sequential data, we exploit two ways to extend the connectivity of convolutional networks across temporal domain, in which the spatiotemporal deformations are modeled in views of facial appearance and geometry separately. Besides, to overcome the shortcomings of limited and imbalanced training samples, two temporal data augmentation strategies as well as a balanced loss are jointly used for our deep network. By performing the experiments on three spontaneous micro-expression datasets, we verify the effectiveness of our proposed micro-expression recognition approach compared to the state-of-the-art methods.
Zhaoqiang Xia, Xiaopeng Hong, Xingyu Gao 0001, Xiaoyi Feng, Guoying Zhao 0001
IEEE Trans. Multim.3
2020 Corrections to "Spatiotemporal Recurrent Convolutional Networks for Recognizing Spontaneous Micro-Expressions"
abstract
Presents corrections to the author's information in the above named paper.
Zhaoqiang Xia, Xiaopeng Hong, Xingyu Gao 0001, Xiaoyi Feng, Guoying Zhao 0001
IEEE Trans. Multim.3
2019 Learning Correlation Filter With Detection Response For Visual Tracking
abstract
Correlation Filter (CF) has been a powerful tool for real-time visual object tracking. Although enormous improvements have been made since the first CF tracker was proposed, most of the existing CF tracking schemes adopt a fixed Gaussian label to train correlation filter. We argue that it is not optimal for challenging scenarios. Additionally, existing training label managing strategies are too complex to maintain the speed advantage of CF algorithms. In this work, we propose a novel label to supervise filter training. Our method is capable of being integrated into most CF trackers conveniently without increasing computational complexity. We employ the proposed scheme for four existing CF trackers and conduct extensive experiments on popular benchmark. The encouraging experimental results validate both the effi-ciency and efficacy of our proposed method. What's more, we provide a detailed analysis for challenging applications and hyperparameter setting.
Yu Zhang 0102, Xingyu Gao 0001, Zhenyu Chen 0003, Huicai Zhong
ICIP2
2019 Multimodel Framework for Indoor Localization Under Mobile Edge Computing Environment
abstract
Location estimation technology under the wireless environment has become a vital technology in the field of mobile edge computing. Especially, under the mobile edge of entire networks environment, indoor location estimation is gradually getting the interest research and application topic, due to technical constraints of global positioning system technology for indoor environment and the popularity of the mobile edge computing servers. In this paper, the widely used single-model framework for indoor localization is presented as an introduction, which consists of three stages: 1) sample data collection; 2) model building; and 3) localization estimation. And then, through analyzing of the actual scene of indoor localization, a new framework for indoor localization under mobile edge computing environment, named Multimodel, is proposed from the theoretical perspective. It is mainly based on the observation that the environment of the sample data collection and that of localization data collection may change seriously. In order to make up for the shortcomings of this framework, two combinatorial optimization problems are proposed. Later, we discuss the NP-hardness of them in several different cases. In addition, two heuristic algorithms are given, and the performance of which are illustrated by the corresponding experimental results.
Wenjun Li 0001, Zhenyu Chen 0003, Xingyu Gao 0001, Wei Liu 0010, Jin Wang 0001
IEEE Internet Things J.3
2017 Detecting Uyghur text in complex background images with convolutional neural network
Shancheng Fang, Hongtao Xie 0001, Zhineng Chen, Shiai Zhu, Xiaoyan Gu 0001, Xingyu Gao 0001
Multim. Tools Appl.6
2017 Robust and parallel Uyghur text localization in complex background images
Yun Song, Hongtao Xie 0001, Zhineng Chen, Xingyu Gao 0001
Mach. Vis. Appl.5
2017 Sparse Online Learning of Image Similarity
abstract
Learning image similarity plays a critical role in real-world multimedia information retrieval applications, especially in Content-Based Image Retrieval (CBIR) tasks, in which an accurate retrieval of visually similar objects largely relies on an effective image similarity function. Crafting a good similarity function is very challenging because visual contents of images are often represented as feature vectors in high-dimensional spaces, for example, via bag-of-words (BoW) representations, and traditional rigid similarity functions, for example, cosine similarity, are often suboptimal for CBIR tasks. In this article, we address this fundamental problem, that is, learning to optimize image similarity with sparse and high-dimensional representations from large-scale training data, and propose a novel scheme of Sparse Online Learning of Image Similarity (SOLIS). In contrast to many existing image-similarity learning algorithms that are designed to work with low-dimensional data, SOLIS is able to learn image similarity from large-scale image data in sparse and high-dimensional spaces. Our encouraging results showed that the proposed new technique achieves highly competitive accuracy as compared to the state-of-the-art approaches but enjoys significant advantages in computational efficiency, model sparsity, and retrieval scalability, making it more practical for real-world multimedia retrieval applications.
Xingyu Gao 0001, Steven C. H. Hoi, Yongdong Zhang 0001, Jianshe Zhou, Ji Wan, Zhenyu Chen 0003, Jintao Li 0001, Jianke Zhu
ACM Trans. Intell. Syst. Technol.1
2016 Adaptive weighted imbalance learning with application to abnormal activity recognition
Xingyu Gao 0001, Zhenyu Chen 0003, Sheng Tang, Yongdong Zhang 0001, Jintao Li 0001
Neurocomputing1
2016 Adaptive incremental learning of image semantics with application to social robot
Hong Zhang 0022, Aryel Beck, Zhijun Zhang 0003, Xingyu Gao 0001
Neurocomputing5
2016 A cross-media distance metric learning framework based on multi-view correlation mining and matching
Hong Zhang 0022, Xingyu Gao 0001, Xin Xu 0007
World Wide Web2
2015 Unobtrusive Sensing Incremental Social Contexts Using Fuzzy Class Incremental Learning
abstract
By utilizing captured characteristics of surrounding contexts through widely used Bluetooth sensor, user-centric social contexts can be effectively sensed and discovered by dynamic Bluetooth information. At present, state-of-the-art approaches for building classifiers can basically recognize limited classes trained in the learning phase; however, due to the complex diversity of social contextual behavior, the built classifier seldom deals with newly appeared contexts, which results in degrading the recognition performance greatly. To address this problem, we propose, an OSELM (online sequential extreme learning machine) based class incremental learning method for continuous and unobtrusive sensing new classes of social contexts from dynamic Bluetooth data alone. We integrate fuzzy clustering technique and OSELM to discover and recognize social contextual behaviors by real-world Bluetooth sensor data. Experimental results show that our method can automatically cope with incremental classes of social contexts that appear unpredictably in the real-world. Further, our proposed method have the effective recognition capability for both original known classes and newly appeared unknown classes, respectively.
Zhenyu Chen 0003, Yiqiang Chen 0001, Xingyu Gao 0001, Shuangquan Wang, Lisha Hu, Chenggang Yan 0001, Nicholas D. Lane, Chunyan Miao
ICDM3
2015 Online Learning to Rank for Content-Based Image Retrieval
Ji Wan, Steven C. H. Hoi, Peilin Zhao, Xingyu Gao 0001, Yongdong Zhang 0001, Jintao Li 0001
IJCAI5
2014 SOML: Sparse Online Metric Learning with Application to Image Retrieval
abstract
Image similarity search plays a key role in many multimediaapplications, where multimedia data (such as images and videos) areusually represented in high-dimensional feature space. In thispaper, we propose a novel Sparse Online Metric Learning (SOML)scheme for learning sparse distance functions from large-scalehigh-dimensional data and explore its application to imageretrieval. In contrast to many existing distance metric learningalgorithms that are often designed for low-dimensional data, theproposed algorithms are able to learn sparse distance metrics fromhigh-dimensional data in an efficient and scalable manner. Ourexperimental results show that the proposed method achieves betteror at least comparable accuracy performance than thestate-of-the-art non-sparse distance metric learning approaches, butenjoys a significant advantage in computational efficiency andsparsity, making it more practical for real-world applications.
Xingyu Gao 0001, Steven C. H. Hoi, Yongdong Zhang 0001, Ji Wan, Jintao Li 0001
AAAI1
2014 Boosting cross-media retrieval via visual-auditory feature analysis and relevance feedback
abstract
Different types of multimedia data express high-level semantics from different aspects. How to learn comprehensive high-level semantics from different types of data and enable efficient cross-media retrieval becomes an emerging hot issue. There are abundant statistical and semantic correlations among heterogeneous low-level media content, which makes it challenging to query cross-media data effectively. In this paper, we propose a new cross-media retrieval method based on short-term and long-term relevance feedback. Our method mainly focuses on two typical types of media data, i.e. image and audio. First, we build multimodal representation via statistical canonical correlation between image and audio feature matrices, and define cross-media distance metric for similarity measure; then we propose optimization strategy based on relevance feedback, which fuses short-term learning results and long-term accumulated knowledge into the objective function. Experiments on image-audio dataset have demonstrated the superiority of our method over several existing algorithms.
Hong Zhang 0022, Junsong Yuan 0001, Xingyu Gao 0001, Zhenyu Chen 0003
ACM Multimedia3
2014 A Unified Geolocation Framework for Web Videos
abstract
In this article, we propose a unified geolocation framework to automatically determine where on the earth a web video was shot. We analyze different social, visual, and textual relationships from a real-world dataset and find four relationships with apparent geography clues that can be used for web video geolocation. Then, the geolocation process is formulated as an optimization problem that simultaneously takes the social, visual, and textual relationships into consideration. The optimization problem is solved by an iterative procedure, which can be interpreted as a propagation of the geography information among the web video social network. Extensive experiments on a real-world dataset clearly demonstrate the effectiveness of our proposed framework, with the geolocation accuracy higher than state-of-the-art approaches.
Yicheng Song, Yongdong Zhang 0001, Juan Cao 0001, Jinhui Tang 0001, Xingyu Gao 0001, Jintao Li 0001
ACM Trans. Intell. Syst. Technol.5
2013 GeSoDeck: a geo-social event detection and tracking system
abstract
This demonstration presents a novel geo-social event detection and tracking system based on geographical pattern mining and content analysis, called "GeSoDeck". A user can capture what events happened by our system. Unlike most existing social event detection applications, GeSoDeck aims to detect events with high accuracy and efficiency, and track them as well. Given a geographical area, the system can not only detect diverse social events in this area using the geographical pattern mining and density-based K-means clustering, but also track the representative tweets of the detected event in real time, mining geographical diffusion trajectory on the map and temporal pattern of retweeting process. On a realistic dataset collected from Sina Weibo, the system can outperform the state-of-the-art methods.
Xingyu Gao 0001, Juan Cao 0001, Zhiwei Jin, Xin Li 0118, Jintao Li 0001
ACM Multimedia1
2010 Automatic Generation of Pencil Sketch for 2D Images
Xingyu Gao 0001, Jingye Zhou, Zhenyu Chen 0003, Yiqiang Chen 0001
ICASSP1
2009 Interactive 3D caricature generation based on double sampling
abstract
Recently, 3D caricature generation and applications have attracted wide attention from both the research community and the entertainment industry. This paper proposes a novel interactive approach for various and interesting 3D caricature generation based on double sampling. Firstly, according to user's operation, we obtain a coarse 3D caricature with local features transformation by sampling in well-built principle component analysis (PCA) subspace. Secondly, to utilize information of the 2D caricature dataset, we sample in the local linear embedding (LLE) manifold subspace. Finally, we use the learned 2D caricature information to further refine the coarse caricature by applying Kriging interpolation. The experiments show that the 3D caricature generated by our method can preserve highly artistic styles and also reflect the user's intention.
Jinjing Xie, Yiqiang Chen 0001, Junfa Liu, Chunyan Miao, Xingyu Gao 0001
ACM Multimedia5
2009 Semi-supervised Learning of Caricature Pattern from Manifold Regularization
Junfa Liu, Yiqiang Chen 0001, Jinjing Xie, Xingyu Gao 0001, Wen Gao 0001
MMM4
2009 Semi-Supervised Learning in Reconstructed Manifold Space for 3D Caricature Generation
abstract
Abstract Recently, automatic 3D caricature generation has attracted much attention from both the research community and the game industry. Machine learning has been proven effective in the automatic generation of caricatures. However, the lack of 3D caricature samples makes it challenging to train a good model. This paper addresses this problem by two steps. First, the training set is enlarged by reconstructing 3D caricatures. We reconstruct 3D caricatures based on some 2D caricature samples with a Principal Component Analysis (PCA)‐based method. Secondly, between the 2D real faces and the enlarged 3D caricatures, a regressive model is learnt by the semi‐supervised manifold regularization (MR) method. We then predict 3D caricatures for 2D real faces with the learnt model. The experiments show that our novel approach synthesizes the 3D caricature more effectively than traditional methods. Moreover, our system has been applied successfully in a massive multi‐user educational game to provide human‐like avatars.
Junfa Liu, Yiqiang Chen 0001, Chunyan Miao, Jinjing Xie, Charles Ling 0001, Xingyu Gao 0001, Wen Gao 0001
Comput. Graph. Forum6