VLDB 2026 Research / reviewers in the wild / expert
Jibin Yang
dblp:157/9096
· DBLP profile ↗
16ranked-venue papers
0as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 since 2021Security and privacy · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Non-Iterative Reversible Information Hiding in the Sharing Domain With Adjustable Capacity
Kaili Qi, Jibin Yang, Tieyong Cao, Xiangli Xiao, Xiongwei Zhang, Yushu Zhang 0001, Zehang Wang |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2025 | Uncertainty-aware focal loss for object segmentation
Lei Chen 0086, Yang Wang 0015, Jibin Yang, Tong Han, Tieyong Cao |
Eng. Appl. Artif. Intell. | 3 |
| 2025 | Proton Exchange Membrane Fuel Cells Lifetime Estimation Using GCN-GRU: Simulation and Prospective Tramway ApplicationsabstractProton Exchange Membrane Fuel Cells (PEMFC) are a high-efficiency, clean energy source with significant potential for intelligent transportation systems, such as tramways. However, accurately predicting the lifespan of these fuel cells remains a significant challenge, critical for ensuring reliable and continuous tramway operation. This paper proposes an innovative hybrid neural network model combining Graph Convolutional Networks (GCN) and Gated Recurrent Units (GRU) for precise Remaining Useful Life (RUL) prediction of PEMFC. The model employs a graph learning layer to capture inter-node relationships from PEMFC data, constructing an asymmetric adjacency matrix that reflects the system’s internal directional dependencies. It then utilizes a mix-hop propagation layer to integrate time-series data, effectively capturing the dynamic behaviors and performance variations of PEMFC. Experimental results show that this model outperforms traditional GRU and CNN-GRU models in both short-term and long-term predictions, providing more accurate RUL estimations. The model utilizes high-fidelity test data from a France laboratory to improve the accuracy of fuel cell lifetime predictions, and the potential applications and experimental validation of the model in future intelligent transportation systems such as trams are discussed. This innovative approach provides a robust framework for predictive maintenance, and provides reliable data support and optimization schemes for practical applications in tramways, enhancing the reliability and efficiency of intelligent transportation systems. Jinling Ma, Jiye Zhang 0001, Jibin Yang |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2025 | Introducing Euclidean distance optimization into Softmax loss under neural collapse
Qiang Zhang 0054, Xiongwei Zhang, Jibin Yang, Meng Sun 0001, Tieyong Cao |
Pattern Recognit. | 3 |
| 2025 | Self-distillation salient object detection via generalized diversity loss
Jibin Yang, Haijun Tao, Lei Chen 0086, Tieyong Cao |
Pattern Recognit. | 2 |
| 2025 | A Soft-Contrastive Pseudo Learning Approach Toward Open-World Forged Speech AttributionabstractAnti-spoofing of deepfake or forged speech is an important technique for the security usage of generative artificial intelligence. Beyond binary classification of real and forged speech, method attribution of forged speech is becoming a practical solution of interpretable anti-spoofing strategies. However, existing related methods have poor performance on analyzing speech forgery methods unseen in their training data, which is inefficient in open-world scenarios with emerging new forgery methods. In this paper, Open-World Forged Speech Attribution (OW-FSA) is firstly defined towards the attribution of forged speech on the methods generating it, where the recognized methods are not limited to the seen ones in training data and the properties of the unseen methods should also be depicted adequately. A novel algorithm, Soft-contrastive Pseudo Learning (SPL), is proposed to address the challenges outlined in OW-FSA, which introduces two key innovations: 1) Based on similarities between features at different scales, the proposed similarity-based soft filtering module filters and matches utterances from the same forgery class to enhance the intra-class compactness of features through contrastive learning. 2) The proposed similarity-based soft pseudo-labeling module integrates label-smoothing-like and similarity weighting techniques to mitigate possible errors in pseudo-labeling. Besides, an iterative algorithm based on SPL is proposed to predict the number of unseen classes. Extensive experiments have validated the superiority of the proposed algorithm over other recently proposed methods on the task of OW-FSA with or without the knowledge of the number of unseen classes. Intuitive visualization and ablation studies have also been conducted to illustrate the advantages of the proposed algorithm. The newly defined task OW-FSA and the proposed algorithm SPL in this paper will help advance the research in speech anti-spoofing. Qiang Zhang 0054, Xiongwei Zhang, Meng Sun 0001, Jibin Yang |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2024 | An efficient low-perceptual environmental sound classification adversarial method based on GAN
Jibin Yang, Xiongwei Zhang, Tieyong Cao |
Multim. Tools Appl. | 2 |
| 2023 | Effect of Impulses on Robust Exponential Stability of Delayed Quaternion-Valued Neural Networks
Xiaohui Xu, Jibin Yang, Shulei Sun |
Neural Process. Lett. | 2 |
| 2022 | SO-softmax loss for discriminable embedding learning in CNNs
Qiang Zhang 0054, Jibin Yang, Xiongwei Zhang, Tieyong Cao |
Pattern Recognit. | 2 |
| 2020 | Further research on exponential stability for quaternion-valued neural networks with mixed delays
Xiaohui Xu, Quan Xu 0002, Jibin Yang, Huanbin Xue, Yanhai Xu |
Neurocomputing | 3 |
| 2016 | Adaptive extraction of repeating non-negative temporal patterns for single-channel speech enhancementabstractEstimating unknown background noise from single-channel noisy speech is a key yet challenging problem for speech enhancement. Given the fact that the background noises typically have the repeating property and the foreground speech is sparse and time-variant, many literatures decompose the noisy spectrogram directly in an unsupervised fashion when there is no isolated training example of the target speaker or particular noise types beforehand. However, recently proposed methods suffer from un-interpretable decomposed patterns, neglecting the temporal structure of the background noise or being constrained by the pre-fixed parameters. To settle these issues, we propose a novel method based on autocorrelation technique and convolutive non-negative matrix factorization. The proposed method can adaptively estimate the underlying non-negative repeating temporal patterns from noisy speech and identify the clean speech spectrogram simultaneously. Experiments on NOIZEUS dataset mixed with various real-world background noises showed that the proposed method performs better than some state-of-the-art methods. Yinan Li 0006, Xiongwei Zhang, Meng Sun 0001, Gang Min, Jibin Yang |
ICASSP | 5 |
| 2016 | Joint optimization of audible noise suppression and deep neural networks for single-channel speech enhancementabstractImproving the perceptual quality of speech signals is a key yet challenging problem for many real world applications. Taking into account the good performance of deep learning in signal representation, a novel single-channel speech enhancement technique is presented based on joint Deep Neural Networks and audible noise suppression as a whole network architecture. This new deep neural network jointly trains an audible noise suppression function which is used to estimate the magnitude spectrum of the clean speech and shape the spectrum of the audible noise at the same time. Experimental results on TIMIT with 20 noise types at various noise levels demonstrate the superiority of the proposed method over the baselines, no matter whether the noise conditions are included in the training set or not. Xiongwei Zhang, Gang Min, Meng Sun 0001, Jibin Yang |
ICME | 5 |
| 2016 | A perceptually motivated approach via sparse and low-rank model for speech enhancementabstractA perceptually motivated speech enhancement approach is proposed in this paper. Different from the conventional sparse and low-rank model based approaches, this new approach takes into account the perceptual differences in different frequency bands of the human auditory system, and separates speech from background noises in the Mel spectral domain. After two propositions for the Mel frequency weighted spectrogram are proved, speech enhancement can be modeled as a sparse and low-rank constrained optimization problem, which is solved efficiently by the alternating direction method of multipliers (ADMM). The proposed approach is totally unsupervised, neither the speech nor the noise dictionary needs to be trained beforehand. The experimental results have shown its promising performance under strong background noises. The performance can be further improved by information fusion technique at high input SNRs. Gang Min, Xiongwei Zhang, Jibin Yang, Xia Zou |
ICME | 3 |
| 2016 | Perceptually Weighted Analysis-by-Synthesis Vector Quantization for Low Bit Rate MFCC CodecabstractThis letter presents a perceptually weighted analysis-by-synthesis vector quantization (VQ) algorithm for low bit rate MFCC codec. Different from conventional VQ of mel-frequency cepstral coefficients (MFCCs) vector, this algorithm uses an analysis-by-synthesis technique and aims to minimize the perceptually weighted spectral reconstruction distortion rather than the distortion of MFCCs vector itself. Also, to reduce the computational complexity, we propose a practical suboptimal codebook searching technique and embed it into the split and multistage VQ framework. Objective and subjective experimental results on Mandarin speech show that the proposed algorithm yields intelligible and natural sounding speech for speech coding at 600-2400 bit/s. Compared to current VQ in MFCC codec, the output speech quality is substantially improved in terms of frequency-weighted segmental SNR, short-time objective intelligibility score, perceptual evaluation of speech quality score, and mean opinion score. Gang Min, Xiongwei Zhang, Xia Zou, Jibin Yang |
IEEE Signal Process. Lett. | 4 |
| 2015 | SegBOMP: An efficient algorithm for block non-sparse signal recoveryabstractBlock sparse signal recovery methods have attracted great interests which take the block structure of the nonzero coefficients into account when clustering. Compared with traditional compressive sensing methods, it can obtain better recovery performance with fewer measurements by utilizing the block-sparsity explicitly. In this paper we propose a segmented-version of the block orthogonal matching pursuit algorithm in which it divides any vector into several sparse sub-vectors. By doing this, the original method can be significantly accelerated due to the dimension reduction of measurements for each segmented vector. Experimental results showed that with low complexity the proposed method yielded identical or even better reconstruction performance than the conventional methods which treated the signal in the standard block-sparsity fashion. Furthermore, in the specific case, where not all segments contain nonzero blocks, the performance improvement can be interpreted as a gain in “effective SNR” in noisy environment. Xushan Chen, Xiongwei Zhang, Jibin Yang, Meng Sun 0001 |
ICME | 3 |
| 2015 | Speech reconstruction from mel-frequency cepstral coefficients via ℓ1-norm minimizationabstractThis paper presents a high quality speech reconstruction method from Mel-frequency cepstral coefficients (MFCC). Due to the sparse characteristic of the power spectrum of speech, the ℓ1-norm minimization method is used to tackle the under-determined nature of the speech reconstruction problem. The phase spectrum is recovered by the well-known LSE-ISTFTM algorithm. Experimental results demonstrate that the quality of the reconstructed speech is dramatically improved than the common ℓ2-norm minimization method, it sounds very close to the original speech when using the high-resolution MFCC, the PESQ score reaches 4.0. Gang Min, Xiongwei Zhang, Jibin Yang, Xia Zou |
MMSP | 3 |