EDBT 2026 Demo / reviewers in the wild / expert
Ziqiang Shi
dblp:91/8676
· DBLP profile ↗
43ranked-venue papers
32as first author
17since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 26 · 20 first-author · 12 since 2021Artificial intelligence and machine learning · 23 · 20 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-authorHuman-computer interaction and ubiquitous computing · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Eliminating Object Hallucination in MLLMs via Convex Potential Flow Intervention
Ziqiang Shi, Rujie Liu, Koichi Shirahata |
ICPR (9) | 1 |
| 2026 | Scalpel: Fine-Grained Alignment of Attention Activation Manifolds via Mixture Gaussian Bridges to Mitigate Multimodal HallucinationabstractRapid progress in large vision-language models (LVLMs) has achieved unprecedented performance in vision-language tasks. However, due to the strong prior of large language models (LLMs) and misaligned attention across modalities, LVLMs often generate outputs inconsistent with visual content - termed hallucination. To address this, we propose Scalpel, a method that reduces hallucination by refining attention activation distributions toward more credible regions. Scalpel predicts trusted attention directions for each head in Transformer layers during inference and adjusts activations accordingly. It employs a Gaussian mixture model to capture multi-peak distributions of attention in trust and hallucination manifolds, and uses entropic optimal transport (equivalent to Schrödinger bridge problem) to map Gaussian components precisely. During mitigation, Scalpel dynamically adjusts intervention strength and direction based on component membership and mapping relationships between hallucination and trust activations. Extensive experiments across multiple datasets and benchmarks demonstrate that Scalpel effectively mitigates hallucinations, outperforming previous methods and achieving state-of-the-art performance. Moreover, Scalpel is model-and data-agnostic, requiring no additional computation, only a single decoding step. Ziqiang Shi, Rujie Liu, Satoshi Munakata, Koichi Shirahata |
WACV | 1 |
| 2026 | Physics-informed neural networks for constitutive modeling and multiphysics coupling in viscoelastic materials: Applications to asphalt pavement mechanics
Li'an Shen, Ziqiang Shi |
Neural Networks | 5 |
| 2025 | Attribute Conditional Diffusion-Augmented Person Re-IdentificationabstractDue to privacy and cost issues, the lack of large-scale labeled datasets limits the advancement of person re-identification. Existing methods use generative adversarial networks or game engine rendering for data augmentation to improve re-identification performance. However, these approaches struggle to maintain realistic images. This paper introduces a novel approach called Identity Diffuser, which uses diffusion models to generate synthetic data for the same identity with different poses. Our proposed framework incorporates identity-specific embeddings and target poses into the diffusion process, enabling the generation of realistic and diverse images that consistently preserve identity features. Guided by pretrained re-identification net and target pose heatmap, the framework learns transformation trajectories through forward and backward denoising steps in the diffusion models. This approach effectively maintains key pedestrian attributes across various poses. Experimental results on the Market1501 and DukeMTMC datasets demonstrate a notable improvement in performance, with a 1.73%/0.80% mAp increase in Market1501/DukeMTMC datasets compared with current state-of-the-art method. When less real data is included, the increment can be 5.1%/1.5%, separately. Shijie Nie, Ziqiang Shi, Rujie Liu, Meng Zhang 0042, Mengjiao Wang 0001, Kazuki Osamura, Lina Septiana, Narishige Abe |
ICASSP | 2 |
| 2025 | TrueCount: Improving Open-World Object Counting with Visual-Language Models and Dynamic Multi-Modal Inputs
Ziqiang Shi, Rujie Liu |
ACM Multimedia | 1 |
| 2025 | Selective-SAM: Memory Optimization for Segment Anything Model 2 with Application in Self-Checkout Product CountingabstractThe Segment Anything Model 2 (SAM 2) [1] demonstrates strong capabilities for video object segmentation (VOS) [2]. We present a SAM 2-based framework for self-checkout product counting, where box prompts are generated by a detector selecting optimal tracking initiation frames. Our key contribution is Selective-SAM, a training-free enhancement that improves robustness in complex scenes via a selective memory bank. This mechanism selectively preserves high-quality and diverse features while filtering poor segmentation priors. For final counting, we introduce a novel Mask Overlap Degree metric to analyze object trajectories. By segmenting trajectories based on the mask overlap degrees, we accurately determine the product count. Experiments across various retail scenarios show improvements of 1.95% IDF1 and 3.86% MOTA. Zhongling Liu, Liu Liu 0020, Ziqiang Shi, Rujie Liu |
SMC | 3 |
| 2025 | AugCount: Test-Time Semantic Augmentation via Diffusion for General Open-World Object CountingabstractOpen-world general object counting is a critical task in computer vision, with applications in image understanding, environmental monitoring, and surveillance. Traditional methods rely on large-scale annotated data, which is expensive and time-consuming to obtain, and often fail to manage diverse object categories and complex scenes. To address this, we propose AugCount, a novel framework that enhances open-world object counting by generating high-quality, diverse synthetic data during testing. For the first time, we employ a diffusion model to produce conditional images based on density maps specifying object locations for general object counting. Our framework is highly versatile and adaptable to various counting tasks. Experiments demonstrate that AugCount significantly improves performance on benchmark datasets like FSC-147, reducing the average counting error by 1 per image and achieving a state-of-the-art MAE of 4.70. AugCount effectively addresses data scarcity and model generalization challenges, offering enhanced robustness and adaptability for practical counting systems. Ziqiang Shi, Rujie Liu |
SMC | 1 |
| 2025 | Bayesian Optimal Latent Projection for Noisy Image RestorationabstractIn recent years, image restoration using large-scale la-tent diffusion generative models (DGM) has attracted in-creasing attention and achieved significant progress. Most of these latent DGM-based image restoration methods re-quire predicting the original clean image in each iteration, which is then used to estimate the image for the next iter-ation. However, these predicted original clean images are often inaccurate, leading to errors in the subsequent im-age estimation. In other words, there is a significant devi-ation between the final sampling restoration trajectory and the ground truth trajectory. The purpose of this paper is to narrow the gap between these two trajectories and en-hance the performance of image restoration. We propose the Bayesian Optimal Latent Projection (BOLP) algorithm, which identifies the optimal random noise within the Gaus-sian distribution to iteratively correct the estimated image at each step, thereby minimizing the distance to the ground truth image. Experiments in deblurring, super-resolution, and inpainting on FFHQ and ImageNet datasets demon-strate that the BOLP outperforms the previously established best algorithms and sets a new state of the art. Ziqiang Shi, Rujie Liu, Takuma Yamamoto |
WACV | 1 |
| 2024 | Noisy Image Restoration Based on Conditional Acceleration Score ApproximationabstractIn recent years, score-based generative models (SGM) have achieved state-of-the-art (SOTA) performance in noisy image restoration [1], [2]. But at present, most of these methods are performed in the position space, and there is a lack in modeling of the velocity and acceleration of the image on the restoration path. In this paper, we propose a new image restoration method called conditional acceleration score approximation (CASA), which introduces velocity and acceleration variables on top of the data position along the recovery path. Guided by the degraded image, CASA can effectively and dynamically control the direction and speed of motion along the diffusion path in the reverse-time stochastic differential equation. Therefore, the key to this process is how to inject the degraded image as a guidance into the third-order reverse-time process in this position-velocity-acceleration space, especially in the evolution direction of the diffusion path. We propose a strategy for approximating the conditional acceleration score by decomposing the true posterior CAS into a priori CAS and an observed acceleration score for the measurement at the current moment. Experiments on 3 different datasets and 7 kinds of restoration tasks show that CASA is better than other methods and achieves a new SOTA. Ziqiang Shi, Rujie Liu |
ICASSP | 1 |
| 2024 | Langwave: Realistic Voice Generation Based on High-Order Langevin DynamicsabstractIn recent years, methods based on diffusion generative models have achieved state-of-the-art performances in voice generation. Most of these previous approaches are based on first-order stochastic differential equations or their equivalent diffusion models. This paper attempts to upgrade these first-order methods and propose LangWave, which uses the third-order Langevin dynamical system to generate speech waveforms. LangWave can simultaneously model the position, velocity and acceleration of voice wave diffusion and sampling in the ambient Euclidean space. Thus our vocoder can more precisely and smoothly control the wave evolution from white noise to meaningful waveforms. The experiments on the public data set LJSpeech show that the effect is significant in both objective and subjective evaluation, and achieve the new state-of-the-art MOS of 4.55. Audio samples are available at https://shiziqiang.github.io/langwave. Ziqiang Shi, Rujie Liu |
ICASSP | 1 |
| 2024 | Project, Skate, and Refresh: Improved Schrödinger Bridge Sampler for Image RestorationabstractThe recent advancements in diffusion model-based image restoration (DMIR) have attracted significant attention, particularly the Image-to-Image Schrödinger Bridge ($\mathrm{I}^{2} \mathrm{SB}$) method. This approach has surpassed previous state-of-the-art (SOTA) benchmarks in handling complex data distributions, including ImageNet. However, $I^{2}$ SB faces challenges with three main estimation errors: original image estimation error at current step, next step image estimation error, and persistent random errors. These issues contribute to artifacts in the output. Our research focuses on overcoming these limitations. We introduce a set of three innovative algorithms, named PSR (Project, Skate, and Refresh), designed to efficiently address the estimation errors in $\mathrm{I}^{2} \mathrm{SB}$ without extra training. ‘Project’ aligns the current moment’s image estimation with the original space of the image restoration problem. ‘Skate’ follows the gradient descent on the manifold to correct the next moment’s image estimation error. ‘Refresh’ diminishes random errors at any step. PSR integrates smoothly with $\mathrm{I}^{2} \mathrm{SB}$ model, showing minimal additional computational load. Our tests on ImageNet demonstrate that PSR greatly outperforms the $\mathrm{I}^{2} \mathrm{SB}$ method in tasks like deblurring, super-resolution, and JPEG artifact removal, achieving new SOTA on public benchmarks in metrics such as FID, SSIM, and PSNR. Ziqiang Shi, Rujie Liu |
ICIP | 1 |
| 2024 | Multimedia Generative Modelling with High-Order Langevin DynamicsabstractDiffusion generative models based on stochastic differential equations (SDEs) with score matching have shown remarkable success in data generation. This paper introduces an advanced generative modeling approach, leveraging high-order Langevin dynamics (HOLD) coupled with score matching. Our method substantiated by third-order Langevin dynamics, extends traditional SDEs like variance exploding or variance preserving SDEs for single-variable (data) processes. HOLD uniquely models position, velocity, and acceleration, enhancing both the quality and speed of data generation. Comprising an Ornstein-Uhlenbeck process and two Hamiltonians, HOLD significantly reduces mixing time by approximately two orders of magnitude. Empirical tests on unconditional image generation using the public CIFAR-10 and ImageNet datasets demonstrate notable improvements. The proposed method achieves state-of-the-art Frechet Inception Distances of 1.85 and 1.48 on CIFAR-10 and ImageNet respectively, and also showing substantial gains in negative log-likelihood. Ziqiang Shi, Rujie Liu |
ICME | 1 |
| 2024 | Self-Checkout Product Detection with Occlusion Layer Prediction and Intersection WeightingabstractAutomatic self-checkout based on computer vision is gaining popularity in the field of retail industry, due to the convenience for customers and manpower saving. Thus, retail product detection is vital important in the process of automatic checkout. The task of product detection based on single camera is still challenging, like (1) holding a variety of different products in one or both hands, (2) variable product appearance, (3) intentional fraudulent checkout practices. In this paper, we introduce a third branch on ordinary detectors to predict the occlusion layer of a product and then adopt occlusion layer aware non-max suppression (OLA-NMS) to depress false positives while keeping detection rate. Furthermore, IoU-activate loss is adopted by considering location information in the classification loss. Our third contribution is that we have collected a large-scale of retail checkout images for the target of self-checkout monitoring (SCOM), since there is no dataset or benchmark available for retail product detection under occlusion. Experiments are conducted on SCOM dataset to demonstrate the effectiveness of the proposed method. Zhongling Liu, Ziqiang Shi, Rujie Liu, Liu Liu 0020, Takuma Yamamoto, Daisuke Uchida |
SMC | 2 |
| 2024 | Conditional Velocity Score Estimation for Image RestorationabstractThis paper proposes a new image restoration method by introducing a velocity variable on top of the data position during recovery. Under the guidance of the degraded image, it can effectively and dynamically control the direction of the diffusion path in the reverse-time stochastic differential equation (SDE). So the crucial factor is how to combine the degraded signal as a guide in this second-order reverse process with velocity, especially in the moving direction as a diffusion path. To this end, we propose a conditional velocity score approximation (CVSA) method based on the Bayesian principle to approximate the true posterior conditional velocity score by the sum of a priori conditional velocity score and an observation velocity score of the degraded measurement at the current moment. Our method is versatile from two perspectives. It can be used for both nonblind restoration and blind restoration. At the same time, there is almost no requirement for the degradation operator, and both linear and nonlinear tasks are acceptable. In nonblind restoration, including deblurring, inpainting, superresolution, phase retrieval, and blind restoration, such as deblurring experiments, CVSA is better than other methods and achieves a new state-of-the-art. Ziqiang Shi, Rujie Liu |
WACV | 1 |
| 2024 | RealSinger: Ultra-realistic singing voice generation via stochastic differential equations
Ziqiang Shi, Shoule Wu |
Neurocomputing | 1 |
| 2023 | Semi-Supervised Contrastive Learning with Soft Mask Attention for Facial Action Unit DetectionabstractThis paper presents a novel facial action unit (AU) detection method by simultaneously improving AU feature’s discriminative ability and alleviating the AU data scarcity problem. We design a supervised AU soft mask attention scheme to learn local AU features by integrating prior expert knowledge. To further improve the discriminativeness of AU features, contrastive learning is introduced in both instance-level and prototype-level for each AU. For the data scarcity problem, prototypical pseudo label assignment method is proposed in order to make the potential of unlabeled data, where pseudo-labels are assigned to unlabeled data based on the prototypes of each AU. Overall, our semi-supervised contrastive learning approach employs region learning, contrastive learning and pseudo labeling jointly to enhance the discriminativeness of AU features in the feature space and improve the generalization ability of the model. The effectiveness of the proposed method has been verified by the experiments on benchmark datasets BP4D and DISFA, achieving the state-of-the-art F1-scores of 64.1% and 64.2% respectively. Zhongling Liu, Rujie Liu, Ziqiang Shi, Liu Liu 0020, Xiaoyu Mi, Kentaro Murase |
ICASSP | 3 |
| 2022 | ItôWave: Itô Stochastic Differential Equation is all You Need for Wave GenerationabstractIn this paper, we propose a vocoder based on a pair of forward and reverse-time linear stochastic differential equations (SDE). The solutions of this SDE pair are two stochastic processes, one of which turns the distribution of wave, that we want to generate, into a simple and tractable distribution. The other is the generation procedure that turns this tractable simple signal into the target wave. The model is called ItôWave. It uses the Wiener process as a driver to gradually subtract the excess signal from the noise signal to generate realistic corresponding meaningful audio respectively, under the conditional inputs of original mel spectrogram. The results of the experiment show that the mean opinion scores (MOS) of ItôWave can exceed the current state-of-the-art (SOTA) methods, and reached 4.35±0.115. The generated audio samples are available online1. Shoule Wu, Ziqiang Shi |
ICASSP | 2 |
| 2020 | Speech Separation Based on Multi-Stage Elaborated Dual-Path Deep BiLSTM with Auxiliary Identity LossabstractDeep neural network with dual-path bi-directional long short-term memory (BiLSTM) block has been proved to be very effective in sequence modeling, especially in speech separation. This work investigates how to extend dual-path BiLSTM to result in a new state-of-the-art approach, called TasTas, for multi-talker monaural speech separation (a.k.a cocktail party problem). TasTas introduces two simple but effective improvements, one is an iterative multi-stage refinement scheme, and the other is to correct the speech with imperfect separation through a loss of speaker identity consistency between the separated speech and original speech, to boost the performance of dual-path BiLSTM based networks. TasTas takes the mixed utterance of two speakers and maps it to two separated utterances, where each utterance contains only one speaker's voice. Our experiments on the notable benchmark WSJ0-2mix data corpus result in 20.55dB SDR improvement, 20.35dB SI-SDR improvement, 3.69 of PESQ, and 94.86\% of ESTOI, which shows that our proposed networks can lead to big performance improvement on the speaker separation task. We have open sourced our re-implementation of the DPRNN-TasNet here (this https URL), and our TasTas is realized based on this implementation of DPRNN-TasNet, it is believed that the results in this paper can be reproduced with ease. Ziqiang Shi, Rujie Liu, Jiqing Han 0001 |
INTERSPEECH | 1 |
| 2020 | ATReSN-Net: Capturing Attentive Temporal Relations in Semantic Neighborhood for Acoustic Scene Classification
Liwen Zhang 0001, Jiqing Han 0001, Ziqiang Shi |
INTERSPEECH | 3 |
| 2020 | FurcaNeXt: End-to-End Monaural Speech Separation with Dynamic Gated Dilated Temporal Convolutional Networks
Liwen Zhang 0001, Ziqiang Shi, Jiqing Han 0001, Anyan Shi, Ding Ma 0001 |
MMM (1) | 2 |
| 2020 | Learning Temporal Relations from Semantic Neighbors for Acoustic Scene ClassificationabstractConvolutional networks have achieved the state-of-the-art performance on Acoustic Scene Classification (ASC). Given the Log Mel-Spectrogram of an audio sample, the network can extract useful semantic contents in a certain range receptive field by stacking local convolutional operations. However, the temporal relations between different receptive fields are not captured explicitly. In this letter, we propose an end-to-end 3D Convolutional Neural Network (CNN) for ASC, named SeNoT-Net, which can generate effective audio representations by capturing temporal relations from semantic neighbors of different receptive fields over time. The SeNoT-Net treats the Log-Mel spectrogram as an ordered segment-level sequence. For each segment, the residual block can produce the semantic feature maps, then the semantic neighbors over time (SeNoT) module is applied to capture the relations between each feature point in the feature maps and its top-k semantic neighbors. The proposed SeNoT-Net outperforms most of the state-of-the-art CNN models on both DCASE 2018 and 2019 ASC datasets. Liwen Zhang 0001, Jiqing Han 0001, Ziqiang Shi |
IEEE Signal Process. Lett. | 3 |
| 2020 | Pyramidal Temporal Pooling With Discriminative Mapping for Audio ClassificationabstractAudio signals are temporally-structured data, and learning their discriminative representations containing temporal information is crucial for the audio classification. In this article, we propose an audio representation learning method with a hierarchical pyramid structure called pyramidal temporal pooling (PTP) which aims to capture the temporal information of an entire audio sample. By stacking a global temporal pooling layer on multiple local temporal pooling layers, the PTP can capture the high-level temporal dynamics of the input feature sequence in an unsupervised way. Furthermore, in the top global temporal pooling layer, we jointly optimize a learnable discriminative mapping (DM) and a softmax classifier. Such that, a joint learning method for the discriminative audio representations and the classifier called DM-PTP is also presented. By treating the temporal encoding as a low-level constraint of a bi-level optimization problem, the DM-PTP can produce the discriminative representation while maintaining the temporal information of the whole sequence. For an audio sample with an arbitrary time duration, both our PTP and DM-PTP can encode the input feature sequence with arbitrary length into a fixed-length representation. Without using any data augmentation and ensemble learning methods, both PTP and DM-PTP outperform the state-of-the-art CNNs on the audio event recognition (AER) dataset, and can achieve comparable performance on the DCASE 2018 acoustic scene classification (ASC) dataset compared with other best models in the challenge. Liwen Zhang 0001, Ziqiang Shi, Jiqing Han 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | Link Prediction Adversarial Attack Via Iterative Gradient AttackabstractIncreasing deep neural networks are applied in solving graph evolved tasks, such as node classification and link prediction. However, the vulnerability of deep models can be revealed using carefully crafted adversarial examples generated by various adversarial attack methods. To explore this security problem, we define the link prediction adversarial attack problem and put forward a novel iterative gradient attack (IGA) strategy using the gradient information in the trained graph autoencoder (GAE) model. Not surprisingly, GAE can be fooled by an adversarial graph with a few links perturbed on the clean one. The results on comprehensive experiments of different real-world graphs indicate that most deep models and even the state-of-the-art link prediction algorithms cannot escape the adversarial attack, such as GAE. We can benefit the attack as an efficient privacy protection tool from the link prediction of unknown violations. On the other hand, the adversarial attack is a robust evaluation metric for current link prediction algorithms of their defensibility. Jinyin Chen, Ziqiang Shi, Yi Liu 0024 |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2019 | Furcax: End-to-end Monaural Speech Separation Based on Deep Gated (De)convolutional Neural Networks with Adversarial Example TrainingabstractDeep gated convolutional networks have been proved to be very effective in single channel speech separation. However current state-of-the-art framework often considers training the gated convolutional networks in time-frequency (TF) domain. Such an approach will result in limited perceptual score, such as signal-to-distortion ratio (SDR) upper bound of separated utterances and also fail to exploit an end-to-end framework. In this paper we present an integrated simple and effective end-to-end approach called FurcaX1to monaural speech separation, which consists of deep gated (de)convolutional neural networks (GCNN) that takes the mixed utterance of two speakers and maps it to two separated utterances, where each utterance contains only one speaker's voice. For the objective, we propose to train the network by directly optimizing utterance level SDR in a permutation invariant training (PIT) style. We execute generative adversarial training (GAT) throughout the training, which makes the separated speech indistinguishable from the real one. Our experiments on the the public WSJ0-2mix data corpus demonstrate that this new scheme can produce more discriminative separated utterances and leading to performance improvement on the speaker separation task. Ziqiang Shi, Huibin Lin, Liu Liu 0020, Rujie Liu, Shoji Hayakawa, Jiqing Han 0001 |
ICASSP | 1 |
| 2019 | Robustness Evaluation of Deep Learning Models Based on Local Prediction ConsistencyabstractIt is important to estimate the performance gap of a given deep learning model on the target data set, since discrepancy or bias between source and target domains is a common and fundamental problem in the practice of machine learning techniques. Without any assumptions on data bias, such as label shift or covariate shift and without target data labels, we propose a robustness estimation method based on prediction consistency evaluation between source and target data in the neighborhood of the source samples. Considering outliers and whether the user provided model is fully trained, a variety of variant methods are also tried, including setting neighborhood threshold to average intra-class distance for each category and relative robustness. Furthermore, the time complexity of this method is O(nlogn), which is applicable for large datasets. Experiments on the handwritten digit recognition and Japanese handwriting recognition show that the proposed methods are effective. Ziqiang Shi, Chaoliang Zhong, Yasuto Yokota, Wensheng Xia, Jun Sun 0004 |
ICMLA | 1 |
| 2019 | End-to-End Monaural Speech Separation with Multi-Scale Dynamic Weighted Gated Dilated Convolutional Pyramid Network
Ziqiang Shi, Huibin Lin, Liu Liu 0020, Rujie Liu, Shoji Hayakawa, Shouji Harada, Jiqing Han 0001 |
INTERSPEECH | 1 |
| 2019 | Deep Attention Gated Dilated Temporal Convolutional Networks with Intra-Parallel Convolutional Modules for End-to-End Monaural Speech Separation
Ziqiang Shi, Huibin Lin, Liu Liu 0020, Rujie Liu, Jiqing Han 0001, Anyan Shi |
INTERSPEECH | 1 |
| 2018 | Double Joint Bayesian Modeling of DNN Local I-Vector for Text Dependent Speaker Verification with Random Digit Strings
Ziqiang Shi, Huibin Lin, Liu Liu 0020, Rujie Liu |
INTERSPEECH | 1 |
| 2018 | Joint Learning of J-Vector Extractor and Joint Bayesian Model for Text Dependent Speaker Verification
Ziqiang Shi, Liu Liu 0020, Huibin Lin, Rujie Liu |
INTERSPEECH | 1 |
| 2018 | Latent Factor Analysis of Deep Bottleneck Features for Speaker Verification with Random Digit Strings
Ziqiang Shi, Huibin Lin, Liu Liu 0020, Rujie Liu |
INTERSPEECH | 1 |
| 2017 | Multi-view (Joint) probability linear discrimination analysis for J-vector based text dependent speaker verificationabstractJ-vector has been proved to be very effective in text dependent speaker verification with short-duration speech. However, the current back-end classifiers cannot make full use of such deep features. In this paper, we propose a method to model the multi-faceted information in the j-vector explicitly and jointly. Examples of the multi-faceted information include speaker identity and text content. In our approach, the j-vector was modeled as a result derived by a generative multi-view (joint1) Probability Linear Discriminant Analysis (PLDA) model, which contains multiple kinds of latent variables. The usual PLDA model only considers one single label. However, in practical use, when using multi-task learned network as feature extractor, the extracted feature are always associated with several labels. This type of feature is called multi-view deep feature (e.g. j-vector). With multi-view (joint) PLDA, we are able to explicitly build a model that can combine multiple heterogeneous information from the j-vectors. In verification step, we calculated the likelihood to describe whether the two j-vectors having consistent labels or not. This likelihood is used in the following decision-making. Experiments have been conducted on large scale data corpus of different languages. On the public RSR2015 data corpus, the results showed that our approach can achieve 0.02% EER and 0.09% EER for impostor wrong and impostor correct cases respectively. Ziqiang Shi, Liu Liu 0020, Mengjiao Wang 0001, Rujie Liu |
ASRU | 1 |
| 2017 | Better Worst-Case Complexity Analysis of the Block Coordinate Descent Method for Large Scale Machine LearningabstractThis paper considers the problem of unconstrained minimization of large scale machine learning evolving smooth convex functions having block-coordinate-wise Lipschitz continuous gradients. The Block Coordinate Descent (BCD) method was among the first optimization schemes suggested for solving such problems [1]. In this work, we obtain a new lower (to our best knowledge the lowest currently) bound, which is 16p3 times smaller than the best known on the information-based complexity of BCD method. We achieve this by using an effective technique called Performance Estimation Problem (PEP) approach for analyzing the performance of first-order black box optimization methods. Numerical test confirms our analysis. Ziqiang Shi, Rujie Liu |
ICMLA | 1 |
| 2015 | Online and Stochastic Universal Gradient Methods for Minimizing Regularized Hölder Continuous Finite Sums in Machine Learning
Ziqiang Shi, Rujie Liu |
PAKDD (1) | 1 |
| 2015 | Large Scale Optimization with Proximal Stochastic Newton-Type Gradient Descent
Ziqiang Shi, Rujie Liu |
ECML/PKDD (1) | 1 |
| 2015 | Soft Margin Based Low-Rank Audio Signal Classification
Ziqiang Shi, Jiqing Han 0001, Tieran Zheng |
Neural Process. Lett. | 1 |
| 2013 | Guarantees of Augmented Trace Norm Models in Tensor Recovery
Ziqiang Shi, Jiqing Han 0001, Tieran Zheng |
IJCAI | 1 |
| 2013 | Audio Segment Classification Using Online Learning Based Tensor Representation Feature DiscriminationabstractIn order to naturally combine audio information from different dimensions and build robust audio processing system, a novel framework based on low-rank tensor representation features for audio segment classification is proposed in this paper. The audio signal is first transformed into tensor format data, and then these tensor data are mapped to a low-rank space which is insensitive under certain noises, especially white Gaussian noise and gross corruptions. For these low-rank tensor based features, tensor classification via a linear classifier based on minimization a smooth loss function regularized by the trace norm proposed recently is used. Most previous methods find the weight tensor and bias in batch-mode learning, which makes them inefficient for large-scale problems. In this paper, we propose to address this problem with an online learning algorithm based on the accelerated proximal gradient (APG) method, which scales up gracefully to large data sets. Experiments on simulation and real audio data demonstrate the efficiency of the methods. Ziqiang Shi, Jiqing Han 0001, Tieran Zheng, Shiwen Deng |
IEEE Trans. Speech Audio Process. | 1 |
| 2013 | Identification of Objectionable Audio Segments Based on Pseudo and Heterogeneous Mixture ModelsabstractIn this paper, we generalize the Gaussian Mixture Model (GMM) in two ways: a) by introducing novel distance measures between two vectors based on nonlinear maps to give more general mixture models; b) by building mixture models based on multiple different kinds of distributions. These two generalizations cope with different problems arisen in feature modeling. Mixture model obtained by first method is called pseudo Gaussian Mixture Model (pseudo GMM). Compared to the traditional GMM, pseudo GMM with nonlinear maps have better performance on nonlinear problems, while the computational complexity is almost the same as the Expectation-Maximization (EM) algorithm for traditional GMM according to the iteration procedures. The second generalization considers that in practice the practical learning problem often involves multiple, heterogeneous data sources, while classical mixture models are based on a single kind of distribution. In this work, we consider heterogeneous mixture models (hetMM) based on multiple different kinds of distributions. Different types of distributions in hetMM may have quite different properties and may capture different features of the data. Component classifiers including pseudo and hetMM based classifiers are employed in our task of erotic audio recognition. Experimental results with classifiers built based on pseudo GMM and hetMM for erotic audio recognition demonstrate the effectiveness of the proposed model. Online and off-line experiments show that the proposed approach is highly effective for erotic audio recognition. Ziqiang Shi, Jiqing Han 0001, Tieran Zheng |
IEEE Trans. Speech Audio Process. | 1 |
| 2013 | Audio classification with low-rank matrix representation featuresabstractIn this article, a novel framework based on trace norm minimization for audio classification is proposed. In this framework, both the feature extraction and classification are obtained by solving corresponding convex optimization problem with trace norm regularization. For feature extraction, robust principle component analysis (robust PCA) via minimization a combination of the nuclear norm and the ℓ 1 -norm is used to extract low-rank matrix features which are robust to white noise and gross corruption for audio signal. These low-rank matrix features are fed to a linear classifier where the weight and bias are learned by solving similar trace norm constrained problems. For this linear classifier, most methods find the parameters, that is the weight matrix and bias in batch-mode, which makes it inefficient for large scale problems. In this article, we propose a parallel online framework using accelerated proximal gradient method. This framework has advantages in processing speed and memory cost. In addition, as a result of the regularization formulation of matrix classification, the Lipschitz constant was given explicitly, and hence the step size estimation of the general proximal gradient method was omitted, and this part of computing burden is saved in our approach. Extensive experiments on real data sets for laugh/non-laugh and applause/non-applause classification indicate that this novel framework is effective and noise robust. Ziqiang Shi, Jiqing Han 0001, Tieran Zheng |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2012 | Low-rank Audio Signal Classification Under Soft Margin and Trace Norm Constraints
Ziqiang Shi, Tieran Zheng, Jiqing Han 0001, Shiwen Deng |
INTERSPEECH | 1 |
| 2011 | A Novel Framework Based on Trace Norm Minimization for Audio Event Detection
Ziqiang Shi, Jiqing Han 0001, Tieran Zheng |
ICONIP (2) | 1 |
| 2011 | Real-World Speech/Non-Speech Audio Classification Based on Sparse Representation Features and GPCs
Ziqiang Shi, Jiqing Han 0001, Tieran Zheng |
INTERSPEECH | 1 |
| 2010 | Study on the Recognition of Objectionable AudioabstractIn this paper, a novel method from the feature — porno-sounds recognition — point of view is proposed to detect adult video sequences automatically which may serve as a verification step, a supplementary method or an independent detector. To the specificity of erotic sound, its feature analysis is given. Based on the popular features, histograms and contours are introduced as new sets of features. At the same time due to the complexity of outside data, a general framework called in-class clustering is proposed which selects the most representative subclass for training and classification. All these efforts increase the recall rate and decrease the false positive rate. Experiments on real data from the Internet indicate that the proposed method yields superior performance with 89.17% recall rate and 10.78% false positive rate being achieved. Ziqiang Shi, Boyang Gao, Tieran Zheng, Jiqing Han 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 1 |