EDBT 2026 Demo / reviewers in the wild / expert
Rujie Liu
dblp:18/1353
· DBLP profile ↗
50ranked-venue papers
7as first author
19since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 41 · 4 first-author · 16 since 2021Artificial intelligence and machine learning · 20 · 4 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 5 · 5 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Security and privacy · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Eliminating Object Hallucination in MLLMs via Convex Potential Flow Intervention
Ziqiang Shi, Rujie Liu, Koichi Shirahata |
ICPR (9) | 2 |
| 2026 | Scalpel: Fine-Grained Alignment of Attention Activation Manifolds via Mixture Gaussian Bridges to Mitigate Multimodal HallucinationabstractRapid progress in large vision-language models (LVLMs) has achieved unprecedented performance in vision-language tasks. However, due to the strong prior of large language models (LLMs) and misaligned attention across modalities, LVLMs often generate outputs inconsistent with visual content - termed hallucination. To address this, we propose Scalpel, a method that reduces hallucination by refining attention activation distributions toward more credible regions. Scalpel predicts trusted attention directions for each head in Transformer layers during inference and adjusts activations accordingly. It employs a Gaussian mixture model to capture multi-peak distributions of attention in trust and hallucination manifolds, and uses entropic optimal transport (equivalent to Schrödinger bridge problem) to map Gaussian components precisely. During mitigation, Scalpel dynamically adjusts intervention strength and direction based on component membership and mapping relationships between hallucination and trust activations. Extensive experiments across multiple datasets and benchmarks demonstrate that Scalpel effectively mitigates hallucinations, outperforming previous methods and achieving state-of-the-art performance. Moreover, Scalpel is model-and data-agnostic, requiring no additional computation, only a single decoding step. Ziqiang Shi, Rujie Liu, Satoshi Munakata, Koichi Shirahata |
WACV | 2 |
| 2025 | Attribute Conditional Diffusion-Augmented Person Re-IdentificationabstractDue to privacy and cost issues, the lack of large-scale labeled datasets limits the advancement of person re-identification. Existing methods use generative adversarial networks or game engine rendering for data augmentation to improve re-identification performance. However, these approaches struggle to maintain realistic images. This paper introduces a novel approach called Identity Diffuser, which uses diffusion models to generate synthetic data for the same identity with different poses. Our proposed framework incorporates identity-specific embeddings and target poses into the diffusion process, enabling the generation of realistic and diverse images that consistently preserve identity features. Guided by pretrained re-identification net and target pose heatmap, the framework learns transformation trajectories through forward and backward denoising steps in the diffusion models. This approach effectively maintains key pedestrian attributes across various poses. Experimental results on the Market1501 and DukeMTMC datasets demonstrate a notable improvement in performance, with a 1.73%/0.80% mAp increase in Market1501/DukeMTMC datasets compared with current state-of-the-art method. When less real data is included, the increment can be 5.1%/1.5%, separately. Shijie Nie, Ziqiang Shi, Rujie Liu, Meng Zhang 0042, Mengjiao Wang 0001, Kazuki Osamura, Lina Septiana, Narishige Abe |
ICASSP | 3 |
| 2025 | Human Identification at a Distance: Challenges, Methods and Results on the Competition HID 2025abstractHuman identification at a distance (HID) faces challenges due to the difficulty of acquiring traditional biometric modalities like face and fingerprints. Gait recognition offers a viable solution since it can be captured at a distance. To promote progress in gait recognition and provide a fair evaluation platform, the International Competition on Human Identification at a Distance (HID) has been organized annually since 2020. Since 2023, the competition has adopted the challenging SUSTech-Competition dataset, which includes significant variations in clothing, carried objects, and view angles. No training data is provided, requiring participants to train their models using external datasets. Each year, the competition applies a different random seed to generate distinct evaluation splits, reducing the risk of overfitting and ensuring fair evaluation of cross-domain generalization. Although the previous two competitions (HID 2023 and HID 2024) already utilized this dataset, HID 2025 aimed explicitly to explore whether algorithmic improvements could surpass the accuracy limits observed previously. Despite these heightened challenges, participants again demonstrated significant advancements, with the highest accuracy reaching 94.2%, setting a new benchmark for this dataset. We also analyze key technical trends and outline potential directions for future research on gait recognition. Jingzhe Ma, Jianlong Yu, Zunxiao Xu, Xue Cheng, Zepeng Wang 0002, Kazuki Osamura, Rujie Liu, Narishige Abe, Shunli Zhang 0005, Haojun Xie, Weiming Wu, Wenxiong Kang, Qingshuo Gao, Jiaming Xiong, Xianye Ben, Lei Chen 0095, Lichen Song, Junjian Cui, Haijun Xiong, Junhao Lu, Bin Feng 0001, Baoquan Zhao, Ke Xu 0001, Yongzhen Huang, Liang Wang 0001, Manuel J. Marín-Jiménez, Md. Atiqur Rahman Ahad, Shiqi Yu 0001 |
IJCB | 12 |
| 2025 | TrueCount: Improving Open-World Object Counting with Visual-Language Models and Dynamic Multi-Modal Inputs
Ziqiang Shi, Rujie Liu |
ACM Multimedia | 2 |
| 2025 | Selective-SAM: Memory Optimization for Segment Anything Model 2 with Application in Self-Checkout Product CountingabstractThe Segment Anything Model 2 (SAM 2) [1] demonstrates strong capabilities for video object segmentation (VOS) [2]. We present a SAM 2-based framework for self-checkout product counting, where box prompts are generated by a detector selecting optimal tracking initiation frames. Our key contribution is Selective-SAM, a training-free enhancement that improves robustness in complex scenes via a selective memory bank. This mechanism selectively preserves high-quality and diverse features while filtering poor segmentation priors. For final counting, we introduce a novel Mask Overlap Degree metric to analyze object trajectories. By segmenting trajectories based on the mask overlap degrees, we accurately determine the product count. Experiments across various retail scenarios show improvements of 1.95% IDF1 and 3.86% MOTA. Zhongling Liu, Liu Liu 0020, Ziqiang Shi, Rujie Liu |
SMC | 4 |
| 2025 | AugCount: Test-Time Semantic Augmentation via Diffusion for General Open-World Object CountingabstractOpen-world general object counting is a critical task in computer vision, with applications in image understanding, environmental monitoring, and surveillance. Traditional methods rely on large-scale annotated data, which is expensive and time-consuming to obtain, and often fail to manage diverse object categories and complex scenes. To address this, we propose AugCount, a novel framework that enhances open-world object counting by generating high-quality, diverse synthetic data during testing. For the first time, we employ a diffusion model to produce conditional images based on density maps specifying object locations for general object counting. Our framework is highly versatile and adaptable to various counting tasks. Experiments demonstrate that AugCount significantly improves performance on benchmark datasets like FSC-147, reducing the average counting error by 1 per image and achieving a state-of-the-art MAE of 4.70. AugCount effectively addresses data scarcity and model generalization challenges, offering enhanced robustness and adaptability for practical counting systems. Ziqiang Shi, Rujie Liu |
SMC | 2 |
| 2025 | Bayesian Optimal Latent Projection for Noisy Image RestorationabstractIn recent years, image restoration using large-scale la-tent diffusion generative models (DGM) has attracted in-creasing attention and achieved significant progress. Most of these latent DGM-based image restoration methods re-quire predicting the original clean image in each iteration, which is then used to estimate the image for the next iter-ation. However, these predicted original clean images are often inaccurate, leading to errors in the subsequent im-age estimation. In other words, there is a significant devi-ation between the final sampling restoration trajectory and the ground truth trajectory. The purpose of this paper is to narrow the gap between these two trajectories and en-hance the performance of image restoration. We propose the Bayesian Optimal Latent Projection (BOLP) algorithm, which identifies the optimal random noise within the Gaus-sian distribution to iteratively correct the estimated image at each step, thereby minimizing the distance to the ground truth image. Experiments in deblurring, super-resolution, and inpainting on FFHQ and ImageNet datasets demon-strate that the BOLP outperforms the previously established best algorithms and sets a new state of the art. Ziqiang Shi, Rujie Liu, Takuma Yamamoto |
WACV | 2 |
| 2024 | Noisy Image Restoration Based on Conditional Acceleration Score ApproximationabstractIn recent years, score-based generative models (SGM) have achieved state-of-the-art (SOTA) performance in noisy image restoration [1], [2]. But at present, most of these methods are performed in the position space, and there is a lack in modeling of the velocity and acceleration of the image on the restoration path. In this paper, we propose a new image restoration method called conditional acceleration score approximation (CASA), which introduces velocity and acceleration variables on top of the data position along the recovery path. Guided by the degraded image, CASA can effectively and dynamically control the direction and speed of motion along the diffusion path in the reverse-time stochastic differential equation. Therefore, the key to this process is how to inject the degraded image as a guidance into the third-order reverse-time process in this position-velocity-acceleration space, especially in the evolution direction of the diffusion path. We propose a strategy for approximating the conditional acceleration score by decomposing the true posterior CAS into a priori CAS and an observed acceleration score for the measurement at the current moment. Experiments on 3 different datasets and 7 kinds of restoration tasks show that CASA is better than other methods and achieves a new SOTA. Ziqiang Shi, Rujie Liu |
ICASSP | 2 |
| 2024 | Langwave: Realistic Voice Generation Based on High-Order Langevin DynamicsabstractIn recent years, methods based on diffusion generative models have achieved state-of-the-art performances in voice generation. Most of these previous approaches are based on first-order stochastic differential equations or their equivalent diffusion models. This paper attempts to upgrade these first-order methods and propose LangWave, which uses the third-order Langevin dynamical system to generate speech waveforms. LangWave can simultaneously model the position, velocity and acceleration of voice wave diffusion and sampling in the ambient Euclidean space. Thus our vocoder can more precisely and smoothly control the wave evolution from white noise to meaningful waveforms. The experiments on the public data set LJSpeech show that the effect is significant in both objective and subjective evaluation, and achieve the new state-of-the-art MOS of 4.55. Audio samples are available at https://shiziqiang.github.io/langwave. Ziqiang Shi, Rujie Liu |
ICASSP | 2 |
| 2024 | Face Helps Person Re-Identification: Multi-modality Person Re-Identification Based on Vision-Language ModelsabstractPerson re-identification (ReID), aiming to identify individuals from camera views, often faces challenges such as occlusion and appearance variations by cloth changing. Moreover, due to the long-distance capturing and varying positions of the pedestrians, human face is not always visible thus it is usually neglected in ReID. This paper proposes a novel approach to enhance ReID performance by integrating face and body into a multi-modality ReID framework, particularly improving the behavior in scenarios with occlusion and clothes-changing. Leveraging the visual-linguistic capabilities of the CLIP model, our framework comprises two CLIP-like structures: one dedicated to extracting body appearance features and the other one focused on face features. Furthermore, a feature adapter method is proposed to address the issue of invisible face. Experiments show that state-of-the-art (SOTA) performance is achieved on six popular benchmarks datasets, including Market1501, and LTCC, confirming the superiority of the proposed method. Additionally, we have proposed a multimodality ReID dataset to further verify and analyze the effectiveness of the proposed multi-modality ReID framework. Rujie Liu, Narishige Abe |
IJCB | 2 |
| 2024 | Project, Skate, and Refresh: Improved Schrödinger Bridge Sampler for Image RestorationabstractThe recent advancements in diffusion model-based image restoration (DMIR) have attracted significant attention, particularly the Image-to-Image Schrödinger Bridge ($\mathrm{I}^{2} \mathrm{SB}$) method. This approach has surpassed previous state-of-the-art (SOTA) benchmarks in handling complex data distributions, including ImageNet. However, $I^{2}$ SB faces challenges with three main estimation errors: original image estimation error at current step, next step image estimation error, and persistent random errors. These issues contribute to artifacts in the output. Our research focuses on overcoming these limitations. We introduce a set of three innovative algorithms, named PSR (Project, Skate, and Refresh), designed to efficiently address the estimation errors in $\mathrm{I}^{2} \mathrm{SB}$ without extra training. ‘Project’ aligns the current moment’s image estimation with the original space of the image restoration problem. ‘Skate’ follows the gradient descent on the manifold to correct the next moment’s image estimation error. ‘Refresh’ diminishes random errors at any step. PSR integrates smoothly with $\mathrm{I}^{2} \mathrm{SB}$ model, showing minimal additional computational load. Our tests on ImageNet demonstrate that PSR greatly outperforms the $\mathrm{I}^{2} \mathrm{SB}$ method in tasks like deblurring, super-resolution, and JPEG artifact removal, achieving new SOTA on public benchmarks in metrics such as FID, SSIM, and PSNR. Ziqiang Shi, Rujie Liu |
ICIP | 2 |
| 2024 | Multimedia Generative Modelling with High-Order Langevin DynamicsabstractDiffusion generative models based on stochastic differential equations (SDEs) with score matching have shown remarkable success in data generation. This paper introduces an advanced generative modeling approach, leveraging high-order Langevin dynamics (HOLD) coupled with score matching. Our method substantiated by third-order Langevin dynamics, extends traditional SDEs like variance exploding or variance preserving SDEs for single-variable (data) processes. HOLD uniquely models position, velocity, and acceleration, enhancing both the quality and speed of data generation. Comprising an Ornstein-Uhlenbeck process and two Hamiltonians, HOLD significantly reduces mixing time by approximately two orders of magnitude. Empirical tests on unconditional image generation using the public CIFAR-10 and ImageNet datasets demonstrate notable improvements. The proposed method achieves state-of-the-art Frechet Inception Distances of 1.85 and 1.48 on CIFAR-10 and ImageNet respectively, and also showing substantial gains in negative log-likelihood. Ziqiang Shi, Rujie Liu |
ICME | 2 |
| 2024 | RTAT: A Robust Two-Stage Association Tracker for Multi-object Tracking
Rujie Liu, Narishige Abe |
ICPR (16) | 2 |
| 2024 | Self-Checkout Product Detection with Occlusion Layer Prediction and Intersection WeightingabstractAutomatic self-checkout based on computer vision is gaining popularity in the field of retail industry, due to the convenience for customers and manpower saving. Thus, retail product detection is vital important in the process of automatic checkout. The task of product detection based on single camera is still challenging, like (1) holding a variety of different products in one or both hands, (2) variable product appearance, (3) intentional fraudulent checkout practices. In this paper, we introduce a third branch on ordinary detectors to predict the occlusion layer of a product and then adopt occlusion layer aware non-max suppression (OLA-NMS) to depress false positives while keeping detection rate. Furthermore, IoU-activate loss is adopted by considering location information in the classification loss. Our third contribution is that we have collected a large-scale of retail checkout images for the target of self-checkout monitoring (SCOM), since there is no dataset or benchmark available for retail product detection under occlusion. Experiments are conducted on SCOM dataset to demonstrate the effectiveness of the proposed method. Zhongling Liu, Ziqiang Shi, Rujie Liu, Liu Liu 0020, Takuma Yamamoto, Daisuke Uchida |
SMC | 3 |
| 2024 | Conditional Velocity Score Estimation for Image RestorationabstractThis paper proposes a new image restoration method by introducing a velocity variable on top of the data position during recovery. Under the guidance of the degraded image, it can effectively and dynamically control the direction of the diffusion path in the reverse-time stochastic differential equation (SDE). So the crucial factor is how to combine the degraded signal as a guide in this second-order reverse process with velocity, especially in the moving direction as a diffusion path. To this end, we propose a conditional velocity score approximation (CVSA) method based on the Bayesian principle to approximate the true posterior conditional velocity score by the sum of a priori conditional velocity score and an observation velocity score of the degraded measurement at the current moment. Our method is versatile from two perspectives. It can be used for both nonblind restoration and blind restoration. At the same time, there is almost no requirement for the degradation operator, and both linear and nonlinear tasks are acceptable. In nonblind restoration, including deblurring, inpainting, superresolution, phase retrieval, and blind restoration, such as deblurring experiments, CVSA is better than other methods and achieves a new state-of-the-art. Ziqiang Shi, Rujie Liu |
WACV | 2 |
| 2024 | Texture-Guided Transfer Learning for Low-Quality Face RecognitionabstractAlthough many advanced works have achieved significant progress for face recognition with deep learning and large-scale face datasets, low-quality face recognition remains a challenging problem in real-word applications, especially for unconstrained surveillance scenes. We propose a texture-guided (TG) transfer learning approach under the knowledge distillation scheme to improve low-quality face recognition performance. Unlike existing methods in which distillation loss is built on forward propagation; e.g., the output logits and intermediate features, in this study, the backward propagation gradient texture is used. More specifically, the gradient texture of low-quality images is forced to be aligned to that of its high-quality counterpart to reduce the feature discrepancy between the high- and low-quality images. Moreover, attention is introduced to derive a soft-attention (SA) version of transfer learning, termed as SA-TG, to focus on informative regions. Experiments on the benchmark low-quality face DB's TinyFace and QMUL-SurFace confirmed the superiority of the proposed method, especially more than 6.6% Rank1 accuracy improvement is achieved on TinyFace. Meng Zhang 0042, Rujie Liu, Daisuke Deguchi, Hiroshi Murase |
IEEE Trans. Image Process. | 2 |
| 2023 | Semi-Supervised Contrastive Learning with Soft Mask Attention for Facial Action Unit DetectionabstractThis paper presents a novel facial action unit (AU) detection method by simultaneously improving AU feature’s discriminative ability and alleviating the AU data scarcity problem. We design a supervised AU soft mask attention scheme to learn local AU features by integrating prior expert knowledge. To further improve the discriminativeness of AU features, contrastive learning is introduced in both instance-level and prototype-level for each AU. For the data scarcity problem, prototypical pseudo label assignment method is proposed in order to make the potential of unlabeled data, where pseudo-labels are assigned to unlabeled data based on the prototypes of each AU. Overall, our semi-supervised contrastive learning approach employs region learning, contrastive learning and pseudo labeling jointly to enhance the discriminativeness of AU features in the feature space and improve the generalization ability of the model. The effectiveness of the proposed method has been verified by the experiments on benchmark datasets BP4D and DISFA, achieving the state-of-the-art F1-scores of 64.1% and 64.2% respectively. Zhongling Liu, Rujie Liu, Ziqiang Shi, Liu Liu 0020, Xiaoyu Mi, Kentaro Murase |
ICASSP | 2 |
| 2022 | Sample-Level and Class-Level Adaptive Training for Face RecognitionabstractMarginal softmax loss function has been widely used for face recognition, where a universal angular margin is added between weight prototypes. However, this method neglects the differences between classes and samples. On class-level, the imbalanced real world training dataset requires different margin for the head and tail classes to equally squeeze each class's feature space. On the sample-level, it's also necessary to assign larger importance for the hard samples during training. In this paper, we address these two issues by combining two strategies: (1) explicitly assign the adaptive margin according to the image quantity so that the margin is enlarged for the tail classes; (2) semantically identify the ‘hard positive, samples and misclassified samples [1] to attach adaptive weights to increase the training emphasis on these samples. Extensive experiments on LFW/CFP/AGEDB and IJB-B/IJB-C show our method's effectiveness. Mengjiao Wang 0001, Rujie Liu, Narishige Abe, Tomoaki Matsunami, Hidetsugu Uchida, Lina Septiana |
ICME | 2 |
| 2020 | Speech Separation Based on Multi-Stage Elaborated Dual-Path Deep BiLSTM with Auxiliary Identity LossabstractDeep neural network with dual-path bi-directional long short-term memory (BiLSTM) block has been proved to be very effective in sequence modeling, especially in speech separation. This work investigates how to extend dual-path BiLSTM to result in a new state-of-the-art approach, called TasTas, for multi-talker monaural speech separation (a.k.a cocktail party problem). TasTas introduces two simple but effective improvements, one is an iterative multi-stage refinement scheme, and the other is to correct the speech with imperfect separation through a loss of speaker identity consistency between the separated speech and original speech, to boost the performance of dual-path BiLSTM based networks. TasTas takes the mixed utterance of two speakers and maps it to two separated utterances, where each utterance contains only one speaker's voice. Our experiments on the notable benchmark WSJ0-2mix data corpus result in 20.55dB SDR improvement, 20.35dB SI-SDR improvement, 3.69 of PESQ, and 94.86\% of ESTOI, which shows that our proposed networks can lead to big performance improvement on the speaker separation task. We have open sourced our re-implementation of the DPRNN-TasNet here (this https URL), and our TasTas is realized based on this implementation of DPRNN-TasNet, it is believed that the results in this paper can be reproduced with ease. Ziqiang Shi, Rujie Liu, Jiqing Han 0001 |
INTERSPEECH | 2 |
| 2019 | Furcax: End-to-end Monaural Speech Separation Based on Deep Gated (De)convolutional Neural Networks with Adversarial Example TrainingabstractDeep gated convolutional networks have been proved to be very effective in single channel speech separation. However current state-of-the-art framework often considers training the gated convolutional networks in time-frequency (TF) domain. Such an approach will result in limited perceptual score, such as signal-to-distortion ratio (SDR) upper bound of separated utterances and also fail to exploit an end-to-end framework. In this paper we present an integrated simple and effective end-to-end approach called FurcaX1to monaural speech separation, which consists of deep gated (de)convolutional neural networks (GCNN) that takes the mixed utterance of two speakers and maps it to two separated utterances, where each utterance contains only one speaker's voice. For the objective, we propose to train the network by directly optimizing utterance level SDR in a permutation invariant training (PIT) style. We execute generative adversarial training (GAT) throughout the training, which makes the separated speech indistinguishable from the real one. Our experiments on the the public WSJ0-2mix data corpus demonstrate that this new scheme can produce more discriminative separated utterances and leading to performance improvement on the speaker separation task. Ziqiang Shi, Huibin Lin, Liu Liu 0020, Rujie Liu, Shoji Hayakawa, Jiqing Han 0001 |
ICASSP | 4 |
| 2019 | End-to-End Monaural Speech Separation with Multi-Scale Dynamic Weighted Gated Dilated Convolutional Pyramid Network
Ziqiang Shi, Huibin Lin, Liu Liu 0020, Rujie Liu, Shoji Hayakawa, Shouji Harada, Jiqing Han 0001 |
INTERSPEECH | 4 |
| 2019 | Deep Attention Gated Dilated Temporal Convolutional Networks with Intra-Parallel Convolutional Modules for End-to-End Monaural Speech Separation
Ziqiang Shi, Huibin Lin, Liu Liu 0020, Rujie Liu, Jiqing Han 0001, Anyan Shi |
INTERSPEECH | 4 |
| 2018 | Discover the Effective Strategy for Face Recognition Model Compression by Improved Knowledge DistillationabstractFor the sake of better accuracy, the face recognition model is becoming larger and larger, which makes them difficult to be deployed on embedded systems. This work proposes an effective model compression method using knowledge distillation, where a fast student model is trained under the guidance of a complex teacher model. Firstly, different loss combinations and network architectures are analyzed through comprehensive experiments to find the most effective approach. To augment the performance, the feature layer is further normalized to make the optimization objective consistent with cosine similarity metric. Moreover, a teacher weighting strategy is proposed to address the issue when teacher provides wrong guidance. Experimental results show that the student model built by our approach can surpass the teacher model while achieving 3× acceleration. Mengjiao Wang 0001, Rujie Liu, Narishige Abe, Hidetsugu Uchida, Tomoaki Matsunami, Shigefumi Yamada |
ICIP | 2 |
| 2018 | Double Joint Bayesian Modeling of DNN Local I-Vector for Text Dependent Speaker Verification with Random Digit Strings
Ziqiang Shi, Huibin Lin, Liu Liu 0020, Rujie Liu |
INTERSPEECH | 4 |
| 2018 | Joint Learning of J-Vector Extractor and Joint Bayesian Model for Text Dependent Speaker Verification
Ziqiang Shi, Liu Liu 0020, Huibin Lin, Rujie Liu |
INTERSPEECH | 4 |
| 2018 | Latent Factor Analysis of Deep Bottleneck Features for Speaker Verification with Random Digit Strings
Ziqiang Shi, Huibin Lin, Liu Liu 0020, Rujie Liu |
INTERSPEECH | 4 |
| 2017 | Multi-view (Joint) probability linear discrimination analysis for J-vector based text dependent speaker verificationabstractJ-vector has been proved to be very effective in text dependent speaker verification with short-duration speech. However, the current back-end classifiers cannot make full use of such deep features. In this paper, we propose a method to model the multi-faceted information in the j-vector explicitly and jointly. Examples of the multi-faceted information include speaker identity and text content. In our approach, the j-vector was modeled as a result derived by a generative multi-view (joint1) Probability Linear Discriminant Analysis (PLDA) model, which contains multiple kinds of latent variables. The usual PLDA model only considers one single label. However, in practical use, when using multi-task learned network as feature extractor, the extracted feature are always associated with several labels. This type of feature is called multi-view deep feature (e.g. j-vector). With multi-view (joint) PLDA, we are able to explicitly build a model that can combine multiple heterogeneous information from the j-vectors. In verification step, we calculated the likelihood to describe whether the two j-vectors having consistent labels or not. This likelihood is used in the following decision-making. Experiments have been conducted on large scale data corpus of different languages. On the public RSR2015 data corpus, the results showed that our approach can achieve 0.02% EER and 0.09% EER for impostor wrong and impostor correct cases respectively. Ziqiang Shi, Liu Liu 0020, Mengjiao Wang 0001, Rujie Liu |
ASRU | 4 |
| 2017 | Learning Residual Images for Face Attribute ManipulationabstractFace attributes are interesting due to their detailed description of human faces. Unlike prior researches working on attribute prediction, we address an inverse and more challenging problem called face attribute manipulation which aims at modifying a face image according to a given attribute value. Instead of manipulating the whole image, we propose to learn the corresponding residual image defined as the difference between images before and after the manipulation. In this way, the manipulation can be operated efficiently with modest pixel modification. The framework of our approach is based on the Generative Adversarial Network. It consists of two image transformation networks and a discriminative network. The transformation networks are responsible for the attribute manipulation and its dual operation and the discriminative network is used to distinguish the generated images from real images. We also apply dual learning to allow transformation networks to learn from each other. Experiments show that residual images can be effectively learned and used for attribute manipulations. The generated images remain most of the details in attribute-irrelevant areas. Rujie Liu |
CVPR | 2 |
| 2017 | Better Worst-Case Complexity Analysis of the Block Coordinate Descent Method for Large Scale Machine LearningabstractThis paper considers the problem of unconstrained minimization of large scale machine learning evolving smooth convex functions having block-coordinate-wise Lipschitz continuous gradients. The Block Coordinate Descent (BCD) method was among the first optimization schemes suggested for solving such problems [1]. In this work, we obtain a new lower (to our best knowledge the lowest currently) bound, which is 16p3 times smaller than the best known on the information-based complexity of BCD method. We achieve this by using an effective technique called Performance Estimation Problem (PEP) approach for analyzing the performance of first-order black box optimization methods. Numerical test confirms our analysis. Ziqiang Shi, Rujie Liu |
ICMLA | 2 |
| 2017 | Color-Introduced Frame-to-Model Registration for 3D Reconstruction
Fei Li 0005, Yunfan Du, Rujie Liu |
MMM (2) | 3 |
| 2015 | Multi-graph multi-instance learning with soft label consistency for object-based image retrievalabstractObject-based image retrieval has been an active research topic in the last decade, in which a user is only interested in some object instead of the whole image. As a promising approach, graph-based multi-instance learning has been paid much attention. Early retrieval methods often conduct learning on one graph in either image or region level. To further improve the performance, some recent methods adopt multi-graph learning, but the relationship between image- and region-level information is not well explored. In this paper, by constructing both image- and region-level graphs, a novel multi-graph multi-instance learning method is proposed. Different from the existing methods, the relationship between each labeled image and its segmented regions is reflected by the consistency of their corresponding soft labels, and it is formulated by the mutual restrictions in an optimization framework. A comprehensive cost function is designed to involve all the available information, and an iterative solution is introduced to solve the problem. Experimental results on the benchmark data set demonstrate the effectiveness of our proposal. Fei Li 0005, Rujie Liu |
ICME | 2 |
| 2015 | Online and Stochastic Universal Gradient Methods for Minimizing Regularized Hölder Continuous Finite Sums in Machine Learning
Ziqiang Shi, Rujie Liu |
PAKDD (1) | 2 |
| 2015 | Large Scale Optimization with Proximal Stochastic Newton-Type Gradient Descent
Ziqiang Shi, Rujie Liu |
ECML/PKDD (1) | 2 |
| 2015 | Fast Interactive Image Segmentation Using Bipartite Graph Based Random Walk with Restart
Yunfan Du, Fei Li 0005, Rujie Liu |
PSIVT | 3 |
| 2013 | Multi-SVM Multi-instance Learning for Object-Based Image Retrieval
Fei Li 0005, Rujie Liu, Takayuki Baba |
CAIP (1) | 2 |
| 2013 | Semi-supervised Learning for Large Scale Image CosegmentationabstractThis paper introduces to use semi-supervised learning for large scale image co segmentation. Different from traditional unsupervised cosegmentation that does not use any segmentation ground truth, semi-supervised cosegmentation exploits the similarity from both the very limited training image foregrounds, as well as the common object shared between the large number of unsegmented images. This would be a much practical way to effectively co segment a large number of related images simultaneously, where previous unsupervised co segmentation work poorly due to the large variances in appearance between different images and the lack of segmentation ground truth for guidance in co segmentation. For semi-supervised co segmentation in large scale, we propose an effective method by minimizing an energy function, which consists of the inter-image distance, the intra-image distance and the balance term. We also propose an iterative updating algorithm to efficiently solve this energy function, which decomposes the original energy minimization problem into sub-problems, and updates each image alternatively to reduce the number of variables in each sub-problem for computation efficiency. Experiment results on iCoseg and Pascal VOC datasets show that the proposed co segmentation method can effectively co segment hundreds of images in less than one minute. And our semi-supervised co segmentation is able to outperform both unsupervised co segmentation as well as fully supervised single image segmentation, especially when the training data is limited. Zhengxiang Wang, Rujie Liu |
ICCV | 2 |
| 2012 | Graph-based dimensionality reduction for KNN-based image annotation
Rujie Liu, Fei Li 0005, Qiong Cao |
ICPR | 2 |
| 2012 | A renewed image annotation baseline by image embedding and tag correlation
Rujie Liu, Yuehong Wang, Satoshi Naoi |
ICPR | 1 |
| 2012 | Multi-graph multi-instance learning for object-based image and video retrievalabstractObject-based image retrieval has been an active research topic in recent years, in which a user is only interested in some object in the images. As one promising approach, graph-based multi-instance learning has attracted many researchers. The existing methods often conduct learning on one graph, either in image level or in region level. While in this paper, by considering both image- and region-level information at the same time, a novel method based on multi-graph multi-instance learning is proposed. Two graphs are constructed in our method, and the relationship between each image and its segmented regions is introduced into an optimization framework. Moreover, our method is further extended to video retrieval. By exploring the relationships between video shots, representative images, and segmented regions, it can deal with the case when training labels are only assigned in shot level. Experimental results on the SIVAL image benchmark and the TRECVID video set demonstrate the effectiveness of our proposal. Fei Li 0005, Rujie Liu |
ICMR | 2 |
| 2011 | Graph-based multiple-instance learning with instance weighting for image retrievalabstractObject-based image retrieval has been an active research topic in recent years, in which user only pays his attention to some object in the images. As one promising approach, multiple-instance learning has attracted many researchers. Most of recently proposed methods either need additional restrictions for instance selection or lead to heavy computational load, so that they are often inconvenient for practical applications. In this paper, a novel method based on weighting regions in positive images is proposed, which mainly includes two steps of graph-based learning. The first step is only conducted on regions in training images, and different weights are efficiently set to each region in positive images based on the learning results. The second step is conducted on regions of all the database images, regions in positive images are fully utilized without selection, and ranking scores for each image are calculated. Experimental results demonstrate the effectiveness of our proposal. Fei Li 0005, Rujie Liu |
ICIP | 2 |
| 2010 | An automatic vehicle detection method based on traffic videosabstractA vision-based vehicle detection method is presented in this paper. The proposed method is composed of two steps, i.e., hypothesis generation and hypothesis verification. An adaptive background modeling and updating method is proposed to detect foreground regions in video sequences. With the prior knowledge of the vehicle appearance, the possible vehicle locations are extracted from the foreground regions and the touched vehicles are separated. Finally, hypothesized regions are verified by comparing their appearances with vehicle model. The performance of the proposed method is verified on videos captured under versatile conditions, and good results are achieved even in heavy traffic conditions. Qiong Cao, Rujie Liu, Fei Li 0005, Yuehong Wang |
ICIP | 2 |
| 2010 | Shape detection from line drawings with local neighborhood structure
Rujie Liu, Yuehong Wang, Takayuki Baba, Daiki Masumoto |
Pattern Recognit. | 1 |
| 2009 | Shape Detection from Line Drawings by Hierarchical Matching
Rujie Liu, Yuehong Wang, Takayuki Baba, Daiki Masumoto |
CAIP | 1 |
| 2008 | Semi-supervised learning by locally linear embedding in kernel spaceabstractGraph based semi-supervised learning methods (SSL) implicitly assume that the intrinsic geometry of the data points can be fully specified by an Euclidean distance based local neighborhood graph, however, this assumption may not always be necessarily true. To overcome this problem, we propose to apply locally linear embedding (LLE) method to characterize the geometric structure of the data points; besides this, the embedding process is performed in the kernel induced feature space rather than the original input space. After embedding, the proposed transductive learning method predicts the labels of the unlabeled data within the regularization framework. Experimental results on image retrieval and pattern recognition verify the performance of the proposed approach. Rujie Liu, Yuehong Wang, Takayuki Baba, Daiki Masumoto |
ICPR | 1 |
| 2008 | An Images-Based 3D Model Retrieval Approach
Yuehong Wang, Rujie Liu, Takayuki Baba, Yusuke Uehara, Daiki Masumoto, Shigemi Nagata |
MMM | 2 |
| 2008 | SVM-based active feedback in image retrieval using clustering and unlabeled data
Rujie Liu, Yuehong Wang, Takayuki Baba, Daiki Masumoto, Shigemi Nagata |
Pattern Recognit. | 1 |
| 2007 | SVM-Based Active Feedback in Image Retrieval Using Clustering and Unlabeled Data
Rujie Liu, Yuehong Wang, Takayuki Baba, Yusuke Uehara, Daiki Masumoto, Shigemi Nagata |
CAIP | 1 |
| 2005 | Similarity-based Partial Image Retrieval System for Engineering DrawingsabstractDesigners of mechanical products frequently refer to engineering drawings which are stored as image data in databases to design a new mechanical product efficiently. Multiple mechanical parts are usually drawn on each engineering drawing. Therefore designers want to find engineering drawings containing parts similar to a query image in the shape of a part drawn on an engineering drawing. In this paper, we propose a novel similarity based partial image retrieval system for engineering drawings. A unique aspect of this system is that a graph representation is utilized to robustly find engineering drawings containing similar parts which are invariant to the size, position, and rotation. We verified the performance for the similarity based partial image retrieval system through experiments using industrial engineering drawings. The results show that the top five similar engineering drawings for every query image are always accurately retrieved by our proposed system. This finding suggests that this system could be useful for the reuse of stored engineering drawings. Takayuki Baba, Rujie Liu, Susumu Endo, Shuichi Shiitani, Yusuke Uehara, Daiki Masumoto, Shigemi Nagata |
ISM | 2 |
| 2004 | Attributed Graph Matching Based Engineering Drawings Retrieval
Rujie Liu, Takayuki Baba, Daiki Masumoto |
Document Analysis Systems | 1 |