VLDB 2026 Research / reviewers in the wild / expert
Zhirong Yang
dblp:85/4582
· DBLP profile ↗
51ranked-venue papers
20as first author
15since 2021 · last 2026
0000-0001-8412-5684ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 45 · 19 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Stochastic approximation to contrastive learningabstractContrastive learning pulls positive samples (similar examples) closer and pushes negative samples (dissimilar examples) away to learn a meaningful mapping from inputs to outputs. It is found in a broad range of applications in computer vision that range from image classification to object detection to video processing and has a similarly broad impact and relevance in natural language processing, audio and speech, graphs, recommendation systems, and multimodal learning. Although contrastive learning done right typically achieves state-of-the-art results, there are major limitations with traditional methods because they rely on arbitrary definitions of positive and negative samples and are not decomposable in minibatch optimization. Thus, they require large batchsizes to effectively manage a tradeoff between positive and negative terms. This approach wastes significant computational resources on negative samples that may have minimal learning signals. To address these limitations, we propose a novel method that reformulates contrastive learning as a matrix approximation problem using I-divergence, a non-normalized variant of Kullback-Leibler divergence. Our objective function is decomposable across instance pairs, enabling efficient stochastic approximation algorithms that perform well with fewer negative samples by leveraging neighbor embeddings. Additionally, we generalize the scaling factor beyond standard normalization to adaptively emphasize positive samples with higher learning potential, reducing computational waste from negative samples. Our ambition is to demonstrate that even with a low batchsize and as few as one negative term per image, our method outperforms existing contrastive learning approaches and is competitive in overall performance to other state-of-the-art approaches on ImageNet which use a higher batchsize. Erland Brandser Olsson, Zhirong Yang |
Neural Networks | 2 |
| 2025 | Sensitive Label Replacement-Assisted Property Inference Attack Against Federated Learning on Server SideabstractFederated Learning (FL), as a distributed machine learning framework, allows each client to exchange only model parameters without uploading local data, effectively alleviating the problem of privacy leakage. However, FL has been proven to be vulnerable to property inference attack. Recent research on property inference attack in FL mainly focuses on clientside attack, assuming that the server side is honest and secure. This assumption is not entirely reliable in real-world applications. Therefore, in this paper, under the scenario where the server is assumed to be the adversary, we propose the Sensitive Label replacement-Assisted Property Inference attack on server side called SLAPI against FL. In this attack framework, we further assume that there is a spy client implementing a sensitive label replacement mechanism to assist the server for the attack. The spy client replaces the labels of the local dataset with those of sensitive property. This replacement of sensitive labels can distort the decision boundary of the global model, thereby inducing benign clients to disclose more information related to sensitive property and enhancing the potential discriminative ability of model parameters. We conduct experimental evaluations on realworld datasets. Compared to state-of-the-art attack schemes, the AUC score has achieved an enhancement of as much as 4 percentage points, fully verifying the effectiveness of the SLAPI attack. Longxing Zou, Tinghang Wang, Zhirong Yang |
HPCC | 4 |
| 2025 | Dynamic robustness evaluation for automated model selection in operationabstractContext: The increasing use of artificial neural network (ANN) classifiers in systems, especially safety-critical systems (SCSs), requires ensuring their robustness against out-of-distribution (OOD) shifts in operation, which are changes in the underlying data distribution from the data training the classifier. However, measuring the robustness of classifiers in operation with only unlabeled data is challenging. Additionally, machine learning engineers may need to compare different models or versions of the same model and switch to an optimal version based on their robustness. Objective: This paper explores the problem of dynamic robustness evaluation for automated model selection. We aim to find efficient and effective metrics for evaluating and comparing the robustness of multiple ANN classifiers using unlabeled operational data. Methods: To quantitatively measure the differences between the model outputs and assess robustness under OOD shifts using unlabeled data, we choose distance-based metrics. An empirical comparison of five such metrics, suitable for higher-dimensional data like images, is performed. The selected metrics include Wasserstein distance (WD), maximum mean discrepancy (MMD), Hellinger distance (HL), Kolmogorov–Smirnov statistic (KS), and Kullback–Leibler divergence (KL), known for their efficacy in quantifying distribution differences. We evaluate these metrics on 20 state-of-the-art models (ten CIFAR10-based models, five CIFAR100-based models, and five ImageNet-based models) from a widely used robustness benchmark ( RobustBench ) using data perturbed with various types and magnitudes of corruptions to mimic real-world OOD shifts. Results: Our findings reveal that the WD metric outperforms others when ranking multiple ANN models for CIFAR10- and CIFAR100-based models, while the KS metric demonstrates superior performance for ImageNet-based models. MMD can be used as a reliable second option for both datasets. Conclusion: This study highlights the effectiveness of distance-based metrics in ranking models’ robustness for automated model selection. It also emphasizes the significance of advancing research in dynamic robustness evaluation. Jingyue Li, Zhirong Yang |
Inf. Softw. Technol. | 3 |
| 2025 | Self-distillation improves self-supervised learning for DNA sequence inferenceabstractSelf-supervised Learning (SSL) has been recognized as a method to enhance prediction accuracy in various downstream tasks. However, its efficacy for DNA sequences remains somewhat constrained. This limitation stems primarily from the fact that most existing SSL approaches in genomics focus on masked language modeling of individual sequences, neglecting the crucial aspect of encoding statistics across multiple sequences. To overcome this challenge, we introduce an innovative deep neural network model, which incorporates collaborative learning between a 'student' and a 'teacher' subnetwork. In this model, the student subnetwork employs masked learning on nucleotides and progressively adapts its parameters to the teacher subnetwork through an exponential moving average approach. Concurrently, both subnetworks engage in contrastive learning, deriving insights from two augmented representations of the input sequences. This self-distillation process enables our model to effectively assimilate both contextual information from individual sequences and distributional data across the sequence population. We validated our approach with preliminary pretraining using the human reference genome, followed by applying it to 20 downstream inference tasks. The empirical results from these experiments demonstrate that our novel method significantly boosts inference performance across the majority of these tasks. Our code is available at https://github.com/wiedersehne/FinDNA. Ruslan Khalitov, Erland Brandser Olsson, Zhirong Yang |
Neural Networks | 5 |
| 2025 | A sparse and wide neural network model for DNA sequencesabstractAccurate modeling of DNA sequences requires capturing distant semantic relationships between the nucleotide acid bases. Most existing deep neural network models face two challenges: (1) they are limited to short DNA fragments and cannot capture long-range interactions, and (2) they require many supervised labels, which is often expensive in practice. We propose a new neural network model called SwanDNA to address the above challenges. By using a sparse and wide network architecture, our model enables inferences over very long DNA sequences. By incorporating the neural network into a self-supervised learning framework, our method can give accurate predictions while using less supervised labels. We evaluate SwanDNA in three DNA sequence inference tasks, human variant effect, open chromatin regions detection in plant genes, and GenomicBenchmarks. SwanDNA outperforms all competitors in the first two tasks and achieves state-of-art in seven of eight datasets in GenomicBenchmarks. Our code is available at https://github.com/wiedersehne/SwanDNA. Ruslan Khalitov, Zhirong Yang |
Neural Networks | 4 |
| 2024 | NLEBench+NorGLM: A Comprehensive Empirical Analysis and Benchmark Dataset for Generative Language Models in NorwegianabstractPeng Liu, Lemei Zhang, Terje Farup, Even W. Lauvrak, Jon Espen Ingvaldsen, Simen Eide, Jon Atle Gulla, Zhirong Yang. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Peng Liu 0025, Lemei Zhang, Terje Nissen Farup, Even W. Lauvrak, Jon Espen Ingvaldsen, Simen Eide, Jon Atle Gulla, Zhirong Yang |
EMNLP | 8 |
| 2024 | Self-supervised Learning for DNA sequences with circular dilated convolutional networks
Ruslan Khalitov, Zhirong Yang |
Neural Networks | 4 |
| 2023 | Simultaneous Wireless Information and Power Transfer in Terahertz Ultra-Massive MIMO SystemsabstractProviding links of rate over 100 Gbps and high energy density at the receiver, Terahertz (THz) simultaneous wireless information and power transfer (SWIPT) has huge potential to support the functionality of Internet of things (IoT) including distributed sensor networks and mobile robot systems. In this paper, a THz SWIPT system with ultra-massive MIMO (UM-MIMO) is proposed based on power splitting structure. By utilizing electromagnetic (EM) analysis, approximated closed-form expressions for energy distribution are derived to characterize the achievable rate and harvested energy of THz SWIPT system. Furthermore, the impact of finite-resolution phase shifter (FRPS) is analyzed, which cause the decline of energy density at the receiver and thus received power. Simulation results are provided to validate the derivations and demonstrate that decline of received power caused by FRPS can be locally compensated by increasing the inter-element spacing of antenna arrays. Moreover, compared with SWIPT at 1 THz and SWIPT at 28 GHz, SWIPT at 300 GHz balances the achievable rate and the harvested power, which allows the user to achieve rate of 200 Gbps and harvest power over 0.01 mW simultaneously. Zhirong Yang, Yongzhi Wu, Chong Han 0001 |
ICC | 1 |
| 2023 | ChordMixer: A Scalable Neural Attention Model for Sequences with Different Length
Ruslan Khalitov, Zhirong Yang |
ICLR | 4 |
| 2023 | Classification of long sequential data using circular dilated convolutional neural networksabstractClassification of long sequential data is an important Machine Learning task and appears in many application scenarios. Recurrent Neural Networks, Transformers, and Convolutional Neural Networks are three major techniques for learning from sequential data. Among these methods, Temporal Convolutional Networks (TCNs) which are scalable to very long sequences have achieved remarkable progress in time series regression. However, the performance of TCNs for sequence classification is not satisfactory because they use a skewed connection protocol and output classes at the last position. Such asymmetry restricts their performance for classification which depends on the whole sequence. In this work, we propose a symmetric multi-scale architecture called Circular Dilated Convolutional Neural Network (CDIL-CNN), where every position has an equal chance to receive information from other positions at the previous layers. Our model gives classification logits in all positions, and we can apply a simple ensemble learning to achieve a better decision. We have tested CDIL-CNN on various long sequential datasets. The experimental results show that our method has superior performance over many state-of-the-art approaches. The model and experiments are available at (https://github.com/LeiCheng-no/CDIL-CNN). Ruslan Khalitov, Zhirong Yang |
Neurocomputing | 5 |
| 2022 | Paramixer: Parameterizing Mixing Links in Sparse Factors Works Better than Dot-Product Self-AttentionabstractSelf-Attention is a widely used building block in neural modeling to mix long-range data elements. Most self-attention neural networks employ pairwise dot-products to specify the attention coefficients. However, these methods require O(N2) computing cost for sequence length N. Even though some approximation methods have been introduced to relieve the quadratic cost, the performance of the dot-product approach is still bottlenecked by the lowrank constraint in the attention matrix factorization. In this paper, we propose a novel scalable and effective mixing building block called Paramixer. Our method factorizes the interaction matrix into several sparse matrices, where we parameterize the non-zero entries by MLPs with the data elements as input. The overall computing cost of the new building block is as low as O(N log N). Moreover, all factorizing matrices in Paramixer are full-rank, so it does not suffer from the low-rank bottleneck. We have tested the new method on both synthetic and various real-world long sequential data sets and compared it with several state-of-the-art attention networks. The experimental results show that Paramixer has better performance in most learning tasks. Ruslan Khalitov, Zhirong Yang |
CVPR | 4 |
| 2022 | Distribution and gradient constrained embedding model for zero-shot learning with fewer seen samples
Jing Zhang 0058, Wen Wang 0019, Wenju Sun, Zhirong Yang, Qingyong Li |
Knowl. Based Syst. | 5 |
| 2022 | Sparse factorization of square matrices with application to neural attention modelingabstractSquare matrices appear in many machine learning problems and models. Optimization over a large square matrix is expensive in memory and in time. Therefore an economic approximation is needed. Conventional approximation approaches factorize the square matrix into a number matrices of much lower ranks. However, the low-rank constraint is a performance bottleneck if the approximated matrix is intrinsically high-rank or close to full rank. In this paper, we propose to approximate a large square matrix with a product of sparse full-rank matrices. In the approximation, our method needs only N(logN)2 non-zero numbers for an N×N full matrix. Our new method is especially useful for scalable neural attention modeling. Different from the conventional scaled dot-product attention methods, we train neural networks to map input data to the non-zero entries of the factorizing matrices. The sparse factorization method is tested for various square matrices, and the experimental results demonstrate that our method gives a better approximation when the approximated matrix is sparse and high-rank. As an attention module, our new method defeats Transformer and its several variants for long sequences in synthetic data sets and in the Long Range Arena benchmarks. Our code is publicly available2. Ruslan Khalitov, Zhirong Yang |
Neural Networks | 4 |
| 2022 | CIR-Net: Automatic Classification of Human Chromosome Based on Inception-ResNet ArchitectureabstractBACKGROUND: In medicine, karyotyping chromosomes is important for medical diagnostics, drug development, and biomedical research. Unfortunately, chromosome karyotyping is usually done by skilled cytologists manually, which requires experience, domain expertise, and considerable manual efforts. Therefore, automating the karyotyping process is a significant and meaningful task. METHOD: This paper focuses on chromosome classification because it is critical for chromosome karyotyping. In recent years, deep learning-based methods are the most promising methods for solving the tasks of chromosome classification. Although the deep learning-based Inception architecture has yielded state-of-the-art performance in the 2015 ILSVRC challenge, it has not been used in chromosome classification tasks so far. Therefore, we develop an automatic chromosome classification approach named CIR-Net based on Inception-ResNet which is an optimized version of Inception. However, the classification performance of origin Inception-ResNet on the insufficient chromosome dataset still has a lot of capacity for improvement. Further, we propose a simple but effective augmentation method called CDA for improving the performance of CIR-Net. RESULTS: The experimental results show that our proposed method achieves 95.98 percent classification accuracy on the clinical G-band chromosome dataset whose training dataset is insufficient. Moreover, the proposed augmentation method CDA improves more than 8.5 percent (from 87.46 to 95.98 percent) classification accuracy comparing to other methods. In this paper, the experimental results demonstrate that our proposed method is recent the most effective solution for solving clinical chromosome classification problems in chromosome auto-karyotyping on the condition of the insufficient training dataset. Code and Dataset are available at https://github.com/CloudDataLab/CIR-Net. Chengchuang Lin, Gansen Zhao, Zhirong Yang, Aihua Yin, Jianxin Wang 0001, Li Guo 0019, Hanbiao Chen, Zhaohui Ma, Haoyu Luo, Bichao Ding, Xiongwen Pang, Qiren Chen |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2021 | A novel chromosome cluster types identification method using ResNeXt WSL model
Chengchuang Lin, Gansen Zhao, Aihua Yin, Zhirong Yang, Li Guo 0019, Hanbiao Chen, Shuangyin Li, Haoyu Luo, Zhaohui Ma |
Medical Image Anal. | 4 |
| 2020 | A Simple and Novel Method to Predict the Hospital Energy Use Based on Machine Learning: A Case Study in Norway
Kai Xue, Yiyu Ding, Zhirong Yang, Natasa Nord, Mael Roger Albert Barillec, Hans Martin Mathisen, Tor Emil Giske, Liv I. Stenstad, Guangyu Cao |
ICONIP (4) | 3 |
| 2020 | Machine Learning for Hydropower Scheduling: State of the Art and Future Research DirectionsabstractThis paper investigates and discusses the current and future role of machine learning (ML) within the hydropower sector. An overview of the main applications of ML in the field of hydropower operations is presented to show the most common topics that have been addressed in the scientific literature in the last years. The objective is to provide recommendations for novel research directions that can be taken in the near future to cover those areas that have not been studied so far. The key contribution of this paper lies in a critical investigation of the state of the art of ML applications in hydropower scheduling. In light of the established literature available in the last years, this study identifies and discusses new roles that can be covered by ML, coupled with cyber-physical systems (CPSs), with a particular focus on short-term hydropower scheduling (STHS) challenges. Chiara Bordin, Hans Ivar Skjelbred, Jiehong Kong, Zhirong Yang |
KES | 4 |
| 2020 | A 91-Channel Hyperspectral LiDAR for Coal/Rock ClassificationabstractDuring the mining operation, it is a critical task in coal mines to significantly improve the safety by precision coal mining sorting and rock classification from different layers. It implies that a technique for rapidly and accurately classifying coal/rock in-site needs to be investigated and established, which is of significance for improving the coal mining efficiency and safety. In this letter, a 91-channel hyperspectral LiDAR (HSL) using an acousto-optic tunable filter (AOTF) as the spectroscopic device is designed, which operates based on the wide-spectrum emission laser source with a 5-nm spectral resolution to tackle this issue. The spectra of four-type coal/rock specimens collected by HSL are used to classify with three multi-label classifiers: naive Bayes (NB), logistic regression (LR), and support vector machine (SVM). Furthermore, we discuss and explore whether Gaussian fitting (GF) method and calibration with the reference whiteboard (RB) can enhance the classification accuracy. The experimental results show that the GF technique not only improves the accuracy of range measurement but also optimizes the classification performance using the spectra collected by the HSL. In addition, calibration with RB can improve classification accuracy as well. In addition, we also discuss methods to improve the calibration-free classification accuracy preliminarily. Yuwei Chen 0005, Zhirong Yang, Changhui Jiang, Wei Li 0095, Haohao Wu, Zhijie Wen, Eetu Puttonen, Juha Hyyppä |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2019 | Learning Image Relations with Contrast Association NetworksabstractInferring the relations between two images is an important class of tasks in computer vision. Examples of such tasks include computing optical flow and stereo disparity. We treat the relation inference tasks as a machine learning problem and tackle it with neural networks. A key to the problem is learning a representation of relations. We propose a new neural network module, contrast association unit (CAU), which explicitly models the relations between two sets of input variables. Due to the non-negativity of the weights in CAU, we adopt a multiplicative update algorithm for learning these weights. Experiments show that neural networks with CAUs are more effective in learning five fundamental image transformations than conventional neural networks. Yao Lu 0027, Zhirong Yang, Juho Kannala, Samuel Kaski |
IJCNN | 2 |
| 2019 | Doubly Stochastic Neighbor Embedding on SpheresabstractStochastic Neighbor Embedding (SNE) methods minimize the divergence between the similarity matrix of a high-dimensional data set and its counterpart from a low-dimensional embedding, leading to widely applied tools for data visualization. Despite their popularity, the current SNE methods experience a crowding problem when the data include highly imbalanced similarities. This implies that the data points with higher total similarity tend to get crowded around the display center. To solve this problem, we introduce a fast normalization method and normalize the similarity matrix to be doubly stochastic such that all the data points have equal total similarities. Furthermore, we show empirically and theoretically that the doubly stochasticity constraint often leads to embeddings which are approximately spherical. This suggests replacing a flat space with spheres as the embedding space. The spherical embedding eliminates the discrepancy between the center and the periphery in visualization, which efficiently resolves the crowding problem. We compared the proposed method (DOSNES) with the state-of-the-art SNE method on three real-world datasets and the results clearly indicate that our method is more favorable in terms of visualization quality. DOSNES is freely available at http://yaolubrain.github.io/dosnes/. Yao Lu 0027, Jukka Corander, Zhirong Yang |
Pattern Recognit. Lett. | 3 |
| 2018 | Word Embedding Based on Low-Rank Doubly Stochastic Matrix Decomposition
Denis Sedov, Zhirong Yang |
ICONIP (3) | 2 |
| 2016 | Low-Rank Doubly Stochastic Matrix Decomposition for Cluster AnalysisabstractCluster analysis by nonnegative low-rank approximations has experienced a remarkable progress in the past decade. However, the majority of such approximation approaches are still restricted to nonnegative matrix factorization (NMF) and suffer from the following two drawbacks: 1) they are unable to produce balanced partitions for large-scale manifold data which are common in real-world clustering tasks; 2) most existing NMF-type clustering methods cannot automatically determine the number of clusters. We propose a new low-rank learning method to address these two problems, which is beyond matrix factorization. Our method approximately decomposes a sparse input similarity in a normalized way and its objective can be used to learn both cluster assignments and the number of clusters. For efficient optimization, we use a relaxed formulation based on Data- Cluster-Data random walk, which is also shown to be equivalent to low-rank factorization of the doubly-stochastically normalized cluster incidence matrix. The probabilistic cluster assignments can thus be learned with a multiplicative majorization-minimization algorithm. Experimental results show that the new method is more accurate both in terms of clustering large-scale manifold data sets and of selecting the number of clusters. Zhirong Yang, Jukka Corander, Erkki Oja |
J. Mach. Learn. Res. | 1 |
| 2015 | Majorization-Minimization for Manifold EmbeddingabstractNonlinear dimensionality reduction by manifold embedding has become a popular and powerful approach both for visualization and as preprocessing for predictive tasks, but more efficient optimization algorithms are still crucially needed. Majorization-Minimization (MM) is a promising approach that monotonically decreases the cost function, but it remains unknown how to tightly majorize the manifold embedding objective functions such that the resulting MM algorithms are efficient and robust. We propose a new MM procedure that yields fast MM algorithms for a wide variety of manifold embedding problems. In our majorization step, two parts of the cost function are respectively upper bounded by quadratic and Lipschitz surrogates, and the resulting upper bound can be minimized in closed form. For cost functions amenable to such QL-majorization, the MM yields monotonic improvement and is efficient: in experiments the newly developed MM algorithms outperform five state-of-the-art optimization approaches in manifold embedding tasks. Zhirong Yang, Jaakko Peltonen, Samuel Kaski |
AISTATS | 1 |
| 2015 | Denoising Cluster Analysis
Ruqi Zhang, Zhirong Yang, Jukka Corander |
ICONIP (3) | 2 |
| 2015 | Understanding emotional impact of images using Bayesian multiple kernel learning
He Zhang 0009, Mehmet Gönen, Zhirong Yang, Erkki Oja |
Neurocomputing | 3 |
| 2015 | Learning the Information DivergenceabstractInformation divergence that measures the difference between two nonnegative matrices or tensors has found its use in a variety of machine learning problems. Examples are Nonnegative Matrix/Tensor Factorization, Stochastic Neighbor Embedding, topic models, and Bayesian network optimization. The success of such a learning task depends heavily on a suitable divergence. A large variety of divergences have been suggested and analyzed, but very few results are available for an objective choice of the optimal divergence for a given task. Here we present a framework that facilitates automatic selection of the best divergence among a given family, based on standard maximum likelihood estimation. We first propose an approximated Tweedie distribution for the β-divergence family. Selecting the best β then becomes a machine learning problem solved by maximum likelihood. Next, we reformulate α-divergence in terms of β-divergence, which enables automatic selection of α by maximum likelihood with reuse of the learning principle for β-divergence. Furthermore, we show the connections between γ- and β-divergences as well as Renyi- and α-divergences, such that our automatic selection framework is extended to non-separable divergences. Experiments on both synthetic and real-world data demonstrate that our method can quite accurately select information divergence across different learning problems and various divergence families. Onur Dikmen, Zhirong Yang, Erkki Oja |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2014 | Optimization Equivalence of Divergences Improves Neighbor EmbeddingabstractVisualization methods that arrange data objects in 2D or 3D layouts have followed two main schools, methods oriented for graph layout and methods oriented for vectorial embedding. We show the two previously separate approaches are tied by an optimization equivalence, making it possible to relate methods from the two approaches and to build new methods that take the best of both worlds. In detail, we prove a theorem of optimization equivalences between beta- and gamma-, as well as alpha- and Renyi-divergences through a connection scalar. Through the equivalences we represent several nonlinear dimensionality reduction and graph drawing methods in a generalized stochastic neighbor embedding setting, where information divergences are minimized between similarities in input and output spaces, and the optimal connection scalar provides a natural choice for the tradeoff between attractive and repulsive forces. We give two examples of developing new visualization methods through the equivalences: 1) We develop weighted symmetric stochastic neighbor embedding (ws-SNE) from Elastic Embedding and analyze its benefits, good performance for both vectorial and network data; in experiments ws-SNE has good performance across data sets of different types, whereas comparison methods fail for some of the data sets; 2) we develop a gamma-divergence version of a PolyLog layout method; the new method is scale invariant in the output space and makes it possible to efficiently use large-scale smoothed neighborhoods. Zhirong Yang, Jaakko Peltonen, Samuel Kaski |
ICML | 1 |
| 2014 | Adaptive multiplicative updates for quadratic nonnegative matrix factorization
He Zhang 0009, Zhirong Yang, Erkki Oja |
Neurocomputing | 2 |
| 2014 | Improving cluster analysis by co-initializations
He Zhang 0009, Zhirong Yang, Erkki Oja |
Pattern Recognit. Lett. | 2 |
| 2013 | Scalable Optimization of Neighbor Embedding for VisualizationabstractNeighbor embedding (NE) methods have found their use in data visualization but are limited in big data analysis tasks due to their O(n^2) complexity for n data samples. We demonstrate that the obvious approach of subsampling produces inferior results and propose a generic approximated optimization technique that reduces the NE optimization cost to O(n log n). The technique is based on realizing that in visualization the embedding space is necessarily very low-dimensional (2D or 3D), and hence efficient approximations developed for n-body force calculations can be applied. In gradient-based NE algorithms the gradient for an individual point decomposes into “forces” exerted by the other points. The contributions of close-by points need to be computed individually but far-away points can be approximated by their “center of mass”, rapidly computable by applying a recursive decomposition of the visualization space into quadrants. The new algorithm brings a significant speed-up for medium-size data, and brings “big data” within reach of visualization. Zhirong Yang, Jaakko Peltonen, Samuel Kaski |
ICML (2) | 1 |
| 2013 | Predicting Emotional States of Images Using Bayesian Multiple Kernel Learning
He Zhang 0009, Mehmet Gönen, Zhirong Yang, Erkki Oja |
ICONIP (3) | 3 |
| 2013 | Affective Abstract Image Classification and Retrieval Using Multiple Kernel Learning
He Zhang 0009, Zhirong Yang, Mehmet Gönen, Markus Koskela, Jorma Laaksonen, Timo Honkela, Erkki Oja |
ICONIP (3) | 2 |
| 2012 | Selecting β-Divergence for Nonnegative Matrix Factorization by Score Matching
Zhiyun Lu, Zhirong Yang, Erkki Oja |
ICANN (2) | 2 |
| 2012 | Pairwise Clustering with t-PLSI
He Zhang 0009, Tele Hao, Zhirong Yang, Erkki Oja |
ICANN (2) | 3 |
| 2012 | Clustering by Low-Rank Doubly Stochastic Matrix Decomposition
Zhirong Yang, Erkki Oja |
ICML | 1 |
| 2012 | Online Projective Nonnegative Matrix Factorization for Large Datasets
Zhirong Yang, He Zhang 0009, Erkki Oja |
ICONIP (3) | 1 |
| 2012 | Adaptive Multiplicative Updates for Projective Nonnegative Matrix Factorization
He Zhang 0009, Zhirong Yang, Erkki Oja |
ICONIP (3) | 2 |
| 2012 | Clustering by Nonnegative Matrix Factorization Using Graph Random WalkabstractNonnegative Matrix Factorization (NMF) is a promising relaxation technique for clustering analysis. However, conventional NMF methods that directly approximate the pairwise similarities using the least square error often yield mediocre performance for data in curved manifolds because they can capture only the immediate similarities between data samples. Here we propose a new NMF clustering method which replaces the approximated matrix with its smoothed version using random walk. Our method can thus accommodate farther relationships between data samples. Furthermore, we introduce a novel regularization in the proposed objective function in order to improve over spectral clustering. The new learning objective is optimized by a multiplicative Majorization-Minimization algorithm with a scalable implementation for learning the factorizing matrix. Extensive experimental results on real-world datasets show that our method has strong performance in terms of cluster purity. Zhirong Yang, Tele Hao, Onur Dikmen, Xi Chen 0017, Erkki Oja |
NIPS | 1 |
| 2012 | Quadratic nonnegative matrix factorization
Zhirong Yang, Erkki Oja |
Pattern Recognit. | 1 |
| 2011 | Kullback-Leibler Divergence for Nonnegative Matrix Factorization
Zhirong Yang, He Zhang 0009, Zhijian Yuan, Erkki Oja |
ICANN (1) | 1 |
| 2011 | Unified Development of Multiplicative Algorithms for Linear and Quadratic Nonnegative Matrix FactorizationabstractMultiplicative updates have been widely used in approximative nonnegative matrix factorization (NMF) optimization because they are convenient to deploy. Their convergence proof is usually based on the minimization of an auxiliary upper-bounding function, the construction of which however remains specific and only available for limited types of dissimilarity measures. Here we make significant progress in developing convergent multiplicative algorithms for NMF. First, we propose a general approach to derive the auxiliary function for a wide variety of NMF problems, as long as the approximation objective can be expressed as a finite sum of monomials with real exponents. Multiplicative algorithms with theoretical guarantee of monotonically decreasing objective function sequence can thus be obtained. The solutions of NMF based on most commonly used dissimilarity measures such as α- and β-divergence as well as many other more comprehensive divergences can be derived by the new unified principle. Second, our method is extended to a nonseparable case that includes e.g., γ-divergence and Rényi divergence. Third, we develop multiplicative algorithms for NMF using second-order approximative factorizations, in which each factorizing matrix may appear twice. Preliminary numerical experiments demonstrate that the multiplicative algorithms developed using the proposed procedure can achieve satisfactory Karush-Kuhn-Tucker optimality. We also demonstrate NMF problems where algorithms by the conventional method fail to guarantee descent at each iteration but those by our principle are immune to such violation. Zhirong Yang, Erkki Oja |
IEEE Trans. Neural Networks | 1 |
| 2010 | Linear and nonlinear projective nonnegative matrix factorizationabstractA variant of nonnegative matrix factorization (NMF) which was proposed earlier is analyzed here. It is called projective nonnegative matrix factorization (PNMF). The new method approximately factorizes a projection matrix, minimizing the reconstruction error, into a positive low-rank matrix and its transpose. The dissimilarity between the original data matrix and its approximation can be measured by the Frobenius matrix norm or the modified Kullback-Leibler divergence. Both measures are minimized by multiplicative update rules, whose convergence is proven for the first time. Enforcing orthonormality to the basic objective is shown to lead to an even more efficient update rule, which is also readily extended to nonlinear cases. The formulation of the PNMF objective is shown to be connected to a variety of existing NMF methods and clustering approaches. In addition, the derivation using Lagrangian multipliers reveals the relation between reconstruction and sparseness. For kernel principal component analysis (PCA) with the binary constraint, useful in graph partitioning problems, the nonlinear kernel PNMF provides a good approximation which outperforms an existing discretization approach. Empirical study on three real-world databases shows that PNMF can achieve the best or close to the best in clustering. The proposed algorithm runs more efficiently than the compared NMF methods, especially for high-dimensional data. Moreover, contrary to the basic NMF, the trained projection matrix can be readily used for newly coming samples and demonstrates good generalization. Zhirong Yang, Erkki Oja |
IEEE Trans. Neural Networks | 1 |
| 2009 | Projective Nonnegative Matrix Factorization with α-Divergence
Zhirong Yang, Erkki Oja |
ICANN (1) | 1 |
| 2009 | Adaptive Regularization for Transductive Support Vector MachineabstractWe discuss the framework of Transductive Support Vector Machine (TSVM) from the perspective of the regularization strength induced by the unlabeled data. In this framework, SVM and TSVM can be regarded as a learning machine without regularization and one with full regularization from the unlabeled data, respectively. Therefore, to supplement this framework of the regularization strength, it is necessary to introduce data-dependant partial regularization. To this end, we reformulate TSVM into a form with controllable regularization strength, which includes SVM and TSVM as special cases. Furthermore, we introduce a method of adaptive regularization that is data dependant and is based on the smoothness assumption. Experiments on a set of benchmark data sets indicate the promising results of the proposed work compared with state-of-the-art TSVM algorithms. Zenglin Xu, Rong Jin 0001, Jianke Zhu, Irwin King, Michael R. Lyu, Zhirong Yang |
NIPS | 6 |
| 2009 | Heavy-Tailed Symmetric Stochastic Neighbor EmbeddingabstractStochastic Neighbor Embedding (SNE) has shown to be quite promising for data visualization. Currently, the most popular implementation, t-SNE, is restricted to a particular Student t-distribution as its embedding distribution. Moreover, it uses a gradient descent algorithm that may require users to tune parameters such as the learning step size, momentum, etc., in finding its optimum. In this paper, we propose the Heavy-tailed Symmetric Stochastic Neighbor Embedding (HSSNE) method, which is a generalization of the t-SNE to accommodate various heavy-tailed embedding similarity functions. With this generalization, we are presented with two difficulties. The first is how to select the best embedding similarity among all heavy-tailed functions and the second is how to optimize the objective function once the heave-tailed function has been selected. Our contributions then are: (1) we point out that various heavy-tailed embedding similarities can be characterized by their negative score functions. Based on this finding, we present a parameterized subset of similarity functions for choosing the best tail-heaviness for HSSNE; (2) we present a fixed-point optimization algorithm that can be applied to all heavy-tailed functions and does not require the user to set any parameters; and (3) we present two empirical studies, one for unsupervised visualization showing that our optimization algorithm runs as fast and as good as the best known t-SNE implementation and the other for semi-supervised visualization showing quantitative superiority using the homogeneity measure as well as qualitative advantage in cluster separation over t-SNE. Zhirong Yang, Irwin King, Zenglin Xu, Erkki Oja |
NIPS | 1 |
| 2008 | Principal whitened gradient for information geometry
Zhirong Yang, Jorma Laaksonen |
Neural Networks | 1 |
| 2007 | Face Recognition Using Parzenfaces
Zhirong Yang, Jorma Laaksonen |
ICANN (2) | 1 |
| 2007 | Approximated Geodesic Updates with Principal Natural GradientsabstractWe propose a novel optimization algorithm which overcomes two drawbacks of Amari's natural gradient updates for information geometry. First, prewhitening the tangent vectors locally converts a Riemannian manifold to an Euclidean space so that the additive parameter update sequence approximates geodesics. Second, we prove that dimensionality reduction of natural gradients is necessary for learning multidimensional linear transformations. Removal of minor components also leads to noise reduction and better computational efficiency. The proposed method demonstrates faster and more robust convergence in the simulations on recovering a Gaussian mixture of artificial data and on discriminative learning of ionosphere data. Zhirong Yang, Jorma Laaksonen |
IJCNN | 1 |
| 2007 | Multiplicative updates for non-negative projections
Zhirong Yang, Jorma Laaksonen |
Neurocomputing | 1 |
| 2007 | Projective Non-Negative Matrix Factorization with Applications to Facial Image ProcessingabstractWe propose a new variant of Non-negative Matrix Factorization (NMF), including its model and two optimization rules. Our method is based on positively constrained projections and is related to the conventional SVD or PCA decomposition. The new model can potentially be applied to image compression and feature extraction problems. Of the latter, we consider processing of facial images, where each image consists of several parts and for each part the observations with different lighting mainly distribute along a straight line through the origin. No regularization terms are required in the objective functions and both suggested optimization rules can easily be implemented by matrix manipulations. The experiments show that the derived base vectors are spatially more localized than those of NMF. In turn, the better part-based representations improve the recognition rate of semantic classes such as the gender or existence of mustache in the facial images. Zhirong Yang, Zhijian Yuan, Jorma Laaksonen |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2006 | A Fast Fixed-Point Algorithm for Two-Class Discriminative Feature Extraction
Zhirong Yang, Jorma Laaksonen |
ICANN (2) | 1 |