VLDB 2026 Research / reviewers in the wild / expert
Stephen S.-T. Yau
dblp:y/StephenSTYau · also Stephen Shing-Toung Yau, Stephen Yau 0001
· DBLP profile ↗
15ranked-venue papers
0as first author
9since 2021 · last 2026
0000-0001-7634-7981ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6 · 5 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Theory of computation · 2Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Nonlinear Bayesian Filtering With Natural Gradient Gaussian ApproximationabstractPractical Bayes filters often assume the state distribution of each time step to be Gaussian for computational tractability, resulting in the so-called Gaussian filters. When facing nonlinear systems, Gaussian filters such as extended Kalman filter (EKF) or unscented Kalman filter (UKF) typically rely on certain linearization techniques, which can introduce large estimation errors. To address this issue, this paper reconstructs the prediction and update steps of Gaussian filtering as solutions to two distinct optimization problems, whose optimal conditions are found to have analytical forms from Stein's lemma. It is observed that the stationary point for the prediction step requires calculating the first two moments of the prior distribution, which is equivalent to that step in existing moment-matching filters. In the update step, instead of linearizing the model to approximate the stationary points, we propose an iterative approach to directly minimize the update step's objective to avoid linearization errors. For the purpose of performing the steepest descent on the Gaussian manifold, we derive its natural gradient that leverages Fisher information matrix to adjust the gradient direction, accounting for the curvature of the parameter space. Combining this update step with moment matching in the prediction step, we introduce a new iterative filter for nonlinear systems called Natural Gradient Gaussian Approximation filter, or NANO filter for short. We prove that NANO filter locally converges to the optimal Gaussian approximation at each time step. Furthermore, the estimation error is proven exponentially bounded for nearly linear measurement equation and low noise levels through constructing a supermartingale-like property across consecutive time steps. Real-world experiments demonstrate that, compared to popular Gaussian filters such as EKF, UKF, iterated EKF, and posterior linearization filter, NANO filter reduces the average root mean square error by approximately 45% while maintaining a comparable computational burden. Wenhan Cao, Zeju Sun, Chang Liu 0002, Stephen S.-T. Yau, Shengbo Eben Li |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | Energy entropy vector: a novel approach for efficient microbial genomic sequence analysis and classificationabstractWith the rapid development of genomic sequencing technologies, there is an increasing demand for efficient and accurate sequence analysis methods. However, existing methods face challenges in handling long, variable-length sequences and large-scale datasets. To address these issues, we propose a novel encoding method-Energy Entropy Vector (EEV). This method encodes gene sequences of arbitrary length into fixed-dimensional vector representations by modeling nucleotide energy characteristics based on information entropy. Experiments conducted on five microbial datasets demonstrate that, compared to traditional alignment-free methods, EEV achieves higher accuracy in convex hull classification and species classification tasks, with improvements of 15% to 30% in family-level classification. In phylogenetic tree construction, EEV significantly accelerates the process relative to multiple sequence alignment methods while maintaining high tree quality, enabling rapid and accurate phylogenetic reconstruction. Moreover, EEV supports flexible dimensional expansion by superimposing nucleotide energies, enhancing its ability to represent complex genomic sequences while effectively alleviating sparsity issues in high-dimensional representations. This study provides an efficient gene encoding strategy for large-scale genomic analysis and evolutionary research. Hao Wang 0262, Guoqing Hu, Stephen S.-T. Yau |
Briefings Bioinform. | 3 |
| 2025 | A new alignment-free method: K-mer Subsequence Natural Vector (K-mer SNV) for classification of fungiabstractAs eukaryotic organisms, fungi play a pivotal role within ecosystems and exert profound influences on agriculture, the pharmaceutical industry, and human health. The classification of fungi in databases has emerged as a crucial and complex issue in the field of biology. In this study, by leveraging the local distribution of k-mer in nucleotide sequences, we introduce a novel alignment-free method, denoted as k-mer SNV, to address this challenge. On a large fungi dataset including 120,140 sequences, our innovative approach has achieved remarkable success in predicting the taxonomic labels of fungi across six hierarchical taxonomic levels: phylum (99.52%), class (98.17%), order (97.20%), family (96.11%), genus (94.14%), and species (93.32%). The approach is also evaluated on the common Taxxi benchmark dataset. Based on these results, it has been convincingly demonstrated that the k-mer SNV method exhibits outstanding performance in processing large-scale fungal sequence data. Lily He, Mochao Huang, Gulinisha Yiming, Ruowei Liu, Jinghan Chen, Stephen S.-T. Yau |
BMC Bioinform. | 7 |
| 2025 | scMFF: a machine learning framework with multiple feature fusion strategies for cell type identificationabstractAccurate cell type classification is critical for downstream analysis in single-cell RNA sequencing (scRNA-seq). Most existing methods rely on a single type of feature representation-such as statistical, information theory, matrix factorization, or deep learning-based features. However, each captures different aspects of the data, and no single feature type can fully represent the complex differences between cell types. Moreover, naïvely concatenating multiple features may introduce redundancy or noise, reducing model performance. To address these challenges, we propose scMFF, which is a multiple feature fusion framework that integrates four features and explores six fusion strategies in combination with various classifiers for single-cell type classification. Comprehensive evaluations on 42 disease-related datasets and an external COVID-19 dataset demonstrate that scMFF outperforms single-feature approaches in terms of performance and stability, providing a reliable and effective solution for scRNA-seq data analysis. Nan Sun 0001, Dengcheng Yang, Rongling Wu, Stephen S.-T. Yau |
BMC Bioinform. | 6 |
| 2024 | CAPE: a deep learning framework with Chaos-Attention net for Promoter EvolutionabstractPredicting the strength of promoters and guiding their directed evolution is a crucial task in synthetic biology. This approach significantly reduces the experimental costs in conventional promoter engineering. Previous studies employing machine learning or deep learning methods have shown some success in this task, but their outcomes were not satisfactory enough, primarily due to the neglect of evolutionary information. In this paper, we introduce the Chaos-Attention net for Promoter Evolution (CAPE) to address the limitations of existing methods. We comprehensively extract evolutionary information within promoters using merged chaos game representation and process the overall information with modified DenseNet and Transformer structures. Our model achieves state-of-the-art results on two kinds of distinct tasks related to prokaryotic promoter strength prediction. The incorporation of evolutionary information enhances the model's accuracy, with transfer learning further extending its adaptability. Furthermore, experimental results confirm CAPE's efficacy in simulating in silico directed evolution of promoters, marking a significant advancement in predictive modeling for prokaryotic promoter strength. Our paper also presents a user-friendly website for the practical implementation of in silico directed evolution on promoters. The source code implemented in this study and the instructions on accessing the website can be found in our GitHub repository https://github.com/BobYHY/CAPE. Ruohan Ren, Hongyu Yu, Jiahao Teng, Sihui Mao, Zixuan Bian, Yangtianze Tao, Stephen S.-T. Yau |
Briefings Bioinform. | 7 |
| 2024 | Neural Projection Filter: Learning Unknown Dynamics Driven by Noisy ObservationsabstractIn this article, we propose the novel neural stochastic differential equations (SDEs) driven by noisy sequential observations called neural projection filter (NPF) under the continuous state-space models (SSMs) framework. The contributions of this work are both theoretical and algorithmic. On the one hand, we investigate the approximation capacity of the NPF, i.e., the universal approximation theorem for NPF. More explicitly, under some natural assumptions, we prove that the solution of the SDE driven by the semimartingale can be well approximated by the solution of the NPF. In particular, the explicit estimation bound is given. On the other hand, as an important application of this result, we develop a novel data-driven filter based on NPF. Also, under certain condition, we prove the algorithm convergence; i.e., the dynamics of NPF converges to the target dynamics. At last, we systematically compare the NPF with the existing filters. We verify the convergence theorem in linear case and experimentally demonstrate that the NPF outperforms existing filters in nonlinear case with robustness and efficiency. Furthermore, NPF could handle high-dimensional systems in real-time manner, even for the 100-D cubic sensor, while the state-of-the-art (SOTA) filter fails to do it. Yangtianze Tao, Jiayi Kang, Stephen S.-T. Yau |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Outlier-Robust Iterative Extended Kalman FilteringabstractIn this paper, we develop OR-IEKF which is a novel outlier-robust iterative extended Kalman filtering (IEKF) framework based on nonlinear regression formulation of update step. A new Kalman-type update step with reweighted prediction covariance and reweighted observation noise covariance are produced under the OR-IEKF framework, which could cut off the large outliers in observations causing by unknown outlier noises. By using various robust cost functions to solve such special nonlinear regression problems, we derive three algorithms. The performances of these new filters are evaluated in a nonlinear system simulation study. Yangtianze Tao, Stephen S.-T. Yau |
IEEE Signal Process. Lett. | 2 |
| 2023 | Recurrent Neural Networks Are Universal Approximators With Stochastic InputsabstractIn this article, we investigate the approximation ability of recurrent neural networks (RNNs) with stochastic inputs in state space model form. More explicitly, we prove that open dynamical systems with stochastic inputs can be well-approximated by a special class of RNNs under some natural assumptions, and the asymptotic approximation error has also been delicately analyzed as time goes to infinity. In addition, as an important application of this result, we construct an RNN-based filter and prove that it can well-approximate finite dimensional filters which include Kalman filter (KF) and Beneš filter as special cases. The efficiency of RNN-based filter has also been verified by two numerical experiments compared with optimal KF. Xiuqiong Chen, Yangtianze Tao, Stephen S.-T. Yau |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2022 | New Genome Sequence Detection via Natural Vector Convex Hull MethodabstractIt remains challenging how to find existing but undiscovered genome sequence mutations or predict potential genome sequence mutations based on real sequence data. Motivated by this, we develop approaches to detect new, undiscovered genome sequences. Because discovering new genome sequences through biological experiments is resource-intensive, we want to achieve the new genome sequence detection task mathematically. However, little literature tells us how to detect new, undiscovered genome sequence mutations mathematically. We form a new framework based on natural vector convex hull method that conducts alignment-free sequence analysis. Our newly developed two approaches, Random-permutation Algorithm with Penalty (RAP) and Random-permutation Algorithm with Penalty and COstrained Search (RAPCOS), use the geometry properties captured by natural vectors. In our experiment, we discover a mathematically new human immunodeficiency virus (HIV) genome sequence using some real HIV genome sequences. Significantly, the proposed methods are applicable to solve the new genome sequence detection challenge and have many good properties, such as robustness, rapid convergence, and fast computation. Ruzhang Zhao, Shaojun Pei, Stephen S.-T. Yau |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2012 | Decentralized Detection in Ad hoc Sensor Networks With Low Data Rate Inter Sensor CommunicationabstractDecentralized binary detection problem in ad-hoc sensor networks where a link between two sensors is on with a certain probability is considered in this paper. We propose a consensus based detection scheme where sensors exchange their local decisions, update their own decisions based on the exchanges and finally reach a consensus about the state of nature. We analyze the error probability and convergence of this decision consensus scheme. We show that with our scheme, the detection performance in ad-hoc networks is asymptotically equivalent to that of a parallel sensor network where all the local decisions are processed by a central node (fusion center) in the sense that the error exponents are the same. The probability distribution of the consensus time is also studied. Simulation and numerical results are given to verify the theoretical results. Yingwei Yao, Mo Deng, Stephen S.-T. Yau |
IEEE Trans. Inf. Theory | 4 |
| 2011 | DNA sequence comparison by a novel probabilistic method
Mo Deng, Stephen S.-T. Yau |
Inf. Sci. | 3 |
| 2008 | Numerical representation of DNA sequences based on genetic code context and its applications in periodicity analysis of genomesabstractThe indispensable prerequisites in characterizing information content of DNA molecules by computational methods are the numerical representations of symbolic DNA sequences. Current numerical representation methods for DNA sequences do not contain the genetic code context information, which may play an important role in defining protein coding regions. We propose a novel numerical representation of DNA sequences based on genetic code context within DNA sequences and explore the feasibility of applying this method to identify protein coding regions in genomes. Computational experiments indicate that incorporating genetic code information into numerical representations is a promising approach in which DNA sequences are uniquely represented and more information is represented so that digital processing tools can be applied to the periodicity analysis in DNA sequences effectively. Changchuan Yin, Stephen S.-T. Yau |
CIBCB | 2 |
| 2007 | Survey on index based homology search algorithms
Xianyang Jiang, Peiheng Zhang, Xinchun Liu, Stephen S.-T. Yau |
J. Supercomput. | 4 |
| 1997 | Contribution to Munuera's problem on the main conjecture of geometric hyperelliptic MDS codesabstractIn coding theory, it is of great interest to know the maximal length of MDS codes. In fact, the main conjecture says that the length of MDS codes over F/sub q/ is less than or equal to q+1 (except for some special cases). Munuera (see ibid., vol.38, p.1573-7, 1992) proposed a new way to attack the main conjecture on MDS codes for geometric codes. In particular, he proved the conjecture for codes arising from curves of genus one or two when the cardinal of the ground field is large enough. He also asked whether a similar theorem can be proved for any hyperelliptic curve. The purpose of this correspondence is to give an affirmative answer. In fact, our method also proves the main conjecture for geometric MDS codes for q=2 if the genus of the hyperelliptic curve is either 1, 2 or 3, and for q=3 if the genus of the curve is 1. Hao Chen 0096, Stephen S.-T. Yau |
IEEE Trans. Inf. Theory | 2 |
| 1996 | A Unified Approach to Iconic Indexing, Retrieval, and Maintenance of Spatial Relationships in Image Databases
Shi-Kuo Chang, Stephen S.-T. Yau |
J. Vis. Commun. Image Represent. | 3 |