VLDB 2026 Research / reviewers in the wild / expert
Neil He
dblp:395/0217
· DBLP profile ↗
4ranked-venue papers
3as first author
4since 2021 · last 2025
0009-0008-3193-2448ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Deep learning architectures and training · 26% Generative modeling · 18% 3D vision · 18% |
Topics — the 12 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision › geometric deep learning
hyperbolic neural networks |
1.7 | 2 | 2025 | Hyperbolic Deep Learning for Foundation Models: A Survey · KDD (2) 2025 Lorentzian Residual Neural Networks · KDD (1) 2025 |
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction › manifold learning › geometric representation learning
hyperbolic representation learning |
1.7 | 2 | 2025 | HELM: Hyperbolic Large Language Models via Mixture-of-Curvature Experts · NeurIPS 2025 Hyperbolic Deep Learning for Foundation Models: A Survey · KDD (2) 2025 |
Machine learning › Generative modeling
diffusion model |
0.9 | 1 | 2025 | Efficient Diffusion Models for Symmetric Manifolds · ICML 2025 |
Machine learning › Efficient and distributed learning
efficient training |
0.9 | 1 | 2025 | Efficient Diffusion Models for Symmetric Manifolds · ICML 2025 |
Machine learning › Deep learning architectures and training
foundation model |
0.9 | 1 | 2025 | Hyperbolic Deep Learning for Foundation Models: A Survey · KDD (2) 2025 |
Natural language and speech › Language models and text generation
large language model |
0.9 | 1 | 2025 | HELM: Hyperbolic Large Language Models via Mixture-of-Curvature Experts · NeurIPS 2025 |
Machine learning › Deep learning architectures and training
mixture of experts |
0.9 | 1 | 2025 | HELM: Hyperbolic Large Language Models via Mixture-of-Curvature Experts · NeurIPS 2025 |
Machine learning › Generative modeling › diffusion model › geometric diffusion model
riemannian diffusion model |
0.9 | 1 | 2025 | Efficient Diffusion Models for Symmetric Manifolds · ICML 2025 |
Machine learning › Deep learning architectures and training
attention mechanism |
0.3 | 1 | 2025 | HELM: Hyperbolic Large Language Models via Mixture-of-Curvature Experts · NeurIPS 2025 |
Machine learning › Deep learning architectures and training
convolutional neural network |
0.3 | 1 | 2025 | Lorentzian Residual Neural Networks · KDD (1) 2025 |
Machine learning › Graph learning
graph neural network |
0.3 | 1 | 2025 | Lorentzian Residual Neural Networks · KDD (1) 2025 |
Machine learning › Deep learning architectures and training
transformer |
0.3 | 1 | 2025 | HELM: Hyperbolic Large Language Models via Mixture-of-Curvature Experts · NeurIPS 2025 |
Methods — techniques the papers use, named apart from their topics
hyperbolic geometry · 1.7weighted lorentzian centroid · 0.9rotary positional encoding · 0.9residual connections · 0.9projection of euclidean brownian motion · 0.9non-euclidean embeddings · 0.9multi-head latent attention · 0.9ito's lemma · 0.9RMS normalization · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Efficient Diffusion Models for Symmetric ManifoldsabstractWe introduce a framework for designing efficient diffusion models for $d$-dimensional symmetric-space Riemannian manifolds, including the torus, sphere, special orthogonal group and unitary group. Existing manifold diffusion models often depend on heat kernels, which lack closed-form expressions and require either $d$ gradient evaluations or exponential-in-$d$ arithmetic operations per training step. We introduce a new diffusion model for symmetric manifolds with a spatially-varying covariance, allowing us to leverage a projection of Euclidean Brownian motion to bypass heat kernel computations. Our training algorithm minimizes a novel efficient objective derived via Ito’s Lemma, allowing each step to run in $O(1)$ gradient evaluations and nearly-linear-in-$d$ ($O(d^{1.19})$) arithmetic operations, reducing the gap between diffusions on symmetric manifolds and Euclidean space. Manifold symmetries ensure the diffusion satisfies an "average-case" Lipschitz condition, enabling accurate and efficient sample generation. Empirically, our model outperforms prior methods in training speed and improves sample quality on synthetic datasets on the torus, special orthogonal group, and unitary group. Oren Mangoubi, Neil He, Nisheeth K. Vishnoi |
ICML | 2 |
| 2025 | Lorentzian Residual Neural NetworksabstractHyperbolic neural networks have emerged as a powerful tool for modeling hierarchical data structures prevalent in real-world datasets. Notably, residual connections, which facilitate the direct flow of information across layers, have been instrumental in the success of deep neural networks. However, current methods for constructing hyperbolic residual networks suffer from limitations such as increased model complexity, numerical instability, and errors due to multiple mappings to and from the tangent space. To address these limitations, we introduce LResNet, a novel Lorentzian residual neural network based on the weighted Lorentzian centroid in the Lorentz model of hyperbolic geometry. Our method enables the efficient integration of residual connections in Lorentz hyperbolic neural networks while preserving their hierarchical representation capabilities. We demonstrate that our method can theoretically derive previous methods while offering improved stability, efficiency, and effectiveness. Extensive experiments on both graph and vision tasks showcase the superior performance and robustness of our method compared to state-of-the-art Euclidean and hyperbolic alternatives. Our findings highlight the potential of LResNet for building more expressive neural networks in hyperbolic embedding space as a generally applicable method to multiple architectures, including CNNs, GNNs, and graph Transformers. Neil He, Menglin Yang 0001, Rex Ying |
KDD (1) | 1 |
| 2025 | Hyperbolic Deep Learning for Foundation Models: A SurveyabstractFoundation models pre-trained on massive datasets, including large language models (LLMs), vision-language models (VLMs), and large multimodal models, have demonstrated remarkable success in diverse downstream tasks. However, recent studies have shown fundamental limitations of these models: (1) limited representational capacity(2) lower adaptability, and (3) diminishing scalability. These shortcomings raise a critical question: is Euclidean geometry truly the optimal inductive bias for all foundation models, or could incorporating alternative geometric spaces enable models to better align with the intrinsic structure of real-world data and improve reasoning processes? Hyperbolic spaces, a class of non-Euclidean manifolds characterized by exponential volume growth with respect to distance, offer a mathematically grounded solution. These spaces enable low-distortion embeddings of hierarchical structures (e.g., trees, taxonomies) and power-law distributions with substantially fewer dimensions compared to Euclidean counterparts. Recent advances have leveraged these properties to enhance foundation models, including improving LLMs' complex reasoning ability, VLMs' zero-shot generalization, and cross-modal semantic alignment, while maintaining parameter efficiency. This paper provides a comprehensive review of hyperbolic neural networks and their recent development for foundation models. We further outline key challenges and research directions to advance the field. Neil He, Hiren Madhu, Ngoc Bui, Menglin Yang 0001, Rex Ying |
KDD (2) | 1 |
| 2025 | HELM: Hyperbolic Large Language Models via Mixture-of-Curvature ExpertsabstractFrontier large language models (LLMs) have shown great success in text modeling and generation tasks across domains. However, natural language exhibits inherent semantic hierarchies and nuanced geometric structure, which current LLMs do not capture completely owing to their reliance on Euclidean operations such as dot-products and norms. Furthermore, recent studies have shown that not respecting the underlying geometry of token embeddings leads to training instabilities and degradation of generative capabilities. These findings suggest that shifting to non-Euclidean geometries can better align language models with the underlying geometry of text. We thus propose to operate fully in $\textit{Hyperbolic space}$, known for its expansive, scale-free, and low-distortion properties. To this end, we introduce $\textbf{HELM}$, a family of $\textbf{H}$yp$\textbf{E}$rbolic Large $\textbf{L}$anguage $\textbf{M}$odels, offering a geometric rethinking of the Transformer-based LLM that addresses the representational inflexibility, missing set of necessary operations, and poor scalability of existing hyperbolic LMs. We additionally introduce a $\textbf{Mi}$xture-of-$\textbf{C}$urvature $\textbf{E}$xperts model, $\textbf{HELM-MiCE}$, where each expert operates in a distinct curvature space to encode more fine-grained geometric structure from text, as well as a dense model, $\textbf{HELM-D}$. For $\textbf{HELM-MiCE}$, we further develop hyperbolic Multi-Head Latent Attention ($\textbf{HMLA}$) for efficient, reduced-KV-cache training and inference. For both models, we further develop essential hyperbolic equivalents of rotary positional encodings and root mean square normalization. We are the first to train fully hyperbolic LLMs at billion-parameter scale, and evaluate them on well-known benchmarks such as MMLU and ARC, spanning STEM problem-solving, general knowledge, and commonsense reasoning. Our results show consistent gains from our $\textbf{HELM}$ architectures – up to 4\% – over popular Euclidean architectures used in LLaMA and DeepSeek with superior semantic hierarchy modeling capabilities, highlighting the efficacy and enhanced reasoning afforded by hyperbolic geometry in large-scale language model pretraining. Neil He, Rishabh Anand, Hiren Madhu, Ali Maatouk, Smita Krishnaswamy, Leandros Tassiulas, Menglin Yang 0001, Rex Ying |
NeurIPS | 1 |