Hiren Madhu

dblp:287/7884 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
6since 2021 · last 2025
0000-0002-6701-6782ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Representation and self-supervised learning · 35% Graph learning · 21% Deep learning architectures and training · 20%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%

Topics — the 15 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction › manifold learning › geometric representation learning
hyperbolic representation learning
1.722025
HELM: Hyperbolic Large Language Models via Mixture-of-Curvature Experts · NeurIPS 2025
Hyperbolic Deep Learning for Foundation Models: A Survey · KDD (2) 2025
Machine learning › Graph learning › geometric learning › topological deep learning
simplicial neural network
1.622025
HiPoNet: A Multi-View Simplicial Complex Network for High Dimensional Point-Cloud and Single-Cell data · NeurIPS 2025
Unsupervised Parameter-free Simplicial Representation Learning with Scattering Transforms · ICML 2024
Machine learning › Deep learning architectures and training
foundation model
0.912025
Hyperbolic Deep Learning for Foundation Models: A Survey · KDD (2) 2025
Computer vision › 3D vision › geometric deep learning
hyperbolic neural networks
0.912025
Hyperbolic Deep Learning for Foundation Models: A Survey · KDD (2) 2025
Natural language and speech › Language models and text generation
large language model
0.912025
HELM: Hyperbolic Large Language Models via Mixture-of-Curvature Experts · NeurIPS 2025
Machine learning › Deep learning architectures and training
mixture of experts
0.912025
HELM: Hyperbolic Large Language Models via Mixture-of-Curvature Experts · NeurIPS 2025
Computer vision › 3D vision › point cloud analysis › point cloud learning
point cloud representation learning
0.912025
HiPoNet: A Multi-View Simplicial Complex Network for High Dimensional Point-Cloud and Single-Cell data · NeurIPS 2025
Machine learning › Representation and self-supervised learning
scattering transform
0.812024
Unsupervised Parameter-free Simplicial Representation Learning with Scattering Transforms · ICML 2024
Machine learning › Representation and self-supervised learning › representation learning
unsupervised representation learning
0.812024
Unsupervised Parameter-free Simplicial Representation Learning with Scattering Transforms · ICML 2024
Machine learning › Graph learning
graph self-supervised learning
0.712023
TopoSRL: Topology preserving self-supervised Simplicial Representation Learning · NeurIPS 2023
Machine learning › Representation and self-supervised learning › representation learning › topological representation learning
topology-preserving representation
0.712023
TopoSRL: Topology preserving self-supervised Simplicial Representation Learning · NeurIPS 2023
Machine learning › Deep learning architectures and training
attention mechanism
0.312025
HELM: Hyperbolic Large Language Models via Mixture-of-Curvature Experts · NeurIPS 2025
Machine learning › Deep learning architectures and training
transformer
0.312025
HELM: Hyperbolic Large Language Models via Mixture-of-Curvature Experts · NeurIPS 2025
Bioinformatics and computational biology
single-cell analysis
0.312025
HiPoNet: A Multi-View Simplicial Complex Network for High Dimensional Point-Cloud and Single-Cell data · NeurIPS 2025
Bioinformatics and computational biology › transcriptomics
spatial transcriptomics
0.312025
HiPoNet: A Multi-View Simplicial Complex Network for High Dimensional Point-Cloud and Single-Cell data · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

simplicial wavelet transform · 1.7differentiable neural network · 1.7hyperbolic geometry · 1.7rotary positional encoding · 0.9non-euclidean embeddings · 0.9multi-head latent attention · 0.9RMS normalization · 0.9random walk matrices · 0.8filter bank networks · 0.8contrastive learning · 0.7
YearPublicationVenuePosition
2025 Hyperbolic Deep Learning for Foundation Models: A Survey
abstract
Foundation models pre-trained on massive datasets, including large language models (LLMs), vision-language models (VLMs), and large multimodal models, have demonstrated remarkable success in diverse downstream tasks. However, recent studies have shown fundamental limitations of these models: (1) limited representational capacity(2) lower adaptability, and (3) diminishing scalability. These shortcomings raise a critical question: is Euclidean geometry truly the optimal inductive bias for all foundation models, or could incorporating alternative geometric spaces enable models to better align with the intrinsic structure of real-world data and improve reasoning processes? Hyperbolic spaces, a class of non-Euclidean manifolds characterized by exponential volume growth with respect to distance, offer a mathematically grounded solution. These spaces enable low-distortion embeddings of hierarchical structures (e.g., trees, taxonomies) and power-law distributions with substantially fewer dimensions compared to Euclidean counterparts. Recent advances have leveraged these properties to enhance foundation models, including improving LLMs' complex reasoning ability, VLMs' zero-shot generalization, and cross-modal semantic alignment, while maintaining parameter efficiency. This paper provides a comprehensive review of hyperbolic neural networks and their recent development for foundation models. We further outline key challenges and research directions to advance the field.
Neil He, Hiren Madhu, Ngoc Bui, Menglin Yang 0001, Rex Ying
KDD (2)2
2025 HELM: Hyperbolic Large Language Models via Mixture-of-Curvature Experts
abstract
Frontier large language models (LLMs) have shown great success in text modeling and generation tasks across domains. However, natural language exhibits inherent semantic hierarchies and nuanced geometric structure, which current LLMs do not capture completely owing to their reliance on Euclidean operations such as dot-products and norms. Furthermore, recent studies have shown that not respecting the underlying geometry of token embeddings leads to training instabilities and degradation of generative capabilities. These findings suggest that shifting to non-Euclidean geometries can better align language models with the underlying geometry of text. We thus propose to operate fully in $\textit{Hyperbolic space}$, known for its expansive, scale-free, and low-distortion properties. To this end, we introduce $\textbf{HELM}$, a family of $\textbf{H}$yp$\textbf{E}$rbolic Large $\textbf{L}$anguage $\textbf{M}$odels, offering a geometric rethinking of the Transformer-based LLM that addresses the representational inflexibility, missing set of necessary operations, and poor scalability of existing hyperbolic LMs. We additionally introduce a $\textbf{Mi}$xture-of-$\textbf{C}$urvature $\textbf{E}$xperts model, $\textbf{HELM-MiCE}$, where each expert operates in a distinct curvature space to encode more fine-grained geometric structure from text, as well as a dense model, $\textbf{HELM-D}$. For $\textbf{HELM-MiCE}$, we further develop hyperbolic Multi-Head Latent Attention ($\textbf{HMLA}$) for efficient, reduced-KV-cache training and inference. For both models, we further develop essential hyperbolic equivalents of rotary positional encodings and root mean square normalization. We are the first to train fully hyperbolic LLMs at billion-parameter scale, and evaluate them on well-known benchmarks such as MMLU and ARC, spanning STEM problem-solving, general knowledge, and commonsense reasoning. Our results show consistent gains from our $\textbf{HELM}$ architectures – up to 4\% – over popular Euclidean architectures used in LLaMA and DeepSeek with superior semantic hierarchy modeling capabilities, highlighting the efficacy and enhanced reasoning afforded by hyperbolic geometry in large-scale language model pretraining.
Neil He, Rishabh Anand, Hiren Madhu, Ali Maatouk, Smita Krishnaswamy, Leandros Tassiulas, Menglin Yang 0001, Rex Ying
NeurIPS3
2025 HiPoNet: A Multi-View Simplicial Complex Network for High Dimensional Point-Cloud and Single-Cell data
abstract
In this paper, we propose HiPoNet, an end-to-end differentiable neural network for regression, classification, and representation learning on high-dimensional point clouds. Our work is motivated by single-cell data which can have very high-dimensionality --exceeding the capabilities of existing methods for point clouds which are mostly tailored for 3D data. Moreover, modern single-cell and spatial experiments now yield entire cohorts of datasets (i.e., one data set for every patient), necessitating models that can process large, high-dimensional point-clouds at scale. Most current approaches build a single nearest-neighbor graph, discarding important geometric and topological information. In contrast, HiPoNet models the point-cloud as a set of higher-order simplicial complexes, with each particular complex being created using a reweighting of features. This method thus generates multiple constructs corresponding to different views of high-dimensional data, which in biology offers the possibility of disentangling distinct cellular processes. It then employs simplicial wavelet transforms to extract multiscale features, capturing both local and global topology from each view. We show that geometric and topological information is preserved in this framework both theoretically and empirically. We showcase the utility of HiPoNet on point-cloud level tasks, involving classification and regression of entire point-clouds in data cohorts. Experimentally, we find that HiPoNet outperforms other point-cloud and graph-based models on single-cell data. We also apply HiPoNet to spatial transcriptomics datasets using spatial coordinates as one of the views. Overall, HiPoNet offers a robust and scalable solution for high-dimensional data analysis.
Siddharth Viswanath, Hiren Madhu, Dhananjay Bhaskar, Jake Kovalic, Dave Johnson 0004, Christopher J. Tape, Ian Adelstein, Rex Ying, Michael Perlmutter, Smita Krishnaswamy
NeurIPS2
2024 Unsupervised Parameter-free Simplicial Representation Learning with Scattering Transforms
abstract
Simplicial neural network models are becoming popular for processing and analyzing higher-order graph data, but they suffer from high training complexity and dependence on task-specific labels. To address these challenges, we propose simplicial scattering networks (SSNs), a parameter-free model inspired by scattering transforms designed to extract task-agnostic features from simplicial complex data without labels in a principled manner. Specifically, we propose a simplicial scattering transform based on random walk matrices for various adjacencies underlying a simplicial complex. We then use the simplicial scattering transform to construct a deep filter bank network that captures high-frequency information at multiple scales. The proposed simplicial scattering transform possesses properties such as permutation invariance, robustness to perturbations, and expressivity. We theoretically prove that including higher-order information improves the robustness of SSNs to perturbations. Empirical evaluations demonstrate that SSNs outperform existing simplicial or graph neural models in many tasks like node classification, simplicial closure, graph classification, trajectory prediction, and simplex prediction while being computationally efficient.
Hiren Madhu, Sravanthi Gurugubelli, Sundeep Prabhakar Chepuri
ICML1
2023 TopoSRL: Topology preserving self-supervised Simplicial Representation Learning
abstract
In this paper, we introduce $\texttt{TopoSRL}$, a novel self-supervised learning (SSL) method for simplicial complexes to effectively capture higher-order interactions and preserve topology in the learned representations. $\texttt{TopoSRL}$ addresses the limitations of existing graph-based SSL methods that typically concentrate on pairwise relationships, neglecting long-range dependencies crucial to capture topological information. We propose a new simplicial augmentation technique that generates two views of the simplicial complex that enriches the representations while being efficient. Next, we propose a new simplicial contrastive loss function that contrasts the generated simplices to preserve local and global information present in the simplicial complexes. Extensive experimental results demonstrate the superior performance of $\texttt{TopoSRL}$ compared to state-of-the-art graph SSL techniques and supervised simplicial neural models across various datasets corroborating the efficacy of $\texttt{TopoSRL}$ in processing simplicial complex data in a self-supervised setting.
Hiren Madhu, Sundeep Prabhakar Chepuri
NeurIPS1
2023 Detecting offensive speech in conversational code-mixed dialogue on social media: A contextual dataset and benchmark experiments
Hiren Madhu, Shrey Satapara, Sandip Modha, Thomas Mandl 0001, Prasenjit Majumder
Expert Syst. Appl.1