VLDB 2026 Research / reviewers in the wild / expert
Shubhendu Trivedi
dblp:97/9735
· DBLP profile ↗
19ranked-venue papers
5as first author
10since 2021 · last 2025
0000-0003-1237-4301ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 2 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-authorHuman-computer interaction and ubiquitous computing · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Symmetry-Based Structured Matrices for Efficient Approximately Equivariant NetworksabstractThere has been much recent interest in designing neural networks (NNs) with relaxed equivariance, which interpolate between exact equivariance and full flexibility for consistent performance gains. In a separate line of work, structured parameter matrices with low displacement rank (LDR)—which permit fast function and gradient evaluation—have been used to create compact NNs, though primarily benefiting classical convolutional neural networks (CNNs). In this work, we propose a framework based on symmetry-based structured matrices to build approximately equivariant NNs with fewer parameters. Our approach unifies the aforementioned areas using Group Matrices (GMs), a forgotten precursor to the modern notion of regular representations of finite groups. GMs allow the design of structured matrices similar to LDR matrices, which can generalize all the elementary operations of a CNN from cyclic groups to arbitrary finite groups. We show GMs can also generalize classical LDR theory to general discrete groups, enabling a natural formalism for approximate equivariance. We test GM-based architectures on various tasks with relaxed symmetry and find that our framework performs competitively with approximately equivariant NNs and other structured matrix-based methods, often with one to two orders of magnitude fewer parameters. Ashwin Samudre, Mircea Petrache, Brian Nord, Shubhendu Trivedi |
AISTATS | 4 |
| 2025 | On Universality Classes of Equivariant NetworksabstractEquivariant neural networks provide a principled framework for incorporating symmetry into learning architectures and have been extensively analyzed through the lens of their *separation power*, that is, the ability to distinguish inputs modulo symmetry. This notion plays a central role in settings such as graph learning, where it is often formalized via the Weisfeilern–Leman hierarchy. In contrast, the *universality* of equivariant models—their capacity to approximate target functions—remains comparatively underexplored.
In this work, we investigate the approximation power of equivariant neural networks beyond separation constraints.
We show that separation power does not fully capture expressivity: models with identical separation power may differ in their approximation ability.
To demonstrate this, we characterize the universality classes of shallow invariant networks, providing a general framework for understanding which functions these architectures can approximate. Since equivariant models reduce to invariant ones under projection, this analysis yields sufficient conditions under which shallow equivariant networks fail to be universal.
Conversely, we identify settings where shallow models do achieve separation-constrained universality.
These positive results, however, depend critically on structural properties of the symmetry group, such as the existence of adequate normal subgroups, which may not hold in important cases like permutation symmetry. Marco Pacini, Gabriele Santin, Bruno Lepri, Shubhendu Trivedi |
NeurIPS | 4 |
| 2024 | Contextualized Sequence Likelihood: Enhanced Confidence Scores for Natural Language GenerationabstractThe advent of large language models (LLMs) has dramatically advanced the state-of-the-art in numerous natural language generation tasks.For LLMs to be applied reliably, it is essential to have an accurate measure of their confidence.Currently, the most commonly used confidence score function is the likelihood of the generated sequence, which, however, conflates semantic and syntactic components.For instance, in question-answering (QA) tasks, an awkward phrasing of the correct answer might result in a lower probability prediction.Additionally, different tokens should be weighted differently depending on the context.In this work, we propose enhancing the predicted sequence probability by assigning different weights to various tokens using attention values elicited from the base LLM.By employing a validation set, we can identify the relevant attention heads, thereby significantly improving the reliability of the vanilla sequence probability confidence measure.We refer to this new score as the Contextualized Sequence Likelihood (CSL).CSL is easy to implement, fast to compute, and offers considerable potential for further improvement with task-specific prompts.Across several QA datasets and a diverse array of LLMs, CSL has demonstrated significantly higher reliability than state-of-the-art baselines in predicting generation quality, as measured by the AUROC or AUARC. Zhen Lin 0001, Shubhendu Trivedi, Jimeng Sun 0001 |
EMNLP | 2 |
| 2024 | Improving Equivariant Model Training via Constraint RelaxationabstractEquivariant neural networks have been widely used in a variety of applications due to their ability to generalize well in tasks where the underlying data symmetries are known. Despite their successes, such networks can be difficult to optimize and require careful hyperparameter tuning to train successfully. In this work, we propose a novel framework for improving the optimization of such models by relaxing the hard equivariance constraint during training: We relax the equivariance constraint of the network's intermediate layers by introducing an additional non-equivariant term that we progressively constrain until we arrive at an equivariant solution. By controlling the magnitude of the activation of the additional relaxation term, we allow the model to optimize over a larger hypothesis space containing approximate equivariant networks and converge back to an equivariant solution at the end of training. We provide experimental results on different state-of-the-art network architectures, demonstrating how this training framework can result in equivariant models with improved generalization performance. Our code is available at https://github.com/StefanosPert/Equivariant_Optimization_CR Stefanos Pertigkiozoglou, Evangelos Chatzipantazis, Shubhendu Trivedi, Kostas Daniilidis |
NeurIPS | 3 |
| 2023 | Taking a Step Back with KCal: Multi-Class Kernel-Based Calibration for Deep Neural Networks
Zhen Lin 0001, Shubhendu Trivedi, Jimeng Sun 0001 |
ICLR | 2 |
| 2023 | Fast Online Value-Maximizing Prediction Sets with Conformal Cost ControlabstractMany real-world multi-label prediction problems involve set-valued predictions that must satisfy specific requirements dictated by downstream usage. We focus on a typical scenario where such requirements, separately encoding *value* and *cost*, compete with each other. For instance, a hospital might expect a smart diagnosis system to capture as many severe, often co-morbid, diseases as possible (the value), while maintaining strict control over incorrect predictions (the cost). We present a general pipeline, dubbed as FavMac, to maximize the value while controlling the cost in such scenarios. FavMac can be combined with almost any multi-label classifier, affording distribution-free theoretical guarantees on cost control. Moreover, unlike prior works, FavMac can handle real-world large-scale applications via a carefully designed online update mechanism, which is of independent interest. Our methodological and theoretical contributions are supported by experiments on several healthcare tasks and synthetic datasets - FavMac furnishes higher value compared with several variants and baselines while maintaining strict cost control. Zhen Lin 0001, Shubhendu Trivedi, Cao Xiao, Jimeng Sun 0001 |
ICML | 2 |
| 2023 | Approximation-Generalization Trade-offs under (Approximate) Group EquivarianceabstractThe explicit incorporation of task-specific inductive biases through symmetry has emerged as a general design precept in the development of high-performance machine learning models. For example, group equivariant neural networks have demonstrated impressive performance across various domains and applications such as protein and drug design. A prevalent intuition about such models is that the integration of relevant symmetry results in enhanced generalization. Moreover, it is posited that when the data and/or the model exhibits only approximate or partial symmetry, the optimal or best-performing model is one where the model symmetry aligns with the data symmetry. In this paper, we conduct a formal unified investigation of these intuitions. To begin, we present quantitative bounds that demonstrate how models capturing task-specific symmetries lead to improved generalization. Utilizing this quantification, we examine the more general question of dealing with approximate/partial symmetries. We establish, for a given symmetry group, a quantitative comparison between the approximate equivariance of the model and that of the data distribution, precisely connecting model equivariance error and data equivariance error. Our result delineates the conditions under which the model equivariance error is optimal, thereby yielding the best-performing model for the given task and data. Mircea Petrache, Shubhendu Trivedi |
NeurIPS | 2 |
| 2022 | Capacity of Group-invariant Linear Readouts from Equivariant Representations: How Many Objects can be Linearly Classified Under All Possible Views?
Matthew Farrell, Blake Bordelon, Shubhendu Trivedi, Cengiz Pehlevan |
ICLR | 3 |
| 2022 | Conformal Prediction with Temporal Quantile AdjustmentsabstractWe develop Temporal Quantile Adjustment (TQA), a general method to construct efficient and valid prediction intervals (PIs) for regression on cross-sectional time series data. Such data is common in many domains, including econometrics and healthcare. A canonical example in healthcare is predicting patient outcomes using physiological time-series data, where a population of patients composes a cross-section. Reliable PI estimators in this setting must address two distinct notions of coverage: cross-sectional coverage across a cross-sectional slice, and longitudinal coverage along the temporal dimension for each time series. Recent works have explored adapting Conformal Prediction (CP) to obtain PIs in the time series context. However, none handles both notions of coverage simultaneously. CP methods typically query a pre-specified quantile from the distribution of nonconformity scores on a calibration set. TQA adjusts the quantile to query in CP at each time $t$, accounting for both cross-sectional and longitudinal coverage in a theoretically-grounded manner. The post-hoc nature of TQA facilitates its use as a general wrapper around any time series regression model. We validate TQA's performance through extensive experimentation: TQA generally obtains efficient PIs and improves longitudinal coverage while preserving cross-sectional coverage. Zhen Lin 0001, Shubhendu Trivedi, Jimeng Sun 0001 |
NeurIPS | 2 |
| 2021 | Locally Valid and Discriminative Prediction Intervals for Deep Learning ModelsabstractCrucial for building trust in deep learning models for critical real-world applications is efficient and theoretically sound uncertainty quantification, a task that continues to be challenging. Useful uncertainty information is expected to have two key properties: It should be valid (guaranteeing coverage) and discriminative (more uncertain when the expected risk is high). Moreover, when combined with deep learning (DL) methods, it should be scalable and affect the DL model performance minimally. Most existing Bayesian methods lack frequentist coverage guarantees and usually affect model performance. The few available frequentist methods are rarely discriminative and/or violate coverage guarantees due to unrealistic assumptions. Moreover, many methods are expensive or require substantial modifications to the base neural network. Building upon recent advances in conformal prediction [13, 33] and leveraging the classical idea of kernel regression, we propose Locally Valid and Discriminative prediction intervals (LVD), a simple, efficient, and lightweight method to construct discriminative prediction intervals (PIs) for almost any DL model. With no assumptions on the data distribution, such PIs also offer finite-sample local coverage guarantees (contrasted to the simpler marginal coverage). We empirically verify, using diverse datasets, that besides being the only locally valid method for DL, LVD also exceeds or matches the performance (including coverage rate and prediction accuracy) of existing uncertainty quantification methods, while offering additional benefits in scalability and flexibility. Zhen Lin 0001, Shubhendu Trivedi, Jimeng Sun 0001 |
NeurIPS | 2 |
| 2018 | On the Generalization of Equivariance and Convolution in Neural Networks to the Action of Compact GroupsabstractConvolutional neural networks have been extremely successful in the image recognition domain because they ensure equivariance with respect to translations. There have been many recent attempts to generalize this framework to other domains, including graphs and data lying on manifolds. In this paper we give a rigorous, theoretical treatment of convolution and equivariance in neural networks with respect to not just translations, but the action of any compact group. Our main result is to prove that (given some natural constraints) convolutional structure is not just a sufficient, but also a necessary condition for equivariance to the action of a compact group. Our exposition makes use of concepts from representation theory and noncommutative harmonic analysis and derives new generalized convolution formulae. Risi Kondor, Shubhendu Trivedi |
ICML | 2 |
| 2018 | Clebsch-Gordan Nets: a Fully Fourier Space Spherical Convolutional Neural NetworkabstractRecent work by Cohen et al. has achieved state-of-the-art results for learning spherical images in a rotation invariant way by using ideas from group representation theory and noncommutative harmonic analysis. In this paper we propose a generalization of this work that generally exhibits improved performace, but from an implementation point of view is actually simpler. An unusual feature of the proposed architecture is that it uses the Clebsch--Gordan transform as its only source of nonlinearity, thus avoiding repeated forward and backward Fourier transforms. The underlying ideas of the paper generalize to constructing neural networks that are invariant to the action of other compact groups. Risi Kondor, Zhen Lin 0001, Shubhendu Trivedi |
NeurIPS | 3 |
| 2014 | Discriminative Metric Learning by Neighborhood Gerrymandering
Shubhendu Trivedi, David A. McAllester, Gregory Shakhnarovich |
NIPS | 1 |
| 2014 | A Consistent Estimator of the Expected Gradient Outerproduct
Shubhendu Trivedi, Samory Kpotufe, Gregory Shakhnarovich |
UAI | 1 |
| 2012 | The real world significance of performance prediction
Zachary A. Pardos, Qing Yang Wang, Shubhendu Trivedi |
EDM | 3 |
| 2012 | Co-Clustering by Bipartite Spectral Graph Partitioning for Out-of-Tutor Prediction
Shubhendu Trivedi, Zachary A. Pardos, Gábor N. Sárközy, Neil T. Heffernan |
EDM | 1 |
| 2012 | Clustered Knowledge Tracing
Zachary A. Pardos, Shubhendu Trivedi, Neil T. Heffernan, Gábor N. Sárközy |
ITS | 2 |
| 2011 | Clustering Students to Generate an Ensemble to Improve Standard Test Score Predictions
Shubhendu Trivedi, Zachary A. Pardos, Neil T. Heffernan |
AIED | 1 |
| 2011 | Spectral Clustering in Educational Data Mining
Shubhendu Trivedi, Zachary A. Pardos, Gábor N. Sárközy, Neil T. Heffernan |
EDM | 1 |