Shib Sankar Dasgupta

dblp:222/9398 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 4 first-author · 5 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Representation and self-supervised learning · 54% Graph learning · 24% Information extraction and text analysis · 10%
Databases, data mining, and information retrieval
2 papers
Knowledge graphs · 79% Machine learning and data management · 21%

Topics — the 13 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Representation and self-supervised learning › representation learning › embedding learning › geometric embedding
box embedding
1.932025
A Geometric Approach to Personalized Recommendation with Set-Theoretic Constraints Using Box Embeddings · ICML 2025
Word2Box: Capturing Set-Theoretic Semantics of Words using Box Embeddings · ACL (1) 2022
Improving Local Identifiability in Probabilistic Box Embeddings · NeurIPS 2020
Machine learning › Graph learning › directed graph learning
directed graph representation
0.812024
Learning Representations for Hierarchies with Minimal Support · NeurIPS 2024
Machine learning › Representation and self-supervised learning › representation learning › embedding learning
hierarchical embedding
0.812024
Learning Representations for Hierarchies with Minimal Support · NeurIPS 2024
Machine learning › Graph learning
network embedding
0.812024
Learning Representations for Hierarchies with Minimal Support · NeurIPS 2024
Natural language and speech › Information extraction and text analysis
temporal information extraction
0.722018
AD3: Attentive Deep Document Dater · EMNLP 2018
Dating Documents using Graph Convolution Networks · ACL (1) 2018
Machine learning › Representation and self-supervised learning › word representation
word embedding
0.612022
Word2Box: Capturing Set-Theoretic Semantics of Words using Box Embeddings · ACL (1) 2022
Machine learning › Representation and self-supervised learning › representation learning › embedding learning
geometric embedding
0.412020
Improving Local Identifiability in Probabilistic Box Embeddings · NeurIPS 2020
Knowledge, reasoning and agents › Knowledge representation and reasoning
knowledge graph
0.412020
Improving Local Identifiability in Probabilistic Box Embeddings · NeurIPS 2020
Knowledge graphs
knowledge graph embedding
0.312018
HyTE: Hyperplane-based Temporally aware Knowledge Graph Embedding · EMNLP 2018
Knowledge graphs › link prediction
temporal knowledge graph completion
0.312018
HyTE: Hyperplane-based Temporally aware Knowledge Graph Embedding · EMNLP 2018
Knowledge graphs › knowledge graph embedding
temporal knowledge graph embedding
0.312018
HyTE: Hyperplane-based Temporally aware Knowledge Graph Embedding · EMNLP 2018
Machine learning and data management
matrix completion
0.312025
A Geometric Approach to Personalized Recommendation with Set-Theoretic Constraints Using Box Embeddings · ICML 2025
Machine learning › Graph learning › graph neural network
graph convolutional network
0.112018
Dating Documents using Graph Convolution Networks · ACL (1) 2018

Methods — techniques the papers use, named apart from their topics

matrix factorization · 1.7jaccard index · 1.7box embedding · 1.7sampling · 0.8node embedding · 0.8gumbel distribution · 0.4gaussian convolution · 0.4neural document dating · 0.3knowledge graph embedding · 0.3hyperplane-based embedding · 0.3graph convolutional network · 0.3attention · 0.3
YearPublicationVenuePosition
2025 A Geometric Approach to Personalized Recommendation with Set-Theoretic Constraints Using Box Embeddings
abstract
Personalized item recommendation typically suffers from data sparsity, which is most often addressed by learning vector representations of users and items via low-rank matrix factorization. While this effectively densifies the matrix by assuming users and movies can be represented by linearly dependent latent features, it does not capture more complicated interactions. For example, vector representations struggle with set-theoretic relationships, such as negation and intersection, e.g. recommending a movie that is “comedy and action, but not romance”. In this work, we formulate the problem of personalized item recommendation as matrix completion where rows are set-theoretically dependent. To capture this set-theoretic dependence we represent each user and attribute by a hyperrectangle or box (i.e. a Cartesian product of intervals). Box embeddings can intuitively be understood as trainable Venn diagrams, and thus not only inherently represent similarity (via the Jaccard index), but also naturally and faithfully support arbitrary set-theoretic relationships. Queries involving set-theoretic constraints can be efficiently computed directly on the embedding space by performing geometric operations on the representations. We empirically demonstrate the superiority of box embeddings over vector-based neural methods on both simple and complex item recommendation queries by up to 30% overall.
Shib Sankar Dasgupta, Michael Boratko, Andrew McCallum
ICML1
2024 Learning Representations for Hierarchies with Minimal Support
abstract
When training node embedding models to represent large directed graphs (digraphs), it is impossible to observe all entries of the adjacency matrix during training. As a consequence most methods employ sampling. For very large digraphs, however, this means many (most) entries may be unobserved during training. In general, observing every entry would be necessary to uniquely identify a graph, however if we know the graph has a certain property some entries can be omitted - for example, only half the entries would be required for a symmetric graph. In this work, we develop a novel framework to identify a subset of entries required to uniquely distinguish a graph among all transitively-closed DAGs. We give an explicit algorithm to compute the provably minimal set of entries, and demonstrate empirically that one can train node embedding models with greater efficiency and performance, provided the energy function has an appropriate inductive bias. We achieve robust performance on synthetic hierarchies and a larger real-world taxonomy, observing improved convergence rates in a resource-constrained setting while reducing the set of training examples by as much as 99%.
Benjamin Rozonoyer, Michael Boratko, Dhruvesh Patel, Wenlong Zhao 0001, Shib Sankar Dasgupta, Andrew McCallum
NeurIPS5
2022 Word2Box: Capturing Set-Theoretic Semantics of Words using Box Embeddings
abstract
Shib Dasgupta, Michael Boratko, Siddhartha Mishra, Shriya Atmakuri, Dhruvesh Patel, Xiang Li, Andrew McCallum. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Shib Sankar Dasgupta, Michael Boratko, Siddhartha Mishra, Shriya Atmakuri, Dhruvesh Patel, Xiang Li 0069, Andrew McCallum
ACL (1)1
2021 Probabilistic Box Embeddings for Uncertain Knowledge Graph Reasoning
abstract
Xuelu Chen, Michael Boratko, Muhao Chen, Shib Sankar Dasgupta, Xiang Lorraine Li, Andrew McCallum. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Xuelu Chen, Michael Boratko, Muhao Chen 0001, Shib Sankar Dasgupta, Xiang Li 0069, Andrew McCallum
NAACL-HLT4
2021 Min/max stability and box distributions
abstract
In representation learning, capturing correlations between the represented elements is paramount. A recent line of work introduces the notion of learning region-based representations, with the objective of being able to better capture these correlations as set interactions. Box models use regions which are products of intervals on $[0,1]$ (i.e., "boxes"), representing joint probability distributions via Lebesgue measure. To mitigate issues with training, a recent work models the endpoints of these intervals using Gumbel distributions, chosen due to their min/max-stability. In this work we analyze min/max-stability on a bounded domain and provide a specific family of such distributions which, replacing Gumbel, allow for stochastic boxes embedded in a finite measure space. This allows for a latent noise model which is a probability measure. Furthermore, we demonstrate an equivalence between this region-based representation and a density representation, where intersection is given by products of densities. We compare our model to previous region-based probability models, and demonstrate it is capable of being trained effectively to modeling correlations.
Michael Boratko, Javier Burroni, Shib Sankar Dasgupta, Andrew McCallum
UAI3
2020 Improving Local Identifiability in Probabilistic Box Embeddings
abstract
Geometric embeddings have recently received attention for their natural ability to represent transitive asymmetric relations via containment. Box embeddings, where objects are represented by n-dimensional hyperrectangles, are a particularly promising example of such an embedding as they are closed under intersection and their volume can be calculated easily, allowing them to naturally represent calibrated probability distributions. The benefits of geometric embeddings also introduce a problem of local identifiability, however, where whole neighborhoods of parameters result in equivalent loss which impedes learning. Prior work addressed some of these issues by using an approximation to Gaussian convolution over the box parameters, however this intersection operation also increases the sparsity of the gradient. In this work we model the box parameters with min and max Gumbel distributions, which were chosen such that the space is still closed under the operation of intersection. The calculation of the expected intersection volume involves all parameters, and we demonstrate experimentally that this drastically improves the ability of such models to learn.
Shib Sankar Dasgupta, Michael Boratko, Luke Vilnis, Xiang Li 0069, Andrew McCallum
NeurIPS1
2018 Dating Documents using Graph Convolution Networks
abstract
Document date is essential for many important tasks, such as document retrieval, summarization, event detection, etc.While existing approaches for these tasks assume accurate knowledge of the document date, this is not always available, especially for arbitrary documents from the Web.Document Dating is a challenging problem which requires inference over the temporal structure of the document.Prior document dating systems have largely relied on handcrafted features while ignoring such documentinternal structures.In this paper, we propose NeuralDater, a Graph Convolutional Network (GCN) based document dating approach which jointly exploits syntactic and temporal graph structures of document in a principled way.To the best of our knowledge, this is the first application of deep learning for the problem of document dating.Through extensive experiments on real-world datasets, we find that NeuralDater significantly outperforms state-of-the-art baseline by 19% absolute (45% relative) accuracy points.
Shikhar Vashishth, Shib Sankar Dasgupta, Swayambhu Nath Ray, Partha P. Talukdar
ACL (1)2
2018 HyTE: Hyperplane-based Temporally aware Knowledge Graph Embedding
abstract
Knowledge Graph (KG) embedding has emerged as an active area of research resulting in the development of several KG embedding methods.Relational facts in KG often show temporal dynamics, e.g., the fact (Cristiano Ronaldo, playsFor, Manchester United) is valid only from 2003 to 2009.Most of the existing KG embedding methods ignore this temporal dimension while learning embeddings of the KG elements.In this paper, we propose HyTE, a temporally aware KG embedding method which explicitly incorporates time in the entity-relation space by associating each timestamp with a corresponding hyperplane.HyTE not only performs KG inference using temporal guidance, but also predicts temporal scopes for relational facts with missing time annotations.Through extensive experimentation on temporal datasets extracted from real-world KGs, we demonstrate the effectiveness of our model over both traditional as well as temporal KG embedding methods.
Shib Sankar Dasgupta, Swayambhu Nath Ray, Partha P. Talukdar
EMNLP1
2018 AD3: Attentive Deep Document Dater
abstract
Knowledge of the creation date of documents facilitates several tasks such as summarization, event extraction, temporally focused information extraction etc.Unfortunately, for most of the documents on the Web, the time-stamp metadata is either missing or can't be trusted.Thus, predicting creation time from document content itself is an important task.In this paper, we propose Attentive Deep Document Dater (AD3), an attention-based neural document dating system which utilizes both context and temporal information in documents in a flexible and principled manner.We perform extensive experimentation on multiple real-world datasets to demonstrate the effectiveness of AD3 over neural and non-neural baselines.
Swayambhu Nath Ray, Shib Sankar Dasgupta, Partha P. Talukdar
EMNLP2