Soham De

dblp:124/9197 · DBLP profile ↗
← Back
23ranked-venue papers
7as first author
11since 2021 · last 2026
0000-0001-7907-1335ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 5 first-author · 6 since 2021Databases, data management, data science and information retrieval · 5 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author
YearPublicationVenuePosition
2026 Question the Questions: Auditing Representation in Online Deliberative Processes
abstract
A central feature of many deliberative processes, such as citizens' assemblies and deliberative polls, is the opportunity for participants to engage directly with experts. While participants are typically invited to propose questions for expert panels, only a limited number can be selected due to time constraints. This raises the challenge of how to choose a small set of questions that best represent the interests of all participants. We introduce an auditing framework for measuring the level of representation provided by a slate of questions, based on the social choice concept known as justified representation (JR). We present the first algorithms for auditing JR in the general utility setting, with our most efficient algorithm achieving a runtime of $O(mn\log n)$, where $n$ is the number of participants and $m$ is the number of proposed questions. We apply our auditing methods to historical deliberations, comparing the representativeness of (a) the actual questions posed to the expert panel (chosen by a moderator), (b) participants' questions chosen via integer linear programming, (c) summary questions generated by large language models (LLMs). Our results highlight both the promise and current limitations of LLMs in supporting deliberative processes. By integrating our methods into an online deliberation platform that has been used for over hundreds of deliberations across more than 50 countries, we make it easy for practitioners to audit and improve representation in future deliberations.
Soham De, Lodewijk Gelauff, Ashish Goel, Smitha Milli, Ariel D. Procaccia, Alice Siu
WWW1
2025 Supernotes: Driving Consensus in Crowd-Sourced Fact-Checking
abstract
X's Community Notes, a crowd-sourced fact-checking system, allows users to annotate potentially misleading posts.Notes rated as helpful by a diverse set of users are prominently displayed below the original post.While demonstrably effective at reducing misinformation's impact when notes are displayed, there is an opportunity for notes to appear on many more posts: for 91% of posts where at least one note is proposed, no notes ultimately achieve sufficient support from diverse users to be shown on the platform.This motivates the development of Supernotes: AI-generated notes that synthesize information from several existing community notes and are written to foster consensus among a diverse set of users.Our framework uses an LLM to generate many diverse Supernote candidates from existing proposed notes.These candidates are then evaluated by a novel scoring model, trained on millions of historical Community Notes ratings, selecting candidates that are most likely to be rated helpful by a diverse set of users.To test our framework, we ran a human subjects experiment in which we asked participants to compare the Supernotes generated by our framework to the best existing community notes for 100 sample posts.We found that participants rated the Supernotes as significantly more helpful, and when asked to choose between the two, preferred the Supernotes 75.2% of the time.Participants also rated the Supernotes more favorably than the best existing notes on quality, clarity, coverage, context, and argumentativeness.Finally, in a follow-up experiment, we asked participants to compare the Supernotes against LLMgenerated summaries and found that the participants rated the Supernotes significantly more helpful, demonstrating that both the LLM-based candidate generation and the consensus-driven scoring play crucial roles in creating notes that effectively build consensus among diverse users.
Soham De, Michiel A. Bakker, Jay Baxter, Martin Saveski
WWW1
2024 Universality of Linear Recurrences Followed by Non-linear Projections: Finite-Width Guarantees and Benefits of Complex Eigenvalues
abstract
Deep neural networks based on linear RNNs interleaved with position-wise MLPs are gaining traction as competitive approaches for sequence modeling. Examples of such architectures include state-space models (SSMs) like S4, LRU, and Mamba: recently proposed models that achieve promising performance on text, genetics, and other data that require long-range reasoning. Despite experimental evidence highlighting these architectures' effectiveness and computational efficiency, their expressive power remains relatively unexplored, especially in connection to specific choices crucial in practice - e.g., carefully designed initialization distribution and potential use of complex numbers. In this paper, we show that combining MLPs with both real or complex linear diagonal recurrences leads to arbitrarily precise approximation of regular causal sequence-to-sequence maps. At the heart of our proof, we rely on a separation of concerns: the linear RNN provides a lossless encoding of the input sequence, and the MLP performs non-linear processing on this encoding. While we show that real diagonal linear recurrences are enough to achieve universality in this architecture, we prove that employing complex eigenvalues near unit disk - i.e., empirically the most successful strategy in S4 - greatly helps the RNN in storing information. We connect this finding with the vanishing gradient issue and provide experiments supporting our claims.
Antonio Orvieto, Soham De, Caglar Gulcehre, Razvan Pascanu, Samuel L. Smith
ICML2
2024 Networks and Influencers in Online Propaganda Events: A Comparative Study of Three Cases in India
abstract
The structure and mechanics of organized outreach around certain issues, such as in propaganda networks, is constantly evolving on social media. We collect tweets on two propaganda events and one non-propaganda event with varying degrees of organized messaging. We then perform a comparative analysis of the user and network characteristics of social media networks around these events and find clearly distinguishable traits across events. We find that influential entities like prominent politicians, digital influencers, and mainstream media prefer to engage more with social media events with lesser degree of propaganda while avoiding events with high degree of propaganda, which are mostly sustained by lesser known but dedicated micro-influencers. We also find that network communities of events with high degree of propaganda are significantly centralized with respect to the influence exercised by their leaders. The methods and findings of this study can pave the way for modeling and early detection of other propaganda events, using their user and community characteristics.
Anirban Sen, Soham De, Joyojeet Pal
Proc. ACM Hum. Comput. Interact.2
2023 Resurrecting Recurrent Neural Networks for Long Sequences
abstract
Recurrent Neural Networks (RNNs) offer fast inference on long sequences but are hard to optimize and slow to train. Deep state-space models (SSMs) have recently been shown to perform remarkably well on long sequence modeling tasks, and have the added benefits of fast parallelizable training and RNN-like fast inference. However, while SSMs are superficially similar to RNNs, there are important differences that make it unclear where their performance boost over RNNs comes from. We show that careful design of deep RNNs using standard signal propagation arguments can recover the impressive performance of deep SSMs on long-range reasoning tasks, while matching their training speed. To achieve this, we analyze and ablate a series of changes to standard RNNs including linearizing and diagonalizing the recurrence, using better parameterizations and initializations, and ensuring careful normalization of the forward pass. Our results provide new insights on the origins of the impressive performance of deep SSMs, and introduce an RNN block called the Linear Recurrent Unit (or LRU) that matches both their performance on the Long Range Arena benchmark and their computational efficiency.
Antonio Orvieto, Samuel L. Smith, Albert Gu, Anushan Fernando, Caglar Gulcehre, Razvan Pascanu, Soham De
ICML7
2022 Note: Picking Sides: The influencer-driven #HijabBan discourse on Twitter
abstract
Social Media is increasingly central to culture wars, and in recent years has been at the center of debates around the weaponization of peoples’ opinions in polarized situations. We examine one such issue, an attempt to allow schools and junior colleges to ban girls from wearing Hijab in India in the southern state of Karnataka, India, by studying Twitter messaging related to the issue in early 2022, when it was in the news. We find that the narrative supporting the ban of the Hijab on Twitter is primarily driven by a minority of highly-influential individuals who are predominantly male and polarised in favour of the ruling party, while the discourse against the ban relies largely on influencers from within the Muslim community. Our findings show that social media can be a useful tool to craft the contours of politically-motivated escalation in the Global South.
Soham De, Anmol Panda, Joyojeet Pal
COMPASS1
2022 Closed Ranks: The Discursive Value of Military Support for Indian Politicians on Social Media
abstract
Influencers play a crucial role in shaping public narratives through information creation and diffusion in the Global South. While public figures from various walks of life and their impact on public discourse have been studied, defence veterans as influencers of the political discourse have been largely overlooked. Veterans matter in the public spehere as a normatively important political lobby. They are also interesting because, unlike active-duty military officers, they are not restricted from taking public sides on politics, so their posts may provide a window into the views of those still in the service. In this work, we systematically analyze the engagement on Twitter of self-described defence-related accounts and politician accounts that post on defence-related issues. We find that self-described defence-related accounts disproportionately engage with the current ruling party in India. We find that politicians promote their closeness to the defence services and nationalist credentials through engagements with defence-related influencers. We briefly consider the institutional implications of these patterns and connections.
Agrima Seth, Soham De, Arshia Arya, Steven Wilkinson, Sushant Singh, Joyojeet Pal
ICTD2
2022 DISMISS: Database of Indian Social Media Influencers on Twitter
Arshia Arya, Soham De, Dibyendu Mishra, Gazal Shekhawat, Anmol Panda, Faisal M. Lalani, Parantak Singh, Ramaravind Kommiya Mothilal, Rynaa Grover, Sachita Nishal, Saloni Dash, Shehla Rashid Shora, Syeda Zainab Akbar, Joyojeet Pal
ICWSM2
2021 Characterizing signal propagation to close the performance gap in unnormalized ResNets
Andrew Brock, Soham De, Samuel L. Smith
ICLR2
2021 On the Origin of Implicit Regularization in Stochastic Gradient Descent
Samuel L. Smith, Benoit Dherin, David G. T. Barrett, Soham De
ICLR4
2021 High-Performance Large-Scale Image Recognition Without Normalization
abstract
Batch normalization is a key component of most image classification models, but it has many undesirable properties stemming from its dependence on the batch size and interactions between examples. Although recent work has succeeded in training deep ResNets without normalization layers, these models do not match the test accuracies of the best batch-normalized networks, and are often unstable for large learning rates or strong data augmentations. In this work, we develop an adaptive gradient clipping technique which overcomes these instabilities, and design a significantly improved class of Normalizer-Free ResNets. Our smaller models match the test accuracy of an EfficientNet-B7 on ImageNet while being up to 8.7x faster to train, and our largest models attain a new state-of-the-art top-1 accuracy of 86.5%. In addition, Normalizer-Free models attain significantly better performance than their batch-normalized counterparts when fine-tuning on ImageNet after large-scale pre-training on a dataset of 300 million labeled images, with our best models obtaining an accuracy of 89.2%.
Andrew Brock, Soham De, Samuel L. Smith, Karen Simonyan
ICML2
2020 The Impact of Neural Network Overparameterization on Gradient Confusion and Stochastic Gradient Descent
abstract
This paper studies how neural network architecture affects the speed of training. We introduce a simple concept called gradient confusion to help formally analyze this. When gradient confusion is high, stochastic gradients produced by different data samples may be negatively correlated, slowing down convergence. But when gradient confusion is low, data samples interact harmoniously, and training proceeds quickly. Through theoretical and experimental results, we demonstrate how the neural network architecture affects gradient confusion, and thus the efficiency of training. Our results show that, for popular initialization techniques, increasing the width of neural networks leads to lower gradient confusion, and thus faster model training. On the other hand, increasing the depth of neural networks has the opposite effect. Our results indicate that alternate initialization techniques or networks using both batch normalization and skip connections help reduce the training burden of very deep networks.
Karthik Abinav Sankararaman, Soham De, Zheng Xu 0002, W. Ronny Huang, Tom Goldstein
ICML2
2020 On the Generalization Benefit of Noise in Stochastic Gradient Descent
abstract
It has long been argued that minibatch stochastic gradient descent can generalize better than large batch gradient descent in deep neural networks. However recent papers have questioned this claim, arguing that this effect is simply a consequence of suboptimal hyperparameter tuning or insufficient compute budgets when the batch size is large. In this paper, we perform carefully designed experiments and rigorous hyperparameter sweeps on a range of popular models, which verify that small or moderately large batch sizes can substantially outperform very large batches on the test set. This occurs even when both models are trained for the same number of iterations and large batches achieve smaller training losses. Our results confirm that the noise in stochastic gradients can enhance generalization. We study how the optimal learning rate schedule changes as the epoch budget grows, and we provide a theoretical account of our observations based on the stochastic differential equation perspective of SGD dynamics.
Samuel L. Smith, Erich Elsen, Soham De
ICML3
2020 Batch Normalization Biases Residual Blocks Towards the Identity Function in Deep Networks
abstract
Batch normalization dramatically increases the largest trainable depth of residual networks, and this benefit has been crucial to the empirical success of deep residual networks on a wide range of benchmarks. We show that this key benefit arises because, at initialization, batch normalization downscales the residual branch relative to the skip connection, by a normalizing factor on the order of the square root of the network depth. This ensures that, early in training, the function computed by normalized residual blocks in deep networks is close to the identity function (on average). We use this insight to develop a simple initialization scheme that can train deep residual networks without normalization. We also provide a detailed empirical study of residual networks, which clarifies that, although batch normalized networks can be trained with larger learning rates, this effect is only beneficial in specific compute regimes, and has minimal benefits when the batch size is small.
Soham De, Samuel L. Smith
NeurIPS1
2020 Modeling Citation Trajectories of Scientific Papers
Dattatreya Mohapatra, Siddharth Pal, Soham De, Ponnurangam Kumaraguru, Tanmoy Chakraborty 0002
PAKDD (2)3
2019 Adversarial Robustness through Local Linearization
abstract
Adversarial training is an effective methodology for training deep neural networks that are robust against adversarial, norm-bounded perturbations. However, the computational cost of adversarial training grows prohibitively as the size of the model and number of input dimensions increase. Further, training against less expensive and therefore weaker adversaries produces models that are robust against weak attacks but break down under attacks that are stronger. This is often attributed to the phenomenon of gradient obfuscation; such models have a highly non-linear loss surface in the vicinity of training examples, making it hard for gradient-based attacks to succeed even though adversarial examples still exist. In this work, we introduce a novel regularizer that encourages the loss to behave linearly in the vicinity of the training data, thereby penalizing gradient obfuscation while encouraging robustness. We show via extensive experiments on CIFAR-10 and ImageNet, that models trained with our regularizer avoid gradient obfuscation and can be trained significantly faster than adversarial training. Using this regularizer, we exceed current state of the art and achieve 47% adversarial accuracy for ImageNet with L-infinity norm adversarial perturbations of radius 4/255 under an untargeted, strong, white-box attack. Additionally, we match state of the art results for CIFAR-10 at 8/255.
Chongli Qin, James Martens, Sven Gowal, Dilip Krishnan, Krishnamurthy Dvijotham, Alhussein Fawzi, Soham De, Robert Stanforth, Pushmeet Kohli
NeurIPS7
2019 Efficient Neural Network Verification with Exactness Characterization
Krishnamurthy Dvijotham, Robert Stanforth, Sven Gowal, Chongli Qin, Soham De, Pushmeet Kohli
UAI5
2017 Automated Inference with Adaptive Batches
abstract
Classical stochastic gradient methods for optimization rely on noisy gradient approximations that become progressively less accurate as iterates approach a solution. The large noise and small signal in the resulting gradients makes it difficult to use them for adaptive stepsize selection and automatic stopping. We propose alternative “big batch” SGD schemes that adaptively grow the batch size over time to maintain a nearly constant signal-to-noise ratio in the gradient approximation. The resulting methods have similar convergence rates to classical SGD, and do not require convexity of the objective. The high fidelity gradients enable automated learning rate selection and do not require stepsize decay. Big batch methods are thus easily automated and can run with little or no oversight.
Soham De, Abhay Kumar Yadav, David Jacobs 0001, Tom Goldstein
AISTATS1
2017 Son of Zorn's lemma: Targeted style transfer using instance-aware semantic segmentation
abstract
Style transfer is an important task in which the style of a source image is mapped onto that of a target image. The method is useful for synthesizing derivative works of a particular artist or specific painting. This work considers targeted style transfer, in which the style of a template image is used to alter only part of a target image. For example, an artist may wish to alter the style of only one particular object in a target image without altering the object's general morphology or surroundings. This is useful, for example, in augmented reality applications (such as the recently released Pokémon go), where one wants to alter the appearance of a single real-world object in an image frame to make it appear as a cartoon. Most notably, the rendering of real-world objects into cartoon characters has been used in a number of films and television show, such as the upcoming series Son of Zorn. We present a method for targeted style transfer that simultaneously segments and stylizes single objects selected by the user. The method uses a Markov random field model to smooth and anti-alias outlier pixels near object boundaries, so that stylized objects naturally blend into their surroundings.
Carlos Domingo Castillo, Soham De, Xintong Han, Abhay Kumar Yadav, Tom Goldstein
ICASSP2
2017 Training Quantized Nets: A Deeper Understanding
abstract
Currently, deep neural networks are deployed on low-power portable devices by first training a full-precision model using powerful hardware, and then deriving a corresponding low-precision model for efficient inference on such systems. However, training models directly with coarsely quantized weights is a key step towards learning on embedded platforms that have limited computing resources, memory capacity, and power consumption. Numerous recent publications have studied methods for training quantized networks, but these studies have mostly been empirical. In this work, we investigate training methods for quantized neural networks from a theoretical viewpoint. We first explore accuracy guarantees for training methods under convexity assumptions. We then look at the behavior of these algorithms for non-convex problems, and show that training algorithms that exploit high-precision representations have an important greedy search phase that purely quantized training methods lack, which explains the difficulty of training using low-precision arithmetic.
Hao Li 0022, Soham De, Zheng Xu 0002, Christoph Studer, Hanan Samet, Tom Goldstein
NIPS2
2016 Efficient Distributed SGD with Variance Reduction
abstract
Stochastic Gradient Descent (SGD) has become one of the most popular optimization methods for training machine learning models on massive datasets. However, SGD suffers from two main drawbacks: (i) The noisy gradient updates have high variance, which slows down convergence as the iterates approach the optimum, and (ii) SGD scales poorly in distributed settings, typically experiencing rapidly decreasing marginal benefits as the number of workers increases. In this paper, we propose a highly parallel method, CentralVR, that uses error corrections to reduce the variance of SGD gradient updates, and scales linearly with the number of worker nodes. CentralVR enjoys low iteration complexity, provably linear convergence rates, and exhibits linear performance gains up to hundreds of cores for massive datasets. We compare CentralVR to state-of-the-art parallel stochastic optimization methods on a variety of models and datasets, and find that our proposed methods exhibit stronger scaling than other SGD variants.
Soham De, Tom Goldstein
ICDM1
2015 Layer-Specific Adaptive Learning Rates for Deep Networks
abstract
The increasing complexity of deep learning architectures is resulting in training time requiring weeks or even months. This slow training is due in part to "vanishing gradients," in which the gradients used by back-propagation are extremely large for weights connecting deep layers (layers near the output layer), and extremely small for shallow layers (near the input layer), this results in slow learning in the shallow layers. Additionally, it has also been shown that in highly non-convex problems, such as deep neural networks, there is a proliferation of high-error low curvature saddle points, which slows down learning dramatically [1]. In this paper, we attempt to overcome the two above problems by proposing an optimization method for training deep neural networks which uses learning rates which are both specific to each layer in the network and adaptive to the curvature of the function, increasing the learning rate at low curvature points. This enables us to speed up learning in the shallow layers of the network and quickly escape high-error low curvature saddle points. We test our method on standard image classification datasets such as MNIST, CIFAR10 and ImageNet, and demonstrate that our method increases accuracy as well as reduces the required training time over standard algorithms.
Soham De, Yangmuzi Zhang, Tom Goldstein, Gavin Taylor
ICMLA2
2012 Plagiarism Detection in Polyphonic Music using Monaural Signal Separation
abstract
Given the large number of new musical tracks released each year, automated approaches to plagiarism detection are essential to help us track potential violations of copyright.Most current approaches to plagiarism detection are based on musical similarity measures, which typically ignore the issue of polyphony in music.We present a novel feature space for audio derived from compositional modelling techniques, commonly used in signal separation, that provides a mechanism to account for polyphony without incurring an inordinate amount of computational overhead.We employ this feature representation in conjunction with traditional audio feature representations in a classification framework which uses an ensemble of distance features to characterize pairs of songs as being plagiarized or not.Our experiments on a database of about 3000 musical track pairs show that the new feature space characterization produces significant improvements over standard baselines.
Soham De, Indradyumna Roy, Tarunima Prabhakar, Kriti Suneja, Sourish Chaudhuri, Rita Singh, Bhiksha Raj
INTERSPEECH1