VLDB 2026 Research / reviewers in the wild / expert
Md. Shamim Hussain
dblp:232/1798
· DBLP profile ↗
3ranked-venue papers
3as first author
3since 2021 · last 2024
0000-0002-0832-913XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Graph learning · 53% Deep learning architectures and training · 28% Efficient and distributed learning · 19% |
Topics — the 8 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Graph learning › graph neural network
graph transformer |
1.3 | 2 | 2024 | Triplet Interaction Improves Graph Transformers: Accurate Molecular Graph Learning with Triplet Graph Transformers · ICML 2024 Global Self-Attention as a Replacement for Graph Convolution · KDD 2022 |
Machine learning › Deep learning architectures and training
transformer |
1.2 | 2 | 2023 | The Information Pathways Hypothesis: Transformers are Dynamic Self-Ensembles · KDD 2023 Global Self-Attention as a Replacement for Graph Convolution · KDD 2022 |
Machine learning › Graph learning
graph neural network |
0.8 | 2 | 2024 | Global Self-Attention as a Replacement for Graph Convolution · KDD 2022 Triplet Interaction Improves Graph Transformers: Accurate Molecular Graph Learning with Triplet Graph Transformers · ICML 2024 |
Machine learning › Graph learning
high-order interaction |
0.8 | 1 | 2024 | Triplet Interaction Improves Graph Transformers: Accurate Molecular Graph Learning with Triplet Graph Transformers · ICML 2024 |
Machine learning › Graph learning › molecular representation learning › molecular graph learning
molecular property prediction |
0.8 | 1 | 2024 | Triplet Interaction Improves Graph Transformers: Accurate Molecular Graph Learning with Triplet Graph Transformers · ICML 2024 |
Machine learning › Efficient and distributed learning › model compression › pruning › structured pruning
attention head pruning |
0.7 | 1 | 2023 | The Information Pathways Hypothesis: Transformers are Dynamic Self-Ensembles · KDD 2023 |
Machine learning › Efficient and distributed learning
model compression |
0.7 | 1 | 2023 | The Information Pathways Hypothesis: Transformers are Dynamic Self-Ensembles · KDD 2023 |
Machine learning › Deep learning architectures and training › attention mechanism
sparse attention |
0.7 | 1 | 2023 | The Information Pathways Hypothesis: Transformers are Dynamic Self-Ensembles · KDD 2023 |
Methods — techniques the papers use, named apart from their topics
triplet attention · 0.8transfer learning · 0.8stochastic subsampling · 0.7self-attention · 0.7ensemble of sub-models · 0.7message passing · 0.6global self-attention · 0.6edge channels · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Triplet Interaction Improves Graph Transformers: Accurate Molecular Graph Learning with Triplet Graph TransformersabstractGraph transformers typically lack third-order interactions, limiting their geometric understanding which is crucial for tasks like molecular geometry prediction. We propose the Triplet Graph Transformer (TGT) that enables direct communication between pairs within a 3-tuple of nodes via novel triplet attention and aggregation mechanisms. TGT is applied to molecular property prediction by first predicting interatomic distances from 2D graphs and then using these distances for downstream tasks. A novel three-stage training procedure and stochastic inference further improve training efficiency and model performance. Our model achieves new state-of-the-art (SOTA) results on open challenge benchmarks PCQM4Mv2 and OC20 IS2RE. We also obtain SOTA results on QM9, MOLPCBA, and LIT-PCBA molecular property prediction benchmarks via transfer learning. We also demonstrate the generality of TGT with SOTA results on the traveling salesman problem (TSP). Md. Shamim Hussain, Mohammed J. Zaki, Dharmashankar Subramanian |
ICML | 1 |
| 2023 | The Information Pathways Hypothesis: Transformers are Dynamic Self-EnsemblesabstractTransformers use the dense self-attention mechanism which gives a lot of flexibility for long-range connectivity. Over multiple layers of a deep transformer, the number of possible connectivity patterns increases exponentially. However, very few of these contribute to the performance of the network, and even fewer are essential. We hypothesize that there are sparsely connected sub-networks within a transformer, called information pathways which can be trained independently. However, the dynamic (i.e., input-dependent) nature of these pathways makes it difficult to prune dense self-attention during training. But the overall distribution of these pathways is often predictable. We take advantage of this fact to propose Stochastically Subsampled self-Attention (SSA) - a general-purpose training strategy for transformers that can reduce both the memory and computational cost of self-attention by 4 to 8 times during training while also serving as a regularization method - improving generalization over dense training. We show that an ensemble of sub-models can be formed from the subsampled pathways within a network, which can achieve better performance than its densely attended counterpart. We perform experiments on a variety of NLP, computer vision and graph learning tasks in both generative and discriminative settings to provide empirical evidence for our claims and show the effectiveness of the proposed method. Md. Shamim Hussain, Mohammed J. Zaki, Dharmashankar Subramanian |
KDD | 1 |
| 2022 | Global Self-Attention as a Replacement for Graph ConvolutionabstractWe propose an extension to the transformer neural network architecture for general-purpose graph learning by adding a dedicated pathway for pairwise structural information, called edge channels. The resultant framework - which we call Edge-augmented Graph Transformer (EGT) - can directly accept, process and output structural information of arbitrary form, which is important for effective learning on graph-structured data. Our model exclusively uses global self-attention as an aggregation mechanism rather than static localized convolutional aggregation. This allows for unconstrained long-range dynamic interactions between nodes. Moreover, the edge channels allow the structural information to evolve from layer to layer, and prediction tasks on edges/links can be performed directly from the output embeddings of these channels. We verify the performance of EGT in a wide range of graph-learning experiments on benchmark datasets, in which it outperforms Convolutional/Message-Passing Graph Neural Networks. EGT sets a new state-of-the-art for the quantum-chemical regression task on the OGB-LSC PCQM4Mv2 dataset containing 3.8 million molecular graphs. Our findings indicate that global self-attention based aggregation can serve as a flexible, adaptive and effective replacement of graph convolution for general-purpose graph learning. Therefore, convolutional local neighborhood aggregation is not an essential inductive bias. Md. Shamim Hussain, Mohammed J. Zaki, Dharmashankar Subramanian |
KDD | 1 |