VLDB 2026 Research / reviewers in the wild / expert
Philip Pham
dblp:263/3069
· DBLP profile ↗
8ranked-venue papers
0as first author
4since 2021 · last 2021
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Deep learning architectures and training · 66% Graph learning · 13% Trustworthy machine learning · 12% | |
| Theoretical computer science
1 paper |
Algorithms and data structures · 100% |
Topics — the 12 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Deep learning architectures and training
transformer |
1.4 | 3 | 2021 | OmniNet: Omnidirectional Representations from Transformers · ICML 2021 Long Range Arena : A Benchmark for Efficient Transformers · ICLR 2021 Big Bird: Transformers for Longer Sequences · NeurIPS 2020 |
Machine learning › Deep learning architectures and training
attention mechanism |
0.5 | 1 | 2021 | OmniNet: Omnidirectional Representations from Transformers · ICML 2021 |
Machine learning › Deep learning architectures and training › attention mechanism › efficient attention
efficient self-attention |
0.5 | 1 | 2021 | OmniNet: Omnidirectional Representations from Transformers · ICML 2021 |
Machine learning › Deep learning architectures and training › transformer
efficient transformer |
0.5 | 1 | 2021 | Long Range Arena : A Benchmark for Efficient Transformers · ICLR 2021 |
Machine learning › Graph learning
graph regularization |
0.5 | 1 | 2021 | Neural Structured Learning: Training Neural Networks with Structured Signals · WSDM 2021 |
Machine learning › Trustworthy machine learning › fairness › fair unsupervised learning
fair clustering |
0.4 | 1 | 2020 | Fair Hierarchical Clustering · NeurIPS 2020 |
Machine learning › Trustworthy machine learning
fairness |
0.4 | 1 | 2020 | Fair Hierarchical Clustering · NeurIPS 2020 |
Natural language and speech › Language models and text generation › language modeling
long-context language modeling |
0.4 | 1 | 2020 | Big Bird: Transformers for Longer Sequences · NeurIPS 2020 |
Machine learning › Deep learning architectures and training › attention mechanism
sparse attention |
0.4 | 1 | 2020 | Big Bird: Transformers for Longer Sequences · NeurIPS 2020 |
Algorithms and data structures › clustering
hierarchical clustering |
0.4 | 1 | 2020 | Fair Hierarchical Clustering · NeurIPS 2020 |
Performance modeling and evaluation
benchmarking |
0.1 | 1 | 2021 | Long Range Arena : A Benchmark for Efficient Transformers · ICLR 2021 |
Machine learning › Representation and self-supervised learning › representation learning
embedding learning |
0.1 | 1 | 2020 | Neural Structured Learning: Training Neural Networks with Structured Signals · KDD 2020 |
Methods — techniques the papers use, named apart from their topics
long-range sequence modeling · 1.0adversarial perturbation · 0.9meta-learning · 0.5low-rank attention · 0.5kernel-based attention · 0.5embedding learning · 0.5big bird attention · 0.5sparse attention · 0.4relative position encoding · 0.4graph regularization · 0.4approximation algorithm · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Long Range Arena : A Benchmark for Efficient Transformers
Yi Tay, Mostafa Dehghani 0001, Samira Abnar, Yikang Shen, Dara Bahri, Philip Pham, Jinfeng Rao, Sebastian Ruder, Donald Metzler |
ICLR | 6 |
| 2021 | OmniNet: Omnidirectional Representations from TransformersabstractThis paper proposes Omnidirectional Representations from Transformers (OMNINET). In OmniNet, instead of maintaining a strictly horizon-tal receptive field, each token is allowed to attend to all tokens in the entire network. This process can also be interpreted as a form of extreme or intensive attention mechanism that has the receptive field of the entire width and depth of the network. To this end, the omnidirectional attention is learned via a meta-learner, which is essentially another self-attention based model. In order to mitigate the computationally expensive costs of full receptive field attention, we leverage efficient self-attention models such as kernel-based, low-rank attention and/or Big Bird as the meta-learner. Extensive experiments are conducted on autoregressive language modeling(LM1B, C4), Machine Translation, Long Range Arena (LRA), and Image Recognition.The experiments show that OmniNet achieves considerable improvements across these tasks, including achieving state-of-the-art performance on LM1B,WMT’14 En-De/En-Fr, and Long Range Arena.Moreover, using omnidirectional representation in Vision Transformers leads to significant improvements on image recognition tasks on both few-shot learning and fine-tuning setups. Yi Tay, Mostafa Dehghani 0001, Vamsi Aribandi, Jai Gupta 0001, Philip Pham, Zhen Qin 0001, Dara Bahri, Da-Cheng Juan, Donald Metzler |
ICML | 5 |
| 2021 | ReadTwice: Reading Very Large Documents with MemoriesabstractYury Zemlyanskiy, Joshua Ainslie, Michiel de Jong, Philip Pham, Ilya Eckstein, Fei Sha. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Yury Zemlyanskiy, Joshua Ainslie, Michiel de Jong, Philip Pham, Ilya Eckstein, Fei Sha |
NAACL-HLT | 4 |
| 2021 | Neural Structured Learning: Training Neural Networks with Structured SignalsabstractWe present Neural Structured Learning (NSL) in TensorFlow [1], a new learning paradigm to train neural networks by leveraging structured signals in addition to feature inputs. Structure can be explicit as represented by a graph, or implicit, either induced by adversarial perturbation or inferred using techniques like embedding learning. NSL is open-sourced as part of the TensorFlow [2] ecosystem and is widely used in Google across many products and services. In this tutorial, we provide an overview of the NSL framework including various libraries, tools, and APIs as well as demonstrate the practical use of NSL in different applications. The NSL website is hosted at www.tensorflow.org/neural_structured_learning, which includes details about the theoretical foundations of the technology, extensive API documentation, and hands-on tutorials. Arjun Gopalan, Da-Cheng Juan, Cesar Ilharco Magalhaes, Chun-Sung Ferng, Allan Heydon, Chun-Ta Lu, Philip Pham, George Yu, Yicheng Fan |
WSDM | 7 |
| 2020 | ETC: Encoding Long and Structured Inputs in TransformersabstractJoshua Ainslie, Santiago Ontanon, Chris Alberti, Vaclav Cvicek, Zachary Fisher, Philip Pham, Anirudh Ravula, Sumit Sanghai, Qifan Wang, Li Yang. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Joshua Ainslie, Santiago Ontañón, Christopher Alberti, Vaclav Cvicek, Zachary Fisher, Philip Pham, Anirudh Ravula, Sumit Sanghai, Qifan Wang 0001 |
EMNLP (1) | 6 |
| 2020 | Neural Structured Learning: Training Neural Networks with Structured SignalsabstractWe present Neural Structured Learning (NSL) in TensorFlow [2], a new learning paradigm to train neural networks by leveraging structured signals in addition to feature inputs. Structure can be explicit as represented by a graph, or implicit, either induced by adversarial perturbation or inferred using techniques like embedding learning. NSL is open-sourced as part of the TensorFlow [3] ecosystem and is widely used in Google across many products and services. In this tutorial, we provide an overview of the NSL framework including various libraries, tools, and APIs as well as demonstrate the practical use of NSL in different applications. The NSL website is hosted at www.tensorflow.org/neural_structured_learning, which includes details about the theoretical foundations of the technology, extensive API documentation, and hands-on tutorials. Arjun Gopalan, Da-Cheng Juan, Cesar Ilharco Magalhaes, Chun-Sung Ferng, Allan Heydon, Chun-Ta Lu, Philip Pham, George Yu |
KDD | 7 |
| 2020 | Fair Hierarchical ClusteringabstractAs machine learning has become more prevalent, researchers have begun to recognize the necessity of ensuring machine learning systems are fair. Recently, there has been an interest in defining a notion of fairness that mitigates over-representation in traditional clustering. In this paper we extend this notion to hierarchical clustering, where the goal is to recursively partition the data to optimize a specific objective. For various natural objectives, we obtain simple, efficient algorithms to find a provably good fair hierarchical clustering. Empirically, we show that our algorithms can find a fair hierarchical clustering, with only a negligible loss in the objective. Sara Ahmadian, Alessandro Epasto, Marina Knittel, Ravi Kumar 0001, Mohammad Mahdian, Benjamin Moseley, Philip Pham, Sergei Vassilvitskii |
NeurIPS | 7 |
| 2020 | Big Bird: Transformers for Longer SequencesabstractTransformers-based models, such as BERT, have been one of the most successful deep learning models for NLP. Unfortunately, one of their core limitations is the quadratic dependency (mainly in terms of memory) on the sequence length due to their full attention mechanism. To remedy this, we propose, BigBird, a sparse attention mechanism that reduces this quadratic dependency to linear. We show that BigBird is a universal approximator of sequence functions and is Turing complete, thereby preserving these properties of the quadratic, full attention model. Along the way, our theoretical analysis reveals some of the benefits of having $O(1)$ global tokens (such as CLS), that attend to the entire sequence as part of the sparse attention mechanism. The proposed sparse attention can handle sequences of length up to 8x of what was previously possible using similar hardware. As a consequence of the capability to handle longer context, BigBird drastically improves performance on various NLP tasks such as question answering and summarization. We also propose novel applications to genomics data. Manzil Zaheer, Guru Guruganesh, Avinava Dubey, Joshua Ainslie, Christopher Alberti, Santiago Ontañón, Philip Pham, Anirudh Ravula, Qifan Wang 0001, Amr Ahmed 0001 |
NeurIPS | 7 |