VLDB 2026 Research / reviewers in the wild / expert
Michael Tang
dblp:69/475
· DBLP profile ↗
8ranked-venue papers
1as first author
3since 2021 · last 2025
0000-0003-4284-7712ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorSecurity and privacy · 1Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Reinforcement learning · 87% Robot manipulation · 9% Learning paradigms · 4% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% | |
| Network and information security
2 papers |
Systems and software security · 54% Privacy and data protection · 36% Usable security · 11% | |
| Software engineering, system software, and programming languages
1 paper |
Program verification · 100% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Medical and health informatics · 86% Bioinformatics and computational biology · 14% |
Topics — the 14 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning › exploration
directed exploration |
0.9 | 1 | 2025 | A Single Goal is All You Need: Skills and Exploration Emerge from Contrastive RL without Rewards, Demonstrations, or Subgoals · ICLR 2025 |
Machine learning › Reinforcement learning
exploration |
0.9 | 1 | 2025 | A Single Goal is All You Need: Skills and Exploration Emerge from Contrastive RL without Rewards, Demonstrations, or Subgoals · ICLR 2025 |
Machine learning › Reinforcement learning › hierarchical reinforcement learning › skill learning
skill discovery |
0.9 | 1 | 2025 | A Single Goal is All You Need: Skills and Exploration Emerge from Contrastive RL without Rewards, Demonstrations, or Subgoals · ICLR 2025 |
Information retrieval › document retrieval
reasoning-intensive retrieval |
0.9 | 1 | 2025 | BRIGHT: A Realistic and Challenging Benchmark for Reasoning-Intensive Retrieval · ICLR 2025 |
Information retrieval
retrieval models |
0.9 | 1 | 2025 | BRIGHT: A Realistic and Challenging Benchmark for Reasoning-Intensive Retrieval · ICLR 2025 |
Systems and software security
vulnerability discovery |
0.6 | 1 | 2022 | Hardening attack surfaces with formally proven binary format parsers · PLDI 2022 |
Program verification › code-level verification
verified parsing |
0.6 | 1 | 2022 | Hardening attack surfaces with formally proven binary format parsers · PLDI 2022 |
Medical and health informatics
computational pathology |
0.4 | 1 | 2019 | Atlas of Digital Pathology: A Generalized Hierarchical Histological Tissue Type-Annotated Database for Deep Learning · CVPR 2019 |
Privacy and data protection
social network privacy |
0.4 | 1 | 2019 | Moving Beyond Set-It-And-Forget-It Privacy Settings on Social Media · CCS 2019 |
Information retrieval
evaluation |
0.3 | 1 | 2025 | BRIGHT: A Realistic and Challenging Benchmark for Reasoning-Intensive Retrieval · ICLR 2025 |
Information retrieval › evaluation › test collection
retrieval benchmark |
0.3 | 1 | 2025 | BRIGHT: A Realistic and Challenging Benchmark for Reasoning-Intensive Retrieval · ICLR 2025 |
Machine learning › Learning paradigms
multi-label classification |
0.1 | 1 | 2019 | Atlas of Digital Pathology: A Generalized Hierarchical Histological Tissue Type-Annotated Database for Deep Learning · CVPR 2019 |
Bioinformatics and computational biology › structural bioinformatics › nucleic acid structure analysis
RNA structural bioinformatics |
0.0 | 1 | 2004 | RAG: RNA-As-Graphs database-concepts, analysis, features · Bioinform. 2004 |
Bioinformatics and computational biology › molecular informatics › cheminformatics
molecular graph representation |
0.0 | 1 | 2004 | RAG: RNA-As-Graphs database-concepts, analysis, features · Bioinform. 2004 |
Methods — techniques the papers use, named apart from their topics
formal proof · 1.1contrastive reinforcement learning · 0.9supervised learning · 0.8convolutional neural network · 0.8user study · 0.4friend interaction features · 0.4classifier · 0.4topological complexity ranking · 0.0graph enumeration · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Single Goal is All You Need: Skills and Exploration Emerge from Contrastive RL without Rewards, Demonstrations, or SubgoalsabstractIn this paper, we present empirical evidence of skills and directed exploration emerging from a simple RL algorithm long before any successful trials are observed. For example, in a manipulation task, the agent is given a single observation of the goal state (see Fig. 1) and learns skills, first for moving its end-effector, then for pushing the block, and finally for picking up and placing the block. These skills emerge before the agent has ever successfully placed the block at the goal location and without the aid of any reward functions, demonstrations, or manually-specified distance metrics. Once the agent has learned to reach the goal state reliably, exploration is reduced. Implementing our method involves a simple modification of prior work and does not require density estimates, ensembles, or any additional hyperparameters. Intuitively, the proposed method seems like it should be terrible at exploration, and we lack a clear theoretical understanding of why it works so effectively, though our experiments provide some hints. Grace Liu, Michael Tang, Benjamin Eysenbach |
ICLR | 2 |
| 2025 | BRIGHT: A Realistic and Challenging Benchmark for Reasoning-Intensive RetrievalabstractExisting retrieval benchmarks primarily consist of information-seeking queries (e.g., aggregated questions from search engines) where keyword or semantic-based retrieval is usually sufficient. However, many complex real-world queries require in-depth reasoning to identify relevant documents that go beyond surface form matching. For example, finding documentation for a coding question requires understanding the logic and syntax of the functions involved. To better benchmark retrieval on such challenging queries, we introduce BRIGHT, the first text retrieval benchmark that requires intensive reasoning to retrieve relevant documents. Our dataset consists of 1,398 real-world queries spanning diverse domains such as economics, psychology, mathematics, coding, and more. These queries are drawn from naturally occurring or carefully curated human data. Extensive evaluation reveals that even state-of-the-art retrieval models perform poorly on BRIGHT. The leading model on the MTEB leaderboard (Muennighoff et al., 2023), which achieves a score of 59.0 nDCG@10,1 produces a score of nDCG@10 of 18.0 on BRIGHT. We show that incorporating explicit reasoning about the query improves retrieval performance by up to 12.2 points. Moreover, incorporating retrieved documents from the top-performing retriever boosts question answering performance by over 6.6 points. We believe that BRIGHT paves the way for future research on retrieval systems in more realistic and challenging settings. Hongjin Su, Howard Yen, Mengzhou Xia, Niklas Muennighoff, Han-yu Wang, Haisu Liu, Zachary S. Siegel, Michael Tang, Ruoxi Sun 0002, Jinsung Yoon, Sercan Ö. Arik, Danqi Chen 0001, Tao Yu 0009 |
ICLR | 10 |
| 2022 | Hardening attack surfaces with formally proven binary format parsersabstractWith an eye toward performance, interoperability, or legacy concerns, low-level system software often must parse binary encoded data formats. Few tools are available for this task, especially since the formats involve a mixture of arithmetic and data dependence, beyond what can be handled by typical parser generators. As such, parsers are written by hand in languages like C, with inevitable errors leading to security vulnerabilities. Nikhil Swamy, Tahina Ramananandro, Aseem Rastogi, Irina Spiridonova, Haobin Ni, Dmitry Malloy, Juan Vazquez, Michael Tang, Omar Cardona, Arti Gupta |
PLDI | 8 |
| 2019 | Moving Beyond Set-It-And-Forget-It Privacy Settings on Social MediaabstractWhen users post on social media, they protect their privacy by choosing an access control setting that is rarely revisited. Changes in users' lives and relationships, as well as social media platforms themselves, can cause mismatches between a post's active privacy setting and the desired setting. The importance of managing this setting combined with the high volume of potential friend-post pairs needing evaluation necessitate a semi-automated approach. We attack this problem through a combination of a user study and the development of automated inference of potentially mismatched privacy settings. A total of 78 Facebook users reevaluated the privacy settings for five of their Facebook posts, also indicating whether a selection of friends should be able to access each post. They also explained their decision. With this user data, we designed a classifier to identify posts with currently incorrect sharing settings. This classifier shows a 317% improvement over a baseline classifier based on friend interaction. We also find that many of the most useful features can be collected without user intervention, and we identify directions for improving the classifier's accuracy. Mainack Mondal, Günce Su Yilmaz, Noah Hirsch, Mohammad Taha Khan, Michael Tang, Christopher Tran 0001, Chris Kanich, Blase Ur, Elena Zheleva |
CCS | 5 |
| 2019 | Atlas of Digital Pathology: A Generalized Hierarchical Histological Tissue Type-Annotated Database for Deep LearningabstractIn recent years, computer vision techniques have made large advances in image recognition and been applied to aid radiological diagnosis. Computational pathology aims to develop similar tools for aiding pathologists in diagnosing digitized histopathological slides, which would improve diagnostic accuracy and productivity amidst increasing workloads. However, there is a lack of publicly-available databases of (1) localized patch-level images annotated with (2) a large range of Histological Tissue Type (HTT). As a result, computational pathology research is constrained to diagnosing specific diseases or classifying tissues from specific organs, and cannot be readily generalized to handle unexpected diseases and organs. In this paper, we propose a new digital pathology database, the ``Atlas of Digital Pathology'' (or ADP), which comprises of 17,668 patch images extracted from 100 slides annotated with up to 57 hierarchical HTTs. Our data is generalized to different tissue types across different organs and aims to provide training data for supervised multi-label learning of patch-level HTT in a digitized whole slide image. We demonstrate the quality of our image labels through pathologist consultation and by training three state-of-the-art neural networks on tissue type classification. Quantitative results support the visually consistency of our data and we demonstrate a tissue type-based visual attention aid as a sample tool that could be developed from our database. Mahdi S. Hosseini, Lyndon Chan, Gabriel Tse, Michael Tang, Sajad Norouzi, Corwyn Rowsell, Konstantinos N. Plataniotis, Savvas Damaskinos |
CVPR | 4 |
| 2004 | RAG: RNA-As-Graphs database-concepts, analysis, featuresabstractMOTIVATION: Understanding RNA's structural diversity is vital for identifying novel RNA structures and pursuing RNA genomics initiatives. By classifying RNA secondary motifs based on correlations between conserved RNA secondary structures and functional properties, we offer an avenue for predicting novel motifs. Although several RNA databases exist, no comprehensive schemes are available for cataloguing the range and diversity of RNA's structural repertoire. RESULTS: Our RNA-As-Graphs (RAG) database describes and ranks all mathematically possible (including existing and candidate) RNA secondary motifs on the basis of graphical enumeration techniques. We represent RNA secondary structures as two-dimensional graphs (networks), specifying the connectivity between RNA secondary structural elements, such as loops, bulges, stems and junctions. We archive RNA tree motifs as 'tree graphs' and other RNAs, including pseudoknots, as general 'dual graphs'. All RNA motifs are catalogued by graph vertex number (a measure of sequence length) and ranked by topological complexity. The RAG inventory immediately suggests candidates for novel RNA motifs, either naturally occurring or synthetic, and thereby might stimulate the prediction and design of novel RNA motifs. AVAILABILITY: The database is accessible on the web at http://monod.biomath.nyu.edu/rna Hin Hark Gan, Daniela Fera, Julie Zorn, Nahum Shiffeldrim, Michael Tang, Uri Laserson, Namhee Kim, Tamar Schlick |
Bioinform. | 5 |
| 2003 | Mining Emerging SubstringsabstractWe introduce a new type of KDD patterns called emerging substrings. In a sequence database, an emerging substring (ES) of a data class is a substring which occurs more frequently in that class rather than in other classes. ESs are important to sequence classification as they capture significant contrasts between data classes and provide insights for the construction of sequence classifiers. We propose a suffix tree-based framework for mining ESs, and study the effectiveness of applying one or more pruning techniques in different stages of our ES mining algorithm. Experimental results show that if the target class is of a small population with respect to the whole database, which is the normal scenario in single-class ES mining, most of the pruning techniques would achieve considerable performance gain. Sarah Chan, Ben Kao, Chi Lap Yip, Michael Tang |
DASFAA | 4 |
| 2000 | Selection of Melody Lines for Music DatabasesabstractOne major approach to music retrieval is to model music as a sequence of features, after which traditional information retrieval techniques are applied on the sequence. Because of the temporal nature of music and the inexactness of user queries, most effort on music retrieval systems focus on issues such as indexing and approximation match. In contrast, the processing of music before feature extraction, such as the identification of a melody track, were often considered easy or done. This may be the case in a controlled environment, such as one for musicology research, where the pieces are carefully analyzed by human beings before being submitted to the database. However, in an environment where large volumes of music is obtained from the Web, manual music analysis is impractical. Since many well-known musical features often pertain to the melody of musical pieces, and users often remember the melody of a song, algorithms that select the melody tracks of a piece are important for Web-based content-based retrieval systems. We describe a number of algorithms for automatic melody track selection in a music retrieval context. We also study the performance of the algorithms by comparing their answers to those judged by human beings. Michael Tang, Chi Lap Yip, Ben Kao |
COMPSAC | 1 |