Michael Tang

dblp:69/475 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
3since 2021 · last 2025
0000-0003-4284-7712ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorSecurity and privacy · 1Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Reinforcement learning · 87% Robot manipulation · 9% Learning paradigms · 4%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%
Network and information security
2 papers
Systems and software security · 54% Privacy and data protection · 36% Usable security · 11%
Software engineering, system software, and programming languages
1 paper
Program verification · 100%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Medical and health informatics · 86% Bioinformatics and computational biology · 14%

Topics — the 14 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning › exploration
directed exploration
0.912025
A Single Goal is All You Need: Skills and Exploration Emerge from Contrastive RL without Rewards, Demonstrations, or Subgoals · ICLR 2025
Machine learning › Reinforcement learning
exploration
0.912025
A Single Goal is All You Need: Skills and Exploration Emerge from Contrastive RL without Rewards, Demonstrations, or Subgoals · ICLR 2025
Machine learning › Reinforcement learning › hierarchical reinforcement learning › skill learning
skill discovery
0.912025
A Single Goal is All You Need: Skills and Exploration Emerge from Contrastive RL without Rewards, Demonstrations, or Subgoals · ICLR 2025
Information retrieval › document retrieval
reasoning-intensive retrieval
0.912025
BRIGHT: A Realistic and Challenging Benchmark for Reasoning-Intensive Retrieval · ICLR 2025
Information retrieval
retrieval models
0.912025
BRIGHT: A Realistic and Challenging Benchmark for Reasoning-Intensive Retrieval · ICLR 2025
Systems and software security
vulnerability discovery
0.612022
Hardening attack surfaces with formally proven binary format parsers · PLDI 2022
Program verification › code-level verification
verified parsing
0.612022
Hardening attack surfaces with formally proven binary format parsers · PLDI 2022
Medical and health informatics
computational pathology
0.412019
Atlas of Digital Pathology: A Generalized Hierarchical Histological Tissue Type-Annotated Database for Deep Learning · CVPR 2019
Privacy and data protection
social network privacy
0.412019
Moving Beyond Set-It-And-Forget-It Privacy Settings on Social Media · CCS 2019
Information retrieval
evaluation
0.312025
BRIGHT: A Realistic and Challenging Benchmark for Reasoning-Intensive Retrieval · ICLR 2025
Information retrieval › evaluation › test collection
retrieval benchmark
0.312025
BRIGHT: A Realistic and Challenging Benchmark for Reasoning-Intensive Retrieval · ICLR 2025
Machine learning › Learning paradigms
multi-label classification
0.112019
Atlas of Digital Pathology: A Generalized Hierarchical Histological Tissue Type-Annotated Database for Deep Learning · CVPR 2019
Bioinformatics and computational biology › structural bioinformatics › nucleic acid structure analysis
RNA structural bioinformatics
0.012004
RAG: RNA-As-Graphs database-concepts, analysis, features · Bioinform. 2004
Bioinformatics and computational biology › molecular informatics › cheminformatics
molecular graph representation
0.012004
RAG: RNA-As-Graphs database-concepts, analysis, features · Bioinform. 2004

Methods — techniques the papers use, named apart from their topics

formal proof · 1.1contrastive reinforcement learning · 0.9supervised learning · 0.8convolutional neural network · 0.8user study · 0.4friend interaction features · 0.4classifier · 0.4topological complexity ranking · 0.0graph enumeration · 0.0
YearPublicationVenuePosition
2025 A Single Goal is All You Need: Skills and Exploration Emerge from Contrastive RL without Rewards, Demonstrations, or Subgoals
abstract
In this paper, we present empirical evidence of skills and directed exploration emerging from a simple RL algorithm long before any successful trials are observed. For example, in a manipulation task, the agent is given a single observation of the goal state (see Fig. 1) and learns skills, first for moving its end-effector, then for pushing the block, and finally for picking up and placing the block. These skills emerge before the agent has ever successfully placed the block at the goal location and without the aid of any reward functions, demonstrations, or manually-specified distance metrics. Once the agent has learned to reach the goal state reliably, exploration is reduced. Implementing our method involves a simple modification of prior work and does not require density estimates, ensembles, or any additional hyperparameters. Intuitively, the proposed method seems like it should be terrible at exploration, and we lack a clear theoretical understanding of why it works so effectively, though our experiments provide some hints.
Grace Liu, Michael Tang, Benjamin Eysenbach
ICLR2
2025 BRIGHT: A Realistic and Challenging Benchmark for Reasoning-Intensive Retrieval
abstract
Existing retrieval benchmarks primarily consist of information-seeking queries (e.g., aggregated questions from search engines) where keyword or semantic-based retrieval is usually sufficient. However, many complex real-world queries require in-depth reasoning to identify relevant documents that go beyond surface form matching. For example, finding documentation for a coding question requires understanding the logic and syntax of the functions involved. To better benchmark retrieval on such challenging queries, we introduce BRIGHT, the first text retrieval benchmark that requires intensive reasoning to retrieve relevant documents. Our dataset consists of 1,398 real-world queries spanning diverse domains such as economics, psychology, mathematics, coding, and more. These queries are drawn from naturally occurring or carefully curated human data. Extensive evaluation reveals that even state-of-the-art retrieval models perform poorly on BRIGHT. The leading model on the MTEB leaderboard (Muennighoff et al., 2023), which achieves a score of 59.0 nDCG@10,1 produces a score of nDCG@10 of 18.0 on BRIGHT. We show that incorporating explicit reasoning about the query improves retrieval performance by up to 12.2 points. Moreover, incorporating retrieved documents from the top-performing retriever boosts question answering performance by over 6.6 points. We believe that BRIGHT paves the way for future research on retrieval systems in more realistic and challenging settings.
Hongjin Su, Howard Yen, Mengzhou Xia, Niklas Muennighoff, Han-yu Wang, Haisu Liu, Zachary S. Siegel, Michael Tang, Ruoxi Sun 0002, Jinsung Yoon, Sercan Ö. Arik, Danqi Chen 0001, Tao Yu 0009
ICLR10
2022 Hardening attack surfaces with formally proven binary format parsers
abstract
With an eye toward performance, interoperability, or legacy concerns, low-level system software often must parse binary encoded data formats. Few tools are available for this task, especially since the formats involve a mixture of arithmetic and data dependence, beyond what can be handled by typical parser generators. As such, parsers are written by hand in languages like C, with inevitable errors leading to security vulnerabilities.
Nikhil Swamy, Tahina Ramananandro, Aseem Rastogi, Irina Spiridonova, Haobin Ni, Dmitry Malloy, Juan Vazquez, Michael Tang, Omar Cardona, Arti Gupta
PLDI8
2019 Moving Beyond Set-It-And-Forget-It Privacy Settings on Social Media
abstract
When users post on social media, they protect their privacy by choosing an access control setting that is rarely revisited. Changes in users' lives and relationships, as well as social media platforms themselves, can cause mismatches between a post's active privacy setting and the desired setting. The importance of managing this setting combined with the high volume of potential friend-post pairs needing evaluation necessitate a semi-automated approach. We attack this problem through a combination of a user study and the development of automated inference of potentially mismatched privacy settings. A total of 78 Facebook users reevaluated the privacy settings for five of their Facebook posts, also indicating whether a selection of friends should be able to access each post. They also explained their decision. With this user data, we designed a classifier to identify posts with currently incorrect sharing settings. This classifier shows a 317% improvement over a baseline classifier based on friend interaction. We also find that many of the most useful features can be collected without user intervention, and we identify directions for improving the classifier's accuracy.
Mainack Mondal, Günce Su Yilmaz, Noah Hirsch, Mohammad Taha Khan, Michael Tang, Christopher Tran 0001, Chris Kanich, Blase Ur, Elena Zheleva
CCS5
2019 Atlas of Digital Pathology: A Generalized Hierarchical Histological Tissue Type-Annotated Database for Deep Learning
abstract
In recent years, computer vision techniques have made large advances in image recognition and been applied to aid radiological diagnosis. Computational pathology aims to develop similar tools for aiding pathologists in diagnosing digitized histopathological slides, which would improve diagnostic accuracy and productivity amidst increasing workloads. However, there is a lack of publicly-available databases of (1) localized patch-level images annotated with (2) a large range of Histological Tissue Type (HTT). As a result, computational pathology research is constrained to diagnosing specific diseases or classifying tissues from specific organs, and cannot be readily generalized to handle unexpected diseases and organs. In this paper, we propose a new digital pathology database, the ``Atlas of Digital Pathology'' (or ADP), which comprises of 17,668 patch images extracted from 100 slides annotated with up to 57 hierarchical HTTs. Our data is generalized to different tissue types across different organs and aims to provide training data for supervised multi-label learning of patch-level HTT in a digitized whole slide image. We demonstrate the quality of our image labels through pathologist consultation and by training three state-of-the-art neural networks on tissue type classification. Quantitative results support the visually consistency of our data and we demonstrate a tissue type-based visual attention aid as a sample tool that could be developed from our database.
Mahdi S. Hosseini, Lyndon Chan, Gabriel Tse, Michael Tang, Sajad Norouzi, Corwyn Rowsell, Konstantinos N. Plataniotis, Savvas Damaskinos
CVPR4
2004 RAG: RNA-As-Graphs database-concepts, analysis, features
abstract
MOTIVATION: Understanding RNA's structural diversity is vital for identifying novel RNA structures and pursuing RNA genomics initiatives. By classifying RNA secondary motifs based on correlations between conserved RNA secondary structures and functional properties, we offer an avenue for predicting novel motifs. Although several RNA databases exist, no comprehensive schemes are available for cataloguing the range and diversity of RNA's structural repertoire. RESULTS: Our RNA-As-Graphs (RAG) database describes and ranks all mathematically possible (including existing and candidate) RNA secondary motifs on the basis of graphical enumeration techniques. We represent RNA secondary structures as two-dimensional graphs (networks), specifying the connectivity between RNA secondary structural elements, such as loops, bulges, stems and junctions. We archive RNA tree motifs as 'tree graphs' and other RNAs, including pseudoknots, as general 'dual graphs'. All RNA motifs are catalogued by graph vertex number (a measure of sequence length) and ranked by topological complexity. The RAG inventory immediately suggests candidates for novel RNA motifs, either naturally occurring or synthetic, and thereby might stimulate the prediction and design of novel RNA motifs. AVAILABILITY: The database is accessible on the web at http://monod.biomath.nyu.edu/rna
Hin Hark Gan, Daniela Fera, Julie Zorn, Nahum Shiffeldrim, Michael Tang, Uri Laserson, Namhee Kim, Tamar Schlick
Bioinform.5
2003 Mining Emerging Substrings
abstract
We introduce a new type of KDD patterns called emerging substrings. In a sequence database, an emerging substring (ES) of a data class is a substring which occurs more frequently in that class rather than in other classes. ESs are important to sequence classification as they capture significant contrasts between data classes and provide insights for the construction of sequence classifiers. We propose a suffix tree-based framework for mining ESs, and study the effectiveness of applying one or more pruning techniques in different stages of our ES mining algorithm. Experimental results show that if the target class is of a small population with respect to the whole database, which is the normal scenario in single-class ES mining, most of the pruning techniques would achieve considerable performance gain.
Sarah Chan, Ben Kao, Chi Lap Yip, Michael Tang
DASFAA4
2000 Selection of Melody Lines for Music Databases
abstract
One major approach to music retrieval is to model music as a sequence of features, after which traditional information retrieval techniques are applied on the sequence. Because of the temporal nature of music and the inexactness of user queries, most effort on music retrieval systems focus on issues such as indexing and approximation match. In contrast, the processing of music before feature extraction, such as the identification of a melody track, were often considered easy or done. This may be the case in a controlled environment, such as one for musicology research, where the pieces are carefully analyzed by human beings before being submitted to the database. However, in an environment where large volumes of music is obtained from the Web, manual music analysis is impractical. Since many well-known musical features often pertain to the melody of musical pieces, and users often remember the melody of a song, algorithms that select the melody tracks of a piece are important for Web-based content-based retrieval systems. We describe a number of algorithms for automatic melody track selection in a music retrieval context. We also study the performance of the algorithms by comparing their answers to those judged by human beings.
Michael Tang, Chi Lap Yip, Ben Kao
COMPSAC1