VLDB 2026 Research / reviewers in the wild / expert
Theodore Lee
dblp:142/9535
· DBLP profile ↗
2ranked-venue papers
0as first author
1since 2021 · last 2023
0000-0003-0482-4819ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Trustworthy machine learning · 91% Information extraction and text analysis · 9% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
content moderation |
0.7 | 1 | 2023 | A Holistic Approach to Undesired Content Detection in the Real World · AAAI 2023 |
Machine learning › Trustworthy machine learning › robustness
overfitting mitigation |
0.7 | 1 | 2023 | A Holistic Approach to Undesired Content Detection in the Real World · AAAI 2023 |
Machine learning › Trustworthy machine learning
robustness |
0.7 | 1 | 2023 | A Holistic Approach to Undesired Content Detection in the Real World · AAAI 2023 |
Natural language and speech › Information extraction and text analysis
text classification |
0.2 | 1 | 2023 | A Holistic Approach to Undesired Content Detection in the Real World · AAAI 2023 |
Methods — techniques the papers use, named apart from their topics
data quality control · 0.7active learning · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | A Holistic Approach to Undesired Content Detection in the Real WorldabstractWe present a holistic approach to building a robust and useful natural language classification system for real-world content moderation. The success of such a system relies on a chain of carefully designed and executed steps, including the design of content taxonomies and labeling instructions, data quality control, an active learning pipeline to capture rare events, and a variety of methods to make the model robust and to avoid overfitting. Our moderation system is trained to detect a broad set of categories of undesired content, including sexual content, hateful content, violence, self-harm, and harassment. This approach generalizes to a wide range of different content taxonomies and can be used to create high-quality content classifiers that outperform off-the-shelf models. Todor Markov, Sandhini Agarwal, Tyna Eloundou, Theodore Lee, Steven Adler, Angela Jiang, Lilian Weng |
AAAI | 5 |
| 2015 | Risk factor detection for heart disease by applying text analytics in electronic medical recordsabstractIn the United States, about 600,000 people die of heart disease every year. The annual cost of care services, medications, and lost productivity reportedly exceeds 108.9 billion dollars. Effective disease risk assessment is critical to prevention, care, and treatment planning. Recent advancements in text analytics have opened up new possibilities of using the rich information in electronic medical records (EMRs) to identify relevant risk factors. The 2014 i2b2/UTHealth Challenge brought together researchers and practitioners of clinical natural language processing (NLP) to tackle the identification of heart disease risk factors reported in EMRs. We participated in this track and developed an NLP system by leveraging existing tools and resources, both public and proprietary. Our system was a hybrid of several machine-learning and rule-based components. The system achieved an overall F1 score of 0.9185, with a recall of 0.9409 and a precision of 0.8972. Manabu Torii, Jungwei Fan 0001, Weili Yang, Theodore Lee, Matthew T. Wiley, Daniel Zisook, Yang Huang 0008 |
J. Biomed. Informatics | 4 |