Tai Nguyen 0005

dblp:340/8047 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Trustworthy machine learning · 75% Transfer learning and domain adaptation · 25%
Software engineering, system software, and programming languages
1 paper
Empirical software engineering · 100%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Transfer learning and domain adaptation
fine-tuning
0.712023
Explanation-based Finetuning Makes Models More Robust to Spurious Cues · ACL (1) 2023
Machine learning › Trustworthy machine learning › robustness › robust learning
noise-tolerant learning
0.712023
Software Entity Recognition with Noise-Robust Learning · ASE 2023
Machine learning › Trustworthy machine learning
robustness
0.712023
Explanation-based Finetuning Makes Models More Robust to Spurious Cues · ACL (1) 2023
Machine learning › Trustworthy machine learning › robustness › spurious correlation
spurious cue robustness
0.712023
Explanation-based Finetuning Makes Models More Robust to Spurious Cues · ACL (1) 2023
Empirical software engineering
mining software repositories
0.212023
Software Entity Recognition with Noise-Robust Learning · ASE 2023
Empirical software engineering › mining software repositories › question and answer sites
stack overflow analysis
0.212023
Software Entity Recognition with Noise-Robust Learning · ASE 2023

Methods — techniques the papers use, named apart from their topics

self-regularization · 1.3BERT · 1.3explanation-based finetuning · 0.7
YearPublicationVenuePosition
2023 Explanation-based Finetuning Makes Models More Robust to Spurious Cues
abstract
Josh Magnus Ludan, Yixuan Meng, Tai Nguyen, Saurabh Shah, Qing Lyu, Marianna Apidianaki, Chris Callison-Burch. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Josh Magnus Ludan, Yixuan Meng, Tai Nguyen 0005, Saurabh Shah, Qing Lyu 0001, Marianna Apidianaki, Chris Callison-Burch
ACL (1)3
2023 Software Entity Recognition with Noise-Robust Learning
abstract
Recognizing software entities such as library names from free-form text is essential to enable many software engineering (SE) technologies, such as traceability link recovery, automated documentation, and API recommendation. While many approaches have been proposed to address this problem, they suffer from small entity vocabularies or noisy training data, hindering their ability to recognize software entities mentioned in sophisticated narratives. To address this challenge, we leverage the Wikipedia taxonomy to develop a comprehensive entity lexicon with 79K unique software entities in 12 fine-grained types, as well as a large labeled dataset of over 1.7M sentences. Then, we propose self-regularization, a noise-robust learning approach, to the training of our software entity recognition (SER) model by accounting for many dropouts. Results show that models trained with self-regularization outperform both their vanilla counterparts and state-of-the-art approaches on our Wikipedia benchmark and two Stack Overflow benchmarks. We release our models11https://huggingface.co/taidng/wikiser-bert-base; https.//huggingface.co/taidng/wikiser-bert-large., data, and code for future research.22https://github.com/taidnguyen/software_entity_recognition
Tai Nguyen 0005, Yifeng Di, Joohan Lee, Muhao Chen 0001, Tianyi Zhang 0001
ASE1