VLDB 2026 Research / reviewers in the wild / expert
Huu Nguyen
dblp:236/0642
· DBLP profile ↗
4ranked-venue papers
0as first author
4since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 4 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Language models and text generation · 50% Face, body and person analysis · 23% Efficient and distributed learning · 17% | |
| Human-computer interaction and pervasive computing
1 paper |
Human-AI interaction · 100% |
Topics — the 10 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning › data curation
training data curation |
1.3 | 2 | 2024 | RedPajama: an Open Dataset for Training Large Language Models · NeurIPS 2024 The BigScience ROOTS Corpus: A 1.6TB Composite Multilingual Dataset · NeurIPS 2022 |
Computer vision › Face, body and person analysis › facial expression analysis
facial expression recognition |
0.9 | 1 | 2025 | EmoNet-Face: An Expert-Annotated Benchmark for Synthetic Emotion Recognition · NeurIPS 2025 |
Machine learning › Probabilistic and Bayesian machine learning
data filtering |
0.8 | 1 | 2024 | RedPajama: an Open Dataset for Training Large Language Models · NeurIPS 2024 |
Natural language and speech › Language models and text generation › large language model
open language model development |
0.8 | 1 | 2024 | RedPajama: an Open Dataset for Training Large Language Models · NeurIPS 2024 |
Natural language and speech › Language models and text generation
alignment |
0.7 | 1 | 2023 | OpenAssistant Conversations - Democratizing Large Language Model Alignment · NeurIPS 2023 |
Natural language and speech › Language models and text generation › alignment › preference alignment
human feedback alignment |
0.7 | 1 | 2023 | OpenAssistant Conversations - Democratizing Large Language Model Alignment · NeurIPS 2023 |
Natural language and speech › Language models and text generation
instruction tuning |
0.7 | 1 | 2023 | OpenAssistant Conversations - Democratizing Large Language Model Alignment · NeurIPS 2023 |
Natural language and speech › Language models and text generation
multilingual dataset |
0.6 | 1 | 2022 | The BigScience ROOTS Corpus: A 1.6TB Composite Multilingual Dataset · NeurIPS 2022 |
Natural language and speech › Language models and text generation
multilingual language models |
0.6 | 1 | 2022 | The BigScience ROOTS Corpus: A 1.6TB Composite Multilingual Dataset · NeurIPS 2022 |
Human-AI interaction › affective computing
affective state recognition |
0.3 | 1 | 2025 | EmoNet-Face: An Expert-Annotated Benchmark for Synthetic Emotion Recognition · NeurIPS 2025 |
Methods — techniques the papers use, named apart from their topics
synthetic face generation · 1.7expert annotation · 1.7quality signal analysis · 0.8ablation study · 0.8supervised fine-tuning · 0.7reinforcement learning from human feedback · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | EmoNet-Face: An Expert-Annotated Benchmark for Synthetic Emotion RecognitionabstractEffective human-AI interaction relies on AI's ability to accurately perceive and interpret human emotions. Current benchmarks for vision and vision-language models are severely limited, offering a narrow emotional spectrum that overlooks nuanced states (e.g., bitterness, intoxication) and fails to distinguish subtle differences between related feelings (e.g., shame vs. embarrassment). Existing datasets also often use uncontrolled imagery with occluded faces and lack demographic diversity, risking significant bias. To address these critical gaps, we introduce EmoNet Face, a comprehensive benchmark suite. EmoNet Face features: (1) A novel 40-category emotion taxonomy, meticulously derived from foundational research to capture finer details of human emotional experiences. (2) Three large-scale, AI-generated datasets (EmoNet HQ, Binary, and Big) with explicit, full-face expressions and controlled demographic balance across ethnicity, age, and gender. (3) Rigorous, multi-expert annotations for training and high-fidelity evaluation. (4) We build Empathic Insight Face, a model achieving human-expert-level performance on our benchmark. The publicly released EmoNet Face suite—taxonomy, datasets, and model—provides a robust foundation for developing and evaluating AI systems with a deeper understanding of human emotions. Christoph Schuhmann, Robert Kaczmarczyk, Gollam Rabby, Maurice Kraus, Felix Friedrich, Huu Nguyen, Krishna Kalyan, Kourosh Nadi, Kristian Kersting, Sören Auer |
NeurIPS | 6 |
| 2024 | RedPajama: an Open Dataset for Training Large Language ModelsabstractLarge language models are increasingly becoming a cornerstone technology in artificial intelligence, the sciences, and society as a whole, yet the optimal strategies for dataset composition and filtering remain largely elusive. Many of the top-performing models lack transparency in their dataset curation and model development processes, posing an obstacle to the development of fully open language models. In this paper, we identify three core data-related challenges that must be addressed to advance open-source language models. These include (1) transparency in model development, including the data curation process, (2) access to large quantities of high-quality data, and (3) availability of artifacts and metadata for dataset curation and analysis. To address these challenges, we release RedPajama-V1, an open reproduction of the LLaMA training dataset. In addition, we release RedPajama-V2, a massive web-only dataset consisting of raw, unfiltered text data together with quality signals and metadata.Together, the RedPajama datasets comprise over 100 trillion tokens spanning multiple domains and with their quality signals facilitate the filtering of data, aiming to inspire the development of numerous new datasets. To date, these datasets have already been used in the training of strong language models used in production, such as Snowflake Arctic, Salesforce's XGen and AI2's OLMo. To provide insight into the quality of RedPajama, we present a series of analyses and ablation studies with decoder-only language models with up to 1.6B parameters. Our findings demonstrate how quality signals for web data can be effectively leveraged to curate high-quality subsets of the dataset, underscoring the potential of RedPajama to advance the development of transparent and high-performing language models at scale. Maurice Weber, Daniel Y. Fu, Quentin Anthony, Yonatan Oren, Shane Adams, Anton Alexandrov, Xiaozhong Lyu, Huu Nguyen, Xiaozhe Yao, Virginia Adams, Ben Athiwaratkun, Rahul Chalamala, Kezhen Chen, Max Ryabinin, Tri Dao, Percy Liang, Christopher Ré, Irina Rish, Ce Zhang 0001 |
NeurIPS | 8 |
| 2023 | OpenAssistant Conversations - Democratizing Large Language Model AlignmentabstractAligning large language models (LLMs) with human preferences has proven to drastically improve usability and has driven rapid adoption as demonstrated by ChatGPT.Alignment techniques such as supervised fine-tuning (\textit{SFT}) and reinforcement learning from human feedback (\textit{RLHF}) greatly reduce the required skill and domain knowledge to effectively harness the capabilities of LLMs, increasing their accessibility and utility across various domains.However, state-of-the-art alignment techniques like \textit{RLHF} rely on high-quality human feedback data, which is expensive to create and often remains proprietary.In an effort to democratize research on large-scale alignment, we release OpenAssistant Conversations, a human-generated, human-annotated assistant-style conversation corpus consisting of 161,443 messages in 35 different languages, annotated with 461,292 quality ratings, resulting in over 10,000 complete and fully annotated conversation trees.The corpus is a product of a worldwide crowd-sourcing effort involving over 13,500 volunteers.Models trained on OpenAssistant Conversations show consistent improvements on standard benchmarks over respective base models.We release our code\footnote{\git} and data\footnote{\data} under a fully permissive licence. Andreas Köpf, Yannic Kilcher, Dimitri von Rütte, Sotiris Anagnostidis, Zhi Rui Tam, Keith Stevens, Abdullah Barhoum, Oliver Stanley, Richárd Nagyfi, Shahul ES, Sameer Suri, David Glushkov, Arnav Dantuluri, Andrew Maguire, Christoph Schuhmann, Huu Nguyen, Alexander Mattick |
NeurIPS | 17 |
| 2022 | The BigScience ROOTS Corpus: A 1.6TB Composite Multilingual DatasetabstractAs language models grow ever larger, the need for large-scale high-quality text datasets has never been more pressing, especially in multilingual settings. The BigScience workshop, a 1-year international and multidisciplinary initiative, was formed with the goal of researching and training large language models as a values-driven undertaking, putting issues of ethics, harm, and governance in the foreground. This paper documents the data creation and curation efforts undertaken by BigScience to assemble the Responsible Open-science Open-collaboration Text Sources (ROOTS) corpus, a 1.6TB dataset spanning 59 languages that was used to train the 176-billion-parameter BigScience Large Open-science Open-access Multilingual (BLOOM) language model. We further release a large initial subset of the corpus and analyses thereof, and hope to empower large-scale monolingual and multilingual modeling projects with both the data and the processing tools, as well as stimulate research around this large multilingual corpus. Hugo Laurençon, Lucile Saulnier, Thomas Wang, Christopher Akiki, Albert Villanova del Moral, Teven Le Scao, Leandro von Werra, Chenghao Mou, Eduardo G. Ponferrada, Huu Nguyen, Jörg Frohberg, Mario Sasko, Quentin Lhoest, Angelina McMillan-Major, Gérard Dupont, Stella Biderman, Anna Rogers, Loubna Ben Allal, Francesco De Toni, Giada Pistilli, Olivier Nguyen, Somaieh Nikpoor, Maraim Masoud, Pierre Colombo, Javier de la Rosa 0001, Paulo Villegas, Tristan Thrush, Shayne Longpre, Sebastian Nagel 0005, Leon Weber-Genzel, Manuel Muñoz, Daniel van Strien, Zaid Alyafeai, Khalid Almubarak, Minh Chien Vu, Itziar Gonzalez-Dios, Aitor Soroa, Kyle Lo, Manan Dey, Pedro Ortiz Suarez, Aaron Gokaslan, Shamik Bose, David Ifeoluwa Adelani, Long Phan, Hieu Tran, Ian Yu, Suhas Pai, Jenny Chim, Violette Lepercq, Suzana Ilic, Margaret Mitchell, Sasha Luccioni, Yacine Jernite |
NeurIPS | 10 |