EDBT 2026 Demo / reviewers in the wild / expert
Suppawong Tuarob
dblp:84/11514
· DBLP profile ↗
21ranked-venue papers
11as first author
11since 2021 · last 2026
0000-0002-5201-5699ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 10 · 6 first-author · 3 since 2021Artificial intelligence and machine learning · 7 · 5 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 2 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | XD-LDL: Improving Generalization in Acne Severity Classification via Cross-Dataset Label Distribution LearningabstractThis study introduces a novel approach to enhance acne severity classification by leveraging cross-dataset training within a Label Distribution Learning (LDL) framework. Traditional acne grading methods are often subjective and inconsistent, while automated systems struggle to generalize across diverse datasets, limiting their real-world applicability. This research addresses this challenge by integrating data from multiple sources (Acne04 and AcneSCU) to train a robust model capable of accurately classifying acne severity despite variations in image quality, lighting, and skin tone. We explore the impact of different backbone architectures (ResNet50, EfficientNet B7, and BEiT) within the LDL framework, which is specifically designed to handle the inherent ambiguity in acne severity labeling. Our results demonstrate that cross-dataset training within the LDL framework significantly improves generalization performance compared to models trained on individual datasets. Additionally, results demonstrate that integrating EfficientNet B7 and BEiT enhances the model’s accuracy and robustness. This novel approach offers a promising solution for developing scalable and reliable automated acne grading systems suitable for telemedicine and dermatology applications, ultimately aiding dermatologists in providing more precise and personalized treatment plans. Sirawich Vachmanus, Teerapong Rattananukrom, Suppawong Tuarob |
ACM Trans. Comput. Heal. | 3 |
| 2025 | Beyond administrative reports: a deep learning framework for classifying and monitoring crime and accidents leveraging large-scale online newsabstractAbstract The escalating prevalence of violent crimes and accidents underscores the urgent need for efficient and timely monitoring systems. Traditional methods reliant on administrative reports often suffer from significant delays. This paper proposes CRIMSON, a novel framework that leverages large-scale online news to provide real-time insights into crime and accident trends. CRIMSON utilizes a multi-label classification technique that leverages a fine-tuned, pre-trained, cross-lingual language model to accurately categorize news articles. Our experimental results, conducted on a substantial dataset of Thai news articles, demonstrate superior performance, achieving an average F1 score of 86%. Beyond classification, CRIMSON aggregates categorized news into real-time statistics, revealing strong correlations between news-reported incidents and official crime data. This study pioneers online news as a reliable and timely crime and accident monitoring source, offering valuable insights for law enforcement, policymakers, and researchers. Suppawong Tuarob, Phonarnun Tatiyamaneekul, Siripen Pongpaichet, Tanisa Tawichsri, Thanapon Noraset |
Neural Comput. Appl. | 1 |
| 2025 | Sprint2Vec: A Deep Characterization of Sprints in Iterative Software DevelopmentabstractIterative approaches like Agile Scrum are commonly adopted to enhance the software development process. However, challenges such as schedule and budget overruns still persist in many software projects. Several approaches employ machine learning techniques, particularly classification, to facilitate decision-making in iterative software development. Existing approaches often concentrate on characterizing a sprint to predict solely productivity. We introduce Sprint2Vec, which leverages three aspects of sprint information – sprint attributes, issue attributes, and the developers involved in a sprint, to comprehensively characterize it for predicting both productivity and quality outcomes of the sprints. Our approach combines traditional feature extraction techniques with automated deep learning-based unsupervised feature learning techniques. We utilize methods like Long Short-Term Memory (LSTM) to enhance our feature learning process. This enables us to learn features from unstructured data, such as textual descriptions of issues and sequences of developer activities. We conducted an evaluation of our approach on two regression tasks: predicting the deliverability (i.e., the amount of work delivered from a sprint) and quality of a sprint (i.e., the amount of delivered work that requires rework). The evaluation results on five well-known open-source projects (Apache, Atlassian, Jenkins, Spring, and Talendforge) demonstrate our approach's superior performance compared to baseline and alternative approaches. Morakot Choetkiertikul, Peerachai Banyongrakkul, Chaiyong Ragkhitwetsagul, Suppawong Tuarob, Khanh Hoa Dam, Thanwadee Sunetnanta |
IEEE Trans. Software Eng. | 4 |
| 2024 | TeReKG: A temporal collaborative knowledge graph framework for software team recommendationabstractSuccessful software development requires a cohesive team with the right mix of technical skills and the ability to collaborate effectively. However, forming a software team that can execute tasks with precision and efficiency requires a deep understanding of each member’s competence, experience, and cooperation history. Previously, automated software team selection has evaluated technical skills, cohesion, and cooperation history. However, the previous method had some limitations. Particularly, local features directly calculated from team members were subjective to the researchers’ views, and the method ignored the temporal aspect of open-source software development. To overcome these limitations, this paper proposes a knowledge-graph software team recommendation framework called TeReKG. This framework encapsulates temporal collaboration patterns and uses a temporal knowledge graph to encode software collaboration history, technical abilities, task dependencies, and project structure. TeReKG was against state-of-the-art team recommendation algorithms using three popular open-source software projects: Moodle, Apache, and Atlassian. The evaluation results show that TeReKG outperforms the state-of-the-art baselines in both single-role and team recommendation tasks. These findings demonstrate that knowledge graph embedding can be effectively utilized in automated recommendation tasks in software engineering. Additionally, this highlights the potential for knowledge graphs to capture global information that can benefit various software development applications, including impact prediction of software repositories, code clone detection, and source code retrieval. Pisol Ruenin, Morakot Choetkiertikul, Akara Supratak, Suppawong Tuarob |
Knowl. Based Syst. | 4 |
| 2023 | FALCoN: Detecting and classifying abusive language in social networks using context features and unlabeled data
Suppawong Tuarob, Manisa Satravisut, Pochara Sangtunchai, Sakunrat Nunthavanich, Thanapon Noraset |
Inf. Process. Manag. | 1 |
| 2022 | Quantifying effectiveness of team recommendation for collaborative software development
Noppadol Assavakamhaenghan, Waralee Tanaphantaruk, Ponlakit Suwanworaboon, Morakot Choetkiertikul, Suppawong Tuarob |
Autom. Softw. Eng. | 5 |
| 2022 | Language-agnostic deep learning framework for automatic monitoring of population-level mental health from social networks
Thanapon Noraset, Krittin Chatrinan, Tanisa Tawichsri, Tipajin Thaipisutikul, Suppawong Tuarob |
J. Biomed. Informatics | 5 |
| 2021 | Automatic team recommendation for collaborative software development
Suppawong Tuarob, Noppadol Assavakamhaenghan, Waralee Tanaphantaruk, Ponlakit Suwanworaboon, Saeed-Ul Hassan, Morakot Choetkiertikul |
Empir. Softw. Eng. | 1 |
| 2021 | DGSD: Distributed graph representation via graph statistical properties
Anwar Said, Saeed-Ul Hassan, Suppawong Tuarob, Raheel Nawaz, Mudassir Shabbir |
Future Gener. Comput. Syst. | 3 |
| 2021 | WabiQA: A Wikipedia-Based Thai Question-Answering System
Thanapon Noraset, Lalita Lowphansirikul, Suppawong Tuarob |
Inf. Process. Manag. | 3 |
| 2021 | Attributed Collaboration Network Embedding for Academic Relationship MiningabstractFinding both efficient and effective quantitative representations for scholars in scientific digital libraries has been a focal point of research. The unprecedented amounts of scholarly datasets, combined with contemporary machine learning and big data techniques, have enabled intelligent and automatic profiling of scholars from this vast and ever-increasing pool of scholarly data. Meanwhile, recent advance in network embedding techniques enables us to mitigate the challenges of large scale and sparsity of academic collaboration networks. In real-world academic social networks, scholars are accompanied with various attributes or features, such as co-authorship and publication records, which result in attributed collaboration networks. It has been observed that both network topology and scholar attributes are important in academic relationship mining. However, previous studies mainly focus on network topology, whereas scholar attributes are overlooked. Moreover, the influence of different scholar attributes are unclear. To bridge this gap, in this work, we present a novel framework of Attributed Collaboration Network Embedding (ACNE) for academic relationship mining. ACNE extracts four types of scholar attributes based on the proposed scholar profiling model, including demographics, research, influence, and sociability. ACNE can learn a low-dimensional representation of scholars considering both scholar attributes and network topology simultaneously. We demonstrate the effectiveness and potentials of ACNE in academic relationship mining by performing collaborator recommendation on two real-world datasets and the contribution and importance of each scholar attribute on scientific collaborator recommendation is investigated. Our work may shed light on academic relationship mining by taking advantage of attributed collaboration network embedding. Wei Wang 0077, Jiaying Liu 0006, Tao Tang 0007, Suppawong Tuarob, Feng Xia 0001, Zhiguo Gong, Irwin King |
ACM Trans. Web | 4 |
| 2020 | Deep Learning-based Extraction of Algorithmic Metadata in Full-Text Scholarly Documents
Iqra Safder, Saeed-Ul Hassan, Anna Visvizi, Thanapon Noraset, Raheel Nawaz, Suppawong Tuarob |
Inf. Process. Manag. | 6 |
| 2020 | Automatic Classification of Algorithm Citation Functions in Scientific LiteratureabstractComputer sciences and related disciplines evolve around developing, evaluating, and applying algorithms. Typically, an algorithm is not developed from scratch, but uses and builds upon existing ones, which often are proposed and published in scholarly articles. The ability to capture this evolution relationship among these algorithms in scientific literature would not only allow us to understand how a particular algorithm is composed, but also shed light on large-scale analysis of algorithmic evolution through different temporal spans and thematic scales. We propose to capture such evolution relationship between two algorithms by investigating the knowledge represented in citation contexts, where authors explain how cited algorithms are used in their works. A set of heterogeneous ensemble machine-learning methods is proposed, where the combination of two base classifiers trained with heterogeneous feature types is used to automatically identify the algorithm usage relationship. The proposed heterogeneous ensemble methods achieve the best average F1 of 0.749 and 0.905 for fine-grained and binary algorithm citation function classification, respectively. The success of this study will allow us to generate a large-scale algorithm citation network from a collection of scholarly documents representing multiple time spans, venues, and fields of study. Such a network will be used as an instrument not only to answer critical questions in algorithm search, such as identifying the most influential and generalizable algorithms, but also to study the evolution of algorithmic development and trends over time. Suppawong Tuarob, Sung Woo Kang, Poom Wettayakorn, Chanathip Pornprasit, Tanakitti Sachati, Saeed-Ul Hassan, Peter Haddawy |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2017 | How are you feeling?: A personalized methodology for predicting mental states from temporally observable physical and behavioral information
Suppawong Tuarob, Conrad S. Tucker, Soundar R. T. Kumara, C. Lee Giles, Aaron L. Pincus, David E. Conroy, Nilam Ram |
J. Biomed. Informatics | 1 |
| 2016 | AlgorithmSeer: A System for Extracting and Searching for Algorithms in Scholarly Big DataabstractAlgorithms are usually published in scholarly articles, especially in the computational sciences and related disciplines. The ability to automatically find and extract these algorithms in this increasingly vast collection of scholarly digital documents would enable algorithm indexing, searching, discovery, and analysis. Recently, AlgorithmSeer, a search engine for algorithms, has been investigated as part of CiteSeer' with the intent of providing a large algorithm database. Currently, over 200,000 algorithms have been extracted from over 2 million scholarly documents. This paper proposes a novel set of scalable techniques used by AlgorithmSeer to identify and extract algorithm representations in a heterogeneous pool of scholarly documents. Specifically, hybrid machine learning approaches are proposed to discover algorithm representations. Then, techniques to extract textual metadata for each algorithm are discussed. Finally, a demonstration version of AlgorithmSeer that is built on Solr/Lucene open source indexing and search system is presented. Suppawong Tuarob, Sumit Bhatia, Prasenjit Mitra 0001, C. Lee Giles |
IEEE Trans. Big Data | 1 |
| 2015 | Modeling Individual-Level Infection Dynamics Using Social Network InformationabstractEpidemic monitoring systems engaged in accurate discovery of infected individuals enable better understanding of the dynamics of epidemics and thus may promote effective disease mitigation or prevention. Currently, infection discovery systems require either physical participation of potential patients or provision of information from hospitals and health-care services. While social media has emerged as an increasingly important knowledge source that reflects multiple real world events, there is only a small literature examining how social media information can be incorporated into computational epidemic models. In this paper, we demonstrate how social media information can be incorporated into and improve upon traditional techniques used to model the dynamics of infectious diseases. Using flu infection histories and social network data collected from 264 students in a college community, we identify social network signals that can aid identification of infected individuals. Extending the traditional SIRS model, we introduce and illustrate the efficacy of an Online-Interaction-Aware Susceptible-Infected-Recovered-Susceptible (OIA-SIRS) model based on four social network signals for modeling infection dynamics. Empirical evaluations of our case study, flu infection within a college community, reveal that the OIA-SIRS model is more accurate than the traditional model, and also closely tracks the real-world infection rates as reported by CDC ILINet and Google Flu Trend. Suppawong Tuarob, Conrad S. Tucker, Marcel Salathé, Nilam Ram |
CIKM | 1 |
| 2015 | A hybrid approach to discover semantic hierarchical sections in scholarly documentsabstractScholarly documents are usually composed of sections, each of which serves a different purpose by conveying specific context. The ability to automatically identify sections would allow us to understand the semantics of what is different in different sections of documents, such as what was in the introduction, methodologies used, experimental types, trends, etc. We propose a set of hybrid algorithms to 1) automatically identify section boundaries, 2) recognize standard sections, and 3) build a hierarchy of sections. Our algorithms achieve an F-measure of 92.38% in section boundary detection, 96% accuracy (average) on standard section recognition, and 95.51% in accuracy in the section positioning task. Suppawong Tuarob, Prasenjit Mitra 0001, C. Lee Giles |
ICDAR | 1 |
| 2015 | PDFMEF: A Multi-Entity Knowledge Extraction Framework for Scholarly Documents and Semantic SearchabstractWe introduce PDFMEF, a multi-entity knowledge extraction framework for scholarly documents in the PDF format. It is implemented with a framework that encapsulates open-source extraction tools. Currently, it leverages PDFBox and TET for full text extraction, the scholarly document filter described in [5] for document classification, GROBID for header extraction, ParsCit for citation extraction, PDFFigures for figure and table extraction, and algorithm extraction [27]. While it can be run as a whole, the extraction tool in each module is highly customizable. Users can substitute default extractors with other extraction tools they prefer by writing a thin wrapper to implement the abstracts. The framework is designed to be scalable and is capable of running in parallel using a multi-processing technique in Python. Experiments indicate that the system with default setups is CPU bounded, and leaves a small footprint in the memory, which makes it best to run on a multi-core machine. The best performance using a dedicated server of 16 cores takes 1.3 seconds on average to process one PDF document. It is used to index extracted information and help users to quickly locate relevant results in published scholarly documents and to efficiently construct a large knowledge base in order to build a semantic scholarly search engine. Part of it is running on CiteSeerX digital library search engine. Jian Wu 0006, Jason Killian, Huaiyu Yang, Kyle Williams 0001, Sagnik Ray Choudhury, Suppawong Tuarob, Cornelia Caragea, C. Lee Giles |
K-CAP | 6 |
| 2014 | An ensemble heterogeneous classification methodology for discovering health-related knowledge in social media messages
Suppawong Tuarob, Conrad S. Tucker, Marcel Salathé, Nilam Ram |
J. Biomed. Informatics | 1 |
| 2013 | Discovering health-related knowledge in social media using ensembles of heterogeneous featuresabstractSocial media is emerging as a powerful source of communication, information dissemination and mining. Being colloquial and ubiquitous in nature makes it easier for users to express their opinions and preferences in a seamless, dynamic manner. Epidemic surveillance systems that utilize social media to detect the emergence of diseases have been proposed in the literature. These systems mostly employ traditional document classification techniques that represent a document with a bag of N-grams. However, such techniques are not optimal for social media where sparsity and noise are norms. The authors address the limitations posed by the traditional N-gram based methods and propose to use features that represent different semantic aspects of the data in combination with ensemble machine learning techniques to identify health-related messages in a heterogenous pool of social media data. Furthermore, the results reveal significant improvement in identifying health related social media content which can be critical in the emergence of a novel, unknown disease epidemic. Suppawong Tuarob, Conrad S. Tucker, Marcel Salathé, Nilam Ram |
CIKM | 1 |
| 2013 | Automatic Detection of Pseudocodes in Scholarly Documents Using Machine LearningabstractA significant number of scholarly articles in computer science and other disciplines contain algorithms that provide concise descriptions for solving a wide variety of computational problems. For example, Dijkstra's algorithm describes how to find the shortest paths between two nodes in a graph. Automatic identification and extraction of these algorithms from scholarly digital documents would enable automatic algorithm indexing, searching, analysis and discovery. An algorithm search engine, which identifies pseudocodes in scholarly documents and makes them searchable, has been implemented as a part of the CiteSeerX suite. Here, we illustrate the limitations of start-of-the-art rule based pseudocode detection approach, and present a novel set of machine learning based techniques that extend previous methods. Suppawong Tuarob, Sumit Bhatia, Prasenjit Mitra 0001, C. Lee Giles |
ICDAR | 1 |