Wallapak Tavanapong

dblp:79/5821 · DBLP profile ↗
← Back
57ranked-venue papers
6as first author
10since 2021 · last 2025
0000-0002-7799-2015ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 21 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 16 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 13 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 2 first-author · 2 since 2021Computer networks · 8 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 7 · 1 first-author · 1 since 2021Systems, architecture and hardware · 2Software engineering, systems software and programming languages · 1 · 1 first-author
YearPublicationVenuePosition
2025 CellFMCount: A Fluorescence Microscopy Dataset, Benchmark, and Methods for Cell Counting
abstract
Accurate cell counting is essential in various biomedical research and clinical applications, including cancer diagnosis, stem cell research, and immunology. Manual counting is laborintensive and error-prone, motivating automation through deep learning techniques. However, training reliable deep learning models requires large amounts of high-quality annotated data, which is difficult and time-consuming to produce manually. Consequently, existing cell-counting datasets are often limited, frequently containing fewer than 500 images. In this work, we introduce a large-scale annotated dataset comprising 3,023 images from immunocytochemistry experiments related to cellular differentiation, containing over 430,000 manually annotated cell locations. The dataset presents significant challenges: high cell density, overlapping and morphologically diverse cells, a long-tailed distribution of cell count per image, and variation in staining protocols. We benchmark three categories of existing methods: regression-based, crowd-counting, and cell-counting techniques on a test set with cell counts ranging from 10 to 2,126 cells per image. We also evaluate how the Segment Anything Model (SAM) can be adapted for microscopy cell counting using only dot-annotated datasets. As a case study, we implement a density-map-based adaptation of SAM (SAM-Counter) and report a mean absolute error (MAE) of 22.12, which outperforms existing approaches (second-best MAE of 27.46). Our results underscore the value of the dataset and the benchmarking framework for driving progress in automated cell counting and provide a robust foundation for future research and development.
Abdurahman Ali Mohammed, Catherine Fonder, Wallapak Tavanapong, Donald S. Sakaguchi, Qi Li 0012, Surya K. Mallapragada
ICDM4
2025 Reducing Domain Gap with Diffusion-Based Domain Adaptation for Cell Counting
abstract
Generating realistic synthetic microscopy images is critical for training deep learning models in label-scarce environments, such as cell counting with many cells per image. However, traditional domain adaptation methods often struggle to bridge the domain gap when synthetic images lack the complex textures and visual patterns of real samples. In this work, we adapt the Inversion-Based Style Transfer (InST) framework originally designed for artistic style transfer to biomedical microscopy images. Our method combines latent-space Adaptive Instance Normalization with stochastic inversion in a diffusion model to transfer the style from real fluorescence microscopy images to synthetic ones, while weakly preserving content structure.We evaluate the effectiveness of our InST-based synthetic dataset for downstream cell counting by pre-training and finetuning EfficientNet-B0 models on various data sources, including real data, hard-coded synthetic data, and the public Cell200-s dataset. Models trained with our InST-synthesized images achieve up to 37% lower Mean Absolute Error (MAE) compared to models trained on hard-coded synthetic data, and a 52% reduction in MAE compared to models trained on Cell200-s (from 53.70 to 25.95 MAE). Notably, our approach also outperforms models trained on real data alone (25.95 vs. 27.74 MAE). Further improvements are achieved when combining InST-synthesized data with lightweight domain adaptation techniques such as DACS with CutMix. These findings demonstrate that InST-based style transfer most effectively reduces the domain gap between synthetic and real microscopy data. Our approach offers a scalable path for enhancing cell counting performance while minimizing manual labeling effort. The source code and resources are publicly available at: https://github.com/MohammadDehghan/InST-Microscopy.
Mohammad Dehghanmanshadi, Wallapak Tavanapong
ICMLA2
2025 Synthesized Image Training Techniques: On Improving Model Performance Using Confusion
abstract
The performance of supervised deep learning image classifiers has significantly improved with large, labeled datasets and increased computing power. However, obtaining large, labeled image datasets in areas like medicine is expensive. This study seeks to improve model performance on limited labeled datasets by reducing confusion. We observed that misclassification (or confusion) between classes is usually more prevalent between specific classes. Thus, we developed a synthesized image training technique (SIT2), a novel confusion-based training framework that identifies pairs of classes with high confusion and synthesizes not-sure images from these pairs. The not-sure images are utilized in three new training strategies as follows: (1) the not-sure training strategy pretrains a model using not-sure images and the original training images, (2) the sure-or-not strategy pretrains with synthesized sure or not-sure images, and (3) the multi-label strategy pretrains with synthesized images but predicts the original class(es) of the synthesized images. Finally, the pretrained model is fine-tuned on the original dataset. An extensive evaluation was conducted on five medical and nonmedical datasets. Several improvements are statistically significant, which shows the promising future of our confusion-based training framework.
Azeez Idris, Mohammed Khaleel, Wallapak Tavanapong, Piet C. de Groen
ACM Trans. Multim. Comput. Commun. Appl.3
2024 How Much Do Prompting Methods Help LLMs on Quantitative Reasoning with Irrelevant Information?
abstract
Real-world quantitative reasoning problems are complex, often including extra information irrelevant to the question (or "IR noise" for short). State-of-the-art (SOTA) prompting methods have increased the Large Language Model's ability for quantitative reasoning on grade-school Math Word Problems (MWPs). To assess how well these SOTA methods handle IR noise, we constructed four new datasets with IR noise, each consisting of 300 problems from each of the four public datasets: MAWPS, ASDiv, SVAMP, and GSM8K, with added IR noise. We called the collection of these new datasets "MPN"--Math Word Problems with IR Noise. We evaluated SOTA prompting methods using MPN. We propose Noise Reduction Prompting (NRP) and its variant (NRP+) to reduce the impact of IR noise. Findings: Our IR noise significantly degrades the performance of Chain-of-Thought (CoT) Prompting on three different backend models: ChatGPT (gpt-3.5-turbo-0613), PaLM2, and Llama3-8B-instruct. Among them, ChatGPT offers the best accuracy on MPN with and without IR noise. With IR noise, performances of CoT, Least-To-Most Prompting, Progressive-Hint Prompting, and Program-aided Language Models with ChatGPT were significantly impacted, each with an average accuracy drop of above 12%. NRP is least impacted by the noise, with a drop in average accuracy to only around 1.9%. Our NRP+ and NRP perform comparably in the presence of IR noise.
Seok Hwan Song, Wallapak Tavanapong
CIKM2
2024 VisActive: Visual-concept-based Active Learning for Image Classification under Class Imbalance
abstract
Active learning methods recommend the most informative images from a large unlabeled dataset for manual labeling. These methods improve the performance of an image classifier while minimizing manual labeling efforts. We propose VisActive, a visual-concept-based active learning method for image classification under class imbalance. VisActive learns a visual concept, a generalized representation that holds the most important image characteristics for class prediction, and then recommends for each class four sets of unlabeled images with different visual concepts to increase the diversity and enlarge the training dataset. Experimental results on four datasets show that VisActive outperforms the state-of-the-art deep active learning methods.
Mohammed Khaleel, Azeez Idris, Wallapak Tavanapong, Jacob Pratt, Jung-Hwan Oh 0001, Piet C. de Groen
ACM Trans. Multim. Comput. Commun. Appl.3
2023 Interpreting Deep Text Quantification Models
YunQi Bang, Mohammed Khaleel, Wallapak Tavanapong
DEXA (2)3
2023 IDCIA: Immunocytochemistry Dataset for Cellular Image Analysis
abstract
We present a new annotated microscopic cellular image dataset to improve the effectiveness of machine learning methods for cellular image analysis. Cell counting is an important step in cell analysis. Typically, domain experts manually count cells in a microscopic image. Automated cell counting can potentially eliminate this tedious, time-consuming process. However, a good, labeled dataset is required for training an accurate machine learning model. Our dataset includes microscopic images of cells, and for each image, the cell count and the location of individual cells. The data were collected as part of an ongoing study investigating the potential of electrical stimulation to modulate stem cell differentiation and possible applications for neural repair. Compared to existing publicly available datasets, our dataset has more images of cells stained with more variety of antibodies (protein components of immune responses against invaders) typically used for cell analysis. The experimental results on this dataset indicate that none of the five existing models under this study are able to achieve sufficiently accurate count to replace the manual methods. The dataset is available at https://figshare.com/articles/dataset/Dataset/21970604.
Abdurahman Ali Mohammed, Catherine Fonder, Donald S. Sakaguchi, Wallapak Tavanapong, Surya K. Mallapragada, Azeez Idris
MMSys4
2022 Training Strategy for Limited Labeled Data by Learning from Confusion
abstract
A major problem of training deep learning image classifiers is the limited availability of domain-specific labeled datasets. This problem is particularly pressing in medicine due to the cost and expertise needed for annotation. This paper proposes a training strategy to reduce model misclassification by 1) using weakly supervised learning to learn class confusion, 2) creating a new class with synthetic training data highlighting the confusion, and 3) training with the expanded training data and using transfer learning. We tested our new training strategy using the open medical (Kvasir) and non-medical (CIFAR10) datasets. The proposed training strategy improves classification accuracy, precision, and recall when applied to a small subset of the training data.
Azeez Idris, Mohammed Khaleel, Wallapak Tavanapong, Jung-Hwan Oh 0001, Piet C. de Groen
ICIP3
2022 Artificial Intelligence for Colonoscopy: Past, Present, and Future
abstract
During the past decades, many automated image analysis methods have been developed for colonoscopy. Real-time implementation of the most promising methods during colonoscopy has been tested in clinical trials, including several recent multi-center studies. All trials have shown results that may contribute to prevention of colorectal cancer. We summarize the past and present development of colonoscopy video analysis methods, focusing on two categories of artificial intelligence (AI) technologies used in clinical trials. These are (1) analysis and feedback for improving colonoscopy quality and (2) detection of abnormalities. Our survey includes methods that use traditional machine learning algorithms on carefully designed hand-crafted features as well as recent deep-learning methods. Lastly, we present the gap between current state-of-the-art technology and desirable clinical features and conclude with future directions of endoscopic AI technology development that will bridge the current gap.
Wallapak Tavanapong, Jung-Hwan Oh 0001, Michael Riegler 0001, Mohammed Khaleel, Bhuvan Mittal, Piet C. de Groen
IEEE J. Biomed. Health Informatics1
2021 Hierarchical Visual Concept Interpretation for Medical Image Classification
abstract
Most state-of-the-art local interpretation methods explain the behavior of deep learning classification models by assigning importance scores to image pixels based on how influential each pixel was towards the final decision. These interpretations are unable to provide further details to aid understanding of a complex concept in a domain such as medicine. We propose a novel Hierarchical Visual Concept (HVC) interpretation framework for CNN-based image classification models. As an explanation of the classification decision of a given image, HVC presents a concept hierarchy of most relevant visual concepts at multiple semantic levels. These concepts are automatically learned during training such that the lower-level concepts in the hierarchy support the corresponding higher-level concepts. Our quantitative and qualitative evaluation of the interpretation of VGG16 and ResNet50 classifiers on public and private colonoscopy image datasets shows very promising results.
Mohammed Khaleel, Wallapak Tavanapong, Johnny S. Wong, Jung-Hwan Oh 0001, Piet C. de Groen
CBMS2
2020 Real-Time Feedback for Colonoscopy in a Multicenter Clinical Trial
abstract
We report the technical challenges, solutions, and lessons learned from deploying real-time feedback systems in three hospitals as part of a multi-center controlled clinical trial to improve quality of colonoscopy. Previous clinical trials were conducted in one center. The technical challenges for our multicenter clinical trial include 1) reducing additional work by the endoscopists to utilize real-time feedback, 2) handling different colonoscopy practices at different hospitals, and 3) training an effective CNN-based classification model with a large variety of patterns of data in day-to-day colonoscopy practice. We report performance of our real-time systems over a period of 20 weeks at each hospital. We conclude that CNN-based classification can achieve very good performance in real-world deployment when trained with high quality data.
Wallapak Tavanapong, Jung-Hwan Oh 0001, Gavin Kijkul, Jacob Pratt, Johnny S. Wong, Piet C. de Groen
CBMS1
2020 A Framework for Deep Quantification Learning
Mohammed Khaleel, Wallapak Tavanapong, Adisak Sukul, David A. M. Peterson
ECML/PKDD (1)3
2019 CMAIR: content and mask-aware image retargeting
Hon-Hang Chang, Timothy K. Shih, Carl K. Chang, Wallapak Tavanapong
Multim. Tools Appl.4
2018 Similarity-Based Active Learning for Image Classification Under Class Imbalance
abstract
Many image classification tasks (e.g., medical image classification) have a severe class imbalance problem. Convolutional neural network (CNN) is currently a state-of-the-art method for image classification. CNN relies on a large training dataset to achieve high classification performance. However, manual labeling is costly and may not even be feasible for medical domain. In this paper, we propose a novel similarity-based active deep learning framework (SAL) that deals with class imbalance. SAL actively learns a similarity model to recommend unlabeled rare class samples for experts' manual labeling. Based on similarity ranking, SAL recommends high confidence unlabeled common class samples for automatic pseudo-labeling without experts' labeling effort. To the best of our knowledge, SAL is the first active deep learning framework that deals with a significant class imbalance. Our experiments show that SAL consistently outperforms two other recent active deep learning methods on two challenging datasets. What's more, SAL obtains nearly the upper bound classification performance (using all the images in the training dataset) while the domain experts labeled only 5.6% and 7.5% of all images in the Endoscopy dataset and the Caltech-256 dataset, respectively. SAL significantly reduces the experts' manual labeling efforts while achieving near optimal classification performance.
Chuanhai Zhang, Wallapak Tavanapong, Gavin Kijkul, Johnny S. Wong, Piet C. de Groen, Jung-Hwan Oh 0001
ICDM2
2017 Identifying Policy Agenda Sub-Topics in Political Tweets based on Community Detection
abstract
The explosive use of twitter in the political landscape presents new avenues for tracking political conversations at federal and state level. Tweets are used by state and federal government bodies to present citizens with information about future and present policies. It is also used by political candidates to express their views on policy changes, laws and to campaign for legislative body elections, the most recent example being the 2016 US presidential elections. In this paper, we use supervised learning, textual semantic similarity and community detection techniques to find actively discussed policy agenda sub-topics among political tweets within a certain time period. Specifically, we target tweets pertaining to major policy agendas published by state representatives in US, to try and discern the major policy sub-topics that they address using their twitter accounts. Using our method, we demonstrate how we achieve a high accuracy in terms of Topic Recall and Order Recall, by comparing the output of our proposed method with sub-topic annotations done by domain experts.
Rohit Iyer, Johnny S. Wong, Wallapak Tavanapong, David A. M. Peterson
ASONAM3
2017 Social Media in State Politics: Mining Policy Agendas Topics
abstract
Twitter is a popular online microblogging service that has become widely used by politicians to communicate with their constituents. Gaining understanding of the influence of Twitter in state politics in the United States cannot be achieved without proper computational tools. We present the first attempt to automatically classify tweets of state legislatures (policy makers at the state level) into major policy agenda topics defined by Policy Agendas Project (PAP), which was initiated to group national policies. We investigated the effectiveness of three popular machine learning algorithms, Support Vector Machine (SVM), Convolutional Neural Networks (CNN), and Long Short-Term Memory Network (LSTM). We proposed a new synthetic data augmentation method to further improve classification performance. Our experimental results show that CNN provides the best F1 score of 78.3%. The new data augmentation method improves the classification perfromance by about 2%. Our tool provides a good prediction of the top three popular PAP topics in each month, which is useful for tracking popular PAP topics over time and across states and for comparing with national policy agendas.
Rihui Li, Johnny S. Wong, Wallapak Tavanapong, David A. M. Peterson
ASONAM4
2017 Online video ad measurement for political science research
abstract
With over 1 billion users and over 4 billion video views a day, YouTube is the most popular online video sharing website. Over a million advertisers use the Google ad platform which provides the ability to target specific groups of users. Political campaigns have used YouTube to reach voters. However, political science literature does not have information about how online video ads are targeted to voters via YouTube. This research made two contributions. (1) We developed “CyAds Meter,” a unique web bot platform that collects information about video and image ads delivered on YouTube. Our bot system adheres to YouTube's terms of service and avoids financial cost to online advertisers. Our two-tailed Wilcoxon-Mann-Whitney test verified that the bot and the human with the same online profile see the same number of ads on average with 5% significant level. In other words, our bot does not significantly miss ads nor see more ads than the human does. (2) We used CyAds Meter to provide evidence for political science scholars that the numbers of ads aired in the battleground and non-battleground states prior to the 2016 presidential election is statistically significant different at a 5% significant level. CyAds Meter is a potentially useful for political science research to study micro-targeting tactics of online ad election campaigns.
Adisak Sukul, Baskar Gopalakrishnan, Wallapak Tavanapong, David A. M. Peterson
IEEE BigData3
2017 Real-Time Instrument Scene Detection in Screening GI Endoscopic Procedures
abstract
We describe a new and effective real-time solution for detecting video segments showing an instrument used during diagnostic or therapeutic operations in endoscopic procedures. In addition, we present a new method to collect a large training dataset: similarity-based data augmentation. This method automates most of the creation of a large training dataset and prevents extensive manual effort to collect and annotate training data by domain experts. Convolutional Neural Network (CNN) analysis using the training data collected with similarity-based data augmentation yields an average F1 score within 1% of that of the CNN analysis using a large manually collected training dataset.
Chuanhai Zhang, Wallapak Tavanapong, Johnny S. Wong, Piet C. de Groen, Jung-Hwan Oh 0001
CBMS2
2016 Automated Coding of Political Video Ads for Political Science Research
abstract
With the advent of new media technology and the ability to identify more information about potential voters, political campaigns have aggressively changed their campaign strategies. Election campaigns increasingly rely on online video advertising to reach voters. Until now, the contents of these ads are manually coded for political science research to study campaign strategies. Manual coding is tremendously time consuming and not scalable to handle the expected increase in online ads. We make the first attempt to investigate automated coding of the content of political video ads for political science research. Specifically, we focus on the problem of classifying a political ad into one of these categories: attack ads, promoting ads, and contrast ads. Together with the domain expert, we introduce a concrete definition for each of these categories. We made available the ground truth labels of 773 political ads of the 2016 primary presidential election. We investigate the effectiveness of several classifiers using single modality and two modalities. The best average F1 score is 0.845 using text features from audio and embedded text in image frames.
Chuanhai Zhang, Adisak Sukul, Wallapak Tavanapong, David A. M. Peterson
ISM4
2016 Multimedia and Medicine: Teammates for Better Disease Detection and Survival
abstract
Health care has a long history of adopting technology to save lives and improve the quality of living. Visual information is frequently applied for disease detection and assessment, and the established fields of computer vision and medical imaging provide essential tools. It is, however, a misconception that disease detection and assessment are provided exclusively by these fields and that they provide the solution for all challenges. Integration and analysis of data from several sources, real-time processing, and the assessment of usefulness for end-users are core competences of the multimedia community and are required for the successful improvement of health care systems. We have conducted initial investigations into two use cases surrounding diseases of the gastrointestinal (GI) tract, where the detection of abnormalities provides the largest chance of successful treatment if the initial observation of disease indicators occurs before the patient notices any symptoms. Although such detection is typically provided visually by applying an endoscope, we are facing a multitude of new multimedia challenges that differ between use cases. In real-time assistance for colonoscopy, we combine sensor information about camera position and direction to aid in detecting, investigate means for providing support to doctors in unobtrusive ways, and assist in reporting. In the area of large-scale capsular endoscopy, we investigate questions of scalability, performance and energy efficiency for the recording phase, and combine video summarization and retrieval questions for analysis.
Michael Riegler 0001, Mathias Lux, Carsten Griwodz, Concetto Spampinato, Thomas de Lange, Sigrun Losada Eskeland, Konstantin Pogorelov, Wallapak Tavanapong, Peter Thelin Schmidt, Cathal Gurrin, Dag Johansen, Håvard D. Johansen, Pål Halvorsen
ACM Multimedia8
2014 Abnormal image detection in endoscopy videos using a filter bank and local binary patterns
Ruwan Dharshana Nawarathna, Jung-Hwan Oh 0001, Jayantha Muthukudage, Wallapak Tavanapong, Johnny S. Wong, Piet C. de Groen, Shou Jiang Tang
Neurocomputing4
2014 Part-Based Multiderivative Edge Cross-Sectional Profiles for Polyp Detection in Colonoscopy
abstract
This paper presents a novel technique for automated detection of protruding polyps in colonoscopy images using edge cross-section profiles (ECSP). We propose a part-based multiderivative ECSP that computes derivative functions of an edge cross-section profile and segments each of these profiles into parts. Therefore, we can model or extract features suitable for each part. Our features obtained from the parts can effectively describe complex properties of protruding polyps including the shape of the parts, texture, and protrusion and smoothness of the polyp surface. We evaluated our method against two existing polyp image detection techniques on 42 different polyps, including those with little protrusion. Each polyp has a large variation of appearance in viewing angles, light conditions, and scales in different images. The evaluation showed that our technique outperformed the existing techniques in both accuracy and analysis time. Our method has a higher area under the free-response receiver operating characteristic curve. For instance, when both techniques have a true positive rate for polyp image detection of 81.4%, the average number of false regions per image of our technique is 0.32 compared to 1.8 of the best existing technique under study. Additionally, our technique can precisely mark edges of candidate polyp regions as visual feedback. These results altogether indicate that our technique is promising to provide visual feedback of polyp regions in clinical practice.
Yi Wang 0012, Wallapak Tavanapong, Johnny S. Wong, Jung-Hwan Oh 0001, Piet C. de Groen
IEEE J. Biomed. Health Informatics2
2013 Near Real-Time Retroflexion Detection in Colonoscopy
abstract
Colonoscopy is the most popular screening tool for colorectal cancer. Recent studies reported that retroflexion during colonoscopy helped to detect more polyps. Retroflexion is an endoscope maneuver that enables visualization of internal mucosa along the shaft of the endoscope, enabling visualization of the mucosa area that is difficult to see with typical forward viewing. This paper describes our new method that detects the retroflexion during colonoscopy. We propose region shape and location (RSL) features and edgeless edge cross-section profile (ECSP) features that encapsulate important properties of endoscope appearance and edge information during retroflexion. Our experimental results on 50 colonoscopy test videos show that a simple ensemble classifier using both ECSP and RSL features can effectively identify retroflexion in terms of analysis time and detection rate.
Yi Wang 0012, Wallapak Tavanapong, Johnny S. Wong, Jung-Hwan Oh 0001, Piet C. de Groen
IEEE J. Biomed. Health Informatics2
2011 SAPPHIRE middleware and software development kit for medical video analysis
abstract
This paper presents SAPPHIRE — a novel middleware and software development kit, developed to reduce implementation efforts in utilizing task parallelism to speed up execution time of analysis of a stream of data, such as medical video. As a case study, we implemented a real-time quality measurement of colonoscopy using SAPPHIRE. We increased the number of threads in our case study from 4 (prior to our middleware) to 28 with ease. As a result, through better load-balancing and taking better advantage of available processor cores, we increased the case study's maximum processing rate from 40 frames per second to 90 frames per second.
Sean Stanek, Wallapak Tavanapong, Johnny S. Wong, Jung-Hwan Oh 0001, Ruwan Dharshana Nawarathna, Jayantha Muthukudage, Piet C. de Groen
CBMS2
2011 Computer-aided detection of retroflexion in colonoscopy
abstract
Colonoscopy is the most popular screening tool for colorectal cancer. Recent studies reported that retroflexion during colonoscopy improved polyp yields. Retroflexion is an endoscope maneuver that enables visualization of internal mucosa along the shaft of the endoscope, enabling visualization of the mucosa area that is difficult to see with typical forward viewing. This paper describes our new method that detects endoscopic images showing retroflexion. This problem has not been investigated in the literature. We propose new region features that encapsulate important properties of endoscope appearance during retroflexion. Our experimental results on 25 colonoscopy videos show that trained Decision Tree classifiers can effectively identify retroflexion in the rectum at 92.0% accuracy and 94.4% precision.
Yi Wang 0012, Wallapak Tavanapong, Johnny S. Wong, Jung-Hwan Oh 0001, Piet C. de Groen
CBMS2
2011 Color Based Stool Region Detection in Colonoscopy Videos for Quality Measurements
Jayantha Muthukudage, Jung-Hwan Oh 0001, Wallapak Tavanapong, Johnny S. Wong, Piet C. de Groen
PSIVT (1)3
2009 3D Reconstruction of Colon Segments from Colonoscopy Images
abstract
A new algorithm for reconstruction of a 3D virtual colon segment from an individual image captured from colonoscopy is presented. Colonoscopy is currently the gold standard method for detection and prevention of colorectal cancer. However, the protective effect depends on the amount of colon mucosa that is actually seen by the endoscopist. The proposed algorithm takes contours of colon folds in the image as input and calculates the depth and the slant angle of each fold. Finally, the colon mucosa is created using Cubic Bezier Curve interpolation between the folds. The proposed algorithm is an important step toward 3D reconstruction of the virtual colon from video of an entire colonoscopy procedure. The reconstruction is potentially useful to estimate percent of colon mucosa inspected by the endoscopist.
DongHo Hong, Wallapak Tavanapong, Johnny S. Wong, Jung-Hwan Oh 0001, Piet C. de Groen
BIBE2
2008 Modeling of End-to-End Available Bandwidth in Wide Area Network
abstract
Modeling the available bandwidth of a path using a known stochastic process is one possible method for estimating future available bandwidth along the path without explicit support from network routers. Our two hypotheses for the stochastic process are as follows. First, an auto-regressive integrated moving-average process (ARIMA) is a suitable model for the available bandwidth over time of a path. Second, the available bandwidth over time of a path can be modeled as a self-similar process. We verify both hypotheses using R statistical software and available bandwidth data sets published by Stanford Linear Accelerator Center (SLAC). Our results indicate that the available bandwidth over time of an end-to-end path can be modeled as fractional Gaussian Noise (FGN) and seasonal fractional ARIMA (SFARIMA) processes. On the other hand, we found that an ARIMA process is not a good model for available bandwidth over time of an end-to-end path.
Wanida Putthividhya, Arka P. Ghosh, Wallapak Tavanapong
ISPA3
2008 Safe-Time: Distributed Real-Time Monitoring of cKNN in Mobile Peer-to-Peer Networks
abstract
A continuous k nearest neighbor (cKNN) query is a query that continuously returns a set of k nearest moving objects (mobile hosts) to a given query point. For example, report three nearest moving sensors to a given location continuously. Most existing research efforts focus on centralized solutions. In a mobile peer-to-peer network (M-P2P), a centralized approach incurs expensive communication cost. In this paper, we propose Safe-Time - a distributed solution for cKNN given a stationary query point for M-P2P. The two key features are as follows. 1) Actual execution of a cKNN query is not needed during a safe-time period since the query result is guaranteed to remain the same during this period. 2) Once the safe-time expires, execution of a cKNN query involves only objects in a circular band of width equal to the estimated distance between the kthand the k + 1thnearest neighbors. To further reduce communication cost for dense queries, we introduce Unite-Safe-Time that executes one virtual query derived from nearby queries instead of executing each of them separately. Our simulation result shows that the proposed distributed solutions outperform a centralized solution under a range of conditions. Safe-Time incurs up to 2/3 less communication cost compared to a centralized solution. Unite-Safe-Time shows up to 1/3 less communication cost than Safe-Time in our study.
Ying Cai 0001, Wallapak Tavanapong
MDM3
2008 Caching collaboration and cache allocation in peer-to-peer video systems
Ying Cai 0001, Wallapak Tavanapong
Multim. Tools Appl.3
2008 Deadline-constrained media uploading systems
Mu Zhang 0004, Johnny S. Wong, Wallapak Tavanapong, Jung-Hwan Oh 0001, Piet C. de Groen
Multim. Tools Appl.3
2007 Polyp Detection in Colonoscopy Video using Elliptical Shape Feature
abstract
Early detection of polyps and cancers is one of the most important goals of colonoscopy. Computer-based analysis of video files using texture features, as has been proposed for polyps of the stomach and colon, has two major limitations: this method uses a fixed size analysis window and relies heavily on a training set of images for accuracy. To overcome these limitations, we propose a new technique focusing on shape instead of texture in this paper. The proposed polyp region detection method is based on the elliptical shape that is common for nearly all small colon polyps.
Sae Hwang, Jung-Hwan Oh 0001, Wallapak Tavanapong, Johnny S. Wong, Piet C. de Groen
ICIP (2)3
2007 Informative frame classification for endoscopy video
Jung-Hwan Oh 0001, Sae Hwang, JeongKyu Lee, Wallapak Tavanapong, Johnny S. Wong, Piet C. de Groen
Medical Image Anal.4
2007 A double patching technique for efficient bandwidth sharing in video-on-demand systems
Ying Cai 0001, Wallapak Tavanapong, Kien A. Hua
Multim. Tools Appl.2
2007 OCS: An effective caching scheme for video streaming on overlay networks
Minh Tran 0002, Wallapak Tavanapong, Wanida Putthividhya
Multim. Tools Appl.2
2006 Manual Annotation of Colonoscopy Videos: A First Step towards Automation
Piet C. de Groen, Wallapak Tavanapong, Jung-Hwan Oh 0001, Johnny S. Wong
AMIA2
2006 Detecting Malicious Peers in Overlay Multicast Streaming
abstract
Overlay multicast streaming is built out of loosely coupled end-hosts (peers) that contribute resources to stream media to other peers. Peers, however, can be malicious. They may intentionally wish to disrupt the multicast service or cause confusions to other peers. We propose two new schemes to detect malicious peers in overlay multicast streaming. These schemes compute a level of trust for each peer in the network. Peers with a trust value below a threshold are considered to be malicious. Results from our simulations indicate that the proposed schemes can detect malicious peers with medium to high accuracy, depending on cheating patterns and malicious peer percentages
Samarth Shetty, Patricio A. Galdames, Wallapak Tavanapong, Ying Cai 0001
LCN3
2006 Sharing Location Dependent Experiences in MANET
abstract
This paper investigates a new problem of sharing location dependent experiences among mobile hosts in mobile adhoc networks. An experience is location and observer dependent. In other words, experiences of different people witnessing the same event may be quite different. The ability to retrieve prior experiences observed in a given area in advance is very useful for newcomers wishing to enter the same vicinity. Solving this problem is vital for important applications such as hurricane rescue missions, combat missions, and deep space or deep sea exploration. In this paper, we propose a distributed solution that lets mobile hosts share their experiences efficiently. Our simulation results show that our approach significantly outperforms a centralized approach.
Ying Cai 0001, Wallapak Tavanapong
MDM3
2006 A scalable cost-effective video broadcasting system for on-demand video services
Simon Sheu, Wallapak Tavanapong, Kien A. Hua
Multim. Tools Appl.2
2005 A priority forwarding technique for efficient and fast flooding in wireless ad hoc networks
abstract
In this paper, we propose a new technique, called priority forwarding, for efficient and fast flooding operations in wireless ad hoc networks. The new scheme is featured by dynamic delay and priority checking. The former feature allows a host to wait as long as possible to refrain from retransmission, minimizing retransmission overhead. The priority checking feature, on the other hand, allows a flooding packet to be propagated as quickly as possible, keeping flooding latency low. Unlike many location-aided flooding techniques, priority forwarding requires each host to know only the distance of its 1-hop neighbors, instead of their exact locations. Therefore, it has low implementation cost. For performance evaluation, we compare priority forwarding with some existing techniques using simulation. Our results show that under most scenarios, the new technique performs many times better in reducing packet retransmissions and flooding latency.
Ying Cai 0001, Wallapak Tavanapong
ICCCN3
2005 Peers-assisted Dynamic Content Distribution Networks
abstract
Content distribution networks (CDNs) have been proposed to primarily distribute Web and some limited streaming audio/video content over the Internet. Current CDNs consist of fixed content nodes placed at strategic locations on the Internet to improve the service latency and to reduce server and network load. Although the approach of delivering content with a fixed CDN has enjoyed some successes, it is often difficult to upgrade and extend a CDN to serve more users due to a very high deployment cost. In this paper, we propose that a CDN be agile while delivering content. We make the case for a dynamic CDN architecture in which a CDN leverages local peers to deliver the CDN's content to local users. We define a formal problem for using peers to increase the capacity of a CDN and we propose an algorithm to help a CDN to dynamically recruit local peers to be part of the CDN. The CDN and the recruited local peers cooperate to serve content to other local users. Our performance evaluation shows promising results with the new dynamic CDN architecture. Our simple CDN growing strategies dramatically improve the average service latency in streaming video content
Minh Tran 0002, Wallapak Tavanapong
LCN2
2005 Automatic measurement of quality metrics for colonoscopy videos
abstract
Colonoscopy is the accepted screening method for detection of colorectal cancer or its precursor lesions, colorectal polyps. Indeed, colonoscopy has contributed to a decline in the number of colorectal cancer related deaths. However, not all cancers or large polyps are detected at the time of colonoscopy, and methods to investigate why this occurs are needed. We present a new computer-based method that allows automated measurement of a number of metrics that likely reflect the quality of the colonoscopic procedure. The method is based on analysis of a digitized video file created during colonoscopy, and produces information regarding insertion time, withdrawal time, images at the time of maximal intubation, the time and ratio of clear versus blurred or non-informative images, and a first estimate of effort performed by the endoscopist. As these metrics can be obtained automatically, our method allows future quality control in the day-to-day medical practice setting on a large scale. In addition, our method can be adapted to other healthcare procedures. Last but not least, our method may be useful to assess progress during colonoscopy training, or as part of endoscopic skills assessment evaluations.
Sae Hwang, Jung-Hwan Oh 0001, JeongKyu Lee, Yu Cao 0002, Wallapak Tavanapong, Danyu Liu, Johnny S. Wong, Piet C. de Groen
ACM Multimedia5
2005 Video Management in Peer-to-Peer Systems
abstract
Providing scalable video services in a peer-to-peer (P2P) environment is challenging. Since videos are typically large and require high communication bandwidth for delivery, many peers may be unwilling to cache them in whole to serve others. In this paper, the authors addressed two fundamental research problems in providing scalable P2P video services, namely (1) how a host can find enough video pieces, which may scatter among the whole system, to assemble a complete video, and (2) given a limited buffer size, what part of a video a host should cache. A new distributed video management technique was proposed. This scheme organizes hosts into a number of cells, each of which is a distinct set of hosts which together can supply a video in its entirety. A client looking for a video can stop its search as soon as it finds a host that caches any part of the video. Caching operations can be coordinated within each cell to balance data redundancy in the system. The extensive study on a Gnutella-like simulation network shows convincingly the performance advantage of the new scheme.
Ying Cai 0001, Wallapak Tavanapong
Peer-to-Peer Computing3
2004 Distributed core selection with QoS support
abstract
Core-based routing with quality of service (QoS) support is essential in facilitating multi-sender multimedia multicast applications such as video conferencing and virtual collaboration applications. In this paper, we introduce a new distributed core selection protocol that has the following desirable properties. First, the protocol utilizes a new distributed primary core selection algorithm that selects as many primary cores per multicast group as necessary, to maximize the number of group members with satisfied QoS requirements. Second, the protocol is distributed, preventing a single router from becoming a hot spot and a single point of failure during core selection. Lastly, the protocol employs a distributed backup core selection algorithm to provide quick recovery should some primary cores fail. Our analytical experiments show that the proposed protocol significantly satisfies more group members with noticeably less communication overhead than a recent core selection algorithm with QoS support using a single core.
Wanida Putthividhya, Wallapak Tavanapong, Minh Tran 0002, Johnny S. Wong
ICC2
2004 Providing scalable on-demand video services for heterogeneous receivers
abstract
To provide scalable video-on-demand services, precious system bandwidth must be shared among video requests. Although many efficient bandwidth-sharing techniques have been proposed, they are all designed for homogeneous receivers, i.e., the clients are assumed to have the same receiving bandwidth. For networks with client heterogeneity, these techniques either cannot work or have to compromise their performance. We address this problem and propose an efficient solution allowing each client to use its whole receiving bandwidth to download data. To serve a client, the server first selects a set of serving channels for the client according to its actual receiving bandwidth. The video data needed for this client are then scheduled and dynamically adjusted for delivery over these channels. We evaluate the performance of the new technique using simulation and compare it with an existing scheme. Our study shows that by effectively leveraging client heterogeneity, the new technique achieves significantly smaller service latency.
Ying Cai 0001, Wallapak Tavanapong, Johnny S. Wong
ICME3
2004 A framework for parsing colonoscopy videos for semantic units
abstract
Colonoscopy is an important screening procedure for colorectal cancer. During this procedure, the endoscopist visually inspects the colon. Currently, there is no content-based analysis and retrieval system that automatically analyzes videos captured from colonoscopic procedures and provides a user-friendly and efficient access to important content. Such a system will be valuable for endoscopic research and education. The first necessary step for the analysis is parsing for semantic units. Since the characteristics of colonoscopy videos differ from those of videos studied in the literature, we introduce a new video parsing framework that includes: (i) a new scene definition and a new video parsing paradigm; (ii) a novel scene segmentation algorithm using audio analysis and finite state automata to recognize scenes and associated boundaries. Our experimental results show average precision and recall of 95% and 81%, respectively, for parsing scenes. The framework is extensible to videos captured from other endoscopic procedures such as upper gastrointestinal endoscopy, enteroscopy, cystoscopy, and laparoscopy.
Yu Cao 0002, Wallapak Tavanapong, Johnny S. Wong, Jung-Hwan Oh 0001, Piet C. de Groen
ICME2
2004 Parsing and browsing tools for colonoscopy videos
abstract
Colonoscopy is an important screening tool for colorectal cancer. During a colonoscopic procedure, a tiny video camera at the tip of the endoscope generates a video signal of the internal mucosa of the colon. The video data are displayed on a monitor for real-time analysis by the endoscopist. We call videos captured from colonoscopic procedures colonoscopy videos. Because these videos possess unique characteristics, new types of semantic units and parsing techniques are required. In this paper, we define new semantic units called operation shots, each is a segment of visual and audio data that correspond to a therapeutic or biopsy operation. We introduce a new spatio-temporal analysis technique to detect operation shots. Our experiments on colonoscopy videos demonstrate that the technique does not miss any meaningful operation shots and incurs a small number of false operation shots. Our prototype parsing software implements the operation shot detection technique along with our other techniques previously developed for colonoscopy videos. Our browsing tool enables users to quickly locate operation shots of interest. The proposed technique and software are useful (1) for post-procedure reviews and analyses for causes of complications due to biopsy or therapeutic operations, (2) for developing an effective content-based retrieval system for colonoscopy videos to facilitate endoscopic research and education, and (3) for development of a systematic approach to assess endoscopists' procedural skills.
Yu Cao 0002, Dalei Li, Wallapak Tavanapong, Jung-Hwan Oh 0001, Johnny S. Wong, Piet C. de Groen
ACM Multimedia3
2004 Video delivery technologies for large-scale deployment of multimedia applications
abstract
Deployment of a large-scale multimedia streaming application requires an enormous amount of server and network resources. The simplest delivery technique allocates server resources for each specific request. This technique is very expensive and is not scalable to support a very large user community such as the Internet. Hence, the past decade has witnessed tremendous research efforts to facilitate cost-effective, large-scale deployment of multimedia streaming applications. In this paper, we describe three complementary research approaches: server transmission schemes using multicast, streaming strategies with application layer multicast, and proxy caching techniques. We discuss pros and cons of these technologies and provide our observations on current business solutions.
Kien A. Hua, Mounir A. Tantaoui, Wallapak Tavanapong
Proc. IEEE3
2004 Shot clustering techniques for story browsing
abstract
Automatic video segmentation is the first and necessary step for organizing a long video file into several smaller units. The smallest basic unit is a shot. Relevant shots are typically grouped into a high-level unit called a scene. Each scene is part of a story. Browsing these scenes unfolds the entire story of a film, enabling users to locate their desired video segments quickly and efficiently. Existing scene definitions are rather broad, making it difficult to compare the performance of existing techniques and to develop a better one. This paper introduces a stricter scene definition for narrative films and presents ShotWeave, a novel technique for clustering relevant shots into a scene using the stricter definition. The crux of ShotWeave is its feature extraction and comparison. Visual features are extracted from selected regions of representative frames of shots. These regions capture essential information needed to maintain viewers' thought in the presence of shot breaks. The new feature comparison is developed based on common continuity-editing techniques used in film making. Experiments were performed on full-length films with a wide range of camera motions and a complex composition of shots. The experimental results show that ShotWeave outperforms two recent techniques utilizing global visual features in terms of segmentation accuracy and time.
Wallapak Tavanapong
IEEE Trans. Multim.1
2003 Image Retrieval Based on Regions of Interest
abstract
Query-by-example is the most popular query model in recent content-based image retrieval (CBIR) systems. A typical query image includes relevant objects (e.g., Eiffel Tower), but also irrelevant image areas (including background). The irrelevant areas limit the effectiveness of existing CBIR systems. To overcome this limitation, the system must be able to determine similarity based on relevant regions alone. We call this class of queries region-of-interest (ROI) queries and propose a technique for processing them in a sampling-based matching framework. A new similarity model is presented and an indexing technique for this new environment is proposed. Our experimental results confirm that traditional approaches, such as Local Color Histogram and Correlogram, suffer from the involvement of irrelevant regions. Our method can handle ROI queries and provide significantly better performance. We also assessed the performance of the proposed indexing technique. The results clearly show that our retrieval procedure is effective for large image data sets.
Khanh Vu, Kien A. Hua, Wallapak Tavanapong
IEEE Trans. Knowl. Data Eng.3
2002 Video caching network for on-demand video streaming
abstract
Previous years have seen a tremendous growth in streaming continuous media such as videos over the Internet, resulting in an enormous increase in the demand on various server and networking resources. To minimize service delays and reduce the loads placed on these resources, we propose a video caching network (VCN) that utilizes an aggregated cache space of distributed systems along the delivery path between the server and the users for caching popular videos. VCN is set up and adjusted dynamically according to users' locations and request patterns. Our simulation results indicate that VCN outperforms the video proxy approach by a significant margin.
Wallapak Tavanapong, Minh Tran 0002, Srikanth Krishnamohan
GLOBECOM1
2001 Design and implementation of a video browsing system for the Internet
abstract
Abstract Recent advances in multimedia processing technologies, internet‐working technologies, and the World Wide Web phenomenon have resulted in a vast creation and use of digital videos. Due to this reason, an efficient technique to locate and retrieve a desired video from a remote video archive is needed. A trial‐and‐error approach popularly used in current Web search engines is not applicable for searching for a desired video segment since the technique incurs intolerable delays. In this paper, we present the design and implementation of a video browsing system that lets the user view a summary of a selected video and search within the video while being downloaded so that the user can determine the relevance of the video as early as possible. The system is inexpensive and scalable, making it suitable for large‐scale distributed systems such as the Internet. The browsing system consists of two major software components: a video server, and a video browser and player calledVideoCenterimplemented using Microsoft DirectShow multimedia development kit. The implementation of VideoCenter enables us to assess the ease of use of DirectShow as well as its drawbacks in developing multimedia applications. Copyright © 2001 John Wiley & Sons, Ltd.
Wallapak Tavanapong, Kien A. Hua
Softw. Pract. Exp.1
1999 Performance of Load Balancing Techniques for Join Operations in Shared-Noting Database Management Systems
Kien A. Hua, Wallapak Tavanapong, Yu-lung Lo
J. Parallel Distributed Comput.2
1999 2PSM: An Efficient Framework for Searching Video Information in a Limited-Bandwidth Environment
Kien A. Hua, Wallapak Tavanapong, James Zijun Wang
Multim. Syst.2
1997 A Framework for Supporting Previewing and VCR Operations in a Low Bandwidth Environment
abstract
We propose a novel delivery mechanism called 2-Phase Service Model to deliver video data to home users connected to the Internet through a low-bandwidth device such as a modem.In our scheme, non-adjacent fragments of the requested video file are first downloaded to the client during Initialization Phase.The missing fragments are transmitted to the client as the video is being played out, using a novel pipelining technique.This scheme offers several benefits as follows.First, it allows the user to perform a quick preview through the video with minimal delay.Second, it naturally supports VCR functionality with almost no delay as demonstrated by the simulation results shown in the paper.Finally, our mathematical analysis shows that despite the desirable features it offers, 2-Phase Service Model does not incur any more initialization delay than-that of the conventional pipelining technique.
Wallapak Tavanapong, Kien A. Hua, James Zijun Wang
ACM Multimedia1
1996 Scheduling Queries for Parallel Execution on Multicomputer Database Management Systems
Yu-lung Lo, Kien A. Hua, Wallapak Tavanapong
DEXA3
1995 A Performance Evaluation of Load Balancing Techniques for Join Operations on Multicomputer Database Systems
abstract
There has been a wealth of research in the area of parallel join algorithms. Among them, hash-based algorithms are particularly suitable for shared-nothing database systems. The effectiveness of these techniques depends on the uniformity in the distribution of the join attribute values. When this condition is not met, a severe fluctuation may occur among the bucket sizes, causing uneven workload for the processing nodes. Many parallel join algorithms with load balancing capability have been proposed to address this problem. Among them, the sampling and incremental approaches have been shown to provide an improvement over the more conventional methods. The comparison between these two approaches, however, has not been investigated. In this paper, we improve these techniques and implement them on an nCUBE/2 parallel computer to compare their performance. Our study indicates that the sampling technique is the better approach.>
Kien A. Hua, Wallapak Tavanapong, Honesty C. Young
ICDE2