VLDB 2026 Research / reviewers in the wild / expert
Ting-Chuen Pong
dblp:25/1058
· DBLP profile ↗
60ranked-venue papers
6as first author
8since 2021 · last 2025
0000-0002-1699-8587ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 30 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 25 · 4 first-authorHuman-computer interaction and ubiquitous computing · 10 · 1 first-author · 3 since 2021Systems, architecture and hardware · 4Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | eLIVE: e-Learning Laboratory for Immersive Virtual EnvironmentsabstractThe adoption of virtual laboratories has seen significant growth, particularly within tertiary education, driven by the increasing need for flexible, accessible, and scalable learning environments. A significant development in this area is the integration of Virtual Reality (VR) technologies, leading to the emergence of VR Labs, which offer immersive, interactive simulations that enhance experiential learning. Recognizing the potential of these VR Labs to transform education, several Massive Open Online Courses (MOOCs) have begun incorporating them as part of the course content, providing students with hands-on experience that transcends geographical and physical limitations. Some of these VR Labs use avatar recordings to demonstrate lab procedures, but research on their design for effective learning and live learner monitoring remains limited. In this work, we examined how experiment complexity, equipment placement, and instructor avatar positioning affect learning in a VR Lab. By analyzing an existing avatar recording-based VR Lab, we designed an improved VR Lab by using the Polymerase Chain Reaction (PCR) test for COVID-19 in biomedical engineering as an example experiment. We conducted a user study with 10 users to evaluate the learning experience and gathered feedback on live monitoring and educational use of VR Labs from a learner's perspective. The findings are summarized as novel design guidelines for creating enhanced avatar recording-based VR Labs, focusing on the experiment steps, the equipment placement, the instructor avatar placement, a monitoring interface, and recommendations on utilizing VR Labs as a pedagogical tool in standalone experiences and a course component scenario. Pak Ming Fan, Santawat Thanyadit, Ting-Chuen Pong |
EDUCON | 3 |
| 2025 | Exploring Spatial Hybrid User Interface for Visual SensemakingabstractWe built a spatial hybrid system that combines a personal computer (PC) and virtual reality (VR) for visual sensemaking, addressing limitations in both environments. Although VR offers immense potential for interactive data visualization (e.g., large display space and spatial navigation), it can also present challenges such as imprecise interactions and user fatigue. At the same time, a PC offers precise and familiar interactions but has limited display space and interaction modality. Therefore, we iteratively designed a spatial hybrid system (PC+VR) to complement these two environments by enabling seamless switching between PC and VR environments. To evaluate the system's effectiveness and user experience, we compared it to using a single computing environment (i.e., PC-only and VR-only). Our study results (N=18) showed that spatial PC+VR could combine the benefits of both devices to outperform user preference for VR-only without a negative impact on performance from device switching overhead. Finally, we discussed future design implications. Wai Tong, Haobo Li 0003, Meng Xia 0002, Kamkwai Wong, Ting-Chuen Pong, Huamin Qu, Yalong Yang 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2025 | VisTellAR: Embedding Data Visualization to Short-Form Videos Using Mobile Augmented RealityabstractWith the rise of short-form video platforms and the increasing availability of data, we see the potential for people to share short-form videos embedded with data in situ (e.g., daily steps when running) to increase the credibility and expressiveness of their stories. However, creating and sharing such videos in situ is challenging since it involves multiple steps and skills (e.g., data visualization creation and video editing), especially for amateurs. By conducting a formative study (N=10) using three design probes, we collected the motivations and design requirements. We then built VisTellAR, a mobile AR authoring tool, to help amateur video creators embed data visualizations in short-form videos in situ. A two-day user study shows that participants (N=12) successfully created various videos with data visualizations in situ and they confirmed the ease of use and learning. AR pre-stage authoring was useful to assist people in setting up data visualizations in reality with more designs in camera movements and interaction with gestures and physical objects to storytelling. Wai Tong, Kento Shigyo, Linping Yuan, Mingming Fan 0001, Ting-Chuen Pong, Huamin Qu, Meng Xia 0002 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2024 | VisTA-LIVE: A Visualization Tool for Assessment of Laboratories in Virtual EnvironmentsabstractA Virtual Reality Laboratory (VR Lab) experiment refers to an experiment session that is being conducted in the virtual environment through Virtual Reality (VR) and aims to deliver procedural knowledge to students similar to that in a physical lab environment. While VR Lab is becoming more popular among education institutes as a learning tool for students, existing designs are mostly considered from a student's perspective. Instructors could only receive limited information on how the students are performing and could not provide useful feedback to aid the students' learning and evaluate their performance. This motivated us to create VisTA-LIVE: a Visualization Tool for Assessment of Laboratories In Virtual Environments. In this article, we present in detail the design thinking approach that was applied to create VisTA-LIVE. The tool is deployed in an Extended Reality (XR) environment, and we report the evaluation results with domain experts and discuss issues related to monitoring and assessing a live VR lab session which lay potential directions for future work. We also describe how the resulting design of the tool could be used as a reference for other education developers who wish to develop similar applications. Pak Ming Fan, Santawat Thanyadit, Ting-Chuen Pong |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2023 | Towards an Understanding of Distributed Asymmetric Collaborative Visualization on Problem-solvingabstractThis paper provided empirical knowledge of the user experience for using collaborative visualization in a distributed asymmetrical setting through controlled user studies. With the ability to access various computing devices, such as Virtual Reality (VR) head-mounted displays, scenarios emerge when collaborators have to or prefer to use different computing environments in different places. However, we still lack an understanding of using VR in an asymmetric setting for collaborative visualization. To get an initial understanding and better inform the designs for asymmetric systems, we first conducted a formative study with 12 pairs of participants. All participants collaborated in asymmetric (PC-VR) and symmetric settings (PC-PC and VR-VR). We then improved our asymmetric design based on the key findings and observations from the first study. Another ten pairs of participants collaborated with enhanced PC-VR and PC-PC conditions in a follow-up study. We found that a well-designed asymmetric collaboration system could be as effective as a symmetric system. Surprisingly, participants using PC perceived less mental demand and effort in the asymmetric setting (PC-VR) compared to the symmetric setting (PC-PC). We provided fine-grained discussions about the trade-offs between different collaboration settings. Wai Tong, Meng Xia 0002, Kamkwai Wong, Doug A. Bowman, Ting-Chuen Pong, Huamin Qu, Yalong Yang 0001 |
VR | 5 |
| 2023 | GestureLens: Visual Analysis of Gestures in Presentation VideosabstractAppropriate gestures can enhance message delivery and audience engagement in both daily communication and public presentations. In this article, we contribute a visual analytic approach that assists professional public speaking coaches in improving their practice of gesture training through analyzing presentation videos. Manually checking and exploring gesture usage in the presentation videos is often tedious and time-consuming. There lacks an efficient method to help users conduct gesture exploration, which is challenging due to the intrinsically temporal evolution of gestures and their complex correlation to speech content. In this article, we propose GestureLens, a visual analytics system to facilitate gesture-based and content-based exploration of gesture usage in presentation videos. Specifically, the exploration view enables users to obtain a quick overview of the spatial and temporal distributions of gestures. The dynamic hand movements are firstly aggregated through a heatmap in the gesture space for uncovering spatial patterns, and then decomposed into two mutually perpendicular timelines for revealing temporal patterns. The relation view allows users to explicitly explore the correlation between speech content and gestures by enabling linked analysis and intuitive glyph designs. The video view and dynamic view show the context and overall dynamic movement of the selected gestures, respectively. Two usage scenarios and expert interviews with professional presentation coaches demonstrate the effectiveness and usefulness of GestureLens in facilitating gesture exploration and analysis of presentation videos. Haipeng Zeng, Xingbo Wang 0001, Yong Wang 0021, Aoyu Wu, Ting-Chuen Pong, Huamin Qu |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2022 | XR-LIVE: Enhancing Asynchronous Shared-Space Demonstrations with Spatial-temporal Assistive Toolsets for Effective Learning in Immersive Virtual LaboratoriesabstractAn immersive virtual laboratory (VL) could offer flexibility of time and space, as well as safety, for remote students to conduct laboratory activities through online experiential learning. Recording an instructor's demonstration inside a VL is an approach that allows students to learn directly from a demonstration. However, students have to learn from a recording while controlling the playback, which requires attention spent on additional spatial and temporal cues. This additional cognitive load could lead to errors during the laboratory procedure. To address these challenges, we have identified four design requirements to reduce attention load in VLs; namely, organized learning steps, improved student sense of co-presence, reduction of task-instructor split-attention, and learning independent of interpersonal distance. Based on these requirements, we have designed and implemented spatial-temporal assistive toolsets for laboratories in a virtual environment, namely XR-LIVE, to reduce cognitive load and enhance learning in an asynchronous shared-space demonstration, implemented based on the setup of a standard civil engineering laboratory. We also analyzed students' behavior in the VL demonstration to design guidelines applicable to generic VLs. Santawat Thanyadit, Parinya Punpongsanon, Thammathip Piumsomboon, Ting-Chuen Pong |
Proc. ACM Hum. Comput. Interact. | 4 |
| 2021 | EmotionCues: Emotion-Oriented Visual Summarization of Classroom VideosabstractAnalyzing students' emotions from classroom videos can help both teachers and parents quickly know the engagement of students in class. The availability of high-definition cameras creates opportunities to record class scenes. However, watching videos is time-consuming, and it is challenging to gain a quick overview of the emotion distribution and find abnormal emotions. In this article, we propose EmotionCues, a visual analytics system to easily analyze classroom videos from the perspective of emotion summary and detailed analysis, which integrates emotion recognition algorithms with visualizations. It consists of three coordinated views: a summary view depicting the overall emotions and their dynamic evolution, a character view presenting the detailed emotion status of an individual, and a video view enhancing the video analysis with further details. Considering the possible inaccuracy of emotion recognition, we also explore several factors affecting the emotion analysis, such as face size and occlusion. They provide hints for inferring the possible inaccuracy and the corresponding reasons. Two use cases and interviews with end users and domain experts are conducted to show that the proposed system could be useful and effective for analyzing emotions in the classroom videos. Haipeng Zeng, Xinhuan Shu, Yanbang Wang, Yong Wang 0021, Liguo Zhang 0002, Ting-Chuen Pong, Huamin Qu |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2020 | DFSeer: A Visual Analytics Approach to Facilitate Model Selection for Demand ForecastingabstractSelecting an appropriate model to forecast product demand is critical to the manufacturing industry. However, due to the data complexity, market uncertainty and users' demanding requirements for the model, it is challenging for demand analysts to select a proper model. Although existing model selection methods can reduce the manual burden to some extent, they often fail to present model performance details on individual products and reveal the potential risk of the selected model. This paper presents DFSeer, an interactive visualization system to conduct reliable model selection for demand forecasting based on the products with similar historical demand. It supports model comparison and selection with different levels of details. Besides, it shows the difference in model performance on similar products to reveal the risk of model selection and increase users' confidence in choosing a forecasting model. Two case studies and interviews with domain experts demonstrate the effectiveness and usability of DFSeer. Dong Sun 0001, Zezheng Feng, Yuanzhe Chen, Yong Wang 0021, Mingxuan Yuan, Ting-Chuen Pong, Huamin Qu |
CHI | 7 |
| 2020 | ViSeq: Visual Analytics of Learning Sequence in Massive Open Online CoursesabstractThe research on massive open online courses (MOOCs) data analytics has mushroomed recently because of the rapid development of MOOCs. The MOOC data not only contains learner profiles and learning outcomes, but also sequential information about when and which type of learning activities each learner performs, such as reviewing a lecture video before undertaking an assignment. Learning sequence analytics could help understand the correlations between learning sequences and performances, which further characterize different learner groups. However, few works have explored the sequence of learning activities, which have mostly been considered aggregated events. A visual analytics system called ViSeq is introduced to resolve the loss of sequential information, to visualize the learning sequence of different learner groups, and to help better understand the reasons behind the learning behaviors. The system facilitates users in exploring learning sequences from multiple levels of granularity. ViSeq incorporates four linked views: the projection view to identify learner groups, the pattern view to exhibit overall sequential patterns within a selected group, the sequence view to illustrate the transitions between consecutive events, and the individual view with an augmented sequence chain to compare selected personal learning sequences. Case studies and expert interviews were conducted to evaluate the system. Qing Chen 0001, Xuanwu Yue, Xavier Plantaz, Yuanzhe Chen, Conglei Shi, Ting-Chuen Pong, Huamin Qu |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2020 | PlanningVis: A Visual Analytics Approach to Production Planning in Smart FactoriesabstractProduction planning in the manufacturing industry is crucial for fully utilizing factory resources (e.g., machines, raw materials and workers) and reducing costs. With the advent of industry 4.0, plenty of data recording the status of factory resources have been collected and further involved in production planning, which brings an unprecedented opportunity to understand, evaluate and adjust complex production plans through a data-driven approach. However, developing a systematic analytics approach for production planning is challenging due to the large volume of production data, the complex dependency between products, and unexpected changes in the market and the plant. Previous studies only provide summarized results and fail to show details for comparative analysis of production plans. Besides, the rapid adjustment to the plan in the case of an unanticipated incident is also not supported. In this paper, we propose PlanningVis, a visual analytics system to support the exploration and comparison of production plans with three levels of details: a plan overview presenting the overall difference between plans, a product view visualizing various properties of individual products, and a production detail view displaying the product dependency and the daily production details in related factories. By integrating an automatic planning algorithm with interactive visual explorations, PlanningVis can facilitate the efficient optimization of daily production planning as well as support a quick response to unanticipated incidents in manufacturing. Two case studies with real-world data and carefully designed interviews with domain experts demonstrate the effectiveness and usability of PlanningVis. Dong Sun 0001, Renfei Huang, Yuanzhe Chen, Yong Wang 0021, Mingxuan Yuan, Ting-Chuen Pong, Huamin Qu |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2019 | Designing Narrative Slideshows for Learning AnalyticsabstractThe practical power of data visualization is currently attracting much attention in the e-learning domain. A growing number of studies have been conducted in recent years to help instructors better analyze learner behavior and reflect on their teaching. However, current e-learning dashboards and visualization systems usually require a lot of time and effort into the exploration process. Moreover, the lack of communication power of existing systems constrains users from organizing the narrative of information pieces into a compelling data story. In this paper, we have proposed a narrative visualization approach with an interactive slideshow that helps instructors and education experts explore potential learning patterns and convey data stories. This approach contains three key components: guided-tour concept, drill-down path, and dig-in exploration dimension. The use cases further demonstrate the potential of employing this visual narrative approach in the e-learning context. Qing Chen 0001, Zhen Li 0044, Ting-Chuen Pong, Huamin Qu |
PacificVis | 3 |
| 2019 | ObserVAR: Visualization System for Observing Virtual Reality Users using Augmented RealityabstractWhile virtual reality (VR) tools provide an immersive learning experience for students, it is difficult for an instructor to observe the students' learning activities in a virtual environment (VE). Thus, it hinders interactions that could occur between the instructor and students, which are usually required in a classroom environment to understand how each student learns. Previous work has added virtual awareness cues that can help a small group of students to collaborate in a VE. However, when the number of students increases, such virtual awareness cues can cause visual clutter and confuse the instructor. We propose ObserVAR, a visualization system that allows the instructor to observe students in a VE at scale. ObserVAR uses augmented reality techniques to visualize each student's gaze in a VE and improves the instructor's awareness of the entire class. The visualizations are then optimized to reduce visual clutter in the scene using a force-directed graph drawing algorithm. In designing ObserVAR, we first investigated visualizations that can provide the instructor with an overall awareness of the VE that can be scaled up as the number of users increases. Second, we optimized the visualization of students by leveraging a graph drawing algorithm to reduce the visual clutter in the class scene. We compared the performance of our prototype with some commercially available user interfaces for VE classrooms. In our study, ObserVAR has demonstrated improvement and flexibility in several application scenarios. Santawat Thanyadit, Parinya Punpongsanon, Ting-Chuen Pong |
ISMAR | 3 |
| 2019 | Effective Feature Learning with Unsupervised Learning for Improving the Predictive Models in Massive Open Online CoursesabstractThe effectiveness of learning in massive open online courses (MOOCs) can be significantly enhanced by introducing personalized intervention schemes which rely on building predictive models of student learning behaviors such as some engagement or performance indicators. A major challenge that has to be addressed when building such models is to design handcrafted features that are effective for the prediction task at hand. In this paper, we make the first attempt to solve the feature learning problem by taking the unsupervised learning approach to learn a compact representation of the raw features with a large degree of redundancy. Specifically, in order to capture the underlying learning patterns in the content domain and the temporal nature of the clickstream data, we train a modified auto-encoder (AE) combined with the long short-term memory (LSTM) network to obtain a fixed-length embedding for each input sequence. When compared with the original features, the new features that correspond to the embedding obtained by the modified LSTM-AE are not only more parsimonious but also more discriminative for our prediction task. Using simple supervised learning models, the learned features can improve the prediction accuracy by up to 17% compared with the supervised neural networks and reduce overfitting to the dominant low-performing group of students, specifically in the task of predicting students' performance. Our approach is generic in the sense that it is not restricted to a specific supervised learning model nor a specific prediction task for MOOC learning analytics. Mucong Ding, Dit-Yan Yeung, Ting-Chuen Pong |
LAK | 4 |
| 2019 | Investigating Visualization Techniques for Observing a Group of Virtual Reality Users Using Augmented RealityabstractAs an emerging technology, virtual reality (VR) has been used increasingly as a learning tool to explore “outside the classroom experiences” inside the classroom. While VR provides an immersive experience to the students, it is difficult for the instructor to monitor the students' activities in the VR. Thus, it hinders interactions between the instructor and students. To solve this challenge, we investigated a technique that allows the instructor to observe VR users at scale using Augmented Reality. Augmented Reality techniques are used to visualize the gazes of the VR users in the virtual environment, and improve the instructor's awareness. Santawat Thanyadit, Parinya Punpongsanon, Ting-Chuen Pong |
VR | 3 |
| 2018 | Efficient Information Sharing Techniques between Workers of Heterogeneous Tasks in 3D CVEabstractCollaboration between a helper and a worker in a 3D collaborative virtual environment usually requires real-time information sharing, since the worker relies on the timely assistance from the helper. In contrast, collaboration between workers requires them to shift their attention between independent tasks and dependent tasks. In worker-worker collaborations, a real-time updating technique could create excess information, which may be a distraction. In this paper, we compare different information sharing techniques and determine an efficient technique for the collaboration between workers. In our user experiment, participants performed a floor plan design task in a designer and engineer pairing on a desktop VR environment. The results showed that the proposed information sharing technique, in which objects are updated based on local users' actions, is more suitable than real-time updates. In addition, we discuss design implications that can be applied to different collaborative scenarios. Santawat Thanyadit, Parinya Punpongsanon, Ting-Chuen Pong |
Proc. ACM Hum. Comput. Interact. | 3 |
| 2017 | Desktop VR using a Mirror Metaphor for Natural User InterfaceabstractThe main objective of this research work is to create a desktop VR environment which enables users to interact naturally with virtual objects positioned both in front and behind the screen. We propose a mirror metaphor that simulates a physical stereoscopic screen with the properties of a mirror. In addition to allowing users to interact with virtual objects positioned in front of the stereoscopic screen using their virtual hands, the virtual hands can be transferred inside the virtual mirror to interact with objects behind the screen. When the virtual hands are operating inside the virtual mirror, they are transformed like the reflection in a real mirror. This effectively doubles the interactable space and creates an interactive space that could facilitate collaborative tasks. Our user study shows that users could interact through the mirror approach as effectively as similar interaction techniques, hence demonstrating that the mirror technique is a viable interface in certain VR setups. Santawat Thanyadit, Ting-Chuen Pong |
ISS | 2 |
| 2008 | Structuring low-quality videotaped lectures for cross-reference browsing by video text analysis
Feng Wang 0036, Chong-Wah Ngo, Ting-Chuen Pong |
Pattern Recognit. | 3 |
| 2008 | Simulating a Smartboard by Real-Time Gesture Detection in Lecture VideosabstractGesture plays an important role for recognizing lecture activities in video content analysis. In this paper, we propose a real-time gesture detection algorithm by integrating cues from visual, speech and electronic slides. In contrast to the conventional ldquocomplete gesturerdquo recognition, we emphasize detection by the prediction from ldquoincomplete gesturerdquo. Specifically, intentional gestures are predicted by the modified hidden Markov model (HMM) which can recognize incomplete gestures before the whole gesture paths are observed. The multimodal correspondence between speech and gesture is exploited to increase the accuracy and responsiveness of gesture detection. In lecture presentation, this algorithm enables the on-the-fly editing of lecture slides by simulating appropriate camera motion to highlight the intention and flow of lecturing. We develop a real-time application, namely simulated smartboard, and demonstrate the feasibility of our prediction algorithm using hand gesture and laser pen with simple setup without involving expensive hardware. Feng Wang 0036, Chong-Wah Ngo, Ting-Chuen Pong |
IEEE Trans. Multim. | 3 |
| 2007 | Lecture Video Enhancement and Editing by Integrating Posture, Gesture, and TextabstractThis paper describes a novel framework for automatic lecture video editing by gesture, posture, and video text recognition. In content analysis, the trajectory of hand movement is tracked and the intentional gestures are automatically extracted for recognition. In addition, head pose is estimated through overcoming the difficulties due to the complex lighting conditions in classrooms. The aim of recognition is to characterize the flow of lecturing with a series of regional focuses depicted by human postures and gestures. The regions of interest (ROIs) in videos are semantically structured with text recognition and the aid of external documents. By tracing the flow of lecturing, a finite state machine (FSM) which incorporates the gestures, postures, ROIs, general editing rules and constraints, is proposed to edit videos with novel views. The FSM is designed to generate appropriate simulated camera motion and cutting effects that suit the pace of a presenter's gestures and postures. To remedy the undesirable visual effects due to poor lighting conditions, we also propose approaches to automatically enhance the visibility and readability of slides and whiteboard images in the edited videos Feng Wang 0036, Chong-Wah Ngo, Ting-Chuen Pong |
IEEE Trans. Multim. | 3 |
| 2006 | A Multi-Layer MRF Model for Video Object Segmentation
Zoltan Kato, Ting-Chuen Pong |
ACCV (2) | 2 |
| 2006 | Prediction-Based Gesture Detection in Lecture Videos by Combining Visual, Speech and Electronic SlidesabstractThis paper presents an efficient algorithm for gesture detection in lecture videos by combining visual, speech and electronic slides. Besides accuracy, response time is also considered to cope with the efficiency requirements of real-time applications. Candidate gestures are first detected by visual cue. Then we modify HMM models for complete gestures to predict and recognize incomplete gestures before the whole gestures paths are observed. Gesture recognition is used to verify the results of gesture detection. The relations between visual, speech and slides are analyzed. The correspondence between speech and gesture is employed to improve the accuracy and the responsiveness of gesture detection Feng Wang 0036, Chong-Wah Ngo, Ting-Chuen Pong |
ICME | 3 |
| 2006 | A Markov random field image segmentation model for color textured images
Zoltan Kato, Ting-Chuen Pong |
Image Vis. Comput. | 2 |
| 2005 | Exploiting self-adaptive posture-based focus estimation for lecture video editingabstractHead pose plays a special role in estimating a presenter's focuses and actions for lecture video editing. This paper presents an efficient and robust head pose estimation algorithm to cope with the new challenges arising in the content management of lecture videos. These challenges include speed requirement, low video quality, variant presenting styles and complex settings in modern classrooms. Our algorithm is based on a robust hierarchical representation of skin color clustering and a set of pose templates that are automatically trained. Contextual information is also considered to refine pose estimation. Most importantly, we propose an online learning approach to deal with different presenting styles, which has not been addressed before. We show that the proposed approach can significantly improve the performance of pose estimation. In addition, we also describe how posture is used in focus estimation for lecture video editing by integrating with gesture. Feng Wang 0036, Chong-Wah Ngo, Ting-Chuen Pong |
ACM Multimedia | 3 |
| 2003 | Unsupervised segmentation of color textured images using a multilayer MRF modelabstractHerein, we propose a novel multilayer Markov random field (MRF) image segmentation model which aims at combining color and texture features: each feature is associated to a so called feature layer, where an MRF model is defined using only the corresponding feature. A special layer is assigned to the combined MRF model. This layer interacts with each feature layer and provides the segmentation based on the combination of different features. The model is quite generic and isn't restricted to a particular texture feature. Herein we will test the algorithm using Gabor and MRSAR texture features. Furthermore, the algorithm automatically estimates the number of classes at each layer (there can be different classes at different layers) and the associated model parameters. Zoltan Kato, Ting-Chuen Pong, Song Guo Qiang |
ICIP (1) | 2 |
| 2003 | Synchronization of lecture videos and electronic slides by video text analysisabstractAn essential goal of structuring lecture videos captured in live presentation is to provide a synchronized view of video clips and electronic slides. This paper presents an automatic approach to match video clips and slides based on the analysis of text embedded in lecture videos. We describe a method to reconstruct high-resolution video texts from multiple keyframes for robust OCR recognition. A two-stage matching algorithm based on the title and content similarity measures between video clips and slides is also proposed. Feng Wang 0036, Chong-Wah Ngo, Ting-Chuen Pong |
ACM Multimedia | 3 |
| 2003 | Motion analysis and segmentation through spatio-temporal slices processingabstractThis paper presents new approaches in characterizing and segmenting the content of video. These approaches are developed based upon the pattern analysis of spatio-temporal slices. While traditional approaches to motion sequence analysis tend to formulate computational methodologies on two or three adjacent frames, spatio-temporal slices provide rich visual patterns along a larger temporal scale. We first describe a motion computation method based on a structure tensor formulation. This method encodes visual patterns of spatio-temporal slices in a tensor histogram, on one hand, characterizing the temporal changes of motion over time, on the other hand, describing the motion trajectories of different moving objects. By analyzing the tensor histogram of an image sequence, we can temporally segment the sequence into several motion coherent subunits, in addition, spatially segment the sequence into various motion layers. The temporal segmentation of image sequences expeditiously facilitates the motion annotation and content representation of a video, while the spatial decomposition of image sequences leads to a prominent way of reconstructing background panoramic images and computing foreground objects. Chong-Wah Ngo, Ting-Chuen Pong, HongJiang Zhang |
IEEE Trans. Image Process. | 2 |
| 2002 | Detection of slide transition for topic indexingabstractThis paper presents an automatic and novel approach in detecting the transitions of slides for video sequences of technical lectures. Our approach adopts a foreground vs background segmentation algorithm to separate a presenter from the projected electronic slides. Once a background template is generated, text captions are detected and analyzed. The segmented caption regions as well as background templates together provide salient visual cues to decide whether a slide is flipped and replaced. The partitioning of videos according to slide changes not only structure the content of video according to topics, but also facilitate the synchronization of video, audio and electronic slides for effective indexing, browsing and retrieval. Chong-Wah Ngo, Ting-Chuen Pong, Thomas S. Huang |
ICME (2) | 2 |
| 2002 | Motion-Based Video Representation for Scene Change Detection
Chong-Wah Ngo, Ting-Chuen Pong, HongJiang Zhang |
Int. J. Comput. Vis. | 2 |
| 2002 | On clustering and retrieval of video shots through temporal slices analysisabstractBased on the analysis of temporal slices, we propose novel approaches for clustering and retrieval of video shots. Temporal slices are a set of two-dimensional (2-D) images extracted along the time dimension of an image volume. They encode rich set of visual patterns for similarity measure. In this paper, we first demonstrate that tensor histogram features extracted from temporal slices are suitable for motion retrieval. Subsequently, we integrate both tensor and color histograms for constructing a two-level hierarchical clustering structure. Each cluster in the top level contains shots with similar color while each cluster in bottom level consists of shots with similar motion. The constructed structure is then used for the cluster-based retrieval. The proposed approaches are found to be useful particularly for sports games, where motion and color are important visual cues when searching and browsing the desired video shots. Chong-Wah Ngo, Ting-Chuen Pong, HongJiang Zhang |
IEEE Trans. Multim. | 2 |
| 2001 | A Markov Random Field Image Segmentation Model Using Combined Color and Texture Features
Zoltan Kato, Ting-Chuen Pong |
CAIP | 2 |
| 2001 | On clustering and retrieval of video shotsabstractClustering of video data is an important issue in video abstraction, browsing and retrieval. In this paper, we propose a two-level hierarchical clustering approach by aggregating shots with similar motion and color features. Motion features are computed directly from 2D tensor histograms, while color features are represented by 3D color histograms. Cluster validity analysis is further applied to automatically determine the number of clusters at each level. Video retrieval can then be done directly based on the result of clustering. The proposed approach is found to be useful particularly for sports games, where motion and color are important visual cues when searching and browsing the desired video shots. Since most games involve two teams, clsssification and retrieval of teams becomes an interesting topic. To achieve these goals, nevertheless, an initial as well as critical step is to isolate team players from background regions. Thus, we also introduce approach to segment foreground objects (players) prior to classification and retrieval. Chong-Wah Ngo, Ting-Chuen Pong, HongJiang Zhang |
ACM Multimedia | 2 |
| 2001 | Exploiting image indexing techniques in DCT domain
Chong-Wah Ngo, Ting-Chuen Pong, Roland T. Chin |
Pattern Recognit. | 2 |
| 2001 | Color image segmentation and parameter estimation in a markovian framework
Zoltan Kato, Ting-Chuen Pong, John Chung-Mong Lee |
Pattern Recognit. Lett. | 2 |
| 2001 | Video partitioning by temporal slice coherencyabstractWe present a novel approach for video partitioning by detecting three essential types of camera breaks, namely cuts, wipes, and dissolves. The approach is based on the analysis of temporal slices which are extracted from the video by slicing through the sequence of video frames and collecting temporal signatures. Each of these slices contains both spatial and temporal information from which coherent regions are indicative of uninterrupted video partitions separated by camera breaks. Properties could further be extracted from the slice for both the detection and classification of camera breaks. For example, cut and wipes are detected by color-texture properties, while dissolves are detected by statistical characteristics. The approach has been tested by extensive experiments. Chong-Wah Ngo, Ting-Chuen Pong, Roland T. Chin |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2000 | Motion-Based Video Representation for Scene Change DetectionabstractWe present a scheme for automatically partitioning videos into scenes. A scene is generally referred to as a group of shots taken at the same site. We first propose a motion annotation algorithm based on the analysis of spatiotemporal image volumes. The algorithm characterizes the motions within shots by extracting and analyzing the motion trajectories encoded in the temporal slices of image volumes. A motion-based keyframe computing and selection strategy is thus proposed to compactly represent the content of shots. With these techniques, we further present a scene change detection algorithm by measuring the similarity of the representative keyframes in shots. Chong-Wah Ngo, Ting-Chuen Pong, Roland T. Chin, HongJiang Zhang |
ICPR | 2 |
| 1999 | Detection of Gradual Transitions through Temporal Slice AnalysisabstractIn this paper, we present approaches for detecting camera cuts, wipes and dissolves based on the analysis of spatio-temporal slices obtained from videos. These slices are composed of spatially and temporally coherent regions which can be perceived as shots. In the proposed methods, camera breaks are located by performing color-texture segmentation and statistical analysis on these video slices. In addition to detecting camera breaks, our methods can classify the detected breaks as camera cuts, wipes and dissolves in an efficient manner. Chong-Wah Ngo, Ting-Chuen Pong, Roland T. Chin |
CVPR | 2 |
| 1999 | An Efficient Parallel Algorithm for Computing the Gaussian Convolution of Multi-dimensional Image Data
Hoi Man Yip, Ishfaq Ahmad 0001, Ting-Chuen Pong |
J. Supercomput. | 3 |
| 1998 | Motion Compensated Color Video Classification Using Markov Random Fields
Zoltan Kato, Ting-Chuen Pong, John Chung-Mong Lee |
ACCV (1) | 2 |
| 1998 | Enhancing Object Recognition Using Recency And Co-Occurrence Huristics
John Chung-Mong Lee, Ting-Chuen Pong, Albert C. Esterline |
Pattern Recognit. | 2 |
| 1996 | Detection of moving objects using a spatiotemporal representationabstractThe problem of moving object detection using a moving observer is addressed. The problem is investigated for long image sequences. The frames of an image sequence are stacked together to form a 3D spatiotemporal cube. A set of 2-D temporal slices which are termed epipolar plane images (EPIs) is then extracted from the cube along specific direction determined by a tracking process. For static scenes, each EPI is a structured image in which the trajectories are constrained by mathematical functions of time. We show that the problem of moving object detection is equivalent to the inspection of violation of such constraints. Experimental results on real sequences reveal that this approach is robust and reliable. Moreover, the algorithm is capable of detecting multiple moving objects in the scene. Hoi Man Yip, Ting-Chuen Pong |
ICPR | 2 |
| 1996 | Cooperative fusion of stereo and motion
Anthony Yuk-Kwan Ho, Ting-Chuen Pong |
Pattern Recognit. | 2 |
| 1995 | Building Surfaces from Three-Dimensional Image Data on the Intel Paragon
Hoi Man Yip, Sai Chung Yeung, Ishfaq Ahmad 0001, Ting-Chuen Pong |
ICPP (3) | 4 |
| 1995 | A note on parsing pattern languages
Oscar H. Ibarra, Ting-Chuen Pong, Stephen M. Sohn |
Pattern Recognit. Lett. | 2 |
| 1994 | An experimental study of an object recognition system that learns
John Chung-Mong Lee, Ting-Chuen Pong, James R. Slagle, Albert C. Esterline |
Pattern Recognit. | 2 |
| 1993 | Multi-layer surface segmentation using energy minimizationabstractA multilayer, depth planes approach to image segmentation is proposed, where pixels may have arisen from a single smooth surface in the scene are represented in a common layer. Two types of output are produced at each pixel, i.e., a layer number and a vector of depth values, one value for each layer. The layer assignment performs image partitioning based on surface properties. The depth value assignment for each layer either represents the input data with the noise removed or interpolates between data values to fill-in nonvisible parts of the scene. The disjoint surfaces due to occlusion or transparency are also grouped together if they form a smooth surface.> Suthep Madarasmi, Daniel J. Kersten, Ting-Chuen Pong |
CVPR | 3 |
| 1992 | The Computation of Stereo Disparity for Transparent and for Opaque Surfaces
Suthep Madarasmi, Daniel J. Kersten, Ting-Chuen Pong |
NIPS | 3 |
| 1991 | Parallel Regognition and Parsing on the HypercubeabstractThe authors present parallel algorithms for recognizing and parsing context-free languages on the hypercube. This algorithm is both time-wise and space-wise optimal with respect to the usual sequential dynamic programming algorithm. Also, the number of nonoverlapping interprocessor data transmissions for the recognition phase is small. It is noted that this is desirable since communication cost in reality is a function of the number of transmissions as well as transmission length. The authors present another recognition algorithm that achieves the same time and space bounds but employs a dynamic loading balancing technique to increase processor utilization. The results of implementing these algorithms on a 64-node NCUBE/7 MIMD hypercube machine are also given. The experimental evidence indicates that, while both recognition algorithms exhibit acceptable speedups, using load balancing results in significantly better performance. The authors obtain parallel algorithms with the same time and space bounds as above for the polygon triangulation problem and the matrix product chain problem.> Oscar H. Ibarra, Ting-Chuen Pong, Stephen M. Sohn |
IEEE Trans. Computers | 2 |
| 1990 | Detecting moving objects
William B. Thompson, Ting-Chuen Pong |
Int. J. Comput. Vis. | 2 |
| 1990 | A Knowledge-Based System for the Image Correspondence ProblemabstractThe image correspondence problem has generally been considered the most difficult step in both stereo and temporal vision. Most existing approaches match area features or linear features extracted from an image pair. The approach described in this paper is novel in that it uses an expert system shell to develop an image correspondence knowledge-based system for the general image correspondence problem. The knowledge it uses consists of both physical properties and spatial relationships of the edges and regions in images for every edge or region matching. A computation network is used to represent this knowledge. It allows the computation of the likelihood of matching two edges or regions with logical and heuristic operators. Heuristics for determining the correspondences between image features and the problem of handling missing information will be discussed. The values of the individual matching results are used to direct the traversal and pruning of the global matching process. The problem of parallelizing the entire process will be discussed. Experimental results on real-world images show that all matching edges and regions have been identified correctly. John Chung-Mong Lee, Ting-Chuen Pong, James R. Slagle |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 1990 | Motion analysis of long image sequence flow
Martin Kenner, Ting-Chuen Pong |
Pattern Recognit. Lett. | 2 |
| 1989 | Shape from shading using the facet model
Ting-Chuen Pong, Robert M. Haralick, Linda G. Shapiro |
Pattern Recognit. | 1 |
| 1989 | Matching topographic structures in stereo vision
Ting-Chuen Pong, Robert M. Haralick, Linda G. Shapiro |
Pattern Recognit. Lett. | 1 |
| 1989 | A hierarchical approach to the correspondence problemabstractAn algorithm for feature-point matching using a feature-point hierarchy (edges directed by corners) is presented. Psychophysical evidence indicates that higher-level image structures, such as line segments or corners, can aid the matching of low-level structures. This observation, applied to the correspondence problem in machine vision, directly motivates the implementation of the hierarchical matcher described. In addition, edge and corner detectors developed for the matcher and based on the facet model are described. The matcher gives an accurate disparity assignment for real images and performs better than the simpler nonhierarchical relaxation matcher on which it is based.> Ting-Chuen Pong, B. G. Kaiser |
IEEE Trans. Syst. Man Cybern. | 1 |
| 1988 | Two-Dimensional Convolution on a Pyramid ComputerabstractAn algorithm for convolving a k*k window of weighting coefficients with an n*n image matrix on a pyramid computer of O(n/sup 2/) processors in time O(logn+k/sup 2/), excluding the time to load the image matrix, is presented. If k= Omega ( square root log n), which is typical in practice, the algorithm has a processor-time product O(n/sup 2/ k/sup 2/) which is optimal with respect to the usual sequential algorithm. A feature of the algorithm is that the mechanism for controlling the transmission and distribution of data in each processor is finite state, independent of the values of n and k. Thus, for convolving two (0, 1)-valued matrices using Boolean operations rather than the typical sum and product operations, the processors of the pyramid computer are finite-state.> Jik H. Chang, Oscar H. Ibarra, Ting-Chuen Pong, Stephen M. Sohn |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 1987 | Two-Dimensional Convolution on a Pyramid Computer
Jik H. Chang, Oscar H. Ibarra, Ting-Chuen Pong, Stephen M. Sohn |
ICPP | 3 |
| 1985 | Shape estimation from topographic primal sketch
Ting-Chuen Pong, Linda G. Shapiro, Robert M. Haralick |
Pattern Recognit. | 1 |
| 1984 | Experiments in segmentation using a facet model region grower
Ting-Chuen Pong, Linda G. Shapiro, Layne T. Watson, Robert M. Haralick |
Comput. Vis. Graph. Image Process. | 1 |
| 1983 | Texture analysis of aerial photographs
Ronald Lumia, Robert M. Haralick, Oscar A. Zuniga, Linda G. Shapiro, Ting-Chuen Pong, Far-Peing Wang |
Pattern Recognit. | 5 |
| 1983 | The application of image analysis techniques to mineral processing
Ting-Chuen Pong, Robert M. Haralick, James R. Craig, Roe-Hoan Yoon, Woo-Zin Choi |
Pattern Recognit. Lett. | 1 |