VLDB 2026 Research / reviewers in the wild / expert
Irene Cheng 0001
dblp:c/LIreneCheng · also L. Irene Cheng
· DBLP profile ↗
85ranked-venue papers
22as first author
16since 2021 · last 2026
0000-0001-9699-4895ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 51 · 18 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 18 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 15 · 2 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 12 · 2 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dual-branch non-negative matrix factorization guided by information decoupling for multi-view clustering
Mingxia Gong, Chengcai Leng, Zhao Pei, Jinye Peng 0001, Irene Cheng 0001, Anup Basu |
Neurocomputing | 6 |
| 2025 | Block information strategy for multi-modal remote sensing image registration
Yameng Hong, Chengcai Leng, Beihua Liu, Jinye Peng 0001, Irene Cheng 0001, Anup Basu |
Eng. Appl. Artif. Intell. | 5 |
| 2025 | Orthogonal Diversity Nonnegative Matrix Factorization for multi-view clustering
Xinling Zhang, Chengcai Leng, Jinye Peng 0001, Irene Cheng 0001, Anup Basu |
Eng. Appl. Artif. Intell. | 4 |
| 2025 | MFEL-YOLO for small object detection in UAV aerial images
Ting Hou, Chengcai Leng, Zhao Pei, Jinye Peng 0001, Irene Cheng 0001, Anup Basu |
Expert Syst. Appl. | 6 |
| 2025 | Scale- and Shape-Aware Network With Prediction Decoupling for Building Fine-Grained Change DetectionabstractBuilding change detection (BCD) is a hot topic in geoscience and remote sensing (RS) with widespread applications. However, most existing BCD methods only focus on areas where changes have occurred, but ignore the change statuses. To address this problem, a building fine-grained change detection (BFCD) task is further explored in this work, which aims to judge the time-related “disappeared”, “appeared”, and “rebuilt” change types of buildings. Meanwhile, a scale- and shape-aware network (S2Net) with prediction decoupling is designed. Firstly, a prediction decoupling framework with dual decoders is built to ensure the prediction consistency with the temporal order of bi-temporal images. Secondly, considering the rebuilt type is the changes between building instances, which are often reflected in the scale and shape differences of the buildings. Thereby, a scale-aware module (ScAM) and a shape-aware module (ShAM) are designed. These two modules help extract the discriminative features of buildings with different scales and shapes for subsequent change detection (CD). In addition, two BCD datasets widely used, LEVIR-CD+ and WHU-CD, are relabeled in this work to support the study of BFCD. Experimental results show that S2Net achieves competitive performance, and its effectiveness is confirmed. The code and datasets will be publicly available at https://github.com/ptdoge/S2Net. Chengcai Leng, Xi Li 0001, Irene Cheng 0001, Anup Basu, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Demo: Blockchain Shield - Advanced Threat Detection & Forensic Analysis PlatformabstractIn the rapidly evolving landscape of blockchain technology, security emerges as a paramount concern. This paper introduces an innovative blockchain security threat awareness platform, designed to comprehensively address the multifaceted security challenges within blockchain networks, particularly focusing on Ethereum contracts. Central to the platform is a dual-database architecture, blending a NoSQL database with a graph database, enhancing data management, and enabling intricate transaction network visualizations. The platform's Threat Detection module, utilizing Large Language Models (LLMs) in conjunction with traditional methods, offers a novel approach to identifying and categorizing vulnerabilities in Ethereum smart contracts. Complementing this, the Threat Evidence Collection module provides detailed post-attack analysis, tracing transactions to their sources and evaluating address risks. This module's capabilities extend to producing statistical reports, including the transactional history and risk evaluation of individual addresses. Demonstrated on the Ethereum blockchain, the platform showcases its proficiency in handling complex data, rapid threat detection, and extensive forensic analysis, presenting a robust solution to fortifying blockchain security and offering a proactive defense mechanism for users and developers in the blockchain environment. Ningbo Zhu, Jinghan Sun, Xinyao Sun, Irene Cheng 0001 |
ICDCS | 5 |
| 2024 | Incremental semi-supervised graph learning NMF with block-diagonal
Xue Lv, Chengcai Leng, Jinye Peng 0001, Zhao Pei, Irene Cheng 0001, Anup Basu |
Eng. Appl. Artif. Intell. | 5 |
| 2024 | Dual-Graph Global and Local Concept Factorization for Data ClusteringabstractConsidering a wide range of applications of nonnegative matrix factorization (NMF), many NMF and their variants have been developed. Since previous NMF methods cannot fully describe complex inner global and local manifold structures of the data space and extract complex structural information, we propose a novel NMF method called dual-graph global and local concept factorization (DGLCF). To properly describe the inner manifold structure, DGLCF introduces the global and local structures of the data manifold and the geometric structure of the feature manifold into CF. The global manifold structure makes the model more discriminative, while the two local regularization terms simultaneously preserve the inherent geometry of data and features. Finally, we analyze convergence and the iterative update rules of DGLCF. We illustrate clustering performance by comparing it with latest algorithms on four real-world datasets. Chengcai Leng, Irene Cheng 0001, Anup Basu, Licheng Jiao |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Active Learning for Multi-Class Vehicle Categorization and Traffic Analysis in complex environmentsabstractThis paper presents a novel approach designed for the study of vehicles, with a primary focus on enhancing the assessment of goods and their value. The framework aims to improve the comprehension of vehicular traffic dynamics on municipalities, thereby enabling improved route planning and inspection strategies. Our proposed closed-loop system integrates deep learning, conventional image processing and computer vision to detect, track, count, timestamp, and estimate the direction of travel for vehicles, thus laying the groundwork for in-depth traffic flow analysis and optimization. The proposed framework incorporates a unique data processing mechanism within a crowdsourcing environment, enhancing the scalability of our system. For multiclass object detection we proposed a single stage and two-stage pipelines using YOLOv8, YOLOv6, YOLOv5 and RT-DETR-LR models. Our tracking stage computes cumulative average confidence scores per estimated class over a vehicle’s lifespan, enhancing class prediction robustness. Our method achieved 0.891 mAP score with data augmentation strategies. Experimental results demonstrate the effectiveness, efficiency, and robustness of the proposed system on challenge scenes and adaptability with active learning for vehicular analysis. Gabriel Lugo Bustillo, Joey Quinlan, Lingrui Zhou, Md Nahid Sadik, Irene Cheng 0001 |
ISM | 5 |
| 2023 | Exploring terrestrial point clouds with Google Street View for discovery and fine-grained catalog of urban objectsabstractLocalization and annotation of objects of interest can become an exhaustive task in dense point clouds without color information. Previous studies used point cloud data with aligned optical images for better scene understanding, however, are limited by huge data acquisition and system setup. Geospatial technologies are increasingly relevant to automatic land survey, but their use may be limited by cost and domain expertise. Technologies like Google Street View are appealing because they are unconfined and easy to operate on a web browser. We proposed a module to interconnect Google Street View client and The Visualization Toolkit (VTK) for terrestrial point cloud exploration. We evaluate the proposed pipeline with large-scale terrestrial laser scanned point-clouds with over a billion points in complex outdoor scenes. The integration of Google Street View and point cloud data provide better scene understanding and efficiency point cloud annotation than annotating point cloud without color information. Our results demonstrate that interactive modules using Street View may be suitable for self-referencing a surveyor within the point cloud and assist survey tasks more efficiently than software-based and field surveys. Gabriel Lugo Bustillo, Rutvik Chauhan, Irene Cheng 0001 |
ISM | 3 |
| 2023 | Visible-Infrared Features Fusion Based Object DetectionabstractFusion techniques are frequently utilized in the realm of multimodal object detection tasks. While many current studies showcase their proficiency in generating visually pleasing fused images, only a limited number of them focused on the object detection performance. This study addresses the issue by presenting an end-to-end framework for object detection through the fusion of visible and infrared features (VIFF). Specifically, our approach involves the use of two distinct processing units that independently extract features from visible and infrared images, followed by the fusion of these features using a novel fusion strategy. While the visible feature processing unit preserves the direction of the gradient of visible images, the infrared feature processing unit focuses on extracting the contrast and semantic features of infrared images. Both features are aggregated by attention mechanisms and then fed into the backbone of the object detection networks. Our fusion network achieved superior object detection accuracy compared to existing state-of-the-art approaches on various datasets. We have also demonstrated that the proposed visible feature and infrared feature processing units are capable of enhancing the performance of various object detection models. Irene Cheng 0001 |
SMC | 2 |
| 2022 | LiSurveying: A high-resolution TLS-LiDAR benchmarkabstractOutdoor point-cloud object localization is an essential processing step for urban scene analysis and modeling in numerous applications, especially in land surveying and site analysis. Given the increasingly use of Terrestrial Laser Scanning (TLS) and LiDAR 3D data acquisition, numerous annotated point-cloud datasets are available and can be used to evaluate computer vision and machine learning based algorithms. Nevertheless, current point-cloud datasets mainly focus on object detection and classification in autonomous driving or urban planning type of applications, and share redundant object classes, e.g., trees, vehicles and pedestrians, which limit their usefulness for land surveying and site analysis. This paper introduces a novel 3D benchmark dataset LiSurveying, which is a large-scale point-cloud dataset with over a billion points and uncommon urban object categories in complex outdoor environments. Our dataset incorporates more urban object classes than existing datasets. Its instances have diverse point densities, shapes and dimensions, which also impose a challenge for point-cloud detection and classification algorithms. We conducted baselines experiments for point-cloud classification using machine learning classifiers and deep learning methods on different subsets of the LiSurveying dataset, and we are able to demonstrate that the various types of object classes, number of instances per class, distribution of object points, and variety of complex scenes, make this LiSurveying benchmark dataset suitable for evaluating 3D point-cloud classification, semantic segmentation, and object detection algorithms. Gabriel Lugo Bustillo, Yunwei Li 0001, Rutvik Chauhan, Palak Tiwary, Utkarsh Pandey, Archi Patel, Steve Rombough, Rod Schatz, Irene Cheng 0001 |
Comput. Graph. | 10 |
| 2022 | Multi-step implicit Adams predictor-corrector network for fire detectionabstractAbstract Fire detection methods based on the Convolutional Neural Networks (CNN) have advantages of high accuracy, wide coverage and robustness, receiving significant attention from researchers. Among CNN‐based methods, ResNet has achieved better performance than other CNN frameworks in fire detection system, since it uses stacked residual blocks to enlarge the receptive field to overcome the vanishing gradient problem with residual learning. The merits of ResNet can be attributed to the similarity between ResNet and the single‐step explicit solver for Ordinary Differential Equations (ODEs), for example, the Euler method. Motivated by the theory of numerical ODE that a multi‐step implicit solver has higher accuracy than a single‐step explicit solver, the Multi‐step Implicit Adams predictor‐corrector (MIAPC) network for fire detection is proposed. The MIAPC method is first mapped to a corresponding predictor‐corrector Adams block which achieves higher accuracy than a single‐step explicit solver. Then, Adaptive Feature Fusion (AFF) and the Spatial Attention Layer (SAL) are utilized to extract hierarchical features from stacked predictor‐corrector Adams blocks, forming the corresponding Adams module. Finally, the 4 Adams modules which are made of 4, 6, 8, 10 predictor‐corrector Adams blocks and followed by AFF and SAL form the crucial ODE‐based approximation part in the proposed network. By adding a simple feature extraction and detection in front of and after the ODE‐based approximation part, the MIAPC network is built. Experiments demonstrate that the method achieves 87% accuracy in the challenging test dataset, outperforming existing methods by at least 6%. Besides, the 5.3M model size with inference speed of 4.7 frames/second in CPU and 65.7 frames/second in GPU enables the proposed method to be used in practical applications. Zhen Deng, Shuhao Hu, Shibai Yin, Yibin Wang 0001, Anup Basu, Irene Cheng 0001 |
IET Image Process. | 6 |
| 2022 | Find Small Objects in UAV Images by Feature Mining and AttentionabstractWith the increasing popularity of Unmanned Aerial Vehicles (UAVs), the accuracy of detecting small objects in large-view images is also expected to increase. However, accurate small object detection is still a challenging problem. Currently, Image Pyramid Network, Feature Pyramid Network (FPN), rich training strategies and data augmentation are widely used to address this problem. To accurately detect small objects, the most important thing is to mine for more feature information. We propose Widened Residual Block (WRB) to break through the bottleneck of residual information gain to extract more feature information. The second is to emphasize or suppress features to prevent small objects from being overwhelmed by a broad background. We introduce an attention mechanism into PANet and propose Enhanced Attention PANet (EA-PANet), which consists of two parts: Context Attention Module (COAM) and Attention Enhancement Module (AEM). COAM outputs attention heatmaps with context, and AEM fuses features from the channel attention module (CAM) and COAM to avoid distraction from a vast background. In addition, we design a lightweight Decoupled Attention Head (DA-head) to dynamically compute important regions for specific tasks and achieve reliable predictions. Experiments show that our method outperforms state-of-the-art (SOTA) detectors. The source code for this work is available at https://github.com/liuxiaolei111/FindSmallObjects. Chengcai Leng, Xiaoming Niu, Zhao Pei, Irene Cheng 0001, Anup Basu |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2021 | A Real-Time Mobile Application for Cattle Tracking using Video Captured from a DroneabstractIn many countries, governments have issued laws on cattle traceability, which require farmers to keep track of their livestock. While there are many algorithms to monitor cattle in confined indoor environments, monitoring cattle of a large number in outdoor pastures is still an open research problem. In recent years, with the advancement of unmanned aerial vehicles, e.g., drones, cattle monitoring can be based on images captured by drones. However, the challenges include image analysis in real-time, and keeping track of a dynamic scene (cattle movements) based on a moving sensor device. In this work, we develop an iOS application that can send captured herd images to our deep-learning-based server for cow segmentation and counting. Furthermore, the app can guide the operator to control the drone flight route and viewing perspectives. Chuyang Liu, Zihao Jian, Minshan Xie, Irene Cheng 0001 |
ISNCC | 4 |
| 2021 | An Unsupervised Generative Neural Approach for InSAR Phase Filtering and Coherence EstimationabstractPhase filtering and pixel quality (coherence) estimation is critical in producing digital elevation models (DEMs) from interferometric synthetic aperture radar (InSAR) images, as it removes spatial inconsistencies (residues) and immensely improves the subsequent unwrapping. Large amount of InSAR data facilitates wide area monitoring (WAM) over geographical regions. Advances in parallel computing have accelerated convolutional neural networks (CNNs), giving them advantages over human performance on visual pattern recognition, which makes CNNs a good choice for WAM. Nevertheless, this research is largely unexplored. We thus propose “GenInSAR,” a CNN-based generative model for joint phase filtering and coherence estimation that directly learns the InSAR data distribution. GenInSAR’s unsupervised training on satellite and simulated noisy InSAR images outperforms other five related methods in total residue reduction (over 16(1/2)% better on average) with less over-smoothing/artifacts around branch cuts. GenInSAR’s phase and coherence root-mean-squared-error and phase cosine error have average improvements of 0.54, 0.07, and 0.05, respectively compared to the related methods. Subhayan Mukherjee, Aaron Zimmer, Xinyao Sun, Parwant Ghuman, Irene Cheng 0001 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2020 | Parkinson's Disease Detection Using Ensemble Architecture from MR Images*abstractParkinson's Disease(PD) is one of the major nervous system disorders that affect people over 60. PD can cause cognitive impairments. In this work, we explore various approaches to identify Parkinson's using Magnetic Resonance (MR) T1 images of the brain. We experiment with ensemble architectures combining some winning Convolutional Neural Network models of ImageNet Large Scale Visual Recognition Challenge (ILSVRC) and propose two architectures. We find that detection accuracy increases drastically when we focus on the Gray Matter (GM) and White Matter (WM) regions from the MR images instead of using whole MR images. We achieved an average accuracy of 94.7% using smoothed GM and WM extracts and one of our proposed architectures. We also perform occlusion analysis and determine which brain areas are relevant in the architecture decision making process. Tahjid Ashfaque Mostafa, Irene Cheng 0001 |
BIBE | 2 |
| 2019 | Parkinson's Disease Mid-Brain Assessment using MR T2 ImagesabstractThe reduction of dopamine generating neurons in the brain regions known as substantia nigra (SN) is the reason for Parkinson's Disease (PD). To detect such symptom, for each subject, our algorithm only needs to analyze 3 slices around the center of a MRI DICOM volume, i.e., mid-brain area. In each slice, a window covering the SN becomes the region of interest (ROI) for further analysis. The ROIs are pre-processed by denoising and removing intensity non-uniformity. Local Binary Pattern (LBP) and Histogram Oriented Gradient (HOG) are used for feature extraction. Random Forest (RF) and Support Vector Machine (SVM) are used as classifiers with Principle Component Analysis (PCA) as feature reduction method. For evaluation, we use MRI T2 scans from the Parkinson's Progression Markers Initiative (PPMI) data set. We conducted experiments to illustrate the different classification capabilities of LBP, HOG and the fusion of these features for PD prognosis. Analysis shows that the SVM classifier with fusion feature descriptors has the most accurate classification outcome for PD assessment. Sara Soltaninejad, Pengda Xu, Irene Cheng 0001 |
BIBE | 3 |
| 2018 | Body Movement Monitoring for Parkinson's Disease Patients Using A Smart Sensor Based Non-Invasive TechniqueabstractThere have been increasing interests in recent years on using smart sensor technology, e.g., Kinect and Leap Motion, to capture and analyze human body movements, with the goal to benefit not only games, but also health care and rehab applications. We propose a non-invasive approach using movement data captured from Kinect to monitor motor deficits of Parkinson's disease (PD) patients. We captured and evaluated simple exercises, normally performed in rehabilitation sessions by physical therapist: Stride Length, Tremor and Timed Up & Go (TUG). The standard medical UPDRS scale is used by a physical therapist to determine the level of severity as the ground truth. The general framework after getting the motion data includes two steps feature extraction from the kinematic motion data, and classification using random forest (RF) (for the stride length and tremor data) and K-means (for the TUG data). Our technique was validated by inviting a group of subjects whose kinematic data are used for PD motion analysis. The experimental results demonstrate the high accuracy of our approach in the assessment of PD using kinematic motion data. Our technique is also suitable in a remote monitoring environment, where data collected can be transmitted to experts for assessment. Sara Soltaninejad, Andres Rosales-Castellanos, Fang Ba, Mario Alberto Ibarra-Manzano, Irene Cheng 0001 |
HealthCom | 5 |
| 2018 | Learning Local Distortion Visibility from Image QualityabstractAccurate prediction of local distortion visibility thresholds is critical in many image and video processing applications. Existing methods require an accurate modeling of the human visual system, and are derived through pshycophysical experiments with simple, artificial stimuli. These approaches, however, are difficult to generalize to natural images with complex types of distortion. In this paper, we explore a different perspective, and we investigate whether it is possible to learn local distortion visibility from image quality scores. We propose a convolutional neural network based optimization framework to infer local detection thresholds in a distorted image. Our model is trained on multiple quality datasets, and the results are correlated with empirical visibility thresholds collected on complex stimuli in a recent study. Our results are comparable to state-of-the-art mathematical models that were trained on phsycovisual data directly. This suggests that it is possible to predict psychophysical phenomena from visibility information embedded in image quality scores. Navaneeth K. Kottayil, Irene Cheng 0001, Giuseppe Valenzise, Frédéric Dufaux |
ICIP | 2 |
| 2018 | Adaptive Resolution Optimization and Tracklet Reliability Assessment for Efficient Multi-Object TrackingabstractRecent digital acquisition systems can acquire high-resolution videos, generating a large amount of dynamic data and leading to higher computational cost in online target tracking and learning, especially for complex scenes. We introduce an efficient and robust approach to improve the performance of multi-object online tracking and learning. Prior methods saved on computational cost by scaling down each video frame to a fixed smaller resolution, without considering the image features. Our algorithm computes the optimal image resolution adaptively by exploiting the correlation between an image's gray-value distribution and resolution. This dimensionality reduction step significantly improves the time performance in subsequent online tracking and learning, while preserving high tracking accuracy. Since a small detection error in one frame can cause cumulative error in the video sequence leading to incorrect labeling and tracking, we introduce a new tracklet reliability assessment metric to eliminate incorrect samples. Experimental results show that our approach can successfully track multiple objects in real time with both high precision and recall. Ruixing Yu, Irene Cheng 0001, Sweta Bedmutha, Anup Basu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2018 | Blind Quality Estimation by Disentangling Perceptual and Noisy Features in High Dynamic Range ImagesabstractHigh dynamic range (HDR) image visual quality assessment in the absence of a reference image is challenging. This research topic has not been adequately studied largely due to the high cost of HDR display devices. Nevertheless, HDR imaging technology has attracted increasing attention, because it provides more realistic content, consistent to what the human visual system perceives. We propose a new no-reference image quality assessment (NR-IQA) model for HDR data based on convolutional neural networks. The proposed model is able to detect visual artifacts, taking into consideration perceptual masking effects, in a distorted HDR image without any reference. The error and perceptual masking values are measured separately, yet sequentially, and then processed by a mixing function to predict the perceived quality of the distorted image. Instead of using simple stimuli and psychovisual experiments, perceptual masking effects are computed from a set of annotated HDR images during our training process. Experimental results demonstrate that our proposed NR-IQA model can predict HDR image quality as accurately as state-of-the-art full-reference IQA methods. Navaneeth K. Kottayil, Giuseppe Valenzise, Frédéric Dufaux, Irene Cheng 0001 |
IEEE Trans. Image Process. | 4 |
| 2017 | Facial expression recognition using SVM classification on mic-macro patternsabstractThe identification of facial expressions is a fundamental topic in the area of human computer interaction and pattern recognition. The research has gained significant attention in recent years. However many challenges still exist. This is because an individual might display different expressions at different times for the same mood. Expressions can also be influenced by health. Our proposed framework aims to capture unique information related to expressions from salient patches. We extract representative feature patterns at both micro and macro levels within a pixel-patch, and use a support vector machine (SVM) classifier to label expressions. Our experimental results using the Japanese facial expression (JAFEE) and Cohn-Kanade (CK) datasets achieve high recognition rate and efficient computation time, outperforming existing work. Housam Khalifa Bashier Babiker, Randy Goebel, Irene Cheng 0001 |
ICIP | 3 |
| 2017 | A content-specific IQM augmentation technique for better assessment of image qualityabstractIn this paper, we propose a computational strategy to enhance the performance of Image Quality Metrics (IQM) by using content specific features of an image. We do this by creating Visual Error Importance (VEI) map that is applied to the error maps computed by the IQM. A global optimization can be used to compute the VEI map that is optimal for any given IQM. We demonstrate this concept by categorizing the image content into three classes, generating and applying VEI to different IQM's and showing performance improvement in all of the cases. The performance evaluation was conducted on CSIQ dataset. Navaneeth K. Kottayil, Irene Cheng 0001 |
SMC | 2 |
| 2017 | Subjective and Objective Visual Quality Assessment of Textured 3D MeshesabstractObjective visual quality assessment of 3D models is a fundamental issue in computer graphics. Quality assessment metrics may allow a wide range of processes to be guided and evaluated, such as level of detail creation, compression, filtering, and so on. Most computer graphics assets are composed of geometric surfaces on which several texture images can be mapped to make the rendering more realistic. While some quality assessment metrics exist for geometric surfaces, almost no research has been conducted on the evaluation of texture-mapped 3D models. In this context, we present a new subjective study to evaluate the perceptual quality of textured meshes, based on a paired comparison protocol. We introduce both texture and geometry distortions on a set of 5 reference models to produce a database of 136 distorted models, evaluated using two rendering protocols. Based on analysis of the results, we propose two new metrics for visual quality assessment of textured mesh, as optimized linear combinations of accurate geometry and texture quality measurements. These proposed perceptual metrics outperform their counterparts in terms of correlation with human opinion. The database, along with the associated subjective scores, will be made publicly available online. Jinjiang Guo, Vincent Vidal 0002, Irene Cheng 0001, Anup Basu, Atilla Baskurt, Guillaume Lavoué |
ACM Trans. Appl. Percept. | 3 |
| 2016 | Highlighting objects of interest in an image by integrating saliency and depthabstractStereo images have been captured primarily for 3D reconstruction in the past. However, the depth information acquired from stereo can also be used along with saliency to highlight certain objects in a scene. This approach can be used to make still images more interesting to look at, and highlight objects of interest in the scene. We introduce this novel direction in this paper, and discuss the theoretical framework behind the approach. Even though we use depth from stereo in this work, our approach is applicable to depth data acquired from any sensor modality. Experimental results on both indoor and outdoor scenes demonstrate the benefits of our algorithm. Subhayan Mukherjee, Irene Cheng 0001, Anup Basu |
ICIP | 2 |
| 2016 | Robust Human Animation Skeleton Extraction Using Compatibility and Correctness ConstraintsabstractThe ability to automatically animate arbitrary 3D characters based on motion capture (MoCap) data has many applications in simulation, entertainment and multimedia transmission. However, defining trajectory key-points in human figures for animation without any manual intervention remains a challenging problem that makes complete automation difficult. To animate an articulated 3D character an animation skeleton needs to be extracted from, or be embedded into, a 3D model for deformation during animation. In conventional animation software, this process is mostly done manually by expert animators, which makes it a very tedious and time consuming step. The automatic rigging approaches proposed in the literature require a front facing model with neutral T-pose to accurately extract or embed an animation skeleton. We propose a fully automatic skeleton extraction approach based on optimization of constraints on human shape that can generate the animation skeleton, regardless of the model's orientation and position. Experimental results demonstrate the effectiveness of our approach. Robust skeleton extraction followed by efficient MoCap data compression can greatly improve the fidelity of 3D animated model transmission. Nasim Hajari, Irene Cheng 0001, Anup Basu |
ISM | 2 |
| 2016 | Spatio-Temporally Optimized Multi-sensor Motion FusionabstractThe latest advances in smart sensor technology, e.g., Leap Motion Sensor, has increased the precision in tracking fully articulated human hand and finger movements, without the need for placing electrical or optical markers. A remaining challenge is finger occlusion, which can affect tracking accuracy. In this paper, we introduce a spatio-temporal optimization technique for motion data generated from multiple sensors. We demonstrate that our algorithm can produce a fused stream of probabilistic optimal hand poses, by improving local spatial domain analysis and proposing a fast and effective flow analysis technique in the temporal domain, which computes how well the hand pose estimation in the current frame fits the movement flow within a time segment. By using an artificial hand to represent the hand pose ground truth at selected time steps, experimental results demonstrate that our spatio-temporal optimization algorithm increases the estimation accuracy by 6% compared to the reference method, achieving an overall accuracy of 91.29%. Our proposed method can be used offline or in real-time, and can benefit a wide range of applications, including surgical planning and training, where hand motion is the focus of performance efficiency and assessment. Xinyao Sun, Irene Cheng 0001, Anup Basu |
ISM | 2 |
| 2016 | Optimized per-joint compression of hand motion dataabstractMotion data is quickly expanding its application scope, following the recent advancements in smart sensing technology. In particular, it has been shown to be helpful for objective measurement and assessment of surgical dexterity among users at different levels of training. The goal is to allow trainees to evaluate their performance based on a reference set of hand movements. Similar to other multimedia data types, recording motion can produce a substantial amount of data, some of which are redundant for the application. Compression methods aim to optimize storage and transmission of motion capture (MoCap) data by taking advantage of temporal and spatial correlation. Hand motion data is a special sub-type of MoCap and is the focus of many applications, where hand movement evaluation is important. In this paper, we propose a lossy but visually indifferent, compression method that exploits redundancy found in hand motion data. Since individual joint movements have different impacts on the motion sequence, our technique is designed to minimize the overall distortion by providing a per-joint compression. We are able to demonstrate that our approach offers a quantitative gain for different compression ratios, while preserving visual quality. Antonio Carlos Furtado, Xinyao Sun, Anup Basu, Irene Cheng 0001 |
SMC | 4 |
| 2016 | Robust Lung Segmentation combining adaptive Concave Hulls with Active ContoursabstractLung segmentation is an important first step towards an automated CAD (Computer Aided Detection) system for a variety of medical applications. These applications range from lung nodule detection for identifying cancerous tumors to acinar shadow detection for identifying Tuberculosis. In our prior work we had used the Concave Hull algorithm for lung segmentation. However, our results showed over segmentation. In this work we introduce “Adaptive” concave hulls, combine it with Adaptive Median Filtering, and finally apply an Active Contour Model to make the results much more robust and eliminate the over segmentation and under segmentation problem. Our technique is especially useful for automated detection of Juxtapleural pulmonary nodules that are attached to the chest wall. Experimental results demonstrate the improvements achieved by our new algorithm. Sara Soltaninejad, Irene Cheng 0001, Anup Basu |
SMC | 2 |
| 2016 | A Multisensor Technique for Gesture Recognition Through Intelligent Skeletal Pose AnalysisabstractRecent advances in smart sensor technology and computer vision techniques have made the tracking of unmarked human hand and finger movements possible with high accuracy and at sampling rates of over 120 Hz. However, these new sensors also present challenges for real-time gesture recognition due to the frequent occlusion of fingers by other parts of the hand. We present a novel multisensor technique that improves the pose estimation accuracy during real-time computer vision gesture recognition. A classifier is trained offline, using a premeasured artificial hand, to learn which hand positions and orientations are likely to be associated with higher pose estimation error. During run-time, our algorithm uses the prebuilt classifier to select the best sensor-generated skeletal pose at each time step, which leads to a fused sequence of optimal poses over time. The artificial hand used to establish the ground truth is configured in a number of commonly used hand poses such as pinches and taps. Experimental results demonstrate that this new technique can reduce total pose estimation error by over 30% compared with using a single sensor, while still maintaining real-time performance. Our evaluations also demonstrate that our approach significantly outperforms many other alternative approaches such as weighted averaging of hand poses. An analysis of our classifier performance shows that the offline training time is insignificant, and our configuration achieves about 90.8% optimality for the dataset used. Our method effectively increases the robustness of touchless display interactions, especially in high-occlusion situations by analyzing skeletal poses from multiple views. Nathaniel Rossol, Irene Cheng 0001, Anup Basu |
IEEE Trans. Hum. Mach. Syst. | 2 |
| 2015 | Foveated High Efficiency Video Coding for Low Bit Rate TransmissionabstractThis work describes the design and subjective performance of Foveated High Efficiency Video Coding (FHEVC). Even though foveation has been widely used for various forms of compression since the early 1990s, we believe its use to improve HEVC is new. We consider the application of, possibly moving, foveated compression in this work and evaluate scenarios where it can be used to improve perceptual quality of videos under constrained transmission resources, e.g., bandwidth. A new method to reduce artifacts during remapping is also proposed. The preliminary implementation considers a single fovea only. Experiments summarizing user evaluations are presented to validate our implementation. Irene Cheng 0001, Masha Mohammadkhani, Anup Basu, Frédéric Dufaux |
ISM | 1 |
| 2015 | Normalized Gaussian Distance Graph Cuts for Image SegmentationabstractThis paper presents a novel, fast image segmentation method based on normalized Gaussian distance on nodes in conjunction with normalized graph cuts. We review the equivalence between kernel k-means and normalized cuts. Then we extend the framework of efficient spectral clustering and avoid choosing weights in the weighted graph cuts approach. Experiments on synthetic data sets and real-world images demonstrate that the proposed method is effective and accurate. Chengcai Leng, Wei Xu 0009, Irene Cheng 0001, Zhihui Xiong, Anup Basu |
ISM | 3 |
| 2015 | Perceptually motivated LSPIHT for motion capture data compression
Irene Cheng 0001, Amirhossein Firouzmanesh, Anup Basu |
Comput. Graph. | 1 |
| 2015 | Graph Matching Based on Stochastic PerturbationabstractThis paper presents a novel perspective on characterizing the spectral correspondence between the nodes of weighted graphs for image matching applications. The algorithm is based on the principal feature components obtained by stochastic perturbation of a graph. There are three areas of contributions in this paper. First, a stochastic normalized Laplacian matrix of a weighted graph is obtained by perturbing the matrix of a sensed graph model. Second, we obtain the eigenvectors based on an eigen-decomposition approach, where representative elements of each row of this matrix can be considered to be the feature components of a feature point. Third, correct correspondences are determined in a low-dimensional principal feature component space between the graphs. In order to further enhance image matching, we also exploit the random sample consensus algorithm, as a post-processing step, to eliminate mismatches in feature correspondences. The experiments on synthetic and real-world images demonstrate the effectiveness and accuracy of the proposed method. Chengcai Leng, Wei Xu 0009, Irene Cheng 0001, Anup Basu |
IEEE Trans. Image Process. | 3 |
| 2013 | Evaluation of 3D Model Segmentation Techniques Based on Animal Anatomyabstract3D model decomposition is a challenging and important problem in computer graphics. Several semantically based approaches have been proposed in the literature, however, due to the lack of proper evaluation criteria, comparison of these techniques is almost impossible. In this paper we suggest to use animal anatomy as the ground truth and compare the result of different segmentation techniques based on that. Differing from previous approaches which perform the evaluation based on ground truth databases created subjectively by human observers, we consider expert knowledge on anatomy of various animals. Based on this knowledge we specify the ground truth for different animals and compare alternative algorithms. Nasim Hajari, Irene Cheng 0001, Anup Basu, Guillaume Lavoué |
SMC | 2 |
| 2013 | Perceptual Quality Metrics for 3D Meshes: Towards an Optimal Multi-attribute Computational Modelabstract3D graphical data, commonly represented using triangular meshes, are deployed in a wide range of application processes including compression, filtering, watermarking, and simplification. These processes often introduce geometric distortions which affect the visual quality of the ultimate data visualization. In order to accurately evaluate perceptual impacts caused by the distortions, assessment metrics on 3D Mesh Visual Quality (MVQ) have been extensively discussed in the literature. Researchers recommended various metrics to predict the adverse effects that visual artifacts can have in applications. Most of these metrics are based on geometric attributes, conventional geometric distance, Laplacian coordinates, different types of curvature computation, and dihedral angles. We hypothesize that an optimal combination of multiple attributes associated with a 3D mesh surface can contribute to better perceptual prediction than single attributes used separately. In this paper, we use two user studies to validate our hypothesis. Our contributions are: (1) providing a detailed analysis of the most relevant geometric attributes for mesh quality assessment, and (2) introducing a new perceptual evaluation metric based on multiple attributes, with the optimal combination determined through machine learning techniques. Statistical quantitative analysis shows that our metric delivers better results than other state-of-the-art approaches. The proposed method is simple to implement and fast in execution. Moreover, our framework can easily be expanded to accommodate additional surface attributes. Guillaume Lavoué, Irene Cheng 0001, Anup Basu |
SMC | 2 |
| 2013 | QoE-Based Multi-Exposure Fusion in Hierarchical Multivariate Gaussian CRFabstractMany state-of-the-art fusion methods, combining details in images taken under different exposures into one well-exposed image, can be found in the literature. However, insufficient study has been conducted to explore how perceptual factors can provide viewers better quality of experience on fused images. We propose two perceptual quality measures: perceived local contrast and color saturation, which are embedded in our novel hierarchical multivariate Gaussian conditional random field model, to illustrate improved performance for multi-exposure fusion. We show that our method generates images with better quality than existing methods for a variety of scenes. Rui Shen 0002, Irene Cheng 0001, Anup Basu |
IEEE Trans. Image Process. | 2 |
| 2012 | A general optimal pixel aspect ratio model for stereo-based 3D reconstruction and visualizationabstractWe propose a mathematical model for optimizing pixel aspect ratio for the best 3D estimation in both stereo-based 3D reconstruction and 3D viewing applications. We analyze the 3D reconstruction through a 3D display medium and reduce the whole process to a single stereo system so that a unified model can be applied to both direct reconstruction from stereo images and indirect reconstruction through stereo content presented on a 3D display. We use this unified model to extend our earlier work to a general solution for determining the optimal pixel aspect ratio for both applications. Unlike earlier work, the solution proposed here relates the optimal pixel aspect ratio to the device-specific parameters, rather than the stereo configuration parameters, which makes it more easily applicable in design and manufacture of stereo capture and display devices. In general, our mathematical model and subjective user studies suggest that, for a given total resolution, a finer horizontal discretization with a ratio of about 0.6 leads to a more accurate 3D reconstruction and a better 3D visual experience. Hossein Azari, Irene Cheng 0001, Anup Basu |
SMC | 2 |
| 2012 | Hand and face tracking under occlusion with anthropomorphic constraintsabstractWe propose a graphical model for a decentralized, simultaneous detection and tracking algorithm for efficient localization of hands from a sequence in color and range images. We deduce the location of key-points using a Bayesian framework. We use anthropomorphic constraints for modelling body part articulation. Furthermore, our algorithm reasons about occlusion and preserves data association to deal with ambiguities. Experimental results demonstrate that our system tracks face and hands more accurately in video, compared to prior research. Abhishek Sen, Irene Cheng 0001, Anup Basu |
SMC | 2 |
| 2012 | HW/SW co-design of an embedded omni-imaging systemabstractOmni-imaging can be used in many practical applications that need a wide field of view, therefore a real-time and high-definition embedded system design and implementation of omni-imaging is desired. In this study, we propose a hardware/software co-design method for the design and implementation of embedded omni-imaging systems. In order to achieve real-time and high-definition goals, we perform hardware/software partitioning based on the analysis of functional modules in a basic embedded omni-imaging system. In the experiments, the proposed hardware/software co-design omni-imaging system is implemented in a FPGA (Field Programmable Gate Array) plus DSP (Digital Signal Processor) system architecture. Results indicate that the omni-imaging speed achieved is 39fps with the imaging resolution set at 1024×768 for the original omni-image and 1280×288 for the unwarped image. Zhihui Xiong, Irene Cheng 0001, Maojun Zhang, Anup Basu |
SMC | 2 |
| 2012 | Optimal pixel aspect ratio for enhanced 3D TV visualization
Hossein Azari, Irene Cheng 0001, Kostas Daniilidis, Anup Basu |
Comput. Vis. Image Underst. | 2 |
| 2012 | Perceptually Coded Transmission of Arbitrary 3D Objects over Burst Packet Loss Channels Enhanced with a Generic JND FormulationabstractIn this work we propose a new approach to account for burst packet loss during transmission of 3D objects represented by texture and mesh over unreliable networks. Our strategy includes applying stripification on the 3D mesh following the valence-driven algorithm and distributing nearby vertices into different packets, combined with an interleaving technique that does not need texture or mesh packets to be re-transmitted. The perceptually-driven technique is able to successfully interpolate lost mesh features even under severe packet loss. The reconstructed mesh is further improved by applying our curvature-driven probabilistic strategy to safeguard visually significant structures on the 3D surface. Experimental results show that smoothness on the object surface is preserved even at 50% packet loss. At 75% packet loss, smoothness on the object surface deteriorates but the overall shape of an object is still preserved. We also define a Quality of Experience (QoE) metric to formulate the Just-Noticeable-Difference (JND) concept, to quantify the qualitative findings obtained from earlier subjective user studies, which provides flexibility to applications for reducing the transmission of visually redundant data. Irene Cheng 0001, Lihang Ying, Anup Basu |
IEEE J. Sel. Areas Commun. | 1 |
| 2012 | Efficient omni-image unwarping using geometric symmetry
Zhihui Xiong, Irene Cheng 0001, Anup Basu, Wei Wang 0068, Wei Xu 0019, Maojun Zhang |
Mach. Vis. Appl. | 2 |
| 2011 | Optimized point splatting based on a fully-balanced hierarchical structureabstractWe propose an improved two-layer hierarchical point-based 3D model representation for interactive 3D rendering. A tree structure called multi-section tree is introduced that allows creating a fully-balanced hierarchy with a desired branching factor for any arbitrary number of points. The proximity problem, the challenge of keeping close-by points together in equi-partitioning process, is addressed by applying an initial bottom-up point-grouping process that divides the 3D model surface into small, nearly equal-size groups of neighboring points. Each group is represented by a small multi-section subtree but treated as a single point in the main top-down equi-splitting process. The top-down process creates the main balanced multi-section tree which holds the whole structure together. Results show that our two-step description leads to a very compact quantized representation with enhanced rendering quality. Hossein Azari, Irene Cheng 0001, Anup Basu |
ICME | 2 |
| 2011 | An in-place texture synthesis technique for memory constrained multimedia applicationsabstractDespite the rapid evolution of multimedia content, from 2D to 3D and to stereo on IMAX display, material and texture remain an indispensable component when rendering realistic and appealing graphics and animations. As the demand for high-definition displays increases, so is the necessity for high-resolution textures. Nevertheless, the available texture images very often have low resolution and are inadequate for high-quality rendering. Texture synthesis from examples offers an effective way not only for the creation of high resolution texture, but also useful in interpolating missing data resulted from unreliable transmission and overly compressed data. We present a new memory-efficient technique that facilitates high-quality texture synthesis. We compare the time performance and memory usage between our approach and the commonly used caching techniques. Experimental results show that our method can perform better, given limited memory resources. Alexey Badalov, Irene Cheng 0001, Anup Basu |
ICME | 2 |
| 2011 | Anatomy preserving 3D model decomposition based on robust skeleton-surface node correspondenceabstractIn this work, we present an effective anatomy preserving model decomposition technique. By extracting unit-width curve skeletons, which are robust to noise, and mapping skeleton branches to model surface nodes, our method accurately identifies the topology and geometry information of a 3D model, resulting in more semantically rich segmented components. Experiments on 2194 models from the Princeton Shape Benchmark and A Benchmark for 3D Mesh Segmentation demonstrate the advantage of the proposed technique. Our results preserve better anatomical structures compared to three commonly used model segmentation methods. Irene Cheng 0001, Anup Basu |
ICME | 2 |
| 2011 | Efficient video sequences alignment using unbiased bidirectional dynamic time warping
Cheng Lu 0001, Meghna Singh, Irene Cheng 0001, Anup Basu, Mrinal Mandal 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2011 | A 2-point algorithm for 3D reconstruction of horizontal lines from a single omni-directional image
Irene Cheng 0001, Zhihui Xiong, Anup Basu, Maojun Zhang |
Pattern Recognit. Lett. | 2 |
| 2011 | Generalized Random Walks for Fusion of Multi-Exposure ImagesabstractA single captured image of a real-world scene is usually insufficient to reveal all the details due to under- or over-exposed regions. To solve this problem, images of the same scene can be first captured under different exposure settings and then combined into a single image using image fusion techniques. In this paper, we propose a novel probabilistic model-based fusion technique for multi-exposure images. Unlike previous multi-exposure fusion methods, our method aims to achieve an optimal balance between two quality measures, i.e., local contrast and color consistency, while combining the scene details revealed under different exposures. A generalized random walks framework is proposed to calculate a globally optimal solution subject to the two quality measures by formulating the fusion problem as probability estimation. Experiments demonstrate that our algorithm generates high-quality images at low computational cost. Comparisons with a number of other techniques show that our method generates better results in most cases. Rui Shen 0002, Irene Cheng 0001, Jianbo Shi, Anup Basu |
IEEE Trans. Image Process. | 2 |
| 2011 | Perceptually Guided Fast Compression of 3-D Motion Capture DataabstractA time efficient compression technique, incorporating attention stimulating factors, for motion capture data is proposed. Compression ratios of 25:1 to 30:1 can be achieved with very little noticeable degradation in perceptual quality of animation. Experimental analysis shows that the proposed algorithm is much faster than comparable approaches using wavelets, thereby making our approach feasible for motion capture, transmission, and real-time synthesis on mobile devices, where processing power and memory capacity are limited. Amirhossein Firouzmanesh, Irene Cheng 0001, Anup Basu |
IEEE Trans. Multim. | 2 |
| 2010 | Fully automatic brain tumor segmentation using a normalized Gaussian Bayesian Classifier and 3D Fluid Vector FlowabstractBrain tumor segmentation from Magnetic Resonance Images (MRIs) is an important task to measure tumor responses to treatments. However, automatic segmentation is very challenging. This paper presents an automatic brain tumor segmentation method based on a Normalized Gaussian Bayesian classification and a new 3D Fluid Vector Flow (FVF) algorithm. In our method, a Normalized Gaussian Mixture Model (NGMM) is proposed and used to model the healthy brain tissues. Gaussian Bayesian Classifier is exploited to acquire a Gaussian Bayesian Brain Map (GBBM) from the test brain MRIs. GBBM is further processed to initialize the 3D FVF algorithm, which segments the brain tumor. This algorithm has two major contributions. First, we present a NGMM to model healthy brains. Second, we extend our 2D FVF algorithm to 3D space and use it for brain tumor segmentation. The proposed method is validated on a publicly available dataset. Irene Cheng 0001, Anup Basu |
ICIP | 2 |
| 2010 | An Improved Fluid Vector Flow for Cavity Segmentation in Chest RadiographsabstractFluid vector flow (FVF) is a recently developed edge-based parametric active contour model for segmentation. By keeping its merits of large capture range and ability to handle acute concave shapes, we improved the model from two aspects: edge leakage and control point selection. Experimental results of cavity segmentation in chest radiographs show that the proposed method provides at least 8% improvement over the original FVF method. Irene Cheng 0001, Mrinal Mandal 0001 |
ICPR | 2 |
| 2010 | Automatic segmentation of spinal cord mri using symmetric boundary tracingabstractWe develop an adaptive active contour tracing algorithm for extraction of spinal cord from MRI that is fully automatic, unlike existing approaches that need manually chosen seeds. We can accurately extract the target spinal cord and construct the volume of interest to provide visual guidance for strategic rehabilitation surgery planning. Dipti Prasad Mukherjee, Irene Cheng 0001, Nilanjan Ray, Vivian Mushahwar, R. Marc Lebel, Anup Basu |
IEEE Trans. Inf. Technol. Biomed. | 2 |
| 2009 | Distortion metric for robust 3D point cloud transmissionabstractThis paper discusses using a forward error correction (FEC) algorithm to protect the transmission of progressively compressed 3D point clouds against packets loss. We design a metric to evaluate each layer's quality contribution to the decoding result of the progressively compressed model. With this metric, we minimize the expected distortion when applying an Unequal Error Protection (UEP) strategy to allocate channel bits to different layers of the model. The performance of employing UEP and Equal Error Protection (EEP) are compared with respect to the expected distortion. Experimental results show that by incorporating our distortion estimation metric with UEP, the rendering quality of a reconstructed 3D model degrades more gracefully as the packet-loss rate increases. Feng Chen 0003, Irene Cheng 0001, Anup Basu |
ICME | 2 |
| 2009 | A multimedia item authoring framework for computer-based educationabstractPerceptually inspired curriculum using multimedia content connects students to subject matter and deepens their understanding of abstract concepts. This approach has become increasingly attractive in education. Multimedia content can arouse user engagement through interactivity and immersion, and thus inspire a student to learn. Differing from multiple-choice, multimedia items require diverse screen layouts. Non-standard templates bring challenges to techers, who either do not have the programming skills or cannot afford the time outside their primary duties to study complex templates. Inflexibility in item creation may cause hesitation in adopting new technologies. In order to support a smooth transition from conventional to multimedia item creation, we introduce a Multimedia Item Generator (MIG) for educators. Portability, reusability, scalability and interoperability are the characteristics of MIG. In this paper, we describe the design and implementation, and present our future plan. Irene Cheng 0001, Alexey Badalov |
ICME | 1 |
| 2009 | ACM 2009 workshop on ambient media computing (AMC'09) overviewabstractNo abstract available. Howard Leung, Cha Zhang, Qing Li 0001, Rynson W. H. Lau, Benjamin W. Wah, Abdulmotaleb El Saddik, K. Selçuk Candan, Irene Cheng 0001 |
ACM Multimedia | 8 |
| 2009 | Interactive Graphics for Computer Adaptive TestingabstractAbstract Interactive graphics are commonly used in games and have been shown to be successful in attracting the general audience. Instead of computer games, animations, cartoons, and videos being used only for entertainment, there is now an interest in using interactive graphics for ‘innovative testing’. Rather than traditional pen‐and‐paper tests, audio, video and graphics are being conceived as alternative means for more effective testing in the future. In this paper, we review some examples of graphics item types for testing. As well, we outline how games can be used to interactively test concepts; discuss designing chemistry item types with interactive 3D graphics; suggest approaches for automatically adjusting difficulty level in interactive graphics based questions; and propose strategies for giving partial marks for incorrect answers. We study how to test different cognitive skills, such as music, using multimedia interfaces; and also evaluate the effectiveness of our model. Methods for estimating difficulty level of a mathematical item type using Item Response Theory (IRT) and a molecule construction item type using Graph Edit Distance are discussed. Evaluation of the graphics item types through extensive testing on some students is described. We also outline the application of using interactive graphics over cell phones. All of the graphics item types used in this paper are developed by members of our research group. Irene Cheng 0001, Anup Basu |
Comput. Graph. Forum | 1 |
| 2008 | Optimization of Symmetric Transfer Error for Sub-frame Video Synchronization
Meghna Singh, Irene Cheng 0001, Mrinal Mandal 0001, Anup Basu |
ECCV (2) | 2 |
| 2008 | Self-Tutoring, Teaching and Testing: An Intelligent Process AnalyzerabstractMastery of basic concepts and logical units is a prerequisite for solving complex problems. The divide-and-conquer strategy has been used successfully in a variety of problem solving situations, including algorithm design and software engineering. In this paper, we adopt a similar approach and propose a process analyzer in education. The goal is to help students improve their problem solving skills, as well as to assist teachers to monitor student response so that guidance can be provided as required. Complex and tedious processes are often encountered in curricula such as physics, chemistry, and mathematics, where the final answer is built upon the results of numerous smaller processes. Therefore the process analyzer defines the top-level process as a hierarchy composed of smaller integral parts so that students who are not able to solve the problem on their own are able to look at lower-level simpler processes and follow the hints leading to the correct answer. Question designers can define hints, as well as the score in each step. Responses are recorded for modeling student performance. The students interact with the Process Analyzer through a graphical interface, which provides an engaging and motivating environment. The positive feedback on our process analyzer shows the feasibility of our approach. We describe the design and implementation of the process analyzer, and present our future plan. Irene Cheng 0001, Nathaniel Rossol, Randy Goebel |
ICALT | 1 |
| 2008 | An Algorithm for Automatic Difficulty Level Estimation of Multimedia Mathematical Test ItemsabstractThe use of multimedia in learning and testing has been gaining popularity over the last few years. In this paper, we describe a method for estimating difficulty level of a multimedia mathematical item type using Item Response Theory (IRT). Experimental results on determining the parameters of an IRT model through tests on a group of students are also presented. Linear regression equations are fitted based on test data collected on a group of students. Results show a high degree of correlation between observed difficulty levels and fitted values. Irene Cheng 0001, Rui Shen 0002, Anup Basu |
ICALT | 1 |
| 2008 | Assessing rhythm recognition skills in a multimedia environmentabstractMultimedia content is an effective tool for enhancing learning and testing, and it is more effective than the traditional paper and pencil presentation format. In order to make educational materials more intuitive, more interactive and more inspiring to students, image, audio, video, graphics, animation and 3D items are widely used in current educational applications. However, they are designed mainly for learning and not for testing. Those designed for testing focus on using simple, text-based formats (e.g. multiple-choice and True/False questions) to assess a studentpsilas subject knowledge rather than on evaluating a studentpsilas cognitive skills and problem-solving ability. In this paper, we propose an interactive framework that uses an innovative test format enriched with audios and videos, for evaluating a studentpsilas skills, in this case musical rhythms recognition skills. Experiments were conducted with human observers and the results verify the feasibility of our approach. The goal of this paper is to advance multimedia research in education by extending its capabilities beyond simple knowledge retention and reproduction, and to model the acquisition of cognitive skills in a dynamic context. Irene Cheng 0001, Chris Kerr, Walter F. Bischof |
ICME | 1 |
| 2008 | Stereo matching using random walksabstractThis paper presents a novel two-phase stereo matching algorithm using the random walks framework. At first, a set of reliable matching pixels is extracted with prior matrices defined on the penalties of different disparity configurations and Laplacian matrices defined on the neighbourhood information of pixels. Following this, using the reliable set as seeds, the disparities of unreliable regions are determined by solving a Dirichlet problem. The variance of illumination across different images is taken into account when building the prior matrices and the Laplacian matrices, which improves the accuracy of the resulting disparity maps. Even though random walks have been used in other applications, our work is the first application of random walks in stereo matching. The proposed algorithm demonstrates good performance using the Middlebury stereo datasets. Rui Shen 0002, Irene Cheng 0001, Xiaobo Li 0001, Anup Basu |
ICPR | 2 |
| 2007 | Contrast Enhancement from Multiple Panoramic ImagesabstractIn this work we discuss an efficient strategy for combining multiple panoramic scans of the same scene to create a single higher contrast image. Prior research papers mainly consider that precise exposure times for multiple images are known and create a new high dynamic range representation for each pixel. We simply consider multiple scans of the same scene where the exposure times are not known, and no special filters or image detectors are used in the image acquisition process. In this unrestricted scenario, the problem is how to select different parts of a scene from the "best" image among a collection of images in order to optimize the overall clarity or contrast of a single final image. Experimental results are shown, and steps to improve results using modified contrast measures for color images are described. The results are enhanced using morphological filtering and image blending. Irene Cheng 0001, Anup Basu |
ICCV | 1 |
| 2007 | Multimedia Adaptive Computer based Testing: An OverviewabstractInstead of computer games, animations, cartoons, and videos being used only for entertainment by kids, there is now an interest in using multimedia for "innovative testing." Rather than traditional paper-and-pencil tests, audio, video and graphics are being conceived as alternative means for more effective testing in the future. In this paper we review some examples of multimedia item types for testing. As well, we will outline research on testing the seven types of intelligence through multimedia; describe how games can be used to test physics concepts; discuss designing chemistry item types with interactive multimedia; consider architecture for supporting online multimedia testing; and suggest approaches for automatically determining difficulty level in interactive mathematical questions. Detailed description on various topics will be given in the other papers to follow in the special session. Anup Basu, Irene Cheng 0001, Mun Prasad, Gautam Rao |
ICME | 2 |
| 2007 | Multimedia Item Type Design for Assessing Human Cognitive SkillsabstractMultimedia content has been used in education applications, e.g. in distance learning, to make learning more intuitive, more interactive and more effective than the traditional presentation formats. Although image, audio, video, graphics, animation and 3D representation can be found in current multimedia implementations, they are designed mainly for learning and not for testing. Most educational tests rely on simple, text-based items (e.g. multiple-choice questions), which focus on assessing a student's knowledge rather than on evaluating the student's cognitive skills and problem-solving abilities. In this paper, we propose a novel design that uses innovative test item types enriched with multimedia content, for evaluating a student's cognitive skills including linguistic, logical-mathematical, spatial, bodily-kinesthetic, musical, interpersonal and intrapersonal. Irene Cheng 0001, Walter F. Bischof |
ICME | 1 |
| 2007 | An Effective Multimedia Item Shell Design for Individualized EducationabstractThe advantages of creating multimedia item types and applying computer-based adaptive testing in education are: First, the capability to motivate learning by making the learners feel more engaged and interactive, as well as a better representation of concepts, which are not possible when using conventional multiple choice tests. Second, instead of following a curriculum designed for average students, an individual is given a customized curriculum suitable for his or her learning capability. However, the issue to address when achieving these goals is the enormous amount of item types required to transform the current multiple choice questions into multimedia formats, and the criteria used to determine the difficulty level of a multimedia question item. In this paper we propose a multimedia item shell design that not only reduces the number of item types required, but also computes difficulty level of an item automatically. The concept of question seed is also introduced to make content creation more cost-effective. Irene Cheng 0001, Anup Basu |
ICME | 1 |
| 2007 | Multimedia Games for Learning and Testing PhysicsabstractMany difficult concepts can often be best explained or understood through simple games. In order to motivate the understanding of movements and trajectories of projectiles, and other Physics concepts, we propose using educational games for learning and testing physics so that the students can feel more engaged and rewarding, and thus are motivated to acquire new knowledge. Our novel approach includes automatically generating the next game episode at a different difficulty level based on how the student performs in the current episode. This adaptive approach is associated with a computer scoring system which can evaluate the student's understanding of physics concepts and computational skills, as well as strategic planning. Saul D. Rodriguez, Irene Cheng 0001, Anup Basu |
ICME | 2 |
| 2007 | An Interactive 3D Environment for Computer Based EducationabstractIndividual learning capabilities can vary from gifted to exceptionally slow; some students may take longer to understand a concept and may not be able to achieve the expected standard. However, in a non-discriminative education system, one of the objectives is to motivate individuals to acquire new knowledge and understand concepts, no matter what their initial skills and knowledge levels may be. Students get motivated and will continue to learn only if they feel engaged and rewarded. In this paper we propose an interactive environment, using 3D rendering, for computer-based learning and testing. A novel approach to automatically assign difficulty levels to questions and to adaptively evaluate individual student performance is discussed along with some examples. Irene Cheng 0001 |
ICME | 2 |
| 2007 | Perceptually Optimized 3-D Transmission Over Wireless NetworksabstractMany protocols optimized to transmissions over wireless networks have been proposed. However, one issue that has not been looked into is considering human perception in deciding a transmission strategy for three-dimensional (3D) objects. Several factors, such as the number of vertices and the resolution of texture, can affect the display quality of 3D objects. When the resources of a graphics system are not sufficient to render the ideal image, degradation is inevitable. It is therefore important to study how individual factors affect the overall quality, and how the degradation can be controlled given limited bandwidth resources and possibility of data loss. In this paper, the essential factors determining the display quality are reviewed. We provide an overview of our research on designing a 3D perceptual quality metric integrating two important ones, resolution of texture and resolution of mesh, that control the transmission bandwidth requirements. A review of robust mesh transmission considering packet loss is presented, followed by a discussion of the difference of existing literature with our problem and approach. We then suggest alternative strategies for packet transmission of both 3D texture and mesh. These strategies are then compared with respect to preserving 3D perceptual quality under packet loss Irene Cheng 0001, Anup Basu |
IEEE Trans. Multim. | 1 |
| 2006 | Packet Loss Modeling for Perceptually Optimized 3D TransmissionabstractTransmissions over wireless and other unreliable networks can lead to packet loss. An area that has received limited research attention is how to tailor multimedia information taking into account the way packets are lost. We provide a brief overview of our research on designing a 3D perceptual quality metric integrating two important factors, resolution of texture and resolution of mesh, which control transmission bandwidth. We then suggest alternative strategies for packet 3D transmission of both texture and mesh. These strategies are then compared with respect to preserving 3D perceptual quality under packet loss in ad hoc wireless networks. Experiments are conducted to study how the time between consecutive packet transmission and packet size affects loss in wireless channels. A preliminary model for estimating the optimal packet size is then proposed Irene Cheng 0001, Lihang Ying, Anup Basu |
ICME | 1 |
| 2006 | Improving Multimedia Innovative Item Types for Computer Based TestingabstractInstead of computer games, animations, cartoons, and videos being used only for entertainment by kids, there is now an interest in using some of these media for educational purposes as well. Along with content creation, multimedia has potential for use in "innovative testing". Rather than traditional paper-and-pencil tests, audio, video and graphics are being conceived as alternative means for more effective testing in the future [1,17,21,28,29,30,33,42,44,49,50]. For example, we would like to use animation and games to help in learning concepts; consider how image, graphics and audio tools can be used for innovative testing; and develop techniques for measuring the impact of multimedia in improving performance or arousing interest in students. In this paper we discuss some examples of multimedia item types for testing, followed by a strategy for adaptive testing using those item types. We also show how techniques for perceptual evaluations can be used to improve strategies for adaptive testing Irene Cheng 0001, Anup Basu |
ISM | 1 |
| 2006 | Perceptually Enhanced Multimedia Processing, Visualization and TransmissionabstractData reduction has long been a method for adaptation to limited computational and network resources. But one major concern is the tradeoff between preserving visual quality and reducing data size. Furthermore, the presence of multi-modal data, e.g. visual and aural, is common and therefore distributing competing resources among multi-modal data to achieve optimal visual quality becomes a major. Since humans are typically the penultimate viewer of multimedia data, it is reasonable to take human perception into consideration during the data reduction and resource distribution process, in order to estimate and control the resulting visual quality. Psychophysical experiments reported in the literature have shown that better performance can be achieved in multimedia processing, visualization and transmission, by incorporating perceptual factors. This paper gives an overview of how human perception plays a role in the development of multimedia applications, so as to inspire and inform future research in this direction Irene Cheng 0001, Randy Goebel |
ISM | 1 |
| 2006 | Perceptual Analysis of Level-of-Detail: The JND ApproachabstractMultimedia content is becoming more widely used in applications resulting from the easy accessibility of high-speed networks in the public domain. An important component in multimedia content is 3D geometry, which in the past had low resolution due to acquisition, computational and network limitation, and was not able to approximate 3D surfaces realistically. Although processing speed and network capacity have been greatly increased in the last decade, the increase in demands for multimedia content surpass the increase in resources. Consequently, techniques for data simplification especially for 3D mesh data is inevitable in order to achieve shorter latency and satisfactory interactivity in applications. This paper presents a perceptual analysis to evaluate the visual quality associated with a change in level-of-detail. Our analysis is consistent to how the human visual system evaluates 3D objects in the real world and is based on the just-noticeable-difference methodology. Experimental results show that our approach presents an accurate estimation of visual quality and thus provides a systematic method to evaluate the performance of different simplification algorithms Irene Cheng 0001, Rui Shen 0002, Xing-Dong Yang, Pierre Boulanger |
ISM | 1 |
| 2006 | Adaptive online transmission of 3-D TexMesh using scale-space and visual perception analysisabstractEfficient online visualization of three-dimensional (3-D) mesh, mapped with photo realistic texture, is essential for a variety of applications such as museum exhibits and medical images. In these applications synthetic texture or color per vertex loses authenticity and resolution. An image-based view dependent approach requires too much overhead to generate a 360/spl deg/ display for online applications. We propose using a mesh simplification algorithm based on scale-space analysis of the feature point distribution, combined with an associated visual perception analysis of the surface texture, to address the needs of adaptive online transmission of high quality 3-D objects. The premise of the proposed textured mesh (TexMesh)simplification, taking the human visual system into consideration, is the following: given limited bandwidth, texture quality in low feature density surfaces can be reduced, without significantly affecting human perception. The advantage of allocating higher bandwidth, and thus higher quality, to dense feature density surfaces, is to improve the overall visual fidelity. Statistics on feature point distribution and their associated texture fragments are gathered during preprocessing. Online transmission is based on these statistics,which can be retrieved in constant time. Using an initial estimated bandwidth,a scaled mesh is first transmitted. Starting from a default texture quality,we apply an efficient Harmonic Time Compensation Algorithm based on the current bandwidth and a time limit, to adaptively adjust the texture quality of the next fragment to be transmitted. Properties of the algorithm are proved. Experimental results show the usefulness of our approach. Irene Cheng 0001, Pierre Boulanger |
IEEE Trans. Multim. | 1 |
| 2005 | Balanced incomplete designs for 3D perceptual quality estimationabstractMany factors, such as the number of vertices and the resolution of texture, can affect the display quality of 3D objects. When the resources of a graphics system are not sufficient to render the ideal image, degradation is inevitable. It is therefore important to study how individual factors affect the overall quality, and how the degradation can be controlled given limited resources. The essential factors determining the display quality are reviewed and a 3D perceptual estimation method is described. One of the major concerns in designing perceptual experiments is the large number of evaluations to be performed by judges, which results in fatigue and errors in judgement. To reduce judging fatigue and increase reliability of evaluations we propose using a statistical approach, following the design of experiments technique of balanced incomplete block design (BIBD). We develop models following BIBD for perceptual experiments and validate our model with experimental results. Even though the BIBD framework for perceptual evaluations is described in the context of 3D quality estimation, the approach can be used in other perceptual evaluation scenarios as well. Anup Basu, Irene Cheng 0001 |
ICIP (1) | 2 |
| 2005 | Feature extraction on 3-D TexMesh using scale-space analysis and perceptual evaluationabstractEfficient online visualization of three-dimensional (3-D) textured models is essential for a variety of applications including not only games and e-commerce, but also heritage and medicine. To visualize 3-D objects online, it is necessary to quickly adapt both mesh and texture to the available computational or network resources. Earlier research showed that after reaching a minimum required mesh density, high-resolution texture has more impact on human perception than a denser mesh. Given limited bandwidth, an important issue is how to extract features that best represent the original object, and how to allocate resources between mesh and texture data to achieve optimal perceptual quality. In this paper, we propose a textured mesh (TexMesh) model, which applies scale-space analysis and perceptual evaluation to extract 3-D features for textured mesh simplification and transmission. Texture data is divided into fragments to facilitate quality and bandwidth adaptation. Texture quality assignment is based on feature point distribution. Online transmission is based on statistics gathered during preprocessing, which are stored in a priority queue and lookup tables. Quality of service requested by a client site can be met by applying an efficient adaptive algorithm to ensure optimal use of the specified time and available bandwidth, and at the same time preserving satisfactory quality. Our TexMesh framework integrates feature extraction, mesh simplification, texture reduction, bandwidth adaptation, and perceptual evaluation into a multiscale visualization framework. Irene Cheng 0001, Pierre Boulanger |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2005 | Quality metric for approximating subjective evaluation of 3-D objectsabstractMany factors, such as the number of vertices and the resolution of texture, can affect the display quality of three-dimensional (3-D) objects. When the resources of a graphics system are not sufficient to render the ideal image, degradation is inevitable. It is, therefore, important to study how individual factors will affect the overall quality, and how the degradation can be controlled given limited resources. In this paper, the essential factors determining the display quality are reviewed. We then integrate two important ones, resolution of texture and resolution of wireframe, and use them in our model as a perceptual metric. We assess this metric using statistical data collected from a 3-D quality evaluation experiment. The statistical model and the methodology to assess the display quality metric are discussed. A preliminary study of the reliability of the estimates is also described. The contribution of this paper lies in: 1) determining the relative importance of wireframe versus texture resolution in perceptual quality evaluation and 2) proposing an experimental strategy for verifying and fitting a quantitative model that estimates 3-D perceptual quality. The proposed quantitative method is found to fit closely to subjective ratings by human observers based on preliminary experimental results. Yixin Pan, Irene Cheng 0001, Anup Basu |
IEEE Trans. Multim. | 2 |
| 2004 | Scale-space 3D TexMesh simplificationabstractEfficient online 3D visualization is essential for a variety of applications, including not only games and e-commerce, but also heritage and medicine. For efficient online visualization, it is necessary to adapt 3D models (both mesh and texture) quickly to the available computational or network resources. We propose using 3D model simplification based on a scale-space analysis of the surface curvature variations combined with an associated scale-space analysis of the surface texture to reduce the size of texture files, and facilitate distributed transmission. The premise of the proposed simplification is that minor variations in texture can be ignored in relatively smooth regions of a 3D surface, without significantly affecting human perception. Statistics of feature points and their associated texture fragments are gathered during preprocessing. On-line transmission and rendering for the next higher resolution scale is based on the statistics, which can be retrieved in constant time. Quality of service (QoS) can be provided based on the time limit, or the number of vertices, or faces, requested by the viewer. Experimental results showing the simplified models demonstrate the feasibility of our approach. Irene Cheng 0001, Pierre Boulanger |
ICME | 1 |
| 2003 | Perceptual quality metric for qualitative 3D scene evaluationabstractVarious factors affect the quality of 3D images, such as the number of vertices and the resolution of texture. In this report, we discuss the factors determining the quality of 3D images. Among these factors two important ones, most relevant for bandwidth constrained online applications, are combined to model a perceptual metric. We estimate this metric from the statistical data collected in a 3D quality evaluation experiment. The theory behind modeling a quality metric and the details of experiments performed for deriving the parameters of this metric are described. Yixin Pan, Irene Cheng 0001, Anup Basu |
ICIP (3) | 2 |
| 2003 | Efficient 3D object simplification and fragmented texture scaling for online visualizationabstractVisualization of 3D images is becoming more commonplace for a variety of applications including online games and e-commerce. For efficient online visualization of 3D objects it is necessary to quickly adapt 3D models (including the wireframe and texture) to the available computational or network resources. In this paper we propose enhancements to 3D model simplification based on simplification envelopes (ESE) and curvature variations (CPD) on the surface of an object. We compare the two approaches and demonstrate that the CPD method is much faster than ESE while producing comparable results. A technique for texture simplification driven by model simplification is also proposed. Experimental results comparing the simplified models and the computational time demonstrate the feasibility of our approach. Irene Cheng 0001 |
ICME | 1 |
| 2003 | QoS based video delivery with foveation and bandwidth monitoring
Irene Cheng 0001, Anup Basu |
Pattern Recognit. Lett. | 1 |
| 2003 | Optimal adaptive bandwidth monitoring for QoS based retrievalabstractNetwork aware multimedia delivery applications are a class of applications that provide certain level of quality of service (QoS) guarantees to end users while not assuming underlying network resource reservations. These applications guarantee QoS parameters like media object transmission time limit by actively monitoring the available bandwidth of the network and adapting the object to a target size that can be transmitted within a given time limit. A critical problem is how to obtain an accurate enough estimation of available bandwidth while not wasting too much time in bandwidth testing. In this paper, we present an algorithm to determine optimal amount of bandwidth testing given a probabilistic confidence level for network-aware multimedia object retrieval applications. The model treats the bandwidth testing as sampling from an actual bandwidth population. It uses statistical estimation method to quantify the benefit of each new bandwidth-testing sample, which is used to determine the optimal amount of bandwidth testing by balancing the benefit with the cost of each sample. Our implementation and experiments shows the algorithm determines the optimal amount of bandwidth testing effectively with minimum computation overhead. Yinzhe Yu, Irene Cheng 0001, Anup Basu |
IEEE Trans. Multim. | 2 |
| 2001 | QoS based video delivery with foveationabstractSpatially varying sensing (foveation) was first used as a means for image compression in our past research. We extend previous work to address the advantages of foveation in improving the performance of MPEG compression over bandwidth limited channels, such as the Internet. Unlike other approaches to foveating MPEG which used multiresolution representations, we use continuously spatially varying resolution and demonstrate that this approach is indeed advantageous over others. Two parameters, scaling and distortion, are used to allow us to adapt MPEG video to various compression ratios. Experimental results are presented, and can be viewed on the Web, validating our approach. Irene Cheng 0001, Anup Basu |
ICIP (1) | 1 |
| 2000 | DISIMA: An Object-Oriented Approach to Developing an Image Database SystemabstractMost image database prototypes and products focus mainly on similarity searches over syntactic features of images. DISIMA aims at providing querying on both syntactic and semantic features of images. The content of an image is viewed as a set of salient objects (regions of interest). Salient objects are organized into two levels: physical salient objects that store syntactic features and logical salient objects that give the semantics. DISIMA integrates a declarative query language (MOQL) and a visual query language (VisualMOQL). Vincent Oria, M. Tamer Özsu, Paul Iglinski, Irene Cheng 0001 |
ICDE | 5 |