Sen-Ching S. Cheung

dblp:c/SenChingSCheung · also Sen-Ching Samson Cheung, Sen-ching Samson Cheung · DBLP profile ↗
← Back
74ranked-venue papers
14as first author
13since 2021 · last 2026
0000-0002-9207-5514ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 57 · 11 first-author · 8 since 2021Artificial intelligence and machine learning · 9 · 4 since 2021Security and privacy · 6 · 3 first-authorSystems, architecture and hardware · 4 · 2 since 2021Computer networks · 2Applied, interdisciplinary, general and emerging computing · 2
YearPublicationVenuePosition
2026 Seeing Is Believing: Rich-Context Hallucination Detection for MLLMs via Backward Visual Grounding
abstract
Multimodal Large Language Models (MLLMs) have unlocked powerful cross-modal capabilities, but still significantly suffer from hallucinations. As such, accurate detection of hallucinations in MLLMs is imperative for ensuring their reliability in practical applications. To this end, guided by the principle of “Seeing is Believing”, we introduce VBackChecker, a novel reference-free hallucination detection framework that verifies the consistency of MLLM-generated responses with visual inputs, by leveraging a pixel-level Grounding LLM equipped with reasoning and referring segmentation capabilities. This referencefree framework not only effectively handles rich-context scenarios, but also offers interpretability. To facilitate this, an innovative pipeline is accordingly designed for generating instruction-tuning data (R-Instruct), featuring richcontext descriptions, grounding masks, and hard negative samples. We further establish R 2 -HalBench, a new hallucination benchmark for MLLMs, which, unlike previous benchmarks, encompasses real-world, rich-context descriptions from 18 MLLMs with high-quality annotations, spanning diverse object-, attribute-, and relationship-level details. VBackChecker outperforms prior complex frameworks and achieves state-of-the-art performance on R^2 -HalBench, even rivaling GPT-4o’s capabilities in hallucination detection. It also surpasses prior methods in the pixel-level grounding task, achieving over a 10% improvement.
Pinxue Guo, Chongruo Wu, Xinyu Zhou 0006, Lingyi Hong, Zhaoyu Chen 0001, Kaixun Jiang, Sen-Ching S. Cheung, Wei Zhang 0016
AAAI8
2026 Empowering Source-Free Domain Adaptation via MLLM-Guided Reliability-Based Curriculum Learning
Dongjie Chen, Kartik Patwari, Zhengfeng Lai, Xiaoguang Zhu, Sen-Ching S. Cheung, Chen-Nee Chuah
WACV5
2026 SCIA-GAN: Robust image watermarking via spatial-channel interaction attention and feature preservation
Lingchen Gu, Jun Wang 0061, Wenbo Wan, Jiande Sun 0001, Sen-Ching S. Cheung
Expert Syst. Appl.7
2025 A Warmer Start to Active Learning with Adaptive Gaussian Mixture Models for Skin Lesion Segmentation
abstract
Active learning is a promising strategy for reducing annotation burdens in medical image segmentation, particularly for tasks like skin lesion segmentation, where expert annotations are costly and time-intensive. However, existing methods suffer from cold-start issues and inefficient sample selection. This paper introduces a novel active learning framework called Task-Aligned Iterative Active Learning (TAIAL) that employs clustering and entropy ranking on a progressively refined feature space to select active samples that balance diversity, informativeness, and uncertainty. Coupled with a self-supervised initialization step, TAIAL provides an effective solution for both the cold-start problem and sample selection. Extensive experiments on the ISIC17 dataset demonstrate that TAIAL achieves early-stage sample selection performance, representing a 32% improvement over random sampling and an average improvement of 27% over other active learning schemes. In the later stage, it reaches 98.7% of fully supervised performance with only 38.4% labeled data, outperforming baseline methods. Our approach provides a scalable and efficient active learning paradigm for annotation-constrained medical imaging applications.
Lakmali Nadeesha Kumari, Chanaka Thushitha Bandara, Chen-Nee Chuah, Sen-Ching S. Cheung
ICIP4
2025 Dual Prototypes-Based Personalized Federated Adversarial Cross-Modal Hashing
abstract
With the rapid advances in wireless communication and IoT platforms, it is increasingly difficult to analyze relevant multi-modal data distributed across geographically diverse and heterogeneous platforms. One promising approach is to rely on federated learning to build compact cross-modal hash codes. However, existing federated learning methods easily exhibit degenerative performance in the global model due to the distributed data being derived from diverse domains. In addition, directly forcing each client to adopt the same global parameters as local parameters, without effective local training, significantly reduces the performance of each client. To overcome these challenges, we propose a novel federated adversarial cross-modal hashing, called Dual Prototypes-based personalized Federated Adversarial (DP-FeAd), which provides iterated training of shared dual prototypes. Specifically, aiming to expand local hashing models beyond their knowledge realms, DP-FeAd enables participating clients to engage in cooperative learning through two constructions: cluster prototypes and unbiased prototypes, instead of the traditional global prototypes, ensuring both generalization and stability. Specifically, the cluster prototypes are derived from local class-level prototypes and adversarially trained with local approximate hash codes to align their distributions. The unbiased prototypes are averaged from cluster prototypes and integrated into the training of local hashing models to maintain consistency across different local class-level prototypes further. The experiments conducted on two benchmark datasets demonstrate that our proposed method significantly enhances the performance of deep cross-modal hashing models in both IID (Independent and Identically Distributed) and non-IID scenarios.
Lingchen Gu, Xiaojuan Shen, Jiande Sun 0001, Jing Li 0046, Zhihui Li 0001, Sen-Ching S. Cheung, Wenbo Wan
IEEE Trans. Circuits Syst. Video Technol.7
2025 Attention Guidance by Cross-Domain Supervision Signals for Scene Text Recognition
abstract
Despite recent advances, scene text recognition remains a challenging problem due to the significant variability, irregularity and distortion in text appearance and localization. Attention-based methods have become the mainstream due to their superior vocabulary learning and observation ability. Nonetheless, they are susceptible to attention drift which can lead to word recognition errors. Most works focus on correcting attention drift in decoding but completely ignore the error accumulated during the encoding process. In this paper, we propose a novel scheme, called the Attention Guidance by Cross-Domain Supervision Signals for Scene Text Recognition (ACDS-STR), which can mitigate the attention drift at the feature encoding stage. At the heart of the proposed scheme is the cross-domain attention guidance and feature encoding fusion module (CAFM) that uses the core areas of characters to recursively guide attention to learn in the encoding process. With precise attention information sourced from CAFM, we propose a non-attention-based adaptive transformation decoder (ATD) to guarantee decoding performance and improve decoding speed. In the training stage, we fuse manual guidance and subjective learning to learn the core areas of characters, which notably augments the recognition performance of the model. Experiments are conducted on public benchmarks and show the state-of-the-art performance. The source will be available at https://github.com/xuefanfu/ACDS-STR.
Fanfu Xue, Jiande Sun 0001, Yaqi Xue, Qiang Wu 0009, Lei Zhu 0002, Xiaojun Chang, Sen-Ching S. Cheung
IEEE Trans. Image Process.7
2024 Localizing Moments of Actions in Untrimmed Videos of Infants with Autism Spectrum Disorder
abstract
Autism Spectrum Disorder (ASD) presents significant challenges in early diagnosis and intervention, impacting children and their families. With prevalence rates rising, there is a critical need for accessible and efficient screening tools. Leveraging machine learning (ML) techniques, in particular Temporal Action Localization (TAL), holds promise for automating ASD screening. This paper introduces a self-attention based TAL model designed to identify ASD-related behaviors in infant videos. Unlike existing methods, our approach simplifies complex modeling and emphasizes efficiency, which is essential for practical deployment in real-world scenarios. Importantly, this work underscores the importance of developing computer vision methods capable of operating in naturilistic environments with little equipment control, addressing key challenges in ASD screening. This study is the first to conduct end-to-end temporal action localization in untrimmed videos of infants with ASD, offering promising avenues for early intervention and support. We report baseline results of behavior detection using our TAL model. We achieve 70% accuracy for look face, 79% accuracy for look object, 72% for smile and 65% for vocalization.
Halil Ismail Helvaci, Chen-Nee Chuah, Sally Ozonoff, Sen-Ching S. Cheung
ICIP4
2023 He-Gan: Differentially Private Gan Using Hamiltonian Monte Carlo Based Exponential Mechanism
abstract
Differentially-private (DP) Generative Adversarial Networks (GAN) can be used to protect the privacy of training data and support public downstream learning tasks with synthetic data. However, typical DP mechanisms add noise to the training process and can lead to various convergence problems. We propose HE-GAN, a DP generative framework that eliminates noise addition by using Exponential Mechanism (EM) on the privacy-factor-adjusted posterior predictive distribution of a classifier trained on the private data. EM is more general than many other DP mechanisms including Laplacian and Gaussian mechanisms. EM’s reliance on sampling the output space also prevents the DP noise from corrupting the training process. However, there are two challenges: first, sampling the posterior distribution of the private discriminative classifier may not be able to produce high-quality synthetic samples. Instead, we sample from the latent space of a publicly-trained GAN to optimize the private posterior. Second, we use the highly effective Hamiltonian Monte Carlo (HMC) method for latent space sampling. We perform experiments on MNIST and Fashion-MNIST under public-private splits. Results show that HE-GAN can achieve downstream classification accuracy on par with or better than state-of-the-art scheme over a wide range of privacy budgets.
Usman Hassan, Dongjie Chen, Sen-Ching S. Cheung, Chen-Nee Chuah
ICASSP3
2022 Smoothed Adaptive Weighting for Imbalanced Semi-Supervised Learning: Improve Reliability Against Unknown Distribution Data
abstract
Despite recent promising results on semi-supervised learning (SSL), data imbalance, particularly in the unlabeled dataset, could significantly impact the training performance of a SSL algorithm if there is a mismatch between the expected and actual class distributions. The efforts on how to construct a robust SSL framework that can effectively learn from datasets with unknown distributions remain limited. We first investigate the feasibility of adding weights to the consistency loss and then we verify the necessity of smoothed weighting schemes. Based on this study, we propose a self-adaptive algorithm, named Smoothed Adaptive Weighting (SAW). SAW is designed to enhance the robustness of SSL by estimating the learning difficulty of each class and synthesizing the weights in the consistency loss based on such estimation. We show that SAW can complement recent consistency-based SSL algorithms and improve their reliability on various datasets including three standard datasets and one gigapixel medical imaging application without making any assumptions about the distribution of the unlabeled set.
Zhengfeng Lai, Chao Wang 0067, Henrry Gunawan, Sen-Ching S. Cheung, Chen-Nee Chuah
ICML4
2021 Privacy-Protected Denoising for Signals on Graphs from Distributed Systems
abstract
The fast-growing networked computing devices create many distributed systems and generate new signals on a large scale. Typical applications include peer-to-peer streaming of multimedia data, crowd- sourcing, and measurement by sensor networks. Therefore, the massive amount of networked data is a form of big data, calling for new data structures and algorithms different from classical ones suitable for small data sizes. We consider a vital data format for recording information from networked distributed systems: signals on graphs. A significant concern is to protect the privacy of large scales of signals when processed at third parties, such as cloud data centers. A de-facto solution is to outsource encrypted data before they arrive at the third-parties. We propose a novel and efficient privacy-protected outsourced denoising algorithm based on the information-theoretic secure multi-party computation (secure MPC). Among the operations of signals on graphs, denoising is useful before further meaningful processing can occur. We experiment with our algorithms in a popular platform of secure MPC and compare it with Paillier's homomorphic encryption approach. The results demonstrate a better efficiency of our approach.
Zhaohong Wang, Sen-Ching S. Cheung
ISCAS2
2021 Augmented Reality Circuit Learning
abstract
Building electronic circuits is one of the most common hands-on activities in learning STEM subjects. Beginning students often find it difficult to translate a circuit schematic into the construction of the physical circuit on a breadboard. In this paper, we describe an Augmented Reality circuit learning software that can provide students with step-by-step instructions by placing virtual circuit components on a physical breadboard. We propose a novel image processing pipeline that can robustly identify the planar structure of a breadboard, even in the presence of occluding circuit components on the breadboard. Using a commercially available library of 3D circuit component models, the estimated 3D structure of the breadboard allows us to render arbitrary circuit components on it in real time. Experimental results demonstrate that our algorithms are accurate and can produce realistic-looking virtual circuits.
Hao Wang 0183, Sen-Ching S. Cheung
ISCAS2
2021 Enhanced 3D Human Pose Estimation from Videos by Using Attention-Based Neural Network with Dilated Convolutions
Ruixu Liu, Ju Shen, Chen Chen 0001, Sen-Ching S. Cheung, Vijayan K. Asari
Int. J. Comput. Vis.5
2021 Predicting ASD diagnosis in children with synthetic and image-based eye gaze data
Sidrah Liaqat, Chongruo Wu, Prashanth Reddy Duggirala, Sen-Ching S. Cheung, Chen-Nee Chuah, Sally Ozonoff, Gregory Young
Signal Process. Image Commun.4
2020 Attention Mechanism Exploits Temporal Contexts: Real-Time 3D Human Pose Reconstruction
abstract
We propose a novel attention-based framework for 3D human pose estimation from a monocular video. Despite the general success of end-to-end deep learning paradigms, our approach is based on two key observations: (1) temporal incoherence and jitter are often yielded from a single frame prediction; (2) error rate can be remarkably reduced by increasing the receptive field in a video. Therefore, we design an attentional mechanism to adaptively identify significant frames and tensor outputs from each deep neural net layer, leading to a more optimal estimation. To achieve large temporal receptive fields, multi-scale dilated convolutions are employed to model long-range dependencies among frames. The architecture is straightforward to implement and can be flexibly adopted for real-time applications. Any off-the-shelf 2D pose estimation system, e.g. Mocap libraries, can be easily integrated in an ad-hoc fashion. We both quantitatively and qualitatively evaluate our method on various standard benchmark datasets (e.g. Human3.6M, HumanEva). Our method considerably outperforms all the state-of-the-art algorithms up to 8% error reduction (average mean per joint position error: 34.7) as compared to the best-reported results. Code is available at: (https://github.com/lrxjason/Attention3DHumanPose)
Ruixu Liu, Ju Shen, Chen Chen 0001, Sen-Ching S. Cheung, Vijayan K. Asari
CVPR5
2020 Machine Learning Based Autism Spectrum Disorder Detection from Videos
abstract
Early diagnosis of Autism Spectrum Disorder (ASD) is crucial for best outcomes to interventions. In this paper, we present a machine learning (ML) approach to ASD diagnosis based on identifying specific behaviors from videos of infants of ages 6 through 36 months. The behaviors of interest include directed gaze towards faces or objects of interest, positive affect, and vocalization. The dataset consists of 2000 videos of 3-minute duration with these behaviors manually coded by expert raters. Moreover, the dataset has statistical features including duration and frequency of the above mentioned behaviors in the video collection as well as independent ASD diagnosis by clinicians. We tackle the ML problem in a two-stage approach. Firstly, we develop deep learning models for automatic identification of clinically relevant behaviors exhibited by infants in a one-on-one interaction setting with parents or expert clinicians. We report baseline results of behavior classification using two methods: (1) image based model (2) facial behavior features based model. We achieve 70% accuracy for smile, 68% accuracy for look face, 67% for look object and 53% accuracy for vocalization. Secondly, we focus on ASD diagnosis prediction by applying a feature selection process to identify the most significant statistical behavioral features and a over and under sampling process to mitigate the class imbalance, followed by developing a baseline ML classifier to achieve an accuracy of 82% for ASD diagnosis.
Chongruo Wu, Sidrah Liaqat, Halil Ismail Helvaci, Sen-Ching S. Cheung, Chen-Nee Chuah, Sally Ozonoff, Gregory Young
HealthCom4
2019 Quannet: Joint Image Compression and Classification Over Channels with Limited Bandwidth
abstract
The performance of cloud based image classification depends critically on its allocated bandwidth. Traditional data compression methods can negatively impact classification accuracy under limited bandwidth. We investigate the design of bandwidth efficient quantization for image encoding and compression with minimum classification accuracy loss. This work develops a simple neural network framework for joint quantization and classification. The proposed 'QuanNet' can optimize the quantization intervals of JPEG2000 encoder to minimize the classification loss. We show that our quantizer optimization can achieve significant accuracy improvement for a given channel bandwidth. Similarly, significant bandwidth can be saved to achieve a desired accuracy for cloud based image classification.
Lahiru D. Chamain, Sen-Ching S. Cheung, Zhi Ding 0001
ICME2
2019 Motion and appearance based background subtraction for freely moving cameras
Hasan Sajid, Sen-Ching S. Cheung, Nathan Jacobs
Signal Process. Image Commun.2
2019 Human body reshaping and its application using multiple RGB-D sensors
Wanxin Xu, Po-Chang Su, Sen-Ching S. Cheung
Signal Process. Image Commun.3
2019 People Counting in Dense Crowd Images Using Sparse Head Detections
abstract
People counting in extremely dense crowds is a challenging problem due to severe occlusions, few pixels per head, cluttered environments, and skewed camera perspectives. In this paper, we present a novel algorithm for people counting in highly dense crowd images. Our approach relies on the fact that the head is the most visible part of an individual in a dense crowd. As such, a head detector can be used to estimate the spatially varying head size, which is the key feature used in our head counting procedure. We leverage the state-of-the art convolutional neural network for the sparse head detection in a dense crowd. After sub-dividing the image into rectangular patches, we first use an speeded-up robust features-based support vector machine binary classifier to label each patch as crowd/not-crowd and eliminate all not-crowd patches. Regression is then performed on each crowd patch to estimate average head size. The number of individuals in each patch is estimated by dividing the patch area with the estimated head size. For the crowd patches where no heads are detected, the counts are estimated based on distance-based weighted averaging over the counts from neighboring patches. Finally, the individual patch counts are summed up to obtain the total count. We evaluate our approach on three publicly available datasets for extremely dense crowds: UCF_CC_50, ShanghaiTech, and AHU-Crowd. Our approach gives comparable results on these challenging datasets to other state of the art algorithms but, unlike other algorithms, our proposed method does not require the laborious task of obtaining labeled training data of dense crowd images.
Mamoona Birkhez Shami, Salman Maqbool, Hasan Sajid, Yasar Ayaz, Sen-Ching S. Cheung
IEEE Trans. Circuits Syst. Video Technol.5
2018 Learning Sensitive Images Using Generative Models
abstract
The sheer amount of personal data being transmitted to cloud services and the ubiquity of cellphones cameras and various sensors, have provoked a privacy concern among many people. On the other hand, the recent phenomenal growth of deep learning that brings advancements in almost every aspect of human life is heavily dependent on the access to data, including sensitive images, medical records, etc. Therefore, there is a need for a mechanism that transforms sensitive data in such a way as to preserves the privacy of individuals, yet still be useful for deep learning algorithms. This paper proposes the use of Generative Adversarial Networks (GANs) as one such mechanism, and through experimental results, shows its efficacy.
Sen-Ching S. Cheung, Herb Wildfeuer, Mehdi Nikkhah, Wai-tian Tan
ICIP1
2017 Information-Theoretic Secure Multi-Party Computation With Collusion Deterrence
abstract
Secure multi-party computation (MPC) has been established as the de facto paradigm for protecting privacy in distributed computation. Among many secure MPC primitives, Shamir's secret sharing (SSS) has the advantages of having low complexity and information-theoretic security. However, SSS requires multiple honest participants and is susceptible to collusion attacks. In this paper, we provide a detailed analysis of different types of collusion attacks and propose novel mechanisms to deter such attacks in a fully distributed manner. Focusing on outsourced computing environments where secret data owners can collaborate on a public computing platform, we study collusion attacks using game theory. For those attacks where the thefts are detectable, we show that they can be effectively deterred by an explicit retaliation mechanism between data owners. The result is based on a comprehensive analysis that takes into account the cost of collusion, the privacy preference, and the associated uncertainty. For those attacks where the thefts cannot be detected, we expand the analysis to include the computing platform and provide deterrence through deceptive collusion requests as well as a novel cryptographic censorship protocol. The correctness and the privacy of the protocols are proved under the rational adversarial model. Our SSS-based protocols are shown to outperform the state-of-the-art garbled circuit systems, while our simulation results validate the proposed mechanism designs in deterring collusion.
Zhaohong Wang, Sen-Ching S. Cheung, Ying Luo 0008
IEEE Trans. Inf. Forensics Secur.2
2017 Universal Multimode Background Subtraction
abstract
In this paper, we present a complete change detection system named multimode background subtraction. The universal nature of system allows it to robustly handle multitude of challenges associated with video change detection, such as illumination changes, dynamic background, camera jitter, and moving camera. The system comprises multiple innovative mechanisms in background modeling, model update, pixel classification, and the use of multiple color spaces. The system first creates multiple background models of the scene followed by an initial foreground/background probability estimation for each pixel. Next, the image pixels are merged together to form mega-pixels, which are used to spatially denoise the initial probability estimates to generate binary masks for both RGB and YCbCr color spaces. The masks generated after processing these input images are then combined to separate foreground pixels from the background. Comprehensive evaluation of the proposed approach on publicly available test sequences from the CDnet and the ESI data sets shows superiority in the performance of our system over other state-of-the-art algorithms.
Hasan Sajid, Sen-Ching S. Cheung
IEEE Trans. Image Process.2
2016 On privacy preference in collusion-deterrence games for secure multi-party computation
abstract
Secure multi-party computation (MPC) has been established as the de facto paradigm for protecting privacy in distributed computation. Information-theoretic secure MPC protocols, though more efficient than their computationally secure counterparts, require at least three computational parties and are prone to collusion attacks. Previous work has used mechanism designs to deter collusion. An important element missing is the consideration of how different players value privacy. In this paper, we provide a detailed analysis of possible outcomes under different privacy preferences based on the relative cost of collusion attacks over loss of privacy. We explicitly calculate the conditions under which honesty is the solution. Simulation results provide further evidence to demonstrate the validity of our mechanism design.
Zhaohong Wang, Sen-Ching S. Cheung
ICASSP2
2016 Human pose estimation using two RGB-D sensors
abstract
Accurate human pose estimation plays an important role in various applications such as sports analysis, health care and gaming. Even though recent approaches have shown that 3D positions of body joints can be estimated from a single depth sensor, the depth data often suffer from sensing noise and self-occlusion. In this paper, we present a novel pipeline to estimate the pose of a human body by using two depth sensors. The two sensors simultaneously capture the front and back of the body's movement. Using a wide-baseline RGB-D camera calibration algorithm, the two 3D scans are first geometrically aligned, and then registered to a generic human template using a Gaussian-mixture-model based point set registration procedure with local structure constraints. The new pose of person is finally estimated by a rigid bone-based pose transformation. Experimental results demonstrate the effectiveness of our system in estimating the body pose over other state-of-the-arts techniques.
Wanxin Xu, Po-Chang Su, Sen-Ching S. Cheung
ICIP3
2016 RoboMirror: Simulating a mirror with a robotic camera
abstract
Simulated mirror display systems (SMDs) provide augmented rendering of mirror images. Many SMDs use multiple static cameras to create viewpoint-dependent rendering. Unfortunately, the quality of the rendering around the face of the viewer is typically poor due to visual distortion from warping and camera view misalignment. We propose the RoboMirror SMD to provide high-quality mirror rendering of the frontal face by using a robotic camera to track the viewer's head. An inclined one-way mirror allows the viewer's face to be captured by the robotic camera, and reflects the output from a projector to the viewer. Novel mirror calibration and depth-based image warping are proposed to produce an accurate mirror rendering of all objects based on the perspective of the viewer. To compensate for the blockage of light by the oneway mirror, a channel-equalization based contrast enhancement scheme is proposed that outperforms other image enhancement schemes for this system.
Nkiruka Uzuegbunam, Wanxin Xu, Sen-Ching S. Cheung
ICIP4
2016 Appearance based background subtraction for PTZ cameras
Hasan Sajid, Sen-Ching S. Cheung, Nathan Jacobs
Signal Process. Image Commun.2
2015 Background subtraction for static & moving camera
abstract
Background subtraction is one of the most commonly used components in machine vision systems. Despite the numerous algorithms proposed in the literature and used in practical applications, key challenges remain in designing a single system that can handle diverse environmental conditions. In this paper we present Multiple Background Model based Background Subtraction Algorithm as such a candidate. The algorithm was originally designed for handling sudden illumination changes. The new version has been refined with changes at different steps of the process, specifically in terms of selecting optimal color space, clustering of training images for Background Model Bank and parameter for each channel of color space. This has allowed the algorithm's applicability to wide variety of challenges associated with change detection including camera jitter, dynamic background, Intermittent Object Motion, shadows, bad weather, thermal, night videos etc. Comprehensive evaluation demonstrates the superiority of algorithm against state of the art.
Hasan Sajid, Sen-Ching S. Cheung
ICIP2
2015 Affect-preserving privacy protection of video
abstract
The prevalence of wireless networks and the convenience of mobile cameras enable many new video applications other than security and entertainment. From behavioral diagnosis to wellness monitoring, cameras are increasing used for observations in various educational and medical settings. Videos collected for such applications are considered protected health information under privacy laws in many countries. At the same time, there is an increasing need to share such video data across a wide spectrum of stakeholders including professionals, therapists and families facing similar challenges. Visual privacy protection techniques, such as blurring or object removal, can be used to mitigate privacy concern, but they also obliterate important visual cues of affect and social behaviors that are crucial for the target applications. In this paper, we propose a method of manipulating facial expression and body shape to conceal the identity of individuals while preserving the underlying affect states. The experiment results demonstrate the effectiveness of our method.
Wanxin Xu, Sen-Ching S. Cheung, Neelkamal Soares
ICIP2
2015 MEBook: Kinect-based self-modeling intervention for children with autism
abstract
Autism spectrum disorder (ASD) is a chronic developmental disorder that impairs the development of social and communication skills. Multiple studies have shown that children with ASD prefer images of self over others. These studies explain the effectiveness of video self-modeling (VSM), an evidence-based autism intervention in which one learns by watching oneself performing a target behavior in video. VSM content is difficult to create as target behaviors are sporadic, but advances in sensing and graphics enable synthesis of such behaviors. In this paper, we propose the MEBook system which uses Kinect sensor to inject self-images into a social narrative game to teach students with ASD proper greeting behaviors. The social narrative is an animated story about the main character meeting and greeting different cartoon characters in a clinic. Self-modeling is achieved by first replacing the main characters face with an image of the subject, and then animating the subject to match the narration. The second component is a positive reinforcement game in which the subject is prompted to greet different cartoon characters. Through depth-based body posture tracking, proper greeting behaviors are recognized and immediately rewarded with praises and visual confetti. A multiple-baseline single-subject study has been conducted and the preliminary results show that MEBook is effective in teaching greeting behaviors to children with ASD.
Nkiruka Uzuegbunam, Wing-Hang Wong, Sen-Ching S. Cheung, Lisa Ruble
ICME3
2015 Automatic video self modeling for voice disorder
Ju Shen, Changpeng Ti, Anusha Raghunathan, Sen-Ching S. Cheung, Rita R. Patel
Multim. Tools Appl.4
2014 Efficient multi-party computation with collusion-deterred secret sharing
abstract
Many secure multiparty computation (SMC) protocols use Shamir's Secret Sharing (SSS) scheme as a building block. Unlike other cryptographic SMC techniques such as garbled circuits (GC), SSS requires no data expansion and achieves information theoretic security. A weakness of SSS is the possibility of collusion attacks from participants. In this paper, we propose an evolutionary game-theoretic (EGT) approach to deter collusion in SSS-based protocols. First, we consider the possibility of detecting the leak of secret data caused by collusion, devise an explicit retaliation mechanism, and show that the evolutionary stable strategy of this game is not to collude if the technology to detect the leakage of secret is readily available. Then, we consider the situation in which data-owners are unaware of the leakage and thereby unable to retaliate. Such behaviors are deterred by injecting occasional fake collusion requests, and detected by a censorship scheme that destroys subliminal communication. Comparison results show that our collusion-deterred SSS system significantly outperforms GC, while game simulations confirm the validity of our EGT framework on modeling collusion behaviors.
Zhaohong Wang, Ying Luo 0008, Sen-Ching S. Cheung
ICASSP3
2014 Background subtraction under sudden illumination change
abstract
In this paper, we propose a Multiple Background Model based Background Subtraction (MB2S) algorithm that is robust against sudden illumination changes in indoor environment. It uses multiple background models of expected illumination changes followed by both pixel and frame based background subtraction on both RGB and YCbCr color spaces. The masks generated after processing these input images are then combined in a framework to classify background and foreground pixels. Evaluation of proposed approach on publicly available test sequences show higher precision and recall than other state-of-the-art algorithms.
Hasan Sajid, Sen-Ching S. Cheung
MMSP2
2014 Extrinsic calibration for wide-baseline RGB-D camera network
abstract
In the recent years, color and depth camera systems have attracted intensive attention because of its wide applications in image-based rendering, 3D model reconstruction, and human tracking and pose estimation. These applications often require multiple color and depth cameras to be placed with wide separation so as to capture the scene objects from different prospectives. The difference in modality and the wide baseline make calibration a challenging problem. In this paper, we present an algorithm that simultaneously and automatically calibrates the extrinsics across multiple color and depth cameras across the network. Rather than using the standard checkerboard, we use a sphere as a calibration object to identify the correspondences across different views. We experimentally demonstrate that our calibration framework can seamlessly integrate different views with wide baselines that outperforms other techniques in the literature.
Ju Shen, Wanxin Xu, Ying Luo 0008, Po-Chang Su, Sen-Ching S. Cheung
MMSP5
2014 Human segmentation by geometrically fusing visible-light and thermal imageries
Jian Zhao 0003, Sen-Ching S. Cheung
Multim. Tools Appl.2
2013 Layer Depth Denoising and Completion for Structured-Light RGB-D Cameras
abstract
The recent popularity of structured-light depth sensors has enabled many new applications from gesture-based user interface to 3D reconstructions. The quality of the depth measurements of these systems, however, is far from perfect. Some depth values can have significant errors, while others can be missing altogether. The uncertainty in depth measurements among these sensors can significantly degrade the performance of any subsequent vision processing. In this paper, we propose a novel probabilistic model to capture various types of uncertainties in the depth measurement process among structured-light systems. The key to our model is the use of depth layers to account for the differences between foreground objects and background scene, the missing depth value phenomenon, and the correlation between color and depth channels. The depth layer labeling is solved as a maximum a-posteriori estimation problem, and a Markov Random Field attuned to the uncertainty in measurements is used to spatially smooth the labeling process. Using the depth-layer labels, we propose a depth correction and completion algorithm that outperforms other techniques in the literature.
Ju Shen, Sen-Ching S. Cheung
CVPR2
2013 A robust RGB-D SLAM system for 3D environment with planar surfaces
abstract
With the increasing popularity of RGB-depth (RGB-D) sensors such as the Microsoft Kinect, there have been much research on capturing and reconstructing 3D environments using a movable RGB-D sensor. The key process behind these kinds of simultaneous location and mapping (SLAM) systems is the iterative closest point or ICP algorithm, which is an iterative algorithm that can estimate the rigid movement of the camera based on the captured 3D point clouds. While ICP is a well-studied algorithm, it is problematic when it is used in scanning large planar regions such as wall surfaces in a room. The lack of depth variations on planar surfaces makes the global alignment an ill-conditioned problem. In this paper, we present a novel approach for registering 3D point clouds by combining both color and depth information. Instead of directly searching for point correspondences among 3D data, the proposed method first extracts features from the RGB images, and then back-projects the features to the 3D space to identify more reliable correspondences. These color correspondences form the initial input to the ICP procedure which then proceeds to refine the alignment. Experimental results show that our proposed approach can achieve better accuracy than existing SLAMs in reconstructing indoor environments with large planar surfaces.
Po-Chang Su, Ju Shen, Sen-Ching S. Cheung
ICIP3
2013 Guest Editorial: Special issue on privacy and trust management in cloud and distributed systems
abstract
The 13 papers in this special issue cover three major areas including privacy enhanced technology, trust and reputation, as well as applications in cloud computing environments.
Sen-Ching S. Cheung, Karl Aberer, Jayant R. Haritsa, Bill G. Horne, Kai Hwang 0001
IEEE Trans. Inf. Forensics Secur.1
2013 Virtual Mirror Rendering With Stationary RGB-D Cameras and Stored 3-D Background
abstract
Mirrors are indispensable objects in our lives. The capability of simulating a mirror on a computer display, augmented with virtual scenes and objects, opens the door to many interesting and useful applications from fashion design to medical interventions. Realistic simulation of a mirror is challenging as it requires accurate viewpoint tracking and rendering, wide-angle viewing of the environment, as well as real-time performance to provide immediate visual feedback. In this paper, we propose a virtual mirror rendering system using a network of commodity structured-light RGB-D cameras. The depth information provided by the RGB-D cameras can be used to track the viewpoint and render the scene from different prospectives. Missing and erroneous depth measurements are common problems with structured-light cameras. A novel depth denoising and completion algorithm is proposed in which the noise removal and interpolation procedures are guided by the foreground/background label at each pixel. The foreground/background label is estimated using a probabilistic graphical model that considers color, depth, background modeling, depth noise modeling, and spatial constraints. The wide viewing angle of the mirror system is realized by combining the dynamic scene, captured by the static camera network with a 3-D background model created off-line, using a color-depth sequence captured by a movable RGB-D camera. To ensure a real-time response, a scalable client-and-server architecture is used with the 3-D point cloud processing, the viewpoint estimate, and the mirror image rendering are all done on the client side. The mirror image and the viewpoint estimate are then sent to the server for final mirror view synthesis and viewpoint refinement. Experimental results are presented to show the accuracy and effectiveness of each component and the entire system.
Ju Shen, Po-Chang Su, Sen-Ching S. Cheung, Jian Zhao 0003
IEEE Trans. Image Process.3
2012 Automatic lip-synchronized video-self-modeling intervention for voice disorders
abstract
Video self-modeling (VSM) is a behavioral intervention technique in which a learner models a target behavior by watching a video of him- or herself. In the field of speech language pathology, the approach of VSM has been successfully used for treatment of language in children with Autism and in individuals with fluency disorder of stuttering. Technical challenges remain in creating VSM contents that depict previously unseen behaviors. In this paper, we propose a novel system that synthesizes new video sequences for VSM treatment of patients with voice disorders. Starting with a video recording of a voice-disorder patient, the proposed system replaces the coarse speech with a clean, healthier speech that bears resemblance to the patient's original voice. The replacement speech is synthesized using either a text-to-speech engine or selecting from a database of clean speeches based on a voice similarity metric. To realign the replacement speech with the original video, a novel audiovisual algorithm that combines audio segmentation with lip-state detection is proposed to identify corresponding time markers in the audio and video tracks. Lip synchronization is then accomplished by using an adaptive video re-sampling scheme that minimizes the amount of motion jitter and preserves the spatial sharpness. Experimental evaluations on a dataset with 31 subjects demonstrate the effectiveness of the proposed techniques.
Ju Shen, Changpeng Ti, Sen-Ching S. Cheung, Rita R. Patel
Healthcom3
2012 An efficient protocol for private iris-code matching by means of garbled circuits
abstract
Biometric-based access control is receiving increasing attention due to its security and ease-of-use. However, concerns are often raised regarding the protection of the privacy of enrolled users. Signal processing in the encrypted domain has been proposed as a viable solution to protect biometric templates and the privacy of the users. In particular, several solutions have been proposed to protect the privacy of the biometric probe during the authentication process. In this paper we focus on privacy-preserving iris-based authentication. The main innovations compared to the prior art include: i) an iris masking technique that simplifies the operations on the encrypted data without sacrificing the recognition rate; ii) the adoption of a matching protocol based only on garbled circuits which offers longer term security over existing solutions based on homomorphic encryption or hybrid techniques. The computational and communication complexity of the on-line phase of the proposed protocol is extremely low, thus opening the way to its exploitation in practical applications.
Ying Luo 0008, Sen-Ching S. Cheung, Tommaso Pignata, Riccardo Lazzeretti, Mauro Barni
ICIP2
2012 Privacy protected image denoising with secret shares
abstract
The proliferation of digital cameras, wireless networks and distributed computing make sharing of visual data easier than ever. Such casual exchange of data, however, has increasingly raised questions on how sensitive visual information can be protected. Encrypted-domain signal processing techniques based on homomorphic encryption and garbled circuits are increasingly applied for such applications. Their high computation and communication complexity, however, are not suitable for pixel-level processing. In this paper, we propose an alternative approach of using information-theoretically secure protocols over multiple non-colluding semi-honest computing agents. The proposed protocols are based on classical Shamir's secret sharing scheme which supports multiplication and addition in the random-share domain. We extend the sharing scheme to handle other fundamental signal processing operations and use them to develop a novel privacy-protected wavelet denoising scheme over three computing agents. Our experimental results demonstrate the viability of using information-theoretic secure protocols to safeguard privacy in distributed pixel-level processing.
Sayed M. SaghaianNejadEsfahani, Ying Luo 0008, Sen-Ching S. Cheung
ICIP3
2011 Automatic content generation for video self modeling
abstract
Video self modeling (VSM) is a behavioral intervention technique in which a learner models a target behavior by watching a video of him or herself. Its effectiveness in rehabilitation and education has been repeatedly demonstrated but technical challenges remain in creating video contents that depict previously unseen behaviors. In this paper, we propose a novel system that re-renders new talking-head sequences suitable to be used for VSM treatment of patients with voice disorder. After the raw footage is captured, a new speech track is either synthesized using text-to-speech or selected based on voice similarity from a database of clean speeches. Voice conversion is then applied to match the new speech to the original voice. Time markers extracted from the original and new speech track are used to re-sample the video track for lip synchronization. We use an adaptive re-sampling strategy to minimize motion jitter, and apply bilinear and optical-flow based interpolation to ensure the image quality. Both objective measurements and subjective evaluations demonstrate the effectiveness of the proposed techniques.
Ju Shen, Anusha Raghunathan, Sen-Ching S. Cheung, Rita R. Patel
ICME3
2010 Eye tracking based perceptual image inpainting quality analysis
abstract
The objective of image inpainting is to perform a seamless completion of missing areas in images. Evaluating the perceptual quality of an inpainting algorithm must rely on features of the Human Visual System. Using eye-tracking experiments, we show that there is a strong correlation between inpainting quality and visual attention. By comparing gaze densities within and outside the hole regions of inpainted images, we show that discernible artifacts due to inpainting attract an unusual amount of visual attention. The gaze density within the hole, normalized with the gaze density of the same region from the unmodified image, provides a useful measure in comparing different inpainting processes and corroborates well with subjective rankings.
M. Vijay Venkatesh, Sen-Ching S. Cheung
ICIP2
2010 Anonymous subject identification in privacy-aware video surveillance
abstract
The widespread deployment of surveillance cameras has raised serious privacy concerns. Many privacy-enhancing schemes have been recently proposed to identify selected individuals and redact their images in the surveillance video. To identify individuals, the best known approach is to use biometric signals as they are immutable and highly discriminative. If misused, these characteristics of biometrics can seriously defeat the goal of privacy protection. In this paper, we propose an anonymous subject identification system based on homo-morphic encryption (HE). It matches the biometric signals in encrypted domain to provide anonymity to users. To make the HE-based protocols computationally scalable, we propose a complexity-privacy tradeoff called k-Anonymous Quantization (kAQ) which narrows the plaintext search to a small cell before running the intensive encrypted-domain processing within the cell. We validate a key assumption in kAQ that privacy is better preserved by grouping biometric patterns far apart into the same cell. We also improve the matching success rate by replacing the original bounding boxes with e-balls as basic units for grouping. Experimental results on a public iris biometric database demonstrate the validity of our framework.
Ying Luo 0008, Shuiming Ye, Sen-Ching S. Cheung
ICME3
2009 Manifold Estimation in View-Based Feature Space for Face Synthesis across Poses
Xinyu Huang 0001, Jizhou Gao, Sen-Ching S. Cheung, Ruigang Yang
ACCV (1)3
2009 Anonymous Biometric Access Control based on homomorphic encryption
abstract
In this paper, we consider the problem of incorporating privacy protection in a biometric access control system. An anonymous biometric access control (ABAC) system is proposed to verify the membership status of a user using biometric signals without knowing his or her true identity. This system is useful in protecting the privacy of authorized users while keeping out potential imposters and attackers. In our proposed system, we employ homomorphic encryption to protect the probe biometric and propose a secure similarity search (SSS) algorithm to authenticate the probe in an anonymous way. We demonstrate our proposed system using a gallery of iris biometrics and show the efficiency of the implemented anonymous biometric matching performed in encrypted domain.
Ying Luo 0008, Sen-Ching S. Cheung, Shuiming Ye
ICME2
2009 Audio-visual privacy protection for video conference
abstract
Group video-conferencing systems are routinely used in major corporations, hospitals and universities for meetings, tele-medicine and distance learning among participants from very distant locations. As the use of video-conferencing becomes widely prevalent, the privacy concern's raised by this technology becomes an important issue to be addressed. In this paper we propose a real-time privacy preserving video conferencing system which protects the visual and audio privacy of selected individuals. In our proposed system we differentiate between the general participants and private participants (PP) whose privacy needs to be protected. We further divide the private participants into two different categories and provide a varying level of privacy protection based on the requirements. Specifically, among private participants, we have Active Private Participants (APP) who interactively participate in the meeting and Passive Private Participants (PPP) who play a passive observatory role. The video and audio privacy of the APP are protected by obfuscating their visual information by simple black boxing and real-time pitch modification process respectively. For the PPP, we completely protect their privacy by continuously detecting their presence and erasing them with a real-time adaptive background replacement process.
M. Vijay Venkatesh, Jian Zhao 0003, L. Profitt, Sen-Ching S. Cheung
ICME4
2009 Optimal Visual Sensor Planning
abstract
Visual sensor networks are becoming more and more common. They have a wide-range of commercial and military applications from video surveillance to smart home and from traffic monitoring to anti-terrorism. The design of such a visual sensor network is a challenging problem due to the complexity of the environment, self and mutual occlusion of moving objects, diverse sensor properties and a myriad of performance metrics for different applications. As such, there is a need to develop a flexible sensor-planning framework that can incorporate all the aforementioned modeling details, and derive the sensor configuration that simultaneously optimizes the target performance and minimizes the cost. In this paper, we tackle this optimal sensor problem by developing a general visibility model for visual sensor networks and solving the optimization problem via Binary Integer Programming (BIP). Our proposed visibility model supports arbitrary-shaped 3D environments and incorporates realistic camera models, occupant traffic models, self occlusion and mutual occlusion. Using this visibility model, a novel BIP algorithms are proposed to find the optimal camera placement for tracking visual tags in multiple cameras. Experimental performance analysis is performed using Monte-Carlo simulations.
Jian Zhao 0003, Sen-Ching S. Cheung
ISCAS2
2009 Enhancing Privacy Protection in Multimedia Systems
Sen-Ching S. Cheung, Deepa Kundur, Andrew W. Senior
EURASIP J. Inf. Secur.1
2009 Video Data Hiding for Managing Privacy Information in Surveillance Systems
abstract
From copyright protection to error concealment, video data hiding has found usage in a great number of applications. In this work, we introduce the detailed framework of using data hiding for privacy information preservation in a video surveillance environment. To protect the privacy of individuals in a surveillance video, the images of selected individuals need to be erased, blurred, or re-rendered. Such video modifications, however, destroy the authenticity of the surveillance video. We propose a new rate-distortion-based compression-domain video data hiding algorithm for the purpose of storing that privacy information. Using this algorithm, we can safeguard the original video as we can reverse the modification process if proper authorization can be established. The proposed data hiding algorithm embeds the privacy information in optimal locations that minimize the perceptual distortion and bandwidth expansion due to the embedding of privacy data in the compressed domain. Both reversible and irreversible embedding techniques are considered within the proposed framework and extensive experiments are performed to demonstrate the effectiveness of the techniques.
Jithendra K. Paruchuri, Sen-Ching S. Cheung, Michael W. Hail
EURASIP J. Inf. Secur.2
2009 Anonymous Biometric Access Control
Shuiming Ye, Ying Luo 0008, Jian Zhao 0003, Sen-Ching S. Cheung
EURASIP J. Inf. Secur.4
2009 Efficient object-based video inpainting
M. Vijay Venkatesh, Sen-Ching S. Cheung, Jian Zhao 0003
Pattern Recognit. Lett.2
2008 Managing privacy data in pervasive camera networks
abstract
Privacy protection of visual information is increasingly important as pervasive camera networks becomes more prevalent. The proposed scheme addresses the problem of preserving and controlling of the privacy visual data through two innovations. First, unlike the existing centralized control of privacy data, the proposed system allows individual users to make the final decision on every access to their privacy data. As such, it offers a much stronger form of privacy protection as the user no longer needs to trust, adhere or register his/her privacy preferences with a server. The second innovation is the development of a secure reversible data hiding scheme for embedding all the ownership information and privacy data into the obfuscated video bitstream. Not only has it resulted in an efficient design of protocols, the reversible data hiding allows perfect reconstruction of original data and supports arbitrary types of video obfuscation techniques. Impact of data hiding on bitrate and distortion is minimized through a rate-distortion optimization procedure and experimental results are provided to demonstrate its efficiency.
Sen-Ching S. Cheung, Jithendra K. Paruchuri, Thinh P. Nguyen
ICIP1
2008 Joint optimization of data hiding and video compression
abstract
From copyright protection to error concealment, video data hiding has found usage in a great number of applications. Recently proposed applications such as privacy data preservation require huge amount of information to be hidden inside a compressed video bitstream. Since data hiding disturbs the underlying statistical patterns of the source data, it adversely affects the performance of compression which are designed based on the statistical properties of the data. As such, it is imperative to design a data hiding scheme that is compatible with the compression algorithm and at the same time, introduces as little perceptual distortion as possible. In this paper, we propose a novel compression-domain video data-hiding algorithm that determines the optimal embedding strategy to minimize both the output perceptual distortion and the output bit rate. The hidden data is embedded into selective discrete cosine transform (DCT) coefficients which are found in most video compression standards. The coefficients are selected based on minimizing a cost function that combines both distortion and bit rate via a user-controlled weighting. Two methods are proposed - exhaustive search and fast Lagrangian approximation. While the former produces optimal results, the latter approach is significantly faster and amenable to real-time implementation.
Jithendra K. Paruchuri, Sen-Ching S. Cheung
ISCAS2
2008 Efficient Multimedia Distribution in Source Constraint Networks
abstract
In recent years, the number of Peer-to-Peer (P2P) applications has increased significantly. One important problem in many P2P applications is how to efficiently disseminate data from a single source to multiple receivers on the Internet. A successful model used for analyzing this problem is a graph consisting of nodes and edges, with a capacity assigned to each edge. In some situations however, it is inconvenient to use this model. To that end, we propose to study the problem of efficient data dissemination in asource constraintnetwork. A source constraint network is modeled as a graph in which, the capacity is associated with a node, rather than an edge. The contributions of this paper include (a) a quantitative data dissemination in any source constraint network, (b) a set of topologies suitable for data dissemination in P2P networks, and (c) an architecture and implementation of a P2P system based on the proposed optimal topologies. We will present the experimental results of our P2P system deployed on PlanetLab nodes demonstrating that our approach achieves near optimal throughput while providing scalability, low delay and bandwidth fairness among peers.
Thinh P. Nguyen, Krishnan Kolazhi, Rohit Kamath, Sen-Ching S. Cheung, Duc A. Tran
IEEE Trans. Multim.4
2008 Multimedia streaming using multiple TCP connections
abstract
In recent years, multimedia applications over the Internet become increasingly popular. However, packet loss, delay, and time-varying bandwidth of the Internet have remained the major problems for multimedia streaming applications. As such, a number of approaches, including network infrastructure and protocol, source and channel coding, have been proposed to either overcome or alleviate these drawbacks of the Internet. In this article, we propose the MultiTCP system, a receiver-driven, TCP-based system for multimedia streaming over the Internet. Our proposed algorithm aims at providing resilience against short term insufficient bandwidth by using multiple TCP connections for the same application. Our proposed system enables the application to achieve and control the desired sending rate during congested periods, which cannot be achieved using traditional TCP. Finally, our proposed system is implemented at the application layer, and hence, no kernel modification to TCP is necessary. We analyze the proposed system, and present simulation and experimental results to demonstrate its advantages over the traditional single-TCP-based approach.
Sunand Tullimas, Thinh P. Nguyen, Rich Edgecomb, Sen-Ching S. Cheung
ACM Trans. Multim. Comput. Commun. Appl.4
2007 A New Security Model for Secure Thresholding
abstract
The goal of secure computation is for distrusted parties on a network to collaborate with each other without disclosing private information. In this paper, we focus on secure thresholding, or comparing two secret numbers, which is a key step in pattern recognition. Existing cryptographic protocols are too complex to be used in real-time signal processing. We propose a new security model based on noninvertible functions called quasi-information-theoretic security. Using this model, we develop a novel secure thresholding protocol that is both secure and computationally efficient. The proposed protocol hides information in a carefully-designed random polynomial and in a lower-rank subspace based on Chebyshev's polynomials.
Sen-Ching S. Cheung
ICASSP (2)2
2007 Peer-to-Peer Streaming with Hierarchical Network Coding
abstract
In recent years, Content Delivery Networks (CDN) and Peer-to-Peer (P2P) networks have emerged as two effective paradigms for delivering multimedia contents over the Internet. An important feature in CDN and P2P networks is the data redundancy across multiple servers/peers which enables efficient media delivery. In this paper, we propose a network coding framework for efficient media streaming in either content delivery networks or P2P networks in which, multiple servers/peers are employed to simultaneously stream a video to a single receiver. Unlike previous multi-sender schemes, we show that network coding technique can (a) reduce the redundancy storage, (b) eliminate the need for tight synchronization between the senders, and (c) be integrated easily with TCP. Furthermore, we propose the Hierarchical Network Coding (HNC) technique to be used with scalable video bit stream to combat bandwidth fluctuation on the Internet. Simulation results demonstrate that our proposed scheme can result in bandwidth saving up to 40% for many cases over the traditional schemes.
Kien Nguyen 0004, Thinh P. Nguyen, Sen-Ching S. Cheung
ICME3
2007 Secure Multiparty Computation between Distrusted Networks Terminals
abstract
One of the most important problems facing any distributed application over a heteroge-neous network is the protection of private sensitive information in local terminals. A subfield of cryptography called Secure Multiparty Computation (SMC) is the study of such distributed computation protocols that allow distrusted parties to perform joint computation without dis-closing private data. SMC is increasingly used in diverse fields from data mining to computer vision. This paper provides a tutorial on SMC for non-experts in cryptography and surveys some of the latest advances in this exciting area including various schemes for reducing commu-nication and computation complexity of SMC protocols, doubly homomorphic encryption and private information retrieval. The proliferation of capturing and storage devices as well as the ubiquitous presence of com-puter networks make sharing of data easier than ever. Such pervasive exchange of data, however, has increasingly raised questions on how sensitive and private information can be protected. For example, it is now commonplace to send private photographs or videos to the hundreds of online photo processing stores for storage, development and enhancement like sharpening and red-eye removal. Few companies provide any protection of the personal pictures they receive. Hackers or
Sen-Ching S. Cheung, Thinh P. Nguyen
EURASIP J. Inf. Secur.1
2006 Efficient Object-Based Video Inpainting
abstract
Video inpainting describes the process of removing a portion of a video and filling in the missing part (hole) in a visually consistent manner. Most existing video inpainting techniques are computationally intensive and cannot handle large holes. In this paper, we propose a complete and efficient video in-painting system. Our system applies different strategies to handle static and dynamic portions of the hole. To inpaint the static portion, our system uses background replacement and image inpainting techniques. To inpaint moving objects in the hole, we utilizes background subtraction and object segmentation to extract a set of object templates and perform optimal object interpolation using dynamic programming. We evaluate the performance of our system based on a set of indoor surveillance sequences with different types of occlusions.
Sen-Ching S. Cheung, Jian Zhao 0003, M. Vijay Venkatesh
ICIP1
2006 Secure Image Filtering
abstract
In today's heterogeneous network environment, there is a growing demand for distrusted parties to jointly execute distributed algorithms on private data whose secrecy needed to be safeguarded. Protocols that support such kind of joint computation without complete sharing of information are called secure multiparty computation (SMC) protocols. Applying SMC protocols in image processing is a challenging problem. Most of the existing SMC protocols are implemented based on cryptographic primitives like oblivious transfer that are too computational intensive for pixel-based operations. In this paper, we develop two efficient SMC protocols for distributed linear image filtering between two parties, one party with the original image and the other with the image filter. The first protocol is based on a combination of rank reduction and random permutation. The second one uses random perturbation with the help of a non-colluding third party. Experimental results show that both of them execute significantly faster than oblivious-transfer based techniques.
Sen-Ching S. Cheung, Thinh P. Nguyen
ICIP2
2006 Symmetric Shape Completion Under Severe Occlusions
abstract
In this paper, we propose a novel algorithm for completing rotationally symmetrical shapes under severe occlusions. The intuitive idea is to use the existing contour, under a carefully estimated similarity transform, to fill in the missing portion of a symmetric object due to occlusions. Our algorithm exploits the invariant nature of the curvature under similarity transform and the periodicity of the curvature of a symmetric object contour. To arrive at the appropriate transform, we first estimate the fundamental period in the curvature. We use the fundamental period and the harmonic components to estimate the fundamental angle of rotation and the centroid of the unoccluded shape, which in turn establish different modes of symmetry. By following each mode of symmetry we compute the corresponding transform and select the ones that best complete the missing portion of the contour.
M. Vijay Venkatesh, Sen-Ching S. Cheung
ICIP2
2006 Efficient Video Dissemination in Structured Hybrid P2P Networks
abstract
In this paper, we propose a structured hybrid P2P mesh for optimal video dissemination from a single source node to multiple receivers in a bandwidth-asymmetric network such as digital subscriber line (DSL) access network. Our hybrid P2P structured mesh consists of one or more supernodes responsible for node and mesh management and a large number of streaming nodes, Peers. The peers are interconnected in a special manner designed for streaming and real-time video dissemination and are responsible for the actual data delivery. Our proposed hybrid P2P structured mesh is designed to achieve scalability, low delay and high throughput. Our experimental Internet-wide system consisting of PlanetLab nodes demonstrates the aforementioned qualities
Thinh P. Nguyen, Krishnan Kolazhi, Rohit Kamath, Sen-Ching S. Cheung
ICME4
2005 Offline Generation of High Quality Background Subtraction Data
abstract
Ground truth is important not only for performance evaluation but also for a principled development of computer vision algorithms. Unfortunately obtaining ground truth data is difficult and often very labor intensive. This is particularly true of video analysis due to the immense cost of producing pixel-wise ground truth in potentially thousands of frames. In this paper, we propose a method to produce foreground/background segmentation for video sequences captured by a stationary camera, that requires very little human labor as compared to complete manual segmentation, while still producing high quality results. Given a sequence, we use a few hand labeled images and Adaboost to train a classifier that segments the rest of the sequence. We demonstrate the effectiveness of our approach on two sequences and discuss the new horizons opened by these encouraging results. 1
Etienne Grossmann, Amit A. Kale, Christopher O. Jaynes, Sen-Ching S. Cheung
BMVC4
2005 Mining arbitrary-length repeated patterns in television broadcast
abstract
Mining repeated patterns in television broadcast is important to advertisers in tracking a large number of television commercials. It can also benefit long-term archival of television because historically significant events are usually marked by repeated airing of the same video clips or sound-bytes. In this paper, we describe a system that can efficiently mine repeated patterns of arbitrary lengths from television broadcast. Compared with existing work, our system has two main innovations: first, our system is robust against minor temporal variations among repeated patterns. This is important as broadcasters often perform temporal editing on commercials so as to fit them into different time slots. Second, our system does not rely on any temporal segmentation algorithm, which may lead to over- or under-segmentation of important patterns. Instead, our system scans the television broadcast with a fixed-size sliding window, summarizes each window into a hash value, and maintains a running frequency count and a reference time-stamp on each hash value. The boundaries of a repeated pattern are identified by the changes in frequency counts and reference time-stamps. Initial experiments show that our system is very efficient in identifying all the repeated commercials from 12 hours of television broadcast.
Sen-Ching S. Cheung, Thinh P. Nguyen
ICIP (3)1
2005 Hiding privacy information in video surveillance system
abstract
This paper proposes a detailed framework of storing privacy information in surveillance video as a watermark. Authorized personnel is not only removed from the surveillance video as in J. Wickramasuriya et al. (2004) but also embedded into the video itself, which can only be retrieved with a secrete key. A perceptual-model-based compressed domain video watermarking scheme is proposed to deal with the huge payload problem in the proposed surveillance system. A signature is also embedded into the header of the video as in M. Pramateftakis et al. (2004) for authentication. Simulation results have shown that the proposed algorithm can embed all the privacy information into the video without affecting its visual quality. As a result, the proposed video surveillance system can monitor the unauthorized persons in a restricted environment, protect the privacy of the authorized persons but, at the same time, allow the privacy information to be revealed in a secure and reliable way.
Wei Zhang 0013, Sen-Ching S. Cheung, Minghua Chen 0001
ICIP (3)2
2005 Multimedia streaming using multiple TCP connections
abstract
As broadband Internet becomes widely available, multimedia applications over the Internet become increasingly popular. However, packet loss, delay, and time-varying bandwidth of the Internet have remained the major problems for multimedia streaming applications. As such, a number of approaches, including network infrastructure and protocol, source and channel coding have been proposed to either overcome or alleviate these drawbacks of the Internet. In this paper, we propose the MultiTCP system, a receiver-driven, TCP-based system for multimedia streaming over the Internet. Our proposed algorithm aims at providing resilience against SHORT TERM insufficient bandwidth by using MULTIPLE TCP connections for the same application. Furthermore, our proposed system enables the application to achieve and control the desired sending rate during congested periods, which cannot be achieved using traditional TCP. Finally, our proposed system is implemented at the application layer, and hence, no kernel modification to TCP is necessary. We analyze the proposed system, and present simulation results to demonstrate its advantages over the traditional single TCP based approach.
Thinh P. Nguyen, Sen-Ching S. Cheung
IPCCC2
2005 Fast similarity search and clustering of video sequences on the world-wide-web
abstract
We define similar video content as video sequences with almost identical content but possibly compressed at different qualities, reformatted to different sizes and frame-rates, undergone minor editing in either spatial or temporal domain, or summarized into keyframe sequences. Building a search engine to identify such similar content in the World-Wide Web requires: 1) robust video similarity measurements; 2) fast similarity search techniques on large databases; and 3) intuitive organization of search results. In a previous paper, we proposed a randomized technique called the video signature (ViSig) method for video similarity measurement. In this paper, we focus on the remaining two issues by proposing a feature extraction scheme for fast similarity search, and a clustering algorithm for identification of similar clusters. Similar to many other content-based methods, the ViSig method uses high-dimensional feature vectors to represent video. To warrant a fast response time for similarity searches on high dimensional vectors, we propose a novel nonlinear feature extraction scheme on arbitrary metric spaces that combines the triangle inequality with the classical Principal Component Analysis (PCA). We show experimentally that the proposed technique outperforms PCA, Fastmap, Triangle-Inequality Pruning, and Haar wavelet on signature data. To further improve retrieval performance, and provide better organization of similarity search results, we introduce a new graph-theoretical clustering algorithm on large databases of signatures. This algorithm treats all signatures as an abstract threshold graph, where the distance threshold is determined based on local data statistics. Similar clusters are then identified as highly connected regions in the graph. By measuring the retrieval performance against a ground-truth set, we show that our proposed algorithm outperforms simple thresholding, single-link and complete-link hierarchical clustering techniques.
Sen-Ching S. Cheung, Avideh Zakhor
IEEE Trans. Multim.1
2004 Robust techniques for background subtraction in urban traffic video
abstract
Identifying moving objects from a video sequence is a fundamental and critical task in many computer-vision applications. A common approach is to perform background subtraction, which identifies moving objects from the portion of a video frame that differs significantly from a background model. There are many challenges in developing a good background subtraction algorithm. First, it must be robust against changes in illumination. Second, it should avoid detecting non-stationary background objects such as swinging leaves, rain, snow, and shadow cast by moving objects. Finally, its internal background model should react quickly to changes in background such as starting and stopping of vehicles. In this paper, we compare various background subtraction algorithms for detecting moving vehicles and pedestrians in urban traffic video sequences. We consider approaches varying from simple techniques such as frame differencing and adaptive median filtering, to more sophisticated probabilistic modeling techniques. While complicated techniques often produce superior performance, our experiments show that simple techniques such as adaptive median filtering can produce good results with much lower computational complexity.
Sen-Ching S. Cheung, Chandrika Kamath 0001
VCIP1
2003 Fast similarity search on video signatures
abstract
Video signatures are compact representations of video sequences designed for efficient similarity measurement. In this paper, we propose a feature extraction technique to support fast similarity search on large databases of video signatures. Our proposed technique transforms the high dimensional video signatures into low dimensional vectors where similarity search can be efficiently performed. We exploit both the upper and lower bounds of the triangle inequalities in approximating the high-dimensional metric, and combine this approximation with the classical PCA to achieve the target dimension. Experimental results on a large set of Web video sequences show that our technique outperforms fastmap, Haar wavelet, PCA, and triangle-inequality pruning.
Sen-Ching S. Cheung, Avideh Zakhor
ICIP (2)1
2003 Efficient video similarity measurement with video signature
abstract
The proliferation of video content on the Web makes similarity detection an indispensable tool in Web data management, searching, and navigation. We propose a number of algorithms to efficiently measure video similarity. We define video as a set of frames, which are represented as high dimensional vectors in a feature space. Our goal is to measure ideal video similarity (IVS), defined as the percentage of clusters of similar frames shared between two video sequences. Since IVS is too complex to be deployed in large database applications, we approximate it with Voronoi video similarity (VVS), defined as the volume of the intersection between Voronoi cells of similar clusters. We propose a class of randomized algorithms to estimate VVS by first summarizing each video with a small set of its sampled frames, called the video signature (ViSig), and then calculating the distances between corresponding frames from the two ViSigs. By generating samples with a probability distribution that describes the video statistics, and ranking them based upon their likelihood of making an error in the estimation, we show analytically that ViSig can provide an unbiased estimate of IVS. Experimental results on a large dataset of Web video and a set of MPEG-7 test sequences with artificially generated similar versions are provided to demonstrate the retrieval performance of our proposed techniques.
Sen-Ching S. Cheung, Avideh Zakhor
IEEE Trans. Circuits Syst. Video Technol.1
2002 Efficient video similarity measurement with video signature
abstract
The video signature method has previously been proposed as a technique to summarize video efficiently for visual similarity measurements (see Cheung, S.-C. and Zakhor, A., Proc. SPIE, vol.3964, p.34-6, 2000; ICIP2000, vol.1, p.85-9, 2000; ICIP2001, vol.1, p.649-52, 2001). We now develop the necessary theoretical framework to analyze this method. We define our target video similarity measure based on the fraction of similar clusters shared between two video sequences. This measure is too computationally complex to be deployed in database applications. By considering this measure geometrically on the image feature space, we find that it can be approximated by the volume of the intersection between Voronoi cells of similar clusters. In the video signature method, sampling is used to estimate this volume. By choosing an appropriate distribution to generate samples, and ranking the samples based upon their distances to the boundary between Voronoi cells, we demonstrate that our target measure can be well approximated by the video signature method. Experimental results on a large dataset of Web video and a set of MPEG-7 test sequences with artificially generated similar versions are used to demonstrate the retrieval performance of our proposed techniques.
Sen-Ching S. Cheung, Avideh Zakhor
ICIP (1)1
2001 Video similarity detection with video signature clustering
abstract
The proliferation of video content on the Web makes similarity detection an indispensable tool in Web data management, searching, and navigation. We have previously proposed a compact representation of video clips, called video signature, for retrieving similar video clips in large databases. In this paper, we propose a new signature clustering algorithm to further improve retrieval performance. The algorithm treats all the signatures as an abstract threshold graph, where the threshold is determined based on local data statistics. Similar clusters are identified as highly connected regions in the graph. This algorithm outperforms simple thresholding and hierarchical clustering techniques in identifying a set of manually-determined similar clusters from a dataset of 46,356 Web video clips. At 95% precision, our algorithm attains 85% recall while simple thresholding and complete-link hierarchical scheme attain 67% and 75% recall respectively. Applying our algorithm to the entire dataset, 6,900 similar clusters are identified, with an average cluster size of 2.81 video clips. The distribution of cluster sizes follows a power-law distribution, which has been shown to describe many Web phenomena.
Sen-Ching S. Cheung, Avideh Zakhor
ICIP (2)1
2000 Efficient Video Similarity Measurement and Search
abstract
We consider the use of meta-data and/or video-domain methods to detect similar videos on the Web. Meta-data is extracted from the textual and hyperlink information associated with each video clip. In the video domain, we apply an efficient similarity detection algorithm called video signature. The idea is to form a signature for each clip by selecting a small number of its frames that are most similar to a set of random seed images. We then apply a statistical pruning algorithm to allow fast detection on very large databases. Using a small ground-truth set, we achieve 90% recall and 95% precision using only 8% of the total number of operations required without pruning. For a database of around 46,000 video clips crawled from the Web, the video signature technique significantly outperforms meta-data in precision and recall. We show that even better performance can be achieved by combining them together. Based on our measurements, each video clip in our database has, on average, 1.53 similar copies.
Sen-Ching S. Cheung, Avideh Zakhor
ICIP1