Ngai-Man Cheung

dblp:82/3605 · DBLP profile ↗
← Back
135ranked-venue papers
16as first author
29since 2021 · last 2026
0000-0003-0135-3791ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 110 · 16 first-author · 19 since 2021Artificial intelligence and machine learning · 40 · 21 since 2021Databases, data management, data science and information retrieval · 8Computer networks · 5Applied, interdisciplinary, general and emerging computing · 3Systems, architecture and hardware · 2Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 SAVER: Mitigating Hallucinations in Large Vision-Language Models via Style-Aware Visual Early Revision
abstract
Large Vision-Language Models (LVLMs) recently achieve significant breakthroughs in understanding complex visual-textual contexts. However, hallucination issues still limit their real-world applicability. Although previous mitigation methods effectively reduce hallucinations in photographic images, they largely overlook the potential risks posed by stylized images, which play crucial roles in critical scenarios such as game scene understanding, art education, and medical analysis. In this work, we first construct a dataset comprising photographic images and their corresponding stylized versions with carefully annotated caption labels. We then conduct head-to-head comparisons on both discriminative and generative tasks by benchmarking 13 advanced LVLMs on the collected datasets. Our findings reveal that stylized images tend to induce significantly more hallucinations than their photographic counterparts. To address this issue, we propose Style-Aware Visual Early Revision (SAVER), a novel mechanism that dynamically adjusts LVLMs' final outputs based on the token-level visual attention patterns, leveraging early-layer feedback to mitigate hallucinations caused by stylized images. Extensive experiments demonstrate that SAVER achieves state-of-the-art performance in hallucination mitigation across various models, datasets, and tasks.
Zhaoxu Li, Chenqi Kong, Yi Yu 0011, Qiangqiang Wu, Xinghao Jiang, Ngai-Man Cheung, Bihan Wen, Alex Chichung Kot, Xudong Jiang 0001
AAAI6
2024 Referring Expression Counting
abstract
Existing counting tasks are limited to the class level, which don't account for fine-grained details within the class. In real applications, it often requires in-context or referring human input for counting target objects. Take urban analysis as an example, fine-grained information such as traffic flow in different directions, pedestrians and vehicles waiting or moving at different sides of the junction, is more beneficial. Current settings of both class-specific and class-agnostic counting treat objects of the same class indifferently, which pose limitations in real use cases. To this end, we propose a new task named Referring Expression Counting (REC) which aims to count objects with different attributes within the same class. To evaluate the REC task, we create a novel dataset named REC-8K which contains 8011 images and 17122 referring expressions. Experiments on REC-8K show that our proposed method achieves state-of-the-art performance compared with several text-based counting methods and an open-set object detection model. We also outperform prior models on the class agnostic counting (CAC) benchmark [36]for the zero-shot setting, and perform on par with the few-shot methods. Code and dataset is available at https://github.com/sydai/referring-expression-counting.
Siyang Dai, Jun Liu 0036, Ngai-Man Cheung
CVPR3
2024 Model Inversion Robustness: Can Transfer Learning Help?
abstract
Model Inversion (MI) attacks aim to reconstruct private training data by abusing access to machine learning models. Contemporary MI attacks have achieved impressive attack performance, posing serious threats to privacy. Meanwhile, all existing MI defense methods rely on regularization that is in direct conflict with the training objective, resulting in noticeable degradation in model utility. In this work, we take a different perspective, and propose a novel and simple Transfer Learning-based Defense against Model Inversion (TL-DMI) to render MI-robust models. Particularly, by leveraging TL, we limit the number of layers encoding sensitive information from private training dataset, thereby degrading the performance of MI attack. We conduct an analysis using Fisher Information to justify our method. Our defense is remarkably simple to implement. Without bells and whistles, we show in extensive experiments that TL-DMI achieves state-of-the-art (SOTA) MI robustness. Our code, pre-trained models, demo and inverted data are available at: https://hosytuyen.github.io/projectsITL-DMI
Sy-Tuyen Ho, Koh Jun Hao, Keshigeyan Chandrasegaran, Ngoc-Bao Nguyen, Ngai-Man Cheung
CVPR5
2024 On the Vulnerability of Skip Connections to Model Inversion Attacks
Koh Jun Hao, Sy-Tuyen Ho, Ngoc-Bao Nguyen, Ngai-Man Cheung
ECCV (81)4
2024 Frequency Masking for Universal Deepfake Detection
abstract
We study universal deepfake detection. Our goal is to detect synthetic images from a range of generative AI approaches, particularly from emerging ones which are unseen during training of the deepfake detector. Universal deepfake detection requires outstanding generalization capability. Motivated by recently proposed masked image modeling which has demonstrated excellent generalization in self-supervised pre-training, we make the first attempt to explore masked image modeling for universal deepfake detection. We study spatial and frequency domain masking in training deepfake detectors. Based on empirical analysis, we propose a novel deepfake detector via frequency masking. Our focus on frequency domain is different from the majority, which primarily target spatial domain detection. Our comparative analyses reveal substantial performance gains over existing methods. Code and models are publicly available1.
Chandler Timm C. Doloriel, Ngai-Man Cheung
ICASSP2
2024 Masked Signal Modeling for Plastic Waste Resin Classification
abstract
Near-infrared spectroscopy is a viable option for plastic material detection. The lack of diversity and scale of such data makes it difficult to translate the research outcomes into industrial settings. In this paper, we collect a diverse plastic dataset, named AI Recycling Plastics (AIRP), with different compositions of clean and contaminated plastics. Moreover, we aim to apply self-supervised pre-training using masked signal modeling (MSM) to improve the performance of deep learning models in detecting the plastic resin types from the spectral signature data. Our results show that MSM, as a pre-training task, teaches the model to learn rich features that generalize well from unlabeled spectra data. Fine-tuning performance in the detection of plastic resin type outperforms the baseline supervised techniques. Our proposed approach shows promising results that can be applied industrially at material recovery facilities to sort recyclable plastics and less common polymer types.
S. Ebrahimkhani, A. C. Y. Ngo, Ngai-Man Cheung
ICIP4
2024 Vision Transformer Neural Architecture Search for Out-of-Distribution Generalization: Benchmark and Insights
abstract
While Vision Transformer (ViT) have achieved success across various machine learning tasks, deploying them in real-world scenarios faces a critical challenge: generalizing under Out-of-Distribution (OoD) shifts. A crucial research gap remains in understanding how to design ViT architectures – both manually and automatically – to excel in OoD generalization. **To address this gap,** we introduce OoD-ViT-NAS, the first systematic benchmark for ViT Neural Architecture Search (NAS) focused on OoD generalization. This comprehensive benchmark includes 3,000 ViT architectures of varying model computational budgets evaluated on common large-scale OoD datasets. With this comprehensive benchmark at hand, we analyze the factors that contribute to the OoD generalization of ViT architecture. Our analysis uncovers several key insights. Firstly, we show that ViT architecture designs have a considerable impact on OoD generalization. Secondly, we observe that In-Distribution (ID) accuracy might not be a very good indicator of OoD accuracy. This underscores the risk that ViT architectures optimized for ID accuracy might not perform well under OoD shifts. Thirdly, we conduct the first study to explore NAS for ViT’s OoD robustness. Specifically, we study 9 Training-free NAS for their OoD generalization performance on our benchmark. We observe that existing Training-free NAS are largely ineffective in predicting OoD accuracy despite their effectiveness at predicting ID accuracy. Moreover, simple proxies like #Param or #Flop surprisingly outperform more complex Training-free NAS in predicting ViTs OoD accuracy. Finally, we study how ViT architectural attributes impact OoD generalization. We discover that increasing embedding dimensions of a ViT architecture generally can improve the OoD generalization. We show that ViT architectures in our benchmark exhibit a wide range of OoD accuracy, with up to 11.85% for some OoD shift, prompting the importance to study ViT architecture design for OoD. We firmly believe that our OoD-ViT-NAS benchmark and our analysis can catalyze and streamline important research on understanding how ViT architecture designs influence OoD generalization. **Our OoD-NAS-ViT benchmark and code are available at [https://hosytuyen.github.io/projects/OoD-ViT-NAS](https://hosytuyen.github.io/projects/OoD-ViT-NAS)**
Sy-Tuyen Ho, Tuan Van Vo, Somayeh Ebrahimkhani, Ngai-Man Cheung
NeurIPS4
2024 FairQueue: Rethinking Prompt Learning for Fair Text-to-Image Generation
abstract
Recently, prompt learning has emerged as the state-of-the-art (SOTA) for fair text-to-image (T2I) generation. Specifically, this approach leverages readily available reference images to learn inclusive prompts for each target Sensitive Attribute (tSA), allowing for fair image generation. In this work, we first reveal that this prompt learning-based approach results in degraded sample quality. Our analysis shows that the approach's training objective--which aims to align the embedding differences of learned prompts and reference images-- could be sub-optimal, resulting in distortion of the learned prompts and degraded generated images. To further substantiate this claim, **as our major contribution**, we deep dive into the denoising subnetwork of the T2I model to track down the effect of these learned prompts by analyzing the cross-attention maps. In our analysis, we propose a novel prompt switching analysis: I2H and H2I. Furthermore, we propose new quantitative characterization of cross-attention maps. Our analysis reveals abnormalities in the early denoising steps, perpetuating improper global structure that results in degradation in the generated samples. Building on insights from our analysis, we propose two ideas: (i) *Prompt Queuing* and (ii) *Attention Amplification* to address the quality issue. Extensive experimental results on a wide range of tSAs show that our proposed method outperforms SOTA approach's image generation quality, while achieving competitive fairness. More resources at FairQueue Project site: https://sutd-visual-computing-group.github.io/FairQueue
Christopher T. H. Teo, Milad Abdollahzadeh, Xinda Ma, Ngai-Man Cheung
NeurIPS4
2024 Uncertainty-Aware Pedestrian Crossing Prediction via Reinforcement Learning
abstract
Pedestrian safety is a huge concern for deploying autonomous vehicles in urban environments. Accidents involving pedestrians pose a higher degree of severity, sometimes causing serious injuries and fatalities [1]. It’s a challenging task to predict whether a pedestrian will cross the road since they can move in any direction and change motion suddenly. The inherent uncertainty in pedestrian motion has been addressed with probabilistic models in previous works. However, these models are too computationally expensive for real-time predictions. In this paper, we propose a novel reinforcement learning (RL) framework which produces soft labels for the training dataset in order to address the observed data uncertainty. We formulate novel state representations incorporating predictive uncertainty to learn more informative soft labels that improve the model performance and reliability. Finally, we validate the proof of concept with two benchmark datasets and show with extensive experiments on competitive prediction models that our method (even using fewer input modalities) significantly improves the accuracy and f1 score by up to 12% and 13% respectively. We also show that soft labeling as a form of regularization increases model reliability where the model is more accurate when the confidence level is high and more aware of its limitations with indication of low confidence.
Siyang Dai, Jun Liu 0036, Ngai-Man Cheung
IEEE Trans. Circuits Syst. Video Technol.3
2023 Fair Generative Models via Transfer Learning
abstract
This work addresses fair generative models. Dataset biases have been a major cause of unfairness in deep generative models. Previous work had proposed to augment large, biased datasets with small, unbiased reference datasets. Under this setup, a weakly-supervised approach has been proposed, which achieves state-of-the-art quality and fairness in generated samples. In our work, based on this setup, we propose a simple yet effective approach. Specifically, first, we propose fairTL, a transfer learning approach to learn fair generative models. Under fairTL, we pre-train the generative model with the available large, biased datasets and subsequently adapt the model using the small, unbiased reference dataset. We find that our fairTL can learn expressive sample generation during pre-training, thanks to the large (biased) dataset. This knowledge is then transferred to the target model during adaptation, which also learns to capture the underlying fair distribution of the small reference dataset. Second, we propose fairTL++, where we introduce two additional innovations to improve upon fairTL: (i) multiple feedback and (ii) Linear-Probing followed by Fine-Tuning (LP-FT). Taking one step further, we consider an alternative, challenging setup when only a pre-trained (potentially biased) model is available but the dataset that was used to pre-train the model is inaccessible. We demonstrate that our proposed fairTL and fairTL++ remain very effective under this setup. We note that previous work requires access to the large, biased datasets and is incapable of handling this more challenging setup. Extensive experiments show that fairTL and fairTL++ achieve state-of-the-art in both quality and fairness of generated samples. The code and additional resources can be found at bearwithchris.github.io/fairTL/.
Christopher T. H. Teo, Milad Abdollahzadeh, Ngai-Man Cheung
AAAI3
2023 Re-Thinking Model Inversion Attacks Against Deep Neural Networks
abstract
Model inversion (MI) attacks aim to infer and reconstruct private training data by abusing access to a model. MI attacks have raised concerns about the leaking of sen-sitive information (e.g. private face images used in training a face recognition system). Recently, several algorithms for MI have been proposed to improve the attack performance. In this work, we revisit MI, study two fundamental issues pertaining to all state-of-the-art (SOTA) MI algorithms, and propose solutions to these issues which lead to a significant boost in attack performance for all SOTA MI. In particular, our contributions are two-fold: 1) We ana-lyze the optimization objective of SOTA MI algorithms, ar-gue that the objective is sub-optimal for achieving MI, and propose an improved optimization objective that boosts attack performance significantly. 2) We analyze “MI overfitting”, show that it would prevent reconstructed images from learning semantics of training data, and propose a novel “model augmentation” idea to overcome this issue. Our proposed solutions are simple and improve all SOTA MI attack accuracy significantly. E.g., in the standard CelebA benchmark, our solutions improve accuracy by 11.8% and achieve for the first time over 90% attack accuracy. Our findings demonstrate that there is a clear risk of leaking sensitive information from deep learning models. We urge serious consideration to be given to the privacy im-plications. Our code, demo, and models are available at https://ngoc-nguyen-0.github.io/re-thinking_mode1_inversion_attacks/.
Ngoc-Bao Nguyen, Keshigeyan Chandrasegaran, Milad Abdollahzadeh, Ngai-Man Cheung
CVPR4
2023 Exploring Incompatible Knowledge Transfer in Few-shot Image Generation
abstract
Few-shot image generation (FSIG) learns to generate diverse and high-fidelity images from a target domain using a few (e.g., 10) reference samples. Existing FSIG methods select, preserve and transfer prior knowledge from a source generator (pretrained on a related domain) to learn the target generator. In this work, we investigate an underexplored issue in FSIG, dubbed as incompatible knowledge transfer, which would significantly degrade the realisticness of synthetic samples. Empirical observations show that the issue stems from the least significant filters from the source generator. To this end, we propose knowledge truncation to mitigate this issue in FSIG, which is a complementary operation to knowledge preservation and is implemented by a lightweight pruning-based method. Extensive experiments show that knowledge truncation is simple and effective, consistently achieving state-of-the-art performance, including challenging setups where the source and target domains are more distant. Project Page: yunqing-me.github.io/RICK.
Yunqing Zhao, Milad Abdollahzadeh, Tianyu Pang, Shuicheng Yan, Ngai-Man Cheung
CVPR7
2023 Label-Only Model Inversion Attacks via Knowledge Transfer
abstract
In a model inversion (MI) attack, an adversary abuses access to a machine learning (ML) model to infer and reconstruct private training data. Remarkable progress has been made in the white-box and black-box setups, where the adversary has access to the complete model or the model's soft output respectively. However, there is very limited study in the most challenging but practically important setup: Label-only MI attacks, where the adversary only has access to the model's predicted label (hard label) without confidence scores nor any other model information. In this work, we propose LOKT, a novel approach for label-only MI attacks. Our idea is based on transfer of knowledge from the opaque target model to surrogate models. Subsequently, using these surrogate models, our approach can harness advanced white-box attacks. We propose knowledge transfer based on generative modelling, and introduce a new model, Target model-assisted ACGAN (T-ACGAN), for effective knowledge transfer. Our method casts the challenging label-only MI into the more tractable white-box setup. We provide analysis to support that surrogate models based on our approach serve as effective proxies for the target model for MI. Our experiments show that our method significantly outperforms existing SOTA Label-only MI attack by more than 15% across all MI benchmarks. Furthermore, our method compares favorably in terms of query budget. Our study highlights rising privacy threats for ML models even when minimal information (i.e., hard labels) is exposed. Our study highlights rising privacy threats for ML models even when minimal information (i.e., hard labels) is exposed. Our code, demo, models and reconstructed data are available at our project page: https://ngoc-nguyen-0.github.io/lokt/
Ngoc-Bao Nguyen, Keshigeyan Chandrasegaran, Milad Abdollahzadeh, Ngai-Man Cheung
NeurIPS4
2023 On Measuring Fairness in Generative Models
abstract
Recently, there has been increased interest in fair generative models. In this work, we conduct, for the first time, an in-depth study on fairness measurement, a critical component in gauging progress on fair generative models. We make three contributions. First, we conduct a study that reveals that the existing fairness measurement framework has considerable measurement errors, even when highly accurate sensitive attribute (SA) classifiers are used. These findings cast doubts on previously reported fairness improvements. Second, to address this issue, we propose CLassifier Error-Aware Measurement (CLEAM), a new framework which uses a statistical model to account for inaccuracies in SA classifiers. Our proposed CLEAM reduces measurement errors significantly, e.g., 4.98%→0.62% for StyleGAN2 w.r.t. Gender. Additionally, CLEAM achieves this with minimal additional overhead. Third, we utilize CLEAM to measure fairness in important text-to-image generator and GANs, revealing considerable biases in these models that raise concerns about their applications. Code and more resources: https: //sutd-visual-computing-group.github.io/CLEAM/.
Christopher T. H. Teo, Milad Abdollahzadeh, Ngai-Man Cheung
NeurIPS3
2023 On Evaluating Adversarial Robustness of Large Vision-Language Models
abstract
Large vision-language models (VLMs) such as GPT-4 have achieved unprecedented performance in response generation, especially with visual inputs, enabling more creative and adaptable interaction than large language models such as ChatGPT. Nonetheless, multimodal generation exacerbates safety concerns, since adversaries may successfully evade the entire system by subtly manipulating the most vulnerable modality (e.g., vision). To this end, we propose evaluating the robustness of open-source large VLMs in the most realistic and high-risk setting, where adversaries have only black-box system access and seek to deceive the model into returning the targeted responses. In particular, we first craft targeted adversarial examples against pretrained models such as CLIP and BLIP, and then transfer these adversarial examples to other VLMs such as MiniGPT-4, LLaVA, UniDiffuser, BLIP-2, and Img2Prompt. In addition, we observe that black-box queries on these VLMs can further improve the effectiveness of targeted evasion, resulting in a surprisingly high success rate for generating targeted responses. Our findings provide a quantitative understanding regarding the adversarial vulnerability of large VLMs and call for a more thorough examination of their potential security flaws before deployment in practice. Our project page: https://yunqing-me.github.io/AttackVLM/.
Yunqing Zhao, Tianyu Pang, Xiao Yang 0028, Chongxuan Li, Ngai-Man Cheung
NeurIPS6
2023 FS-BAN: Born-Again Networks for Domain Generalization Few-Shot Classification
abstract
Conventional Few-shot classification (FSC) aims to recognize samples from novel classes given limited labeled data. Recently, domain generalization FSC (DG-FSC) has been proposed with the goal to recognize novel class samples from unseen domains. DG-FSC poses considerable challenges to many models due to the domain shift between base classes (used in training) and novel classes (encountered in evaluation). In this work, we make two novel contributions to tackle DG-FSC. Our first contribution is to propose Born-Again Network (BAN) episodic training and comprehensively investigate its effectiveness for DG-FSC. As a specific form of knowledge distillation, BAN has been shown to achieve improved generalization in conventional supervised classification with a closed-set setup. This improved generalization motivates us to study BAN for DG-FSC, and we show that BAN is promising to address the domain shift encountered in DG-FSC. Building on the encouraging findings, our second (major) contribution is to propose Few-Shot BAN (FS-BAN), a novel BAN approach for DG-FSC. Our proposed FS-BAN includes novel multi-task learning objectives: Mutual Regularization, Mismatched Teacher, and Meta-Control Temperature, each of these is specifically designed to overcome central and unique challenges in DG-FSC, namely overfitting and domain discrepancy. We analyze different design choices of these techniques. We conduct comprehensive quantitative and qualitative analysis and evaluation over six datasets and three baseline models. The results suggest that our proposed FS-BAN consistently improves the generalization performance of baseline models and achieves state-of-the-art accuracy for DG-FSC. Project Page: yunqing-me.github.io/Born-Again-FS/.
Yunqing Zhao, Ngai-Man Cheung
IEEE Trans. Image Process.2
2023 Multimodal Mutual Information Maximization: A Novel Approach for Unsupervised Deep Cross-Modal Hashing
abstract
In this article, we adopt the maximizing mutual information (MI) approach to tackle the problem of unsupervised learning of binary hash codes for efficient cross-modal retrieval. We proposed a novel method, dubbed cross-modal info-max hashing (CMIMH). First, to learn informative representations that can preserve both intramodal and intermodal similarities, we leverage the recent advances in estimating variational lower bound of MI to maximizing the MI between the binary representations and input features and between binary representations of different modalities. By jointly maximizing these MIs under the assumption that the binary representations are modeled by multivariate Bernoulli distributions, we can learn binary representations, which can preserve both intramodal and intermodal similarities, effectively in a mini-batch manner with gradient descent. Furthermore, we find out that trying to minimize the modality gap by learning similar binary representations for the same instance from different modalities could result in less informative representations. Hence, balancing between reducing the modality gap and losing modality-private information is important for the cross-modal retrieval tasks. Quantitative evaluations on standard benchmark datasets demonstrate that the proposed method consistently outperforms other state-of-the-art cross-modal retrieval methods.
Tuan Hoang, Thanh-Toan Do, Tam V. Nguyen 0002, Ngai-Man Cheung
IEEE Trans. Neural Networks Learn. Syst.4
2022 Graph-Wise Common Latent Factor Extraction for Unsupervised Graph Representation Learning
abstract
Unsupervised graph-level representation learning plays a crucial role in a variety of tasks such as molecular property prediction and community analysis, especially when data annotation is expensive. Currently, most of the best-performing graph embedding methods are based on Infomax principle. The performance of these methods highly depends on the selection of negative samples and hurt the performance, if the samples were not carefully selected. Inter-graph similarity-based methods also suffer if the selected set of graphs for similarity matching is low in quality. To address this, we focus only on utilizing the current input graph for embedding learning. We are motivated by an observation from real-world graph generation processes where the graphs are formed based on one or more global factors which are common to all elements of the graph (e.g., topic of a discussion thread, solubility level of a molecule). We hypothesize extracting these common factors could be highly beneficial. Hence, this work proposes a new principle for unsupervised graph representation learning: Graph-wise Common latent Factor EXtraction (GCFX). We further propose a deep model for GCFX, deepGCFX, based on the idea of reversing the above-mentioned graph generation process which could explicitly extract common latent factors from an input graph and achieve improved results on downstream tasks to the current state-of-the-art. Through extensive experiments and analysis, we demonstrate that, while extracting common latent factors is beneficial for graph-level tasks to alleviate distractions caused by local variations of individual nodes or local neighbourhoods, it also benefits node-level tasks by enabling long-range node dependencies, especially for disassortative graphs.
Thilini Cooray, Ngai-Man Cheung
AAAI2
2022 A Closer Look at Few-shot Image Generation
abstract
Modern GANs excel at generating high quality and diverse images. However, when transferring the pretrained GANs on small target data (e.g., 10-shot), the generator tends to replicate the training samples. Several methods have been proposed to address this few-shot image generation task, but there is a lack of effort to analyze them under a unified framework. As our first contribution, we propose a framework to analyze existing methods during the adaptation. Our analysis discovers that while some methods have disproportionate focus on diversity preserving which impede quality improvement, all methods achieve similar quality after convergence. Therefore, the better methods are those that can slow down diversity degradation. Furthermore, our analysis reveals that there is still plenty of room to further slow down diversity degradation. Informed by our analysis and to slow down the diversity degradation of the target generator during adaptation, our second contribution proposes to apply mutual information (MI) maximization to retain the source domain's rich multi-level diversity information in the target domain generator. We propose to perform MI maximization by contrastive loss (CL), leverage the generator and discriminator as two feature encoders to extract different multi-level features for computing CL. We refer to our method as Dual Contrastive Learning (DCL). Extensive experiments on several public datasets show that, while leading to a slower diversity-degrading generator during adaptation, our proposed DCL brings visually pleasant quality and state-of-the-art quantitative performance.
Yunqing Zhao, Henghui Ding, Houjing Huang, Ngai-Man Cheung
CVPR4
2022 Discovering Transferable Forensic Features for CNN-Generated Images Detection
Keshigeyan Chandrasegaran, Ngoc-Trung Tran, Alexander Binder, Ngai-Man Cheung
ECCV (15)4
2022 Revisiting Label Smoothing and Knowledge Distillation Compatibility: What was Missing?
abstract
This work investigates the compatibility between label smoothing (LS) and knowledge distillation (KD). Contemporary findings addressing this thesis statement take dichotomous standpoints: Muller et al. (2019) and Shen et al. (2021b). Critically, there is no effort to understand and resolve these contradictory findings, leaving the primal question \text{-} to smooth or not to smooth a teacher network? \text{-} unanswered. The main contributions of our work are the discovery, analysis and validation of systematic diffusion as the missing concept which is instrumental in understanding and resolving these contradictory findings. This systematic diffusion essentially curtails the benefits of distilling from an LS-trained teacher, thereby rendering KD at increased temperatures ineffective. Our discovery is comprehensively supported by large-scale experiments, analyses and case studies including image classification, neural machine translation and compact student distillation tasks spanning across multiple datasets and teacher-student architectures. Based on our analysis, we suggest practitioners to use an LS-trained teacher with a low-temperature transfer to achieve high performance students. Code and models are available at https://keshik6.github.io/revisiting-ls-kd-compatibility/
Keshigeyan Chandrasegaran, Ngoc-Trung Tran, Yunqing Zhao, Ngai-Man Cheung
ICML4
2022 Few-shot Image Generation via Adaptation-Aware Kernel Modulation
abstract
Few-shot image generation (FSIG) aims to learn to generate new and diverse samples given an extremely limited number of samples from a domain, e.g., 10 training samples. Recent work has addressed the problem using transfer learning approach, leveraging a GAN pretrained on a large-scale source domain dataset and adapting that model to the target domain based on very limited target domain samples. Central to recent FSIG methods are knowledge preserving criteria, which aim to select a subset of source model's knowledge to be preserved into the adapted model. However, a major limitation of existing methods is that their knowledge preserving criteria consider only source domain/source task, and they fail to consider target domain/adaptation task in selecting source model's knowledge, casting doubt on their suitability for setups of different proximity between source and target domain. Our work makes two contributions. As our first contribution, we re-visit recent FSIG works and their experiments. Our important finding is that, under setups which assumption of close proximity between source and target domains is relaxed, existing state-of-the-art (SOTA) methods which consider only source domain/source task in knowledge preserving perform no better than a baseline fine-tuning method. To address the limitation of existing methods, as our second contribution, we propose Adaptation-Aware kernel Modulation (AdAM) to address general FSIG of different source-target domain proximity. Extensive experimental results show that the proposed method consistently achieves SOTA performance across source/target domains of different proximity, including challenging setups when source and target domains are more apart. Project Page: https://yunqing-me.github.io/AdAM/
Yunqing Zhao, Keshigeyan Chandrasegaran, Milad Abdollahzadeh, Ngai-Man Cheung
NeurIPS4
2022 Shell Theory: A Statistical Model of Reality
abstract
The foundational assumption of machine learning is that the data under consideration is separable into classes; while intuitively reasonable, separability constraints have proven remarkably difficult to formulate mathematically. We believe this problem is rooted in the mismatch between existing statistical techniques and commonly encountered data; object representations are typically high dimensional but statistical techniques tend to treat high dimensions a degenerate case. To address this problem, we develop a dedicated statistical framework for machine learning in high dimensions. The framework derives from the observation that object relations form a natural hierarchy; this leads us to model objects as instances of a high dimensional, hierarchal generative processes. Using a distance based statistical technique, also developed in this paper, we show that in such generative processes, instances of each process in the hierarchy, are almost-always encapsulated by a distinctive-shell that excludes almost-all other instances. The result is shell theory, a statistical machine learning framework in which separability constraints (distinctive-shells) are formally derived from the assumed generative process.
Wen-Yan Lin, Changhao Ren, Ngai-Man Cheung, Hongdong Li, Yasuyuki Matsushita
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 Toward Scalable and Unified Example-Based Explanation and Outlier Detection
abstract
When neural networks are employed for high-stakes decision-making, it is desirable that they provide explanations for their prediction in order for us to understand the features that have contributed to the decision. At the same time, it is important to flag potential outliers for in-depth verification by domain experts. In this work we propose to unify two differing aspects of explainability with outlier detection. We argue for a broader adoption of prototype-based student networks capable of providing an example-based explanation for their prediction and at the same time identify regions of similarity between the predicted sample and the examples. The examples are real prototypical cases sampled from the training set via a novel iterative prototype replacement algorithm. Furthermore, we propose to use the prototype similarity scores for identifying outliers. We compare performance in terms of the classification, explanation quality and outlier detection of our proposed network with baselines. We show that our prototype-based networks extending beyond similarity kernels deliver meaningful explanations and promising outlier detection results without compromising classification accuracy.
Penny Chong, Ngai-Man Cheung, Yuval Elovici, Alexander Binder
IEEE Trans. Image Process.2
2021 A Closer Look at Fourier Spectrum Discrepancies for CNN-Generated Images Detection
abstract
CNN-based generative modelling has evolved to produce synthetic images indistinguishable from real images in the RGB pixel space. Recent works have observed that CNN-generated images share a systematic shortcoming in replicating high frequency Fourier spectrum decay attributes. Furthermore, these works have successfully exploited this systematic shortcoming to detect CNN-generated images reporting up to 99% accuracy across multiple state-of-the-art GAN models.In this work, we investigate the validity of assertions claiming that CNN-generated images are unable to achieve high frequency spectral decay consistency. We meticulously construct a counterexample space of high frequency spectral decay consistent CNN-generated images emerging from our handcrafted experiments using DCGAN, LSGAN, WGAN-GP and StarGAN, where we empirically show that this frequency discrepancy can be avoided by a minor architecture change in the last upsampling operation. We subsequently use images from this counterexample space to successfully bypass the recently proposed forensics detector which leverages on high frequency Fourier spectrum decay attributes for CNN-generated image detection.Through this study, we show that high frequency Fourier spectrum decay discrepancies are not inherent characteristics for existing CNN-based generative models—contrary to the belief of some existing work—, and such features are not robust to perform synthetic image detection. Our results prompt re-thinking of using high frequency Fourier spectrum decay attributes for CNN-generated image detection. Code and models are available at https://keshik6.github.io/Fourier-Discrepancies-CNN-Detection/
Keshigeyan Chandrasegaran, Ngoc-Trung Tran, Ngai-Man Cheung
CVPR3
2021 Revisit Multimodal Meta-Learning through the Lens of Multi-Task Learning
abstract
Multimodal meta-learning is a recent problem that extends conventional few-shot meta-learning by generalizing its setup to diverse multimodal task distributions. This setup makes a step towards mimicking how humans make use of a diverse set of prior skills to learn new skills. Previous work has achieved encouraging performance. In particular, in spite of the diversity of the multimodal tasks, previous work claims that a single meta-learner trained on a multimodal distribution can sometimes outperform multiple specialized meta-learners trained on individual unimodal distributions. The improvement is attributed to knowledge transfer between different modes of task distributions. However, there is no deep investigation to verify and understand the knowledge transfer between multimodal tasks. Our work makes two contributions to multimodal meta-learning. First, we propose a method to quantify knowledge transfer between tasks of different modes at a micro-level. Our quantitative, task-level analysis is inspired by the recent transference idea from multi-task learning. Second, inspired by hard parameter sharing in multi-task learning and a new interpretation of related work, we propose a new multimodal meta-learner that outperforms existing work by considerable margins. While the major focus is on multimodal meta-learning, our work also attempts to shed light on task interaction in conventional meta-learning. The code for this project is available at https://miladabd.github.io/KML.
Milad Abdollahzadeh, Touba Malekzadeh, Ngai-Man Cheung
NeurIPS3
2021 InfoMax-GAN: Improved Adversarial Image Generation via Information Maximization and Contrastive Learning
abstract
While Geerative Adversarial Networks (GANs) are fundamental to many generative modelling applications, they suffer from numerous issues. In this work, we propose a principled framework to simultaneously mitigate two fundamental issues in GANs: catastrophic forgetting of the discriminator and mode collapse of the generator. We achieve this by employing for GANs a contrastive learning and mutual information maximization approach, and perform extensive analyses to understand sources of improvements. Our approach significantly stabilizes GAN training and improves GAN performance for image synthesis across five datasets under the same training and evaluation conditions against state-of-the-art works. In particular, compared to the state-of-the-art SSGAN, our approach does not suffer from poorer performance on image domains such as faces, and instead improves performance significantly. Our approach is simple to implement and practical: it involves only one auxiliary objective, has low computational cost, and performs robustly across a wide range of training settings and datasets without any hyperparameter tuning. For reproducibility, our code is available in the open-source GAN library, Mimicry [34].
Kwot Sin Lee, Ngoc-Trung Tran, Ngai-Man Cheung
WACV3
2021 Joint estimation of low-rank components and connectivity graph in high-dimensional graph signals: Application to brain imaging
Rui Liu 0034, Ngai-Man Cheung
Signal Process.2
2021 On Data Augmentation for GAN Training
abstract
Recent successes in Generative Adversarial Networks (GAN) have affirmed the importance of using more data in GAN training. Yet it is expensive to collect data in many domains such as medical applications. Data Augmentation (DA) has been applied in these applications. In this work, we first argue that the classical DA approach could mislead the generator to learn the distribution of the augmented data, which could be different from that of the original data. We then propose a principled framework, termed Data Augmentation Optimized for GAN (DAG), to enable the use of augmented data in GAN training to improve the learning of the original distribution. We provide theoretical analysis to show that using our proposed DAG aligns with the original GAN in minimizing the Jensen-Shannon (JS) divergence between the original distribution and model distribution. Importantly, the proposed DAG effectively leverages the augmented data to improve the learning of discriminator and generator. We conduct experiments to apply DAG to different GAN models: unconditional GAN, conditional GAN, self-supervised GAN and CycleGAN using datasets of natural images and medical images. The results show that DAG achieves consistent and considerable improvements across these models. Furthermore, when DAG is used in some GAN models, the system establishes state-of-the-art Fréchet Inception Distance (FID) scores. Our code is available (https://github.com/tntrung/dag-gans).
Ngoc-Trung Tran, Viet-Hung Tran, Ngoc-Bao Nguyen, Trung-Kien Nguyen, Ngai-Man Cheung
IEEE Trans. Image Process.5
2020 Attention-Based Context Aware Reasoning for Situation Recognition
abstract
Situation Recognition (SR) is a fine-grained action recognition task where the model is expected to not only predict the salient action of the image, but also predict values of all associated semantic roles of the action. Predicting semantic roles is very challenging: a vast variety of possibilities can be the match for a semantic role. Existing work has focused on dependency modelling architectures to solve this issue. Inspired by the success achieved by query-based visual reasoning (e.g., Visual Question Answering), we propose to address semantic role prediction as a query-based visual reasoning problem. However, existing query-based reasoning methods have not considered handling of inter-dependent queries which is a unique requirement of semantic role prediction in SR. Therefore, to the best of our knowledge, we propose the first set of methods to address inter-dependent queries in query-based visual reasoning. Extensive experiments demonstrate the effectiveness of our proposed method which achieves outstanding performance on Situation Recognition task. Furthermore, leveraging query inter-dependency, our methods improve upon a state-of-the-art method that answers queries separately. Our code: https://github.com/thilinicooray/context-aware-reasoning-for-sr.
Thilini Cooray, Ngai-Man Cheung, Wei Lu 0011
CVPR2
2020 Attentive Weights Generation for Few Shot Learning via Information Maximization
abstract
Few shot image classification aims at learning a classifier from limited labeled data. Generating the classification weights has been applied in many meta-learning methods for few shot image classification due to its simplicity and effectiveness. In this work, we present Attentive Weights Generation for few shot learning via Information Maximization (AWGIM), which introduces two novel contributions: i) Mutual information maximization between generated weights and data within the task; this enables the generated weights to retain information of the task and the specific query sample. ii) Self-attention and cross-attention paths to encode the context of the task and individual queries. Both two contributions are shown to be very effective in extensive experiments. Overall, AWGIM is competitive with state-of-the-art. Code is available at https://github.com/Yiluan/AWGIM.
Yiluan Guo, Ngai-Man Cheung
CVPR2
2020 Explanation-Guided Training for Cross-Domain Few-Shot Classification
abstract
Cross-domain few-shot classification task (CD-FSC) combines few-shot classification with the requirement to generalize across domains represented by datasets. This setup faces challenges originating from the limited labeled data in each class and, additionally, from the domain shift between training and test sets. In this paper, we introduce a novel training approach for existing FSC models. It leverages on the explanation scores, obtained from existing explanation methods when applied to the predictions of FSC models, computed for intermediate feature maps of the models. Firstly, we tailor the layer-wise relevance propagation (LRP) method to explain the predictions of FSC models. Secondly, we develop a model-agnostic explanation-guided training strategy that dynamically finds and emphasizes the features which are important for the predictions. Our contribution does not target a novel explanation method but lies in a novel application of explanations for the training phase. We show that explanation-guided training effectively improves the model generalization. We observe improved accuracy for three different FSC models: RelationNet, cross attention network, and a graph neural network-based formulation, on five few-shot learning datasets: miniImagenet, CUB, Cars, Places, and Plantae. The source code is available https://github.com/SunJiamei/few-shot-lrp-guided.
Jiamei Sun, Sebastian Lapuschkin, Wojciech Samek, Yunqing Zhao, Ngai-Man Cheung, Alexander Binder
ICPR5
2020 Direct Quantization for Training Highly Accurate Low Bit-width Deep Neural Networks
abstract
This paper proposes two novel techniques to train deep convolutional neural networks with low bit-width weights and activations. First, to obtain low bit-width weights, most existing methods obtain the quantized weights by performing quantization on the full-precision network weights. However, this approach would result in some mismatch: the gradient descent updates full-precision weights, but it does not update the quantized weights. To address this issue, we propose a novel method that enables direct updating of quantized weights with learnable quantization levels to minimize the cost function using gradient descent. Second, to obtain low bit-width activations, existing works consider all channels equally. However, the activation quantizers could be biased toward a few channels with high-variance. To address this issue, we propose a method to take into account the quantization errors of individual channels. With this approach, we can learn activation quantizers that minimize the quantization errors in the majority of channels. Experimental results demonstrate that our proposed method achieves state-of-the-art performance on the image classification task, using AlexNet, ResNet and MobileNetV2 architectures on CIFAR-100 and ImageNet datasets.
Tuan Hoang, Thanh-Toan Do, Tam V. Nguyen 0002, Ngai-Man Cheung
IJCAI4
2020 Reproducibility Companion Paper: Selective Deep Convolutional Features for Image Retrieval
abstract
In this companion paper, firstly, we briefly summarize the contributions of our main manuscript: Selective Deep Convolutional Features for Image Retrieval, published in ACM MultiMedia 2017. In addition, we provide detail instructions together with pre-configured MATLAB scripts which allow experiments to be executed and to reproduce the results reported in our main manuscript effortlessly. The source code is available at https://github.com/hnanhtuan/selectiveConvFeatures_ACMMM_reproducibility.
Tuan Hoang, Thanh-Toan Do, Ngai-Man Cheung, Michael Riegler 0001, Jan Zahálka
ACM Multimedia3
2020 Simultaneous compression and quantization: A joint approach for efficient unsupervised hashing
Tuan Hoang, Thanh-Toan Do, Huu Le, Dang-Khoa Le Tan, Ngai-Man Cheung
Comput. Vis. Image Underst.5
2020 TEAGS: time-aware text embedding approach to generate subgraphs
Saeid Hosseini, Saeed Najafi Pour, Ngai-Man Cheung, Hongzhi Yin, Mohammadreza Kangavari, Xiaofang Zhou 0001
Data Min. Knowl. Discov.3
2020 Unsupervised Deep Cross-modality Spectral Hashing
abstract
This paper presents a novel framework, namely Deep Cross-modality Spectral Hashing (DCSH), to tackle the unsupervised learning problem of binary hash codes for efficient cross-modal retrieval. The framework is a two-step hashing approach which decouples the optimization into (1) binary optimization and (2) hashing function learning. In the first step, we propose a novel spectral embedding-based algorithm to simultaneously learn single-modality and binary cross-modality representations. While the former is capable of well preserving the local structure of each modality, the latter reveals the hidden patterns from all modalities. In the second step, to learn mapping functions from informative data inputs (images and word embeddings) to binary codes obtained from the first step, we leverage the powerful CNN for images and propose a CNN-based deep architecture to learn text modality. Quantitative evaluations on three standard benchmark datasets demonstrate that the proposed DCSH method consistently outperforms other state-of-the-art methods.
Tuan Hoang, Thanh-Toan Do, Tam V. Nguyen 0002, Ngai-Man Cheung
IEEE Trans. Image Process.4
2020 Compact Hash Code Learning With Binary Deep Neural Network
abstract
Learning compact binary codes for image retrieval problem using deep neural networks has recently attracted increasing attention. However, training deep hashing networks is challenging due to the binary constraints on the hash codes. In this paper, we propose deep network models and learning algorithms for learning binary hash codes given image representations under both unsupervised and supervised manners. The novelty of our network design is that we constrain one hidden layer to directly output the binary codes. This design has overcome a challenging problem in some previous works: optimizing non-smooth objective functions because of binarization. In addition, we propose to incorporate independence and balance properties in the direct and strict forms into the learning schemes. We also include a similarity preserving property in our objective functions. The resulting optimizations involving these binary, independence, and balance constraints are difficult to solve. To tackle this difficulty, we propose to learn the networks with alternating optimization and careful relaxation. Furthermore, by leveraging the powerful capacity of convolutional neural networks, we propose an end-to-end architecture that jointly learns to extract visual features and produce binary hash codes. Experimental results for the benchmark datasets show that the proposed methods compare favorably or outperform the state of the art.
Thanh-Toan Do, Tuan Hoang, Dang-Khoa Le Tan, Anh-Dzung Doan, Ngai-Man Cheung
IEEE Trans. Multim.5
2019 Improving GAN with Neighbors Embedding and Gradient Matching
abstract
We propose two new techniques for training Generative Adversarial Networks (GANs) in the unsupervised setting. Our objectives are to alleviate mode collapse in GAN and improve the quality of the generated samples. First, we propose neighbor embedding, a manifold learning-based regularization to explicitly retain local structures of latent samples in the generated samples. This prevents generator from producing nearly identical data samples from different latent samples, and reduces mode collapse. We propose an inverse t-SNE regularizer to achieve this. Second, we propose a new technique, gradient matching, to align the distributions of the generated samples and the real samples. As it is challenging to work with high-dimensional sample distributions, we propose to align these distributions through the scalar discriminator scores. We constrain the difference between the discriminator scores of the real samples and generated ones. We further constrain the difference between the gradients of these discriminator scores. We derive these constraints from Taylor approximations of the discriminator function. We perform experiments to demonstrate that our proposed techniques are computationally simple and easy to be incorporated in existing systems. When Gradient matching and Neighbour embedding are applied together, our GN-GAN achieves outstanding results on 1D/2D synthetic, CIFAR-10 and STL-10 datasets, e.g. FID score of 30.80 for the STL-10 dataset. Our code is available at: https://github.com/tntrung/gan
Ngoc-Trung Tran, Tuan-Anh Bui, Ngai-Man Cheung
AAAI3
2019 SDRSAC: Semidefinite-Based Randomized Approach for Robust Point Cloud Registration Without Correspondences
abstract
This paper presents a novel randomized algorithm for robust point cloud registration without correspondences. Most existing registration approaches require a set of putative correspondences obtained by extracting invariant descriptors. However, such descriptors could become unreliable in noisy and contaminated settings. In these settings, methods that directly handle input point sets are preferable. Without correspondences, however, conventional randomized techniques require a very large number of samples in order to reach satisfactory solutions. In this paper, we propose a novel approach to address this problem. In particular, our work enables the use of randomized methods for point cloud registration without the need of putative correspondences. By considering point cloud alignment as a special instance of graph matching and employing an efficient semi-definite relaxation, we propose a novel sampling mechanism, in which the size of the sampled subsets can be larger-than-minimal. Our tight relaxation scheme enables fast rejection of the outliers in the sampled sets, resulting in high quality hypotheses. We conduct extensive experiments to demonstrate that our approach outperforms other state-of-the-art methods. Importantly, our proposed method serves as a generic framework which can be extended to problems with known correspondences.
Huu Le, Thanh-Toan Do, Tuan Hoang, Ngai-Man Cheung
CVPR4
2019 Deep Clustering by Gaussian Mixture Variational Autoencoders With Graph Embedding
abstract
We propose DGG: Deep clustering via a Gaussian-mixture variational autoencoder (VAE) with Graph embedding. To facilitate clustering, we apply Gaussian mixture model (GMM) as the prior in VAE. To handle data with complex spread, we apply graph embedding. Our idea is that graph information which captures local data structures is an excellent complement to deep GMM. Combining them facilitates the network to learn powerful representations that follow global model and local structural constraints. Therefore, our method unifies model-based and similarity-based approaches for clustering. To combine graph embedding with probabilistic deep GMM, we propose a novel stochastic extension of graph embedding: we treat samples as nodes on a graph and minimize the weighted distance between their posterior distributions. We apply Jenson-Shannon divergence as the distance. We combine the divergence minimization with the log-likelihood maximization of the deep GMM. We derive formulations to obtain an unified objective that enables simultaneous deep representation learning and clustering. Our experimental results show that our proposed DGG outperforms recent deep Gaussian mixture methods (model-based) and deep spectral clustering (similarity-based). Our results highlight advantages of combining model-based and similarity-based clustering as proposed in this work. Our code is published here: https:// github.com/dodoyang0929/DGG.git.
Linxiao Yang, Ngai-Man Cheung, Jiaying Li 0001, Jun Fang 0001
ICCV2
2019 Few-Shot and Many-Shot Fusion Learning in Mobile Visual Food Recognition
abstract
Mobile visual food recognition is emerging as an important application in food logging and dietary monitoring in recent years. Existing food recognition methods use conventional many-shot learning to train a large backbone network, which refers to the use of sufficient number of training data to train the network. However, these methods firstly do not consider the cases where certain food categories have limited training data. Therefore, they cannot use the conventional training using many-shot learning. Further, existing solutions focus on improving the food recognition performance by implementing state-of-the-art large full networks, and do not pay much attention to reduce the size and computational cost of the network. As a result, they are not amenable for deployment on mobile devices. In this paper, we address these issues by proposing a new few-shot and many-shot fusion learning for mobile visual food recognition, it has a compact framework and is able to learn from existing dataset categories, and also new food categories given only a few sample images. We construct a new Indian food dataset called NTU-IndianFood107 in order to evaluate the performance of the proposed method. The dataset has two parts: (i) a Base Dataset of 83 classes of Indian food images with over 600 images per class to perform many-shot learning, and (ii) a Food Diary of 24 classes captured in restaurants with limited number to simulate the few-shot learning on new food categories. The proposed fusion method achieves a Top-1 classification accuracy of 72.0% on the new dataset.
Kim-Hui Yap, Alex Chichung Kot, Ling-Yu Duan, Ngai-Man Cheung
ISCAS5
2019 A Neural Attention Model for Real-Time Network Intrusion Detection
abstract
The diversity and ever-evolving nature of network intrusion attacks has made defense a real challenge for security practitioners. Recent research in the domain of Network-based Intrusion Detection System has mainly focused on adopting a flow-based approach when extracting features from raw packets. One drawback of this is that attack detection can only be carried out after the flow has ended. In this work, we present a new technique based on the neural attention mechanism; unlike many existing solutions, our technique can be applied for real-time attack detection since it uses time slot-based features. The proposed solution is a modified version of the transformer model which has been proposed and used in the language translation domain. We conduct experiments on a dataset extracted from a recent repository network traffic containing several kinds of network attack. We use the "bidirectional LSTM" and "conditional random fields" models as baseline for comparison and our performance results demonstrate that the proposed solution significantly outperforms the two baselines in terms of precision, recall, and false positive rates. In addition, we show that our solution is more computationally efficient than the bidirectional LSTM model as a result of the removal of recurrent layers.
Mengxuan Tan, Alfonso Iacovazzi, Ngai-Man Cheung, Yuval Elovici
LCN3
2019 Self-supervised GAN: Analysis and Improvement with Multi-class Minimax Game
abstract
Self-supervised (SS) learning is a powerful approach for representation learning using unlabeled data. Recently, it has been applied to Generative Adversarial Networks (GAN) training. Specifically, SS tasks were proposed to address the catastrophic forgetting issue in the GAN discriminator. In this work, we perform an in-depth analysis to understand how SS tasks interact with learning of generator. From the analysis, we identify issues of SS tasks which allow a severely mode-collapsed generator to excel the SS tasks. To address the issues, we propose new SS tasks based on a multi-class minimax game. The competition between our proposed SS tasks in the game encourages the generator to learn the data distribution and generate diverse samples. We provide both theoretical and empirical analysis to support that our proposed SS tasks have better convergence property. We conduct experiments to incorporate our proposed SS tasks into two different GAN baseline models. Our approach establishes state-of-the-art FID scores on CIFAR-10, CIFAR-100, STL-10, CelebA, Imagenet $32\times32$ and Stacked-MNIST datasets, outperforming existing works by considerable margins in some cases. Our unconditional GAN model approaches performance of conditional GAN without using labeled data. Our code: \url{https://github.com/tntrung/msgan}
Ngoc-Trung Tran, Viet-Hung Tran, Ngoc-Bao Nguyen, Linxiao Yang, Ngai-Man Cheung
NeurIPS5
2019 Binary Constrained Deep Hashing Network for Image Retrieval Without Manual Annotation
abstract
Learning compact binary codes for image retrieval task using deep neural networks has attracted increasing attention recently. However, training deep hashing networks for the task is challenging due to the binary constraints on the hash codes, the similarity preserving property, and the requirement for a vast amount of labelled images. To the best of our knowledge, none of the existing methods has tackled all of these challenges completely in a unified framework. In this work, we propose a novel end-to-end deep learning approach for the task, in which the network is trained to produce binary codes directly from image pixels without the need o f manual annotation. In particular, to deal with the non-smoothness of binary constraints, we propose a novel pairwise constrained loss function, which simultaneously encodes the distances between pairs of hash codes, and the binary quantization error. In order to train the network with the proposed loss function, we propose an efficient parameter learning algorithm. In addition, to provide similar / dissimilar training images to train the network, we exploit 3D models reconstructed from unlabelled images for automatic generation of enormous training image pairs. The extensive experiments on image retrieval benchmark datasets demonstrate the improvements of the proposed method over the state-of-the-art compact representation methods on the image retrieval problem.
Thanh-Toan Do, Tuan Hoang, Dang-Khoa Le Tan, Trung Pham, Huu Le, Ngai-Man Cheung, Ian D. Reid 0001
WACV6
2019 Simultaneous Feature Aggregating and Hashing for Compact Binary Code Learning
abstract
Representing images by compact hash codes is an attractive approach for large-scale content-based image retrieval. In most state-of-the-art hashing-based image retrieval systems, for each image, local descriptors are first aggregated as a global representation vector. This global vector is then subjected to a hashing function to generate a binary hash code. In previous works, the aggregating and the hashing processes are designed independently. Hence, these frameworks may generate suboptimal hash codes. In this paper, we first propose a novel unsupervised hashing framework in which feature aggregating and hashing are designed simultaneously and optimized jointly. Specifically, our joint optimization generates aggregated representations that can be better reconstructed by some binary codes. This leads to more discriminative binary hash codes and improved retrieval accuracy. In addition, the proposed method is flexible. It can be extended for supervised hashing. When the data label is available, the framework can be adapted to learn binary codes which minimize the reconstruction loss with respect to label vectors. Furthermore, we also propose a fast version of the state-of-the-art hashing method Binary Autoencoder to be used in our proposed frameworks. Extensive experiments on benchmark datasets under various settings show that the proposed methods outperform the state-of-the-art unsupervised and supervised hashing methods.
Thanh-Toan Do, Khoa Le, Tuan Hoang, Huu Le, Tam V. Nguyen 0002, Ngai-Man Cheung
IEEE Trans. Image Process.6
2019 On-Device Scalable Image-Based Localization via Prioritized Cascade Search and Fast One-Many RANSAC
abstract
We present the design of an entire on-device system for large-scale urban localization using images. The proposed design integrates compact image retrieval and 2D-3D correspondence search to estimate the location in extensive city regions. Our design is GPS agnostic and does not require network connection. In order to overcome the resource constraints of mobile devices, we propose a system design that leverages the scalability advantage of image retrieval and accuracy of 3D model-based localization. Furthermore, we propose a new hashing-based cascade search for fast computation of 2D-3D correspondences. In addition, we propose a new one-many RANSAC for accurate pose estimation. The new one-many RANSAC addresses the challenge of repetitive building structures (e.g. windows and balconies) in urban localization. Extensive experiments demonstrate that our 2D-3D correspondence search achieves the state-of-the-art localization accuracy on multiple benchmark datasets. Furthermore, our experiments on a large Google street view image dataset show the potential of large-scale localization entirely on a typical mobile device.
Ngoc-Trung Tran, Dang-Khoa Le Tan, Anh-Dzung Doan, Thanh-Toan Do, Tuan-Anh Bui, Mengxuan Tan, Ngai-Man Cheung
IEEE Trans. Image Process.7
2019 From Selective Deep Convolutional Features to Compact Binary Representations for Image Retrieval
abstract
In the large-scale image retrieval task, the two most important requirements are the discriminability of image representations and the efficiency in computation and storage of representations. Regarding the former requirement, Convolutional Neural Network is proven to be a very powerful tool to extract highly discriminative local descriptors for effective image search. Additionally, to further improve the discriminative power of the descriptors, recent works adopt fine-tuned strategies. In this article, taking a different approach, we propose a novel, computationally efficient, and competitive framework. Specifically, we first propose various strategies to compute masks, namely, SIFT-masks , SUM-mask , and MAX-mask , to select a representative subset of local convolutional features and eliminate redundant features. Our in-depth analyses demonstrate that proposed masking schemes are effective to address the burstiness drawback and improve retrieval accuracy. Second, we propose to employ recent embedding and aggregating methods that can significantly boost the feature discriminability. Regarding the computation and storage efficiency, we include a hashing module to produce very compact binary image representations. Extensive experiments on six image retrieval benchmarks demonstrate that our proposed framework achieves the state-of-the-art retrieval performances.
Thanh-Toan Do, Tuan Hoang, Dang-Khoa Le Tan, Huu Le, Tam V. Nguyen 0002, Ngai-Man Cheung
ACM Trans. Multim. Comput. Commun. Appl.6
2019 Leveraging multi-aspect time-related influence in location recommendation
Saeid Hosseini, Hongzhi Yin, Xiaofang Zhou 0001, Shazia Sadiq, Mohammadreza Kangavari, Ngai-Man Cheung
World Wide Web6
2018 Adaptive Quantization for Deep Neural Network
abstract
In recent years Deep Neural Networks (DNNs) have been rapidly developed in various applications, together with increasingly complex architectures. The performance gain of these DNNs generally comes with high computational costs and large memory consumption, which may not be affordable for mobile platforms. Deep model quantization can be used for reducing the computation and memory costs of DNNs, and deploying complex DNNs on mobile equipment. In this work, we propose an optimization framework for deep model quantization. First, we propose a measurement to estimate the effect of parameter quantization errors in individual layers on the overall model prediction accuracy. Then, we propose an optimization process based on this measurement for finding optimal quantization bit-width for each layer. This is the first work that theoretically analyse the relationship between parameter quantization errors of individual layers and model accuracy. Our new quantization algorithm outperforms previous quantization optimization methods, and achieves 20-40% higher compression rate compared to equal bit-width quantization at the same model prediction accuracy.
Yiren Zhou, Seyed-Mohsen Moosavi-Dezfooli, Ngai-Man Cheung, Pascal Frossard
AAAI3
2018 Efficient and Deep Person Re-Identification Using Multi-Level Similarity
abstract
Person Re-Identification (ReID) requires comparing two images of person captured under different conditions. Existing work based on neural networks often computes the similarity of feature maps from one single convolutional layer. In this work, we propose an efficient, end-to-end fully convolutional Siamese network that computes the similarities at multiple levels. We demonstrate that multi-level similarity can improve the accuracy considerably using low-complexity network structures in ReID problem. Specifically, first, we use several convolutional layers to extract the features of two input images. Then, we propose Convolution Similarity Network to compute the similarity score maps for the inputs. We use spatial transformer networks (STNs) to determine spatial attention. We propose to apply efficient depth-wise convolution to compute the similarity. The proposed Convolution Similarity Networks can be inserted into different convolutional layers to extract visual similarities at different levels. Furthermore, we use an improved ranking loss to further improve the performance. Our work is the first to propose to compute visual similarities at low, middle and high levels for ReID. With extensive experiments and analysis, we demonstrate that our system, compact yet effective, can achieve competitive results with much smaller model size and computational complexity.
Yiluan Guo, Ngai-Man Cheung
CVPR2
2018 Exploiting Reshaping Subgraphs from Bilateral Propagation Graphs
Saeid Hosseini, Hongzhi Yin, Ngai-Man Cheung, Kan Pak Leng, Yuval Elovici, Xiaofang Zhou 0001
DASFAA (1)3
2018 Dist-GAN: An Improved GAN Using Distance Constraints
Ngoc-Trung Tran, Tuan-Anh Bui, Ngai-Man Cheung
ECCV (14)3
2018 Fine-Grained Wound Tissue Analysis Using Deep Neural Network
abstract
Tissue assessment for chronic wounds is the basis of wound grading and selection of treatment approaches. While several image processing approaches have been proposed for automatic wound tissue analysis, there has been a shortcoming in these approaches for clinical practices. In particular, seemingly, all previous approaches have assumed only 3 tissue types in the chronic wounds, while these wounds commonly exhibit 7 distinct tissue types that presence of each one changes the treatment procedure. In this paper, for the first time, we investigate the classification of 7 wound tissue types. We work with wound professionals to build a new database of 7 types of wound tissue. We propose to use pre-trained deep neural networks for feature extraction and classification at the patch-level. We perform experiments to demonstrate that our approach outperforms other state-of-the-art. We will make our database publicly available to facilitate research in wound assessment.
Hossein Nejati, Hamed Alizadeh Ghazijahani, Milad Abdollahzadeh, Touba Malekzadeh, Ngai-Man Cheung, Kheng Hock Lee, Lian Leng Low
ICASSP5
2018 DOPING: Generative Data Augmentation for Unsupervised Anomaly Detection with GAN
abstract
Recently, the introduction of the generative adversarial network (GAN) and its variants has enabled the generation of realistic synthetic samples, which has been used for enlarging training sets. Previous work primarily focused on data augmentation for semi-supervised and supervised tasks. In this paper, we instead focus on unsupervised anomaly detection and propose a novel generative data augmentation framework optimized for this task. By using a GAN variant known as the adversarial autoencoder (AAE), we impose a distribution on the latent space of the dataset and systematically sample the latent space to generate artificial samples. To the best of our knowledge, our method is the first data augmentation technique focused on improving performance in unsupervised anomaly detection. We validate our method by demonstrating consistent improvements across several real-world datasets.
Swee Kiat Lim, Yi Loo, Ngoc-Trung Tran, Ngai-Man Cheung, Gemma Roig, Yuval Elovici
ICDM4
2018 Supervised Hashing with End-to-End Binary Deep Neural Network
abstract
Image hashing is a popular technique applied to large scale content-based visual retrieval due to its compact and efficient binary codes. Our work proposes a new end-to-end deep network architecture for supervised hashing which directly learns binary codes from input images and maintains hashing properties, namely similarity preservation, independence, and balancing. Furthermore, we also propose a new learning scheme that copes with the binary constrained loss function. The proposed algorithm not only is scalable for learning over large-scale datasets but also outperforms state-of-the-art supervised hashing methods, which are illustrated throughout extensive experiments from various image retrieval benchmarks.
Dang-Khoa Le Tan, Thanh-Toan Do, Ngai-Man Cheung
ICIP3
2018 Deep Adaptive Temporal Pooling for Activity Recognition
abstract
Deep neural networks have recently achieved competitive accuracy for human activity recognition. However, there is room for improvement, especially in modeling of long-term temporal importance and determining the activity relevance of different temporal segments in a video. To address this problem, we propose a learnable and differentiable module: Deep Adaptive Temporal Pooling (DATP). DATP applies a self-attention mechanism to adaptively pool the classification scores of different video segments. Specifically, using frame-level features, DATP regresses importance of different temporal segments, and generates weights for them. Remarkably, DATP is trained using only the video-level label. There is no need of additional supervision except video-level activity class label. We conduct extensive experiments to investigate various input features and different weight models. Experimental results show that DATP can learn to assign large weights to key video segments. More importantly, DATP can improve training of frame-level feature extractor. This is because relevant temporal segments are assigned large weights during back-propagation. Overall, we achieve state-of-the-art performance on UCF101, HMDB51 and Kinetics datasets.
Sibo Song, Ngai-Man Cheung, Vijay Chandrasekhar 0001, Bappaditya Mandal
ACM Multimedia2
2018 Embedding Based on Function Approximation for Large Scale Image Search
abstract
The objective of this paper is to design an embedding method that maps local features describing an image (e.g., SIFT) to a higher dimensional representation useful for the image retrieval problem. First, motivated by the relationship between the linear approximation of a nonlinear function in high dimensional space and the state-of-the-art feature representation used in image retrieval, i.e., VLAD, we propose a new approach for the approximation. The embedded vectors resulted by the function approximation process are then aggregated to form a single representation for image retrieval. Second, in order to make the proposed embedding method applicable to large scale problem, we further derive its fast version in which the embedded vectors can be efficiently computed, i.e., in the closed-form. We compare the proposed embedding methods with the state of the art in the context of image search under various settings: when the images are represented by medium length vectors, short vectors, or binary vectors. The experimental results show that the proposed embedding methods outperform existing the state of the art on the standard public image retrieval benchmarks.
Thanh-Toan Do, Ngai-Man Cheung
IEEE Trans. Pattern Anal. Mach. Intell.2
2018 Computation and Memory Efficient Image Segmentation
abstract
In this paper, we address the segmentation problem under limited computation and memory resources. Given a segmentation algorithm, we propose a framework that can reduce its computation time and memory requirementsimultaneously, while preserving its accuracy. The proposed framework uses standard pixel-domain downsampling and includes two main steps.Coarse segmentationis first performed on the downsampled image.Refinementis then applied to the coarse segmentation results. We make two novel contributions to enable competitive accuracy using this simple framework. First, we rigorously examine the effect of downsampling on segmentation using a signal processing analysis. The analysis helps to determine theuncertain regions, which are small image regions where pixel labels are uncertain after the coarse segmentation. Second, we propose an efficient minimum spanning tree-based algorithm to propagate the labels into the uncertain regions. We perform extensive experiments using several standard data sets. The experimental results show that our segmentation accuracy is comparable to state-of-the-art methods, while requiring much less computation time and memory than those methods.
Yiren Zhou, Thanh-Toan Do, Haitian Zheng, Ngai-Man Cheung, Lu Fang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2018 Accessible Melanoma Detection Using Smartphones and Mobile Image Analysis
abstract
We investigate the design of an entire mobile imaging system for early detection of melanoma. Different from previous work, we focus on smartphone-captured visible light images. Our design addresses two major challenges. First, images acquired using a smartphone under loosely-controlled environmental conditions may be subject to various distortions, and this makes melanoma detection more difficult. Second, processing performed on a smartphone is subject to stringent computation and memory constraints. In our work, we propose a detection system that is optimized to run entirely on the resource-constrained smartphone. Our system intends to localize the skin lesion by combining a lightweight method for skin detection with a hierarchical segmentation approach using two fast segmentation methods. Moreover, we study an extensive set of image features and propose new numerical features to characterize a skin lesion. Furthermore, we propose an improved feature selection algorithm to determine a small set of discriminative features used by the final lightweight system. In addition, we study the human-computer interface (HCI) design to understand the usability and acceptance issues of the proposed system. Our extensive evaluation on an image dataset provided by National Skin Center - Singapore (117 benign nevi and 67 malignant melanoma) confirms the effectiveness of the proposed system for melanoma detection: 89.09% sensitivity at specificity ≥90%.
Thanh-Toan Do, Tuan Hoang, Victor Pomponiu, Yiren Zhou, Zhao Chen 0005, Ngai-Man Cheung, Dawn Chin-Ing Koh, Aaron Tan, Suat-Hoon Tan
IEEE Trans. Multim.6
2017 Simultaneous Feature Aggregating and Hashing for Large-Scale Image Search
Thanh-Toan Do, Dang-Khoa Le Tan, Trung T. Pham, Ngai-Man Cheung
CVPR4
2017 Detection of Internet Traffic Anomalies Using Sparse Laplacian Component Analysis
abstract
We consider the problem of anomaly detection in network traffic. It is a challenging problem because of high-dimensional and noisy nature of network traffic. A popularly used technique is subspace analysis. Principal component analysis (PCA) and its improvements have been applied for subspace analysis. In this work, we take a different approach to determine the subspace, and propose to capture the essence of the traffic using the eigenvectors of graph Laplacian, which we refer as Laplacian components (LCs). Our main contribution is to propose a regression framework to compute LCs followed by its application in anomaly detection. This framework provides much flexibility in incorporating different properties into the LCs, notably LCs with sparse loadings, which we exploit in detail. Furthermore, different from previous work that uses sample graphs to preserve local structure, we advocate modelling with a dual-input feature graph that encodes the correlation of the time series data and prior information. Therefore, the proposed model can readily incorporate the 'physics' of some applications as prior information to improve the analysis. We perform experiments on volume anomaly detection using only link-based traffic measurements. We demonstrate that the proposed model can correctly uncover the essential low-dimensional principal subspace containing the normal Internet traffic and achieve outstanding detection performance.
Manas Khatua, Seyed Hamid Safavi, Ngai-Man Cheung
GLOBECOM3
2017 Simultaneous low-rank component and graph estimation for high-dimensional graph signals: Application to brain imaging
abstract
We propose an algorithm to uncover the intrinsic low-rank component of a high-dimensional, graph-smooth and grossly-corrupted dataset, under the situations that the underlying graph is unknown. Based on a model with a low-rank component plus a sparse perturbation, and an initial graph estimation, our proposed algorithm simultaneously learns the low-rank component and refines the graph. The refined graph improves the effectiveness of the graph smoothness constraint and increases the accuracy of the low-rank estimation. We derive the learning steps using ADMM. Our evaluations using synthetic and real brain imaging data in a supervised classification task demonstrate encouraging performance.
Rui Liu 0034, Hossein Nejati, Seyed Hamid Safavi, Ngai-Man Cheung
ICASSP4
2017 On classification of distorted images with deep convolutional neural networks
abstract
Image blur and image noise are common distortions during image acquisition. In this paper, we systematically study the effect of image distortions on the deep neural network (DNN) image classifiers. First, we examine the DNN classifier performance under four types of distortions. Second, we propose two approaches to alleviate the effect of image distortion: re-training and fine-tuning with noisy images. Our results suggest that, under certain conditions, fine-tuning with noisy images can alleviate much effect due to distorted inputs, and is more practical than re-training.
Yiren Zhou, Sibo Song, Ngai-Man Cheung
ICASSP3
2017 Non-rigid Object Tracking via Deformable Patches Using Shape-Preserved KCF and Level Sets
abstract
Part-based trackers are effective in exploiting local details of the target object for robust tracking. In contrast to most existing part-based methods that divide all kinds of target objects into a number of fixed rectangular patches, in this paper, we propose a novel framework in which a set of deformable patches dynamically collaborate on tracking of non-rigid objects. In particular, we proposed a shape-preserved kernelized correlation filter (SP-KCF) which can accommodate target shape information for robust tracking. The SP-KCF is introduced into the level set framework for dynamic tracking of individual patches. In this manner, our proposed deformable patches are target-dependent, have the capability to assume complex topology, and are deformable to adapt to target variations. As these deformable patches properly capture individual target subregions, we exploit their photometric discrimination and shape variation to reveal the trackability of individual target subregions, which enables the proposed tracker to dynamically take advantage of those subregions with good trackability for target likelihood estimation. Finally the shape information of these deformable patches enables accurate object contours to be computed as the tracking output. Experimental results on the latest public sets of challenging sequences demonstrate the effectiveness of the proposed method.
Xin Sun 0003, Ngai-Man Cheung, Hongxun Yao, Yiluan Guo
ICCV2
2017 Deep neural networks on graph signals for brain imaging analysis
abstract
Brain imaging data such as EEG or MEG are high-dimensional spatiotemporal data often degraded by complex, non-Gaussian noise. For reliable analysis of brain imaging data, it is important to extract discriminative, low-dimensional intrinsic representation of the recorded data. This work proposes a new method to learn the low-dimensional representations from the noise-degraded measurements. In particular, our work proposes a new deep neural network design that integrates graph information such as brain connectivity with fully-connected layers. Our work leverages efficient graph filter design using Chebyshev polynomial and recent work on convolutional nets on graph-structured data. Our approach exploits graph structure as the prior side information, localized graph filter for feature extraction and neural networks for high capacity learning. Experiments on real MEG datasets show that our approach can extract more discriminative representations, leading to improved accuracy in a supervised classification task.
Yiluan Guo, Hossein Nejati, Ngai-Man Cheung
ICIP3
2017 Enhancing feature discrimination for unsupervised hashing
abstract
We introduce a novel approach to improve unsupervised hashing. Specifically, we propose a very efficient embedding method: Gaussian Mixture Model embedding (Gemb). The proposed method, using Gaussian Mixture Model, embeds feature vector into a low-dimensional vector and, simultaneously, enhances the discriminative property of features before passing them into hashing. Our experiment shows that the proposed method boosts the hashing performance of many state-of-the-art, e.g. Binary Autoencoder (BA) [1], Iterative Quantization (ITQ) [2], in standard evaluation metrics for the three main benchmark datasets.
Tuan Hoang, Thanh-Toan Do, Dang-Khoa Le Tan, Ngai-Man Cheung
ICIP4
2017 Selective Deep Convolutional Features for Image Retrieval
abstract
Convolutional Neural Network (CNN) is a very powerful approach to extract discriminative local descriptors for effective image search. Recent work adopts fine-tuned strategies to further improve the discriminative power of the descriptors. Taking a different approach, in this paper, we propose a novel framework to achieve competitive retrieval performance. Firstly, we propose various masking schemes, namely SIFT-mask, SUM-mask, and MAX-mask, to select a representative subset of local convolutional features and remove a large number of redundant features. We demonstrate that this can effectively address the burstiness issue and improve retrieval accuracy. Secondly, we propose to employ recent embedding and aggregating methods to further enhance feature discriminability. Extensive experiments demonstrate that our proposed framework achieves state-of-the-art retrieval accuracy.
Tuan Hoang, Thanh-Toan Do, Dang-Khoa Le Tan, Ngai-Man Cheung
ACM Multimedia4
2017 Streaming Mobile Cloud Gaming Video Over TCP With Adaptive Source-FEC Coding
abstract
Cloud gaming has emerged as a promising application to enable high-end game playing with thin clients. Transmission control protocol (TCP) is pervasively adopted as the transport-layer protocol in the mainstream cloud gaming systems for video communication. However, streaming mobile cloud gaming video using the TCP is challenged with several key technical barriers: (1) the performance limitations of wireless networks in bandwidth and reliability; (2) the high throughput demand and stringent delay constraint imposed by high-quality gaming video transmission; and (3) the deadline violations and throughput fluctuations caused by the packet retransmission and congestion control mechanisms in the TCP. To address these critical problems, this paper proposes an application-layer source-forward error correction (FEC) coding framework dubbed adaptive source-FEC coding over TCP (ESCOT). First, we analytically formulate the optimization problem of joint source-FEC coding to minimize the end-to-end distortion of real-time video communication over TCP. Second, we develop a heuristic solution for effective loss rate approximation, source rate control, and FEC coding adaptation. ESCOT is distinct from existing source-FEC coding schemes in proactively analyzing and leveraging the TCP characteristics. The proposed solution is able to effectively mitigate both consecutive and sporadic video frame drops caused by congestion and random packet losses. We conduct the performance evaluation through extensive emulations in the Exata platform using real-time gaming video encoded by the H.264 codec. Experimental results show that the ESCOT advances the state of the art with noticeable improvements in video peak signal-to-noise ratio, end-to-end delay, goodput, and frame success rate.
Jiyan Wu, Chau Yuen, Ngai-Man Cheung, Junliang Chen 0001, Chang Wen Chen
IEEE Trans. Circuits Syst. Video Technol.3
2016 Efficient Environmental Temperature Monitoring Using Compressed Sensing
abstract
This paper proposes a novel compressed sensing (CS) framework that exploits (i) the sparsity of the samples in some representation domain, and (ii) the governing structure of the samples, with a focus on temperature monitoring using Wireless Sensor Networks (WSN). WSN's have been used for temperature monitoring for automation and surveillance. Despite the benefits, WSN's have their own limitations in spatial and temporal resolution. We propose to tackle these problems by using CS for spatiotemporal temperature source reconstruction. We propose a Diffusive Compressive Sensing (DCS) [1] framework to leverage the domain knowledge to increase efficiency of classic CS.
Ali Hashemi 0002, Ngai-Man Cheung
DCC3
2016 Learning to Hash with Binary Deep Neural Network
Thanh-Toan Do, Anh-Dzung Doan, Ngai-Man Cheung
ECCV (5)3
2016 Binary Hashing with Semidefinite Relaxation and Augmented Lagrangian
Thanh-Toan Do, Anh-Dzung Doan, Duc Thanh Nguyen, Ngai-Man Cheung
ECCV (2)4
2016 Egocentric activity recognition with multimodal fisher vector
abstract
With the increasing availability of wearable devices, research on egocentric activity recognition has received much attention recently. In this paper, we build a Multimodal Egocentric Activity dataset which includes egocentric videos and sensor data of 20 fine-grained and diverse activity categories. We present a novel strategy to extract temporal trajectory-like features from sensor data. We propose to apply the Fisher Kernel framework to fuse video and temporal enhanced sensor features. Experiment results show that with careful design of feature extraction and fusion algorithm, sensor data can enhance information-rich video data. We make publicly available the Multimodal Egocentric Activity dataset to facilitate future research.
Sibo Song, Ngai-Man Cheung, Vijay Chandrasekhar 0001, Bappaditya Mandal, Jie Lin 0001
ICASSP2
2016 Deepmole: Deep neural networks for skin mole lesion classification
abstract
Nowadays, the occurrence of skin cancer cases has grown worldwide due to the extended exposure to the harmful radiation from the Sun. Most common approach to detect the malignancy of skin moles is by visual inspection performed by an expert dermatologist, using a set of specific clinical rules. Computer-aided diagnosis, based on skin mole imaging, is another concurrent method which has experienced major advancements due to improvement of imaging sensors and processing power. However, these schemes use hand-crafted features which are difficult to tune and perform poorly on new cases due to lack of generalization power. In this study we present a method that use a pretrained deep neural network (DNN) to automatically extract a set of representative features that can be later used to diagnose a sample of skin lesion for malignancy. The experimental tests carried out on a clinical dataset show that the classification performance using DNN-based features performs better than the state-of-the-art techniques.
Victor Pomponiu, Hossein Nejati, Ngai-Man Cheung
ICIP3
2016 Dimensionality reduction of brain imaging data using graph signal processing
abstract
Brain imaging data such as EEG or MEG is high-dimensional spatiotemporal measurements that commonly require dimensionality reduction before being used for further analysis or applications. This paper presents a new dimensionality reduction method based on the recent graph signal processing theory. Specifically, we focus on a task to classify the brain imaging signals recording the cortical activities in response to visual stimuli. We propose to use the resting-state measurements (i.e., before onset of the stimulus) of the subjects to build a connectivity graph. The graph Laplacian and Graph-Based Filtering (GBF) are then applied to learn the low-dimensional linear subspace for the task-state measurements (i.e., after onset of the stimulus). We investigate different techniques to build the connectivity graph suitable for this application. Compared with other dimensionality reduction methods such as Principle Component Analysis (PCA), GBF-based dimensionality reduction leverages the connectivity graph as side information to analyze the degraded, non-Gaussian noise corrupted measurements. Experimental results with synthetic and real MEG datasets are presented.
Rui Liu 0034, Hossein Nejati, Ngai-Man Cheung
ICIP3
2016 Trading Delay for Distortion in One-Way Video Communication Over the Internet
abstract
We study the problem of one-way video communication in a single-source, multiple-destination scenario over the lossy Internet. Forward error correction (FEC) coding is commonly adopted for data protection in implementing loss-resilient video transmission systems. However, the burst packet losses over the Internet frequently degrade FEC performance and induce video quality deteriorations. To address the challenging problem, we propose a novel transmission scheme dubbed trading delay for distortion (TELFORD) that includes three components: 1) adaptive multidestination status estimation; 2) delay-constrained transmission rate assignment; and 3) differentiated FEC packet spreading. We analytically formulate and solve the problem of FEC packet allocation and scheduling to minimize the end-to-end video distortion. The proposed TELFORD is able to cope with multiple-destination scheduling separately. We conduct performance evaluation through semiphysical emulations in Exata using real-time H.264 video streaming. Experimental results show that TELFORD outperforms existing error control transmission schemes in improving the video peak signal-to-noise ratio and mitigating packet transmission impairments.
Jiyan Wu, Bo Cheng 0001, Chau Yuen, Ngai-Man Cheung, Junliang Chen 0001
IEEE Trans. Circuits Syst. Video Technol.4
2016 Layered Coding for Mobile Cloud Gaming Using Scalable Blinn-Phong Lighting
abstract
In a mobile cloud gaming, high-quality, high-frame-rate game images of immense data size need to be delivered to the clients over wireless networks under stringent delay requirement. For good gaming experience, reducing the transmission bit rate of the game images is necessary. Most existing cloud gaming platforms simply employ standard, off-the-shelf video codecs for game image compression. In this paper, we propose the layered coding scheme to reduce transmission bandwidth and latency. We leverage the rendering computation of modern mobile devices to render a low-quality local game image, or the base layer (BL). Instead of sending a high-quality game image, cloud servers can send enhancement layer information, which clients can utilize to improve the quality of the BL. Central to the layered coding scheme is the design of a complexity-scalable BL rendering pipeline that can be executed on a range of power-constrained mobile devices. In this paper, we focus on the lighting stage in modern graphics rendering and propose a method to scale the popular Blinn-Phong lighting for the use in BL rendering. We derive an information-theoretic model on the Blinn-Phong lighting to estimate the rendered image entropy. The analytic model informs the optimal BL rendering design that can lead to maximum bandwidth saving subject to the constraint on the computation capability of the client. We show that the information rate of the enhancement layer could be much less than that of the high-quality game image, while the BL can be generated with only a very small amount of computation. Experiment results suggest that our analytic model is accurate in estimating. For layered coding scheme, up to 84% reduction in bandwidth usage can be achieved by sending the enhancement layer information instead of the original high-quality game images compressed by H.264/AVC.
Seong-Ping Chuah, Ngai-Man Cheung, Chau Yuen
IEEE Trans. Image Process.2
2016 Merge Frame Design for Video Stream Switching Using Piecewise Constant Functions
abstract
The ability to efficiently switch from one pre-encoded video stream to another (e.g., for bitrate adaptation or view switching) is important for many interactive streaming applications. Recently, stream-switching mechanisms based on distributed source coding (DSC) have been proposed. In order to reduce the overall transmission rate, these approaches provide a merge mechanism, where information is sent to the decoder, such that the exact same frame can be reconstructed given that any one of a known set of side information (SI) frames is available at the decoder (e.g., each SI frame may correspond to a different stream from which we are switching). However, the use of bit-plane coding and channel coding in many DSC approaches leads to complex coding and decoding. In this paper, we propose an alternative approach for merging multiple SI frames, using a piecewise constant (PWC) function as the merge operator. In our approach, for each block to be reconstructed, a series of parameters of these PWC merge functions are transmitted in order to guarantee identical reconstruction given the known SI blocks. We consider two different scenarios. In the first case, a target frame is first given, and then merge parameters are chosen, so that this frame can be reconstructed exactly at the decoder. In contrast, in the second scenario, the reconstructed frame and the merge parameters are jointly optimized to meet a rate-distortion criteria. Experiments show that for both scenarios, our proposed merge techniques can outperform both a recent approach based on DSC and the SP-frame approach in H.264, in terms of compression efficiency and decoder complexity.
Wei Dai 0002, Gene Cheung, Ngai-Man Cheung, Antonio Ortega, Oscar C. Au
IEEE Trans. Image Process.3
2016 Estimation of Virtual View Synthesis Distortion Toward Virtual View Position
abstract
We propose an analytical model to estimate the depth-error-induced virtual view synthesis distortion (VVSD) in 3D video, taking the distance between reference and virtual views (virtual view position) into account. In particular, we start with a comprehensive preanalysis and discussion over several possible VVSD scenarios. Taking intrinsic characteristic of each scenario into consideration, we specifically classify them into four clusters: 1) overlapping region; 2) disocclusion and boundary region; 3) edge region; and 4) infrequent region. We propose to model VVSD as the linear combination of the distortion under different scenarios (DDSs) weighted by the probability under different scenarios (PDSs). We show analytically that DDS and PDS can be related to the virtual view position using quadratic/biquadratic models and linear models, respectively. Experimental results verify that the proposed model is capable of estimating the relationship between VVSD and the distance between reference and virtual views. Therefore, our model can be used to inform a reference view setup for capturing, or distortion at certain virtual view positions, when depth information is compressed.
Lu Fang 0001, Yijian Xiang, Ngai-Man Cheung, Feng Wu 0001
IEEE Trans. Image Process.3
2016 Delay-Constrained High Definition Video Transmission in Heterogeneous Wireless Networks with Multi-Homed Terminals
abstract
Delivering high-quality mobile video with the limited radio resources is challenging due to the time-varying channel status and stringent Quality of Service (QoS) requirements. Multi-homing support enables mobile terminals to establish multiple simultaneous associations for enhancing transmission performance. In this paper, we study the multi-homed communication of delay-constrained High Definition (HD) video in heterogeneous wireless networks. The low-delay encoded HD video streaming consists exclusively of Intra (I) and Predicted (P) frames. In the capacity-limited wireless networks, it is highly possible the large-size I frames experience deadline violations and induce severe quality degradations. To address the challenging problem, we propose a novel scheduling framework dubbed delAy Stringent COded Transmission (ASCOT) that featured by frame-level data protection and allocation overmultiple wireless access networks. First, we perform online video distortion estimation to simulate multi-homed transmission impairments according to the feedback channel status and input video data. Second, we control the frame protection level by adapting the Forward Error Correction (FEC) coding redundancy and video data allocation to achieve target quality. The performance of the proposed ASCOT is evaluated through semi-physical emulations in Exata using real-time H.264 video streaming. Experimental results show that ASCOT outperforms existing transmission schemes in improving the video PSNR (Peak Signal-to-Noise Ratio), reducing the end-to-end delay, and increasing the goodput. Or conversely, ASCOTachieves the same video quality with approximately 20 percent bandwidth conservation.
Jiyan Wu, Chau Yuen, Ngai-Man Cheung, Junliang Chen 0001
IEEE Trans. Mob. Comput.3
2016 Modeling and Optimization of High Frame Rate Video Transmission Over Wireless Networks
abstract
High frame rate (HFR) video is emerging as a new paradigm in popular multimedia applications (e.g., cloud gaming) to achieve smooth viewing experience perceived by end-users. In the context of HFR streaming video, end-to-end distortion and sending frame rate are equally important to the perceptual quality. This study presents a modeling-based approach to optimizing the HFR video transmission over wireless networks. First, we develop an analytical model dubbed FRIED (Frame Rate versus vIdEo Distortion) to characterize the tradeoff between sending frame rate and end-to-end video distortion. Second, we propose a Joint frAme Selection and FEC (Forward Error Correction) cOding (JASCO) approach based on the FRIED model to optimize the transmission performance. The efficacy of the proposed JASCO is evaluated through extensive semi-physical emulations in Exata involving H.264 video streaming. Experimental results show that JASCO outperforms the reference approaches in improving video peak signal-to-noise ratio (PSNR) at the same frame rate. Or conversely, JASCO is able to achieve higher received frame rate while guaranteeing the same video PSNR.
Jiyan Wu, Chau Yuen, Ngai-Man Cheung, Junliang Chen 0001, Chang Wen Chen
IEEE Trans. Wirel. Commun.3
2015 FAemb: A function approximation-based embedding method for image retrieval
abstract
The objective of this paper is to design an embedding method mapping local features describing image (e.g. SIFT) to a higher dimensional representation used for image retrieval problem. By investigating the relationship between the linear approximation of a nonlinear function in high dimensional space and state-of-the-art feature representation used in image retrieval, i.e., VLAD, we first introduce a new approach for the approximation. The embedded vectors resulted by the function approximation process are then aggregated to form a single representation used in the image retrieval framework. The evaluation shows that our embedding method gives a performance boost over the state of the art in image retrieval, as demonstrated by our experiments on the standard public image retrieval benchmarks.
Thanh-Toan Do, Quang D. Tran, Ngai-Man Cheung
CVPR3
2015 Depth Error Induced Virtual View Synthesis Distortion Estimation for 3D Video Coding
abstract
We propose an analytical model to estimate the depth-error-induced virtual view synthesis distortion (VVSD) in 3D video, taking into account the configuration of the cameras. Focusing on view synthesis under depth error, we carefully analyze the merging operations under different situations that affect pixel availability: overlapping region, disocclusion and boundary region, disparity error region, and infrequent region. The analysis leads to quadratic/biquadratic models and linear models that explicitly relate the distance between camera positions (reference/virtual view) to Distortion under Different Situations (DDS) and Probability under Different Situations (PDS), respectively. We also show that VVSD is the linear combination of DDS weighted by PDS. Our careful analysis results in state-of-the-art estimation accuracy. Experimental results verify that the proposed model is capable to produce accurate estimates of VVSD based on the distance between reference/virtual views. Therefore, our model can effectively inform camera setup for capturing, in particular, the set-up of the cameras in situation where depth information will be compressed subsequently.
Yijian Xiang, Lu Fang 0001, Ngai-Man Cheung
DCC4
2015 On the Efficiency of View Synthesis Prediction for 3D Video Coding
abstract
We study the efficiency of view synthesis prediction (VSP). The proposed spectral domain analysis relates the power spectral density of the VSP error to the probability density function of the warping error. The analysis takes into account the warping error induced by (i) depth coding and (ii) disparity rounding at integer-pel, half-pel and quarter-pel warping accuracy. The interaction between VSP efficiency and interpolation filter is also studied. We validate our proposed model with empirical data. Using the proposed model, we discuss the interaction between prediction efficiency, depth image distortion, warping accuracy and interpolation filter. The proposed model provides theoretical insights of VSP in 3D-HEVC. It could be used to guide optimization of VSP designs.
Ngai-Man Cheung, Lu Yu 0003
DCC2
2015 The efficiency of view synthesis prediction for 3D video coding: A spectral domain analysis
abstract
We study the coding efficiency of view synthesis prediction (VSP) in 3D video coding. Our spectral domain analysis relates the power spectral density (PSD) of the VSP prediction error to the probability density function (pdf) of the warping error. Our analysis takes into account the warping error induced by (i) depth coding and (ii) rounding error at integer-pel, half-pel and quarter-pel warping accuracy. We also study the interaction between depth coding error and warping accuracy. Our model suggests that the coding gain with using higher warping accuracy diminishes as the depth coding error increases. Our analysis results are validated with empirical data.
Ngai-Man Cheung, Lu Yu 0003
ICASSP2
2015 Open Designettes, Flowcharts, and Pseudocodes in Python Programming with the Aid of Finch
Cheah Huei Yoong, Hyowon Lee 0001, Ngai-Man Cheung
ICCE3
2015 Information-theoretic analysis of Blinn-Phong lighting with applicationto mobile cloud gaming
abstract
In mobile cloud gaming, images are rendered, compressed, and delivered to mobile devices from cloud servers. Joint optimization of rendering and coding aims to reduce data rate of the delivery. Among the tasks in a rendering pipeline, Blinn-Phong lighting is the most widely adopted lighting model and also has the most impact on the visual information of the rendered images. In this paper, we analyze and model the information content of an image rendered using the Blinn-Phong lighting model. Our analytic model estimates the entropy generated by the Blinn-Phong reflections on the rendered images. In the context of cloud gaming, we derive the solution for the joint rendering and coding optimization, and demonstrate how the entropy estimators can facilitate fast and on-the-fly solution to the optimization. Experiments results validate our information-theoretic analyses, and show substantial bitrate reduction by the joint optimization of rendering and coding.
Seong-Ping Chuah, Ngai-Man Cheung, Chau Yuen
ICIP2
2015 Enabling Adaptive High-Frame-Rate Video Streaming in Mobile Cloud Gaming Applications
abstract
High-frame-rate (HFR) video is emerging in popular gaming applications to enhance the smooth experience perceived by end users. However, it is challenging to guarantee the delivery quality of HFR video in mobile cloud gaming scenarios because of the high transmission rate and limited wireless resources. To address this critical problem, we develop a novel transmission scheduling framework dubbed AdaPtive HFR vIdeo Streaming (APHIS). The term adaptive indicates this scheme's capability in dynamically adjusting the video traffic load and forward error correction (FEC) coding. First, we propose an online video frame selection algorithm to minimize the total distortion based on the network status, input video data, and delay constraint. Second, we introduce an unequal FEC coding scheme to provide differentiated protection for Intra (I) and Predicted (P) frames with low-latency cost. The proposed APHIS framework is able to appropriately filter video frames and adjust data protection levels to optimize the quality of HFR video streaming. We conduct extensive emulations in Exata involving HFR video encoded with H.264 codec. Experimental results show that APHIS outperforms the reference transmission schemes in terms of video peak signal-to-noise ratio, end-to-end delay, and goodput. Therefore, we recommend APHIS for delivering HFR video streaming in mobile cloud gaming systems.
Jiyan Wu, Chau Yuen, Ngai-Man Cheung, Junliang Chen 0001, Chang Wen Chen
IEEE Trans. Circuits Syst. Video Technol.3
2014 Bandwidth efficient mobile cloud gaming with layered coding and scalable phong lighting
abstract
In mobile cloud gaming, one of the main challenges is to deliver high-quality game images over wireless networks under stringent delay requirement. To reduce the bit-rate of game images, we propose Layered Coding, which leverages the graphic rendering capability of modern mobile devices to reduce transmission bit-rate. Specifically, we render a low-quality local game image, or the base layer, on the power-constrained mobile client. Instead of sending the high quality game image, the cloud server sends enhancement layer information, which the client utilizes to improve the quality of the base layer. Central to the proposed layered coding is the design of base layer (BL) rendering. We discuss BL design and propose a computationally-scalable Phong lighting that can be used in BL rendering. We performed experiments to compare our layered coding with state-of-the-art, which uses H.264/AVC inter-frame coding to compress game images. With game sequences of different model complexity and motion, our results suggest that layered coding requires substantially lower data-rate. We made available game video test sequences to stimulate future research.
Seong-Ping Chuah, Ngai-Man Cheung
ICIP2
2014 Analytical model for camera distance related 3D virtual view distortion estimation
abstract
We propose an analytical model to estimate the depth-error-induced synthesis distortion in 3D video, taking into account the configuration of the cameras. In particular, the model mathematically relates the Distance between camera positions (reference view and virtual view) to the Virtual View Distortion (VVD), thus it is denoted as DVVD model. Specifically, the DVVD model accounts for two modules: distribution of disparity errors and shift-induced distortion. The former one is derived under a Laplacian distribution assumption of depth errors, and the latter one is estimated under a Quadratic model. We further propose a linear Steady-State model by performing Taylor series approximation of the DVVD model over a region of practical interest. Experiment results demonstrate that both the DVVD and Steady-State models are capable of estimating the relationship between VVD and the distance between virtual/reference view. Therefore, our model can effectively inform camera setup for capturing, in particular, the setup of the cameras in situation where depth information will be compressed subsequently.
Yijian Xiang, Ngai-Man Cheung, Juyong Zhang, Lu Fang 0001
ICIP2
2014 DeepCAPTCHA: an image CAPTCHA based on depth perception
abstract
Over the past decade, text-based CAPTCHA (TBC) have become popular in preventing adversarial attacks and spam in many websites and applications including emails services, social platforms, web-based market places, and recommendation systems. However, in addition to several problems with TBC, it has become increasingly difficult to solve in recent years, to keep up with OCR technologies. Image-based CAPTCHA (IBC), on the other hand, is a relatively new concept that promises to overcome key limitations of TBC. In this paper we present an innovative IBC, DeepCAPTCHA, based on design guidelines, psychological theory and empirical experiments. DeepCAPTCHA exploits the human ability of depth preception. In our IBC users should arrange 3D objects in terms of size (or depth). In our framework for DeepCAPTCHA, we automatically mine 3D models, and use a human-machine Merge Sort algorithm to order these unknown objects. We then create new appearances for these objects at multiplication factor of 200, and present these new images to the end-users for sorting (as CAPTCHA tasks). Humans are able to apply their rapid and reliable object recognition and comparison (arise from years experience with the physical environment) to solve DeepCAPTCHA, while machines are still unable to complete these tasks. Experimental results show that humans can solve DeepCAPTCHA with a high accuracy (~84%) and ease, while machines perform dismally.
Hossein Nejati, Ngai-Man Cheung, Ricardo Sosa, Dawn Chin-Ing Koh
MMSys2
2014 An Analytical Model for Synthesis Distortion Estimation in 3D Video
abstract
We propose an analytical model to estimate the synthesized view quality in 3D video. The model relates errors in the depth images to the synthesis quality, taking into account texture image characteristics, texture image quality, and the rendering process. Especially, we decompose the synthesis distortion into texture-error induced distortion and depth-error induced distortion. We analyze the depth-error induced distortion using an approach combining frequency and spatial domain techniques. Experiment results with video sequences and coding/rendering tools used in MPEG 3DV activities show that our analytical model can accurately estimate the synthesis noise power. Thus, the model can be used to estimate the rendering quality for different system designs.
Lu Fang 0001, Ngai-Man Cheung, Dong Tian, Anthony Vetro, Huifang Sun, Oscar C. Au
IEEE Trans. Image Process.2
2013 Low Bit-Rate Subpixel-Based Color Image Compression
abstract
We propose a novel low bit-rate compression scheme with sub pixel-based down-sampling and reconstruction (SPDR) for full color images. In the encoder stage, a decoder-dependent multi-channel sub pixel-based down-sampling is proposed, which is more effective in retaining high frequency detail than conventional pixel-based process. The decoder first decompresses the low-resolution image and then up-converts it to the original resolution using encoder dependent sub pixel-based reconstruction scheme by jointly considering the sub pixel-based down-sampling effect and the compression degradation. Compared to existing algorithms with comparable encoder and decoder complexity, the proposed SPDR offers complete standard compliance, competitive rate-distortion performance, and superior subjective quality.
Lu Fang 0001, Ngai-Man Cheung, Oscar C. Au, Houqiang Li, Ketan Tang
DCC2
2013 From Boxes to bees: Active learning in freshmen calculus
abstract
The vehicles of education have seen significant broadening with the proliferation of new technologies such as social media, microblogs, online references, multimedia, and interactive teaching tools. This paper summarizes research on the effect of using active learning methods to facilitate student learning and describes our experiences implementing group activities for a calculus course for first year university students at Singapore University of Technology and Design, a new design-centric university established in collaboration with Massachusetts Institute of Technology (MIT). We describe the educational impact of different pedagogical techniques, such as real-time response tools, hands on activities, mathematical modeling, visualization activities and motivational competitions on students with differing learning preferences in a unique cohort classroom setting. Based on faculty reflection and survey data, we provide guidelines on how to adopt the right set of active and group learning techniques to handle the changing learning preferences in the current and future generation of students.
Flora S. Tsai, Karthik Natarajan, Selin Damla Ahipasaoglu, Chau Yuen, Hyowon Lee 0001, Ngai-Man Cheung, Justin Ruths, Shisheng Huang, Thomas L. Magnanti
EDUCON6
2013 Compressed sensing of diffusion fields under heat equation constraint
abstract
Reconstructing a diffusion field from spatiotemporal measurements is an important problem in engineering and physics with applications in temperature flow, pollution dispersion, and disease epidemic dynamics. In such applications, sensor networks are used as spatiotemporal sampling devices and a relatively large number of spatiotemporal measurements may be required for accurate source field reconstruction. Consequently, due to limitations on the number of nodes in the sensor networks as well as hardware limitations of each sensor, situations may arise where the available spatiotemporal sampling density does not allow for recovery of field details. In this paper, the above limitation is resolved by means of using compressed sensing (CS). We propose to exploit the intrinsic property of diffusive fields as side information to improve the reconstruction results of classic CS which we call diffusive compressed sensing (DCS). Experimental results demonstrate the effectiveness and usefulness of the proposed method in substantial data savings while producing estimates of higher accuracy, as compared to classic CS-base estimates.
Ngai-Man Cheung, Tony Q. S. Quek
ICASSP2
2013 Rate-distortion optimized merge frame using piecewise constant functions
abstract
The ability to efficiently switch from one pre-encoded video stream to another is a valuable attribute for a variety of interactive streaming applications, such as switching among streams of the same video encoded in different bit-rates for real-time bandwidth adaptation, or view-switching among videos capturing the same dynamic 3D scene but from different viewpoints. It is well known that intra-coded I-frames can be used at switch boundaries to facilitate stream-switching. However, the size of an I-frame is large, making frequent insertion impractical. A recent proposal towards a more efficient stream-switching mechanism is distributed source coding (D-SC), which exploits worst-case correlation between a set of potential predictor frames in the decoder buffer (called side information (SI) frames) and a target frame to lower encoding rate. However, the conventional use of bit-plane and channel coding means the encoding and decoding complexity of DSC frames is large. In this paper, we pursue a novel approach to the stream-switching problem based on the concept of “signal merging”, using piecewise constant (p-wc) function as the merge operator. Specifically, we propose a new merge mode for a code block, where for each k-th transform coefficient in the block, we encode appropriate step size and horizontal shift parameters at the encoder, so that the resulting floor function at the decoder can map corresponding coefficients from any SI frame to the same reconstructed value, resulting in an identically merged signal. The selection of shift parameter per coefficient, as well as coding modes between intra and merge per block, are optimized in a rate-distortion (RD) optimal manner. Experiments show encouraging coding gain over a previous implementation of DSC frame at low-to mid-bitrates at reduced computation complexity.
Wei Dai 0002, Gene Cheung, Ngai-Man Cheung, Antonio Ortega, Oscar C. Au
ICIP3
2013 Synthesis distortion estimation in 3D video using frequency and spatial analysis
abstract
We propose an analytical model to estimate the synthesized view quality in 3D video. Specifically, we estimate the depth-error induced distortion using an approach that combines frequency and spatial domain analysis. We also propose to decompose the spatial-variant video signals into gradient-based representations to capture the interaction between image gradients, depth errors and synthesis distortion. Experiment results with video sequences and coding/rendering tools used in MPEG 3DV activities show that our analytical model can accurately estimate the synthesis noise power.
Lu Fang 0001, Ngai-Man Cheung, Dong Tian, Anthony Vetro, Huifang Sun, Lu Yu 0003
ICIP2
2013 Novel distortion metric for depth coding of 3D video
abstract
In state-of-the-art HEVC-based 3D video codec, multiview video plus associated depth maps are used. In order to achieve better coding performance, instead of the conventional sum of squared errors (SSE), view synthesis optimization (VSO) is proposed and included in the anchor encoder software to calculate view synthesis distortion in rate-distortion optimization (RDO) of depth coding. The anchor VSO achieves high rate-distortion (RD) performance. However, it requires partial rendering and is quite complex and time-consuming. On the other hand, simple SSE metric is fast but RD performance is low. In this paper, we propose a new distortion metric to be used in RDO for depth coding. The complexity of the proposed method is slightly higher than SSE, while its RD performance remains competitive. With a good trade-off between complexity and performance, the proposed method can replace the conventional SSE metric in RDO for depth coding, and can be used as a low-complexity alternative to the anchor VSO.
Ngai-Man Cheung, Oscar C. Au, Dong Tian
ICIP2
2013 Luma-Chroma Space Filter Design for Subpixel-Based Monochrome Image Downsampling
abstract
In general, subpixel-based downsampling can achieve higher apparent resolution of the down-sampled images on LCD or OLED displays than pixel-based downsampling. With the frequency domain analysis of subpixel-based downsampling, we discover special characteristics of the luma-chroma color transform choice for monochrome images. With these, we model the anti-aliasing filter design for subpixel-based monochrome image downsampling as a human visual system-based optimization problem with a two-term cost function and obtain a closed-form solution. One cost term measures the luminance distortion and the other term measures the chrominance aliasing in our chosen luma-chroma space. Simulation results suggest that the proposed method can achieve sharper down-sampled gray/font images compared with conventional pixel and subpixel-based methods, without noticeable color fringing artifacts.
Lu Fang 0001, Oscar C. Au, Ngai-Man Cheung, Aggelos K. Katsaggelos, Houqiang Li, Feng Zou 0006
IEEE Trans. Image Process.3
2012 On modeling the rendering error in 3D video
abstract
We propose an analytical model to estimate the rendering quality in 3D video. The model relates errors in the depth images to the rendering quality, taking into account texture image characteristics, texture image quality, the camera configuration and the rendering process. Specifically, we derive position (disparity) errors from the depth errors, and the probability distribution of the position errors is used to calculate the power spectral density of the rendering errors. Experiment results with video sequences and coding/rendering tools used in MPEG 3DV activities show that the model can accurately estimate the synthesis noise up to a constant offset. Thus, the model can be used to estimate the change in rendering quality for different system designs.
Ngai-Man Cheung, Dong Tian, Anthony Vetro, Huifang Sun
ICIP1
2012 Analytical study of RGB vertical stripe and RGBX square-shaped subpixel arrangements
abstract
The frequency characteristics of subpixel-based decimation with RGB vertical stripe and RGBX square-shaped subpixel arrangements are studied. To achieve higher apparent resolution than pixel-based decimation, the sampling locations are specially chosen for each of two subpixel arrangements, resulting in relatively small magnitudes of horizontal and vertical aliasing spectra in frequency domain. Thanks to 2-D RGBX square-shaped subpixel arrangement, all the horizontal, vertical, diagonal and anti-diagonal aliasing spectra merely contain low-frequency information, indicating that subpixel-based decimation with RGBX square-shaped panel is more effective in retaining original high frequency details than RGB vertical stripe subpixel arrangement.
Lu Fang 0001, Oscar C. Au, Jingjing Dai, Hanli Wang, Ngai-Man Cheung
ICIP5
2011 Quality-controlled view interpolation for multiview video
abstract
Multiview video coding systems are characterized by their high encoding complexity. In this paper, we propose to reduce the encoding complexity by omitting frames at the encoder in a pattern that keeps the ability to reconstruct an interpolated version of each omitted frame, by motion-compensated and inter-view interpolation. Since our goal is maintaining good visual quality, we propose a method to estimate the quality of the interpolated frames at the decoder by transmitting a small amount of error control information in lieu of an omitted frame. This information is obtained from a projection of the frame on a suitable low-dimensional basis. The projection coefficients can be compressed by conventional techniques or, more efficiently, by Slepian-Wolf coding. Based on the quality estimate, the decoder can adaptively detect the interpolation technique that results in better visual quality. Moreover, it can suppress occasional frames for which interpolation yields noticeable artifacts. Experimental results demonstrate that our approach eliminates most interpolation artifacts and achieves much better visual quality at a very small increase in bit-rate.
Mina Makar, Yao-Chung Lin, Ngai-Man Cheung, Derek Pang, Bernd Girod
ICIP3
2011 Synchronization of presentation slides and lecture videos using bit rate sequences
abstract
The temporal synchronization of presentation slides and lecture videos enables us to enhance the user experience of online lecture viewing systems. In this work, we present a novel approach to robustly and reliably detect and recognize slide changes, which is based on bit rate sequences. By exploiting readily available information, the computational complexity of the approach is very low. In contrast to prior work, no requirements on the amount or type of texture, motion of foreground objects, or text size on the slide are imposed. Experimental results show the ability to detect even minor slide changes and demonstrate the robustness against occlusions, foreground motion, and camera motion.
Georg Schroth, Ngai-Man Cheung, Eckehard G. Steinbach, Bernd Girod
ICIP2
2011 ClassX: an open source interactive lecture StreamingSystem
abstract
The ClassX open source project is a free experimental interactive video streaming platform designed for educators, researchers and software developers. With minimal infra-structure set-up, ClassX offers educational communities a cost-effective solution for online lecture delivery. Our goal is to encourage contributions from other researchers, developers and educators in building an open, cost-effective and state-of-the-art online education video viewing system for the general public.
Sherif A. Halawa, Derek Pang, Ngai-Man Cheung, Bernd Girod
ACM Multimedia3
2011 ClassX Mobile: region-of-interest video streaming to mobile devices with multi-touch interaction
abstract
Small screen sizes, limited bandwidth, and low computational power often prohibit streaming of high-resolution videos to mobile devices over a wireless network. Recent advances in interactive region-of-interest (IRoI) video streaming technology enable users to interactively control pan/tilt/ zoom, while provide bit-rate and complexity savings. One recent application of IRoI video streaming is ClassX developed at Stanford University. It offers an open-source experimental platform for interactive online lecture streaming. In this technical demonstration, we present ClassX Mobile, which extends the current ClassX system and delivers high-quality interactive video to smartphones and tablets with multi-touch screens.
Derek Pang, Sherif A. Halawa, Ngai-Man Cheung, Bernd Girod
ACM Multimedia3
2011 The stanford mobile visual search data set
abstract
We survey popular data sets used in computer vision literature and point out their limitations for mobile visual search applications. To overcome many of the limitations, we propose the Stanford Mobile Visual Search data set. The data set contains camera-phone images of products, CDs, books, outdoor landmarks, business cards, text documents, museum paintings and video clips. The data set has several key characteristics lacking in existing data sets: rigid objects, widely varying lighting conditions, perspective distortion, foreground and background clutter, realistic ground-truth reference data, and query data collected from heterogeneous low and high-end camera phones. We hope that the data set will help push research forward in the field of mobile visual search.
Vijay Chandrasekhar 0001, David M. Chen, Sam S. Tsai, Ngai-Man Cheung, Huizhong Chen, Gabriel Takacs, Yuriy A. Reznik, Ramakrishna Vedantham, Radek Grzeszczuk, Jeff Bach, Bernd Girod
MMSys4
2011 Interactive Streaming of Stored Multiview Video Using Redundant Frame Structures
abstract
While much of multiview video coding focuses on the rate-distortion performance of compressing all frames of all views for storage or non-interactive video delivery over networks, we address the problem of designing a frame structure to enable interactive multiview streaming, where clients can interactively switch views during video playback. Thus, as a client is playing back successive frames (in time) for a given view, it can send a request to the server to switch to a different view while continuing uninterrupted temporal playback. Noting that standard tools for random access (i.e., I-frame insertion) can be bandwidth-inefficient for this application, we propose a redundant representation of I-, P-, and "merge" frames, where each original picture can be encoded into multiple versions, appropriately trading off expected transmission rate with storage, to facilitate view switching. We first present ad hoc frame structures with good performance when the view-switching probabilities are either very large or very small. We then present optimization algorithms that generate more general frame structures with better overall performance for the general case. We show in our experiments that we can generate redundant frame structures offering a range of tradeoff points between transmission and storage, e.g., outperforming simple I-frame insertion structures by up to 45% in terms of bandwidth efficiency at twice the storage cost.
Gene Cheung, Antonio Ortega, Ngai-Man Cheung
IEEE Trans. Image Process.3
2010 Comparison of local feature descriptors for mobile visual search
abstract
We evaluate the performance of MPEG-7 image signatures, Compressed Histogram of Gradients descriptor (CHoG) and Scale Invariant Feature Transform (SIFT) descriptors for mobile visual search applications. We observe that SIFT and CHoG outperform MPEG-7 image signatures greatly in terms of feature-level Receiver Operating Characteristic (ROC) performance and image-level matching. Moreover, CHoG descriptors demonstrate such gains while being comparable with MPEG-7 image signatures in bit-rate.
Vijay Chandrasekhar 0001, David M. Chen, Andy Lin, Gabriel Takacs, Sam S. Tsai, Ngai-Man Cheung, Yuriy A. Reznik, Radek Grzeszczuk, Bernd Girod
ICIP6
2010 Dynamic selection of a feature-rich query frame for mobile video retrieval
abstract
In this paper, we focus on a new application of mobile visual search: snapping a photo with a mobile device of a video playing on a TV screen to automatically retrieve and stream the remainder of the video to the mobile device. When the user takes a photo of the video, the captured query frame may contain too few useful features for good retrieval performance. We design and implement a new algorithm for mobile video retrieval to accurately select a feature-rich frame from a sequence of viewfinder frames in a very short temporal window determined by the user-initiated query event. Fast and accurate selection using efficiently computed Hessian scores is developed for real-time operation on mobile devices. Viewfinder frames captured before the query starts are pre-processed, while the number of viewfinder frames captured afterwards is minimized by a probabilistic optimization process. Evaluated on a large video database of 10 million frames, dynamic query frame selection provides a substantial increase in retrieval accuracy with very low search latency.
David M. Chen, Ngai-Man Cheung, Sam S. Tsai, Vijay Chandrasekhar 0001, Gabriel Takacs, Ramakrishna Vedantham, Radek Grzeszczuk, Bernd Girod
ICIP2
2010 Rate-distortion based reconstruction optimization in distributed source coding for interactive multiview video streaming
abstract
Interactive multiview video streaming (IMVS) is an application where, as the streaming multiview video is played back in time, an observer iteratively requests one of many available views at the server. In response, the server sends the appropriate pre-encoded data to the observer, with data chosen for transmission depending on the specific transmitted data available in the observer's cache. The primary challenge in IMVS is to design a structure for the pre-encoded multiview data, so that during an IMVS streaming session, the transmission rate is appropriately traded off with the pre-encoded data storage size. Previously, we have developed novel distributed source coding (DSC) based frame configurations to optimize the said tradeoff, outperforming periodical insertions of I-frames in both transmission and storage costs. In this paper, we show that by exploiting the freedom to choose a target decoded frame at a view switching point for DSC, further performance gains can be achieved: by up to 0.7 dB in our experiments.
Ngai-Man Cheung, Antonio Ortega, Gene Cheung
ICIP1
2010 Restoration of out-of-focus lecture video by automatic slide matching
abstract
Restoring the fine detail in the slide area of a defocused lecture video is a challenging task. In this work, we propose to use clean images of slides available along with the defocused lecture video to help the restoration. Our proposed method uses local feature descriptors and multiple defocused slide decks to automatically identify the slide that is displayed in the defocused frame. We then use the matching slide as side information to estimate the parameters for deconvolution and bilateral filtering. Experimental results show that the proposed algorithm compares favorably to a computationally-intensive iterative deconvolution algorithm that does not employ any side information. In particular, it can recover small drawings and text that are severely blurred in a poorly focused lecture video
Ngai-Man Cheung, David M. Chen, Vijay Chandrasekhar 0001, Sam S. Tsai, Gabriel Takacs, Sherif A. Halawa, Bernd Girod
ACM Multimedia1
2010 Mobile product recognition
abstract
We present a mobile product recognition system for the camera-phone. By snapping a picture of a product with a camera-phone, the user can retrieve online information of the product. The product is recognized by an image-based retrieval system located on a remote server. Our database currently comprises more than one million entries, primarily products packaged in rigid boxes with printed labels, such as CDs, DVDs, and books. We extract low bit-rate descriptors from the query image and compress the location of the descriptors using location histogram coding on the camera-phone. We transmit the compressed query features, instead of a query image, to reduce the transmission delay. We use inverted index compression and fast geometric re-ranking on our database to provide a low delay image recognition response for large scale databases. Experimental timing results on different parts of the mobile product recognition system is reported in this work.
Sam S. Tsai, David M. Chen, Vijay Chandrasekhar 0001, Gabriel Takacs, Ngai-Man Cheung, Ramakrishna Vedantham, Radek Grzeszczuk, Bernd Girod
ACM Multimedia5
2010 On media data structures for interactive streaming in immersive applications
abstract
Interactive media streaming is the communication paradigm where an observer periodically requests new desired subsets from the streaming sender in real-time, upon which the sender sends the appropriate media data, corresponding to the received requests, for immediate decoding and display. This is in contrast to non-interactive media streaming, e.g., TV broadcast, where the entire media set is compressed and delivered to the observer before the observer interacts with the data (such as switching TV channels). Examples of interactive streaming abound in different media modalities: interactive browsing of JPEG2000 images, interactive light field or multiview video streaming, etc. Interactive media streaming has the obvious advantage of bandwidth efficiency: only the media subsets corresponding to observer's requests are transmitted. This is important when an observer only views a small subset out of a very large media data set during a typical streaming session. The technical challenge is how to structure media data such that good compression efficiency can be achieved by exploiting correlation among media subsets (thus inducing a particular decoding order if correlation is exploited during encoding), while providing sufficient flexibility for the observer to freely navigate the media data set in his/her desired unique order. In this overview paper, we survey different proposals in the literature that simultaneously achieve the conflicting objectives of compression efficiency and decoding flexibility.
Gene Cheung, Antonio Ortega, Ngai-Man Cheung, Bernd Girod
VCIP3
2010 Successive refinement based Wyner-Ziv video compression
Xiaopeng Fan 0001, Oscar C. Au, Ngai-Man Cheung, Yan Chen 0007, Jiantao Zhou 0001
Signal Process. Image Commun.3
2010 Transform-Domain Adaptive Correlation Estimation (TRACE) for Wyner-Ziv Video Coding
abstract
Wyner-Ziv video coding (WZVC) is a newly emerged video coding scheme which compresses the input video frames with the side information (SI) frames only available at the decoder. WZVC exploits the statistics between the source frame and the SI frame at the decoder by utilizing their correlation information. This correlation information is important but also difficult to estimate due to the absence of the SI frame at the encoder, and the lack of the source frame at the decoder. In this paper, we focus on this problem and propose a novel transform-domain adaptive correlation estimation method called TRACE for WZVC. In TRACE, the correlation information is progressively learned during the decoding process of each frame. Within TRACE, we also propose a convex optimization based band-level correlation estimation method which is optimal in the sense of minimizing the theoretical bit rate. Experiments suggest that, when applied in motion compensated interpolation-based low complexity WZVC, TRACE yields competitive results against the state-of-the-art correlation estimation algorithms. More importantly, different from the existing coefficient-level correlation estimation algorithms, the proposed TRACE can be applied in many other WZVC schemes and can provide considerable gain over the popular band-level correlation estimation methods.
Xiaopeng Fan 0001, Oscar C. Au, Ngai-Man Cheung
IEEE Trans. Circuits Syst. Video Technol.3
2009 Transcoding based robust streaming of compressed video
abstract
A variety of techniques have been proposed to enhance the error robustness of the video streaming system. However, most of them improves the error resilience during compression rather than after compression. In this paper, we propose a novel transcoding based scheme called lossless inter frame transcoding (LIFT) scheme to improve the error resilience of existing compressed video stream. In the LIFT scheme, inter coded blocks are selectively transcoded into new kind of blocks called ‘L-block’. At the decoder, the L-block can be transcoded back to the original P-block when the prediction is available and can also be robustly decoded as I-block when the prediction is unavailable. By offline transcoding and online adjusting the ratio of P-blocks and L-blocks, the proposed streaming server achieves error robustness scalability. Experimental results demonstrate the correctness and effectiveness of the proposed method.
Xiaopeng Fan 0001, Oscar C. Au, Mengyao Ma, Ling Hou, Jiantao Zhou 0001, Ngai-Man Cheung
ICASSP6
2009 Parallel rate-distortion optimized intra mode decision on multi-core graphics processors using greedy-based encoding orders
abstract
Rate-distortion (RD) optimized intra-prediction mode selection can lead to significant improvement in coding efficiency in intra-frame encoding. However, it would incur considerable increase in encoding complexity. In this paper, we investigate how multi-core Graphics Processing Units (GPUs) can be efficiently utilized to undertake the task of RD optimized intra mode selection in AVS and H.264 video encoding. Achieving efficient GPU-based intra mode decision, however, could be non-trivial. It is because the mode decision of the current block would depend on the reconstructed data of the neighboring blocks. Therefore, the coding modes of neighboring blocks would need to be computed first before that of the current block can be determined. This dependency poses challenge to computation on multi-core GPUs, which rely heavily on parallel data processing to achieve superior speedups. To address this issue, we analyze the data dependency in intra mode decision, and propose novel greedy-based encoding orders to achieve highly parallel processing. We also prove that the proposed greedy-based orders are optimal in terms of execution time. Experimental results suggest that the proposed GPU-based intra mode decision compares favorably to the counterpart implemented on a single-core CPU.
Ngai-Man Cheung, Oscar C. Au, Man Cheung Kung, Xiaopeng Fan 0001
ICIP1
2009 Optimized frame structure using distributed source coding for interactive multiview video streaming
abstract
While multiview video coding typically focuses on the rate-distortion performance of compressing all frames of all views, we address the problem of designing a pre-encoded frame structure for a streaming server to enable a new functionality-interactive multiview switching, where a streaming client can send requests periodically to a server to switch to different views while continuing uninterrupted temporal playback of streaming video. We observe that providing bandwidth-efficient interactive view switching usually comes at the price of additional overall storage. Thus, our goal is to find a frame structure that minimizes the expected transmission rate during interactive multiview streaming, subject to a storage constraint. Noting that standard tools for random access (i.e., I-frame insertion) can be bandwidth-inefficient for this functionality, we propose to automatically generate a structure, combining I-frames, redundant P-frames and Distributed Source Coded (DSC) frames, in a near-optimal fashion to facilitate view switching. We present three new DSC techniques for view switching and discuss how these techniques can be integrated into an optimization framework. We show experimentally that near-optimal coding structures using DSC frames, in addition to I- and P-frames, reduce transmission cost over structures using I-frames only for view switching by up to 28%, and over structures using I- and P-frames only by up to 20% for the same storage cost.
Gene Cheung, Ngai-Man Cheung, Antonio Ortega
ICIP2
2009 Adaptive correlation estimation for general Wyner-Ziv video coding
abstract
Wyner-Ziv video coding (WZVC) is a new paradigm for video compression with the prediction frames possibly only available at the decoder. It exploits the statistics between the source frame and the prediction frame at the decoder by utilizing their correlation information. This correlation information is important but also difficult to estimate due to the absolute absence of the prediction frame at the encoder, and the lack of the source frame at the decoder. In this paper, we focus on this issue and derive a coefficient-level adaptive correlation model for general Wyner-Ziv video coding. Based on this model, we propose an online transform-domain adaptive correlation estimation (TRACE) approach, in which the correlation information is progressively learned during the decoding process. In our experiments, the proposed approach outperforms the existing approaches up to 4 dB. More importantly, different from existing coefficient-level variance estimation approaches, the proposed on-line TRACE is applicable for not only low complexity WZVC but also other WZVCs such as flexible WZVC as demonstrated in the experiments.
Xiaopeng Fan 0001, Oscar C. Au, Ngai-Man Cheung
ICIP3
2009 On Improving the Robustness of Compressed Video by Slepian-Wolf based Lossless Transcoding
abstract
A variety of techniques have been proposed to enhance the error robustness of the video streaming system. However, most of them improve the error resilience during compression rather than after compression. In this paper, we propose a novel transcoding based scheme called Slepian-Wolf based inter frame transcoding (SWIFT) to improve the error resilience of existing compressed video stream. In the SWIFT scheme, inter coded blocks are selectively transcoded into new kind of blocks called ‘X-block’. At the decoder, the X-block can be transcoded back to the original P-block when there is no error in the prediction, and can also be robustly decoded as I-block when there are errors in the prediction. In the experiments, the proposed SWIFT scheme does not introduce transcoding distortion as expected, and always improves the robustness of the compressed video at all packet loss rate. Compared with the H.264 based transcoder, SWIFT achieves better RD performance and error resilience performance.
Xiaopeng Fan 0001, Oscar C. Au, Mengyao Ma, Ling Hou, Jiantao Zhou 0001, Ngai-Man Cheung
ISCAS6
2009 Distributed source coding techniques for interactive multiview video streaming
abstract
We investigate coding tools for interactive multiview streaming (IMVS), where clients interactively request desired views for successive video frames, and in response the server sends the appropriate pre-compressed video data to the clients. Solution based on using only I-frames to support view switching would incur high transmission cost, while for that based on using only P-frames to encode every possible traversal, although it can minimize transmission cost, prohibitive server's storage may be required. Therefore, efficient solutions for IMVS need to consider the trade-off between transmission and storage cost. In this paper, we study the potential use of distributed source coding (DSC) in IMVS. Specifically, we propose two DSC constructions that could achieve good transmission-storage trade-offs. Central to these constructions is a method that can efficiently encode the least significant bits (LSB) of a frame to be decoded, leading to competitive storage and transmission requirements. Experiment results demonstrate these constructions compare favorably to existing tools, and could be valuable for interactive multiview streaming.
Ngai-Man Cheung, Antonio Ortega, Gene Cheung
PCS1
2009 Highly Parallel Rate-Distortion Optimized Intra-Mode Decision on Multicore Graphics Processors
abstract
Rate-distortion (RD)-based mode selections are important techniques in video coding. In these methods, an encoder may compute the RD costs for all the possible coding modes, and select the one which achieves the best trade-off between encoding rate and compression distortion. Previous papers have demonstrated that RD-based mode selections can lead to significant improvements in coding efficiency. RD-based mode selections, however, would incur considerable increases in encoding complexity, since these methods require computing the RD costs for numerous candidate coding modes. In this paper, we consider the scenario where software-based video encoding is performed on personal computers or game consoles, and investigate how multicore graphics processing units (GPUs) may be efficiently utilized to undertake the task of RD optimized intra-prediction mode selections in audio and video coding standards and H.264 video encoding. Achieving efficient GPU-based intra-mode decisions, however, could be nontrivial for two reasons. First, intra-mode decision tends to be sequential. Specifically, the mode decision of the current block would depend on thereconstructed dataof the neighboring blocks. Therefore, the coding modes of neighboring blocks would need to be computed first before that of the current block can be determined. This dependency poses challenges to GPU-based computation, which relies heavily on parallel data processing to achieve superior speedups. Second, RD-based intra-mode decision may require conditional branchings to determine the encoding bit-rate, and these branching operations may incur substantial performance penalties when being executed on GPUs due to pipeline architectural designs. To address these issues, we analyze the data dependency in intra-mode decision, and propose novel greedy-based encoding orders to achieve highly parallel processing of data blocks. We also prove that the proposed greedy-based orders are optimal in our problem, i.e., they require the minimum number of iterations to process a video frame given the dependency constraints. In addition, we propose a method to estimate the coding rate suitable for GPU implementation. Experimental results suggest our proposed solution can be more than 50 times faster than the previously proposed parallel intra-prediction, since our work can efficiently exploit the massive parallel opportunity in GPUs.
Ngai-Man Cheung, Oscar C. Au, Man Cheung Kung, Peter Hon-Wah Wong, Chun-Hung Liu
IEEE Trans. Circuits Syst. Video Technol.1
2008 Sampling-Based Correlation Estimation for Distributed Source Coding Under Rate and Complexity Constraints
abstract
In many practical distributed source coding (DSC) applications, correlation information has to be estimated at the encoder in order to determine the encoding rate. Coding efficiency depends strongly on the accuracy of this correlation estimation. While error in estimation is inevitable, the impact of estimation error on compression efficiency has not been sufficiently studied for the DSC problem. In this paper,we study correlation estimation subject to rate and complexity constraints, and its impact on coding efficiency in a DSC framework for practical distributed image and video applications. We focus on, in particular, applications where binary correlation models are exploited for Slepian-Wolf coding and sampling techniques are used to estimate the correlation, while extensions to other correlation models would also be briefly discussed. In the first part of this paper, we investigate the compression of binary data. We first propose a model to characterize the relationship between the number of samples used in estimation and the coding rate penalty, in the case of encoding of a single binary source. The model is then extended to scenarios where multiple binary sources are compressed, and based on the model we propose an algorithm to determine the number of samples allocated to different sources so that the overall rate penalty can be minimized, subject to a constraint on the total number of samples. The second part of this paper studies compression of continuous valued data. We propose a model-based estimation for the particular but important situations where binary bit-planes are extracted from a continuous-valued input source, and each bit-plane is compressed using DSC. The proposed model-based method first estimates the source and correlation noise models using continuous valued samples, and then uses the models to derive the bit-plane statistics analytically. We also extend the model-based estimation to the cases when bit-planes are extracted based on the significance of the data, similar to those commonly used in wavelet-based applications. Experimental results, including some based on hyperspectral image compression, demonstrate the effectiveness of the proposed algorithms.
Ngai-Man Cheung, Huisheng Wang, Antonio Ortega
IEEE Trans. Image Process.1
2007 Flexible Video Decoding: A Distributed Source Coding Approach
abstract
We investigate video compression techniques to address problems that requireflexible video decoding. In these, the encoder has access to a number of candidate predictors that allow it to exploit source signal correlation, but only a subset of these predictors will be available at the decoder. Crucially, the encoderdoes notknow which predictors will be available. Flexible decoding is important in a number of applications including frame-by-frame forward and backward video playback, multiview video, bitstreams switching, robust video transmission, etc. The main challenge to support flexible decoding is that the encoder needs to compress a current frame under the uncertainty on the predictor at decoder. An approach based on conventional "closed loop" prediction, e.g., motion-compensated predictive (MCP) coding in the case of video, could be developed by including multiple possible prediction residues in the bitstream, but this would lead to a considerable coding performance penalty, if all possible predictor combinations are supported, or to drifting, if only some combinations are. Moreover, it is not possible in general to guarantee that decoded versions under different prediction scenarios will be identical. In this paper, we propose a distributed source coding (DSC) based algorithm to tackle the problem. The main novelties of the proposed algorithm are that it incorporates different macroblock modes and significance coding within the DSC framework. This, combined with a judicious exploitation of correlation statistics, allows us to achieve competitive coding performance. Using forward/backward video playback as an example, we demonstrate the proposed algorithm can outperform a solution based on MCP coding.
Ngai-Man Cheung, Antonio Ortega
MMSP1
2006 A Model-Based Approach to Correlation Estimation Inwavelet-Based Distributed Source Coding with Application to Hyperspectral Imagery
abstract
In many practical distributed source coding (DSC) applications correlation information has to be obtained at the encoder in order to determine the encoding rate. Coding efficiency depends strongly on the accuracy of this correlation estimation, which often has to be performed under rate and complexity constraints. In this paper we focus on correlation estimation for wavelet-based DSC. We extend our previously proposed model-based estimation techniques, which provided accurate estimates of bit-plane level correlation under rate constraints, in the simple case where bit-planes are generated from the binary representation of the sources. To extend the model-based approach to wavelet-based DSC, we need to address two issues. Firstly, in order to improve coding efficiency, bit-planes are typically generated by more sophisticated algorithms in wavelet-based DSC (e.g., by deciding on the bitplane scan order based on coefficient "significance"), which makes model-based estimation more challenging. Secondly, certain wavelet subbands may not have enough coefficients for reliable model estimation, so that model-based techniques alone may not be sufficiently accurate. We propose solutions to these problems and, using a DSC-based hyperspectral image system as example, we demonstrate that model-based estimation can lead to efficient system implementation with lower computational and data exchange requirements, and improved parallelism, while incurring only small degradation in coding efficiency.
Ngai-Man Cheung, Antonio Ortega
ICIP1
2006 Efficient wavelet-based predictive Slepian-Wolf coding for hyperspectral imagery
Ngai-Man Cheung, Caimu Tang, Antonio Ortega, Cauligi S. Raghavendra
Signal Process.1
2005 Efficient Inter-Band Prediction and Wavelet Based Compression for Hyperspectral Imagery: A Distributed Source Coding Approach
abstract
Hyperspectral images have correlation at the level of pixels; moreover, images from neighboring frequency bands are also closely correlated. In this paper, we propose to use distributed source coding to exploit this correlation with an eye to a more efficient hardware implementation. Slepian-Wolf and Wyner-Ziv based correlated coding theorems have quantified how much additional rate reduction can be obtained. In order to better exploit these correlations, we first propose a prediction model to align images. This model is based on linear prediction techniques and it is simple and shown to be effective for hyperspectral images. We then propose a coding scheme to exploit these correlations. A set-partitioning approach is used on wavelet transformed data to extract bitplanes. Under our correlation model, bitplanes from neighboring bands are correlated and we then use a low-density parity-check based Slepian-Wolf code to exploit this bitplane level correlation. This scheme is appealing for hardware implementation as it is easy to parallelize and it has modest memory requirements. As for coding performance, our preliminary results for high correlation spectral bands from the NASA AVIRIS dataset show, at medium to high reconstructed qualities, gains of about a factor of 3 in compression efficiency as compared to encoding the spectral bands independently using SPIHT.
Caimu Tang, Ngai-Man Cheung, Antonio Ortega, Cauligi S. Raghavendra
DCC2
2005 Correlation estimation for distributed source coding under information exchange constraints
abstract
Distributed source coding (DSC) depends strongly on accurate knowledge of correlation between sources. Previous works have reported capacity-approaching code constructions when exact knowledge of correlation is available at the encoder. However, in many applications exact correlation information may not be available, and correlation estimation is necessary. While error in estimation is inevitable, the impact of estimation error on compression efficiency has not been sufficiently studied for the DSC problem. In this paper we study correlation estimation subject to complexity constraints, and its impact on coding efficiency in a DSC framework. In particular, we consider the case where estimation entails information exchange between spatially separate sources and thus correlation estimation is subject to rate constraints. We first derive optimal strategies for information exchange that minimize the rate penalty due to inaccurate estimation, under constraints on the number of bits that can be exchanged between sources. Experimental results show that significant gain is possible by optimally exchanging information. We then derive analytical expressions to quantify the rate penalty, and analyze how rate penalty changes with a priori knowledge of correlation. In addition, we present a model-based estimation method which can achieve more accurate estimation results compared to directly inspecting the data.
Ngai-Man Cheung, Huisheng Wang, Antonio Ortega
ICIP (2)1
2002 Configurable variable length code with genetic algorithms
abstract
The variable length code (VLC) tables in MPEG-1/2/4 and H.26X are fixed and optimized for a limited range of bit-rates, and they cannot handle a variety of applications. The universal variable length code (UVLC) is a new scheme to encode syntax elements and has some configurable capabilities. It is being considered in the ITU-T H.26L. However, the configurable feature of UVLC has not been well explored. In this paper we describe a VLC scheme which uses configurations as parameters to adapt UVLC to different symbol distributions of different applications. We also propose a method to automatically determine the configuration parameters based on a genetic algorithm (GA). Experimental results show that our method can achieve very good coding efficiency while drastically simplifying the encoding and decoding process, and is applicable to a variety of applications.
Ngai-Man Cheung, Yuji Itoh
ICIP (3)1
2001 Configurable variable length code for video coding
abstract
The variable length code (VLC) tables in MPEG-1/2/4 and H.263 are fixed and optimized for a limited range of bit-rates, and they cannot handle a variety of applications. The universal variable length code (UVLC) is a new scheme to encode syntax elements and has some configurable capabilities. It is also being considered in ITU-T H.26L. However, the configurable feature of the UVLC has not been well explored. We propose configuring the UVLC with the additional code configuration (ACC). The ACC is used to adapt UVLC to different symbol distributions by adjusting the partitioning of the symbols into different categories, and the code size assignment to different categories. Experimental results show that the UVLC with ACC outperforms the current proposed scheme in H.26L and the VLC tables of existing standards, while drastically simplifying the encoding and decoding process, and is applicable to a variety of applications.
Ngai-Man Cheung, Yuji Itoh
ICASSP1
2000 Universal Variable Length Code for DCT Coding
abstract
A new technique for entropy coding named universal VLC (i.e., UVLC) was developed and applied to the motion vector. It was shown that UVLC for motion vector coding provides good performance in terms of coding efficiency as well as error resiliency. In this paper UVLC for DCT coefficient coding is presented. The features of the scheme are: (1) adaptable by parametric representation, (2) support of very low bitrate mode and (3) extensible to error resilient mode with bi-directionally decodable code. This favors a flexible image and video coding that should handle a variety of picture types and applications. The simulation results show that the proposed scheme slightly outperforms several VLC tables of existing standards in terms of coding efficiency. It is also expected that the proposed scheme can simplify the implementation and processing of VLC decoder. It shall also be noted that UVLC is being considered as a candidate component technique for ITU-T H.26L, i.e., next generation video coding algorithm standard.
Yuji Itoh, Ngai-Man Cheung
ICIP2
1999 Software MPEG-2 video decoder on a DSP Enhanced Memory Module
abstract
This paper describes our effort in implementing software MPEG-2 video decoder using TI 'C6201 DSP and Basava Technology. The hardware device, DSP Enhanced Memory Module (DSP-MM), leverages the advantages of high computation performance from 'C6201 DSP, as well as high bandwidth memory access and efficient memory usage from Basava Technology. A prototype with 'C6201 DSP and 32MB SDRAM shared memory is functioning in PC environment running Windows 95 operating system. We will describe the video decoder in detail.
Ngai-Man Cheung, Tsugio Kawashima, Yoshihide Iwata, Steven Trautmann, Jeff Tsay, Raj Pawate
MMSP1
1998 Head-related transfer function modeling in 3-D sound systems with genetic algorithms
abstract
Head-related transfer functions (HRTFs) describe the spectral filtering that occurs between a source sound and the listener's eardrum. Since HRTFs vary as a function of the relative source location and subject, practical implementation of 3D audio must take into account a large set of HRTFs for different azimuths and elevations. Previous work has proposed several HRTF models for data reduction. This paper describes our work in applying genetic algorithms to find a set of HRTF basis spectra, and the normal equation method to compute the optimal combination of linear weights to represent the individual HRTFs at different azimuths and elevations. The genetic algorithm selects the basis spectra from the set of original HRTF amplitude responses, using an average relative spectral error as the fitness function. Encouraging results from the experiments suggest that genetic algorithms provide an effective approach to this data reduction problem.
Ngai-Man Cheung, Steven Trautmann, Andrew Horner
ICASSP1
1997 Genetic algorithm approach to head-related transfer functions modeling in 3-D sound system
abstract
Head-related transfer functions, or HRTFs, refer to the spectral filtering from sound sources to listeners' eardrums. It is an important cue to spatial hearing. Since HRTFs vary as a function of relative source locations and subjects, practical implementation of 3D audio always races a large set of HRTFs for different azimuths and elevations (even if non-individualized HRTFs are used). Previous works have proposed several models to represent the HRTFs in order to achieve data reduction. In this paper, we describe our work In applying Genetic Algorithm (GA) to HRTF modeling. Based on a linear combination model, our system uses a GA to find the basis spectra and the normal equation method to compute the weights matrix. Individual HRTFs at different azimuths and elevations are represented as a weighted combination of the basis spectra. A set of HRTFs from the MIT Media Lab's KEMAR measurements was used as the source data.
Ngai-Man Cheung, Steven Trautmann
MMSP1
1997 Wavetable music synthesis for multimedia and beyond
abstract
Music production is a key element of any complete multimedia system and therefore an important component of multimedia chip set solution. Wavetable synthesis is a set of techniques which allow real time music production with high polyphony on current DSPs while exceeding the quality of simple FM synthesis. However it is important to allow other synthesis techniques to further enhance the musical product. Some of these techniques, such as spectral modeling synthesis and physical modeling of musical instruments can produce higher quality than wavetable but are more expensive computationally. Thus even when real-time performance is possible, there exists a trade-off between polyphony and quality. In order to provide superior quality as well as flexibility and scaleability, we have created an overall structure that can utilize various synthesis techniques in different situations with minimum overhead. Synthesis techniques such as FM, wavetable, spectral and physical modeling and so on will thus be used together to create a more compelling musical output. This is achieved with a control structure that dynamically chooses and controls the different synthesis techniques. The first set of synthesis methods incorporated in this system is a combination of wavetable and sampling techniques.
Steven Trautmann, Ngai-Man Cheung
MMSP2