EDBT 2026 Demo / reviewers in the wild / expert
Fengqing Zhu 0001
dblp:47/3260 · also Fengqing Maggie Zhu
· DBLP profile ↗
62ranked-venue papers
4as first author
41since 2021 · last 2026
0000-0002-3863-3220ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 57 · 3 first-author · 38 since 2021Artificial intelligence and machine learning · 10 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PANDA - Patch and Distribution-Aware Augmentation for Long-Tailed Exemplar-Free Continual LearningabstractExemplar-Free Continual Learning (EFCL) restricts the storage of previous task data and is highly susceptible to catastrophic forgetting. While pre-trained models (PTMs) are increasingly leveraged for EFCL, existing methods often overlook the inherent imbalance of real-world data distributions. We discovered that real-world data streams commonly exhibit dual-level imbalances, dataset-level distributions combined with extreme or reversed skews within individual tasks, creating both intra-task and inter-task disparities that hinder effective learning and generalization. To address these challenges, we propose PANDA, a Patch-and-Distribution-Aware Augmentation framework that integrates seamlessly with existing PTM-based EFCL methods. PANDA amplifies low-frequency classes by using a CLIP encoder to identify representative regions and transplanting those into frequent-class samples within each task. Furthermore, PANDA incorporates an adaptive balancing strategy that leverages prior task distributions to smooth inter-task imbalances, reducing the overall gap between average samples across tasks and enabling fairer learning with frozen PTMs. Extensive experiments and ablation studies demonstrate PANDA's capability to work with existing PTM-based CL methods, improving accuracy and reducing catastrophic forgetting. Siddeshwar Raghavan, Jiangpeng He, Fengqing Zhu 0001 |
AAAI | 3 |
| 2026 | Food Image Generation on Multi-Noun CategoriesabstractGenerating realistic food images for categories with multiple nouns is surprisingly challenging. For instance, the prompt "egg noodle" may result in images that incorrectly contain both eggs and noodles as separate entities. Multi-noun food categories are common in real-world datasets and account for a large portion of entries in benchmarks such as UEC-256. These compound names often cause generative models to misinterpret the semantics, producing unintended ingredients or objects. This is due to insufficient multi-noun category related knowledge in the text encoder and misinterpretation of multi-noun relationships, leading to incorrect spatial layouts. To overcome these challenges, we propose FoCULR (Food Category Understanding and Layout Refinement) which incorporates food domain knowledge and introduces core concepts early in the generation process. Experimental results demonstrate that the integration of these techniques improves image generation performance in the food domain. Xinyue Pan, Yuhao Chen 0001, Jiangpeng He, Fengqing Zhu 0001 |
WACV | 4 |
| 2026 | QARV++: An Improved Hierarchical VAE for Learned Image Compressionabstractvalues, stabilizing variable-rate training. Extensive experiments demonstrate that QARV++ achieves superior rate-distortion (R-D) performance among HVAE-based LIC models, exhibiting -12.20% -16.34% -15.23% BD-Rate against VVC Intra mode on the Kodak, Tecnick, and CLIC2020 test datasets, respectively. Our approach also generalizes effectively to existing LICs, delivering substantial improvements. Yuning Huang, Fengqing Zhu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | Long-Tailed Continual Learning for Visual Food RecognitionabstractDeep learning-based food recognition has made significant progress in predicting food types from eating occasion images. However, two key challenges hinder real-world deployment: (1) continuously learning new food classes without forgetting previously learned ones, and (2) handling the long-tailed distribution of food images, where a few common classes and many more rare classes. To address these, food recognition methods should focus on long-tailed continual learning. In this work, We introduce a dataset that encompasses 186 American foods along with comprehensive annotations. We also introduce three new benchmark datasets, VFN186-LT, VFN186-INSULIN and VFN186-T2D, which reflect real-world food consumption for healthy populations, insulin takers and individuals with type 2 diabetes without taking insulin. We propose a novel end-to-end framework that improves the generalization ability for instance-rare food classes using a knowledge distillation-based predictor to avoid misalignment of representation during continual learning. Additionally, we introduce an augmentation technique by integrating class-activation-map (CAM) and CutMix to improve generalization on instance-rare food classes. Our method, evaluated on Food101-LT, VFN-LT, VFN186-LT, VFN186-INSULIN, and VFN186-T2DM, shows significant improvements over existing methods. An ablation study highlights further performance enhancements, demonstrating its potential for real-world food recognition applications. Jiangpeng He, Luotao Lin, Jack Ma, Heather A. Eicher-Miller, Fengqing Zhu 0001 |
IEEE Trans. Multim. | 6 |
| 2025 | CL-LoRA: Continual Low-Rank Adaptation for Rehearsal-Free Class-Incremental LearningabstractClass-Incremental Learning (CIL) aims to learn new classes sequentially while retaining the knowledge of previously learned classes. Recently, pre-trained models (PTMs) combined with parameter-efficient fine-tuning (PEFT) have shown remarkable performance in rehearsal-free CIL without requiring exemplars from previous tasks. However, existing adapter-based methods, which incorporate lightweight learnable modules into PTMs for CIL, create new adapters for each new task, leading to both parameter redundancy and failure to leverage shared knowledge across tasks. In this work, we propose ContinuaL Low-Rank Adaptation (CL-LoRA), which introduces a novel dual-adapter architecture combining task-shared adapters to learn cross-task knowledge and task-specific adapters to capture unique features of each new task. Specifically, the shared adapters utilize random orthogonal matrices and leverage knowledge distillation with gradient reassignment to preserve essential shared knowledge. In addition, we introduce learnable block-wise weights for task-specific adapters, which mitigate inter-task interference while maintaining the model’s plasticity. We demonstrate CL-LoRA consistently achieves promising performance under multiple benchmarks with reduced training and inference computation, establishing a more efficient and scalable paradigm for continual learning with pre-trained models. Jiangpeng He, Zhihao Duan, Fengqing Zhu 0001 |
CVPR | 3 |
| 2025 | Balanced Rate-Distortion Optimization in Learned Image CompressionabstractLearned image compression (LIC) using deep learning architectures has seen significant advancements, yet standard rate-distortion (R-D) optimization often encounters imbalanced updates due to diverse gradients of the rate and distortion objectives. This imbalance can lead to suboptimal optimization, where one objective dominates, thereby reducing overall compression efficiency. To address this challenge, we reformulate R-D optimization as a multi-objective optimization (MOO) problem and introduce two balanced R-D optimization strategies that adaptively adjust gradient updates to achieve more equitable improvements in both rate and distortion. The first proposed strategy utilizes a coarse-to-fine gradient descent approach along standard RD optimization trajectories, making it particularly suitable for training LIC models from scratch. The second proposed strategy analytically addresses the reformulated optimization as a quadratic programming problem with an equality constraint, which is ideal for fine-tuning existing models. Experimental results demonstrate that both proposed methods enhance the R-D performance of LIC models, achieving around a 2% BD-Rate reduction with acceptable additional training cost, leading to a more balanced and efficient optimization process. Code will be available at https://gitlab.com/viper-purdue/Balanced-RD. Zhihao Duan, Yuning Huang, Fengqing Zhu 0001 |
CVPR | 4 |
| 2025 | Confidence-Aware Agglomeration Classification And Segmentation Of 2D Microscopic Food Crystal Images*abstractFood crystal agglomeration is a phenomenon occurs during crystallization which traps water between crystals and affects food product quality. Manual annotation of agglomeration in 2D microscopic images is particularly difficult due to the transparency of water bonding and the limited perspective focusing on a single slide of the imaged sample. To address this challenge, we first propose a supervised baseline model to generate segmentation pseudo-labels for the coarsely labeled classification dataset. Next, an instance classification model that simultaneously performs pixel-wise segmentation is trained. Both models are used in the inference stage to combine their respective strengths in classification and segmentation. To preserve crystal properties, a post processing module is designed and included to both steps. Our method improves true positive agglomeration classification accuracy and size distribution predictions compared to other existing methods. Given the variability in confidence levels of manual annotations, our proposed method is evaluated under two confidence levels and successfully classifies potential agglomerated instances. Xiaoyu Ji 0004, Ali Shakouri, Fengqing Zhu 0001 |
ICIP | 3 |
| 2025 | Low-Rank Adaptation of Pre-Trained Vision Backbones for Energy-Efficient Image Coding For MachinesabstractImage Coding for Machines (ICM) focuses on optimizing image compression for AI-driven analysis rather than human perception. Existing ICM frameworks often rely on separate codecs for specific tasks, leading to significant storage requirements, training overhead, and computational complexity. To address these challenges, we propose an energy-efficient framework that leverages pre-trained vision backbones to extract robust and versatile latent representations suitable for multiple tasks. We introduce a task-specific low-rank adaptation mechanism, which refines the pre-trained features to be both compressible and tailored to downstream applications. This design minimizes trainable parameters and reduces energy costs for multi-task scenarios. By jointly optimizing task performance and entropy minimization, our method enables efficient adaptation to diverse tasks and datasets without full fine-tuning, achieving high coding efficiency. Extensive experiments demonstrate that our framework significantly outperforms traditional codecs and pre-processors, offering an energy-efficient and effective solution for ICM applications. The code and the supplementary materials will be available at: https://gitlab.com/viper-purdue/efficient-compression. Zhihao Duan, Yuning Huang, Fengqing Zhu 0001 |
ICIP | 4 |
| 2025 | ICP-3DGS: SFM-Free 3D Gaussian Splatting for Large-Scale Unbounded ScenesabstractIn recent years, neural rendering methods such as NeRFs and 3D Gaussian Splatting (3DGS) have made significant progress in scene reconstruction and novel view synthesis. However, they heavily rely on preprocessed camera poses and 3D structural priors from structure-from-motion (SfM), which are challenging to obtain in out door scenarios. To address this challenge, we propose to incorporate Iterative Closest Point (ICP) with optimization-based refinement to achieve accurate camera pose estimation under large camera movements. Additionally, we introduce a voxel-based scene densification approach to guide the reconstruction in large-scale scenes. Experiments demonstrate that our approach ICP-3DGS outperforms existing methods in both camera pose estimation and novel view synthesis across indoor and outdoor scenes of various scales. Source code is available at https://github.com/Chenhao-Z/ICP-3DGS. Yezhi Shen, Fengqing Zhu 0001 |
ICIP | 3 |
| 2024 | Deep Hierarchical Video CompressionabstractRecently, probabilistic predictive coding that directly models the conditional distribution of latent features across successive frames for temporal redundancy removal has yielded promising results. Existing methods using a single-scale Variational AutoEncoder (VAE) must devise complex networks for conditional probability estimation in latent space, neglecting multiscale characteristics of video frames. Instead, this work proposes hierarchical probabilistic predictive coding, for which hierarchal VAEs are carefully designed to characterize multiscale latent features as a family of flexible priors and posteriors to predict the probabilities of future frames. Under such a hierarchical structure, lightweight networks are sufficient for prediction. The proposed method outperforms representative learned video compression models on common testing videos and demonstrates computational friendliness with much less memory footprint and faster encoding/decoding. Extensive experiments on adaptation to temporal patterns also indicate the better generalization of our hierarchical predictive mechanism. Furthermore, our solution is the first to enable progressive decoding that is favored in networked video applications with packet loss. Ming Lu 0003, Zhihao Duan, Fengqing Zhu 0001, Zhan Ma 0001 |
AAAI | 3 |
| 2024 | Another Way to the Top: Exploit Contextual Clustering in Learned Image CodingabstractWhile convolution and self-attention are extensively used in learned image compression (LIC) for transform coding, this paper proposes an alternative called Contextual Clustering based LIC (CLIC) which primarily relies on clustering operations and local attention for correlation characterization and compact representation of an image. As seen, CLIC expands the receptive field into the entire image for intra-cluster feature aggregation. Afterward, features are reordered to their original spatial positions to pass through the local attention units for inter-cluster embedding. Additionally, we introduce the Guided Post-Quantization Filtering (GuidedPQF) into CLIC, effectively mitigating the propagation and accumulation of quantization errors at the initial decoding stage. Extensive experiments demonstrate the superior performance of CLIC over state-of-the-art works: when optimized using MSE, it outperforms VVC by about 10% BD-Rate in three widely-used benchmark datasets; when optimized using MS-SSIM, it saves more than 50% BD-Rate over VVC. Our CLIC offers a new way to generate compact representations for image compression, which also provides a novel direction along the line of LIC development. Zhihao Duan, Ming Lu 0003, Dandan Ding, Fengqing Zhu 0001, Zhan Ma 0001 |
AAAI | 5 |
| 2024 | Towards Backward-Compatible Continual Learning of Image CompressionabstractThis paper explores the possibility of extending the capa-bility of pre-trained neural image compressors (e.g., adapting to new data or target bitrates) without breaking back-ward compatibility, the ability to decode bitstreams encoded by the original model. We refer to this problem as continual learning of image compression. Our initial findings show that baseline solutions, such as end-to-end fine-tuning, do not preserve the desired backward compatibility. To tackle this, we propose a knowledge replay training strategy that effectively addresses this issue. We also design a new model architecture that enables more effective continual learning than existing baselines. Experiments are conducted for two scenarios: data-incremental learning and rate-incremental learning. The main conclusion of this paper is that neural image compressors can be fine-tuned to achieve better per-formance (compared to their pre-trained version) on new data and rates without compromising backward compati-bility. The code is publicly available online. Zhihao Duan, Ming Lu 0003, Justin Yang, Jiangpeng He, Zhan Ma 0001, Fengqing Zhu 0001 |
CVPR | 6 |
| 2024 | Structured Pruning and Quantization for Learned Image CompressionabstractThe high computational costs associated with large deep learning models significantly hinder their practical deployment. Model pruning has been widely explored in deep learning literature to reduce their computational burden, but its application has been largely limited to computer vision tasks such as image classification and object detection. In this work, we propose a structured pruning method targeted for Learned Image Compression (LIC) models that aims to reduce the computational costs associated with image compression while maintaining the rate-distortion performance. We employ a Neural Architecture Search (NAS) method based on the rate-distortion loss for computing the pruning ratio for each layer of the network. We compare our pruned model with the uncompressed LIC Model with same network architecture and show that it can achieve model size reduction without any BD-Rate performance drop. We further show that our pruning method can be integrated with model quantization to achieve further model compression while maintaining similar BD-Rate performance. We have made the source code available at gitlab.com/viper-purdue/lic-pruning. Md Adnan Faisal Hossain, Fengqing Zhu 0001 |
ICIP | 2 |
| 2024 | On Efficient Neural Network Architectures for Image CompressionabstractRecent advances in learning-based image compression typically come at the cost of high complexity. Designing computationally efficient architectures remains an open challenge. In this paper, we empirically investigate the impact of different network designs in terms of rate-distortion performance and computational complexity. Our experiments involve testing various transforms, including convolutional neural networks and transformers, as well as various context models, including hierarchical, channel-wise, and space-channel context models. Based on the results, we present a series of efficient models, the final model of which has comparable performance to recent best-performing methods but with significantly lower complexity. Extensive experiments provide insights into the design of architectures for learned image compression and potential direction for future research. The code is available at https://gitlab.com/viper-purdue/efficient-compression. Zhihao Duan, Fengqing Zhu 0001 |
ICIP | 3 |
| 2024 | Flexible Mixed Precision Quantization for Learne Image CompressionabstractDespite its improvements in coding performance compared to traditional codecs, Learned Image Compression (LIC) suffers from large computational costs for storage and deployment. Model quantization offers an effective solution to reduce the computational complexity of LIC models. However, most existing works perform fixed-precision quantization which suffers from sub-optimal utilization of resources due to the varying sensitivity to quantization of different layers of a neural network. In this paper, we propose a Flexible Mixed Precision Quantization (FMPQ) method that assigns different bit-widths to different layers of the quantized network using the fractional change in rate-distortion loss as the bit-assignment criterion. We also introduce an adaptive search algorithm which reduces the time-complexity of searching for the desired distribution of quantization bit-widths given a fixed model size. Evaluation of our method shows improved BD-Rate performance under similar model size constraints compared to other works on quantization of LIC models. We have made the source code available at gitlab.com/viper-purdue/fmpq. Md Adnan Faisal Hossain, Zhihao Duan, Fengqing Zhu 0001 |
ICME | 3 |
| 2024 | Theoretical Bound-Guided Hierarchical Vae For Neural Image CodecsabstractRecent studies reveal a significant theoretical link between variational autoencoders (VAEs) and rate-distortion theory, notably in utilizing VAEs to estimate the theoretical upper bound of the information rate-distortion function of images. Such estimated theoretical bounds substantially exceed the performance of existing neural image codecs (NICs). To narrow this gap, we propose a theoretical bound-guided hierarchical VAE (BG-VAE) for NIC. The proposed BG-VAE leverages the theoretical bound to guide the NIC model towards enhanced performance. We implement the BG-VAE using Hierarchical VAEs and demonstrate its effectiveness through extensive experiments. Along with advanced neural network blocks, we provide a versatile, variable-rate NIC that outperforms existing methods when considering both ratedistortion performance and computational complexity. The code is available at $\color{magenta}{\text{BG - VAE}}$. Zhihao Duan, Yuning Huang, Fengqing Zhu 0001 |
ICME | 4 |
| 2024 | Efficient Microscopic Image Instance Segmentation for Food Crystal Quality ControlabstractThis paper is directed towards the food crystal quality control area for manufacturing, focusing on efficiently predicting food crystal counts and size distributions. Previously, manufacturers used the manual counting method on microscopic images of food liquid products, which requires substantial human effort and suffers from inconsistency issues. Food crystal segmentation is a challenging problem due to the diverse shapes of crystals and their surrounding hard mimics. To address this challenge, we propose an efficient instance segmentation method based on object detection. Experimental results show that the predicted crystal counting accuracy of our method is comparable with existing segmentation methods, while being five times faster. Based on our experiments, we also define objective criteria for separating hard mimics and food crystals, which could benefit manual annotation tasks on similar dataset. Xiaoyu Ji 0004, Jan P. Allebach, Ali Shakouri, Fengqing Zhu 0001 |
MMSP | 4 |
| 2024 | Compression of Self-Supervised Representations for Machine VisionabstractCoding for machines is an emerging research topic aiming at compressing visual signals for machine analysis. Most existing research on it consider supervised feature compression, i.e., compressing features for a particular task and dataset. This paper explores a more flexible scenario in which the machine vision task (e.g., the classes to predict) is unknown during training and encoding (i.e., only decided during decoding). To achieve this goal, we compress self-supervised learning (SSL) features, which are not tied to a particular dataset but can be used for various tasks without re-training. Empirical studies are provided to analyze and derive an SSL feature compression system. Despite its simplicity, we show that linear transform coding achieves comparable or better rate-accuracy performance for SSL features compared to more advanced techniques. Although SSL feature compression performs slightly worse than its supervised counterpart, it generalizes well for out-of-distribution datasets. We highlight the use cases of various feature compression schemes and provide insights into developing future work. The code is made publicly available11https://gitlab.com/viper-purdue/dino-compress. Zhihao Duan, Fengqing Zhu 0001 |
MMSP | 2 |
| 2024 | FMiFood: Multi-Modal Contrastive Learning for Food Image ClassificationabstractFood image classification is the fundamental step in image-based dietary assessment, which aims to estimate participants' nutrient intake from eating occasion images. A common challenge of food images is the intra-class diversity and inter-class similarity, which can significantly hinder classification performance. To address this issue, we introduce a novel multimodal contrastive learning framework called FMiFood, which learns more discriminative features by integrating additional contextual information, such as food category text descriptions, to enhance classification accuracy. Specifically, we propose a flexible matching technique that improves the similarity matching between text and image embeddings to focus on multiple key information. Furthermore, we incorporate the classification objectives into the framework and explore the use of GPT-4 to enrich the text descriptions and provide more detailed context. Our method demonstrates improved performance on both the UPMC-101 and VFN datasets compared to existing methods. Xinyue Pan, Jiangpeng He, Fengqing Zhu 0001 |
MMSP | 3 |
| 2024 | Pose Guided Portrait View Interpolation from Dual Cameras with a Long BaselineabstractWe introduce a novel method for interpolating views between two fixed cameras to create a free-viewpoint experience comparable to advanced telepresence systems. Inspired by Video Frame Interpolation (VFI) works, our proposed method spatially interpolates between camera views, treating our setup as a unique frame interpolation scenario. Due to the occlusion that exists in areas like the face contour from two cameras with a long baseline, which results in serious artifacts with previous VFI methods. To mitigate this issue, we integrate face pose information of pitch, yaw and roll angles in our method as a robust guidance prior to improve pixel mapping in occluded regions. Furthermore, we synthesize a large portrait-focused multi-view dataset facilitating training of our proposed model. Our multi-stage flow refinement model progressively refines the quality of the bi-directional flows which eventually results in improved interpolated view details. With the flexible setup and fast inference speed of our proposed method, a practical and cost-effective solution can be implemented for telepresence systems without the need for expensive hardware or extensive computational resources. Weichen Xu 0002, Yezhi Shen, Qian Lin 0001, Jan P. Allebach, Fengqing Zhu 0001 |
MMSP | 5 |
| 2024 | Message from the MMSP 2024 General and Technical Program ChairsabstractThe 26th IEEE International Workshop on Multimedia Signal Processing (MMSP 2024), organized by the Multimedia Signal Processing Technical Committee (MMSP-TC) of IEEE Signal Processing Society (SPS), was held at Purdue University, West Lafayette, Indiana, U.S.A., from October 2 - 4, 2024. Fengqing Zhu 0001, Nikolaos Thomos, Balu Adsumilli, Enrico Magli |
MMSP | 1 |
| 2024 | Probing Image Compression for Class-Incremental LearningabstractImage compression emerges as a pivotal tool in the efficient handling and transmission of digital images. Its ability to substantially reduce file size not only facilitates enhanced data storage capacity but also potentially brings advantages to the development of continual machine learning (ML) systems, which learn new knowledge incrementally from sequential data. Continual ML systems often rely on storing representative samples, also known as exemplars, within a limited memory constraint to maintain the performance on previously learned data. These methods are known as memory replay-based algorithms and have proven effective at mitigating the detrimental effects of catastrophic forgetting. Nonetheless, the limited memory buffer size often falls short of adequately representing the entire data distribution. In this paper, we explore the use of image compression as a strategy to enhance the buffer's capacity, thereby increasing exemplar diversity. However, directly using compressed exemplars introduces domain shift during continual ML, marked by a discrepancy between compressed training data and uncompressed testing data. Additionally, it is essential to determine the appropriate compression algorithm and select the most effective rate for continual ML systems to balance the trade-off between exemplar quality and quantity. To this end, we introduce a new framework to incorporate image compression for continual ML including a pre-processing data compression step and an efficient compression rate/algorithm selection method. We conduct extensive experiments on CIFAR-100 and ImageNet datasets and show that our method significantly improves image classification accuracy in continual ML settings. Justin Yang, Zhihao Duan, Andrew Peng, Yuning Huang, Jiangpeng He, Fengqing Zhu 0001 |
PCS | 6 |
| 2024 | Online Class-Incremental Learning For Real-World Food Image ClassificationabstractFood image classification is essential for monitoring health and tracking dietary in image-based dietary assessment methods. However, conventional systems often rely on static datasets with fixed classes and uniform distribution. In contrast, real-world food consumption patterns, shaped by cultural, economic, and personal influences, involve dynamic and evolving data. Thus, require the classification system to cope with continuously evolving data. Online Class Incremental Learning (OCIL) addresses the challenge of learning continuously from a single-pass data stream while adapting to the new knowledge and reducing catastrophic forgetting. Experience Replay (ER) based OCIL methods store a small portion of previous data and have shown encouraging performance. However, most existing OCIL works assume that the distribution of encountered data is perfectly balanced, which rarely happens in real-world scenarios. In this work, we explore OCIL for real-world food image classification by first introducing a probabilistic framework to simulate realistic food consumption scenarios. Subsequently, we present an attachable Dynamic Model Update (DMU) module designed for existing ER methods, which enables the selection of relevant images for model training, addressing challenges arising from data repetition and imbalanced sample occurrences inherent in realistic food consumption patterns within the OCIL framework. Our performance evaluation demonstrates significant enhancements compared to established ER methods, showing great potential for lifelong learning in real-world food image classification scenarios. The code of our method is publicly accessible at https://gitlab.com/viper-purdue/OCIL-real-world-food-image-classification Siddeshwar Raghavan, Jiangpeng He, Fengqing Zhu 0001 |
WACV | 3 |
| 2024 | QARV: Quantization-Aware ResNet VAE for Lossy Image CompressionabstractThis paper addresses the problem of lossy image compression, a fundamental problem in image processing and information theory that is involved in many real-world applications. We start by reviewing the framework of variational autoencoders (VAEs), a powerful class of generative probabilistic models that has a deep connection to lossy compression. Based on VAEs, we develop a new scheme for lossy image compression, which we name quantization-aware ResNet VAE (QARV). Our method incorporates a hierarchical VAE architecture integrated with test-time quantization and quantization-aware training, without which efficient entropy coding would not be possible. In addition, we design the neural network architecture of QARV specifically for fast decoding and propose an adaptive normalization operation for variable-rate compression. Extensive experiments are conducted, and results show that QARV achieves variable-rate compression, high-speed decoding, and better rate-distortion performance than existing baseline methods. Zhihao Duan, Ming Lu 0003, Jack Ma, Yuning Huang, Zhan Ma 0001, Fengqing Zhu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | An Improved Upper Bound on the Rate-Distortion Function of ImagesabstractRecent work has shown that Variational Autoencoders (VAEs) can be used to upper-bound the information rate-distortion (R-D) function of images, i.e., the fundamental limit of lossy image compression. In this paper, we report an improved upper bound on the R-D function of images implemented by (1) introducing a new VAE model architecture, (2) applying variable-rate compression techniques, and (3) proposing a novel smoothing function to stabilize training. We demonstrate that at least 30% BD-rate reduction w.r.t. the intra prediction mode in VVC codec is achievable, suggesting that there is still great potential for improving lossy image compression. Code is made publicly available1. Zhihao Duan, Jack Ma, Jiangpeng He, Fengqing Zhu 0001 |
ICIP | 4 |
| 2023 | Single-Stage Heavy-Tailed Food ClassificationabstractDeep learning based food image classification has enabled more accurate nutrition content analysis for image-based dietary assessment by predicting the types of food in eating occasion images. However, there are two major obstacles to apply food classification in real life applications. First, real life food images are usually heavy-tailed distributed, resulting in severe class-imbalance issue. Second, it is challenging to train a single-stage (i.e. end-to-end) framework under heavy-tailed data distribution, which cause the over-predictions towards head classes with rich instances and under-predictions towards tail classes with rare instance. In this work, we address both issues by introducing a novel single-stage heavy-tailed food classification framework. Our method is evaluated on two heavy-tailed food benchmark datasets, Food101-LT and VFN-LT, and achieves the best performance compared to existing work with over 5% improvements for top-1 accuracy. Jiangpeng He, Fengqing Zhu 0001 |
ICIP | 2 |
| 2023 | Efficient Joint Video Denoising and Super-ResolutionabstractDenoising and super-resolution are two important tasks for video enhancement. Despite recent progress for each task, there are very few works that target both tasks simultaneously. In this paper, we propose an efficient noise-robust video super-resolution method that is trained end-to-end for an input video containing observable noises. We investigate current approaches to address this joint denoising and super-resolution task and compare them to our proposed method. Experimental results show that our method achieves competitive reconstruction performance with existing solutions on various datasets while maintaining a low computation cost and a small model size which prove the effectiveness of our joint model design and training. Our code is available at "https://github.com/Eventhyn/EVDSRNet.". Yuning Huang, Qian Lin 0001, Jan P. Allebach, Fengqing Zhu 0001 |
ICIP | 5 |
| 2023 | An End-to-End Food Portion Estimation Framework Based on Shape Reconstruction from Monocular ImageabstractDietary assessment is a key contributor to monitoring health status. Existing self-report methods are tedious and time-consuming with substantial biases and errors. Image-based food portion estimation aims to estimate food energy values directly from food images, showing great potential for automated dietary assessment solutions. Existing image-based methods either use a single-view image or incorporate multi-view images and depth information to estimate the food energy, which either has limited performance or creates user burdens. In this paper, we propose an end-to-end deep learning framework for food energy estimation from a monocular image through 3D shape reconstruction. We leverage a generative model to reconstruct the voxel representation of the food object from the input image to recover the missing 3D information. Our method is evaluated on a publicly available food image dataset Nutrition5k, resulting a Mean Absolute Error (MAE) of 40.05 kCal and Mean Absolute Percentage Error (MAPE) of 11.47% for food energy estimation. Our method uses RGB image as the only input at the inference stage and achieves competitive results compared to the existing method requiring both RGB and depth information. Zeman Shao, Gautham Vinod, Jiangpeng He, Fengqing Zhu 0001 |
ICME | 4 |
| 2023 | Lossy Image Compression with Quantized Hierarchical VAEsabstractRecent work has shown a strong theoretical connection between variational autoencoders (VAEs) and the rate distortion theory. Motivated by this, we consider the problem of lossy image compression from the perspective of generative modeling. Starting from ResNet VAEs, which are originally designed for data (image) distribution modeling, we redesign their latent variable model using a quantization-aware posterior and prior, enabling easy quantization and entropy coding for image compression. Along with improved neural network blocks, we present a powerful and efficient class of lossy image coders, outperforming previous methods on natural image (lossy) compression. Our model compresses images in a coarse-to-fine fashion and supports parallel encoding and decoding, leading to fast execution on GPUs. Code is made available online. Zhihao Duan, Ming Lu 0003, Zhan Ma 0001, Fengqing Zhu 0001 |
WACV | 4 |
| 2023 | Unified Architecture Adaptation for Compressed Domain Semantic InferenceabstractAdvances in both lossy image compression and semantic content understanding have been greatly fueled by deep learning techniques, yet these two tasks have been developed separately for the past decades. In this work, we address the problem of directly executing semantic inference from quantized latent features in the deep compressed domain without pixel reconstruction. Although different methods have been proposed for this problem setting, they either are restrictive to a specific architecture, or are sub-optimal in terms of compressed domain task accuracy. In contrast, we propose a lightweight, plug-and-play solution which is generally compliant with popular learned image coders and deep vision models, making it attractive to vast applications. Our method adapts prevalent pixel domain neural models that are deployed for various vision tasks to directly accept quantized latent features (other than pixels). We further suggest training the compressed domain model by transferring knowledge from its corresponding pixel domain counterpart. Experiments show that our method is compliant with popular learned image coders and vision task models. Under fair comparison, our approach outperforms a baseline method by a) more than 3% top-1 accuracy for compressed domain classification, and b) more than 7% mIoU for compressed domain semantic segmentation, at various data rates. Zhihao Duan, Zhan Ma 0001, Fengqing Zhu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | 3DG-STFM: 3D Geometric Guided Student-Teacher Feature Matching
Runyu Mao, Yatong An, Fengqing Zhu 0001 |
ECCV (28) | 4 |
| 2022 | Exemplar-Free Online Continual LearningabstractTargeted for real world scenarios, online continual learning aims to learn new tasks from sequentially available data under the condition that each data is observed only once by the learner. Though recent works have made remarkable achievements by storing part of learned task data as exemplars for knowledge replay, the performance is greatly relied on the size of stored exemplars while the storage consumption is a significant constraint in continual learning. In addition, storing exemplars may not always be feasible for certain applications due to privacy concerns. In this work, we propose a novel exemplar-free method by leveraging nearest-class-mean (NCM) classifier where the class mean is estimated during training phase on all data seen so far through online mean update criteria. We focus on image classification task and conduct extensive experiments on benchmark datasets including CIFAR-100 and Food-1k. The results demonstrate that our method without using any exemplar outperforms state-of-the-art exemplar-based approaches with large margins under standard protocol (20 exemplars per class) and is able to achieve competitive performance even with larger exemplar size (100 exemplars per class). Jiangpeng He, Fengqing Zhu 0001 |
ICIP | 2 |
| 2022 | Opening the Black Box of Learned Image CodersabstractEnd-to-end learned lossy image coders (LICs), as opposed to hand-crafted image codecs, have shown increasing superiority in terms of the rate-distortion performance. However, they are mainly treated as black-box systems and their interpretability is not well studied. In this paper, we show that LICs learn a set of basis functions to transform input image for its compact representation in the latent space, as analogous to the orthogonal transforms used in image coding standards. Our analysis provides insights to help understand how learned image coders work and could benefit future design and development. Zhihao Duan, Ming Lu 0003, Zhan Ma 0001, Fengqing Zhu 0001 |
PCS | 4 |
| 2022 | Efficient Feature Compression for Edge-Cloud SystemsabstractOptimizing computation in an edge-cloud system is an important yet challenging problem. In this paper, we consider a three way trade-off between bit rate, classification accuracy, and encoding complexity in an edge-cloud image classification system. Our method includes a new training strategy and an efficient encoder architecture to improve the rate-accuracy performance. Our design can also be easily scaled according to different computation resources on the edge device, taking a step towards achieving rate-accuracy-complexity (RAC) trade-off. Under various settings, our feature coding system consistently outperforms previous methods in terms of the RAC performance. Code is made publicly available1. Zhihao Duan, Fengqing Zhu 0001 |
PCS | 2 |
| 2022 | Online Continual Learning Via Candidates VotingabstractContinual learning in online scenario aims to learn a sequence of new tasks from data stream using each data only once for training, which is more realistic than in offline mode assuming data from new task are all available. However, this problem is still under-explored for the challenging class-incremental setting in which the model classifies all classes seen so far during inference. Particularly, performance struggles with increased number of tasks or additional classes to learn for each task. In addition, most existing methods require storing original data as exemplars for knowledge replay, which may not be feasible for certain applications with limited memory budget or privacy concerns. In this work, we introduce an effective and memory-efficient method for online continual learning under class-incremental setting through candidates selection from each learned task together with prior incorporation using stored feature embeddings instead of original data as exemplars. Our proposed method implemented for image classification task achieves the best results under different benchmark datasets for online continual learning including CIFAR-10, CIFAR-100 and CORE-50 while requiring much less memory resource compared with existing works. Jiangpeng He, Fengqing Zhu 0001 |
WACV | 2 |
| 2021 | Turkey Behavior Identification Using Video Analytics And Object TrackingabstractIn this paper, we propose a method to identify behavior of experimental turkeys by automatically analyzing video recordings. Monitoring turkey health during production is crucial for improved turkey production. Turkey health can be reflected through their common behavior, and changes in the frequency and duration of their behavior can be used to detect sick turkeys early. Video recordings can be manually annotated to assist identifying turkey behaviors, but this is both time consuming and labor intensive. In this paper, we monitor and detect changes in turkey behavior using video analytics. Behaviors of interest include eating, drinking, preening, and pecking. Identifying these behaviors requires accurate estimates of turkeys’ and turkey heads’ locations. Re-identification of each turkey is crucial after significant shape deformation such as wing flapping and fast walking. Therefore, our system integrates a state-of-the-art turkey tracker and a head tracker with a behavior identification module to identify turkey behavior. Results demonstrate that our system is effective and accurate at estimating the spatial location of turkeys and their heads, and identifying all behaviors of interest with high recall. Shengtai Ju, Marisa A. Erasmus, Fengqing Zhu 0001, Amy R. Reibman |
ICIP | 3 |
| 2021 | Improving Dietary Assessment Via Integrated Hierarchy Food ClassificationabstractImage-based dietary assessment refers to the process of determining what someone eats and how much energy and nutrients are consumed from visual data. Food classification is the first and most crucial step. Existing methods focus on improving accuracy measured by the rate of correct classification based on visual information alone, which is very challenging due to the high complexity and inter-class similarity of foods. Further, accuracy in food classification is conceptual as description of a food can always be improved. In this work, we introduce a new food classification framework to improve the quality of predictions by integrating the information from multiple domains while maintaining the classification accuracy. We apply a multi-task network based on a hierarchical structure that uses both visual and nutrition domain specific information to cluster similar foods. Our method is validated on the modified VIPER-FoodNet (VFN) food image dataset by including associated energy and nutrient information. We achieve comparable classification accuracy with existing methods that use visual information only, but with less error in terms of energy and nutrient values for the wrong predictions. Runyu Mao, Jiangpeng He, Luotao Lin, Zeman Shao, Heather A. Eicher-Miller, Fengqing Zhu 0001 |
MMSP | 6 |
| 2021 | Towards Learning Food Portion From Monocular Images With Cross-Domain Feature AdaptationabstractWe aim to estimate food portion size, a property that is strongly related to the presence of food object in 3D space, from single monocular images under real life setting. Specifically, we are interested in end-to-end estimation of food portion size, which has great potential in the field of personal health management. Unlike image segmentation or object recognition where annotation can be obtained through large scale crowd sourcing, it is much more challenging to collect datasets for portion size estimation since human cannot accurately estimate the size of an object in an arbitrary 2D image without expert knowledge. To address such challenge, we introduce a real life food image dataset collected from a nutrition study where the groundtruth food energy (calorie) is provided by registered dietitians, and will be made available to the research community. We propose a deep regression process for portion size estimation by combining features estimated from both RGB and learned energy distribution domains. Our estimates of food energy achieved state-of-the-art with a MAPE of 11.47%, significantly outperforms non-expert human estimates by 27.56%. Zeman Shao, Shaobo Fang, Runyu Mao, Jiangpeng He, Janine Wright, Deborah A. Kerr, Carol J. Boushey, Fengqing Zhu 0001 |
MMSP | 8 |
| 2021 | Saliency-Aware Class-Agnostic Food Image SegmentationabstractAdvances in image-based dietary assessment methods have allowed nutrition professionals and researchers to improve the accuracy of dietary assessment, where images of food consumed are captured using smartphones or wearable devices. These images are then analyzed using computer vision methods to estimate energy and nutrition content of the foods. Food image segmentation, which determines the regions in an image where foods are located, plays an important role in this process. Current methods are data dependent and thus cannot generalize well for different food types. To address this problem, we propose a class-agnostic food image segmentation method. Our method uses a pair of eating scene images, one before starting eating and one after eating is completed. Using information from both the before and after eating images, we can segment food images by finding the salient missing objects without any prior information about the food class. We model a paradigm of top-down saliency that guides the attention of the human visual system based on a task to find the salient missing objects in a pair of images. Our method is validated on food images collected from a dietary study that showed promising results. Sri Yarlagadda, Daniel Mas Montserrat, David Guera, Carol J. Boushey, Deborah A. Kerr, Fengqing Zhu 0001 |
ACM Trans. Comput. Heal. | 6 |
| 2021 | Advances in Video Compression System Using Deep Neural Network: A Review and Case StudiesabstractSignificant advances in video compression systems have been made in the past several decades to satisfy the near-exponential growth of Internet-scale video traffic. From the application perspective, we have identified three major functional blocks, including preprocessing, coding, and postprocessing, which have been continuously investigated to maximize the end-user quality of experience (QoE) under a limited bit rate budget. Recently, artificial intelligence (AI)-powered techniques have shown great potential to further increase the efficiency of the aforementioned functional blocks, both individually and jointly. In this article, we review recent technical advances in video compression systems extensively, with an emphasis on deep neural network (DNN)-based approaches, and then present three comprehensive case studies. On preprocessing, we show a switchable texture-based video coding example that leverages DNN-based scene understanding to extract semantic areas for the improvement of a subsequent video coder. On coding, we present an end-to-end neural video coding framework that takes advantage of the stacked DNNs to efficiently and compactly code input raw videos via fully data-driven learning. On postprocessing, we demonstrate two neural adaptive filters to, respectively, facilitate the in-loop and postfiltering for the enhancement of compressed frames. Finally, a companion website hosting the contents developed in this work can be accessed publicly at https://purdueviper.github.io/dnn-coding/. Dandan Ding, Zhan Ma 0001, Qingshuang Chen, Zoe Liu, Fengqing Zhu 0001 |
Proc. IEEE | 6 |
| 2021 | A progressive CNN in-loop filtering approach for inter frame coding
Dandan Ding, Lingyi Kong, Fengqing Zhu 0001 |
Signal Process. Image Commun. | 4 |
| 2020 | Incremental Learning in Online ScenarioabstractModern deep learning approaches have achieved great success in many vision applications by training a model using all available task-specific data. However, there are two major obstacles making it challenging to implement for real life applications: (1) Learning new classes makes the trained model quickly forget old classes knowledge, which is referred to as catastrophic forgetting. (2) As new observations of old classes come sequentially over time, the distribution may change in unforeseen way, making the performance degrade dramatically on future data, which is referred to as concept drift. Current state-of-the-art incremental learning methods require a long time to train the model whenever new classes are added and none of them takes into consideration the new observations of old classes. In this paper, we propose an incremental learning framework that can work in the challenging online learning scenario and handle both new classes data and new observations of old classes. We address problem (1) in online mode by introducing a modified cross-distillation loss together with a two-step learning technique. Our method outperforms the results obtained from current state-of-the-art offline incremental learning methods on the CIFAR-100 and ImageNet-1000 (ILSVRC 2012) datasets under the same experiment protocol but in online scenario. We also provide a simple yet effective method to mitigate problem (2) by updating exemplar set using the feature of each new observation of old classes and demonstrate a real life application of online food image classification based on our complete framework using the Food-101 dataset. Jiangpeng He, Runyu Mao, Zeman Shao, Fengqing Zhu 0001 |
CVPR | 4 |
| 2020 | Learning Eating Environments Through Scene ClusteringabstractIt is well known that dietary habits have a significant influence on health. While many studies have been conducted to understand this relationship, little is known about the relationship between eating environments and health. Yet researchers and health agencies around the world have recognized the eating environment as a promising context for improving diet and health. In this paper, we propose an image clustering method to automatically extract the eating environments from eating occasion images captured during a community dwelling dietary study. Specifically, we are interested in learning how many different environments an individual consumes food in. Our method clusters images by extracting features at both global and local scales using a deep neural network. The variation in the number of clusters and images captured by different individual makes this a very challenging problem. Experimental results show that our method performs significantly better compared to several existing clustering approaches. Sri Yarlagadda, Sriram Baireddy, David Guera, Carol J. Boushey, Deborah A. Kerr, Fengqing Zhu 0001 |
ICASSP | 6 |
| 2019 | Semi-Automatic Crowdsourcing Tool for Online Food Image Collection and AnnotationabstractAssessing dietary intake accurately remains an open and challenging research problem. In recent years, image-based approaches have been developed to automatically estimate food intake by capturing eat occasions with mobile devices and wearable cameras. To build a reliable machine-learning models that can automatically map pixels to calories, successful image-based systems need large collections of food images with high quality groundtruth labels to improve the learned models. In this paper, we introduce a semi-automatic system for online food image collection and annotation. Our system consists of a web crawler, an automatic food detection method and a web-based crowdsoucing tool. The web crawler is used to download large sets of online food images based on the given food labels. Since not all retrieved images contain foods, we introduce an automatic food detection method to remove irrelevant images. We designed a web-based crowdsourcing tool to assist the crowd or human annotators to locate and label all the foods in the images. The proposed semi-automatic online food image collection system can be used to build large food image datasets with groundtruth labels efficiently from scratch. Zeman Shao, Runyu Mao, Fengqing Zhu 0001 |
IEEE BigData | 3 |
| 2019 | Pixel-level Texture Segmentation Based AV1 Video CompressionabstractModern video coding standards use hybrid coding techniques to remove spatial and temporal redundancy. However, efficient exploitation of statistical dependencies measured by a mean squared error (MSE) does not always produce the best psychovisual result. In this paper, we propose a pixel-level texture segmentation approach based on visual relevancy to improve the coding efficiency of newly developed AV1 video codec. Our method performs semantic segmentation and combines regions with similar texture in a video frame. These texture regions are then reconstructed using a motion model at the decoder instead of inter-frame prediction. A Convolutional Neural Networks based semantic segmentation combined with post-processing generates pixel-level texture masks that are more accurate compared to block-based texture masks in our previous work. We show that for many standard test sets, the proposed method achieves significant data rate reductions with improved visual quality. Qingshuang Chen, Fengqing Zhu 0001 |
ICASSP | 3 |
| 2019 | Shadow Removal Detection and Localization for Forensics AnalysisabstractThe recent advancements in image processing and computer vision allow realistic photo manipulations. In order to avoid the distribution of fake imagery, the image forensics community is working towards the development of image authenticity verification tools. Methods based on shadow analysis are particularly reliable since they are part of the physical integrity of the scene, thus detecting forgeries is possible whenever inconsistencies are found (e.g., shadows not coherent with the light direction). An attacker can easily delete inconsistent shadows and replace them with correctly cast shadows in order to fool forensics detectors based on physical analysis. In this paper, we propose a method to detect shadow removal done with state-of-the-art tools. The proposed method is based on a conditional generative adversarial network (cGAN) specifically trained for shadow removal detection. Sri Yarlagadda, David Guera, Daniel Mas Montserrat, Fengqing Zhu 0001, Edward J. Delp, Paolo Bestagini, Stefano Tubaro |
ICASSP | 4 |
| 2018 | Single-View Food Portion Estimation: Learning Image-to-Energy Mappings Using Generative Adversarial NetworksabstractDue to the growing concern of chronic diseases and other health problems related to diet, there is a need to develop accurate methods to estimate an individual's food and energy intake. Measuring accurate dietary intake is an open research problem. In particular, accurate food portion estimation is challenging since the process of food preparation and consumption impose large variations on food shapes and appearances. In this paper, we present a food portion estimation method to estimate food energy (kilocalories) from food images using Generative Adversarial Networks (GAN). We introduce the concept of an “energy distribution” for each food image. To train the GAN, we design a food image dataset based on ground truth food labels and segmentation masks for each food image as well as energy information associated with the food image. Our goal is to learn the mapping of the food image to the food energy. We can then estimate food energy based on the energy distribution. We show that an average energy estimation error rate of 10.89% can be obtained by learning the image-to-energy mapping. Shaobo Fang, Zeman Shao, Runyu Mao, Chichen Fu, Edward J. Delp, Fengqing Zhu 0001, Deborah A. Kerr, Carol J. Boushey |
ICIP | 6 |
| 2018 | Reliability Map Estimation for CNN-Based Camera Model AttributionabstractAmong the image forensic issues investigated in the last few years, great attention has been devoted to blind camera model attribution. This refers to the problem of detecting which camera model has been used to acquire an image by only exploiting pixel information. Solving this problem has great impact on image integrity assessment as well as on authenticity verification. Recent advancements that use convolutional neural networks (CNNs) in the media forensic field have enabled camera model attribution methods to work well even on small image patches. These improvements are also important for determining forgery localization. Some patches of an image may not contain enough information related to the camera model (e.g., saturated patches). In this paper, we propose a CNN-based solution to estimate the camera model attribution reliability of a given image patch. We show that we can estimate a reliabilitymap indicating which portions of the image contain reliable camera traces. Testing using a well known dataset confirms that by using this information, it is possible to increase small patch camera model attribution accuracy by more than 8% on a single patch. David Guera, Fengqing Zhu 0001, Sri Yarlagadda, Stefano Tubaro, Paolo Bestagini, Edward J. Delp |
WACV | 2 |
| 2018 | Context based image analysis with application in dietary assessment and evaluation
Yu Wang 0034, Ye He 0001, Carol J. Boushey, Fengqing Zhu 0001, Edward J. Delp |
Multim. Tools Appl. | 4 |
| 2017 | Weakly supervised food image segmentation using class activation mapsabstractFood image segmentation plays a crucial role in image-based dietary assessment and management. Successful methods for object segmentation generally rely on a large amount of labeled data on the pixel level. However, such training data are not yet available for food images and expensive to obtain. In this paper, we describe a weakly supervised convolutional neural network (CNN) which only requires image level annotation. We propose a graph based segmentation method which uses the class activation maps trained on food datasets as a top-down saliency model. We evaluate the proposed method for both classification and segmentation tasks. We achieve competitive classification accuracy compared to the previously reported results. Yu Wang 0034, Fengqing Zhu 0001, Carol J. Boushey, Edward J. Delp |
ICIP | 2 |
| 2016 | A comparison of food portion size estimation using geometric models and depth imagesabstractSix of the ten leading causes of death in the United States, including cancer, diabetes, and heart disease, can be directly linked to diet. Dietary intake, the process of determining what someone eats during the course of a day, provides valuable insights for mounting intervention programs for prevention of many of the above chronic diseases. Measuring accurate dietary intake is considered to be an open research problem in the nutrition and health fields. In this paper we compare two techniques to estimating food portion size from images of food. The techniques are based on 3D geometric models and depth images. An expectation-maximization based technique is developed to detect the reference plane in depth images, which is essential for portion size estimation using depth images. Our experimental results indicate that volume estimation based on geometric model is more accurate for objects with well-defined 3D shapes compared to estimation using depth images. Shaobo Fang, Fengqing Zhu 0001, Chufan Jiang, Song Zhang 0002, Carol J. Boushey, Edward J. Delp |
ICIP | 2 |
| 2016 | Efficient superpixel based segmentation for food image analysisabstractIn this paper, we propose a segmentation method based on normalized cut and superpixels. The method relies on color and texture cues for fast computation and efficient use of memory. The method is used for food image segmentation as part of a mobile food record system we have developed for dietary assessment and management. The accurate estimate of nutrients relies on correctly labelled food items and sufficiently well-segmented regions. Our method achieves competitive results using the Berkeley Segmentation Dataset and outperforms some of the most popular techniques in a food image dataset. Yu Wang 0034, Chang Liu 0004, Fengqing Zhu 0001, Carol J. Boushey, Edward J. Delp |
ICIP | 3 |
| 2015 | Single-View Food Portion Estimation Based on Geometric ModelsabstractIn this paper we present a food portion estimation technique based on a single-view food image used for the estimation of the amount of energy (in kilocalories) consumed at a meal. Unlike previous methods we have developed, the new technique is capable of estimating food portion without manual tuning of parameters. Although single-view 3D scene reconstruction is in general an ill-posed problem, the use of geometric models such as the shape of a container can help to partially recover 3D parameters of food items in the scene. Based on the estimated 3D parameters of each food item and a reference object in the scene, the volume of each food item in the image can be determined. The weight of each food can then be estimated using the density of the food item. We were able to achieve an error of less than 6% for energy estimation of an image of a meal assuming accurate segmentation and food classification. Shaobo Fang, Chang Liu 0004, Fengqing Zhu 0001, Edward J. Delp, Carol J. Boushey |
ISM | 3 |
| 2015 | Multiple Hypotheses Image Segmentation and Classification With Application to Dietary AssessmentabstractWe propose a method for dietary assessment to automatically identify and locate food in a variety of images captured during controlled and natural eating events. Two concepts are combined to achieve this: a set of segmented objects can be partitioned into perceptually similar object classes based on global and local features; and perceptually similar object classes can be used to assess the accuracy of image segmentation. These ideas are implemented by generating multiple segmentations of an image to select stable segmentations based on the classifier's confidence score assigned to each segmented image region. Automatic segmented regions are classified using a multichannel feature classification system. For each segmented region, multiple feature spaces are formed. Feature vectors in each of the feature spaces are individually classified. The final decision is obtained by combining class decisions from individual feature spaces using decision rules. We show improved accuracy of segmenting food images with classifier feedback. Fengqing Zhu 0001, Marc Bosch, Nitin Khanna, Carol J. Boushey, Edward J. Delp |
IEEE J. Biomed. Health Informatics | 1 |
| 2011 | Combining global and local features for food identification in dietary assessmentabstractMany chronic diseases, such as heart diseases, diabetes, and obesity, can be related to diet. Hence, the need to accurately measure diet becomes imperative. We are developing methods to use image analysis tools for the identification and quantification of food consumed at a meal. In this paper we describe a new approach to food identification using several features based on local and global measures and a "voting" based late decision fusion classifier to identify the food items. Experimental results on a wide variety of food items are presented. Marc Bosch, Fengqing Zhu 0001, Nitin Khanna, Carol J. Boushey, Edward J. Delp |
ICIP | 2 |
| 2011 | Integrated database system for mobile dietary assessment and analysisabstractOf the 10 leading causes of death in the US, 6 are related to diet. Unfortunately, methods for real-time assessment and proactive health management of diet do not currently exist. There are only minimally successful tools for historical analysis of diet and food consumption available. In this paper, we present an integrated database system that provides a unique perspective on how dietary assessment can be accomplished. We have designed three interconnected databases: an image database that contains data generated by food images, an experiments database that contains data related to nutritional studies and results from the image analysis, and finally an enhanced version of a nutritional database by including both nutritional and visual descriptions of each food. We believe that these databases provide tools to the healthcare community and can be used for data mining to extract diet patterns of individuals and/or entire social groups. Marc Bosch, TusaRebecca Schap, Fengqing Zhu 0001, Nitin Khanna, Carol J. Boushey, Edward J. Delp |
ICME | 3 |
| 2010 | An image analysis system for dietary assessment and evaluationabstractThere is a growing concern about chronic diseases and other health problems related to diet including obesity and cancer. Dietary intake provides valuable insights for mounting intervention programs for prevention of chronic diseases. Measuring accurate dietary intake is considered to be an open research problem in the nutrition and health fields. In this paper, we describe a novel mobile telephone food record that provides a measure of daily food and nutrient intake. Our approach includes the use of image analysis tools for identification and quantification of food that is consumed at a meal. Images obtained before and after foods are eaten are used to estimate the amount and type of food consumed. The mobile device provides a unique vehicle for collecting dietary information that reduces the burden on respondents that are obtained using more classical approaches for dietary assessment. We describe our approach to image analysis that includes the segmentation of food items, features used to identify foods, a method for automatic portion estimation, and our overall system architecture for collecting the food intake information. Fengqing Zhu 0001, Marc Bosch, Carol J. Boushey, Edward J. Delp |
ICIP | 1 |
| 2009 | Perceptual quality evaluation for texture and motion based video codingabstractOne approach that can be used to increase compression efficiency beyond the data rates achievable by state-of-the-art video codecs is to use content-based methods whereby not all the pixels are conventionally encoded. An approach to reduce the data rate is to use different coding methods for pixels belonging to areas containing large amount of detail that are costly to encode, for example textures. This can be extended by focusing on the semantic meaning of objects represented in the video sequence and also taking into consideration human visual system properties. The goal is to determine where ¿detail-irrelevant¿ regions are located in the frame and synthesize them with acceptable perceptual quality. In this paper, we discuss the effects and trade-offs of these techniques based on a set of perceptual experiments and analyze how these areas can influence the viewer's attention. Marc Bosch, Fengqing Zhu 0001, Edward J. Delp |
ICIP | 2 |
| 2009 | An overviewof texture and motion based video coding at Purdue UniversityabstractIn recent years there has been a growing interest in developing novel techniques for increasing the coding efficiency of video compression methods. We approach the problem by not encoding all the pixels, in particular, regions belonging to areas that the viewer will not perceive the specific details in the scene could be skipped or encoded at a much lower data rate. This approach can also be expanded by considering a model of the human visual perception system. In this paper we review some of the approaches we have investigated at Purdue University. The goal is to determine where ldquodetail-irrelevantrdquo regions in the frame are located and not encode them. We will also discuss a set of subjective quality evaluation experiments to determine what is the overall perceptual quality of these approaches. Marc Bosch, Fengqing Zhu 0001, Edward J. Delp |
PCS | 2 |
| 2008 | Video coding using motion classificationabstractIn this paper we present a video coding approach similar to texture- based methods but based on motion models. We consider motion perception properties instead of spatial texture properties of the video sequence. We integrate a motion classification algorithm to separate foreground objects containing noticeable motion from the background. These background areas are labeled as skipped areas that are not encoded. After decoding, frame reconstruction is performed by inserting the skipped background into the decoded frames. We are able to show as much as 15% an improvement over previous texture- based implementations in terms of video compression efficiency. Marc Bosch, Fengqing Zhu 0001, Edward J. Delp |
ICIP | 2 |
| 2007 | Spatial Texture Models for Video CompressionabstractIn this paper we integrate several spatial texture tools into a texture-based video coding scheme. We implemented texture techniques and segmentation strategies in order to detect texture regions in video sequences. These textures are analyzed using temporal motion techniques and are labeled as skipped areas that are not encoded. After the decoding process, frame reconstruction is performed by inserting the skipped texture areas into the decoded frames. We are able to show an improvement over previous texture-based implementations in terms of compression efficiency. Marc Bosch, Fengqing Zhu 0001, Edward J. Delp |
ICIP (1) | 2 |
| 2007 | Spatial and temporal models for texture-based video codingabstractIn this paper, we investigate spatial and temporal models for texture analysis and synthesis. The goal is to use these models to increase the coding efficiency for video sequences containing textures. The models are used to segment texture regions in a frame at the encoder and synthesize the textures at the decoder. These methods can be incorporated into a conventional video coder (e.g. H.264) where the regions to be modeled by the textures are not coded in a usual manner but texture model parameters are sent to the decoder as side information. We showed that this approach can reduce the data rate by as much as 15%. Fengqing Zhu 0001, Ka Ki Ng, Golnaz Abdollahian, Edward J. Delp |
VCIP | 1 |