Edmund Y. Lam

dblp:87/5852 · DBLP profile ↗
← Back
97ranked-venue papers
4as first author
50since 2021 · last 2026
0000-0001-6268-950XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 46 · 4 first-author · 30 since 2021Applied, interdisciplinary, general and emerging computing · 24 · 7 since 2021Artificial intelligence and machine learning · 22 · 15 since 2021Human-computer interaction and ubiquitous computing · 12 · 1 since 2021Systems, architecture and hardware · 9 · 3 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021Theory of computation · 5 · 5 since 2021Computer networks · 2
YearPublicationVenuePosition
2026 CAT++: Enhancing 3D Annotations with Hierarchical-Interleaved Encoding and Attention-Conditioned Implicit Representation
Xiaoyan Qian 0001, Chang Liu 0094, Xiaojuan Qi 0001, Siew-Chong Tan, Edmund Y. Lam, Ngai Wong 0001
Int. J. Comput. Vis.5
2026 Image Restoration Learning via Noisy Supervision in Fourier Domain
abstract
Noisy supervision refers to supervising network learning with targets corrupted by noise, encompassing both weakly supervised learning with noisy targets and fully unsupervised denoising using unpaired noisy images. It alleviates the data collection burden and enhances the practical applicability of deep learning techniques. Existing methods face two main limitations: they are ineffective at handling noise with long-range correlations, commonly found in real-world scenarios such as low-light imaging and remote sensing, and rely on pixel-wise loss functions that offer limited supervision for image deblurring and super-resolution. This work addresses these challenges by leveraging the Fourier domain, where spatially correlated noise exhibits sparsity and independence, and Fourier coefficients capture global information that enables stronger supervision. We prove that Fourier coefficients of a wide range of noise converge in distribution to the Gaussian distribution and establish a statistical equivalence between learning with clean and noisy targets in the Fourier domain. Based on these insights, we develop a weakly supervised framework for image restoration learning with noisy targets, and construct a fully unsupervised denoising method tailored to stripe-wise noise. Extensive experiments show that our approaches achieve superior performance in both quantitative metrics and perceptual quality.
Haosen Liu 0001, Tan Shan, Edmund Y. Lam
IEEE Trans. Image Process.4
2026 Dark-EvGS: Event Camera as an Eye for Radiance Field in the Dark
abstract
In low-light environments, conventional cameras often struggle to capture clear multi-view images of objects due to dynamic range limitations and motion blur caused by long exposure. Event cameras, with their high-dynamic range and high-speed properties, have the potential to mitigate these issues. Additionally, 3D Gaussian Splatting (GS) enables radiance field reconstruction, facilitating bright frame synthesis from multiple viewpoints in low-light conditions. However, naively using an event-assisted 3D GS approach still faced challenges because, in low lights, events are noisy, frames lack quality, and the color tone may be inconsistent. To address these issues, we propose Dark-EvGS, the first event-assisted 3D GS framework that enables the reconstruction of bright frames from arbitrary viewpoints along the camera trajectory. Triplet-level supervision is proposed to gain holistic knowledge, granular details, and sharp scene rendering. The color tone matching block is proposed to guarantee the color consistency of the rendered frames. Furthermore, we introduce the first real-captured dataset for the event-guided bright frame synthesis task via 3D GS-based radiance field reconstruction. Experiments demonstrate that our method achieves better results than existing methods, conquering radiance field reconstruction under challenging low-light conditions. The code and sample data are included in the supplementary material.
Jingqian Wu, Peiqi Duan 0002, Zongqiang Wang, Changwei Wang 0001, Boxin Shi, Edmund Y. Lam
IEEE Trans. Image Process.6
2026 EventTracer: Fast Path Tracing-Based Event Stream Rendering
abstract
Simulating event streams from 3D scenes has become a common practice in event-based vision research, as it meets the demand for large-scale, high temporal frequency data without setting up expensive hardware devices or undertaking extensive data collections. Yet existing methods in this direction typically work with noiseless RGB frames that are costly to render, and therefore their simulations are often unrealistically low in temporal resolution. In this work, we propose EventTracer, a path tracing-based rendering pipeline that simulates high-fidelity event sequences from complex 3D scenes in an efficient and physics-aware manner. Specifically, we speed up the rendering process via low sample-per-pixel (SPP) path tracing, and train a lightweight event spiking network to denoise the resulting RGB videos into realistic event sequences. Our EventTracerpipeline runs at a speed of $\sim$∼1 minutes per second of 360p video, and it inherits the merit of accurate spatiotemporal modeling from its path tracing backbone. We show through the Real2Sim and Sim2Real tests that EventTracercaptures higher-fidelity scene details and demonstrates a greater similarity to real-world event data than alternative event simulators, which establishes it as a potential tool for creating large-scale event-RGB datasets, narrowing the sim-to-real gap in event-based vision, and boosting various downstream applications.
Xiaoyang Bai, Jinfan Lu, Pengfei Shen, Edmund Y. Lam, Yifan Peng 0001
IEEE Trans. Vis. Comput. Graph.5
2025 ColorGPT: Automatic Colorization with Generative Prompts and Transformer
abstract
Automatic image colorization, where grayscale images are transformed into plausible color images, is a critical task with applications in film restoration, digital archiving, and artistic creation. Despite progress, challenges persist in achieving realistic and contextually accurate colorization due to the ambiguity in mapping grayscale intensities to colors and the need for semantic understanding. We design ColorGPT, a framework integrating generative language models and advanced colorization techniques to address these challenges. Our approach features three key modules: (1) a Prompt Generation Module using large language models to create diverse and context-aware color descriptions, (2) a Score Estimation Module employing CLIP to rank and select the most semantically relevant descriptions, and (3) a Colorization Module with transformer for high-quality image generation. Experiments on Set5 and Set14 demonstrate that ColorGPT outperforms state-of-the-art methods, achieving PSNR improvements of 0.7 dB on Set5 and 1.2 dB on Set14.
Juntao Qiu, Edmund Y. Lam
ICIP3
2025 Eigen-Component Analysis: A Quantum Theory-Inspired Linear Model
abstract
In modern machine learning circuits and systems, the need for centralized, standardized data poses significant challenges, especially in privacy-sensitive, resource-constrained, real-time, or asynchronous environments. We introduce Eigen-Component Analysis (ECA), a quantum-inspired linear model that circumvents these requirements by leveraging intrinsic data structures. ECA uses an orthogonal transformation to extract meaningful features from decentralized, non-standardized data, ensuring high classification accuracy without extensive preprocessing. Unlike traditional linear models, ECA adapts to complex data distributions and shows superior performance in feature extraction and classification across various datasets. Experiments indicate that ECA achieves robust accuracy with fewer parameters, highlighting its potential where data centralization and standardization are impractical. Our work extends the feasibility of linear models in diverse scenarios, enhancing interpretability and efficiency in challenging data environments.
Rongzhou Chen, Hanghang Liu, Haohan Xu, Edmund Y. Lam
ISCAS6
2025 Laconic Neuromorphic Vision
abstract
In this paper, we introduce laconic neuromorphic vision, a novel visual intelligence system that encodes scene dynamics into sparse spike representations and processes them using spiking neural networks, mimicking the sparse activation patterns of biological neurons. This approach drastically reduces the input data required while maintaining high accuracy in image classification. Experiments on the MNIST and Fashion-MNIST datasets demonstrate that our system achieves comparable results to traditional methods with significantly fewer input pixels and spikes. This work provides a fresh perspective on energy-efficient, biologically inspired vision systems, with potential applications in edge computing and autonomous systems. Code and data will be released upon manuscript acceptance.
Edmund Y. Lam
ISCAS2
2025 SweepEvGS: Event-Based 3D Gaussian Splatting for Macro and Micro Radiance Field Rendering From a Single Sweep
abstract
Recent advancements in 3D Gaussian Splatting (3D-GS) have demonstrated the potential of using 3D Gaussian primitives for high-speed, high-fidelity, and cost-efficient novel view synthesis from continuously calibrated input views. However, conventional methods require high-frame-rate dense and high-quality sharp images, which are time-consuming and inefficient to capture, especially in dynamic environments. Event cameras, with their high temporal resolution and ability to capture asynchronous brightness changes, offer a promising alternative for more reliable scene reconstruction without motion blur. In this paper, we propose SweepEvGS, a novel hardware-integrated method that leverages event cameras for robust and accurate novel view synthesis across various imaging settings from a single sweep. SweepEvGS utilizes the initial static frame with dense event streams captured during a single camera sweep to effectively reconstruct detailed scene views. We also introduce different real-world hardware imaging systems for real-world data collection and evaluation for future research. We validate the robustness and efficiency of SweepEvGS through experiments in three different imaging settings: synthetic objects, real-world macro-level, and real-world micro-level view synthesis. Our results demonstrate that SweepEvGS surpasses existing methods in visual rendering quality, rendering speed, and computational efficiency, highlighting its potential for dynamic practical applications.
Jingqian Wu, Chutian Wang, Boxin Shi, Edmund Y. Lam
IEEE Trans. Circuits Syst. Video Technol.5
2025 Neuromorphic Imaging With Super-Resolution
abstract
Neuromorphic imaging is an emerging technique that imitates the human retina to sense variations in dynamic scenes. It responds to pixel-level brightness changes by asynchronous streaming events and boasts microsecond temporal precision over a high dynamic range, yielding blur-free recordings under extreme illumination. Nevertheless, this modality falls short in spatial resolution and leads to a low level of visual richness and clarity. Pursuing hardware upgrades is expensive and might cause compromised performance due to more burdens on computational requirements. Another option is to harness offline, plug-in-play super-resolution solutions. However, existing ones, which demand substantial sample volumes for lengthy training on massive computing resources, are largely restricted by real data availability owing to the current imperfect high-resolution devices, as well as the randomness and variability of motion. To tackle these challenges, we introduce the first self-supervised neuromorphic super-resolution prototype. It can be self-adaptive to per input source from any low-resolution camera to estimate an optimal, high-resolution counterpart of any scale, without the need of side knowledge and prior training. Evaluated on downstream tasks, such a simple yet effective method can obtain competitive results against the state-of-the-arts, significantly promoting flexibility but not sacrificing accuracy. It also delivers enhancements for inferior natural images and optical micrographs acquired under non-ideal imaging conditions, breaking through the limitations that are challenging to overcome with frame-based techniques. In the current landscape where the use of high-resolution cameras for event-based sensing remains an open debate, our solution is a cost-efficient and practical alternative, paving the way for more intelligent imaging systems.
Chutian Wang, Edmund Y. Lam
IEEE Trans. Circuits Syst. Video Technol.5
2025 Segment Anything Model Is a Good Teacher for Local Feature Learning
abstract
Local feature detection and description play an important role in many computer vision tasks, which are designed to detect and describe keypoints in any scene and any downstream task. Data-driven local feature learning methods need to rely on pixel-level correspondence for training. However, a vast number of existing approaches ignored the semantic information on which humans rely to describe image pixels. In addition, it is not feasible to enhance generic scene keypoints detection and description simply by using traditional common semantic segmentation models because they can only recognize a limited number of coarse-grained object classes. In this paper, we propose SAMFeat to introduce SAM (segment anything model), a foundation model trained on 11 million images, as a teacher to guide local feature learning. SAMFeat learns additional semantic information brought by SAM and thus is inspired by higher performance even with limited training samples. To do so, first, we construct an auxiliary task of Attention-weighted Semantic Relation Distillation (ASRD), which adaptively distillates feature relations with category-agnostic semantic information learned by the SAM encoder into a local feature learning network, to improve local feature description using semantic discrimination. Second, we develop a technique called Weakly Supervised Contrastive Learning Based on Semantic Grouping (WSC), which utilizes semantic groupings derived from SAM as weakly supervised signals, to optimize the metric space of local descriptors. Third, we design an Edge Attention Guidance (EAG) to further improve the accuracy of local feature detection and description by prompting the network to pay more attention to the edge region guided by SAM. SAMFeat's performance on various tasks, such as image matching on HPatches, and long-term visual localization on Aachen Day-Night showcases its superiority over previous local features. The release code is available at https://github.com/vignywang/SAMFeat.
Jingqian Wu, Rongtao Xu, Zach Wood-Doughty, Changwei Wang 0001, Shibiao Xu, Edmund Y. Lam
IEEE Trans. Image Process.6
2024 Unsupervised Disparity Estimation for Light Field Videos
abstract
Light field (LF) videos contain not only the spatial-angular information but also the temporal information, which are useful for disparity estimation. The existing work on disparity estimation for LF videos relies on supervised training with disparity labels. To overcome this reliance, we develop an unsupervised disparity estimation framework for LF videos, which consists of a matching branch to perform feature matching and a refinement branch to refine the disparity maps. Our framework also includes a cross-feature fusion module with self-attention and cross-attention to fuse the multi-frame features, and a cost aggregation transformer with cross-depth self-attention blocks to explore the global depth dependencies. Moreover, we propose a left-right consistency strategy to estimate the occlusion regions for the input views and introduce a occlusion-aware photometric loss to solve the occlusion issue. Experimental results demonstrate that our method achieves superior performance compared to the existing supervised and unsupervised methods.
Shansi Zhang, Edmund Y. Lam
ICASSP2
2024 SASA: Saliency-Aware Self-Adaptive Snapshot Compressive Imaging
abstract
The ability of snapshot compressive imaging (SCI) systems to efficiently capture high-dimensional (HD) data depends on the advent of novel optical designs to sample the HD data as two-dimensional (2D) compressed measurements. Nonetheless, the traditional SCI scheme is fundamentally limited, due to the complete disregard for high-level information in the sampling process. To tackle this issue, in this paper, we pave the first mile toward the advanced design of adaptive coding masks for SCI. Specifically, we propose an efficient and effective algorithm to generate coding masks with the assistance of saliency detection, in a low-cost and low-power fashion. Experiments demonstrate the effectiveness and efficiency of our approach. Code is available at: https://github.com/IndigoPurple/SASA.
Edmund Y. Lam
ICASSP2
2024 Cross-Camera Human Motion Transfer by Time Series Analysis
abstract
With advances in optical sensor technology, heterogeneous camera systems are increasingly used for high-resolution (HR) video acquisition and analysis. However, motion transfer across multiple cameras poses challenges. To address this, we propose an algorithm based on time series analysis that identifies motion seasonality and constructs an additive model to extract transferable patterns. Validated on real-world data, our algorithm demonstrates effectiveness and interpretability. Notably, it improves pose estimation in low-resolution videos by leveraging patterns derived from HR counterparts, enhancing practical utility. Code is available at: https://github.com/IndigoPurple/TSAMT.
Guanghan Li, Edmund Y. Lam
ICASSP3
2024 Controllable Unsupervised Event-Based Video Generation
abstract
The advent of event cameras, with their unique asynchronous sensing capabilities to capture the edge details of moving objects, has sparked new directions in video generation. So far, the challenge of integrating event-based data for controllable video generation remains largely unexplored. Addressing this gap, we introduce a framework that leverages the edge information from events and combines it with textual descriptions to synthesize videos without the requirement of extensive training. The framework marks a pioneering venture into event-based video generation using diffusion models. Comprehensive evaluations demonstrate the superior performance of our framework compared to existing methods. Code is available at: https://github.com/IndigoPurple/CUBE.
Chutian Wang, Edmund Y. Lam
ICIP4
2024 A field deployable imaging system for detecting microplastics in the aquatic environment
abstract
Microplastic (MP) pollutants, which are widely distributed and resistant to decomposition, pose serious threats to both ecosystems and human health. Developing efficient, field deployable, and low-cost MP detection systems is of paramount importance for promoting sustainable development. In this work, we design an integrated imaging system for detecting MPs in various aquatic environments. To address the challenge of image degradation caused by scattering effects, a de-scattering algorithm based on synthetic polarization holography is proposed. Calibration results demonstrate a substantial enhancement in image contrast and spatial resolution. Furthermore, our imaging system enables simultaneous capture of amplitude, phase, and polarization information, which is advantageous for the quantification and identification of MPs. Consequently, this field deployable imaging system offers an efficient solution for realtime monitoring of MPs in turbid aquatic environments.
Jianqing Huang, Yanmin Zhu 0005, Edmund Y. Lam
ISCAS4
2024 Neuromorphic imaging and classification with graph learning
Chutian Wang, Edmund Y. Lam
Neurocomputing3
2024 Attention-Guided Low-Rank Tensor Completion
abstract
Low-rank tensor completion (LRTC) aims to recover missing data of high-dimensional structures from a limited set of observed entries. Despite recent significant successes, the original structures of data tensors are still not effectively preserved in LRTC algorithms, yielding less accurate restoration results. Moreover, LRTC algorithms often incur high computational costs, which hinder their applicability. In this work, we propose an attention-guided low-rank tensor completion (AGTC) algorithm, which can faithfully restore the original structures of data tensors using deep unfolding attention-guided tensor factorization. First, we formulate the LRTC task as a robust factorization problem based on low-rank and sparse error assumptions. Low-rank tensor recovery is guided by an attention mechanism to better preserve the structures of the original data. We also develop implicit regularizers to compensate for modeling inaccuracies. Then, we solve the optimization problem by employing an iterative technique. Finally, we design a multistage deep network by unfolding the iterative algorithm, where each stage corresponds to an iteration of the algorithm; at each stage, the optimization variables and regularizers are updated by closed-form solutions and learned deep networks, respectively. Experimental results for high dynamic range imaging and hyperspectral image restoration show that the proposed algorithm outperforms state-of-the-art algorithms.
Truong Thanh Nhat Mai, Edmund Y. Lam, Chul Lee
IEEE Trans. Pattern Anal. Mach. Intell.2
2024 Light Field Image Restoration via Latent Diffusion and Multi-View Attention
abstract
Light field (LF) images contain information for multiple views. The restoration of degraded LF images is of great significance for various LF applications. Inspired by the recent achievement of denoising diffusion models, we propose a LF image restoration method based on latent diffusion (LD). We design a LDUNet with efficient cross-attention modules to integrate the features of conditional input, and propose a two-stage training strategy, where the LDUNet is first trained on the individual views and then fine-tuned on the LF images with injected prior noise. A refinement module is jointly trained in the second stage to enhance the spatial-angular structures. It consists of multi-view attention blocks with patch-based angular self-attention to fuse the global view information. Moreover, we introduce an enhanced noise loss for better noise prediction and an auxiliary image loss to obtain high-quality images. We evaluate our method on LF image deraining task and low-light LF image enhancement task. Our method demonstrates superior performance on both tasks compared to the existing methods.
Shansi Zhang, Edmund Y. Lam
IEEE Signal Process. Lett.2
2024 Unsupervised Light Field Depth Estimation via Multi-View Feature Matching With Occlusion Prediction
abstract
Depth estimation from light field (LF) images is a fundamental step for numerous applications. Recently, learning-based methods have achieved higher accuracy and efficiency than the traditional methods. However, it is costly to obtain sufficient depth labels for supervised training. In this paper, we propose an unsupervised framework to estimate depth from LF images. First, we design a disparity estimation network (DispNet) with a coarse-to-fine structure to predict disparity maps from different view combinations. It explicitly performs multi-view feature matching to learn the correspondences effectively. As occlusions may cause the violation of photo-consistency, we introduce an occlusion prediction network (OccNet) to predict the occlusion maps, which are used as the element-wise weights of photometric loss to solve the occlusion issue and assist the disparity learning. With the disparity maps estimated by multiple input combinations, we then propose a disparity fusion strategy based on the estimated errors with effective occlusion handling to obtain the final disparity map with higher accuracy. Experimental results demonstrate that our method achieves superior performance on both the dense and sparse LF images, and also shows better robustness and generalization on the real-world LF images compared to the other methods.
Shansi Zhang, Nan Meng, Edmund Y. Lam
IEEE Trans. Circuits Syst. Video Technol.3
2024 Progressive Self-Supervised Pretraining for Hyperspectral Image Classification
abstract
Self-supervised learning has demonstrated considerable success in hyperspectral image (HSI) classification when limited labeled data is available. However, inherent dissimilarities among HSIs require self-supervised pre-training from scratch for each HSI dataset. Pre-training on a large amount of unlabeled data can be time consuming. Additionally, the poor quality of some HSIs can limit the performance of self-supervised learning algorithms. To address these issues, we propose to enhance self-supervised pre-training on HSIs with transfer learning. We introduce a progressive self-supervised pre-training framework that acquires strong initialization for the final pre-training on the target HSI dataset by sequentially performing self-supervised pre-training on datasets that are increasingly similar to the target HSI, specifically, first on a large general vision dataset and then on a related HSI dataset. This sequential strategy enables the model to progressively learn from domain-general vision knowledge to target-specific hyperspectral knowledge. To mitigate the catastrophic forgetting in sequential training, we develop a regularization method, called self-supervised elastic weight consolidation, to impose adaptive constraints on the changes to model parameters. Thorough classification experiments on various HSI datasets demonstrate that our framework significantly and consistently improves the self-supervised pre-training on HSIs in terms of both convergence speed and representation quality. Furthermore, our framework exhibits high generalizability and can be applied to various self-supervised learning algorithms. Transfer learning continues to prove its usefulness in self-supervised settings.
Peiyan Guan, Edmund Y. Lam
IEEE Trans. Geosci. Remote. Sens.2
2024 Deep Unfolding Tensor Rank Minimization With Generalized Detail Injection for Pansharpening
abstract
Pansharpening aims to generate a high-resolution multispectral (HRMS) image by merging a low-resolution multispectral (LRMS) image with a high-resolution panchromatic (PAN) image. While traditional model-based pansharpening algorithms have strong theoretical foundations, their performance and generalizability are limited by handcrafted formulations. In contrast, recent deep learning approaches outperform model-based algorithms but do not effectively consider the physical properties of multispectral (MS) images, such as their spatial and spectral dependencies. These physical properties facilitate the exploitation of the actual imaging process, leading to enhanced spatial and spectral fidelities. In this work, we propose a deep unfolded tensor rank minimization framework with generalized detail injection for pansharpening to overcome the weaknesses of both model- and learning-based approaches while leveraging their advantages. Specifically, we first formulate the pansharpening task as a tensor rank minimization problem to exploit the low-rankness of MS images, providing a robust theoretical foundation on the physical properties of MS data. We also develop a generalized detail injection component, which effectively exploits the information in the PAN images, and incorporate it into the optimization to improve generalizability and representation capability. Then, we define a data-driven regularizer to compensate for modeling inaccuracies in the low-rank model and solve the optimization problem using an iterative technique. Finally, the iterative algorithm is unfolded into a multistage deep network, in which the optimization variables are solved by closed-form solutions and a data-driven regularizer in each stage. Experimental results on various MS image datasets demonstrate that the proposed algorithm achieves better pansharpening performance and interpretability than state-of-the-art algorithms.
Truong Thanh Nhat Mai, Edmund Y. Lam, Chul Lee
IEEE Trans. Geosci. Remote. Sens.2
2024 Neuromorphic Imaging With Joint Image Deblurring and Event Denoising
abstract
Neuromorphic imaging reacts to per-pixel brightness changes of a dynamic scene with high temporal precision and responds with asynchronous streaming events as a result. It also often supports a simultaneous output of an intensity image. Nevertheless, the raw events typically involve a large amount of noise due to the high sensitivity of the sensor, while capturing fast-moving objects at low frame rates results in blurry images. These deficiencies significantly degrade human observation and machine processing. Fortunately, the two information sources are inherently complementary - events with microsecond-level temporal resolution, which are triggered by the edges of objects recorded in a latent sharp image, can supply rich motion details missing from the blurry one. In this work, we bring the two types of data together and introduce a simple yet effective unifying algorithm to jointly reconstruct blur-free images and noise-robust events in an iterative coarse-to-fine fashion. Specifically, an event-regularized prior offers precise high-frequency structures and dynamic features for blind deblurring, while image gradients serve as a kind of faithful supervision in regulating neuromorphic noise removal. Comprehensively evaluated on real and synthetic samples, such a synergy delivers superior reconstruction quality for both images with severe motion blur and raw event streams with a storm of noise, and also exhibits greater robustness to challenging realistic scenarios such as varying levels of illumination, contrast and motion magnitude. Meanwhile, it can be driven by much fewer events and holds a competitive edge at computational time overhead, rendering itself preferable as available computing resources are limited. Our solution gives impetus to the improvement of both sensing data and paves the way for highly accurate neuromorphic reasoning and analysis.
Haosen Liu 0001, Zhou Ge, Chutian Wang, Edmund Y. Lam
IEEE Trans. Image Process.5
2024 Semi-Supervised Semantic Segmentation for Light Field Images Using Disparity Information
abstract
Light field (LF) images enable numerous applications due to their ability to capture information for multiple views. Semantic segmentation is an essential task for LF scene understanding. However, existing supervised methods heavily rely on a large number of pixel-wise annotations. To relieve this problem, we propose a semi-supervised LF semantic segmentation method that requires only a small subset of labeled data and harnesses the LF disparity information. First, we design an unsupervised disparity estimation network, which can determine the disparity map for every view. With the estimated disparity maps, we generate pseudo-labels along with their weight maps for the peripheral views when only the labels of central views are available. We then merge the predictions from multiple views to obtain more reliable pseudo-labels for unlabeled data, and introduce a disparity-semantics consistency loss to enforce structure similarity. Moreover, we develop a comprehensive contrastive learning scheme that includes a pixel-level strategy to enhance feature representations and an object-level strategy to improve segmentation for individual objects. Our method demonstrates state-of-the-art performance on the benchmark LF semantic segmentation dataset under a variety of training settings and achieves comparable performance to supervised methods when trained under 1/2 protocol.
Shansi Zhang, Edmund Y. Lam
IEEE Trans. Image Process.3
2023 Context-Aware Transformer for 3D Point Cloud Automatic Annotation
abstract
3D automatic annotation has received increased attention since manually annotating 3D point clouds is laborious. However, existing methods are usually complicated, e.g., pipelined training for 3D foreground/background segmentation, cylindrical object proposals, and point completion. Furthermore, they often overlook the inter-object feature correlation that is particularly informative to hard samples for 3D annotation. To this end, we propose a simple yet effective end-to-end Context-Aware Transformer (CAT) as an automated 3D-box labeler to generate precise 3D box annotations from 2D boxes, trained with a small number of human annotations. We adopt the general encoder-decoder architecture, where the CAT encoder consists of an intra-object encoder (local) and an inter-object encoder (global), performing self-attention along the sequence and batch dimensions, respectively. The former models intra-object interactions among points and the latter extracts feature relations among different objects, thus boosting scene-level understanding. Via local and global encoders, CAT can generate high-quality 3D box annotations with a streamlined workflow, allowing it to outperform existing state-of-the-arts by up to 1.79% 3D AP on the hard task of the KITTI test set.
Xiaoyan Qian 0001, Chang Liu 0094, Xiaojuan Qi 0001, Siew-Chong Tan, Edmund Y. Lam, Ngai Wong 0001
AAAI5
2023 Solving Inverse Problems in Compressive Imaging with Score-Based Generative Models
abstract
Snapshot Compressive Imaging (SCI) is a technique for capturing high-dimensional data through snapshot measurements using a two-dimensional (2D) detector. This approach is accomplished via coded aperture compressive temporal imaging (CACTI), which involves applying a temporally variant mask to spatially encode each sequential signal before aggregating the encoded information into a single compressed measurement. The objective of our work is to develop algorithms capable of reconstructing each video frame as a 3D data cube from its 2D measurement. To achieve this goal, we introduce multiple approaches that utilize unconditional and pre-trained 2D score models for video frame reconstruction. Our method involves modeling both the forward perturbation process of the data distribution and its reverse process as stochastic differential equations (SDEs). We also employ score-based deep learning models to estimate the scores of the data distribution across different time steps. Differing from many applications, our sampling process relies on the observed measurement, which directly corresponds to pixel values, rather than class labels. We demonstrate that employing traditional score-based generative methods with 2D score models in SCI, or integrating them into the plug-and-play framework as a deep generative prior, presents challenges. Furthermore, we propose ideas to address these limitations for future research.
Zhen Yuen Chong, Edmund Y. Lam
DSAA4
2023 Adaptive Compressed Sensing for Real-Time Video Compression, Transmission, and Reconstruction
abstract
The real-time transmission of videos with both high resolution and high frame rate is challenging, due to the limited storage space and significant communication overhead. To meet the real-time requirement, these issues are usually tackled by video quality reduction, which compromises the user experience. While previous methods, such as video compressed sensing, have attempted to address these issues, they often employ a fixed compression rate without considering the varying channel gain and do not adequately address the real-time transmission requirements. To mitigate these shortcomings, we propose an adaptive compressed sensing framework that optimizes the compression rate based on the channel state. This approach equivalently optimizes the video quality while ensuring real-time transmission by reducing communication overhead and thus latency. The feasibility and performance of our method are validated and discussed through extensive experiments on both classic and custom datasets.
Qunsong Zeng, Edmund Y. Lam
DSAA3
2023 PIDRo: Parallel Isomeric Attention with Dynamic Routing for Text-Video Retrieval
abstract
Text-video retrieval is a fundamental task with high practical value in multi-modal research. Inspired by the great success of pre-trained image-text models with large-scale data, such as CLIP, many methods are proposed to transfer the strong representation learning capability of CLIP to text-video retrieval. However, due to the modality difference between videos and images, how to effectively adapt CLIP to the video domain is still underexplored. In this paper, we investigate this problem from two aspects. First, we enhance the transferred image encoder of CLIP for fine-grained video understanding in a seamless fashion. Second, we conduct fine-grained contrast between videos and texts from both model improvement and loss design. Particularly, we propose a fine-grained contrastive model equipped with parallel isomeric attention and dynamic routing, namely PIDRo, for text-video retrieval. The parallel isomeric attention module is used as the video encoder, which consists of two parallel branches modeling the spatial-temporal information of videos from both patch and frame levels. The dynamic routing module is constructed to enhance the text encoder of CLIP, generating informative word representations by distributing the fine-grained information to the related word tokens within a sentence. Such model design provides us with informative patch, frame and word representations. We then conduct token-wise interaction upon them. With the enhanced encoders and the token-wise loss, we are able to achieve finer-grained text-video alignment and more accurate retrieval. PIDRo obtains state-of-the-art performance over various text-video retrieval benchmarks, including MSR-VTT, MSVD, LSMDC, DiDeMo and ActivityNet.
Peiyan Guan, Renjing Pei, Jianzhuang Liu, Weimian Li, Jiaxi Gu, Hang Xu 0004, Songcen Xu, Youliang Yan, Edmund Y. Lam
ICCV10
2023 Improving Video Colorization by Test-Time Tuning
abstract
With the advancements in deep learning, video colorization by propagating color information from a colorized reference frame to a monochrome video sequence has been well explored. However, the existing approaches often suffer from overfitting the training dataset and sequentially lead to suboptimal performance on colorizing testing samples. To address this issue, we propose an effective method, which aims to enhance video colorization through test-time tuning. By exploiting the reference to construct additional training samples during testing, our approach achieves a performance boost of 1 ∼ 3 dB in PSNR on average compared to the baseline. Code is available at: https://github.com/IndigoPurple/T3.
Haitian Zheng, Jiebo Luo 0001, Edmund Y. Lam
ICIP4
2023 Point Cloud Denoising Via Momentum Ascent in Gradient Fields
abstract
To achieve point cloud denoising, traditional methods heavily rely on geometric priors, and most learning-based approaches suffer from outliers and loss of details. Recently, the gradient-based method was proposed to estimate the gradient fields from the noisy point clouds using neural networks, and refine the position of each point according to the estimated gradient. However, the predicted gradient could fluctuate, leading to perturbed and unstable solutions, as well as a long inference time. To address these issues, we develop the momentum gradient ascent method that leverages the information of previous iterations in determining the trajectories of the points, thus improving the stability of the solution and reducing the inference time. Experiments demonstrate that the proposed method outperforms state-of-the-art approaches with a variety of point clouds, noise types, and noise levels. Code is available at: https://github.com/IndigoPurple/MAG.
Haitian Zheng, Jiebo Luo 0001, Edmund Y. Lam
ICIP5
2023 LRT: An Efficient Low-Light Restoration Transformer for Dark Light Field Images
abstract
Light field (LF) images containing information for multiple views have numerous applications, which can be severely affected by low-light imaging. Recent learning-based methods for low-light enhancement have some disadvantages, such as a lack of noise suppression, complex training process and poor performance in extremely low-light conditions. To tackle these deficiencies while fully utilizing the multi-view information, we propose an efficient Low-light Restoration Transformer (LRT) for LF images, with multiple heads to perform intermediate tasks within a single network, including denoising, luminance adjustment, refinement and detail enhancement, achieving progressive restoration from small scale to full scale. Moreover, we design an angular transformer block with an efficient view-token scheme to model the global angular dependencies, and a multi-scale spatial transformer block to encode the multi-scale local and global information within each view. To address the issue of insufficient training data, we formulate a synthesis pipeline by simulating the major noise sources with the estimated noise parameters of LF camera. Experimental results demonstrate that our method achieves the state-of-the-art performance on low-light LF restoration with high efficiency.
Shansi Zhang, Nan Meng, Edmund Y. Lam
IEEE Trans. Image Process.3
2022 PATE: Property, Amenities, Traffic and Emotions Coming Together for Real Estate Price Prediction
abstract
Real estate prices have a significant impact on individuals, families, businesses, and governments. The general objective of real estate price prediction is to identify and exploit socioeconomic patterns arising from real estate transactions over multiple aspects, ranging from the property itself to other contributing factors. However, price prediction is a challenging multidimensional problem that involves estimating many characteristics beyond the property itself. In this paper, we use multiple sources of data to evaluate the economic contribution of different socioeconomic characteristics such as surrounding amenities, traffic conditions and social emotions. Our experiments were conducted on 28,550 houses in Beijing, China and we rank each characteristic by its importance. Since the use of multisource information improves the accuracy of predictions, the aforementioned characteristics can be an invaluable resource to assess the economic and social value of real estate. Code and data are available at: https://github.com/IndigoPurple/PATE.
Ramgopal Ravi, Shuhui Shi, Edmund Y. Lam, Jichang Zhao
DSAA5
2022 H4M: Heterogeneous, Multi-source, Multi-modal, Multi-view and Multi-distributional Dataset for Socioeconomic Analytics in the Case of Beijing
abstract
The study of socioeconomic status has been reformed by the availability of digital records containing data on real estate, points of interest, traffic and social media trends such as micro-blogging. In this paper, we describe a heterogeneous, multi-source, multi-modal, multi-view and multi-distributional dataset named "H4M". The mixed dataset contains data on real estate transactions, points of interest, traffic patterns and micro-blogging trends from Beijing, China. The unique composition of H4M makes it an ideal test bed for methodologies and approaches aimed at studying and solving problems related to real estate, traffic, urban mobility planning, social sentiment analysis etc. The dataset is available at: https://indigopurple.github.io/H4M/index.html.
Shuhui Shi, Ramgopal Ravi, Edmund Y. Lam, Jichang Zhao
DSAA5
2022 Improving Source Localization by Perturbing Graph Diffusion
abstract
Graph diffusion has quite common phenomenons in our daily life, such as misinformation propagation. As the inverse problem of graph diffusion, the goal of source localization is to identify those nodes of the network from which the information started to spread. Though graph diffusion has been well explored in the literature, the emerging source localization problem is important yet challenging because of its intrinsic ill-posed characteristics. While graph neural networks (GNN) are recently utilized to implement source localization and achieve state-of-the-art performance, a general GNN framework consists of two stages: feature construction and label propagation. Typically, a neural network is pretrained in the feature construction, and then combine with additional functions to jointly perform finetuning for source localization. However, those emerging methods have risks in overfitting the feature construction task, which usually has a gap with the target downstream task of source localization. Such a gap is neglected by previous methods and leads to suboptimal performance. To address this issue, we propose a very simple yet effective method to help better finetune feature construction on the source localization task by adding some noise to the parameters of the feature construction model before finetuning. More specifically, we utilize a matrix-wise perturbing method that adds different uniform noises to different parameter matrices, and design the noise considering the variances and magnitude of network weights. Extensive experiments on six real-world datasets show the proposed method can consistently empower the finetuning of different pretrained feature construction models on the downstream source localization task. Moreover, we conduct an ablation study to investigate the performance with different noise types and intensities. Code is available at: https://github.com/IndigoPurple/PGD.
Edmund Y. Lam
DSAA3
2022 Multimodal Transformer for Automatic 3D Annotation and Object Detection
Chang Liu 0094, Xiaoyan Qian 0001, Binxiao Huang, Xiaojuan Qi 0001, Edmund Y. Lam, Siew-Chong Tan, Ngai Wong 0001
ECCV (38)5
2022 Manet: Improving Video Denoising with a Multi-Alignment Network
abstract
In video denoising, the adjacent frames often provide very useful information, but accurate alignment is needed before such information can be harnassed. In this work, we present a multi-alignment network, which generates multiple flow proposals followed by attention-based averaging. It serves to mimic the non-local mechanism, suppressing noise by averaging multiple observations. Our approach can be applied to various state-of-the-art models that are based on flow estimation. Experiments on a large-scale video dataset demonstrate that our method improves the denoising baseline model by 0.2 dB, and further reduces the parameters by 47% with model distillation. Code is available at https://github.com/IndigoPurple/MANet.
Haitian Zheng, Jiebo Luo 0001, Edmund Y. Lam
ICIP5
2022 MAP-Gen: An Automated 3D-Box Annotation Flow with Multimodal Attention Point Generator
abstract
Manually annotating 3D point clouds is laborious and costly, limiting the training data preparation for deep learning in real-world object detection. While a few previous studies tried to automatically generate 3D bounding boxes from weak labels such as 2D boxes, the quality is sub-optimal compared to human annotators. This work proposes a novel autolabeler, called multimodal attention point generator (MAP-Gen), that generates high-quality 3D labels from weak 2D boxes. It leverages dense image information to tackle the sparsity issue of 3D point clouds, thus improving label quality. For each 2D pixel, MAP-Gen predicts its corresponding 3D coordinates by referencing context points based on their 2D semantic or geometric relationships. The generated 3D points densify the original sparse point clouds, followed by an encoder to regress 3D bounding boxes. Using MAP-Gen, object detection networks that are weakly supervised by 2D boxes can achieve 94 ~ 99% performance of those fully supervised by 3D annotations. It is hopeful this newly proposed MAP-Gen autolabeling flow can shed new light on utilizing multimodal information for enriching sparse point clouds.
Chang Liu 0094, Xiaoyan Qian 0001, Xiaojuan Qi 0001, Edmund Y. Lam, Siew-Chong Tan, Ngai Wong 0001
ICPR4
2022 A Deep Retinex Framework for Light Field Restoration under Low-light Conditions
abstract
Light field (LF) images can record the scene from multiple directions and have many applications, such as refocusing and depth estimation. However, these applications can be heavily influenced by poor light condition and noise. This work aims to recover the high-quality LF images from their lowlight detection. First, a decomposition network is employed to decompose each LF image into its reflectance and illumination with the Retinex theory. Then, two enhancement networks are designed to denoise the reflectance and enhance the illumination, respectively. They adopt alternate spatial-angular feature extractions and process all the views synchronously with high efficiency. A parallel dual attention mechanism is integrated to both the spatial and angular feature extractions, to encode more important information. Moreover, a discriminator is introduced during the training to generate more realistic LF images by making judgment according to both the spatial and angular characteristics. Experimental results have demonstrated the superior performance of our method, which can restore the content, luminance, color and geometric structures of LF images effectively.
Shansi Zhang, Edmund Y. Lam
ICPR2
2022 Three-Branch Multilevel Attentive Fusion Network for Hyperspectral Pansharpening
abstract
In this paper, we propose a three-branch multilevel attentive fusion network (TMA-Net) for hyperspectral pansharpening, which aims to merge low-resolution hyperspectral images (LR-HSIs) and high-resolution panchromatic images (HR-PANs) to obtain HSIs with high resolution. We construct three branches to extract rich features of the two images and the correlation between them, which enables us to capture abundant useful information for pansharpening. We merge the multilevel features extracted by each branch in multiple steps to fully fuse the useful information. An attentive fusion module (AFM) is designed to guide the fusion procedure. It explores the relation between different features and employs attention mechanism to refine them adaptively. The experimental results illustrate the superiority of the TMA-Net.
Peiyan Guan, Edmund Y. Lam
IGARSS2
2022 Spatial-Spectral Contrastive Learning for Hyperspectral Image Classification
abstract
In spite of being widely used in hyperspectral image (HSI) classification, most deep learning algorithms require plenty of labeled samples to achieve satisfactory performance. However, manual labeling is very time-consuming and laborious in practice. To solve such problem, we propose a method, called spatial-spectral contrastive learning (SSCL), to learn the representations of HSIs suitable for classification in an unsuper-vised manner. We demonstrate that the useful contents (e.g., semantics) are invariant under the spatial and spectral domains while the uninformative ones are usually not. Thus, we learn powerful representations that model domain-invariant information by defining a contrastive prediction task. Specifically, two signals are constructed for an HSI sample to include information of the two domains, and the representations of these two signals are then optimized to be similar, such that the domain-invariant contents are extracted. We conduct classification experiments on the learned representation with very few labels, the results of which verify the superiority of our method over the state-of-the-art techniques.
Peiyan Guan, Edmund Y. Lam
IGARSS2
2022 A Cross-Attention BERT-Based Framework for Continuous Sign Language Recognition
abstract
Continuous sign language recognition (CSLR) is a challenging task involving various signal processing techniques to infer the sequences of glosses performed by signers. Existing approaches in CSLR typically use multiple input modalities such as the raw video data and the extracted hand images to improve their recognition accuracy. However, the large modality differences make it difficult to define an integrative framework to effectively exchange and combine the knowledge obtained from different modalities such that they can complement each other for improving the framework's robustness against the gesture variations and background noises in CSLR. To address this issue, we propose a novel cross-attention deep learning framework named the CA-SignBERT. This framework utilizes multiple Bidirectional Encoder Representations from Transformers (BERT) models to analyze the information from different modalities. Among these BERT models, we introduce a special cross-attention mechanism to ensure an efficient inter-modality knowledge exchange. Besides, an innovative weight control module is proposed to dynamically hybridize their outputs. Experimental results reveal that the CA-SignBERT framework attains state-of-the-art performance in four benchmark CSLR datasets.
Zhenxing Zhou, Vincent W. L. Tam, Edmund Y. Lam
IEEE Signal Process. Lett.3
2022 Multistage Dual-Attention Guided Fusion Network for Hyperspectral Pansharpening
abstract
Deep learning, especially the convolutional neural network, has been widely applied to solve the hyperspectral pansharpening problem. However, most do not explore the intraimage characteristics and the interimage correlation concurrently due to the limited representation ability of the networks, which may lead to insufficient fusion of valuable information encoded in the high-resolution panchromatic images (HR-PANs) and low-resolution hyperspectral images (LR-HSIs). To cope with this problem, we develop a hyperspectral pansharpening method called multistage dual-attention guided fusion network (MDA-Net) to fully extract the important information and accurately fuse them. It employs a three-stream structure, which enables the network to incorporate the intrinsic characteristics of each input and correlation among them simultaneously. In order to combine as much information as possible, we merge the features extracted from three streams in multiple stages, where a dual-attention guided fusion block (DAFB) with spectral and spatial attention mechanisms is utilized to fuse the features efficiently. It identifies the useful components in both spatial and spectral domains, which are beneficial to improving the fusion accuracy. Moreover, we design a multiscale residual dense block (MRDB) to extract dense and hierarchical features, which improves the representation power of the network. Experiments are conducted on both real and simulated datasets. The evaluation results validate the superiority of the MDA-Net.
Peiyan Guan, Edmund Y. Lam
IEEE Trans. Geosci. Remote. Sens.2
2022 Cross-Domain Contrastive Learning for Hyperspectral Image Classification
abstract
Despite the success of deep learning algorithms in hyperspectral image (HSI) classification, most deep learning models require a large amount of labeled data to optimize the numerous parameters. However, it is very expensive and time-consuming to collect a lot of labeled HSI samples. To cope with this problem, we propose a cross-domain contrastive learning (XDCL) framework to learn representations of HSIs in an unsupervised manner. We demonstrate that the features that are valuable for category identification are shared across the spectral and spatial domains, while the less useful contents tend to be independent. The XDCL extracts such domain-invariant information with a cross-domain discrimination task, i.e., predicting which two representations of different domains are matched. With this insight, our method learns semantically meaningful HSI representations. We develop a simple method to construct effective signals representing the two domains, respectively. Moreover, we randomly mask the signals to improve their semantic level and encourage the representations to dig out more useful abstract factors. In order to evaluate the representation quality, we use the learned representations to train a linear classifier on three hyperspectral datasets with limited labeled samples. Experimental results demonstrate that our method surpasses the state-of-the-art methods by a large margin.
Peiyan Guan, Edmund Y. Lam
IEEE Trans. Geosci. Remote. Sens.2
2022 Deep Unrolled Low-Rank Tensor Completion for High Dynamic Range Imaging
abstract
The major challenge in high dynamic range (HDR) imaging for dynamic scenes is suppressing ghosting artifacts caused by large object motions or poor exposures. Whereas recent deep learning-based approaches have shown significant synthesis performance, interpretation and analysis of their behaviors are difficult and their performance is affected by the diversity of training data. In contrast, traditional model-based approaches yield inferior synthesis performance to learning-based algorithms despite their theoretical thoroughness. In this paper, we propose an algorithm unrolling approach to ghost-free HDR image synthesis algorithm that unrolls an iterative low-rank tensor completion algorithm into deep neural networks to take advantage of the merits of both learning- and model-based approaches while overcoming their weaknesses. First, we formulate ghost-free HDR image synthesis as a low-rank tensor completion problem by assuming the low-rank structure of the tensor constructed from low dynamic range (LDR) images and linear dependency among LDR images. We also define two regularization functions to compensate for modeling inaccuracy by extracting hidden model information. Then, we solve the problem efficiently using an iterative optimization algorithm by reformulating it into a series of subproblems. Finally, we unroll the iterative algorithm into a series of blocks corresponding to each iteration, in which the optimization variables are updated by rigorous closed-form solutions and the regularizers are updated by learned deep neural networks. Experimental results on different datasets show that the proposed algorithm provides better HDR image synthesis performance with superior robustness compared with state-of-the-art algorithms, while using significantly fewer training samples.
Truong Thanh Nhat Mai, Edmund Y. Lam, Chul Lee
IEEE Trans. Image Process.2
2022 Dual Alternating Direction Method of Multipliers for Inverse Imaging
abstract
Inverse imaging covers a wide range of imaging applications, including super-resolution, deblurring, and compressive sensing. We propose a novel scheme to solve such problems by combining duality and the alternating direction method of multipliers (ADMM). In addition to a conventional ADMM process, we introduce a second one that solves the dual problem to find the estimated nontrivial lower bound of the objective function, and the related iteration results are used in turn to guide the primal iterations. We call this D-ADMM, and show that it converges to the global minimum when the regularization function is convex and the optimization problem has at least one optimizer. Furthermore, we show how the scheme can give rise to two specific algorithms, called D-ADMM-L2 and D-ADMM-TV, by having different regularization functions. We compare D-ADMM-TV with other methods on image super-resolution and demonstrate comparable or occasionally slightly better quality results. This paves the way of incorporating advanced operators and strategies designed for basic ADMM into the D-ADMM method as well to further improve the performances of those methods.
Zhou Ge, Edmund Y. Lam
IEEE Trans. Image Process.3
2021 A New Approach for Educational Data Analytics with Wearable Devices
abstract
The rapid development of wearable technologies has dramatically promoted the potential usages of wearable devices in educational data analytics. However, the large amount of input data and the various types of educational output labels also increase the difficulties in selecting the useful information and discovering the implicit relations between different input data. To address this issue, this paper proposed a new two-layer approach for conducting educational data analytics automatically. In this approach, there are three key components: input layer, output layer and recognition model. For the input layer, we adopted the newly proposed optimization algorithm: Adaptive Multi-Population Optimization (AMPO) to select the most related input features and suitable model structures. For the output layer, we inserted domain-specific constraints during the searching for all combinations of different output labels to discover a meaningful output strategy with a relatively higher accuracy. Based on the input elements and output strategy provided by the input layer and the output layer, the recognition model will produce the corresponding recognition accuracy. With these three components, our proposed method can find out some connotative information to provide guidance for conducting educational data analytics and drawing meaningful conclusions.
Zhenxing Zhou, Vincent W. L. Tam, King-Shan Lui, Edmund Y. Lam, Runzhi Kong, Xiao Hu 0001, Nancy Law
ICALT4
2021 Ghost-Free HDR Imaging Via Unrolling Low-Rank Matrix Completion
abstract
We propose a ghost-free high dynamic range (HDR) image synthesis algorithm by unrolling low-rank matrix completion. By exploiting the low-rank structure of the irradiance maps from low dynamic range (LDR) images, we formulate ghost-free HDR imaging as a general low-rank matrix completion problem. Then, we solve the problem iteratively using the augmented Lagrange multiplier (ALM) method. At each iteration, the optimization variables are updated by closed-form solutions and the regularizers are updated by learned deep neural networks. Experimental results show that the proposed algorithm provides better image qualities with fewer visual artifacts compared to state-of-the-art algorithms.
Truong Thanh Nhat Mai, Edmund Y. Lam, Chul Lee
ICIP2
2021 Learning to restore light fields under low-light imaging
Shansi Zhang, Edmund Y. Lam
Neurocomputing2
2021 High-Dimensional Dense Residual Convolutional Neural Network for Light Field Reconstruction
abstract
We consider the problem of high-dimensional light field reconstruction and develop a learning-based framework for spatial and angular super-resolution. Many current approaches either require disparity clues or restore the spatial and angular details separately. Such methods have difficulties with non-Lambertian surfaces or occlusions. In contrast, we formulate light field super-resolution (LFSR) as tensor restoration and develop a learning framework based on a two-stage restoration with 4-dimensional (4D) convolution. This allows our model to learn the features capturing the geometry information encoded in multiple adjacent views. Such geometric features vary near the occlusion regions and indicate the foreground object border. To train a feasible network, we propose a novel normalization operation based on a group of views in the feature maps, design a stage-wise loss function, and develop the multi-range training strategy to further improve the performance. Evaluations are conducted on a number of light field datasets including real-world scenes, synthetic data, and microscope light fields. The proposed method achieves superior performance and less execution time comparing with other state-of-the-art schemes.
Nan Meng, Hayden Kwok-Hay So, Xing Sun 0001, Edmund Y. Lam
IEEE Trans. Pattern Anal. Mach. Intell.4
2021 An effective decomposition-enhancement method to restore light field images captured in the dark
Shansi Zhang, Edmund Y. Lam
Signal Process.2
2021 Light Field View Synthesis via Aperture Disparity and Warping Confidence Map
abstract
This paper presents a learning-based approach to synthesize the view from an arbitrary camera position given a sparse set of images. A key challenge for this novel view synthesis arises from the reconstruction process, when the views from different input images may not be consistent due to obstruction in the light path. We overcome this by jointly modeling the epipolar property and occlusion in designing a convolutional neural network. We start by defining and computing the aperture disparity map, which approximates the parallax and measures the pixel-wise shift between two views. While this relates to free-space rendering and can fail near the object boundaries, we further develop a warping confidence map to address pixel occlusion in these challenging regions. The proposed method is evaluated on diverse real-world and synthetic light field scenes, and it shows better performance over several state-of-the-art techniques.
Nan Meng, Jianzhuang Liu, Edmund Y. Lam
IEEE Trans. Image Process.4
2020 High-Order Residual Network for Light Field Super-Resolution
abstract
Plenoptic cameras usually sacrifice the spatial resolution of their SAIs to acquire geometry information from different viewpoints. Several methods have been proposed to mitigate such spatio-angular trade-off, but seldom make use of the structural properties of the light field (LF) data efficiently. In this paper, we propose a novel high-order residual network to learn the geometric features hierarchically from the LF for reconstruction. An important component in the proposed network is the high-order residual block (HRB), which learns the local geometric features by considering the information from all input views. After fully obtaining the local features learned from each HRB, our model extracts the representative geometric features for spatio-angular upsampling through the global residual learning. Additionally, a refinement network is followed to further enhance the spatial details by minimizing a perceptual loss. Compared with previous work, our model is tailored to the rich structure inherent in the LF, and therefore can reduce the artifacts near non-Lambertian and occlusion regions. Experimental results show that our approach enables high-quality reconstruction even in challenging regions and outperforms state-of-the-art single image or LF reconstruction methods with both quantitative measurements and visual evaluation.
Nan Meng, Jianzhuang Liu, Edmund Y. Lam
AAAI4
2020 A Portable Hong Kong Sign Language Translation Platform with Deep Learning and Jetson Nano
abstract
As hearing loss is arousing more and more public concern, different researches have been conducted on translating the sign language into spoken language. However, most of these researches remain in a theoretical level and few of them investigate how to realize a real system. In this paper, we introduce an effective and portable Hong Kong sign language recognition platform which can translate the Hong Kong sign language within a few seconds. In this platform, there are mainly two parts: a mobile application and a Jetson Nano. The mobile application accounts for preprocessing the sign video and transferring the videos to Jetson Nano. Then, Jetson Nano will translate sign videos into spoken language with the pretrained deep learning model and return the results to the mobile application. With this platform, non-disabled people can easily translate and understand the sign performed by deaf people through mobile phones quickly. We believe that this platform can significantly facilitate the daily communication between deaf people and the others in Hong Kong.
Zhenxing Zhou, Yisiang Neo, King-Shan Lui, Vincent W. L. Tam, Edmund Y. Lam, Ngai Wong 0001
ASSETS5
2020 A Sophisticated Platform for Learning Analytics with Wearable Devices
abstract
With the rapid development in wearable technology, wearable devices integrating with various sensors have been broadly applied in different areas. Yet there is seldom any previous study which focuses on applying wearable devices and deep learning in learning analytics. This paper considers a sophisticated real-time learning analytics platform for analyzing students' learning states and learning activities with wearable devices and deep learning. During the experimental period of this platform, students will receive instant notifications from an intelligent mobile application when their heart rate are out of their normal range so that the actual learning activities conducted by students can be collected to train deep learning models for recognizing their learning activities. At the same time, students can enjoy the sleeping monitoring and the exercise monitoring functionalities provided by the smart watches in this platform. The results of the interviews conducted after the experiment for this platform demonstrate that 89% of students think that this platform is useful for their daily lives and 65% of students report that this platform brings positive effects on their learning in different aspects. More importantly, this work sheds lights on the possibility of applying wearable devices in learning analytics to improve the learning effectivenesses and life qualities of students.
Z. X. Zhou, Vincent W. L. Tam, King-Shan Lui, Edmund Y. Lam, Xiao Hu 0001, Allan Hoi Kau Yuen, Nancy Law
ICALT4
2020 MBD-GAN: Model-based image deblurring with a generative adversarial network
abstract
This paper presents a methodology to tackle inverse imaging problems by leveraging the synergistic power of imaging model and deep learning. The premise is that while learning-based techniques have quickly become the methods of choice in various applications, they often ignore the prior knowledge embedded in imaging models. Incorporating the latter has the potential to improve the image estimation. Specifically, we first provide a mathematical basis of using generative adversarial network (GAN) in inverse imaging through considering an optimization framework. Then, we develop the specific architecture that connects the generator and discriminator networks with the imaging model. While this technique can be applied to a variety of problems, from image reconstruction to super-resolution, we take image deblurring as the example here, where we show in detail the implementation and experimental results of what we call the model-based deblurring GAN (MBD-GAN).
Edmund Y. Lam
ICPR2
2020 Applying (3+2+1)D Residual Neural Network with Frame Selection for Hong Kong Sign Language Recognition
abstract
As reported by Hong Kong Government in 2017, there are more than 1.5 million residents suffering from hearing impairment in Hong Kong. Most of them rely on Hong Kong Sign Language for daily communication while there are only 63 registered sign language interpreters in Hong Kong. To address this specific social issue and also facilitate the effective communication between the hearing impaired and other people, this paper introduces a word-level Hong Kong Sign Language(HKSL) dataset which currently includes 45 isolated words and at least 30 sign videos per word performed by different signers(more than 1500 videos in total now and still enlarging). Based on this dataset, this paper systemically compares the performances of various deep learning approaches, including (1) 2D histogram of oriented gradients(HOG) feature/pose estimation/feature extraction with long-short term memory(LSTM) layer; (2) 3D Residual Neural Network(ResNet) (3) (2+1)D Residual Neural Network, in HKSL recognition. Meanwhile, to further improve the accuracy of sign language recognition, this paper proposes a novel method called (3+2+1)D ResNet Model with Frame Selection which adopts blurriness detection with Laplacian kernel to construct high-quality video clips and also combines both (2+1)D and 3D ResNet for recognizing the sign language. At the end, the experimental results show that the proposed method outperforms other deep learning approaches and attains an impressive accuracy of 94.6% in our dataset.
Zhenxing Zhou, King-Shan Lui, Vincent W. L. Tam, Edmund Y. Lam
ICPR4
2020 Holographic Classifier: Deep Learning in Digital Holography for Automatic Micro-objects Classification
abstract
Micro-objects, such as microplastics and particulate pollution, need to be accurately observed and detected by high-precision optical systems. Digital holography is a powerful tool to detect such microscopic objects. However, traditional digital holography requires additional image processing such as phase unwrapping, de-noising, and refocusing, which costs a lot of time and does not have a consistently better performance in micro-object detection. Here, we propose an intelligent holographic classifier, which is a deep learning-based lensless inline digital holography system to detect the micro-object directly on the raw holograms and show the quantitative information of micro-objects for individual hologram by automatic object classification. In a demonstration where we capture the holograms of microplastics particles, which are easily confused with dust particles, we arrive at an accuracy above 97%. Compared with other leading classifiers, our method has shorter training time, faster classification and quantitative analysis, higher accuracy, and better robustness. Furthermore, this intelligent digital holography system, which requires only a light-emitting diode (LED), a sample slide, and a CMOS camera, can be used as a portable low-cost microplastics counting and classification tool, driving the development of microplastics detection in the ecological environment.
Yanmin Zhu 0005, Chok Hang Yeung, Edmund Y. Lam
INDIN3
2019 Spatial and Angular Reconstruction of Light Field Based on Deep Generative Networks
abstract
Light field (LF) cameras often have significant limitations in spatial and angular resolutions due to their design. Many techniques that attempt to reconstruct LF images at a higher resolution only consider either spatial or angular resolution, but not both. We propose a generative network using high-dimensional convolution to improve both aspects. Our experimental results on both synthetic and real-world data demonstrate that the proposed model outperforms existing state-of-the-art methods in terms of both peak signal-to-noise ratio (PSNR) and visual quality. The proposed method can also generate more realistic spatial details with better fidelity.
Nan Meng, Tianjiao Zeng, Edmund Y. Lam
ICIP3
2019 Applying Deep Learning and Wearable Devices for Educational Data Analytics
abstract
With the popularity of wearable devices, smart watches containing various sensors have been widely adopted for many healthcare applications. Yet there is rarely any research study on the possible uses of smart watches for learning analytics, particularly for analyzing students' learning activities through the physiological and/or movement data collected on their smart watches. This paper considers a pioneering and sophisticated learning analytics platform using fine-tuned deep learning models to predict students' learning activities based on the real-time data, including their heart rates, calories, three-axis accelerometer and gyroscope data, captured on wearable devices and then uploaded onto a cloud server for thorough analyses. To validate on the actual activities conducted by each student, an intelligent mobile application is developed to push instant notifications for students to report their own activities whenever the change of heart rates are deviated significantly from their normal values. Based on students' heart rates and calories, a long-short term memory (LSTM) model is built to classify students' learning states as active or not with an impressive prediction accuracy of 95% whereas another hybrid model combining both the LSTM and convolutional neural networks attains the highest prediction accuracy of 74% to predict students' specific learning activities as based on their physiological and movement data. The prototype implementation clearly demonstrates the feasibility of the proposed framework for learning analytics. More importantly, this work shed lights on various directions including the integration of noise filters to preprocess the collected data for further investigation.
Z. X. Zhou, Vincent W. L. Tam, King-Shan Lui, Edmund Y. Lam, Allan Hoi Kau Yuen, Xiao Hu 0001, Nancy Law
ICTAI4
2019 Fringe Pattern Improvement and Super-Resolution Using Deep Learning in Digital Holography
abstract
Digital holographic imaging is a powerful technique that can provide wavefront information of a three-dimensional object for biological and industrial applications. However, due to the constraint and cost of imaging sensors, the acquired digital hologram is limited in terms of pixel count, thus affecting the resolution in holographic reconstruction. To overcome this constraint, in this paper we propose a deep learning-based method to super-resolve holograms and to improve the quality of low-resolution holograms by training a convolutional neural network with large-scale data for resolution enhancement. Moreover, this algorithm can be broadly adapted to enhance the space-bandwidth product of a holographic imaging system without the need of any advanced hardware. We experimentally validate its capability using a lens-free off-axis holographic system, and compare the performance of various loss functions and interpolation methods in training such a network.
Zhenbo Ren, Hayden Kwok-Hay So, Edmund Y. Lam
IEEE Trans. Ind. Informatics3
2019 Large-Scale Multi-Class Image-Based Cell Classification With Deep Learning
abstract
Recent advances in ultra-high-throughput microscopy have enabled a new generation of cell classification methodologies using image-based cell phenotypes alone. In contrast to current single-cell analysis techniques that rely solely on slow and costly genetic/epigenetic analysis, these image-based analyses allow morphological profiling and screening of thousands or even millions of single cells at a fraction of the cost, and have been proven to demonstrate the statistical significance required for understanding the role of cell heterogeneity in diverse biological applications, ranging from cancer screening to drug candidate identification/validation processes. This paper examines the efficacies and opportunities presented by machine learning algorithms in processing large scale datasets with millions of label-free cell images. An automatic single-cell classification framework using convolutional neural network (CNN) has been developed. A comparative analysis of its efficiency in classifying large datasets against conventional k-nearest neighbors (kNN) and support vector machine (SVM) based methods are also presented. Experiments have shown that our proposed framework can efficiently identify multiple types cells with over 99% accuracy based on the phenotypic label-free bright-field images; and CNN-based models perform well and relatively stable against data volume compared with kNN and SVM.
Nan Meng, Edmund Y. Lam, Kevin K. Tsia, Hayden Kwok-Hay So
IEEE J. Biomed. Health Informatics2
2017 Human arm pose modeling with learned features using joint convolutional neural network
Chongguo Li, Nelson H. C. Yung, Xing Sun 0001, Edmund Y. Lam
Mach. Vis. Appl.4
2017 Computationally Efficient Hyperspectral Data Learning Based on the Doubly Stochastic Dirichlet Process
abstract
The Dirichlet process (DP) prior is effective in modeling HSIs (HSI) and identifying land-cover classes. However, modeling a continuously varying intensity of these land covers elegantly and consistently is still a challenge. We propose a doubly stochastic DP (DSDP) as an efficient model of the global topic measurement space, which imposes a weaker assumption compared with the discrete Markov assumption, resulting in a lower computational cost than other DP-prior-based models. We also present a mixture model of DSDP, which is termed the marked sigmoidal Gaussian process (SGP) DSDP mixture model. It can be thinned from a DP mixture without massive auxiliary covariates, and the marked function prior makes the number of land-cover classes consistent, whereas the SGP function prior models the HSI land-cover variation globally. The consistency of the number of land covers is maintained for various HSIs with large-scale geographical areas. Experiments show that the model is robust and consistent on HSI identification with weak or even no supervision.
Xing Sun 0001, Nelson H. C. Yung, Edmund Y. Lam, Hayden Kwok-Hay So
IEEE Trans. Geosci. Remote. Sens.3
2016 High dynamic range imaging via truncated nuclear norm minimization of low-rank matrix
abstract
We propose a ghost-free high dynamic range (HDR) image synthesis algorithm using a rank minimization framework. Based on the linear dependency among irradiance maps from low dynamic range (LDR) images, we formulate ghost-free HDR imaging as a low-rank matrix completion problem. The main contribution is to solve it efficiently via the augmented Lagrange multiplier (ALM) method, where the optimization variables are updated by closed-form solutions. Experiments on real image sets show that the proposed algorithm provides comparable or even better image qualities than state-of-the-art approaches, while demanding lower computational resources.
Chul Lee, Edmund Y. Lam
ICASSP2
2016 Sparse Hierarchical Nonparametric Bayesian learning for light field representation and denoising
abstract
In this paper, we present a sparse hierarchical non-parametric Bayesian (SHNB) model, which is used to represent the data captured by the light field cameras. Specifically, a light field can be represented as a set of sub-aperture views. In order to capture the visual variations of these viewpoints, we propose the so-called “depth flow” features. Then based on the depth flow features, we model these views statistically with a sparse representation in a fully unsupervised manner. While local dictionaries are learned based on each sub-aperture view, all the views with different perspectives share one global dictionary. To show the effectiveness of the proposed model, we apply our model to denoise the light field data. In the experiments, we demonstrate that our method outperforms several state-of-the-art light field denoising approaches.
Xing Sun 0001, Nan Meng, Edmund Y. Lam, Hayden Kwok-Hay So
IJCNN4
2016 Data-driven light field depth estimation using deep Convolutional Neural Networks
abstract
This paper presents a data-driven approach to estimate the object depths from light field data using Convolutional Neural Networks (CNN). By exploring the relationship between the epipolar-plane images (EPI) and the corresponding depth map, we propose an enhanced EPI feature that encodes the depth information of each physical point in the light field and obtains the disparity map of the whole scene in a supervised manner. This work covers two major contributions, namely the extraction of the enhanced EPI features and the light field depth estimation with CNN. The proposed features augment the depth information of the corresponding points in the light field, and then our CNN architecture differentiates them into different depth layers. Forward propagation step of the CNN model allows rapid recognition of the disparity map of the test light field data. In the experiments, we apply our method on the HCI (Heidelberg Col-laboratory for Image Processing) benchmark dataset and demonstrate that it is significantly faster than the state-of-the-art light field depth estimation approaches while achieving satisfactory performance.
Xing Sun 0001, Nan Meng, Edmund Y. Lam, Hayden Kwok-Hay So
IJCNN4
2016 Computationally Efficient Truncated Nuclear Norm Minimization for High Dynamic Range Imaging
abstract
Matrix completion is a rank minimization problem to recover a low-rank data matrix from a small subset of its entries. Since the matrix rank is nonconvex and discrete, many existing approaches approximate the matrix rank as the nuclear norm. However, the truncated nuclear norm is known to be a better approximation to the matrix rank than the nuclear norm, exploiting a priori target rank information about the problem in rank minimization. In this paper, we propose a computationally efficient truncated nuclear norm minimization algorithm for matrix completion, which we call TNNM-ALM. We reformulate the original optimization problem by introducing slack variables and considering noise in the observation. The central contribution of this paper is to solve it efficiently via the augmented Lagrange multiplier (ALM) method, where the optimization variables are updated by closed-form solutions. We apply the proposed TNNM-ALM algorithm to ghost-free high dynamic range imaging by exploiting the low-rank structure of irradiance maps from low dynamic range images. Experimental results on both synthetic and real visual data show that the proposed algorithm achieves significantly lower reconstruction errors and superior robustness against noise than the conventional approaches, while providing substantial improvement in speed, thereby applicable to a wide range of imaging applications.
Chul Lee, Edmund Y. Lam
IEEE Trans. Image Process.2
2016 Unsupervised Tracking With the Doubly Stochastic Dirichlet Process Mixture Model
abstract
We present an unsupervised tracking algorithm for human and car trajectory detection, using what is called the temporal doubly stochastic Dirichlet process (TDSDP) mixture model. The TDSDP captures the global dependence and the variation of human crowds and cars in temporal domains without the Markov assumption, making it particularly suitable for long-term tracking. Moreover, TDSDP prior can estimate the number of trajectories automatically. We first define the TDSDP based on the Cox process and then explain how to construct a TDSDP mixture model from thinning multiple Dirichlet process mixtures (DPMs) with conjugate priors. Next, a Markov chain Monte Carlo sampling inference is presented. Experimental results on synthetic and real-world data demonstrate that the proposed TDSDP mixture is superior to the DPM and the dependent Dirichlet process (DDP) in terms of topic variation modeling. PETS2001 data set experiments show that TDSDP has more robust object tracking capability over DDP based on generalized Polya urn. Low-quality fish data set experiments indicate that the TDSDP excels at solving tracking problems with insufficient features.
Xing Sun 0001, Nelson H. C. Yung, Edmund Y. Lam
IEEE Trans. Intell. Transp. Syst.3
2014 Maximum Likelihood Doppler Frequency Estimation Under Decorrelation Noise for Quantifying Flow in Optical Coherence Tomography
abstract
Recent hardware advances in optical coherence tomography (OCT) have led to ever higher A-scan rates. However, the estimation of blood flow axial velocities is limited by the presence and type of noise. Higher acquisition rates alone do not necessarily enable precise quantification of Doppler velocities, particularly if the estimator is suboptimal. In previous work, we have shown that the Kasai autocorrelation estimator is statistically suboptimal under conditions of additive white Gaussian noise. In addition, for practical OCT measurements of flow, decorrelation noise affects Doppler frequency estimation by broadening the signal spectrum. Here, we derive a general maximum likelihood estimator (MLE) for Doppler frequency estimation that takes into account additive white noise as well as signal decorrelation. We compare the decorrelation MLE with existing techniques using simulated and flow phantom data and find that it has better performance, achieving the Cramer-Rao lower bound. By making an approximation, we also provide an interpretation of this method in the Fourier domain. We anticipate that this estimator will be particularly suited for estimating blood flow in in vivo scenarios.
Aaron C. Chan, Vivek J. Srinivasan, Edmund Y. Lam
IEEE Trans. Medical Imaging3
2013 Applying an Evolutionary Approach for Learning Path Optimization in the Next-Generation E-Learning Systems
abstract
Learning analytics is targeted to better understand and optimize the process of learning and its environments through the measurement, collection and analysis of learners' data and contexts. To advise people's learning in a specific subject, most intelligent e-learning systems would require course instructors to explicitly input some prior knowledge about the subject such as all the pre-requisite requirements between course modules. Yet human experts may sometimes have conflicting views leading to less desirable learning outcomes. In a previous study, we proposed a complete system framework of learning analytics to perform an explicit semantic analysis on the course materials, followed by a heuristic-based concept clustering algorithm to group relevant concepts before finding their relationship measures, and lastly employing a simple yet efficient evolutionary approach to return the optimal learning sequence. In this paper, we carefully consider to enhance the original evolutionary optimizer with the hill-climbing heuristic, and also critically evaluate the impacts of various experts' recommended learning sequences possibly with conflicting views to optimize the learning paths for the next-generation e-learning systems. More importantly, the integration of heuristics can make our proposed framework more self-adaptive to less structured knowledge domains with conflicting views. To demonstrate the feasibility of our prototype, we implemented a prototype of the proposed e-learning system framework for learning analytics. Our empirical evaluation clearly revealed many possible advantages of our proposal with interesting directions for future investigation.
Vincent W. L. Tam, S. T. Fung, Alex Yi, Edmund Y. Lam
ICALT4
2013 Building an Interactive Simulator on a Cloud Computing Platform to Enhance Students' Understanding of Computer Systems
abstract
Cloud computing technologies have been widely adopted to improve the competitiveness and efficiency of core operations in many enterprises through additional computational resources and/or storage as provided on the underlying cloud platforms. Yet there are relatively few studies on how cloud computing may enhance students' understanding of a specific subject in e-learning systems. In a research project awarded by the Microsoft Research Asia, we successfully developed an interactive simulator aimed to enhance the students' understanding of essential concepts related to computer systems through live animations on a cloud computing platform. Essentially, we propose to integrate the latest technologies of cloud computing and learning objects into an efficient, flexible and interactive simulator to deliver powerful computing services for dynamic simulations of various computer systems specified as "reactive" models of learning objects on the cloud storage. More importantly, through adopting the IEEE learning object metadata standard to represent each key concept/component in different computer systems, our proposed simulator can readily facilitate the sharing and reuse of relevant concepts for future e-learning applications. The system design and prototype implementation of our cloud-based interactive simulator is carefully considered with a thorough evaluation plan to investigate on how learners may benefit from our interactive simulator in various ways. And there are many directions for future extensions.
Vincent W. L. Tam, Alex Yi, Edmund Y. Lam
ICALT3
2013 Using Cloud Computing and Mobile Devices to Facilitate Students' Learning through E-Learning Games
abstract
The recent advance in cloud computing and mobile devices empowers many innovative e-learning systems or games with increased interactivity and improved features. In this paper, we consider an innovative framework of cloud-based e-learning games that can be assessed through mobile devices to enhance students' learning anytime and anywhere. Being model-based, our proposal is adaptive and highly portable that can be easily customized to any existing cloud platform. Besides, our proposed framework allows course instructors or game designers to flexibly modify any part of an e-learning game, and continuously monitor the performance of individuals who try to compete with each other to attain better results. It is worth noting that this paper reports an on-going work, namely the iGame@Cloud system, for which a thorough evaluation will be conducted later. After all, our proposal stimulates many interesting directions for further exploration.
Vincent W. L. Tam, Alex Yi, Edmund Y. Lam, Cecilia Ka Yuk Chan, Allan Hoi Kau Yuen
ICALT3
2013 Fast single frame super-resolution using scale-invariant self-similarity
abstract
Example-based super-resolution (SR) attracts great interest due to its wide range of applications. However, these algorithms usually involve patch search in a large database or the input image, which is computationally intensive. In this paper, we propose a scale-invariant self-similarity (SiSS) based super-resolution method. Instead of searching patches, we select the patch according to the SiSS measurement, so that the computational complexity is significantly reduced. Multi-shaped and multi-sized patches are used to collect sufficient patches for high-resolution (HR) image reconstruction and a hybrid weighting method is used to suppress the artifacts. Experimental results show that the proposed algorithm is 20~1,800 times faster than several state-of-the-art approaches and can achieve comparable quality.
Luhong Liang, King Hung Chiu, Edmund Y. Lam
ISCAS3
2013 Comparison of Kasai Autocorrelation and Maximum Likelihood Estimators for Doppler Optical Coherence Tomography
abstract
In optical coherence tomography (OCT) and ultrasound, unbiased Doppler frequency estimators with low variance are desirable for blood velocity estimation. Hardware improvements in OCT mean that ever higher acquisition rates are possible, which should also, in principle, improve estimation performance. Paradoxically, however, the widely used Kasai autocorrelation estimator's performance worsens with increasing acquisition rate. We propose that parametric estimators based on accurate models of noise statistics can offer better performance. We derive a maximum likelihood estimator (MLE) based on a simple additive white Gaussian noise model, and show that it can outperform the Kasai autocorrelation estimator. In addition, we also derive the Cramer Rao lower bound (CRLB), and show that the variance of the MLE approaches the CRLB for moderate data lengths and noise levels. We note that the MLE performance improves with longer acquisition time, and remains constant or improves with higher acquisition rates. These qualities may make it a preferred technique as OCT imaging speed continues to improve. Finally, our work motivates the development of more general parametric estimators based on statistical models of decorrelation noise.
Aaron C. Chan, Edmund Y. Lam, Vivek J. Srinivasan
IEEE Trans. Medical Imaging2
2012 Toward a Complete E-learning System Framework for Semantic Analysis, Concept Clustering and Learning Path Optimization
abstract
Most online e-learning systems often demand the pre-requisite requirements between course modules and/or some relationship measures between involved concepts to be explicitly inputed by the course instructors so that an optimizer can be ultimately used to find an optimal learning sequence of involved concepts or modules for each individual learner after considering his/her past performance, learner's profile, learning style, etc. However, relying solely on the course instructor's input on the relationship among the involved concepts can be imprecise possibly due to the individual biases by human experts. Furthermore, the decision will become more complicated when various instructors hold conflicting views on the relationship among the involved concepts that may hinder any reasonable deduction. Therefore, we propose in this paper a complete system framework that can perform an explicit semantic analysis on the course materials, possibly aided by the relevant Wiki articles for any missing information about the involved concepts, to formulate the individual concepts, and followed by a heuristic-based concept clustering algorithm to group relevant concepts before finding their relationship measures. Lastly, an evolutionary optimizer will be used to return the optimal learning sequence after considering multiple experts' recommended learning sequences possibly containing conflicting views. To demonstrate the feasibility of our prototype, we implemented a prototype of the proposed e-learning system framework. Our empirical evaluation clearly revealed the possible advantages of our proposal with many possible directions for future investigation.
Vincent W. L. Tam, Edmund Y. Lam, S. T. Fung
ICALT2
2012 Nonlinear image reconstruction in block-based compressive imaging
abstract
A block-based compressive imaging (BCI) system with sequential architecture is presented in this paper. Feature measurements are collected using the principal component analysis (PCA) projection vectors. Then, we discuss an object prior learning framework based on the Field-of-Expert (FoE) model, and provide its implementation in the BCI reconstruction problem. Experimental results are used to demonstrate the reconstruction performance of the FoE-based method.
Jun Ke, Edmund Y. Lam
ISCAS2
2012 A Robust Computational Algorithm for Inverse Photomask Synthesis in Optical Projection Lithography
abstract
Inverse lithography technology formulates the photomask synthesis as an inverse mathematical problem. To solve this, we propose a variational functional and develop a robust computational algorithm, where the proposed functional takes into account the process variations and incorporates several regularization terms that can control the mask complexity. We establish the existence of the minimizer of the functional, and in order to optimize it effectively, we adopt an alternating minimization procedure with Chambolle's fast duality projection algorithm. Experimental results show that our proposed algorithm is effective in synthesizing high quality photomasks as compared with existing methods.
Siu-Kai Choy, Ningning Jia, Chong-Sze Tong, Man-Lai Tang, Edmund Y. Lam
SIAM J. Imaging Sci.5
2011 Enhancing Learning Paths with Concept Clustering and Rule-Based Optimization
abstract
Finding a good learning path with respect to existing reference paths of closely related concepts is very challenging yet important for effective course teaching and especially adaptive e-learning systems. There are various approaches including ontology analysis to extract the key concepts which could then be correlated to one another using an implicit or explicit knowledge structure for relevant courses. With the available correlation information, an effective optimizer can ultimately return a good learning path according to its predefined objective function. In this paper, we propose to obtain more thorough correlation information through concept clustering, which will then be passed to our rule-based genetic algorithm to search for better learning path(s). To demonstrate the feasibility of our proposal, a prototype of our ontology analyser enhanced with concept clustering and rule-based optimizer was implemented. Its performance was thoroughly studied and compared favorably against the benchmarking shortest-path optimizer on actual courses. More importantly, our proposal can be easily integrated into existing e-learning systems, and has significant impacts for adaptive or personalized e-learning systems through enhanced ontology analysis.
S. T. Fung, Vincent W. L. Tam, Edmund Y. Lam
ICALT3
2010 Stochastic gradient descent for robust inverse photomask synthesis in optical lithography
abstract
Optical lithography is a critical step in the semiconductor manufacturing process, and one key problem is the design of the photomask for a particular circuit pattern, given the optical aberrations and diffraction effects associated with the small feature size. Inverse lithography synthesizes an optimal mask by treating the design as an image synthesis inverse problem. To date, much effort is dedicated to solving it for some nominal process conditions. However, the small feature size also suggests that the effect of process variations is more pronounced. In this paper, we design a mask that is robust against focus variations within the inverse lithography framework. Each iteration involves more computation than a similar method designed for the nominal conditions, but we simplify the task by using stochastic gradient descent, which is a technique from machine learning. Simulation shows that the proposed algorithm is effective in producing robust masks.
Ningning Jia, Edmund Y. Lam
ICIP2
2010 Sectional image reconstruction in optical scanning holography using compressed sensing
abstract
Optical scanning holography is a form of digital holographic system, which allows us to capture a three-dimensional (3D) object in the two-dimensional (2D) hologram. A post-processing step known as sectional image reconstruction can then be applied to reconstruct a 2D plane of the object. In this paper, we show that the Fourier transform of the hologram corresponds to samples along a semi-spherical surface of the Fourier transform of the 3D object. Since the Fourier basis satisfies the uniform uncertainty principle, we can view the hologram capture as a form of compressed sensing, and reconstruction of the sectional images can then be accomplished by solving a total variation minimization problem.
Xin Zhang 0014, Edmund Y. Lam
ICIP2
2010 Edge detection of three-dimensional objects by manipulating pupil functions in an optical scanning holography system
abstract
The paper presents a technique of manipulating the pupil function in the optical scanning holography (OSH) system to detect the edge of a three-dimensional object. OSH is a fast biomedical imaging system that scans a three-dimensional (3D) object and encodes its intensity information into a two-dimensional hologram. The reconstruction of sectional images from a hologram has been well discussed as post-processing methods. However, manipulating pupil functions as a pre-processing technique has been less investigated. The latter can detect any specific part of an object by manipulating the pupil function. Here, we focus on the edge detection of a 3D object, which is implemented by configuring two pupil functions as a Dirac delta function and the Laplacian of the Gaussian, respectively. The simulated experiment demonstrates that the technique is a real-time pre-processing method to obtain the edge of the 3D object.
Xin Zhang 0014, Edmund Y. Lam
ICIP2
2010 Hyperspectral reconstruction in biomedical imaging using terahertz systems
abstract
Terahertz time-domain spectroscopy (THz-TDS) is an emerging modality for biomedical imaging. It is non-ionizing and can detect differences between water content and tissue density, but the detectors are rather expensive and the scan time tends to be long. Recently, it has been shown that the compressed sensing theory can lead to a radical re-design of the imaging system with lower detector cost and shorter scan time, in exchange for computation in the image reconstruction. We show in this paper that it is in fact possible to make use of the multi-frequency nature of the terahertz pulse to achieve hyperspectral reconstruction. Through effective use of the spatial sparsity, spectroscopic phase information, and correlations across the hyperspectral bands, our method can significantly improve the reconstructed image quality. This is demonstrated through using a set of experimental THz data captured in a single-pixel terahertz system.
Edmund Y. Lam
ISCAS2
2009 Developing an Innovative and Pen-Based Simulator to Enhance Education and Research in Computer Systems
abstract
The recent advance in pen-based computing and mobile devices empowers many innovative e-learning systems with increased interactivity and improved features. In this paper, we proposed an innovative and pen-based COMPAD simulator to enhance both education and research in computer systems. Being model-based, our proposed system is adaptive and different from many commercially available Windows based emulators that are customized for specific computer architectures. Besides, our COMPAD simulator allows learners to flexibly modify any part of a program through pen-based inputs, and instantly visualize the computed results. In this way, the COMPAD simulator can support not only education but also research such as instruction scheduling in computer systems. To demonstrate the feasibility, we built the COMPAD-PRO as its prototype for empirical evaluation. After all, this work stimulates many interesting directions for further exploration.
Johnny Yeung, Vincent W. L. Tam, Edmund Y. Lam, Cheung Hoi Leung
ICALT3
2009 An Improved Direction-of-arrival Estimation via Phase Information of Sparse Solution
abstract
An improved direction-of-arrivals (DOAs) estimation via phase information of sparse solution is presented in this paper. Unlike the conventional sparse source localization approach using the amplitude of sparse solutions only, through a special partition of the receiving data of the sensors, the phase information of the available sparse solutions is also extracted to estimate DOAs. For the true DOAs exactly on the grids which are used to generate the over-complete dictionary, the performance of our method is close to the conventional sparse source localization method. For the true DOAs that are not on the grids, our method is far superior to the conventional method, as demonstrated by several simulation results.
Xiansheng Guo, Qun Wan, Chunqi Chang, Edmund Y. Lam
ISCAS4
2008 Inverse image problem of designing phase shifting masks in optical lithography
abstract
The continual shrinkage of minimum feature size in integrated circuit (IC) fabrication incurs more and more serious distortion in the optical lithography process, generating circuit patterns deviating from the desired ones. Conventional resolution enhancement techniques (RETs) are facing critical challenges in compensating such increasingly severe distortion. The approach of inverse lithography, which is a branch of mask design methodology to treat the design as an inverse image problem, is adopted in this paper. We apply nonlinear optimization techniques to design masks with minimally distorted output. The output patterns so generated have high contrast and low dose sensitivity. We also propose a dynamic program-based initialization scheme to pre-assign phases to the layout.
Stanley H. Chan, Edmund Y. Lam
ICIP2
2008 Handling of multi-reflections in wafer bump 3D reconstruction
abstract
In advanced electronic manufacturing that involves say die-to-die bonding, microscopic surfaces like solder bumps on wafers have to be inspected in 3D. However, because the bumps are of hemispherical shape, light projected onto the bumps could be reflected and illuminate other regions. Such multi-reflections could greatly disturb the intensity distribution in the image data and limit the use of gray level intensities for accurate 3D reconstruction of the bumps. In a previous work, we described a new solution mechanism that was based upon the concept of binary pattern projection, but unlike the traditional mechanisms which use an array of light sources it uses only a single light source. The light source in combination with a binary fringe grating could induce binary pattern on the target surface to be imaged, and the displacements of the binary fringe grating could allow the binary pattern to be varied. In this work, we describe under that solution framework how multi-reflections could be detected and the correct binary signals could be restored. Experimental results on solder bumps validate the feasibility of the proposed approach.
Jun Cheng 0002, Ronald Chung, Edmund Y. Lam, Kenneth S. M. Fung
SMC3
2008 Simultaneous Ultrasound and MRI System for Breast Biopsy: Compatibility Assessment and Demonstration in a Dual Modality Phantom
abstract
Simultaneous capturing of ultrasound (US) and magnetic resonance (MR) images allows fusion of information obtained from both modalities. We propose an MR-compatible US system where MR images are acquired in a known orientation with respect to the US imaging plane and concurrent real-time imaging can be achieved. Compatibility of the two imaging devices is a major issue in the physical setup. Tests were performed to quantify the radio frequency (RF) noise introduced in MR and US images, with the US system used in conjunction with MRI scanner of different field strengths (0.5 T and 3 T). Furthermore, simultaneous imaging was performed on a dual modality breast phantom in the 0.5 T open bore and 3 T close bore MRI systems to aid needle-guided breast biopsy. Fiducial based passive tracking and electromagnetic based active tracking were used in 3 T and 0.5 T, respectively, to establish the location and orientation of the US probe inside the magnet bore. Our results indicate that simultaneous US and MR imaging are feasible with properly-designed shielding, resulting in negligible broadband noise and minimal periodic RF noise in both modalities. US can be used for real time display of the needle trajectory, while MRI can be used to confirm needle placement.
Annie M. Tang, Daniel F. Kacher, Edmund Y. Lam, Kelvin K. Wong, Ferenc A. Jolesz, Edward S. Yang
IEEE Trans. Medical Imaging3
2008 Effective Uses of FPGAs for Brute-Force Attack on RC4 Ciphers
abstract
This paper presents an effective field-programmable gate array (FPGA)-based hardware implementation of a parallel key searching system for the brute-force attack on RC4 encryption. The design employs several novel key scheduling techniques to minimize the total number of cycles for each key search and uses on-chip memories of the FPGA to maximize the number of key searching units per chip. Based on the design, a total of 176 RC4 key searching units can be implemented in a single Xilinx XC2VP20-5 FPGA chip, which currently costs only a few hundred U.S. dollars. Operating at a 47-MHz clock rate, the design can achieve a key searching speed of 1.07 x 107keys per second. Breaking a 40-bit RC4 encryption only requires around 28.5 h.
Sammy H. M. Kwok, Edmund Y. Lam
IEEE Trans. Very Large Scale Integr. Syst.2
2007 Maximizing Angle Coverage in Visual Sensor Networks
abstract
In this paper, we study the angle coverage problem in visual sensor networks where all sensors are equipped with cameras. An object of interest moves around the network and the sensors near the object are responsible for capturing images of it. The angle coverage problem aims to identify a set of sensors that preserve all the angles of view of the object while fulfilling the image resolution requirement. The user is required to specify the minimum acceptable image resolution in the request. Only the images that fulfill the resolution requirement will be considered. In order to save transmission energy, the number of images to be sent should be minimized. We develop a distributed algorithm to identify the minimum set of sensors such that all these images cover the maximum angle of view of the target. Our simulation results show that our protocol can achieve significant reduction in transmission load while preserving the widest angle of view.
Kit-Yee Chow, King-Shan Lui, Edmund Y. Lam
ICC3
2007 Efficient Selective Image Transmission in Visual Sensor Networks
abstract
Wireless sensor networks are under active research for tracking systems, where camera nodes are installed in a large area to take images of a targeted object. In this paper, we consider the scenario where the sensors around the object of interest capture images of it. Since some sensors are in similar viewing directions, the images they capture likely exhibit certain levels of correlation among themselves. It is a waste of transmission energy if we blindly send all images to the sink without checking for redundancy. We develop a protocol for involved sensors to determine how to select and transmit the images to the mobile sink in an energy efficient manner. The simulation results show that our protocol can achieve a significant reduction in energy consumption while preserving most of the viewing directions
Kit-Yee Chow, King-Shan Lui, Edmund Y. Lam
VTC Spring3
2007 Achieving 360° Angle Coverage with Minimum Transmission Cost in Visual Sensor Networks
abstract
In this paper, we study the angle coverage problem in visual sensor network. We consider a tracking system where an object of interest moves around the network, and the sensors surrounding it are responsible for capturing the images of it. We aim at finding the minimum cost cover which preserves all the angles of view with minimum transmission cost. We proved formally that the minimum cost cover problem can be transformed into the shortest path problem. Due to the acyclic nature of graphs generated in the transformation, we can develop a distributed algorithm to solve the problem, which is a lot more efficient than the distance vector protocol. Our simulation results show that our algorithm can successfully save a lot of energy.
Kit-Yee Chow, King-Shan Lui, Edmund Y. Lam
WCNC3
2006 Investigation of Computational Compound-Eye Imaging System With Super-Resolution Reconstruction
abstract
Despite the emerging architectural designs of compound-eye imaging systems, the post-processing algorithms for the reconstruction of the final image from the multiple sub-images is still not fully developed to maturity, resulting in poor quality or low resolution of the reconstructed images. In this paper, we describe and investigate a practical computational compound-eye imaging system with super-resolution reconstruction. This methodology can enhance the image quality by increasing the resolution of the reconstructed image. A virtual compound-eye camera is built to demonstrate the feasibility of the system. Simulation results which investigate the tolerance of the system to lens diversity, for instance focal length and aberrations, are also presented.
Wai-San Chan, Edmund Y. Lam, Michael Kwok-Po Ng
ICASSP (4)2
2006 Feature Selection in Source Camera Identification
abstract
Source camera identification is the process of discerning which camera has been used to capture a particular image. In our previous work, we tackled the problem with a vector of thirty-six features to train and test the classifier. The features include the lens aberration parameters and statistical measurements from pixel intensities. In this paper, we focus on reducing the feature set by stepwise discriminant analysis. Simulation is carried out to evaluate the classifier's performance by using the full feature set, reduced feature sets and randomly selected feature sets. The results show that the reduced feature sets can decrease the processing time while also maintain or even improve the classification accuracy under some circumstances.
Kai San Choi, Edmund Y. Lam, Kenneth K. Y. Wong
SMC2
2006 Curvature Domain Image Stitching
abstract
Digital photograph stitching blends multiple images to form a single one with a wide field of view. However, artifacts may arise, often due to photometric inconsistency and geometric misalignment among the images. Several existing techniques tackle this problem by methods such as pixel selection or pixel blending, which involve the matching and adjustment of intensity, frequency, and gradient values. However, our experience indicates that these methods have yet fully incorporated the uniformity properties of the photometric inconsistency. In this paper, we first explain the causes of inconsistency and its uniformity property. Then, by mathematical analysis, we show that the matching on the intensity and even the gradient domain is insufficient for some non-uniform inconsistencies. Our method thus adds the extra requirement of an optimal matching of curvature. We then explain how its variables can affect the computational and visual performance. Simulations are carried out using our method, with some masks designed with these two concerns. Some real examples show that our method can produce pleasant visual results even when both misalignment and non-uniform inconsistency exist.
Simon T. Y. Suen, Edmund Y. Lam, Kenneth K. Y. Wong
SMC2
2004 A nonlinear image restoration framework using vector quantization
abstract
Vector quantization (VQ) is a powerful method used primarily in signal and image compression. In recent years, it has also been applied to other various image processing tasks, including image classification, histogram modification, and restoration. In this paper, we focus our attention on image restoration using VQ. We present a general framework that incorporates two other methods in the literature, and discuss our method that follows more naturally from this framework. With appropriate training data for the VQ codebook, this method can restore images beyond its diffraction limit.
Edmund Y. Lam
ICIG1
2004 Analysis of the DCT coefficient distributions for document coding
abstract
It is known that the distribution of the discrete cosine transform (DCT) coefficients of most natural images follow a Laplacian distribution, and this knowledge has been employed to improve decoder design. However, such is not the case for text documents. In this letter, we present an analysis of their DCT coefficient distributions, and show that a Gaussian distribution can be a realistic model. Furthermore, we can use a generalized Gaussian model to incorporate the Laplacian distribution found for natural images.
Edmund Y. Lam
IEEE Signal Process. Lett.1
2003 Graylevel alignment between two images using linear programming
abstract
A critical step in defect detection for semiconductor process is to align a test image against a reference. This includes both spatial alignment and grayscale alignment. For the latter, a direct least square approach is not very applicable because the presence of defects would skew the parameters. Instead, we use a linear programming formulation which has the advantage of having a fast algorithm, while at the same time can produce better alignment of the test image to the reference. Furthermore, this is a flexible algorithm capable of incorporating additional constraints, such as ensuring that the aligned pixel values are within the allowable intensity range.
Edmund Y. Lam
ICIP (2)1
2000 A mathematical analysis of the DCT coefficient distributions for images
abstract
Over the past two decades, there have been various studies on the distributions of the DCT coefficients for images. However, they have concentrated only on fitting the empirical data from some standard pictures with a variety of well-known statistical distributions, and then comparing their goodness of fit. The Laplacian distribution is the dominant choice balancing simplicity of the model and fidelity to the empirical data. Yet, to the best of our knowledge, there has been no mathematical justification as to what gives rise to this distribution. We offer a rigorous mathematical analysis using a doubly stochastic model of the images, which not only provides the theoretical explanations necessary, but also leads to insights about various other observations from the literature. This model also allows us to investigate how certain changes in the image statistics could affect the DCT coefficient distributions.
Edmund Y. Lam, Joseph W. Goodman
IEEE Trans. Image Process.1