EDBT 2026 Demo / reviewers in the wild / expert
Chau-Wai Wong
dblp:24/10474
· DBLP profile ↗
26ranked-venue papers
8as first author
16since 2021 · last 2025
0000-0002-3873-7708ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 7 first-author · 3 since 2021Security and privacy · 7 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 6 · 6 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SCALE: Towards Collaborative Content Analysis in Social Science with Large Language Model Agents and Human InterventionabstractContent analysis breaks down complex and unstructured texts into theory-informed numerical categories.Particularly, in social science, this process usually relies on multiple rounds of manual annotation, domain expert discussion, and rule-based refinement.In this paper, we introduce SCALE, 1 a novel multi-agent framework that effectively Simulates Content Analysis via Large language model agEnts.SCALE imitates key phases of content analysis, including text coding, 2 collaborative discussion, and dynamic codebook evolution, capturing the reflective depth and adaptive discussions of human researchers.Furthermore, by integrating diverse modes of human intervention, SCALE is augmented with expert input to further enhance its performance.Extensive evaluations on real-world datasets demonstrate that SCALE achieves human-approximated performance across various complex content analysis tasks, offering an innovative potential for future social science research.* Example: "I love this company's new policy!It's so beneficial for everyone."* Example: "Great job on the recent project!Keep up the good work."-Neutral: Neutral sentiment of users toward the issue/company.* Example: "The company announced a new policy today."* Example: "I heard about the recent changes, but I don't have an opinion yet." Chengshuai Zhao, Zhen Tan 0001, Chau-Wai Wong, Tianlong Chen 0001, Huan Liu 0001 |
ACL (1) | 3 |
| 2025 | PortLLM: Personalizing Evolving Large Language Models with Training-Free and Portable Model PatchesabstractAs large language models (LLMs) increasingly shape the AI landscape, fine-tuning pretrained models has become more popular than in the pre-LLM era for achieving optimal performance in domain-specific tasks. However, pretrained LLMs such as ChatGPT are periodically evolved (i.e., model parameters are frequently updated), making it challenging for downstream users with limited resources to keep up with fine-tuning the newest LLMs for their domain application. Even though fine-tuning costs have nowadays been reduced thanks to the innovations of parameter-efficient fine-tuning such as LoRA, not all downstream users have adequate computing for frequent personalization. Moreover, access to fine-tuning datasets, particularly in sensitive domains such as healthcare, could be time-restrictive, making it crucial to retain the knowledge encoded in earlier fine-tuned rounds for future adaptation. In this paper, we present PORTLLM, a training-free framework that (i) creates an initial lightweight model update patch to capture domain-specific knowledge, and (ii) allows a subsequent seamless plugging for the continual personalization of evolved LLM at minimal cost. Our extensive experiments cover seven representative datasets, from easier question-answering tasks {BoolQ, SST2} to harder reasoning tasks {WinoGrande, GSM8K}, and models including {Mistral-7B,Llama2, Llama3.1, and Gemma2}, validating the portability of our designed model patches and showcasing the effectiveness of our proposed framework. For instance, PORTLLM achieves comparable performance to LoRA fine-tuning with reductions of up to 12.2× in GPU memory usage. Finally, we provide theoretical justifications to understand the portability of our model update patches, which offers new insights into the theoretical dimension of LLMs’ personalization. Rana Muhammad Shahroz, Pingzhi Li, Sukwon Yun, Shahriar Nirjon, Chau-Wai Wong, Tianlong Chen 0001 |
ICLR | 6 |
| 2025 | NTK-DFL: Enhancing Decentralized Federated Learning in Heterogeneous Settings via Neural Tangent KernelabstractDecentralized federated learning (DFL) is a collaborative machine learning framework for training a model across participants without a central server or raw data exchange. DFL faces challenges due to statistical heterogeneity, as participants often possess data of different distributions reflecting local environments and user behaviors. Recent work has shown that the neural tangent kernel (NTK) approach, when applied to federated learning in a centralized framework, can lead to improved performance. We propose an approach leveraging the NTK to train client models in the decentralized setting, while introducing a synergy between NTK-based evolution and model averaging. This synergy exploits inter-client model deviation and improves both accuracy and convergence in heterogeneous settings. Empirical results demonstrate that our approach consistently achieves higher accuracy than baselines in highly heterogeneous settings, where other approaches often underperform. Additionally, it reaches target performance in 4.6 times fewer communication rounds. We validate our approach across multiple datasets, network topologies, and heterogeneity settings to ensure robustness and generalization. Source code for NTK-DFL is available at https://github.com/Gabe-Thomp/ntk-dfl}{https://github.com/Gabe-Thomp/ntk-dfl Gabriel Thompson, Kai Yue, Chau-Wai Wong, Huaiyu Dai |
ICML | 3 |
| 2024 | Muharaf: Manuscripts of Handwritten Arabic Dataset for Cursive Text RecognitionabstractWe present the Manuscripts of Handwritten Arabic (Muharaf) dataset, which is a machine learning dataset consisting of more than 1,600 historic handwritten page images transcribed by experts in archival Arabic. Each document image is accompanied by spatial polygonal coordinates of its text lines as well as basic page elements. This dataset was compiled to advance the state of the art in handwritten text recognition (HTR), not only for Arabic manuscripts but also for cursive text in general. The Muharaf dataset includes diverse handwriting styles and a wide range of document types, including personal letters, diaries, notes, poems, church records, and legal correspondences. In this paper, we describe the data acquisition pipeline, notable dataset features, and statistics. We also provide a preliminary baseline result achieved by training convolutional neural networks using this data. Mehreen Saeed, Adrian Chan, Anupam Mijar, Joseph Moukarzel, Georges Habchi, Carlos Younes, Amin Elias, Chau-Wai Wong, Akram Khater |
NeurIPS | 8 |
| 2024 | Analysis of Coding Gain Due to In-Loop ReshapingabstractReshaping, a point operation that alters the characteristics of signals, has been shown capable of improving the compression ratio in video coding practices. Out-of-loop reshaping that directly modifies the input video signal was first adopted as the supplemental enhancement information (SEI) for the HEVC/H.265 without the need to alter the core design of the video codec. VVC/H.266 further improves the coding efficiency by adopting in-loop reshaping that modifies the residual signal being processed in the hybrid coding loop. In this paper, we theoretically analyze the rate-distortion performance of the in-loop reshaping and use experiments to verify the theoretical result. We prove that the in-loop reshaping can improve coding efficiency when the entropy coder adopted in the coding pipeline is suboptimal, which is in line with the practical scenarios that video codecs operate in. We derive the PSNR gain in a closed form and show that the theoretically predicted gain is consistent with that measured from experiments using standard testing video sequences. Chau-Wai Wong, Chang-Hong Fu 0002, Mengting Xu, Guan-Ming Su |
IEEE Trans. Image Process. | 1 |
| 2024 | Federated Learning via Plurality VoteabstractFederated learning allows collaborative clients to solve a machine-learning problem while preserving data privacy. Recent studies have tackled various challenges in federated learning, but the joint optimization of communication overhead, learning reliability, and deployment efficiency is still an open problem. To this end, we propose a new scheme named federated learning via plurality vote (FedVote). In each communication round of FedVote, clients transmit binary or ternary weights to the server with low communication overhead. The model parameters are aggregated via weighted voting to enhance the resilience against Byzantine attacks. When deployed for inference, the model with binary or ternary weights is resource-friendly to edge devices. Our results demonstrate that the proposed method can reduce quantization error and converges faster compared to the methods directly quantizing the model updates. Kai Yue, Richeng Jin, Chau-Wai Wong, Huaiyu Dai |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Gradient Obfuscation Gives a False Sense of Security in Federated Learning
Kai Yue, Richeng Jin, Chau-Wai Wong, Dror Baron, Huaiyu Dai |
USENIX Security Symposium | 3 |
| 2023 | Analysis of ENF Signal Extraction From Videos Acquired by Rolling ShuttersabstractElectric network frequency (ENF) analysis is a promising forensic technique for authenticating multimedia recordings and detecting tampering. The validity of the ENF analysis heavily relies on the capability of extracting high-quality ENF signals from multimedia recordings. This paper analyzes and compares two representative methods for extracting ENF signals from visual signals acquired by cameras using the rolling-shutter mechanism. The first method proposed in prior work,direct concatenation, ignores the idle period of each frame. The second method proposed in this paper,periodic zeroing-out, inserts zeros to missing sample points instead of ignoring the idle period. Our theoretical analyses of using multirate signal processing reveal and experiments confirm that while the first method can extract ENF signals without knowing the exact value of camera read-out time, there exists some mild distortion to extracted ENF signals. In contrast, the second method taking the read-out time as the additional input is capable of extracting distortion-free ENF signals, and its frequency component of the highest strength is always located at the nominal frequency. Additionally, we examine aliased DC and negative ENF components caused by the two methods and show that their impact on the accuracy of frequency estimation is minimum. This paper facilitates the fundamental understanding of extracting ENF signals from videos. The research findings imply that the periodic zeroing-out method offers more accurate frequency estimates, but the performance improvement is not significant. Jisoo Choi, Chau-Wai Wong, Hui Su, Min Wu 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2023 | Remote Blood Oxygen Estimation From Videos Using Neural NetworksabstractPeripheral blood oxygen saturation (SpO$_{2}$) is an essential indicator of respiratory functionality and received increasing attention during the COVID-19 pandemic. Clinical findings show that COVID-19 patients can have significantly low SpO$_{2}$before any obvious symptoms. Measuring an individual's SpO$_{2}$without having to come into contact with the person can lower the risk of cross contamination and blood circulation problems. The prevalence of smartphones has motivated researchers to investigate methods for monitoring SpO$_{2}$using smartphone cameras. Most prior schemes involving smartphones are contact-based: They require using a fingertip to cover the phone's camera and the nearby light source to capture reemitted light from the illuminated tissue. In this paper, we propose the first convolutional neural network based noncontact SpO$_{2}$estimation scheme using smartphone cameras. The scheme analyzes the videos of an individual's hand for physiological sensing, which is convenient and comfortable for users and can protect their privacy and allow for keeping face masks on. We design explainable neural network architectures inspired by the optophysiological models for SpO$_{2}$measurement and demonstrate the explainability by visualizing the weights for channel combination. Our proposed models outperform the state-of-the-art model that is designed for contact-based SpO$_{2}$measurement, showing the potential of the proposed method to contribute to public health. We also analyze the impact of skin type and the side of a hand on SpO$_{2}$estimation performance. Joshua Mathew, Xin Tian 0018, Chau-Wai Wong, Simon Ho, Donald K. Milton, Min Wu 0001 |
IEEE J. Biomed. Health Informatics | 3 |
| 2022 | Neural Tangent Kernel Empowered Federated LearningabstractFederated learning (FL) is a privacy-preserving paradigm where multiple participants jointly solve a machine learning problem without sharing raw data. Unlike traditional distributed learning, a unique characteristic of FL is statistical heterogeneity, namely, data distributions across participants are different from each other. Meanwhile, recent advances in the interpretation of neural networks have seen a wide use of neural tangent kernels (NTKs) for convergence analyses. In this paper, we propose a novel FL paradigm empowered by the NTK framework. The paradigm addresses the challenge of statistical heterogeneity by transmitting update data that are more expressive than those of the conventional FL paradigms. Specifically, sample-wise Jacobian matrices, rather than model weights/gradients, are uploaded by participants. The server then constructs an empirical kernel matrix to update a global model without explicitly performing gradient descent. We further develop a variant with improved communication efficiency and enhanced privacy. Numerical results show that the proposed paradigm can achieve the same accuracy while reducing the number of communication rounds by an order of magnitude compared to federated averaging. Kai Yue, Richeng Jin, Ryan Pilgrim, Chau-Wai Wong, Dror Baron, Huaiyu Dai |
ICML | 4 |
| 2022 | Invisible Geolocation Signature Extraction From a Single ImageabstractGeotagging images of interest are increasingly important to law enforcement, national security, and journalism. Today, many images do not carry location tags that are trustworthy and resilient to tampering; and landmark-based visual clues may not be readily present in every image, especially in those taken indoors. In this paper, we exploit an environmental signature from the power grid, the electric network frequency (ENF) signal, which can be inherently captured in a sensing stream at the time of recording and carries useful time–location information. Compared to the recent art of extracting ENF traces from audio and video recordings, it is very challenging to extract an ENF trace from a single image. We address this challenge by first mathematically examining the impact of the ENF embedding steps such as electricity to light conversion, scene geometry dilution of radiation, and image sensing. We then incorporate the verified parametric models of the physical embedding process into our proposed entropy minimization method. The optimized results of the entropy minimization are used for creating a two-level ENF presence–classification test for region-of-capturing localization. It identifies whether a single image has an ENF trace; if yes, whether it is at 50 or 60 Hz. We quantitatively study the relationship between the ENF strength and its detectability from a single image. This paper is the first comprehensive work to bring out a unique forensic capability of environmental traces that shed light on an image’s capturing location. Jisoo Choi, Chau-Wai Wong, Adi Hajj-Ahmad, Min Wu 0001, Yanpin Ren |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2021 | Contact Tracing Enhances the Efficiency of Covid-19 Group TestingabstractGroup testing can save testing resources in the context of the ongoing COVID-19 pandemic. In group testing, we are given n samples, one per individual, and arrange them into m < n pooled samples, where each pool is obtained by mixing a subset of the n individual samples. Infected individuals are then identified using a group testing algorithm. In this paper, we use side information (SI) collected from contact tracing (CT) within nonadaptive/single-stage group testing algorithms. We generate data by incorporating CT SI and characteristics of disease spread between individuals. These data are fed into two signal and measurement models for group testing, where numerical results show that our algorithms provide improved sensitivity and specificity. While Nikolopoulos et al. utilized family structure to improve nonadaptive group testing, ours is the first work to explore and demonstrate how CT SI can further improve group testing performance. Ritesh Goenka, Shu-Jie Cao, Chau-Wai Wong, Ajit Rajwade 0001, Dror Baron |
ICASSP | 3 |
| 2021 | Toward Effective Automated Content Analysis via CrowdsourcingabstractMany computer scientists use the aggregated answers of online workers to represent ground truth. Prior work has shown that aggregation methods such as majority voting are effective for measuring relatively objective features. For subjective features such as semantic connotation, online workers, known for optimizing their hourly earnings, tend to deteriorate in the quality of their responses as they work longer. In this paper, we aim to address this issue by proposing a quality-aware semantic data annotation system. We observe that with timely feedback on workers’ performance quantified by quality scores, better informed online workers can maintain the quality of their labeling throughout an extended period of time. We validate the effectiveness of the proposed annotation system through i) evaluating performance based on an expert-labeled dataset, and ii) demonstrating machine learning tasks that can lead to consistent learning behavior with 70%–80% accuracy. Our results suggest that with our system, researchers can collect high-quality answers of subjective semantic features at a large scale. Jiele Wu, Chau-Wai Wong, Xianpeng Liu |
ICME | 2 |
| 2021 | Learning Your Heart Actions From Pulse: ECG Waveform Reconstruction From PPGabstractThis article studies the relation between electrocardiogram (ECG) and photoplethysmogram (PPG) and investigates the inference of the ECG waveforms from the PPG signals that can be obtained from affordable wearable Internet-of-Things (IoT) devices for mobile health. In order to address this inverse problem, a transform is proposed to map the discrete cosine transform (DCT) coefficients of each PPG cycle to those of the corresponding ECG cycle based on the proposed cardiovascular signal model. The proposed method is evaluated with different morphologies of the PPG and ECG signals on three benchmark data sets with a variety of combinations of age, weight, and health conditions under several training setups. The experimental results show that the proposed method can achieve a high prediction accuracy greater than 0.92 in averaged correlation for each data set when the model is trained subjectwise. With a signal processing and learning system that is designed synergistically, we are able to reconstruct ECG signals by exploiting the relation of these two types of cardiovascular measurement. The reconstruction capability of the proposed method can enable low-cost ECG screening from affordable wearable IoT devices for continuous and long-term monitoring. This work opens up a new research direction to transfer the clinical ECG knowledge base to build a knowledge base for PPG and sensing data from wearable devices. Qiang Zhu 0015, Xin Tian 0018, Chau-Wai Wong, Min Wu 0001 |
IEEE Internet Things J. | 3 |
| 2021 | On Microstructure Estimation Using Flatbed Scanners for Paper Surface-Based AuthenticationabstractPaper surfaces under the microscopic view are observed to be formed by intertwisted wood fibers. Such structures of paper surfaces are unique from one location to another and are almost impossible to duplicate. Previous work used microscopic surface normals to characterize such intrinsic structures as a “fingerprint” of paper for security and forensic applications. In this work, we examine several key research questions of feature extraction in both scientific and engineering aspects to facilitate the deployment of paper surface-based authentication when flatbed scanners are used as the acquisition device. We analytically show that, under the unique optical setup of flatbed scanners, the specular reflection does not play a role in norm map estimation. We verify, using a larger dataset than prior work, that the scanner-acquired norm maps, although blurred, are consistent with those measured by confocal microscopes. We confirm that, when choosing an authentication feature, high spatial-frequency subbands of the heightmap are more powerful than the norm map. Finally, we show that it is possible to empirically calculate the physical dimensions of the paper patch needed to achieve a certain authentication performance in equal error rate (EER). We analytically show that log(EER) is decreasing linearly in the edge length of a paper patch. Runze Liu 0004, Chau-Wai Wong |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2021 | Adaptive Multi-Trace Carving for Robust Frequency Tracking in Forensic ApplicationsabstractIn the field of information forensics, many emerging problems involve a critical step that estimates and tracks weak frequency components in noisy signals. It is often challenging for the prior art of frequency tracking to i) achieve a high accuracy under noisy conditions, ii) detect and track multiple frequency components efficiently, or iii) strike a good trade-off of the processing delay versus the resilience and the accuracy of tracking. To address these issues, we propose Adaptive Multi-Trace Carving (AMTC), a unified approach for detecting and tracking one or more subtle frequency components under very low signal-to-noise ratio (SNR) conditions and in near real time. AMTC takes as input a time-frequency representation of the system's preprocessing results (such as the spectrogram), and identifies frequency components through iterative dynamic programming and adaptive trace compensation. The proposed algorithm considers relatively high energy traces sustaining over a certain duration as an indicator of the presence of frequency/oscillation components of interest and track their time-varying trend. Extensive experiments using both synthetic data and real-world forensic data of power signatures and physiological monitoring reveal that the proposed method outperforms representative prior art under low SNR conditions, and can be implemented in near real-time settings. The proposed AMTC algorithm can empower the development of new information forensic technologies that harness very small signals. Qiang Zhu 0015, Mingliang Chen 0001, Chau-Wai Wong, Min Wu 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2019 | ENF Signal Extraction for Rolling-shutter Videos Using Periodic Zero-paddingabstractElectric Network Frequency (ENF) analysis is a promising forensic technique for authenticating digital recordings and detecting tampering within the recordings. The validity of ENF analysis heavily relies on high-quality ENF signals extracted from multimedia recordings. In this paper, we propose an ENF signal extraction method for rolling shutter acquired videos using periodic zero-padding. Our analysis shows that the extracted ENF signals using the proposed method are not distorted and the component with the highest signal-to-noise ratio is located at the intrinsic frequency. The experimental results show that our proposed method can generate more precise ENF signals than those from the state-of-the-art method. Jisoo Choi, Chau-Wai Wong |
ICASSP | 2 |
| 2019 | Factors Affecting ENF Capture in AudioabstractThe electric network frequency (ENF) signal is an environmental signature that can be captured in audiovisual recordings made in locations where there is electrical activity. This signal is influenced by the power grid in which the recording is made, and recent work has shown that it can be useful toward a number of forensics and security applications. An under-studied area of ENF research is the factors that can affect the capture of ENF traces in media recordings. Not all recordings made in the areas of electrical activity will carry prominent ENF traces, and the strengths by which the ENF traces are captured can vary from one recording to another. A thorough understanding of the factors that can affect the capture of ENF traces in recordings is essential to understanding the applicability of ENF-based approaches and can help inform related studies in the future. This paper carried out a study on such factors, with a focus on audio signals. The impact of the characteristics of an audio recorder and the environment and manner of recording on the intrinsically captured ENF are shown and analyzed. Adi Hajj-Ahmad, Chau-Wai Wong, Steven Gambino, Qiang Zhu 0015, Miao Yu 0007, Min Wu 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2018 | Invisible Geo-Location Signature in A Single ImageabstractGeo-tagging images of interest is increasingly important to law enforcement, national security, and journalism. Many images today do not carry location tags that are trustworthy and resilient to tampering; and the landmark-based visual clues may not be readily present in every image, especially in those taken indoors. In this paper, we exploit an invisible signature from the power grid, the Electric Network Frequency (ENF) signal, which can be inherently recorded in a sensing stream at the time of capturing and carries useful location information. It is, however, very challenging to extract an ENF signal from a single image, as compared to the recent art in extracting ENF traces from audio and video. This paper presents novel investigations toward this challenge, by synergistically exploring the rolling shutter effect of CMOS imaging sensors and entropy differences of composite signals. We study quantitatively the relationship between the ENF strength and its detectability from a single image, and bring out a unique forensics capability of invisible traces that shine a light on an image's capturing location. Chau-Wai Wong, Adi Hajj-Ahmad, Min Wu 0001 |
ICASSP | 1 |
| 2017 | Fitness heart rate measurement using face videosabstractRecent studies showed that subtle changes in human's face color due to the heartbeat can be captured by digital video recorders. Most existing work focused on still/rest cases or those with relatively small motions. In this work, we propose a heart-rate monitoring method for fitness exercise videos. We focus on designing a highly precise motion compensation scheme with the help of the optical flow, and use motion information as a cue to adaptively remove ambiguous frequency components for improving the heart rates estimates. Experimental results show that our proposed method can achieve highly precise estimation with an average error of 1.1 beats per minute (BPM) or 0.58% in relative error. Qiang Zhu 0015, Chau-Wai Wong, Chang-Hong Fu 0002, Min Wu 0001 |
ICIP | 2 |
| 2017 | Counterfeit Detection Based on Unclonable Feature of Paper Using Mobile CameraabstractThis paper studies the authentication problem of specific pieces of paper using mobile imaging devices. Prior work showing high matching accuracy has used the normal vector field, which serves as a unique, microscopic, physically unclonable feature of paper surfaces, estimated by consumer grade scanners. Industrial cameras were also used to capture the appearance of the surface rendered after the normal vector field based on the laws of optics under a semi-controlled lighting condition. In comparison, past explorations based on mobile cameras were very limited and have not had substantial success in obtaining consistent appearance images due to the uncontrolled nature of the ambient light. We show in this paper that images captured by mobile cameras can be used for authentication when the camera flashlight is exploited for creating a semi-controlled lighting condition. We have proposed new algorithms to demonstrate that the normal vector field of paper surface can be estimated by using multiple camera-captured images of different viewpoints. Perturbation analysis shows that the proposed method is robust to inaccurate estimates of camera locations, and a matching accuracy of 10-4in equal error rate can be achieved using 6 to 8 images under a lab-controlled ambient light environment. Our findings can relax the restricted imaging setups and enable paper authentication under a more casual, ubiquitous setting with a mobile imaging device, which may facilitate duplicate detection of paper documents and counterfeit mitigation of merchandise packaging. Chau-Wai Wong, Min Wu 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2016 | Impact Analysis of Baseband Quantizer on Coding Efficiency for HDR VideoabstractDigitally acquired high dynamic range (HDR) video baseband signal can take 10-12 bits per color channel. It is economically important to be able to reuse the legacy 8 or 10-bit video codecs to efficiently compress the HDR video. Linear or nonlinear mapping on the intensity can be applied to the baseband signal to reduce the dynamic range before the signal is sent to the codec, and we refer to this range reduction step as a baseband quantization. We show analytically and verify using test sequences that the use of the baseband quantizer lowers the coding efficiency. Experiments show that as the baseband quantizer is strengthened by 1.6 bits, the drop of PSNR at a high bitrate is up to 1.60 dB. Our result suggests that in order to achieve high coding efficiency, information reduction of videos in terms of quantization error should be introduced in the video codec instead of on the baseband signal. Chau-Wai Wong, Guan-Ming Su, Min Wu 0001 |
IEEE Signal Process. Lett. | 1 |
| 2015 | A Study on PUF characteristics for counterfeit detectionabstractLow-cost physically unclonable functions (PUFs) can be deployed with consumer products to deter counterfeiting. An intrinsic physical property - unique textures of paper or other surface - has received strong interest. Extrinsically introduced features, such as randomly positioned bubbles and fiber segments, have also been deployed in the industry to facilitate authentication. This paper carries out a study to gain a better understanding in the factors affecting the authentication performance, with a consideration of the friendliness under mobile imaging. Comparisons are made for paper-based PUFs of different characteristics. It is found that the density of foreground objects have a dominant impact on the authentication performance. Chau-Wai Wong, Min Wu 0001 |
ICIP | 1 |
| 2011 | Transform Kernel Selection Strategy for the H.264/AVC and Future Video Coding StandardsabstractIn this paper, we propose a new discrete cosine transform (DCT)-like kernel IK(5, 7, 3) and revitalize another DCT-like kernel IK(13, 17, 7) for the transform coding process of hybrid video coding. Making use one of these kernels together with the H.264/AVC kernel IK(1, 2, 1), we are able to design new multiple-kernel schemes which give better coding performance over that of the conventional approaches. All these schemes make use of the adaptive kernel mechanism at macroblock-level (MB-AKM), which requires heavy computation during the encoding process. We subsequently discovered that a rate-distortion feature extracted from a pair of kernels gives an intrinsic property that can be used to select a better kernel for a two-kernel MB-AKM system. This is a powerful tool with theoretical interest and practical uses. In order to reduce computation substantially, we make use of this tool to make an analysis and design of a frame-level adaptive kernel mechanism and come up with a simple solution that the kernel IK(1, 2, 1) be used for I-frames and P-frames and the kernel IK(5, 7, 3) be used for B-frames coding. This proposed frame-based AKM gives similar, or even better, performance as the proposed macroblock-based AKM. Furthermore, it substantially reduces computation and certainly gives a good improvement in terms of the PSNR and bitrate compared to those obtained from the H.264/AVC default scheme and other MB-AKM schemes available in the literature. Chau-Wai Wong, Wan-Chi Siu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2011 | Comments on "2-D Order-16 Integer Transforms for HD Video Coding"abstractIn a recent paper, Dong proposed a set of order-16 nonorthogonal integer cosine transforms (NICTs). They proved that the reconstruction error caused by the nonorthogonality is negligible as compared to the error caused by the quantization. However, we would like to point out three problems found in derivations and also give two comments. Nevertheless, the problems are defects only, hence do not affect the overall justifications to the proposed NICT. This letter is to enhance and clarify the proof of Dong 's work. Chau-Wai Wong, Wan-Chi Siu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2011 | Analysis of Dyadic Approximation Error for Hybrid Video Codecs With Integer TransformsabstractIn this paper, we present an analysis of the dyadic approximation error introduced by the integerization of transform coding in H.264/AVC-like codecs. We derive the analytical formulations for dyadic approximation error and nonorthogonality error. We further classify the dyadic approximation error into a "system error" and a "nonflat error," and proposed two models for them. We found that the "nonflat error" has a substantial impact on video quality if the number of shifting bits at decoder side (DQ_BITS) is small. We also give a theoretical justification on why scaling factors at encoder side are better to be adapted to the rescaling factors at decoder side in H.264/AVC-like codecs. Chau-Wai Wong, Wan-Chi Siu |
IEEE Trans. Image Process. | 1 |