EDBT 2026 Demo / reviewers in the wild / expert
Ka-Hou Chan
dblp:166/2335
· DBLP profile ↗
22ranked-venue papers
12as first author
11since 2021 · last 2025
0000-0002-0183-0685ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 7 first-author · 9 since 2021Artificial intelligence and machine learning · 6 · 4 first-author · 3 since 2021Computer networks · 3 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MoEdit: On Learning Quantity Perception for Multi-object Image EditingabstractMulti-object images are prevalent in various real-world scenarios, including augmented reality, advertisement design, and medical imaging. Efficient and precise editing of these images is critical for these applications. With the advent of Stable Diffusion (SD), high-quality image generation and editing have entered a new era. However, existing methods often struggle to consider each object both individually and part of the whole image editing, both of which are crucial for ensuring consistent quantity perception, resulting in suboptimal perceptual performance. To address these challenges, we propose MoEdit, an auxiliaryfree multi-object image editing framework. MoEdit facilitates high-quality multi-object image editing in terms of style transfer, object reinvention, and background regeneration, while ensuring consistent quantity perception between inputs and outputs, even with a large number of objects. To achieve this, we introduce the Feature Compensation (FeCom) module, which ensures the distinction and separability of each object attribute by minimizing the in-between interlacing. Additionally, we present the Quantity Attention (QTTN) module, which perceives and preserves quantity consistency by effective control in editing, without relying on auxiliary tools. By leveraging the SD model, MoEdit enables customized preservation and modification of specific concepts in inputs with high quality. Experimental results demonstrate that our MoEdit achieves State-Of-The-Art (SOTA) performance in multi-object image editing. Data and codes are available at https://github.com/Tear-kitty/MoEdit. Ka-Hou Chan, Yue Sun 0001, Chan-Tong Lam, Tong Tong 0001, Zitong Yu, Keren Fu, Xiaohong Liu 0001, Tao Tan 0002 |
CVPR | 2 |
| 2025 | Contrastive Learning via Randomly Generated Deep SupervisionabstractUnsupervised visual representation learning has gained significant attention in the computer vision community, driven by recent advancements in contrastive learning. Most existing contrastive learning frameworks rely on instance discrimination as a pretext task, treating each instance as a distinct category. However, this often leads to intra-class collision in a large latent space, compromising the quality of learned representations. To address this issue, we propose a novel contrastive learning method that utilizes randomly generated supervision signals. Our framework incorporates two projection heads: one handles conventional classification tasks, while the other employs a random algorithm to generate fixed-length vectors representing different classes. The second head executes a supervised contrastive learning task based on these vectors, effectively clustering instances of the same class and increasing the separation between different classes. Our method, Contrastive Learning via Randomly Generated Supervision(CLRGS), significantly improves the quality of feature representations across various datasets and achieves state-of-the-art performance in contrastive learning tasks. Zili Ma, Ka-Hou Chan, Yue Liu 0001, Tong Tong 0001, Qinquan Gao, Guangtao Zhai, Xiaohong Liu 0001, Tao Tan 0002 |
ICASSP | 3 |
| 2025 | ADAptation: Reconstruction-Based Unsupervised Active Learning for Breast Ultrasound Diagnosis
Yaofei Duan, Yuhao Huang 0001, Xin Yang 0009, Luyi Han, Xinyu Xie, Ka-Hou Chan, Ligang Cui, Sio Kei Im, Dong Ni 0001, Tao Tan 0002 |
MICCAI (16) | 8 |
| 2025 | RefineNet: Elevating Medical Foundation Models Through Quality-Centric Data Curation by MLLM-Annotated Proxy Distillation
Ningyi Zhang, Xin Wang 0121, Ka-Hou Chan, Jian Wu 0033, Chan-Tong Lam, Shanshan Wang 0010, Yue Sun 0001, Sio Kei Im, Tao Tan 0002 |
MICCAI (11) | 4 |
| 2024 | Dynamic estimator selection for double-bit-range estimation in VVC CABAC entropy codingabstractAbstract CABAC is the only entropy coding used in Versatile Video Coding (VVC). This is achieved through multiple estimators approach that provide more accurate predictions by considering different estimated probability results, but CABAC coding requires higher complexity and bit‐range accuracy than other approaches. Therefore, there is more potential to refine the performance from the perspective of bit allocation and architecture design. In this paper, a selection method is proposed to determine which estimator is recommended to dynamically perform the current entropy coding. Taking advantage of the double‐bit‐range architecture, the bits contained in the different estimators are also rearranged based on Most Probable Symbol () determination and Least Probable Symbol () considerations. Experimental reports in the work report that the coding time can be reduced using the proposed method and there is a slight gain in Peak Signal‐to‐Noise Ratio (PSNR) while saving some Rate‐distortion (RD) performance in bitrate. Sio Kei Im, Ka-Hou Chan |
IET Image Process. | 2 |
| 2024 | Local feature-based video captioning with multiple classifier and CARU-attentionabstractAbstract Video captioning aims to identify multiple objects and their behaviours in a video event and generate captions for the current scene. This task aims to generate a detailed description of the current video in real‐time using natural language, which requires deep learning to analyze and determine the relationships between interesting objects in the frame sequence. In practice, existing methods typically involve detecting objects in the frame sequence and then generating captions based on features extracted through object coverage locations. Therefore, the results of caption generation are highly dependent on the performance of object detection and identification. This work proposes an advanced video captioning approach that works in adaptively and effectively addresses the interdependence between event proposals and captions. Additionally, an attention‐based multimodel framework is introduced to capture the main context from the frame and sound in the video scene. Also, an intermediate model is presented to collect the hidden states captured from the input sequence, which performs to extract the main features and implicitly produce multiple event proposals. For caption prediction, the proposed method employs the CARU layer with attention consideration as the primary RNN layer for decoding. Experimental results showed that the proposed work achieves improvements compared to the baseline method and also better performance compared to other state‐of‐the‐art models on the ActivityNet dataset, presenting competitive results in the tasks of video captioning. Sio Kei Im, Ka-Hou Chan |
IET Image Process. | 2 |
| 2023 | Using Four Hypothesis Probability Estimators for CABAC in Versatile Video CodingabstractThis article introduces the key technologies involved in four hypothetical probability estimators for Context-based Adaptive Binary Arithmetic Coding (CABAC). The focus is on the selected adaptation rate performed in these estimators, which are selected based on coding efficiency and memory considerations, and also the relationship with the current size of the coding block. The proposed scheme can linearly realize the quantitative representation of probabilistic prediction and describes the scalability potential for higher accuracy. Besides a description of the design concept, this work also discusses motivation and implementation aspects, which are based on simple operations such as bitwise operations and single subsampling for subinterval updates. The experimental results verify the effectiveness of the proposed CABAC method specified in Versatile Video Coding (VVC). Ka-Hou Chan, Sio Kei Im |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2022 | Double bit range estimation with eight estimators for CABAC in VVCabstractAbstract This work describes the modification of Context‐based Adaptive Binary Arithmetic Coding (CABAC) using the double bit range estimation in the VVC engine and the consideration of range updates by using eight hypothetical probability estimators. The focus is on the selected adaptation rates performed in these proposed estimators, which are chosen based on memory consideration and coding efficiency. An investigation of arithmetic coding engines with multi‐hypothesis probability estimates and their consideration of contextual modeling of entropy coding at the level of transform coefficients. The proposed scheme enables a quantitative representation of probabilistic predictions linearly and describes the scalability potential for higher accuracy. In addition, this work discusses the hardware implementation, which is based on simple operations such as bitwise operations and subinterval updates. The experimental results validate the effectiveness of the proposed approach specified in VTM framework. The improved results show that it provides more significant gains in terms of RA and LD, which is better than the AI configuration. Ka-Hou Chan, Sio Kei Im |
IET Image Process. | 1 |
| 2022 | Multiple classifier for concatenate-designed neural network
Ka-Hou Chan, Sio Kei Im, Wei Ke 0001 |
Neural Comput. Appl. | 1 |
| 2021 | A Self-Weighting Module to Improve Sentiment AnalysisabstractThis article introduces a self-weighting module for filtering meaningless words and normalizing them before RNN encoding, with the purpose of alleviating the long-term dependencies problem. We make use of the concept of weights in our design to analyze the transition of hidden states and indicate the complete architecture for processing the weighted feature and embedded word within the proposed module. In particular, we investigate the conditions that can enhance convergence and show that the proposed classifiers are able to improve the accuracy in the experimental cases significantly, not only giving better performance but also producing faster convergence. Moreover, the proposed module is general and can be applied to all RNN related network models. Ka-Hou Chan, Sio Kei Im |
IJCNN | 1 |
| 2021 | Discrete Tchebichef Transform for Versatile Video CodingabstractThe Discrete Tchebichef Transform (DTT) is a transform method based on discrete orthogonal Tchebichef polynomials, which have applications found in image compression and video coding. Our method is to construct all DTT-related discrete orthogonal transforms in the required size (corresponding to the coding unit supported by H.266/VVC). To investigate the feature of Tchebichef polynomials, we make use of a novel discrete orthogonal matrix generation method with determined DTT roots, and scaling and rounding a DTT that depends on the quantization parameter, instead of integer approximation. We can obtain an accurate integer DTT matrix. Experimental results show that this method can improve the video quality and require fewer bit rates. Ka-Hou Chan, Sio Kei Im |
ICMR | 1 |
| 2020 | Variable-Depth Convolutional Neural Network for Text Classification
Ka-Hou Chan, Sio Kei Im, Wei Ke 0001 |
ICONIP (5) | 1 |
| 2020 | CARU: A Content-Adaptive Recurrent Unit for the Transition of Hidden State in NLP
Ka-Hou Chan, Wei Ke 0001, Sio Kei Im |
ICONIP (1) | 1 |
| 2020 | Higher precision range estimation for context-based adaptive binary arithmetic codingabstractThe Lagrangian rate distortion optimisation is widely employed in modern video encoders, such as high‐efficiency video coding (H.265/HEVC). In this work, the authors propose a more accurate context‐based adaptive binary arithmetic coding look‐up table that can enhance compression quality and provide substantially better accuracy of range estimation, by employing one‐more bit with 64 probability states. For the hardware implementation, they propose a higher precision look‐up table instead of the HEVC Test Model (HM) standard table. The authors also define a new finite‐state machine to handle the probability changing in real‐time. The significant BD‐RATE gain of the proposed context modelling is up to 6.0% for all‐intra mode and 13.0% for inter mode. This finite state machine offers no divergence from the H.265/HEVC standards and can be used in the current systems. Sio Kei Im, Ka-Hou Chan |
IET Image Process. | 2 |
| 2019 | Image resizing enhancement with DCT coefficientsabstractAccording to the principle of scale transformation, signal expansion in the time domain corresponds to compression in the frequency domain, so that information energy is concentrated in the low frequency part. In this paper, an image enlargement algorithm based on the Discrete Cosine Transform (DCT) is proposed, which preserves the low frequency of the image and combines the corresponding enhancement coefficients to realize the resizing operation in the DCT domain. This paper also reach to proving of the enhancement value determinant. Then comparing the experiment on the scaled image with other interpolation algorithms, the result shows that our algorithm performs better than other methods. This method can also be carried out during DCT transformation, and is easier to implement than other methods. Ka-Hou Chan, Sio Kei Im, Wei Ke 0001 |
ICMV | 1 |
| 2018 | Particle-mesh coupling in the interaction of fluid and deformable bodies with screen space refraction renderingabstractAbstract On the basis of the smoothed particle hydrodynamics and finite element method (FEM) model, we propose a method integrating several improvements for the real‐time simulation of fluid interacting with deformable bodies. We improve the particle neighbor search in smoothed particle hydrodynamics, so that the predefined scene containers are no longer needed. This improvement can also be applied to the simulation of fluid interacting with other materials, such as rigid and soft bodies. We also propose a two‐way coupling method for fluid and deformable bodies, where the particle–mesh interaction is obtained by the ray‐traced collision detection method instead of the proxy/ghost particle generation. By using the forward ray‐tracing method for both velocity and position, we are able to calculate the coupling forces based on the conservation of momentum and kinetic energy in the particle–mesh interaction. We use the screen space fluid rendering for fluid, and on the basis of that, we introduce a screen space refraction rendering method to improve the refraction effect. We implement our method in NVIDIA CUDA and OptiX to make use of the full computational power of a graphics processing unit. The simulation results are analyzed and discussed to show the efficiency of our method. Ka-Hou Chan, Wei Ke 0001, Sio Kei Im |
Comput. Animat. Virtual Worlds | 1 |
| 2017 | Fast Binarisation with Chebyshev InequalityabstractIn order to enhance the binarization result of degraded document images with smudged and bleed-through background, we present a fast binarization technique that applies the Chebyshev theory in the image preprocessing. We introduce the Chebyshev filter which uses the Chebyshev inequality in the segmentation of objects and background. Our result shows that the Chebyshev filter is not only effective, but also simple, robust and easy to implement. Because of its simplicity, our method is sufficiently efficient to process live image sequences in real-time. We have implemented and compared with the Document Image Binarization Contest datasets (H-DIBCO 2014) for testing and evaluation. The experimental outcomes have demonstrated that this method achieved good result in this literature. Ka-Hou Chan, Sio Kei Im, Wei Ke 0001 |
DocEng | 1 |
| 2017 | Fast Grid-Based Fluid Dynamics Simulation with Conservation of Momentum and Kinetic Energy on GPU
Ka-Hou Chan, Sio Kei Im |
ICIG (3) | 1 |
| 2017 | Efficient mode decision with enhanced sampling algorithm for HEVCabstractCompared with H.264/AVC, H.265/HEVC achieves better image quality at the same coding bit rate, but requires more complexity in its coding process. In this paper, we propose a fast inter/intra decision method that can reduce the required encoding time with very little increase in required rate and PSNR degradation. We also introduce the weighted sampling and adaptive threshold condition to determine the CU splitting. In addition, a fast CU size selection algorithm based on image complexity for inter-frame coding is proposed. Experiments show that the proposed method provides a significant improvement in computing requirements, and can achieve a reasonable compromise between coding quality and efficiency. Sio Kei Im, Ka-Hou Chan |
WoWMoM | 2 |
| 2016 | Non-integer bit estimation for enhanced inter-picture prediction in H.265/HEVCabstractMany modern video encoders use the Lagrangian rate-distortion optimization (RDO) algorithm for mode decisions during the compression procedure. This paper proposes to increase the accuracy of inter picture prediction in H.265, by computing a non-integer number of bits for arithmetic coding of the syntax elements. During the RDO process the distortion and bit rate of each mode has to be estimated and a more accurate estimation leads to better-optimized decisions. Our method offers a significant gain in RDO performance, with simulations showing that an average bit-rate saving of up to 16.0% can be achieved for the H.265/HEVC codec. Sio Kei Im, Mohammad Mahdi Ghandi, Ka-Hou Chan |
WoWMoM | 3 |
| 2015 | Simulation of Interaction Between Fluid and Deformable Bodies
Ka-Hou Chan, Wei Ke 0001 |
ICIG (3) | 1 |
| 2013 | PredictionIO: a distributed machine learning server for practical software developmentabstractOne of the biggest challenges for software developers to build real-world predictive applications with machine learning is the steep learning curve of data processing frameworks, learning algorithms and scalable system infrastructure. We present PredictionIO, an open source machine learning server that comes with a step-by-step graphical user interface for developers to (i) evaluate, compare and deploy scalable learning algorithms, (ii) tune hyperparameters of algorithms manually or automatically and (iii) evaluate model training status. The system also comes with an Application Programming Interface (API) to communicate with software applications for data collection and prediction retrieval. The whole infrastructure of PredictionIO is horizontally scalable with a distributed computing component based on Hadoop. The demonstration shows a live example and workflows of building real-world predictive applications with the graphical user interface of PredictionIO, from data collection, algorithm tuning and selection, model training and re-training to real-time prediction querying. Simon Chan, Thomas Stone, Kit Pang Szeto, Ka-Hou Chan |
CIKM | 4 |