EDBT 2026 Demo / reviewers in the wild / expert
Sio Kei Im
dblp:18/3364 · also Sio-Kei Im
· DBLP profile ↗
60ranked-venue papers
7as first author
48since 2021 · last 2026
0000-0002-5599-4300ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 27 · 4 first-author · 22 since 2021Artificial intelligence and machine learning · 14 · 12 since 2021Human-computer interaction and ubiquitous computing · 9 · 8 since 2021Computer networks · 4 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 3 since 2021Security and privacy · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Generating pivot Gray codes for spanning trees of complete graphs in constant amortized timeabstractWe present the first known pivot Gray code for spanning trees of complete graphs, listing all spanning trees such that consecutive trees differ by pivoting a single edge around a vertex. This pivot Gray code thus addresses an open problem posed by Knuth in The Art of Computer Programming, Volume 4 (Exercise 101, Section 7.2.1.6, [Knuth 2011]), rated at a difficulty level of 46 out of 50, and imposes stricter conditions than existing revolving-door or edge-exchange Gray codes for spanning trees of complete graphs. Our recursive algorithm generates each spanning tree in constant amortized time using \(O(n^2)\) space. In addition, we provide a novel proof of Cayley’s formula, \(n^{n-2}\), for the number of spanning trees in a complete graph, derived from our recursive approach. We extend the algorithm to generate edge-exchange Gray codes for general graphs with \(n\) vertices, achieving \(O(n^2)\) time per tree using \(O(n^2)\) space. For specific graph classes, the algorithm can be optimized to generate edge-exchange Gray codes for spanning trees in constant amortized time per tree for complete bipartite graphs, \(O(n)\)-amortized time per tree for fan graphs, and \(O(n)\)-amortized time per tree for wheel graphs, all using \(O(n^2)\) space. Bowie Liu, Dennis Wong, Chan-Tong Lam, Sio Kei Im |
SODA | 4 |
| 2026 | Cluster our hairstyles: A novel deep differentiable clustering algorithm for generating three-dimensional representative hairstyles
Pinyan Li, Yapeng Wang 0001, Xu Yang 0010, Sio Kei Im, Jucheng Song, Jie Zhang 0090 |
Eng. Appl. Artif. Intell. | 4 |
| 2026 | Reasoning or not? A comprehensive evaluation of reasoning LLMs for dialogue summarization
Keyan Jin, Yapeng Wang 0001, Leonel Santos, Xu Yang 0010, Sio Kei Im, Hugo Gonçalo Oliveira |
Expert Syst. Appl. | 6 |
| 2026 | GPDPose: Self-supervised transformer with geometry, pose, and depth consistency for multi-view 3D human pose estimation
Jucheng Song, Jie Zhang 0090, Xu Yang 0010, Yapeng Wang 0001, Hao Gao 0005, Haolun Li 0001, Sio Kei Im |
Expert Syst. Appl. | 7 |
| 2026 | A Four-Paradigm Taxonomy and Systematic Survey of Blockchain-Enabled Intrusion Detection Systems for IoT and IIoTabstractTraditional Intrusion Detection Systems (IDS) are increasingly challenged by the distributed, heterogeneous, and rapidly evolving threat landscape in Internet of Things (IoT) and Industrial IoT (IIoT) environments. Blockchain has been explored as a promising foundation for decentralized and trustworthy security mechanisms; however, the existing literature remains fragmented and lacks a clear organizing lens for comparing design choices and evaluation practices. To address this, this paper presents a problem-driven survey of blockchain-enabled IDS for IoT and IIoT. We organize prior work into four integration paradigms, Trusted Rule, ML, DL, and FL—and relate each paradigm to the recurring design tensions it primarily targets. We further distill three fundamental tensions that frequently shape system design, including distributed architectures vs. centralized security management, collaborative information sharing vs. privacy preservation, and real-time detection requirements vs. resource-constrained devices. In addition, we summarize representative frameworks by consolidating datasets, threat models, and reported performance-related metrics, and we discuss common limitations that hinder cross-paper comparability. Finally, we outline a roadmap toward more standardized benchmarking, suggesting candidate evaluation criteria and blockchain-specific KPIs to encourage more transparent and comparable reporting. Overall, this survey aims to provide a structured lens for navigating the design space of blockchain-enabled IDS and to highlight open challenges for future research. Boxi Chen, Yapeng Wang 0001, Leonel Santos, Xu Yang 0010, Sio Kei Im |
IEEE Internet Things J. | 6 |
| 2026 | LogicMix: Sample mixing data augmentation for multi-label image classification with partial labels
Chak Fong Chong, Jielong Guo, Xu Yang 0010, Wei Ke 0001, Pedro H. Abreu, Yapeng Wang 0001, Sio Kei Im |
Pattern Recognit. | 7 |
| 2026 | SEM-UCSNet: A Novel Semantic Maps-Guided Compressive Sensing Framework for Underwater ImagesabstractUnderwater images (UWIs) captured by underwater detectors are essential for underwater detection and exploration. The compressive sensing theory (CS) provides a method for recovering images from few measurements, and it has been proven to be suitable for underwater environments with narrow bandwidth and limited communication channel resources, which may have a significant negative impact on the quality of captured UWIs. However, most existing state-of-art CS methods do not take the characteristics of UWIs into account, so their performance is limited in underwater applications. Compared with on-land images, UWIs have the following characteristics: 1) UWIs contain relatively few semantics, with a large amount of similar feature within the same semantics; 2) The importance of different semantics in UWIs is closely related to the underwater imaging model. In this paper, we combine the underwater imaging model and semantic of UWIs with CS task and propose a novel semantic maps-guided CS framework for UWIs, dubbed SEM-UCSNet, which can improve the performance of sampling and reconstruction, especially under extremely low sampling rate. In the sampling stage, a semantic importance analysis module combined with the imaging model is designed to guide the sampling. In the reconstruction process, a graph-based reconstruction strategy guided by semantic maps is proposed to model all features under the same semantic and mine complementarity between them to improve the reconstruction quality. Simultaneously, we introduce GAN into the underwater CS reconstruction task and use sampled features as conditions to make the reconstructed UWIs have richer details. Experimental results on some real-world UWIs datasets have demonstrated the superiority of our SEM-UCSNet on both objective and subjective metrics. Lihao Zhuang, Xiaochen Yuan, Wei Ke 0001, Chan-Tong Lam, Sio Kei Im |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | TSNN: A Non-Parametric and Interpretable Framework for Traffic Time Series ForecastingabstractAlthough many complex models were proposed to analyze time series data, some studies have demonstrated remarkable performance with simpler structures. A recent study proposed a non-parametric framework for 3D point cloud classification, which has the potential to be adapted for time series forecasting and enable interpretability. Inspired by the previous works, we present TSNN, a non-parametric and interpretable framework for traffic time series forecasting. TSNN consists of multiple layers that decouple the time series by matching the entries in a memory bank, where the memory bank is constructed using a similar matching process within the training set. It leverages the periodicity in traffic data to enhance forecasting accuracy while maintaining a simple model architecture. The proposed model operates without trainable parameters, preserving its inherent interpretability. In the experiments, TSNN achieves competitive performance compared to the typical deep learning models in four real-world traffic flow datasets. We also visualize the decoupling process to show the effectiveness of the components. Finally, we demonstrate the interpretability of the model and illustrate the contribution of each time step within the memory bank. Our code is available athttps://github.com/pzzzzzm/TSNN_release. Bowie Liu, Haijian Lai, Chan-Tong Lam, Junhao Dong 0004, Benjamin K. Ng, Wei Ke 0001, Sio Kei Im |
IEEE Trans. Knowl. Data Eng. | 7 |
| 2026 | Spatial-Temporal Relation Guided Motion Transfer via Diffusion ModelabstractTransferring existing Human-Object Interaction (HOI) motion to novel objects is essential for robotics, virtual reality. Traditional approaches only model spatial surface correspondences between humans and source objects or between source and target objects, ignoring the internal topological structures of humans, the internal topology of objects, the non-surface spatial topological relationships, and temporal motion relations. In this paper, we propose a spatial-temporal relation guided motion transfer framework. Firstly, we define a spatial-temporal relation interaction graph representation(STRIG) to model the human internal topology, object internal topology and human-object global topology together with the temporal motion relation. We propose a STRIGs-guided motion transfer diffusion model for generating spatially, semantically and temporally consistent HOI motions that are adapted to novel objects. To tackle the absence of ground-truth motions after transfer, we introduce a spatial-temporal relation optimization strategy. Extensive experiments demonstrate that our method consistently outperforms other approaches in terms of motion transfer quality, performance, and sequence stability, with particularly robustness under large variations in target object topology. Jian Wu 0033, Runze Fan, Sio Kei Im, Lili Wang 0006 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2026 | Interaction-Aware Shared Scene Synthesis for VR Telepresence
Zhangyao Tan, Qixiang Ma, Runze Fan, Sio Kei Im, Lili Wang 0006 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2026 | Gtfpose: a unified framework with double-chain GCN-transformer fusion for 3D human pose estimation
Junjia Zhang, Jucheng Song, Xu Yang 0010, Yapeng Wang 0001, Sio Kei Im |
Vis. Comput. | 5 |
| 2025 | Generating a Cyclic 2-Gray Code for Lucas Words in Constant Amortized Time
Bowie Liu, Dennis Wong, Chan-Tong Lam, Sio Kei Im |
CPM | 4 |
| 2025 | SSCM: Self-Supervised Critical Model for Reducing Hallucinations in Chinese Financial Text GenerationabstractLarge Language Models (LLMs) show strong performance in natural language processing tasks, but their application in the financial domain is limited. Current methods rely on large datasets and manual prompt engineering, resulting in high data demands, long inference times, and frequent hallucinations. To address these limitations, we propose a novel self-supervised prompt optimization framework tailored for the financial domain. Our approach involves training a critical model that evaluates and ranks generated outputs using both good and bad answers generated from various revised prompts. Experiments on a large Chinese financial corpus show that our framework significantly improves performance on tasks such as summarization and event-based question answering, as evidenced by higher scores on both automated metrics like ROUGE, BLEU, and BERTScore, and also through human evaluations. These results validate the effectiveness of our method in reducing hallucinations and improving the quality of financial text generation. Keyan Jin, Yapeng Wang 0001, Leonel Santos, Xu Yang 0010, Sio Kei Im |
ICASSP | 6 |
| 2025 | DABART: Dynamic Semantic Optimization Framework for Dialogue Summarization via Adaptive Topic Analysis and Semantic BridgingabstractWith the increasing prevalence of online communication and automated services, dialogue summarization technology plays a vital role in meeting minutes, customer service, and online Q&A scenarios. However, existing methods often suffer from insufficient flexibility in topic segmentation, low efficiency in semantic information transfer, and limited role adaptability. To address these challenges, we propose DABART, a dynamic semantic optimization framework. The framework employs a dynamic semantic topic segmentation mechanism to adaptively segment topics based on the distribution characteristics of sentence embeddings within dialogues, effectively identifying key information while overcoming the limitations of fixed-parameter methods in complex dialogue scenarios. Additionally, a dynamic semantic bridging module integrates semantic and positional information, further enhancing the coherence and consistency of dialogue summarization. Experimental results demonstrate that DABART achieves superior performance on widely-used benchmarks such as SAMSum and CSDS. Notably, it surpasses state-of-the-art open-source models on the CSDS dataset in ROUGE and BERTScore metrics, while achieving more balanced and accurate role-oriented summarization. Extensive experimental analyses further validate the robustness and applicability of the DABART across diverse dialogue scenarios. Keyan Jin, Yapeng Wang 0001, Leonel Santos, Xu Yang 0010, Sio Kei Im |
IJCNN | 6 |
| 2025 | ADAptation: Reconstruction-Based Unsupervised Active Learning for Breast Ultrasound Diagnosis
Yaofei Duan, Yuhao Huang 0001, Xin Yang 0009, Luyi Han, Xinyu Xie, Ka-Hou Chan, Ligang Cui, Sio Kei Im, Dong Ni 0001, Tao Tan 0002 |
MICCAI (16) | 10 |
| 2025 | RefineNet: Elevating Medical Foundation Models Through Quality-Centric Data Curation by MLLM-Annotated Proxy Distillation
Ningyi Zhang, Xin Wang 0121, Ka-Hou Chan, Jian Wu 0033, Chan-Tong Lam, Shanshan Wang 0010, Yue Sun 0001, Sio Kei Im, Tao Tan 0002 |
MICCAI (11) | 9 |
| 2025 | An artificial intelligence approach to automatically generate Cantonese meeting minutes for e-governmentabstractWhen artificial intelligence (AI) technology enters the world of acoustics, many projects that require audio processing are automated and no longer require much manual work. In particular, the use of the latest AI technologies to recognize speech and determine speakers makes it possible to generate meeting minutes entirely by machines. In the southeastern region of China (for example Hong Kong and Macao), people use Cantonese as the official language. Internal government meetings are usually conducted in Cantonese, and the executive department will request that the minutes be submitted as soon as possible after the meeting. In addition to the content of the speeches, the minutes must also include the identities of the corresponding speakers. During some periods of intensive meetings on new policy releases, our interpreters faced great pressure to script the exhausting meeting minutes. Due to the presence of local terms and personal names, even state-of-the-art large language models (LLMs) cannot fully suffice. Therefore, we propose a novel approach to solve such problems. This approach is a three-tier software architecture: the data tier (data processing), the service tier (AI models and web services), and the application tier (user interfaces). The implementation work is carried out using modern AI models (OpenAI’s Whisper and Nvidia’s TitaNet) and a dataset we created (Cantonese Policy Address, CPA). Training results (the word error rate is 33.81% and the equal error rate is 0.54) and validation results (confusion matrix up to 97%) show that our proposed approach improves automatic recognition precision, thus helping people understand the spirit of the meeting more effectively and quickly. Pinyan Li, Lap-Man Hoi, Yapeng Wang 0001, Sio Kei Im |
SMC | 4 |
| 2025 | A multi-modal speech emotion recognition method based on graph neural networks
Yapeng Wang 0001, Xu Yang 0010, Lap-Man Hoi, Sio Kei Im |
Appl. Intell. | 5 |
| 2025 | HiSum: Hierarchical Topic-Driven Approach for Role-Oriented Dialogue SummarisationabstractABSTRACT As the volume of information on online communication platforms continues to grow, the task of dialogue summarisation becomes increasingly critical for understanding and extracting key information from diverse conversations. Traditional approaches often struggle to cope with the dynamic nature of dialogues, such as managing perspectives from multiple speakers and seamlessly transitioning between different topics. We propose a novel hierarchical topic‐driven approach to generate role‐oriented dialogue summarisation (HiSum) to address these challenges. First, we utilise VarGMM clustering technology for in‐depth topic segmentation, which enables the model to capture the key topics in a dialogue. Second, we employ a LayerAttn hierarchical attention mechanism to dynamically adjust the focus of dialogue content based on participants' importance and the topics' relevance. Experimental results on three public dialogue summarisation data sets (CSDS, MC and SAMSUM) demonstrate that our method significantly outperforms most existing strong baseline methods across various evaluation metrics and surpasses the current state‐of‐the‐art methods in certain metrics. Detailed analysis demonstrates that HiSum can perform more precise topic segmentation and effectively identify critical information. Our code is publicly available at: https://github.com/kjin0119/HiSum . Keyan Jin, Yapeng Wang 0001, Xu Yang 0010, Sio Kei Im |
Expert Syst. J. Knowl. Eng. | 4 |
| 2025 | An algorithm for improving lower bounds in dynamic time warping
Yuqi Luo, Xinyi Fang, Wei Ke 0001, Chan-Tong Lam, Sio Kei Im, Luís Paquete |
Expert Syst. Appl. | 5 |
| 2025 | HandBrush for Efficient Object Grouping in Virtual Environment with Bare-HandabstractObject grouping task is an important research direction for fast manipulation of a large number of objects. It can help users to improve the efficiency of multi-object manipulation. However, the current research on this aspect is still immature. For this task, in this paper, based on the brush metaphor, we propose a method for grouping objects based on bare hands in virtual reality scenes. We design a number of interactions to facilitate the user's grouping of objects in the three-dimensional virtual space. Object grouping in virtual reality could encompass two subtasks: group generation and group modification. The emphasis of these tasks varies, with the former focusing on creating groups from ungrouped objects and the latter focusing on modifying group members once they are generated. The results of the empirical study show that our method has better performance in accomplishing both sub-tasks compared to the Ray method, Screen method and Cone method. Sichun Huang, Jian Wu 0033, Runze Fan, Sio Kei Im, Lili Wang 0006 |
Int. J. Hum. Comput. Interact. | 4 |
| 2025 | AVICol: Adaptive Visual Instruction for Remote Collaboration Using Mixed RealityabstractThis article describes a mixed reality visual instruction approach for remote collaboration between a trainee and an expert. The expert authors the visual instructions through a virtual reality interface. The instructions are shown to the trainee overlaid onto the workspace using an augmented reality interface. The approach achieves effectiveness and efficiency by addressing three challenges. First, the expert-authored visual instructions are shown to the trainee by taking into account occlusions with the 3D workspace; Second, in addition to abstract visual instructions implemented by arrows, the expert can also author highly suggestive instructions by depicting the target state of the workspace realistically by selecting, copying, pasting, and repositioning workspace objects; Third, multiple instructions can be concatenated in sequences that the trainee executes on their own, without any additional guidance from the expert; The approach has been evaluated in a controlled user study with three experiments. The experiment verification confirms that compared to the conventional instruction, this approach achieves significantly lower error rates, shorter task completion times, and lower rotation angular errors. Moreover, the approach allows the trainee to execute the entire sequence robustly, without real-time instruction from the expert. Lili Wang 0006, Jian Wu 0033, Sio Kei Im, Voicu Popescu |
Int. J. Hum. Comput. Interact. | 5 |
| 2025 | Manipulable cone based bare hand object selection in high occlusion virtual environment
Jian Wu 0033, Sio Kei Im, Runze Fan, Lili Wang 0006 |
Int. J. Hum. Comput. Stud. | 3 |
| 2025 | Category-wise Fine-Tuning: Resisting incorrect pseudo-labels in multi-label image classification with partial labels
Chak Fong Chong, Xinyi Fang, Jielong Guo, Pedro H. Abreu, Yapeng Wang 0001, Xu Yang 0010, Wei Ke 0001, Sio Kei Im |
Neurocomputing | 8 |
| 2025 | Point-FCW: Transposed-FCW Graph Representation for Point Cloud Classification Using TDAabstractDual challenges of computational efficiency and representation effectiveness exist in processing point clouds. Inspired by the TDA (Topological Data Analysis), we propose to convert the point cloud to a transposed fully connected and weighted (t-FCW) graph in order to significantly decrease the computational complexity in the following processing steps. We design a TDA pipeline called Point-FCW with a series of vectorization techniques for the 3D object point cloud feature extraction, which is plugged into the non-parametric classification head. Our experimental results demonstrate that Point-FCW achieves 75.28% accuracy on the ModelNet40 dataset with 512 points, providing a tiny, consistent, and effective representation for TDA. Furthermore, when integrated with the state-of-the-art non-parametric network Point-NN, the mixture model performs better, with an improvement of 4.47% in the OBJ-BG split of the ScanObjectNN dataset. Similarly, when integrating Point-FCW into the parametric network, PointMLP yields a performance improvement of 3.54% in the PB-T50-RS split of the ScanObjectNN dataset. The proposed Point-FCW can serve as a complementary enhancement feature when integrated into the Point-NN and PointMLP models. Moreover, the t-FCW graph representation can be efficiently converted at a rate of 3739 samples/second. Our code is available inhttps://github.com/hawkinglai/Point-FCW. Haijian Lai, Bowie Liu, Chan-Tong Lam, Benjamin K. Ng, Sio Kei Im |
IEEE Signal Process. Lett. | 5 |
| 2025 | TransHFC: Joints Hypergraph Filtering Convolution and Transformer Framework for TemporalForgery LocalizationabstractThe authenticity of audio-visual content is being challenged by advanced multimedia editing technologies inspired by Artificial Intelligence-Generated Content (AIGC). Temporal forgery localization aims to detect suspicious contents by locating forged segments. So far, most of the existing methods are based on Convolutional Neural Networks (CNNs) or Transformers, yet neither of them has fully considered the complex relationships within forged audio-visual content. To address this issue, in this paper, we propose a novel method, named TransHFC, which innovatively introduces hypergraphs to model group relationships among segments while considering point-to-point relationships through Transformers. Through its dual hypergraph filtering convolution branch, TransHFC captures both temporal and spatial level group relationships, enhancing the representation of forged segment features. Furthermore, we propose a new hypergraph filtering convolution Auto-Encoder that uses a multi-frequency filter bank for adaptive signal capture. This design compensates for the limitation of a single hypergraph filter. Our extensive experiments on Lav-DF, TVIL, Psynd, and HAD datasets demonstrate that TransHFC achieves state-of-the-art performance. Xiaochen Yuan, Chan-Tong Lam, Sio Kei Im, Fangyuan Lei, Xiuli Bi |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | PFL-ALP: Personalized Federated Learning Against Backdoor Attacks via Attention-Based Local PurificationabstractFederated learning (FL) enables collaborative model training with local data privacy preserving, but is vulnerable to backdoor attacks from malicious clients. These attacks can manipulate the global model to produce malicious output when encountering specific triggers. Existing defenses, categorized as server-side and client-side approaches, have limitations such as reliance on auxiliary data availability, susceptibility to inference attacks, and instability under non-independent and identically distributed (Non-IID) data. In response to these challenges, we propose a Personalized Federated Learning via Attention-based Local Purification (PFL-ALP) algorithm, a hybrid defense mechanism integrating server-side dynamic clustering and client-side purification enhanced with personalized model knowledge. This approach effectively mitigates bias introduced by Non-IID data on the server side and further purifies the backdoored model on the client side. Specifically, we employ neural attention distillation (NAD) for model purification and enhance it with personalized model knowledge, extending the effectiveness of NAD in Non-IID FL settings. This design makes PFL-ALP compatible with privacy protocols to mitigate inference attacks. Moreover, we establish a convergence guarantee for PFL-ALP and experimentally validate its superior performance in defending against various backdoor attacks compared to multiple state-of-the-art (SOTA) defenses across three datasets. The results show that even with malicious rates ranging from 30% to 90%, PFL-ALP can reduce the attack success rate by more than 69.4 percentage points, with the reduction in main task accuracy less than 12.4 percentage points. Yifeng Jiang 0008, Xiaochen Yuan, Wei Ke 0001, Chan-Tong Lam, Sio Kei Im |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2025 | A Symmetric Self-Embedding Mechanism for High-Fidelity Image Recovery Against TamperingabstractDigital images are inherently fragile and vulnerable to malicious tampering, significantly compromising their authenticity and integrity. Image recovery is crucial for restoring altered content and preserving the reliability of digital images. Traditional fragile watermarking methods achieve high-quality recovery but fail under post-processing attacks, while existing deep learning-based approaches offer some robustness, yet often produce lower-quality recovered images, typically with a PSNR of around 28 dB. To address these challenges, we propose a novel Symmetric Self-embedding Mechanism for High-Fidelity Image Recovery against tampering (SSEM-HIR), which is capable of restoring tampered images with high quality while maintaining some robustness against common attacks. Unlike existing methods that use the fragility of watermarking solely for tampering localization, SSEM-HIR is the first work to integrate fragility with spatial symmetry, enabling high-quality tampering recovery. Specifically, our SSEM-HIR employs a hierarchical watermark embedding module to embed an inverted version of the original image, utilizing spatial symmetry to retrieve lost information from the extracted watermark. To further improve recovery quality, we design a Dual-branch Region-based Self-Recovery module, where a Spatial-based Watermark Extraction block restores tampered regions using embedded watermark information, while a Frequency-assisted Image Repair block compensates for quality degradation in the untampered area. Extensive experiments show that our method achieves an average PSNR of 34.14 dB under common attack scenarios, including noise addition, image scaling, Gaussian blurring, and no post-processing. This represents an improvement of over 5 dB and 18% in recovered image quality compared to state-of-the-art approaches. Tong Liu 0021, Xiaochen Yuan, Wei Ke 0001, Chan-Tong Lam, Sio Kei Im, Pedro Martins 0003 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2025 | Proxy Importance Based Haptic Retargeting With Multiple Props in VRabstractIn virtual reality applications, in addition to visual feedback, real objects can be used as props for virtual objects to provide passive haptic feedback, which greatly enhances user immersion. Usually, real object props are not one-to-one correspondence with virtual objects. Haptic retargeting technique is proposed to establish the virtual-real correspondence by introducing an offset between the virtual hand and the real hand. Sometimes, the offset is too large to cause user discomfort, and it is necessary to introduce a reset between two haptic retargeting operations to force the virtual hand and the real hand to coincide in order to eliminate the offset. However, too many resets can interfere with this immersion. To address this problem, we propose a haptic retargeting method based on proxy importance calculation using multiple props in virtual reality. The concept of proxy importance for props is introduced first, and then a proxy importance based prop selection and placement method for moving virtual objects are proposed. We also improve the performance of our method by using the props' weighted proxy importance strategy for multi-user collaboration. Compared to the state-of-the-art methods, our method significantly reduces the number of resets, the task completion time, hand movement distances, and task load without the cost of cybersickness in the single-user task. In the multi-user collaborative task, our method also achieves significant improvement using the strategy that weights the proxy importance of the props. Jian Wu 0033, Lili Wang 0006, Sio Kei Im |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2025 | SGSG: Stroke-Guided Scene Graph Generationabstract3D scene graph generation is essential for spatial computing in Extended Reality (XR), providing structured semantics for task planning and intelligent perception. However, unlike instance-segmentation-driven setups, generating semantic scene graphs still suffer from limited accuracy due to coarse and noisy point cloud data typically acquired in practice, and from the lack of interactive strategies to incorporate users' spatialized and intuitive guidance. We identify three key challenges: designing controllable interaction forms, involving guidance in inference, and generalizing from local corrections. To address these, we propose SGSG, a Stroke-Guided Scene Graph generation method that enables users to interactively refine 3D semantic relationships and improve predictions in real time. We propose three types of strokes and a lightweight SGstrokes dataset tailored for this modality. Our model integrates stroke guidance representation and injection for spatio-temporal feature learning and reasoning correction, along with intervention losses that combine consistency-repulsive and geometry-sensitive constraints to enhance accuracy and generalization. Experiments and the user study show that SGSG outperforms state-of-the-art methods 3DSSG and SGFN in overall accuracy and precision, surpasses JointSSG in predicate-level metrics, and reduces task load across all control conditions, establishing SGSG as a new benchmark for interactive 3D scene graph generation and semantic understanding in XR. Implementation resources are available at: https://github.com/Sycamore-Ma/SGSG-runtime. Qixiang Ma, Runze Fan, Lizhi Zhao, Jian Wu 0033, Sio Kei Im, Lili Wang 0006 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2025 | Efficient and Comfortable Haptic Retargeting With Reset Point OptimizationabstractPassive haptics utilize the shape of a physical object to convey feedback to the user and enhance immersion in virtual reality. Haptic retargeting is a passive haptic interaction method. Its mapping of physical objects to virtual objects solves the matching problem between virtual and physical objects in the passive haptic method. However, most existing haptic retargeting methods improve efficiency without considering the important factor of user comfort. In this article, we propose an efficient and comfortable haptic retargeting method based on reset point optimization. First, we construct two maps indicating user interaction comfort: the RULA score map and the dominant hand gain map. Subsequently, we propose a reset point optimization algorithm based on these two maps. Moreover, we also optimize the selection of the physical proxy and the placement location when the reset occurs. The user study results show a significant improvement in the efficiency and comfort of our method compared to state-of-the-art methods. Aoxin Sun, Jian Wu 0033, Runze Fan, Sio Kei Im, Lili Wang 0006 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2025 | HFM-GS: Half-Face Mapping 3DGS Avatar Based Real-Time HMD RemovalabstractIn extended reality (XR) applications, enhancing user perception often necessitates head-mounted display (HMD) removal. However, existing methods suffer from low time performance and suboptimal reconstruction quality. In this paper, we propose a half face mapping 3D Gaussian splatting avatar based HMD removal method (HFM-GS), which can perform real-time and high-fidelity online restoration of the complete face in HMD-occluded videos for XR applications after a short un-occluded face registration. We establish a mapping field between the upper and lower face Gaussians to enhance the adaptability to deformation. Then, we introduce correlation weight-based sampling to improve time performance and handle variations in the number of Gaussians. At last, we ensure model robustness through Gaussian Segregation Strategy. Compared to two state-of-the-art methods, our method achieves better quality and time performance. The results of the user study show that fidelity is significantly improved with our method. Kangyu Wang, Jian Wu 0033, Runze Fan, Hongwen Zhang 0001, Sio Kei Im, Lili Wang 0006 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2025 | GaussianHand: Real-Time 3D Gaussian Rendering for Hand Avatar AnimationabstractRendering animatable and realistic hand avatars is pivotal for enhancing user experiences in human-centered AR/VR applications. While recent initiatives have utilized neural radiance fields to forge hand avatars with lifelike appearances, these methods are often hindered by high computational demands and the necessity for extensive training views. In this paper, we introduce GaussianHand, the first Gaussian-based real-time 3D rendering approach that enables efficient free-view and free-pose hand avatar animation from sparse view images. Our approach encompasses two key innovations. We first propose Hand Gaussian Blend Shapes that effectively models hand surface geometry while ensuring consistent appearance across various poses. Second, we introduce the Neural Residual Skeleton, equipped with Residual Skinning Weights, designed to rectify inaccuracies involved in Linear Blend Skinning deformations due to geometry offsets. Experiments demonstrate that our method not only achieves far more realistic rendering quality with as few as 5 or 20 training views, compared to the 139 views required by existing methods, but also excels in efficiency, achieving up to 125 frames per second for real-time rendering and remarkably surpassing recent methods. Lizhi Zhao, Xuequan Lu, Runze Fan, Sio Kei Im, Lili Wang 0006 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2024 | Dynamic estimator selection for double-bit-range estimation in VVC CABAC entropy codingabstractAbstract CABAC is the only entropy coding used in Versatile Video Coding (VVC). This is achieved through multiple estimators approach that provide more accurate predictions by considering different estimated probability results, but CABAC coding requires higher complexity and bit‐range accuracy than other approaches. Therefore, there is more potential to refine the performance from the perspective of bit allocation and architecture design. In this paper, a selection method is proposed to determine which estimator is recommended to dynamically perform the current entropy coding. Taking advantage of the double‐bit‐range architecture, the bits contained in the different estimators are also rearranged based on Most Probable Symbol () determination and Least Probable Symbol () considerations. Experimental reports in the work report that the coding time can be reduced using the proposed method and there is a slight gain in Peak Signal‐to‐Noise Ratio (PSNR) while saving some Rate‐distortion (RD) performance in bitrate. Sio Kei Im, Ka-Hou Chan |
IET Image Process. | 1 |
| 2024 | Local feature-based video captioning with multiple classifier and CARU-attentionabstractAbstract Video captioning aims to identify multiple objects and their behaviours in a video event and generate captions for the current scene. This task aims to generate a detailed description of the current video in real‐time using natural language, which requires deep learning to analyze and determine the relationships between interesting objects in the frame sequence. In practice, existing methods typically involve detecting objects in the frame sequence and then generating captions based on features extracted through object coverage locations. Therefore, the results of caption generation are highly dependent on the performance of object detection and identification. This work proposes an advanced video captioning approach that works in adaptively and effectively addresses the interdependence between event proposals and captions. Additionally, an attention‐based multimodel framework is introduced to capture the main context from the frame and sound in the video scene. Also, an intermediate model is presented to collect the hidden states captured from the input sequence, which performs to extract the main features and implicitly produce multiple event proposals. For caption prediction, the proposed method employs the CARU layer with attention consideration as the primary RNN layer for decoding. Experimental results showed that the proposed work achieves improvements compared to the baseline method and also better performance compared to other state‐of‐the‐art models on the ActivityNet dataset, presenting competitive results in the tasks of video captioning. Sio Kei Im, Ka-Hou Chan |
IET Image Process. | 1 |
| 2024 | Light-Occlusion Text Entry in Mixed RealityabstractText entry is a recurring task in mixed reality (MR) applications, and the ability of eyes-free text entry methods to allow users to enter text without focusing on the input device is ideal and compelling. However, existing eyes-free text entry methods leave much to be desired regarding efficiency and accuracy. In this paper, we propose a new light-occlusion text entry method in MR environment that uses dual thumb typing on a touchscreen. We design a partially visible keyboard as visual feedback to improve user performance. In addition, we optimize the underlying keyboard by collecting eyes-free typing data through a user study. The results show that our method has high typing speed, low error rate, and is very novice-friendly. After a short training period, the average typing speed of the novice group can reach 26.23 WPM (words per minute), while the average typing speed of the potential expert group can reach 30.62 WPM. Aoxin Sun, Lili Wang 0006, Jiaye Leng, Sio Kei Im |
Int. J. Hum. Comput. Interact. | 4 |
| 2024 | VVIR-OM: Efficient Object Manipulation in VR with Variable Virtual Interaction RegionabstractManipulating with virtual objects is a fundamental requirement in virtual environments. A good manipulation method needs to consider efficiency, accuracy, and comfort. This paper proposes VVIR-OM, an object manipulation method in virtual reality (VR) based on a variable virtual interaction region (VVIR). A hand interaction hemisphere region (HIHR) is introduced and constructed in real space, where the user is more comfortable manipulating the objects. Then a VVIR is introduced, and an interaction heat volume (IHV) based method is proposed to update VVIR during the process of the manipulation. At last, a mapping algorithm is proposed to map the user hand position in HIHR to the position in VVIR. Two user studies are designed to evaluate the performance of VVIR-OM. Compared to the state-of-the-art methods, VVIR-OM achieves significant improvements in task completion time, manipulation precision, and a significant reduction in fatigue. Moreover, VVIR-OM outperforms other methods in terms of task load and usability without the cost of cybersickness. Qinwen Zheng, Lili Wang 0006, Wei Ke 0001, Sio Kei Im |
Int. J. Hum. Comput. Interact. | 4 |
| 2024 | Object manipulation based on the head manipulation space in VR
Lili Wang 0006, Wei Ke 0001, Sio Kei Im |
Int. J. Hum. Comput. Stud. | 4 |
| 2024 | EEBA: Efficient and ergonomic Big-Arm for distant object manipulation in VR
Jian Wu 0033, Lili Wang 0006, Sio Kei Im, Chan-Tong Lam |
Int. J. Hum. Comput. Stud. | 3 |
| 2024 | An accurate slicing method for dynamic time warping algorithm and the segment-level early abandoning optimization
Yuqi Luo, Wei Ke 0001, Chan-Tong Lam, Sio Kei Im |
Knowl. Based Syst. | 4 |
| 2024 | MMQW: Multi-Modal Quantum Watermarking SchemeabstractTo address the problem that existing quantum image watermarking schemes have only a single watermarking mode with weak robustness, in this paper we propose a novel multi-modal quantum watermarking (MMQW) scheme using the generalized model of novel enhanced quantum representation. Our scheme provides four quantum watermarking modes (G_G, G_C, C_C, C_G), covering both types of grayscale and color images for the watermark and the carrier image. To enhance the robustness, we propose the Block Bit-plane Centrosymmetric Expansion (BBCE) method, which utilizes controlled quantum gates to extend the watermark, making our method resistant to noise and geometric attacks. Moreover, we propose a Brightness-based Watermarking Mechanism (BWM) for embedding and extraction. By uniform embedding, BWM not only minimizes the impact on the carrier image but also reduces the visual distortion of the extracted watermark. In the proposed MMQW, we implement three adaptive embedding strategies using controlled quantum gates, each of which is adaptively triggered according to the corresponding modalities. Detailed quantum circuits for quantum computing are provided. To evaluate imperceptibility and robustness of the MMQW, we conduct experiments using high-resolution images from the USC-SIPI dataset. The results show that PSNR of the watermarked image ranges from 36 dB to 56 dB, indicating the high visual quality. The PSNR of the extracted watermark is about 34 dB when the noise density is 0.05, while the PSNR is higher than 48 dB under common quantum rotation attacks, which indicate the high robustness against noise addition and geometric attacks. In addition, the proposed MMQW can resist to cropping attack with cropping percentage up to 55%. A comprehensive comparison with existing state-of-the-art works shows that our method has significant advantages. Chan-Tong Lam, Xiaochen Yuan, Sio Kei Im, Penousal Machado |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2024 | Real-scene-constrained virtual scene layout synthesis for mixed reality
Runze Fan, Lili Wang 0006, Xinda Liu, Sio Kei Im, Chan-Tong Lam |
Vis. Comput. | 4 |
| 2024 | SMigraPH: a perceptually retained method for passive haptics-based migration of MR indoor scenes
Qixiang Ma, Lili Wang 0006, Wei Ke 0001, Sio Kei Im |
Vis. Comput. | 4 |
| 2023 | Using Four Hypothesis Probability Estimators for CABAC in Versatile Video CodingabstractThis article introduces the key technologies involved in four hypothetical probability estimators for Context-based Adaptive Binary Arithmetic Coding (CABAC). The focus is on the selected adaptation rate performed in these estimators, which are selected based on coding efficiency and memory considerations, and also the relationship with the current size of the coding block. The proposed scheme can linearly realize the quantitative representation of probabilistic prediction and describes the scalability potential for higher accuracy. Besides a description of the design concept, this work also discusses motivation and implementation aspects, which are based on simple operations such as bitwise operations and single subsampling for subinterval updates. The experimental results verify the effectiveness of the proposed CABAC method specified in Versatile Video Coding (VVC). Ka-Hou Chan, Sio Kei Im |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2022 | Double bit range estimation with eight estimators for CABAC in VVCabstractAbstract This work describes the modification of Context‐based Adaptive Binary Arithmetic Coding (CABAC) using the double bit range estimation in the VVC engine and the consideration of range updates by using eight hypothetical probability estimators. The focus is on the selected adaptation rates performed in these proposed estimators, which are chosen based on memory consideration and coding efficiency. An investigation of arithmetic coding engines with multi‐hypothesis probability estimates and their consideration of contextual modeling of entropy coding at the level of transform coefficients. The proposed scheme enables a quantitative representation of probabilistic predictions linearly and describes the scalability potential for higher accuracy. In addition, this work discusses the hardware implementation, which is based on simple operations such as bitwise operations and subinterval updates. The experimental results validate the effectiveness of the proposed approach specified in VTM framework. The improved results show that it provides more significant gains in terms of RA and LD, which is better than the AI configuration. Ka-Hou Chan, Sio Kei Im |
IET Image Process. | 2 |
| 2022 | Multiple classifier for concatenate-designed neural network
Ka-Hou Chan, Sio Kei Im, Wei Ke 0001 |
Neural Comput. Appl. | 2 |
| 2021 | A Self-Weighting Module to Improve Sentiment AnalysisabstractThis article introduces a self-weighting module for filtering meaningless words and normalizing them before RNN encoding, with the purpose of alleviating the long-term dependencies problem. We make use of the concept of weights in our design to analyze the transition of hidden states and indicate the complete architecture for processing the weighted feature and embedded word within the proposed module. In particular, we investigate the conditions that can enhance convergence and show that the proposed classifiers are able to improve the accuracy in the experimental cases significantly, not only giving better performance but also producing faster convergence. Moreover, the proposed module is general and can be applied to all RNN related network models. Ka-Hou Chan, Sio Kei Im |
IJCNN | 2 |
| 2021 | Discrete Tchebichef Transform for Versatile Video CodingabstractThe Discrete Tchebichef Transform (DTT) is a transform method based on discrete orthogonal Tchebichef polynomials, which have applications found in image compression and video coding. Our method is to construct all DTT-related discrete orthogonal transforms in the required size (corresponding to the coding unit supported by H.266/VVC). To investigate the feature of Tchebichef polynomials, we make use of a novel discrete orthogonal matrix generation method with determined DTT roots, and scaling and rounding a DTT that depends on the quantization parameter, instead of integer approximation. We can obtain an accurate integer DTT matrix. Experimental results show that this method can improve the video quality and require fewer bit rates. Ka-Hou Chan, Sio Kei Im |
ICMR | 2 |
| 2020 | Variable-Depth Convolutional Neural Network for Text Classification
Ka-Hou Chan, Sio Kei Im, Wei Ke 0001 |
ICONIP (5) | 2 |
| 2020 | CARU: A Content-Adaptive Recurrent Unit for the Transition of Hidden State in NLP
Ka-Hou Chan, Wei Ke 0001, Sio Kei Im |
ICONIP (1) | 3 |
| 2020 | Higher precision range estimation for context-based adaptive binary arithmetic codingabstractThe Lagrangian rate distortion optimisation is widely employed in modern video encoders, such as high‐efficiency video coding (H.265/HEVC). In this work, the authors propose a more accurate context‐based adaptive binary arithmetic coding look‐up table that can enhance compression quality and provide substantially better accuracy of range estimation, by employing one‐more bit with 64 probability states. For the hardware implementation, they propose a higher precision look‐up table instead of the HEVC Test Model (HM) standard table. The authors also define a new finite‐state machine to handle the probability changing in real‐time. The significant BD‐RATE gain of the proposed context modelling is up to 6.0% for all‐intra mode and 13.0% for inter mode. This finite state machine offers no divergence from the H.265/HEVC standards and can be used in the current systems. Sio Kei Im, Ka-Hou Chan |
IET Image Process. | 1 |
| 2019 | Image resizing enhancement with DCT coefficientsabstractAccording to the principle of scale transformation, signal expansion in the time domain corresponds to compression in the frequency domain, so that information energy is concentrated in the low frequency part. In this paper, an image enlargement algorithm based on the Discrete Cosine Transform (DCT) is proposed, which preserves the low frequency of the image and combines the corresponding enhancement coefficients to realize the resizing operation in the DCT domain. This paper also reach to proving of the enhancement value determinant. Then comparing the experiment on the scaled image with other interpolation algorithms, the result shows that our algorithm performs better than other methods. This method can also be carried out during DCT transformation, and is easier to implement than other methods. Ka-Hou Chan, Sio Kei Im, Wei Ke 0001 |
ICMV | 2 |
| 2018 | Particle-mesh coupling in the interaction of fluid and deformable bodies with screen space refraction renderingabstractAbstract On the basis of the smoothed particle hydrodynamics and finite element method (FEM) model, we propose a method integrating several improvements for the real‐time simulation of fluid interacting with deformable bodies. We improve the particle neighbor search in smoothed particle hydrodynamics, so that the predefined scene containers are no longer needed. This improvement can also be applied to the simulation of fluid interacting with other materials, such as rigid and soft bodies. We also propose a two‐way coupling method for fluid and deformable bodies, where the particle–mesh interaction is obtained by the ray‐traced collision detection method instead of the proxy/ghost particle generation. By using the forward ray‐tracing method for both velocity and position, we are able to calculate the coupling forces based on the conservation of momentum and kinetic energy in the particle–mesh interaction. We use the screen space fluid rendering for fluid, and on the basis of that, we introduce a screen space refraction rendering method to improve the refraction effect. We implement our method in NVIDIA CUDA and OptiX to make use of the full computational power of a graphics processing unit. The simulation results are analyzed and discussed to show the efficiency of our method. Ka-Hou Chan, Wei Ke 0001, Sio Kei Im |
Comput. Animat. Virtual Worlds | 3 |
| 2017 | Fast Binarisation with Chebyshev InequalityabstractIn order to enhance the binarization result of degraded document images with smudged and bleed-through background, we present a fast binarization technique that applies the Chebyshev theory in the image preprocessing. We introduce the Chebyshev filter which uses the Chebyshev inequality in the segmentation of objects and background. Our result shows that the Chebyshev filter is not only effective, but also simple, robust and easy to implement. Because of its simplicity, our method is sufficiently efficient to process live image sequences in real-time. We have implemented and compared with the Document Image Binarization Contest datasets (H-DIBCO 2014) for testing and evaluation. The experimental outcomes have demonstrated that this method achieved good result in this literature. Ka-Hou Chan, Sio Kei Im, Wei Ke 0001 |
DocEng | 2 |
| 2017 | A teacher's view about introductory programming teaching and learning - Portuguese and Macanese perspectivesabstractThe difficulties faced by students and teachers in learning and teaching introductory programming has been a research issue over the years. Demotivation is common in many novice programming students, who are not able to cope with the natural difficulties associated to programming learning. It is up to the teacher to find strategies to help students and keep them motivated during the course. The objective of our research was to know more about the pedagogical and motivational strategies used by teachers in the author's institutions to promote programming student's motivation and learning. Some time ago we interviewed a few Portuguese teachers with diversified experiences in programming teaching. Recently we had the opportunity to do the same kind of research with Professors from Macao (China). The problems identified, the teachers' motivation to teach programming, the educational and motivational strategies and the student-teacher relationship were specifically addressed. This paper describes the research done, stressing the Macanese teacher's views and relating them with the views previously expressed by their Portuguese colleagues. Anabela Jesus Gomes, Wei Ke 0001, Sio Kei Im, Andrew Siu, António J. Mendes, Maria José Marcelino |
FIE | 3 |
| 2017 | Fast Grid-Based Fluid Dynamics Simulation with Conservation of Momentum and Kinetic Energy on GPU
Ka-Hou Chan, Sio Kei Im |
ICIG (3) | 2 |
| 2017 | Efficient mode decision with enhanced sampling algorithm for HEVCabstractCompared with H.264/AVC, H.265/HEVC achieves better image quality at the same coding bit rate, but requires more complexity in its coding process. In this paper, we propose a fast inter/intra decision method that can reduce the required encoding time with very little increase in required rate and PSNR degradation. We also introduce the weighted sampling and adaptive threshold condition to determine the CU splitting. In addition, a fast CU size selection algorithm based on image complexity for inter-frame coding is proposed. Experiments show that the proposed method provides a significant improvement in computing requirements, and can achieve a reasonable compromise between coding quality and efficiency. Sio Kei Im, Ka-Hou Chan |
WoWMoM | 1 |
| 2016 | Non-integer bit estimation for enhanced inter-picture prediction in H.265/HEVCabstractMany modern video encoders use the Lagrangian rate-distortion optimization (RDO) algorithm for mode decisions during the compression procedure. This paper proposes to increase the accuracy of inter picture prediction in H.265, by computing a non-integer number of bits for arithmetic coding of the syntax elements. During the RDO process the distortion and bit rate of each mode has to be estimated and a more accurate estimation leads to better-optimized decisions. Our method offers a significant gain in RDO performance, with simulations showing that an average bit-rate saving of up to 16.0% can be achieved for the H.265/HEVC codec. Sio Kei Im, Mohammad Mahdi Ghandi, Ka-Hou Chan |
WoWMoM | 1 |
| 2016 | Improved rate-distortion optimized video coding using non-integer bit estimation and multiple Lambda search
Sio Kei Im, Mohammad Mahdi Ghandi |
Frontiers Comput. Sci. | 1 |
| 2006 | An Optimized Mapping Algorithm for Classified Video Transmission with the H.264 Flexible Macroblock OrderingabstractA new mapping algorithm is proposed for categorization of frames' macroblocks into two (or more) classes in flexible macroblock ordering to give error resilient video transmission. The successful transmission of macroblock data not only enhances the quality of the associated pixels, but also improves the quality of the adjacent lost macroblocks by improving the efficiency of error concealment. Therefore, in our scheme, by carefully modeling the decoder error concealment algorithm at the encoder side, we classify the macroblocks according to their eventual influence on picture quality. Within a limited bit rate budget, we employ an optimization algorithm to select the best group of high-priority macroblocks. We show that prioritized transmission of the more important macroblock group will improve the video quality in error situations where our mapping algorithm outperforms the default mappings of the H.264 codec Sio Kei Im, Alan Pearmain |
ICME | 1 |