Wei Ke 0001

dblp:52/7566-1 · DBLP profile ↗
← Back
110ranked-venue papers
3as first author
77since 2021 · last 2027
0000-0003-0952-0961ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 38 · 28 since 2021Artificial intelligence and machine learning · 30 · 22 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 2 first-author · 7 since 2021Computer networks · 10 · 6 since 2021Databases, data management, data science and information retrieval · 8 · 3 since 2021Human-computer interaction and ubiquitous computing · 7 · 5 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 1 since 2021Security and privacy · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2027 MAGen: Multi-agent smart contract generation with automated testing and verification
Lixue Liu, Wei Ke 0001, Haiyang Chi, Junjian Yan
Empir. Softw. Eng.2
2026 Pyramid-Angular-Constraint Network for Light Field Super-Resolution
abstract
Light field (LF) cameras record both intensity and directions of light rays in a scene with a single exposure. Due to the trade-off between spatial and angular dimensions, the spatial resolution of LF images is limited, so super-resolution is widely studied. Pixels follow linear coordinate projection across views in LF images. Hence, auxiliary views nearer to the target view are generally more effective for use in super-resolution. In this paper, an LF-pyramid is proposed based on an angular-distance constraint for discriminatively exploiting auxiliary views. From views of different layers in an LF-pyramid, complementary features of different effectiveness can be extracted. However, shapes of LF-pyramids change for target views with different angular positions. To fully exploit an LF-pyramid, we introduce a pyramid-angular-constraint network for LF super-resolution (LF-PACNet). Specifically, to handle an arbitrary number of views in each layer, an intra-pyramid-layer feature extraction module is designed, which treats all views in the same layer equally in complementary information extraction. Then, to deal with an arbitrary number of layers, a recurrent cross-pyramid-layer feature complementation module is constructed, which discriminatively complements the target view with high-frequency details. Extensive experiments on public datasets demonstrate state-of-the-art performance for our method, both visually and numerically, especially for datasets with large disparities.
Da Yang 0001, Hao Sheng 0001, Wei Ke 0001, Zhang Xiong 0001
Comput. Vis. Media4
2026 Hard constraints and soft learning dual-graph anomaly detection for industrial processes
Ming-Qing Zhang, Wei Ke 0001, Yang Zhang 0032
Eng. Appl. Artif. Intell.4
2026 Knowledge graph augmented meta-learning with condition-sensitive pseudo-labeling for semi-supervised fault diagnosis under multiple working conditions
Ke-Yu Wu, Yuan Xu 0026, Wei Ke 0001, Yang Zhang 0032, Ming-Qing Zhang
Expert Syst. Appl.4
2026 User perception based label layout for efficient target localization in virtual environment
Jian Wu 0033, Shuai Luan, Wei Ke 0001, Lili Wang 0006
Int. J. Hum. Comput. Stud.4
2026 Adaptive semi-supervised meta-learning with pseudo-label for multiple working condition fault diagnosis
Ke-Yu Wu, Yuan Xu 0026, Wei Ke 0001, Yang Zhang 0032, Ming-Qing Zhang
Neurocomputing4
2026 The Latens Patronus: Seamless Model Watermarking for Latent Diffusion Model in IoT Environments
abstract
With the rapid development of the Internet of Things (IoT), generative artificial intelligence has been widely applied in various IoT applications. However, the wide adoption of latent diffusion model (LDM) in such IoT scenarios raises severe risks of copyright infringement and model theft due to the lack of effective protection mechanisms. To address this challenge, we propose Latens Patronus, a seamless model watermarking technique for copyright protection of LDM in IoT environments. Unlike existing model watermarking methods, our method does not require additional watermark input and additional parameters for a specialized embedding network, making it more suitable for deployment in real-world IoT applications. Specifically, we design a Watermark Encoder to integrate image watermark into latent features during the generation process and aWatermark Decoder to accordingly extract the watermark from suspicious images accurately. We further introduce a diminishing training strategy that gradually fades out auxiliary supervision signals, eliminating the need for persistent watermark guidance, and adding no extra overhead to the original model. Extensive experiments on multiple LDM variants demonstrate that Latens Patronus outperforms existing watermarking methods in both invisibility and robustness against image-level and model-level attacks.
Tong Liu 0021, Wei Ke 0001, Guoheng Huang, Xueyuan Gong, Xiaochen Yuan
IEEE Internet Things J.3
2026 Learning Three-Domain Implicit Image Function for Arbitrary-Scale Light Field Super-Resolution
abstract
Various deep learning-based light field image super-resolution methods have attained notable success in recent years. However, most of them focus on encoder design while neglecting the critical role of upsampling process in decoder part. Motivated by the recent progress in single image domain with implicit neural representation, we elaborately propose a spatial-angular-epipolar implicit image function (SAEIIF) in this paper, which can redefine the upsampling process to significantly improve performance and enable arbitrary-scale light field super-resolution. Specifically, it contains two complementary upsampling branches. One branch incorporates spatial implicit image function (SIIF) and angular implicit image function (AIIF) to mine intra-view information in sub-aperture images and inter-view information in macro pixels. The other branch involves epipolar implicit image function (EIIF) to leverage spatial-angular correlation in epipolar plane images. By decomposing SIIF, AIIF and EIIF into horizontal and vertical two-step upsampling to form a perfect match of upsampling scale, SAEIIF introduces a multi-stage feature interaction architecture across two branches to fully merge spatial, angular and epipolar domain information. Furthermore, we optimize feature sampling strategy based on characteristics of sub-aperture images, macro pixels, and epipolar plane images, introducing horizontal-vertical separable local sampling for SIIF and AIIF, as well as dual-source oriented line sampling used for EIIF. The extensive experimental results demonstrate that our SAEIIF can be effectively integrated with most encoders and achieve outstanding performance on both fixed-scale and arbitrary-scale light field spatial super-resolution, angular super-resolution, spatial-angular joint super-resolution.
Ruixuan Cong, Hao Sheng 0001, Da Yang 0001, Zhenglong Cui, Weifeng Lyv, Wei Ke 0001
IEEE Trans. Pattern Anal. Mach. Intell.8
2026 LogicMix: Sample mixing data augmentation for multi-label image classification with partial labels
Chak Fong Chong, Jielong Guo, Xu Yang 0010, Wei Ke 0001, Pedro H. Abreu, Yapeng Wang 0001, Sio Kei Im
Pattern Recognit.4
2026 Cross-Modal Attention Guided Enhanced Fusion Network for RGB-T Tracking
abstract
Visual tracking that combines RGB and thermal infrared modalities (RGB-T) aims to utilize the useful information of each modality to achieve more robust object localization. Most existing tracking methods based on convolutional neural networks (CNNs) and Transformers emphasize integrating multi-modal features through cross-modal attention, but ignore the potential exploitability of complementary information learned by cross-modal attention for enhancing modal features. In this paper, we propose a novel hierarchical progressive fusion network based on cross-modal attention guided enhancement for RGB-T tracking. Specifically, the complementary information generated by cross-modal attention implicitly reflects the consistent regions of interest of important information between different modalities, which is used to enhance modal features in a targeted manner. In addition, a modal feature refinement module and a fusion module are designed based on dynamic routing to perform noise suppression and adaptive integration on the enhanced multi-modal features. Extensive experiments on GTOT, RGBT234, LasHeR and VTUAV show that our method has competitive performance compared with recent state-of-the-art methods.
Jun Liu 0053, Wei Ke 0001, Shuai Wang 0027, Da Yang 0001, Hao Sheng 0001
IEEE Signal Process. Lett.2
2026 STFAENet: Multiattention Enhanced Spatiotemporal Fusion Network for Traffic Flow Prediction
abstract
Traffic flow prediction is a fundamental task in intelligent transportation systems (ITS). Due to the influence of urban functional zones and their neighboring regions, traffic flow data exhibit complex spatiotemporal correlations, making it challenging to effectively capture temporal dependencies and spatial structures for accurate prediction. To address this issue, this article proposes a multiattention enhanced spatiotemporal fusion network (STFAENet) based on the TransUNet architecture. STFAENet consists of an encoder, a skip-connection mechanism, and a decoder, and is designed to jointly learn fine-grained local features and global spatiotemporal dependencies in dynamic traffic scenarios. Specifically, an instance pyramid spatial attention (IPSA) module is introduced in the encoder to enhance multiscale spatial feature representation through hybrid normalization and pyramid attention, enabling the extraction of high-resolution fine-grained features. In the skip-connection stage, a spatiotemporal self-attention (STSA) module is embedded to jointly model spatial and temporal dependencies and reduce the semantic gap between the encoder and decoder. The decoder further fuses local details with global contextual information to generate high-resolution prediction results. Experiments on the TaxiBJ and TaxiCQ datasets demonstrate that STFAENet achieves superior prediction performance compared with several baseline methods.
Yuan Xu 0017, Chen-Yang Yan, Wei Ke 0001, Chongxing Ji, Yang Zhang 0032, Ming-Qing Zhang
IEEE Trans. Comput. Soc. Syst.5
2026 Robust RGB-T Tracking via Multi-Feature Response Adaptive Fusion and Dynamic Selection Recovery
abstract
In RGB and thermal (RGB-T) modalities fusion tracking, the multi-feature responses of each modality contain rich consistency in object localization, which is crucial to enhance tracking robustness. However, existing decision-level fusion paradigms mostly focus on fusing the output of the last layer, ignoring the correlation between multi-feature responses. Moreover, they also lack consideration of tracking failure, which hinders the application of RGB-T tracking in complex environments. To this end, this paper proposes a multi-feature response adaptive fusion model and a dominant-auxiliary dynamic selection recovery mechanism. Specifically, the former achieves joint optimal fusion by mining the correlation between multi-feature responses. The latter flexibly switches between short-term and long-term tracking modes according to the reliability of tracking results, and utilizes the most reliable modality to further improve tracking stability. Experiments on five prevalent RGB-T tracking benchmarks demonstrate the competitive performance of our method compared with the state-of-the-art methods.
Jun Liu 0053, Wei Ke 0001, Hao Sheng 0001
IEEE Trans. Circuits Syst. Video Technol.2
2026 3D-IRMM-Net: An Interpretable Radiologist-Mimicking 3D Multimodal Network for Lesion Detection in ABUS With LLM-Based Guidance and Uncertainty Quantification
abstract
Automated breast ultrasound (ABUS) has emerged as a promising tool for breast lesion detection, but most existing deep learning models for ABUS lack transparency and fail to align with clinical reasoning processes. This limits their interpretability and hinders clinical adoption. Therefore, we propose 3D-IRMM-Net, a radiologist-mimicking 3D multimodal network. It integrates visual, textual, and semantic cues to emulate hierarchical diagnostic reasoning. At its core, the network integrates two novel components: the 3D Scale-Expert Convolution (3D-SEC) block, which enables parameter-efficient, topology-aware feature routing via a shared graph convolutional layer; and the Adaptive Spatial-Gaussian (ASG) block, which enhances lesion localization in ABUS by modeling spatial dependencies through multi-scale Gaussian attention. In addition, we propose LLM-Vision activator, a natural-language-driven mechanism that uses language to guide 3D visual attention activation and enables 3D spatial-perceptual feature learning via a pretrained 2D VLM, boosting training efficiency. By aligning AI inference with clinical workflows, 3D-IRMM-Net enhances both detection rates and interpretability. Extensive experiments on ABUS2025, TDSC-ABUS2023, and FPHXD-ABUS datasets confirm its state-of-the-art performance across multiple metrics. At an FPPI of 0.5, our method achieved a patient-level detection rates of 89.7±1.1% (internal) and 85.1±2.0% (external); at an FPPI of 4, they rose to 96.1±1.6% and 91.5 ±1.9%, respectively. Furthermore, the model provides principled uncertainty quantification, improving trustworthiness in real-world deployment. These results confirm the effectiveness of 3D-IRMM-Net in bridging the gap between AI-driven detection and clinical practice.
Xin Qian 0001, Wei Ke 0001, Luoxi Zhu, Yue Sun 0001, Zhang Xiong 0001, Lingyun Bao, Tao Tan 0002
IEEE Trans. Circuits Syst. Video Technol.2
2026 SEM-UCSNet: A Novel Semantic Maps-Guided Compressive Sensing Framework for Underwater Images
abstract
Underwater images (UWIs) captured by underwater detectors are essential for underwater detection and exploration. The compressive sensing theory (CS) provides a method for recovering images from few measurements, and it has been proven to be suitable for underwater environments with narrow bandwidth and limited communication channel resources, which may have a significant negative impact on the quality of captured UWIs. However, most existing state-of-art CS methods do not take the characteristics of UWIs into account, so their performance is limited in underwater applications. Compared with on-land images, UWIs have the following characteristics: 1) UWIs contain relatively few semantics, with a large amount of similar feature within the same semantics; 2) The importance of different semantics in UWIs is closely related to the underwater imaging model. In this paper, we combine the underwater imaging model and semantic of UWIs with CS task and propose a novel semantic maps-guided CS framework for UWIs, dubbed SEM-UCSNet, which can improve the performance of sampling and reconstruction, especially under extremely low sampling rate. In the sampling stage, a semantic importance analysis module combined with the imaging model is designed to guide the sampling. In the reconstruction process, a graph-based reconstruction strategy guided by semantic maps is proposed to model all features under the same semantic and mine complementarity between them to improve the reconstruction quality. Simultaneously, we introduce GAN into the underwater CS reconstruction task and use sampled features as conditions to make the reconstructed UWIs have richer details. Experimental results on some real-world UWIs datasets have demonstrated the superiority of our SEM-UCSNet on both objective and subjective metrics.
Lihao Zhuang, Xiaochen Yuan, Wei Ke 0001, Chan-Tong Lam, Sio Kei Im
IEEE Trans. Circuits Syst. Video Technol.4
2026 Combining Gaussian and Pixel Representation for Light Field View Reconstruction
abstract
Light field (LF) benefits various applications due to its rich spatial and angular information. To address the technical limitation in terms of imaging resolution, LF view reconstruction becomes a research hotspot. However, relevant methods mainly focus on pixel representation modeling on image plane but ignore the importance of scene geometry modeling. Inspired by powerful geometry description ability embedded in 3D Gaussian Splatting, we construct a network called LFGaussian to perform generalizable LF view reconstruction in this paper. Specifically, owing to the unique composition of cross-view Gaussian attribute deviation under 4D LF imaging setting, we propose disparity-guided feed-forward 2D Gaussian propagation with novel Gaussian primitive definition, subtly implementing Gaussian unprojection-projection operation in camera parameter-free case. On this basis, we introduce a dual-branch workflow including Gaussian representation rendering and pixel representation upsampling to create features of target views from two different levels, which complement each other to jointly realize geometric structure consistency as well as texture detail consistency across all target views. Besides, for the pursuit of high-efficient and high-quality Gaussian representation rendering, we design sub-sampling Gaussian decoding to alleviate Gaussian redundancy and leverage Gaussian splitting to allocate additional Gaussians for complex geometry regions identified by disparity gradient. Experimental results show that the proposed LFGaussian achieves superior performance compared with state-of-the-art methods on both real-world and synthetic LF datasets, proving the effectiveness of introducing Gaussian representation for LF view reconstruction. Furthermore, our LFGaussian supports arbitrary-scale reconstruction, showing high flexibility for the upsampling scale factor.
Ruixuan Cong, Zhenglong Cui, Wei Ke 0001, Weifeng Lyv, Hao Sheng 0001
IEEE Trans. Image Process.4
2026 HRMamba: Fusing Luminance Information for Remote Physiological Measurement in Varied Lighting Conditions
abstract
Camera-based photoplethysmography (cbPPG) represents a non-invasive technique for capturing physiological parameters through facial videos, enabling the extraction of vital signs such as heart rate, respiration rate, and blood oxygen saturation without direct physical contact. Existing deep learning methods face two core challenges when dealing with cbPPG: firstly, extracting weak PPG signals from video segments with large spatial and temporal redundancy and understanding their periodic patterns in long contexts; secondly, accurately extracting PPG signals in complex lighting environments, especially in low-light conditions. To address these issues, this paper proposes an end-to-end method based on Mamba, named HRMamba. This method employs temporal difference mamba to process temporal signals and combines bidirectional state space to enable Mamba to robustly understand the scene and learn the periodic patterns of PPG. Furthermore, a luminance post-processing module is designed to extract luminance information from the video without enhancing lighting or altering the original video data, and embed it into the PPG signal. Experimental results demonstrate that HRMamba achieves state-of-the-art performance, and the designed luminance post-processing module can be applied in various lighting environments, significantly enhancing the performance in dark environments without degrading the performance in normal light scenes.
Nuoer Long, Wei Ke 0001, Chan-Tong Lam, Tao Tan 0002, Zitong Yu, Yue Sun 0001
IEEE J. Biomed. Health Informatics3
2026 TSNN: A Non-Parametric and Interpretable Framework for Traffic Time Series Forecasting
abstract
Although many complex models were proposed to analyze time series data, some studies have demonstrated remarkable performance with simpler structures. A recent study proposed a non-parametric framework for 3D point cloud classification, which has the potential to be adapted for time series forecasting and enable interpretability. Inspired by the previous works, we present TSNN, a non-parametric and interpretable framework for traffic time series forecasting. TSNN consists of multiple layers that decouple the time series by matching the entries in a memory bank, where the memory bank is constructed using a similar matching process within the training set. It leverages the periodicity in traffic data to enhance forecasting accuracy while maintaining a simple model architecture. The proposed model operates without trainable parameters, preserving its inherent interpretability. In the experiments, TSNN achieves competitive performance compared to the typical deep learning models in four real-world traffic flow datasets. We also visualize the decoupling process to show the effectiveness of the components. Finally, we demonstrate the interpretability of the model and illustrate the contribution of each time step within the memory bank. Our code is available athttps://github.com/pzzzzzm/TSNN_release.
Bowie Liu, Haijian Lai, Chan-Tong Lam, Junhao Dong 0004, Benjamin K. Ng, Wei Ke 0001, Sio Kei Im
IEEE Trans. Knowl. Data Eng.6
2026 Multi-Granularity Query Network With Adaptive Category Feature Embedding for Behavior Recognition
abstract
Behavior recognition is a highly challenging task, particularly in scenarios requiring unified recognition across both human and animal subjects. Most existing approaches primarily focus on single-species datasets or rely heavily on prior information such as species labels, positional annotations, or skeletal keypoints, which limits their applicability in real-world scenarios where species labels may be ambiguous or annotations are insufficient. To address these limitations, we propose a query-based Multi-Granularity Behavior Recognition Network that directly mines cross-species shared spatiotemporal behavior patterns from raw video inputs. Specifically, we design a Multi-Granularity Query module to effectively fuse fine-grained and coarse-grained features, thereby enhancing the model's capability in capturing spatiotemporal dynamics at different granularities. Additionally, we introduce a Category Query Decoder that leverages learnable category query vectors to achieve explicit behavior category modeling and mapping. Without relying on any extra annotations, the proposed method achieves unified recognition of multi-species and multi-category behaviors, setting a new state-of-the-art on the Animal Kingdom dataset and demonstrating strong generalization ability on the Charades dataset.
Nuoer Long, Yonghao Dang, Chengpeng Xiong, Shaobin Chen, Tao Tan 0002, Wei Ke 0001, Chan-Tong Lam, Jianqin Yin, Peter H. N. de With, Yue Sun 0001
IEEE Trans. Multim.7
2026 UCGR: Closing the Discretization Gap in Light Field Depth Estimation via Unified Continuous Geometry Representation
abstract
Light field (LF) cameras encode dense spatial-angular information for depth estimation, critical for applications such as 3D reconstruction, refocusing, and virtual reality. However, current deep learning methods for LF depth estimation still face significant challenges due to the discretization gap between the continuous geometry of real scenes and the discrete sampling of digital images, limiting their effectiveness in high-precision application. This gap manifests in two complementary forms: spatial discretization leads to structural ambiguities and information loss, while depth discretization introduces inaccuracies due to fixed, discrete depth sampling. To address these challenges, we propose Unified Continuous Geometry Representation (UCGR), a unified representation that models scene geometry as a continuous field over image coordinates and depth. UCGR treats spatial and depth discretization as two facets of the same problem and realized by two complementary operators: (1) Adaptive Plane Sampling Operator, which learns edge-aware planar priors to preserve geometric details and mitigate spatial discretization. (2) Contextual Depth Correction Operator, which utilizes contextual information for depth correction, ensuring continuous depth estimation and suppressing artifacts. Building on UCGR, we propose a Continuous Geometry Network (CGNet) that collaboratively optimizes both spatial and depth discretization for accurate and consistent LF depth estimation. Extensive experiments on synthetic and real-world LF datasets demonstrate that CGNet achieves state-of-the-art performance, significantly outperforming existing LF depth estimation methods in terms of both accuracy and robustness.
Zexin Sun, Tun Wang, Rongshan Chen, Ruixuan Cong, Wei Ke 0001, Hao Sheng 0001
IEEE Trans. Vis. Comput. Graph.5
2026 Multi-scale feature fusion for cross-modality person re-identification: the MSJLNet approach
Zhixin Tie, Haobiao Fan, Lingbing Tao, Yanbing Chen, Hao Sheng 0001, Wei Ke 0001
Vis. Comput.6
2025 Spatiotemporal Uncertainty-Aware Mamba-Transformer Synergy: Breast Cancer Detection in ABUS
abstract
Breast cancer significantly affects women's health, making early and accurate diagnosis through ultrasound examination essential. However, automated breast ultrasound(ABUS) lesion detection models encounter challenges, including uncertain noise and difficulties in locating small lesions. This paper proposes SUA-MT, a multi-video object detection network using uncertainty-aware transformers and temporal-spatial mamba. To address the performance degradation caused by low-quality frames and noise interference in input data, we propose the uncertainty-gated transformer decoder (UGTD), which dynamically adjusts attention weights to focus on high-confidence regions while suppressing attention to redundant areas. The spatialtemporal(ST) mamba module is designed to model the long-term dependencies and 3D spatial features of different plane video frames, making full use of temporal and spatial information. In general, SUA-MT not only supports dynamic video length input, but also combines multidimensional video modeling of lesions (transverse, sagittal and coronal plane), making full use of the complementarity of multi-view information and spatiotemporal information to enhance the ability to locate small lesions. Experimental results on a combined set of internal and publicly datasets demonstrate that the SUA-MT method achieves state-of-the-art performance compared to existing video detection approaches. Specifically, SUA-MT attains a mean precision of 80.1 % and mean recall of 84.2 %, providing an efficient and robust solution for lesion detection in ABUS.
Xin Qian 0001, Luoxi Zhu, Zhang Xiong 0001, Wei Ke 0001, Lingyun Bao, Tao Tan 0002
BIBM4
2025 GSHOI Denoiser: Denoising Gaussian Hand-Object Interaction for Photorealistic Rendering
abstract
Many VR/AR applications require the photorealistic rendering of hand-object interactions. Virtual hands are driven by users' hand poses captured via motion tracking to interact with virtual objects. The driven pose can be very noisy due to the constraints of tracking hardware and computation accuracy. This noise may lead to distorted hand poses and penetration artifacts during rendering. In this paper, we introduce the Gaussian Hand-Object Interaction Denoiser, the Gaussian splatting-based hand-object interaction denoising method, which effectively denoises the input twisted and penetrated hand poses to produce photorealistic results. We first propose the innovative joint-to-Gaussian surface representation, which accurately models the spatial relationships between hand skeleton joints and object Gaussians while highlighting hand-object penetrations and generalizing well to new hand poses and objects. Then, we propose a geometry-aware de-penetration algorithm that eliminates penetrations by detecting intersections between skeleton bones and object Gaussians and reposing any penetrated fingers onto the estimated underlying surface of the object. Experiments demonstrate that our method not only effectively reduces hand-object penetration depth but also produces more realistic rendering quality compared to the state-of-the-art methods MANUS+GEARS, MANUS+GeneOH, and$2 \text{DGS}+\text{Gene} \text{OH}$. The user study results show that our method significantly improves the users' visual perceptual experience regarding penetration and stability metrics. Project page: https://github.com/ZhaoLizz/GSHOIDenoiser
Lizhi Zhao, Xuequan Lu, Wei Ke 0001, Lili Wang 0006
ISMAR4
2025 FDF-VQVAE: A Frequency Disentanglement and Fusion Learning Framework for Multi-sequence MRI Enhancement
Xinghe Xie, Luyi Han, Yue Sun 0001, Chi Kin Lam, Jian Zheng 0001, Tong Tong 0001, Wei Ke 0001, Chan-Tong Lam, Tao Tan 0002
MICCAI (3)7
2025 Cold-Start Microservice Workload Prediction by Dynamically Annealed Graph-Regularized Matrix Factorization
Xiaoxuan Luo, Hong Shen 0001, Wei Ke 0001
PDCAT3
2025 Pyramid Mamba with Multi-directional Scanning for Light Field Image Super-Resolution
Wenqi Lyu, Wei Ke 0001, Hao Sheng 0001
PDCAT2
2025 Robust Aggregation on Federated Distillation in Adversarial Environments via Entropy and Semantic-Aware Logit Filtering
Hong Shen 0001, Wei Ke 0001, Wenqi Lyu
PDCAT3
2025 Enhance Privacy Protection by Reducing Information Exchanges in Multi-agent System Consensus
Hong Shen 0001, Wei Ke 0001
PDCAT3
2025 Stereo matching on epipolar plane image for light field depth estimation via oriented structure
Rongshan Chen, Hao Sheng 0001, Ruixuan Cong, Da Yang 0001, Zhenglong Cui, Wei Ke 0001
Eng. Appl. Artif. Intell.7
2025 Spatio-temporal attention based collaborative local-global learning for traffic flow prediction
Haiyang Chi, Yuhuan Lu 0001, Can Xie, Wei Ke 0001, Bidong Chen
Eng. Appl. Artif. Intell.4
2025 Latent temporal smoothness-induced Schatten-p norm factorization for sequential subspace clustering
Zhen-Zhen Zhao, Tong-Wei Lu, Wei Ke 0001, Yang Zhang 0032, Ming-Qing Zhang
Eng. Appl. Artif. Intell.4
2025 Regression loss-assisted conditional style generative adversarial network for virtual sample generation with small data in soft sensing
Xue-Yu Zhang, Wei Ke 0001, Ming-Qing Zhang, Yuan Xu 0026
Eng. Appl. Artif. Intell.3
2025 An algorithm for improving lower bounds in dynamic time warping
Yuqi Luo, Xinyi Fang, Wei Ke 0001, Chan-Tong Lam, Sio Kei Im, Luís Paquete
Expert Syst. Appl.3
2025 Category-wise Fine-Tuning: Resisting incorrect pseudo-labels in multi-label image classification with partial labels
Chak Fong Chong, Xinyi Fang, Jielong Guo, Pedro H. Abreu, Yapeng Wang 0001, Xu Yang 0010, Wei Ke 0001, Sio Kei Im
Neurocomputing7
2025 Bi-branch bidirectional coupled interaction fusion network for multi-retinal diseases diagnosis
Shaobin Chen, Tao Tan 0002, Wei Ke 0001, Xiayu Xu, Yanwu Xu 0004, Chan-Tong Lam, Yue Sun 0001
Knowl. Based Syst.4
2025 Dynamic spatiotemporal graph convolutional network collaborative pre-training learning for traffic flow prediction
Haiyang Chi, Yuhuan Lu 0001, Yirong Zhu, Wei Ke 0001, Hanbin Mao
Knowl. Based Syst.4
2025 A GPU-Enabled Framework for Light Field Efficient Compression and Real-Time Rendering
abstract
Real-time rendering offers instantaneous visual feedback, making it crucial for mixed-reality applications. The light field captures both light intensity and direction in a 3D environment, serving as a data-rich medium to enhance mixed-reality experiences. However, two major challenges remain: 1) current light field rendering techniques are unsuitable for real-time computation, and 2) existing real-time methods cannot efficiently process high-dimensional light field data on GPU platforms. To overcome these challenges, we propose an framework utilizing a compact neural representation of light field data, implemented on a GPU platform for real-time rendering. This framework provides both compact storage and high-fidelity real-time computation. Specifically, we introduce a ray global alignment strategy to simplify the framework and improve practicality. This strategy enables the learning of an optimal embedding for all local rays in a globally consistent way, removing the need for camera pose calculations. To achieve effective compression, the neural light field is employed to map each embedded ray to its corresponding color. To enable real-time rendering, we design a novel super-resolution network to enhance rendering speed. Extensive experiments demonstrate that our framework significantly enhances compression efficiency and real-time rendering performance, achieving nearly 50$\mathbf{\times}$compression ratio and 100 FPS rendering.
Mingyuan Zhao 0001, Hao Sheng 0001, Rongshan Chen, Ruixuan Cong, Tun Wang, Zhenglong Cui, Da Yang 0001, Shuai Wang 0027, Wei Ke 0001
IEEE Trans. Computers9
2025 PFL-ALP: Personalized Federated Learning Against Backdoor Attacks via Attention-Based Local Purification
abstract
Federated learning (FL) enables collaborative model training with local data privacy preserving, but is vulnerable to backdoor attacks from malicious clients. These attacks can manipulate the global model to produce malicious output when encountering specific triggers. Existing defenses, categorized as server-side and client-side approaches, have limitations such as reliance on auxiliary data availability, susceptibility to inference attacks, and instability under non-independent and identically distributed (Non-IID) data. In response to these challenges, we propose a Personalized Federated Learning via Attention-based Local Purification (PFL-ALP) algorithm, a hybrid defense mechanism integrating server-side dynamic clustering and client-side purification enhanced with personalized model knowledge. This approach effectively mitigates bias introduced by Non-IID data on the server side and further purifies the backdoored model on the client side. Specifically, we employ neural attention distillation (NAD) for model purification and enhance it with personalized model knowledge, extending the effectiveness of NAD in Non-IID FL settings. This design makes PFL-ALP compatible with privacy protocols to mitigate inference attacks. Moreover, we establish a convergence guarantee for PFL-ALP and experimentally validate its superior performance in defending against various backdoor attacks compared to multiple state-of-the-art (SOTA) defenses across three datasets. The results show that even with malicious rates ranging from 30% to 90%, PFL-ALP can reduce the attack success rate by more than 69.4 percentage points, with the reduction in main task accuracy less than 12.4 percentage points.
Yifeng Jiang 0008, Xiaochen Yuan, Wei Ke 0001, Chan-Tong Lam, Sio Kei Im
IEEE Trans. Inf. Forensics Secur.4
2025 A Symmetric Self-Embedding Mechanism for High-Fidelity Image Recovery Against Tampering
abstract
Digital images are inherently fragile and vulnerable to malicious tampering, significantly compromising their authenticity and integrity. Image recovery is crucial for restoring altered content and preserving the reliability of digital images. Traditional fragile watermarking methods achieve high-quality recovery but fail under post-processing attacks, while existing deep learning-based approaches offer some robustness, yet often produce lower-quality recovered images, typically with a PSNR of around 28 dB. To address these challenges, we propose a novel Symmetric Self-embedding Mechanism for High-Fidelity Image Recovery against tampering (SSEM-HIR), which is capable of restoring tampered images with high quality while maintaining some robustness against common attacks. Unlike existing methods that use the fragility of watermarking solely for tampering localization, SSEM-HIR is the first work to integrate fragility with spatial symmetry, enabling high-quality tampering recovery. Specifically, our SSEM-HIR employs a hierarchical watermark embedding module to embed an inverted version of the original image, utilizing spatial symmetry to retrieve lost information from the extracted watermark. To further improve recovery quality, we design a Dual-branch Region-based Self-Recovery module, where a Spatial-based Watermark Extraction block restores tampered regions using embedded watermark information, while a Frequency-assisted Image Repair block compensates for quality degradation in the untampered area. Extensive experiments show that our method achieves an average PSNR of 34.14 dB under common attack scenarios, including noise addition, image scaling, Gaussian blurring, and no post-processing. This represents an improvement of over 5 dB and 18% in recovered image quality compared to state-of-the-art approaches.
Tong Liu 0021, Xiaochen Yuan, Wei Ke 0001, Chan-Tong Lam, Sio Kei Im, Pedro Martins 0003
IEEE Trans. Inf. Forensics Secur.3
2024 A Parameterized Generative Adversarial Network Using Cyclic Projection for Explainable Medical Image Classifications
abstract
Although current data augmentation methods are successful to alleviate the data insufficiency, conventional augmentation are primarily intra-domain while advanced generative adversarial networks (GANs) generate images remaining uncertain, particularly in small-scale datasets. In this paper, we propose a parameterized GAN (ParaGAN) that effectively controls the changes of synthetic samples among domains and highlights the attention regions for downstream classification. Specifically, ParaGAN incorporates projection distance parameters in cyclic projection and projects the source images to the decision boundary to obtain the class-difference maps. Our experiments show that ParaGAN can consistently outperform the existing augmentation methods with explainable classification on two small-scale medical datasets.
Xiangyu Xiong, Yue Sun 0001, Xiaohong Liu 0001, Chan-Tong Lam, Tong Tong 0001, Hao Chen 0037, Qinquan Gao, Wei Ke 0001, Tao Tan 0002
ICASSP8
2024 MSFGNet: Multi-Scale Features Gathering Network for Change Detection of Remote Sensing Images
abstract
Change detection is an important research area in remote sensing. To achieve accurate results, it is essential to extract multi-scale spatial information from images while filtering out noise. However, existing models lack this capability. Therefore, Multi-Scale Feature Gathering Network (MSFGNet) is proposed. Within MSFGNet, Bi-Temporal Image Multi-Level Fusion Module (BMF) is utilized to fuse bi-temporal remote sensing images. Additionally, Multi-Receptive Field Features Extraction Module (MRFE) is utilized to extract deep features. Within MRFE, Large Receptive Field Features Extraction Module (LRFE) and Multi-Scale Information Fusion Module (MSIF) are designed, which use large kernel convolution and dilated convolution respectively to capture spatial information with large receptive fields. Furthermore, Cross-Dimension Feature Sifting Fusion Module (CDFSF) is designed to sift noise from various dimensions, fusing valuable information. Across multiple public datasets, MSFGNet consistently achieves the best experimental results. The code can be accessed at https://github.com/juncyan/msfgnet.git.
Junqing Huang, Xiaochen Yuan, Chan-Tong Lam, Wei Ke 0001
ICME4
2024 Dual Hypergraph Convolution Networks for Image Forgery Localization
Xiaochen Yuan, Wei Ke 0001, Chan-Tong Lam
ICPR (22)3
2024 Enhancing Federated Learning Robustness in Non-IID Data Environments via MMD-Based Distribution Alignment
Hong Shen 0001, Wenqi Lyu, Wei Ke 0001
PDCAT4
2024 Advancing Evasion: Distributed Backdoor Attacks in Federated Learning
Hong Shen 0001, Wei Ke 0001
PDCAT3
2024 Online 3D behavioral tracking of aquatic model organism with a dual-camera system
Zewei Wu, Wei Zhang 0245, Guodong Sun 0001, Wei Ke 0001, Zhang Xiong 0001
Adv. Eng. Informatics5
2024 VVIR-OM: Efficient Object Manipulation in VR with Variable Virtual Interaction Region
abstract
Manipulating with virtual objects is a fundamental requirement in virtual environments. A good manipulation method needs to consider efficiency, accuracy, and comfort. This paper proposes VVIR-OM, an object manipulation method in virtual reality (VR) based on a variable virtual interaction region (VVIR). A hand interaction hemisphere region (HIHR) is introduced and constructed in real space, where the user is more comfortable manipulating the objects. Then a VVIR is introduced, and an interaction heat volume (IHV) based method is proposed to update VVIR during the process of the manipulation. At last, a mapping algorithm is proposed to map the user hand position in HIHR to the position in VVIR. Two user studies are designed to evaluate the performance of VVIR-OM. Compared to the state-of-the-art methods, VVIR-OM achieves significant improvements in task completion time, manipulation precision, and a significant reduction in fatigue. Moreover, VVIR-OM outperforms other methods in terms of task load and usability without the cost of cybersickness.
Qinwen Zheng, Lili Wang 0006, Wei Ke 0001, Sio Kei Im
Int. J. Hum. Comput. Interact.3
2024 Object manipulation based on the head manipulation space in VR
Lili Wang 0006, Wei Ke 0001, Sio Kei Im
Int. J. Hum. Comput. Stud.3
2024 A simple transformer-based baseline for crowd tracking with Sequential Feature Aggregation and Hybrid Group Training
Zewei Wu, Wei Ke 0001, Zhang Xiong 0001
J. Vis. Commun. Image Represent.3
2024 An accurate slicing method for dynamic time warping algorithm and the segment-level early abandoning optimization
Yuqi Luo, Wei Ke 0001, Chan-Tong Lam, Sio Kei Im
Knowl. Based Syst.2
2024 Light Field Depth Estimation for Non-Lambertian Objects via Adaptive Cross Operator
abstract
Light field (LF) depth estimation is a crucial basis for LF-related applications. Most existing methods are based on the Lambertian assumption and cannot deal with non-Lambertian surfaces represented by transparent objects and mirrors. In this paper, we propose a novel Adaptive-Cross-Operator-based(ACO) depth estimation algorithm for non-Lambertian LF. By analyzing the imaging characteristics of non-Lambertian regions, it is found that the difficulty of depth estimation lies in the photo inconsistency of the center view. Combining with the two-branch structure, we propose ACO with an inter-branch cooperation strategy to adaptively separate depth information with different reflectance coefficients. We discover that the bimodal distribution feature of the operator filtering results can assist in the separation of multi-layer scene information. The first detection branch filters the EPI and implicitly records the severity of multi-layer scene aliasing. According to the identification of bimodal distribution features, the non-Lambertian regions are marked out and the depth of the foreground is estimated. The second branch receives guidance from the first to dynamically adjust the inner weight and infer the background’s depth after weakening the interference from the foreground. Finally, the depth information separation of multi-layer scenes is achieved by extracting the unique X-shaped linear structure. Without the reflection coefficients of the non-Lambertian object, the proposed method can produce high-quality depth estimation under the transparency of 90% to 20%. Experimental results show that the proposed ACO outperforms state-of-the-art LF depth estimation methods in terms of accuracy and robustness.
Zhenglong Cui, Hao Sheng 0001, Da Yang 0001, Rongshan Chen, Wei Ke 0001
IEEE Trans. Circuits Syst. Video Technol.6
2024 Discriminative Feature Learning With Co-Occurrence Attention Network for Vehicle ReID
abstract
Vehicle Re-Identification (ReID) aims to find images of the same vehicle from different videos. It remains a challenging task in the video analysis field due to the huge appearance discrepancy of the same vehicle in cross-view matching and the subtle difference of different similar vehicles in same-view matching. In this paper, we propose a Co-occurrence Attention Net (CAN) to deal with these two challenges. Specifically, CAN consists of two branches, a main branch and an aware branch. The main branch is in charge of extracting global features that are consistent in most views. This feature encodes holistic information such as color and pose, however, it can not handle cross/same-view hard cases, as shown in Fig.1. Therefore, the aware branch is designed to focus on the local details and viewpoint information, which can become an important complement for those hard cases. Considering that the positions of local areas such as wheels and logos change with the viewpoint, Aware Attention Module is introduced to find the hidden relationship among local areas and seamlessly combine the viewpoint information simultaneously. Then, CAN is trained by a partition-and-reunion-based loss, which can narrow the intra-class distance and increase the inter-class distance. Further, an adaptive co-occurrence view emphasize strategy is adopted to fully utilize the learned features. Experimental results on three widely used datasets including VeRi-776, VehicleID and VERI-Wild demonstrate the effectiveness of our method and competitive performance with other state-of-the-art methods.
Hao Sheng 0001, Shuai Wang 0027, Haobo Chen, Da Yang 0001, Wei Ke 0001
IEEE Trans. Circuits Syst. Video Technol.7
2024 Distributed Collaborative Object Retrieval With Blockchain-Based Edge Computing
abstract
In the current industrial informatics society, the numerous cameras deployed in the modern city promote the development of various video services, such as security monitoring and object retrieval. However, traditional methods encounter data leakage risks. Some camera owners are reluctant to share their data since the video contains confidential information. Meanwhile, domain diversities between cameras bring obstacles to practical object retrieval applications. To deal with these dilemmas, we propose a blockchain-based collaborative object retrieval (BCOR) system that can protect privacy as much as possible. BCOR includes two core components: multicamera reidentification framework (MC-ReF) and multicamera collaborative chain (M2C-Chain). Specifically, MC-ReF leverages visual relevance attention net (VRANet) to distinguish object identities in edge nodes. Through domain adaptation gradient optimization, VRANet can adapt to different cameras without the need for private camera data. M2C-Chain is responsible for maintaining the security and trust of the system. Through M2C-Chain, the collaboration among different nodes is transferred into a transaction-based manner, which is validated by a deeply integrated consensus. Finally, we implement a prototype system and deploy it into a real-world outdoor scene. The experiments indicate that BCOR achieves 30%–35% average improvement in domain adaptation on mean average precision and Rank-1 indicators. The performance analysis and security experiments also prove the efficiency and stability of BCOR.
Shuai Wang 0027, Hao Sheng 0001, Dazhi Yang 0003, Da Yang 0001, Yang Zhang 0032, Wei Ke 0001
IEEE Trans. Ind. Informatics7
2024 Triple Consistency for Transparent Cheating Problem in Light Field Depth Estimation
abstract
Depth estimation extracting scenes' structural information is a key step in various light field(LF) applications. However, most existing depth estimation methods are based on the Lambertian assumption, which limits the application in non-Lambertian scenes. In this paper, we discover a unique transparent cheating problem for non-Lambertian scenes which can effectively spoof depth estimation algorithms based on photo consistency. It arises because the spatial consistency and the linear structure superimposed on the epipolar plane image form new spurious lines. Therefore, we propose centrifugal consistency and centripetal consistency for separating the depth information of multi-layer scenes and correcting the error due to the transparent cheating problem, respectively. By comparing the distributional characteristics and the number of minimal values of photo consistency and centrifugal consistency, non-Lambertian regions can be efficiently identified and initial depth estimates obtained. Then centripetal consistency is exploited to reject the projection from different layers and to address transparent cheating. By assigning decreasing weights radiating outward from the central view, pixels with a concentration of colors close to the central viewpoint are considered more significant. The problem of underestimating the depth of background caused by transparent cheating is effectively solved and corrected. Experiments on synthetic and real-world data show that our method can produce high-quality depth estimation under the transparency and the reflectivity of 90% to 20%. The proposed triple-consistency-based algorithm outperforms state-of-the-art LF depth estimation methods in terms of accuracy and robustness.
Zhenglong Cui, Da Yang 0001, Hao Sheng 0001, Rongshan Chen, Ruixuan Cong, Wei Ke 0001
IEEE Trans. Multim.7
2024 VPRF: Visual Perceptual Radiance Fields for Foveated Image Synthesis
abstract
Neural radiance fields (NeRF) has achieved revolutionary breakthrough in the novel view synthesis task for complex 3D scenes. However, this new paradigm struggles to meet the requirements for real-time rendering and high perceptual quality in virtual reality. In this paper, we propose VPRF, a novel visual perceptual based radiance fields representation method, which for the first time integrates the visual acuity and contrast sensitivity models of human visual system (HVS) into the radiance field rendering framework. Initially, we encode both the appearance and visual sensitivity information of the scene into our radiance field representation. Then, we propose a visual perceptual sampling strategy, allocating computational resources according to the HVS sensitivity of different regions. Finally, we propose a sampling weight-constrained training scheme to ensure the effectiveness of our sampling strategy and improve the representation of the radiance field based on the scene content. Experimental results demonstrate that our method renders more efficiently, with higher PSNR and SSIM in the foveal and salient regions compared to the state-of-the-art FoV-NeRF. The results of the user study confirm that our rendering results exhibit high-fidelity visual perception.
Jian Wu 0033, Runze Fan, Wei Ke 0001, Lili Wang 0006
IEEE Trans. Vis. Comput. Graph.4
2024 SMigraPH: a perceptually retained method for passive haptics-based migration of MR indoor scenes
Qixiang Ma, Lili Wang 0006, Wei Ke 0001, Sio Kei Im
Vis. Comput.3
2024 Joint attribute soft-sharing and contextual local: a multi-level features learning network for person re-identification
Wangmeng Wang, Yanbing Chen, Dengwen Wang, Zhixin Tie, Linbing Tao, Wei Ke 0001
Vis. Comput.6
2023 Category-Wise Fine-Tuning for Image Multi-label Classification with Partial Labels
Chak Fong Chong, Xu Yang 0010, Tenglong Wang, Wei Ke 0001, Yapeng Wang 0001
ICONIP (11)4
2023 Light Field Image Super-Resolution via Global-View Information Adaption and Angular Attention Fusion
Wei Zhang 0245, Wei Ke 0001, Hao Sheng 0001
ICONIP (11)2
2023 Locomotion-aware Foveated Rendering
abstract
Optimizing rendering performance improves the user's immersion in virtual scene exploration. Foveated rendering uses the features of the human visual system (HVS) to improve rendering performance without sacrificing perceptual visual quality. We collect and analyze the viewing motion of different locomotion methods, and describe the effects of these viewing motions on HVS's sensitivity, as well as the advantages of these effects that may bring to foveated rendering. Then we propose the locomotion-aware foveated rendering method (LaFR) to further accelerate foveated rendering by leveraging the advantages. In LaFR, we first introduce the framework of LaFR. Secondly, we propose an eccentricity-based shading rate controller that provides the shading rate control of the given region in foveated rendering. Thirdly, we propose a locomotion-aware log-polar mapping method, which controls the foveal average shading rate, the peripheral shading rate decrease speed, and the overall shading quantity with the locomotion-aware coefficients based on the eccentricity-based shading rate controller. LaFR achieves similar perceptual visual quality as the conventional foveated rendering while achieving up to 1.6× speedup. Compared with the full resolution rendering, LaFR achieves up to 3.8× speedup.
Xuehuai Shi, Lili Wang 0006, Jian Wu 0033, Wei Ke 0001, Chan-Tong Lam
VR4
2023 Wearable Real-time Air-writing System Employing KNN and Constrained Dynamic Time Warping
abstract
In the digital world, gesture recognition plays a crucial role in human-computer interaction (HCI). In this paper, we propose an innovative wearable air-writing system that allows users to write the English alphabet and Arabic numerals in free space without using any predefined gestures or rules. Based on an Inertial Measurement Unit (IMU), the proposed air-writing wearable device uses the constrained dynamic time warping (cDTW) algorithm for the distance measure and K-nearest neighbors (KNN) as the classifier. In addition, to increase the recognition accuracy and meet HCI requirements, we develop a novel method that allows users to rapidly switch to correct recognition results when the initial results are erroneous. In the experiment, the accuracy rate is 88.9% for the alphabet and 10 decimal digits in the user-dependent condition, and the recognition is in real-time, consuming only 0.427s for each character, which is superior to many other approaches that employ classic DTW or FastDTW. With the proposed HCI design, character input accuracy of over 95% can be obtained in about 1 second. We also simulated the application scenarios of Parkinson’s disease patients and obtained a high accuracy rate of 85.4%. Besides, we explored the variety of K values in KNN and w values in cDTW, and propose a multi-template system that gives new optimization directions for the KNN-cDTW algorithm.
Yuqi Luo, Wei Ke 0001, Chan-Tong Lam
WCNC2
2023 MFSRNet: spatial-angular correlation retaining for light field super-resolution
Hao Sheng 0001, Da Yang 0001, Zhenglong Cui, Ruixuan Cong, Wei Ke 0001
Appl. Intell.6
2023 Light field super-resolution using complementary-view feature attention
abstract
Light field (LF) cameras record multiple perspectives by a sparse sampling of real scenes, and these perspectives provide complementary information. This information is beneficial to LF super-resolution (LFSR). Compared with traditional single-image super-resolution, LF can exploit parallax structure and perspective correlation among different LF views. Furthermore, the performance of existing methods are limited as they fail to deeply explore the complementary information across LF views. In this paper, we propose a novel network, called the light field complementary-view feature attention network (LF-CFANet), to improve LFSR by dynamically learning the complementary information in LF views. Specifically, we design a residual complementary-view spatial and channel attention module (RCSCAM) to effectively interact with complementary information between complementary views. Moreover, RCSCAM captures the relationships between different channels, and it is able to generate informative features for reconstructing LF images while ignoring redundant information. Then, a maximum-difference information supplementary branch (MDISB) is used to supplement information from the maximum-difference angular positions based on the geometric structure of LF images. This branch also can guide the process of reconstruction. Experimental results on both synthetic and real-world datasets demonstrate the superiority of our method. The proposed LF-CFANet has a more advanced reconstruction performance that displays faithful details with higher SR accuracy than state-of-the-art methods.
Wei Zhang 0245, Wei Ke 0001, Da Yang 0001, Hao Sheng 0001, Zhang Xiong 0001
Comput. Vis. Media2
2023 Hybrid Motion Model for Multiple Object Tracking in Mobile Devices
abstract
For an intelligent transportation system, multiple object tracking (MOT) is more challenging from the traditional static surveillance camera to mobile devices of the Internet of Things (IoT). To cope with this problem, previous works always rely on additional information from multivision, various sensors, or precalibration. Only based on a monocular camera, we propose a hybrid motion model to improve the tracking accuracy in mobile devices. First, the model evaluates camera motion hypotheses by measuring optical flow similarity and transition smoothness to perform robust camera trajectory estimation. Second, along the camera trajectory, smooth dynamic projection is used to map objects from image to world coordinate. Third, to deal with trajectory motion inconsistency, which is caused by occlusion and interaction of long time interval, tracklet motion is described by the multimode motion filter for adaptive modeling. Fourth, in tracklets association, we propose a spatiotemporal evaluation mechanism, which achieves higher discriminability in motion measurement. Experiments on MOT15, MOT17, and KITTI benchmarks show that our proposed method improves the trajectory accuracy, especially in mobile devices and our method achieves competitive results over other state-of-the-art methods.
Yubin Wu, Hao Sheng 0001, Yang Zhang 0032, Shuai Wang 0027, Zhang Xiong 0001, Wei Ke 0001
IEEE Internet Things J.6
2023 Uncertainty-guided joint attention and contextual relation network for person re-identification
Dengwen Wang, Yanbing Chen, Wangmeng Wang, Zhixin Tie, Xian Fang, Wei Ke 0001
J. Vis. Commun. Image Represent.6
2023 AEA-Net:Affinity-supervised entanglement attentive network for person re-identification
Dengwen Wang, Yanbing Chen, Lingbing Tao, Chentao Hu, Zhixin Tie, Wei Ke 0001
Pattern Recognit. Lett.6
2022 Group Guided Data Association for Multiple Object Tracking
Yubin Wu, Hao Sheng 0001, Shuai Wang 0027, Yang Liu 0088, Zhang Xiong 0001, Wei Ke 0001
ACCV (7)6
2022 Online 3D Reconstruction of Zebrafish Behavioral Trajectories within A Holistic Perspective
abstract
Recording activities of zebrafish is a fundamental task in biological research that aims to accurately track individuals and recover their real-world movement trajectories from multiple viewpoint videos. In this paper, we propose a novel online tracking solution based on a holistic perspective that leverages the correlation of appearance and location across views. It first reconstructs the 3D coordinates of targets frame by frame and then tracks them directly in 3D space instead of a 2D image plane. However, it is not trivial to implement such a solution which requires the association of targets across views and neighboring frames under occlusion and parallax distortion. To cope with that, we propose the view-invariant feature representation and the Kalman filter-based 3D state estimation, and combine the advantages of both to generate robust 3D trajectories. Extensive experiments on public datasets verify the efficiency and effectiveness of the approach.
Zewei Wu, Wei Ke 0001, Wei Zhang 0245, Zhang Xiong 0001
BIBM2
2022 Transformation from MVC Applications to Smart Contracts
abstract
A smart contract is a program running on a blockchain platform. Smart devices send data to smart contracts or change their own status based on smart contracts. Businesses want to use smart devices and smart contracts to streamline workflow because smart contracts reduce the need in trusted intermediators and cut down enforcement costs. However, developing smart contract applications is challenging due to different memory models, different interaction models, and a dearth of supporting tools and libraries. When developers are asked to reimplement conventional applications to smart contracts, it is thus ideal to automatically transform them, avoiding manual labor, as well as ensuring reliability and security. This paper contributes a set of rules to transform conventional applications in the Model-View-Controller (MVC) pattern into smart contracts running on Hyperledger Fabric, a blockchain platform preferred by businesses. Major transformations are performed in the model, while in the controller, model calls are replaced by smart contract calls. The source application and the target smart contract are all written in Java. Our rules add read-your-writes consistency that Hyperledger Fabric does not natively support. Runtime pre- and post-condition checking in the original application is supported in the transformed smart contract. We evaluated our rules on CoCoME and other MVC applications, and all code ran correctly and passed unit tests.
Qiqi Gu 0001, Wei Ke 0001, Yilong Yang 0001
EUC2
2022 Partition and Reunion: A Viewpoint-Aware Loss for Vehicle Re-Identification
abstract
Vehicle Re-Identification (ReID) aims to retrieve images of vehicles with the same identity from different scenarios. It is a challenging task due to the large intra-identity discrepancy caused by viewpoint variations and the subtle inter-identity difference produced by similar appearances. In this paper, we propose a Viewpoint-Aware Loss (VAL) function to deal with these challenges. Specifically, we propose partition and reunion operations in VAL, which significantly shrinks the intra-identity distance and acquires viewpoint-invariant representations. In addition, we embed a multi-decision boundary mechanism in VAL. It contributes to enlarging the inter-identity distance. A comprehensive evaluation on two benchmarks shows the superiority of our method in contrast to a series of existing state-of-the-arts.
Haobo Chen, Yang Liu 0088, Wei Ke 0001, Hao Sheng 0001
ICIP4
2022 Mask Guided Spatial-Temporal Fusion Network for Multiple Object Tracking
abstract
Multi-object trackers make the association almost perfectly when no occlusion occurred between two or more targets. However, it is hard to extract reliable features on account of partial occlusion caused by a nearby object, which often leads to tracking failure. In this paper, we utilize mask to guide attention of the neural network in order to focus on the visible part of the target and design a tracklet-level feature extraction method. Then, a tracking framework is proposed based on a mask guided fusion network and multi-hypothesis tracking algorithm. Comprehensive evaluation on the MOT17 dataset shows that our approach achieves competitive results.
Shuangye Zhao, Yubin Wu, Shuai Wang 0027, Wei Ke 0001, Hao Sheng 0001
ICIP4
2022 Data Association with Graph Network for Multi-Object Tracking
Yubin Wu, Hao Sheng 0001, Shuai Wang 0027, Yang Liu 0088, Wei Ke 0001, Zhang Xiong 0001
KSEM (1)5
2022 A Local Rotation Transformation Model for Vehicle Re-Identification
Yanbing Chen, Wei Ke 0001, Hao Sheng 0001, Zhang Xiong 0001
WASA (1)2
2022 Enhancing Efficiency and Quality of Image Caption Generation with CARU
Xuefei Huang, Wei Ke 0001, Hao Sheng 0001
WASA (2)2
2022 Local perspective based synthesis for vehicle re-identification: A transformation state adversarial method
Yanbing Chen, Wei Ke 0001, Hong Lin 0006, Chan-Tong Lam, Kai Lv 0002, Hao Sheng 0001, Zhang Xiong 0001
J. Vis. Commun. Image Represent.2
2022 Multiple classifier for concatenate-designed neural network
Ka-Hou Chan, Sio Kei Im, Wei Ke 0001
Neural Comput. Appl.3
2021 Near-Realtime Face Mask Wearing Recognition Based on Deep Learning
abstract
COVID-19 pandemic has led to serious economic and life losses. Face Masks serve as first infection barrier when used in public spaces. In this paper, we propose a new near-realtime method to automatically recognize face mask wearing that combines human posture recognition with convolutional neural network (CNN). We use the power of human posture recognition to perform background filtering and spatial reduction in the original images. The outcome is then used by a trained CNN model to identify if the subject is wearing a mask. We exploit Openpose to identify the skeleton of human body and locate the facial region thus spatially reducing the area to be processed by the CNN framework. We then adopt supervised learning approach to detect if a face mask is present. The CNN is trained using images, cropped to the supposed face mask covered region. This approach led to a substantial reduction in neural network complexity yet improving the recognition accuracy. The system has been evaluated in a multitude of scenarios using images taken in public places at different time of day and with different angles. Overall, our system achieves a recognition accuracy of 95.8% and 94.6% in daytime and nighttime respectively.
Hong Lin 0006, Rita Tse, Su-Kit Tang, Yanbing Chen, Wei Ke 0001, Giovanni Pau 0001
CCNC5
2021 A Neural Architecture for Detecting Identifier Renaming from Diff
Qiqi Gu 0001, Wei Ke 0001
IDEAL2
2021 Combining Pose Invariant and Discriminative Features for Vehicle Reidentification
abstract
Vehicle reidentification, aiming at identifying vehicles across images, has drawn a lot of attention and has made significant achievements in recent years. However, vehicle reidentification remains a challenging task caused by severe appearance changes due to different orientations. In practice, the result of reidentification is greatly influenced by the pose of vehicles, and we call this influence as a pose barrier problem. One way to address the pose barrier problem is to train a feature representation that is invariant for various vehicle poses. To this end, we present pose robust features (PRFs) that contains two components: 1) pose-invariant features (PIFs) and 2) pose discriminative features (PDFs). On the one hand, PIF is the expert in exploring the overall characteristic of vehicles. When training PIF, we adopt an identity classifier as well as an orientation classifier. In addition, an adversarial loss is deployed in the PIF network. On the other hand, we design a PDF network, which has a similar architecture to the PIF network but can distinguish the difference between local details. The difference between PDF and PIF is that the network of training PDF does not apply the adversarial loss. Finally, by combining PIF and PDF, PRF has the advantages of the two features and can alleviate the influence of the pose barrier problem. Experiments are conducted on the VeRi-776 and VehicleID data sets. We show that PIF and PDF are complementary and that PRF produces competitive performance compared with state-of-the-art approaches.
Hao Sheng 0001, Kai Lv 0002, Yang Liu 0088, Wei Ke 0001, Weifeng Lyu, Zhang Xiong 0001, Wei Li 0022
IEEE Internet Things J.4
2020 Variable-Depth Convolutional Neural Network for Text Classification
Ka-Hou Chan, Sio Kei Im, Wei Ke 0001
ICONIP (5)3
2020 CARU: A Content-Adaptive Recurrent Unit for the Transition of Hidden State in NLP
Ka-Hou Chan, Wei Ke 0001, Sio Kei Im
ICONIP (1)2
2020 Camera Style Guided Feature Generation for Person Re-identification
Hantao Hu, Yang Liu 0088, Kai Lv 0002, Yanwei Zheng, Wei Zhang 0245, Wei Ke 0001, Hao Sheng 0001
WASA (1)6
2020 A Dual Scale Matching Model for Long-Term Association
Yubin Wu, Shuai Wang 0027, Yang Zhang 0032, Yanbing Chen, Wei Ke 0001, Hao Sheng 0001
WASA (1)6
2020 Mining Hard Samples Globally and Efficiently for Person Reidentification
abstract
Person reidentification (ReID) is an important application of Internet of Things (IoT). ReID recognizes pedestrians across camera views at different locations and time, which is usually treated as a ranking task. An essential part of this task is the hard sample mining. Technically, two strategies could be employed, i.e., global hard mining and local hard mining. For the former, hard samples are mined within the entire training set, while for the latter, it is done in mini-batches. In literature, most existing methods operate locally. Examples include batch-hard sample mining and semihard sample mining. The reason for the rare use of global hard mining is the high computational complexity. In this article, we argue that global mining helps to find harder samples that benefit model training. To this end, this article introduces a new system to: 1) efficiently mine hard samples (positive and negative) from the entire training set and 2) effectively use them in training. Specifically, a ranking list network coupled with a multiplet loss is proposed. On the one hand, the multiplet loss makes the ranking list progressively created to avoid the time-consuming initialization. On the other hand, the multiplet loss aims to make effective use of the hard and easy samples during training. In addition, the ranking list makes it possible to globally and effectively mine hard positive and negative samples. In the experiments, we explore the performance of the global and local sample mining methods, and the effects of the semihard, the hardest, and the randomly selected samples. Finally, we demonstrate the validity of our theories using various public data sets and achieve competitive results via a quantitative evaluation.
Hao Sheng 0001, Yanwei Zheng, Wei Ke 0001, Dongxiao Yu, Xiuzhen Cheng, Weifeng Lyu, Zhang Xiong 0001
IEEE Internet Things J.3
2020 Multiplex Labeling Graph for Near-Online Tracking in Crowded Scenes
abstract
In recent years, the demand for intelligent devices related to the Internet of Things (IoT) is rapidly increasing. In the field of computer vision, many algorithms have been preinstalled in IoT devices to achieve higher efficiency, such as face recognition, area detection, target tracking, etc. Tracking is an important but complex task that needs high efficiency solutions in real applications. There is a common assumption that detection can only represent one pedestrian to describe nonoverlapping in physical space. In fact, the pixels of the image do not exactly correspond to the positions in the real world. In order to overcome the limitation of this assumption, we remove this unreasonable assumption and present a novel idea that each detector response can have multiple labels to describe different targets at the same time. Therefore, we propose a graph-based method for near-online tracking in this article. We introduce a detection multiplexing method for tracking in the monocular image and propose a multiplex labeling graph (MLG) model. Each node in MLG has the ability to represent multiple targets. In addition, we improve the shortage of graph-based trackers in using temporal features. We construct long short-term memory networks to model motion and appearance features for MLG optimization. On the public multiobject tracking challenge benchmark, our near-online method gains satisfactory efficiency and achieves state-of-the-art results without additional private detection as well.
Yang Zhang 0032, Hao Sheng 0001, Yubin Wu, Shuai Wang 0027, Wei Ke 0001, Zhang Xiong 0001
IEEE Internet Things J.5
2020 Hypothesis Testing Based Tracking With Spatio-Temporal Joint Interaction Modeling
abstract
Data association is one of the key research in tracking-by-detection framework. Due to frequent interactions among targets, there are various relationships among trajectories in crowded scenes which leads to problems in data association, such as association ambiguity, association omission, etc. To handle these problems, we propose hypothesis-testing based tracking (HTBT) framework to build potential associations between target by constructing and testing hypotheses. In addition, a spatio-temporal interaction graph (STIG) model is introduced to describe the basic interaction patterns of trajectories and test the potential hypotheses. Based on network flow optimization, we formulate offline tracking as a MAP problem. Experimental results show that our tracking framework improves the robustness of tracklet association when detection failure occurs during tracking. On the public MOT16, MOT17 and MOT20 benchmark, our method achieves competitive results compared with other state-of-the-art methods.
Hao Sheng 0001, Yang Zhang 0032, Yubin Wu, Shuai Wang 0027, Weifeng Lyu, Wei Ke 0001, Zhang Xiong 0001
IEEE Trans. Circuits Syst. Video Technol.6
2020 Long-Term Tracking With Deep Tracklet Association
abstract
Recently, most multiple object tracking (MOT) algorithms adopt the idea of tracking-by-detection. Relevant research shows that the performance of the detector obviously affects the tracker, while the improvement of detector is gradually slowing down in recent years. Therefore, trackers using tracklet (short trajectory) are proposed to generate more complete trajectories. Although there are various tracklet generation algorithms, the fragmentation problem still often occurs in crowded scenes. In this paper, we introduce an iterative clustering method that generates more tracklets while maintaining high confidence. Our method shows robust performance on avoiding internal identity switch. Then we propose a deep association method for tracklet association. In terms of motion and appearance, we construct motion evaluation network (MEN) and appearance evaluation network (AEN) to learn long-term features of tracklets for association. In order to explore more robust features of tracklets, a tracklet-based training mechanism is also introduced. Tracklet groups are used as the input of the networks instead of discrete detections. Experimental results show that our training method enhances the performance of the networks. In addition, our tracking framework generates more complete trajectories while maintaining the unique identity of each target as the same time. On the latest MOT 2017 benchmark, we achieve state-of-the-art results.
Yang Zhang 0032, Hao Sheng 0001, Yubin Wu, Shuai Wang 0027, Weifeng Lyu, Wei Ke 0001, Zhang Xiong 0001
IEEE Trans. Image Process.6
2020 Automated Prototype Generation From Formal Requirements Model
abstract
Prototyping is an effective and efficient way of requirements validation to avoid introducing errors in the early stage of software development. However, manually developing a prototype of a software system requires additional efforts, which would increase the overall cost of software development. In this article, we present an approach with a developed tool RM2PT to automated prototype generation from formal requirements models for requirements validation. A requirements model consists of a use case diagram, a conceptual class diagram, use case definitions specified by system sequence diagrams, and the contracts of their system operations. A system operation contract is formally specified by a pair of pre and postconditions in object constraint language. We propose a method with a set of transformation rules to decompose a contract into executable parts and nonexecutable parts. An executable part can be automatically transformed into a sequence of primitive operations by applying their corresponding rules, and a nonexecutable part is not transformable with the rules. The tool RM2PT provides a mechanism for developers to develop a piece of program for each nonexecutable part manually, which can be plugged into the generated prototype source code automatically. We have conducted four case studies with over 50 use cases. The experimental result shows that the 93.65% system operations are executable, and only 6.35% are nonexecutable, which can be implemented by developers manually or invoking the third-party application programming interface (APIs). Overall, the result is satisfactory. Each 1 s generated prototype of four case studies requires approximate one day's manual implementation by a skilled programmer. The proposed approach with the developed computer-aided software engineering tool can be applied to the software industry for requirements engineering.
Yilong Yang 0001, Wei Ke 0001, Zhiming Liu 0001
IEEE Trans. Reliab.3
2019 Image resizing enhancement with DCT coefficients
abstract
According to the principle of scale transformation, signal expansion in the time domain corresponds to compression in the frequency domain, so that information energy is concentrated in the low frequency part. In this paper, an image enlargement algorithm based on the Discrete Cosine Transform (DCT) is proposed, which preserves the low frequency of the image and combines the corresponding enhancement coefficients to realize the resizing operation in the DCT domain. This paper also reach to proving of the enhancement value determinant. Then comparing the experiment on the scaled image with other interpolation algorithms, the result shows that our algorithm performs better than other methods. This method can also be carried out during DCT transformation, and is easier to implement than other methods.
Ka-Hou Chan, Sio Kei Im, Wei Ke 0001
ICMV3
2019 Pedestrian Similarity Extraction to Improve People Counting Accuracy
abstract
Current state-of-the-art single shot object detection pipelines, composed by an object detector such as Yolo, generate multiple detections for each object, requiring a post-processing Non-Maxima Suppression (NMS) algorithm to remove redundant detections. However, this pipeline struggles to achieve high accuracy, particularly in object counting applications, due to a trade-off between precision and recall rates. A higher NMS threshold results in fewer detections suppressed and, consequently, in a higher recall rate, as well as lower precision and accuracy. In this paper, we have explored a new pedestrian detection pipeline which is more flexible, able to adapt to different scenarios and with improved precision and accuracy. A higher NMS threshold is used to retain all true detections and achieve a high recall rate for different scenarios, and a Pedestrian Similarity Extraction (PSE) algorithm is used to remove redundant detentions, consequently improving counting accuracy. The PSE algorithm significantly reduces the detection accuracy volatility and its dependency on NMS thresholds, improving the mean detection accuracy for different input datasets.
Xu Yang 0010, José Gaspar, Wei Ke 0001, Chan-Tong Lam, Yanwei Zheng, Weng Hong Lou, Yapeng Wang 0001
ICPRAM3
2019 Deep Learning for Web Services Classification
abstract
Automated service classification plays a crucial role in service discovery, selection, and composition. Machine learning has been used for service classification in recent years. However, the performance of conventional machine learning methods highly depends on the quality of manual feature engineering. In this paper, we present a deep neural network to automatically abstract low-level representation of service description to high-level features without feature engineering and then predict service classification on 50 service categories. To demonstrate the effectiveness of our approach, we conduct a comprehensive experimental study by comparing 10 machine learning methods on 10,000 real-world web services. The result shows that the proposed deep neural network can achieve higher accuracy than other machine learning methods.
Yilong Yang 0001, Wei Ke 0001, Weiru Wang 0001
ICWS2
2019 Spatio-Temporal Correlation Graph for Association Enhancement in Multi-object Tracking
Hao Sheng 0001, Yang Zhang 0032, Yubin Wu, Jiahui Chen 0001, Wei Ke 0001
KSEM (1)6
2019 RM2PT: Requirements Validation through Automatic Prototyping
abstract
Prototyping is an effective and efficient way of requirements validation to avoid introducing errors in the early stage of software development. Our previous work presents a tool RM2PT to automatically generate prototypes from requirements models. The stakeholders can easily check whether the requirements reflect their real needs by investigating the executions of use cases in the generated prototypes. However, the conflict and contradictory of the requirements are hard to be discovered. In this paper, we enhance RM2PT by introducing consistency checking and state observations in the generated prototypes. Requirements inconsistency can be automatically detected and further fixed through carefully analyzing the contracts of system operations and system state observations. We have conducted four case studies with over 50 use cases. The experimental result shows that 107 requirements inconsistency are founded in requirements validations. Overall, the result is satisfiable, and the enhanced RM2PT can be further applied to the software industry for requirements validation. The tool can be downloaded at http://rm2pt.mydreamy.net and a demo video casting its features is at https://youtu.be/Y7GNa57WGfA.
Yilong Yang 0001, Wei Ke 0001
RE2
2019 An easy-to-hard learning strategy for within-image co-saliency detection
Shaoyue Song, Hongkai Yu, Zhenjiang Miao, Dazhou Guo, Wei Ke 0001, Cong Ma 0004, Song Wang 0002
Neurocomputing5
2019 Iterative Multiple Hypothesis Tracking With Tracklet-Level Association
abstract
This paper proposes a novel iterative maximum weighted independent set (MWIS) algorithm for multiple hypothesis tracking (MHT) in a tracking-by-detection framework. MHT converts the tracking problem into a series of MWIS problems across the tracking time. Previous works solve these NP-hard MWIS problems independently without the use of any prior information from each frame, and they ignore the relevance between adjacent frames. In this paper, we iteratively solve the MWIS problems by using the MWIS solution from the previous frame rather than solving the problem from scratch each time. First, we define five hypothesis categories and a hypothesis transfer model, which explicitly describes the hypothesis relationship between adjacent frames. We also propose a polynomial-time approximation algorithm for the MWIS problem in MHT. In addition to that, we present a confident short tracklet generation method and incorporate tracklet-level association into MHT, which further improves the computational efficiency. Our experiments on both MOT16 and MOT17 benchmarks show that our tracker outperforms all the previously published tracking algorithms on both MOT16 and MOT17 benchmarks. Finally, we demonstrate that the polynomial-time approximate tracker reaches nearly the same tracking performance.
Hao Sheng 0001, Jiahui Chen 0001, Yang Zhang 0032, Wei Ke 0001, Zhang Xiong 0001, Jingyi Yu 0001
IEEE Trans. Circuits Syst. Video Technol.4
2018 Student motivation towards learning to program
abstract
This Research to Practice Full Paper presents a study on student's motivation towards learning to program. Motivation is a key factor in learning. Hence, stimulating student motivation strategies should be present in any pedagogical approach. This is particularly true in courses where a very active student attitude is fundamental. Introductory programming courses in higher education are a good example, which are known to be difficult for many students. To be successful students need to be motivated, as effort and commitment are necessary to overcome the difficulties many of them experience. In our study we analyzed several motivational aspects separately and then we correlated that information with the marks students obtained in introductory programming courses. We used two questionnaires. The Course Interest Survey (CIS) and the Instructional Materials Motivation Survey (IMMS). We could find some interesting correlations that confirm the importance of different motivational aspects to learning. We found other issues that demand more investigation, in order to create the best context to promote student motivation and learning.
Anabela Jesus Gomes, Wei Ke 0001, Chan-Tong Lam, Maria José Marcelino, António J. Mendes
FIE2
2018 Community Evolution Model for Network Flow Based Multiple Object Tracking
abstract
Multiple object tracking is a research hotspot in the artificial intelligent field, and tracking-by-detection is one of the most popular paradigms in recent years. Among these methods, the network flow based tracker is quite popular due to its computational efficiency and optimality, but it still has one main drawback: Object detection is the processing unit, so high-order information is hard to be taken into consideration directly, and it is usually processed hierarchically, which leads to error propagation. To address this problem, we propose community evolution model for network flow based trackers. We introduce a novel community, which maintains detections and tracklets dynamically. The community allows modeling the connectivities of detections and tracklets jointly, which adaptively incorporates all-level correlations among detections and tracklets, including low-level optical flow, mid-level color histogram, and high-level ranking model. We demonstrate the validity of our method on PETS09 dataset and the MOT17 benchmark, and our method achieves competitive results. Our results on the MOT17 benchmark are available on the website.
Jiahui Chen 0001, Hao Sheng 0001, Yang Zhang 0032, Wei Ke 0001, Zhang Xiong 0001
ICTAI4
2018 Combine Coarse and Fine Cues: Multi-grained Fusion Network for Video-Based Person Re-identification
Chao Li 0001, Lei Liu 0016, Kai Lv 0002, Hao Sheng 0001, Wei Ke 0001
KSEM (1)5
2018 Enhancing Network Flow for Multi-target Tracking with Detection Group Analysis
Chao Li 0001, Kun Qian 0007, Jiahui Chen 0001, Guangtao Xue, Hao Sheng 0001, Wei Ke 0001
KSEM (1)6
2018 W-Shaped Selection for Light Field Super-Resolution
Bing Su 0004, Hao Sheng 0001, Shuo Zhang 0003, Da Yang 0001, Nengcheng Chen, Wei Ke 0001
KSEM (1)6
2018 Particle-mesh coupling in the interaction of fluid and deformable bodies with screen space refraction rendering
abstract
Abstract On the basis of the smoothed particle hydrodynamics and finite element method (FEM) model, we propose a method integrating several improvements for the real‐time simulation of fluid interacting with deformable bodies. We improve the particle neighbor search in smoothed particle hydrodynamics, so that the predefined scene containers are no longer needed. This improvement can also be applied to the simulation of fluid interacting with other materials, such as rigid and soft bodies. We also propose a two‐way coupling method for fluid and deformable bodies, where the particle–mesh interaction is obtained by the ray‐traced collision detection method instead of the proxy/ghost particle generation. By using the forward ray‐tracing method for both velocity and position, we are able to calculate the coupling forces based on the conservation of momentum and kinetic energy in the particle–mesh interaction. We use the screen space fluid rendering for fluid, and on the basis of that, we introduce a screen space refraction rendering method to improve the refraction effect. We implement our method in NVIDIA CUDA and OptiX to make use of the full computational power of a graphics processing unit. The simulation results are analyzed and discussed to show the efficiency of our method.
Ka-Hou Chan, Wei Ke 0001, Sio Kei Im
Comput. Animat. Virtual Worlds2
2017 Fast Binarisation with Chebyshev Inequality
abstract
In order to enhance the binarization result of degraded document images with smudged and bleed-through background, we present a fast binarization technique that applies the Chebyshev theory in the image preprocessing. We introduce the Chebyshev filter which uses the Chebyshev inequality in the segmentation of objects and background. Our result shows that the Chebyshev filter is not only effective, but also simple, robust and easy to implement. Because of its simplicity, our method is sufficiently efficient to process live image sequences in real-time. We have implemented and compared with the Document Image Binarization Contest datasets (H-DIBCO 2014) for testing and evaluation. The experimental outcomes have demonstrated that this method achieved good result in this literature.
Ka-Hou Chan, Sio Kei Im, Wei Ke 0001
DocEng3
2017 A teacher's view about introductory programming teaching and learning - Portuguese and Macanese perspectives
abstract
The difficulties faced by students and teachers in learning and teaching introductory programming has been a research issue over the years. Demotivation is common in many novice programming students, who are not able to cope with the natural difficulties associated to programming learning. It is up to the teacher to find strategies to help students and keep them motivated during the course. The objective of our research was to know more about the pedagogical and motivational strategies used by teachers in the author's institutions to promote programming student's motivation and learning. Some time ago we interviewed a few Portuguese teachers with diversified experiences in programming teaching. Recently we had the opportunity to do the same kind of research with Professors from Macao (China). The problems identified, the teachers' motivation to teach programming, the educational and motivational strategies and the student-teacher relationship were specifically addressed. This paper describes the research done, stressing the Macanese teacher's views and relating them with the views previously expressed by their Portuguese colleagues.
Anabela Jesus Gomes, Wei Ke 0001, Sio Kei Im, Andrew Siu, António J. Mendes, Maria José Marcelino
FIE2
2015 Simulation of Interaction Between Fluid and Deformable Bodies
Ka-Hou Chan, Wei Ke 0001
ICIG (3)3
2014 GEARS: A General and Efficient Algorithm for Rendering Shadows
abstract
Abstract We present a soft shadow rendering algorithm that is general, efficient and accurate. The algorithm supports fully dynamic scenes, with moving and deforming blockers and receivers, and with changing area light source parameters. For each output image pixel, the algorithm computes a tight but conservative approximation of the set of triangles that block the light source as seen from the pixel sample. The set of potentially blocking triangles allows estimating visibility between light points and pixel samples accurately and efficiently. As the light source size decreases to a point, our algorithm converges to rendering pixel accurate hard shadows.
Lili Wang 0006, Shiheng Zhou, Wei Ke 0001, Voicu Popescu
Comput. Graph. Forum3
2014 Fast and compact dynamic data compression based on composite rigid body construction
Lili Wang 0006, Wei Ke 0001, Qinping Zhao
Sci. China Inf. Sci.4
2014 Interactive texture design and synthesis from mesh sketches
Lili Wang 0006, Qinglin Qi, Wei Ke 0001, Aimin Hao
Frontiers Comput. Sci.4
2014 Second-Order Feed-Forward Renderingfor Specular and Glossy Reflections
abstract
The feed-forward pipeline based on projection followed by rasterization handles the rays that leave the eye efficiently: these first-order rays are modeled with a simple camera that projects geometry to screen. Second-order rays however, as, for example, those resulting from specular reflections, are challenging for the feed-forward approach. We propose an extension of the feed-forward pipeline to handle second-order rays resulting from specular and glossy reflections. The coherence of second-order rays is leveraged through clustering, the geometry reflected by a cluster is approximated with a depth image, and the color samples captured by the second-order rays of a cluster are computed by intersection with the depth image. We achieve quality specular and glossy reflections at interactive rates in fully dynamic scenes.
Lili Wang 0006, Naiwen Xie, Wei Ke 0001, Voicu Popescu
IEEE Trans. Vis. Comput. Graph.3
2013 Composite Rigid Body Construction for Fast and Compact Dynamic Data Compression
abstract
Compression of 3D dynamic datasets in remote visualization still remains two challenges. One is low time performance due to the grown data and complex computation of compression algorithm. Another is small compression factor because of dynamic scenes without known equations of their motions. In this paper, we propose a fast and compact compression for 3D dynamic datasets. It accelerates compression with KD-tree construction and node-grid mapping for the dynamic data, which allow parallel rigid body decomposition and merging with disjoint union method. To increase the compression factor, composite rigid body is introduced with consideration of temporary motion consistency among rigid bodies. The results of the experiments show that our algorithm can compress dynamic datasets quickly and obtain high compression factor to reduce limitation of bandwidth.
Lili Wang 0006, Xinwe Zhang, Wei Ke 0001, Qinping Zhao
CAD/Graphics4
2013 A graph-based generic type system for object-oriented programs
Wei Ke 0001, Zhiming Liu 0001, Shuling Wang 0003, Liang Zhao 0022
Frontiers Comput. Sci.1
2012 rCOS: a formal model-driven engineering method for component-based software
Wei Ke 0001, Zhiming Liu 0001, Volker Stolz
Frontiers Comput. Sci. China1
2009 A Graph-Based Operational Semantics of OO Programs
Wei Ke 0001, Zhiming Liu 0001, Shuling Wang 0003, Liang Zhao 0022
ICFEM1