VLDB 2026 Research / reviewers in the wild / expert
Wei Ke 0001
dblp:52/7566-1
· DBLP profile ↗
110ranked-venue papers
3as first author
77since 2021 · last 2027
0000-0003-0952-0961ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 38 · 28 since 2021Artificial intelligence and machine learning · 30 · 22 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 2 first-author · 7 since 2021Computer networks · 10 · 6 since 2021Databases, data management, data science and information retrieval · 8 · 3 since 2021Human-computer interaction and ubiquitous computing · 7 · 5 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 1 since 2021Security and privacy · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | MAGen: Multi-agent smart contract generation with automated testing and verification
Lixue Liu, Wei Ke 0001, Haiyang Chi, Junjian Yan |
Empir. Softw. Eng. | 2 |
| 2026 | Pyramid-Angular-Constraint Network for Light Field Super-ResolutionabstractLight field (LF) cameras record both intensity and directions of light rays in a scene with a single exposure. Due to the trade-off between spatial and angular dimensions, the spatial resolution of LF images is limited, so super-resolution is widely studied. Pixels follow linear coordinate projection across views in LF images. Hence, auxiliary views nearer to the target view are generally more effective for use in super-resolution. In this paper, an LF-pyramid is proposed based on an angular-distance constraint for discriminatively exploiting auxiliary views. From views of different layers in an LF-pyramid, complementary features of different effectiveness can be extracted. However, shapes of LF-pyramids change for target views with different angular positions. To fully exploit an LF-pyramid, we introduce a pyramid-angular-constraint network for LF super-resolution (LF-PACNet). Specifically, to handle an arbitrary number of views in each layer, an intra-pyramid-layer feature extraction module is designed, which treats all views in the same layer equally in complementary information extraction. Then, to deal with an arbitrary number of layers, a recurrent cross-pyramid-layer feature complementation module is constructed, which discriminatively complements the target view with high-frequency details. Extensive experiments on public datasets demonstrate state-of-the-art performance for our method, both visually and numerically, especially for datasets with large disparities. Da Yang 0001, Hao Sheng 0001, Wei Ke 0001, Zhang Xiong 0001 |
Comput. Vis. Media | 4 |
| 2026 | Hard constraints and soft learning dual-graph anomaly detection for industrial processes
Ming-Qing Zhang, Wei Ke 0001, Yang Zhang 0032 |
Eng. Appl. Artif. Intell. | 4 |
| 2026 | Knowledge graph augmented meta-learning with condition-sensitive pseudo-labeling for semi-supervised fault diagnosis under multiple working conditions
Ke-Yu Wu, Yuan Xu 0026, Wei Ke 0001, Yang Zhang 0032, Ming-Qing Zhang |
Expert Syst. Appl. | 4 |
| 2026 | User perception based label layout for efficient target localization in virtual environment
Jian Wu 0033, Shuai Luan, Wei Ke 0001, Lili Wang 0006 |
Int. J. Hum. Comput. Stud. | 4 |
| 2026 | Adaptive semi-supervised meta-learning with pseudo-label for multiple working condition fault diagnosis
Ke-Yu Wu, Yuan Xu 0026, Wei Ke 0001, Yang Zhang 0032, Ming-Qing Zhang |
Neurocomputing | 4 |
| 2026 | The Latens Patronus: Seamless Model Watermarking for Latent Diffusion Model in IoT EnvironmentsabstractWith the rapid development of the Internet of Things (IoT), generative artificial intelligence has been widely applied in various IoT applications. However, the wide adoption of latent diffusion model (LDM) in such IoT scenarios raises severe risks of copyright infringement and model theft due to the lack of effective protection mechanisms. To address this challenge, we propose Latens Patronus, a seamless model watermarking technique for copyright protection of LDM in IoT environments. Unlike existing model watermarking methods, our method does not require additional watermark input and additional parameters for a specialized embedding network, making it more suitable for deployment in real-world IoT applications. Specifically, we design a Watermark Encoder to integrate image watermark into latent features during the generation process and aWatermark Decoder to accordingly extract the watermark from suspicious images accurately. We further introduce a diminishing training strategy that gradually fades out auxiliary supervision signals, eliminating the need for persistent watermark guidance, and adding no extra overhead to the original model. Extensive experiments on multiple LDM variants demonstrate that Latens Patronus outperforms existing watermarking methods in both invisibility and robustness against image-level and model-level attacks. Tong Liu 0021, Wei Ke 0001, Guoheng Huang, Xueyuan Gong, Xiaochen Yuan |
IEEE Internet Things J. | 3 |
| 2026 | Learning Three-Domain Implicit Image Function for Arbitrary-Scale Light Field Super-ResolutionabstractVarious deep learning-based light field image super-resolution methods have attained notable success in recent years. However, most of them focus on encoder design while neglecting the critical role of upsampling process in decoder part. Motivated by the recent progress in single image domain with implicit neural representation, we elaborately propose a spatial-angular-epipolar implicit image function (SAEIIF) in this paper, which can redefine the upsampling process to significantly improve performance and enable arbitrary-scale light field super-resolution. Specifically, it contains two complementary upsampling branches. One branch incorporates spatial implicit image function (SIIF) and angular implicit image function (AIIF) to mine intra-view information in sub-aperture images and inter-view information in macro pixels. The other branch involves epipolar implicit image function (EIIF) to leverage spatial-angular correlation in epipolar plane images. By decomposing SIIF, AIIF and EIIF into horizontal and vertical two-step upsampling to form a perfect match of upsampling scale, SAEIIF introduces a multi-stage feature interaction architecture across two branches to fully merge spatial, angular and epipolar domain information. Furthermore, we optimize feature sampling strategy based on characteristics of sub-aperture images, macro pixels, and epipolar plane images, introducing horizontal-vertical separable local sampling for SIIF and AIIF, as well as dual-source oriented line sampling used for EIIF. The extensive experimental results demonstrate that our SAEIIF can be effectively integrated with most encoders and achieve outstanding performance on both fixed-scale and arbitrary-scale light field spatial super-resolution, angular super-resolution, spatial-angular joint super-resolution. Ruixuan Cong, Hao Sheng 0001, Da Yang 0001, Zhenglong Cui, Weifeng Lyv, Wei Ke 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2026 | LogicMix: Sample mixing data augmentation for multi-label image classification with partial labels
Chak Fong Chong, Jielong Guo, Xu Yang 0010, Wei Ke 0001, Pedro H. Abreu, Yapeng Wang 0001, Sio Kei Im |
Pattern Recognit. | 4 |
| 2026 | Cross-Modal Attention Guided Enhanced Fusion Network for RGB-T TrackingabstractVisual tracking that combines RGB and thermal infrared modalities (RGB-T) aims to utilize the useful information of each modality to achieve more robust object localization. Most existing tracking methods based on convolutional neural networks (CNNs) and Transformers emphasize integrating multi-modal features through cross-modal attention, but ignore the potential exploitability of complementary information learned by cross-modal attention for enhancing modal features. In this paper, we propose a novel hierarchical progressive fusion network based on cross-modal attention guided enhancement for RGB-T tracking. Specifically, the complementary information generated by cross-modal attention implicitly reflects the consistent regions of interest of important information between different modalities, which is used to enhance modal features in a targeted manner. In addition, a modal feature refinement module and a fusion module are designed based on dynamic routing to perform noise suppression and adaptive integration on the enhanced multi-modal features. Extensive experiments on GTOT, RGBT234, LasHeR and VTUAV show that our method has competitive performance compared with recent state-of-the-art methods. Jun Liu 0053, Wei Ke 0001, Shuai Wang 0027, Da Yang 0001, Hao Sheng 0001 |
IEEE Signal Process. Lett. | 2 |
| 2026 | STFAENet: Multiattention Enhanced Spatiotemporal Fusion Network for Traffic Flow PredictionabstractTraffic flow prediction is a fundamental task in intelligent transportation systems (ITS). Due to the influence of urban functional zones and their neighboring regions, traffic flow data exhibit complex spatiotemporal correlations, making it challenging to effectively capture temporal dependencies and spatial structures for accurate prediction. To address this issue, this article proposes a multiattention enhanced spatiotemporal fusion network (STFAENet) based on the TransUNet architecture. STFAENet consists of an encoder, a skip-connection mechanism, and a decoder, and is designed to jointly learn fine-grained local features and global spatiotemporal dependencies in dynamic traffic scenarios. Specifically, an instance pyramid spatial attention (IPSA) module is introduced in the encoder to enhance multiscale spatial feature representation through hybrid normalization and pyramid attention, enabling the extraction of high-resolution fine-grained features. In the skip-connection stage, a spatiotemporal self-attention (STSA) module is embedded to jointly model spatial and temporal dependencies and reduce the semantic gap between the encoder and decoder. The decoder further fuses local details with global contextual information to generate high-resolution prediction results. Experiments on the TaxiBJ and TaxiCQ datasets demonstrate that STFAENet achieves superior prediction performance compared with several baseline methods. Yuan Xu 0017, Chen-Yang Yan, Wei Ke 0001, Chongxing Ji, Yang Zhang 0032, Ming-Qing Zhang |
IEEE Trans. Comput. Soc. Syst. | 5 |
| 2026 | Robust RGB-T Tracking via Multi-Feature Response Adaptive Fusion and Dynamic Selection RecoveryabstractIn RGB and thermal (RGB-T) modalities fusion tracking, the multi-feature responses of each modality contain rich consistency in object localization, which is crucial to enhance tracking robustness. However, existing decision-level fusion paradigms mostly focus on fusing the output of the last layer, ignoring the correlation between multi-feature responses. Moreover, they also lack consideration of tracking failure, which hinders the application of RGB-T tracking in complex environments. To this end, this paper proposes a multi-feature response adaptive fusion model and a dominant-auxiliary dynamic selection recovery mechanism. Specifically, the former achieves joint optimal fusion by mining the correlation between multi-feature responses. The latter flexibly switches between short-term and long-term tracking modes according to the reliability of tracking results, and utilizes the most reliable modality to further improve tracking stability. Experiments on five prevalent RGB-T tracking benchmarks demonstrate the competitive performance of our method compared with the state-of-the-art methods. Jun Liu 0053, Wei Ke 0001, Hao Sheng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | 3D-IRMM-Net: An Interpretable Radiologist-Mimicking 3D Multimodal Network for Lesion Detection in ABUS With LLM-Based Guidance and Uncertainty QuantificationabstractAutomated breast ultrasound (ABUS) has emerged as a promising tool for breast lesion detection, but most existing deep learning models for ABUS lack transparency and fail to align with clinical reasoning processes. This limits their interpretability and hinders clinical adoption. Therefore, we propose 3D-IRMM-Net, a radiologist-mimicking 3D multimodal network. It integrates visual, textual, and semantic cues to emulate hierarchical diagnostic reasoning. At its core, the network integrates two novel components: the 3D Scale-Expert Convolution (3D-SEC) block, which enables parameter-efficient, topology-aware feature routing via a shared graph convolutional layer; and the Adaptive Spatial-Gaussian (ASG) block, which enhances lesion localization in ABUS by modeling spatial dependencies through multi-scale Gaussian attention. In addition, we propose LLM-Vision activator, a natural-language-driven mechanism that uses language to guide 3D visual attention activation and enables 3D spatial-perceptual feature learning via a pretrained 2D VLM, boosting training efficiency. By aligning AI inference with clinical workflows, 3D-IRMM-Net enhances both detection rates and interpretability. Extensive experiments on ABUS2025, TDSC-ABUS2023, and FPHXD-ABUS datasets confirm its state-of-the-art performance across multiple metrics. At an FPPI of 0.5, our method achieved a patient-level detection rates of 89.7±1.1% (internal) and 85.1±2.0% (external); at an FPPI of 4, they rose to 96.1±1.6% and 91.5 ±1.9%, respectively. Furthermore, the model provides principled uncertainty quantification, improving trustworthiness in real-world deployment. These results confirm the effectiveness of 3D-IRMM-Net in bridging the gap between AI-driven detection and clinical practice. Xin Qian 0001, Wei Ke 0001, Luoxi Zhu, Yue Sun 0001, Zhang Xiong 0001, Lingyun Bao, Tao Tan 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | SEM-UCSNet: A Novel Semantic Maps-Guided Compressive Sensing Framework for Underwater ImagesabstractUnderwater images (UWIs) captured by underwater detectors are essential for underwater detection and exploration. The compressive sensing theory (CS) provides a method for recovering images from few measurements, and it has been proven to be suitable for underwater environments with narrow bandwidth and limited communication channel resources, which may have a significant negative impact on the quality of captured UWIs. However, most existing state-of-art CS methods do not take the characteristics of UWIs into account, so their performance is limited in underwater applications. Compared with on-land images, UWIs have the following characteristics: 1) UWIs contain relatively few semantics, with a large amount of similar feature within the same semantics; 2) The importance of different semantics in UWIs is closely related to the underwater imaging model. In this paper, we combine the underwater imaging model and semantic of UWIs with CS task and propose a novel semantic maps-guided CS framework for UWIs, dubbed SEM-UCSNet, which can improve the performance of sampling and reconstruction, especially under extremely low sampling rate. In the sampling stage, a semantic importance analysis module combined with the imaging model is designed to guide the sampling. In the reconstruction process, a graph-based reconstruction strategy guided by semantic maps is proposed to model all features under the same semantic and mine complementarity between them to improve the reconstruction quality. Simultaneously, we introduce GAN into the underwater CS reconstruction task and use sampled features as conditions to make the reconstructed UWIs have richer details. Experimental results on some real-world UWIs datasets have demonstrated the superiority of our SEM-UCSNet on both objective and subjective metrics. Lihao Zhuang, Xiaochen Yuan, Wei Ke 0001, Chan-Tong Lam, Sio Kei Im |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Combining Gaussian and Pixel Representation for Light Field View ReconstructionabstractLight field (LF) benefits various applications due to its rich spatial and angular information. To address the technical limitation in terms of imaging resolution, LF view reconstruction becomes a research hotspot. However, relevant methods mainly focus on pixel representation modeling on image plane but ignore the importance of scene geometry modeling. Inspired by powerful geometry description ability embedded in 3D Gaussian Splatting, we construct a network called LFGaussian to perform generalizable LF view reconstruction in this paper. Specifically, owing to the unique composition of cross-view Gaussian attribute deviation under 4D LF imaging setting, we propose disparity-guided feed-forward 2D Gaussian propagation with novel Gaussian primitive definition, subtly implementing Gaussian unprojection-projection operation in camera parameter-free case. On this basis, we introduce a dual-branch workflow including Gaussian representation rendering and pixel representation upsampling to create features of target views from two different levels, which complement each other to jointly realize geometric structure consistency as well as texture detail consistency across all target views. Besides, for the pursuit of high-efficient and high-quality Gaussian representation rendering, we design sub-sampling Gaussian decoding to alleviate Gaussian redundancy and leverage Gaussian splitting to allocate additional Gaussians for complex geometry regions identified by disparity gradient. Experimental results show that the proposed LFGaussian achieves superior performance compared with state-of-the-art methods on both real-world and synthetic LF datasets, proving the effectiveness of introducing Gaussian representation for LF view reconstruction. Furthermore, our LFGaussian supports arbitrary-scale reconstruction, showing high flexibility for the upsampling scale factor. Ruixuan Cong, Zhenglong Cui, Wei Ke 0001, Weifeng Lyv, Hao Sheng 0001 |
IEEE Trans. Image Process. | 4 |
| 2026 | HRMamba: Fusing Luminance Information for Remote Physiological Measurement in Varied Lighting ConditionsabstractCamera-based photoplethysmography (cbPPG) represents a non-invasive technique for capturing physiological parameters through facial videos, enabling the extraction of vital signs such as heart rate, respiration rate, and blood oxygen saturation without direct physical contact. Existing deep learning methods face two core challenges when dealing with cbPPG: firstly, extracting weak PPG signals from video segments with large spatial and temporal redundancy and understanding their periodic patterns in long contexts; secondly, accurately extracting PPG signals in complex lighting environments, especially in low-light conditions. To address these issues, this paper proposes an end-to-end method based on Mamba, named HRMamba. This method employs temporal difference mamba to process temporal signals and combines bidirectional state space to enable Mamba to robustly understand the scene and learn the periodic patterns of PPG. Furthermore, a luminance post-processing module is designed to extract luminance information from the video without enhancing lighting or altering the original video data, and embed it into the PPG signal. Experimental results demonstrate that HRMamba achieves state-of-the-art performance, and the designed luminance post-processing module can be applied in various lighting environments, significantly enhancing the performance in dark environments without degrading the performance in normal light scenes. Nuoer Long, Wei Ke 0001, Chan-Tong Lam, Tao Tan 0002, Zitong Yu, Yue Sun 0001 |
IEEE J. Biomed. Health Informatics | 3 |
| 2026 | TSNN: A Non-Parametric and Interpretable Framework for Traffic Time Series ForecastingabstractAlthough many complex models were proposed to analyze time series data, some studies have demonstrated remarkable performance with simpler structures. A recent study proposed a non-parametric framework for 3D point cloud classification, which has the potential to be adapted for time series forecasting and enable interpretability. Inspired by the previous works, we present TSNN, a non-parametric and interpretable framework for traffic time series forecasting. TSNN consists of multiple layers that decouple the time series by matching the entries in a memory bank, where the memory bank is constructed using a similar matching process within the training set. It leverages the periodicity in traffic data to enhance forecasting accuracy while maintaining a simple model architecture. The proposed model operates without trainable parameters, preserving its inherent interpretability. In the experiments, TSNN achieves competitive performance compared to the typical deep learning models in four real-world traffic flow datasets. We also visualize the decoupling process to show the effectiveness of the components. Finally, we demonstrate the interpretability of the model and illustrate the contribution of each time step within the memory bank. Our code is available athttps://github.com/pzzzzzm/TSNN_release. Bowie Liu, Haijian Lai, Chan-Tong Lam, Junhao Dong 0004, Benjamin K. Ng, Wei Ke 0001, Sio Kei Im |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2026 | Multi-Granularity Query Network With Adaptive Category Feature Embedding for Behavior RecognitionabstractBehavior recognition is a highly challenging task, particularly in scenarios requiring unified recognition across both human and animal subjects. Most existing approaches primarily focus on single-species datasets or rely heavily on prior information such as species labels, positional annotations, or skeletal keypoints, which limits their applicability in real-world scenarios where species labels may be ambiguous or annotations are insufficient. To address these limitations, we propose a query-based Multi-Granularity Behavior Recognition Network that directly mines cross-species shared spatiotemporal behavior patterns from raw video inputs. Specifically, we design a Multi-Granularity Query module to effectively fuse fine-grained and coarse-grained features, thereby enhancing the model's capability in capturing spatiotemporal dynamics at different granularities. Additionally, we introduce a Category Query Decoder that leverages learnable category query vectors to achieve explicit behavior category modeling and mapping. Without relying on any extra annotations, the proposed method achieves unified recognition of multi-species and multi-category behaviors, setting a new state-of-the-art on the Animal Kingdom dataset and demonstrating strong generalization ability on the Charades dataset. Nuoer Long, Yonghao Dang, Chengpeng Xiong, Shaobin Chen, Tao Tan 0002, Wei Ke 0001, Chan-Tong Lam, Jianqin Yin, Peter H. N. de With, Yue Sun 0001 |
IEEE Trans. Multim. | 7 |
| 2026 | UCGR: Closing the Discretization Gap in Light Field Depth Estimation via Unified Continuous Geometry RepresentationabstractLight field (LF) cameras encode dense spatial-angular information for depth estimation, critical for applications such as 3D reconstruction, refocusing, and virtual reality. However, current deep learning methods for LF depth estimation still face significant challenges due to the discretization gap between the continuous geometry of real scenes and the discrete sampling of digital images, limiting their effectiveness in high-precision application. This gap manifests in two complementary forms: spatial discretization leads to structural ambiguities and information loss, while depth discretization introduces inaccuracies due to fixed, discrete depth sampling. To address these challenges, we propose Unified Continuous Geometry Representation (UCGR), a unified representation that models scene geometry as a continuous field over image coordinates and depth. UCGR treats spatial and depth discretization as two facets of the same problem and realized by two complementary operators: (1) Adaptive Plane Sampling Operator, which learns edge-aware planar priors to preserve geometric details and mitigate spatial discretization. (2) Contextual Depth Correction Operator, which utilizes contextual information for depth correction, ensuring continuous depth estimation and suppressing artifacts. Building on UCGR, we propose a Continuous Geometry Network (CGNet) that collaboratively optimizes both spatial and depth discretization for accurate and consistent LF depth estimation. Extensive experiments on synthetic and real-world LF datasets demonstrate that CGNet achieves state-of-the-art performance, significantly outperforming existing LF depth estimation methods in terms of both accuracy and robustness. Zexin Sun, Tun Wang, Rongshan Chen, Ruixuan Cong, Wei Ke 0001, Hao Sheng 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2026 | Multi-scale feature fusion for cross-modality person re-identification: the MSJLNet approach
Zhixin Tie, Haobiao Fan, Lingbing Tao, Yanbing Chen, Hao Sheng 0001, Wei Ke 0001 |
Vis. Comput. | 6 |
| 2025 | Spatiotemporal Uncertainty-Aware Mamba-Transformer Synergy: Breast Cancer Detection in ABUSabstractBreast cancer significantly affects women's health, making early and accurate diagnosis through ultrasound examination essential. However, automated breast ultrasound(ABUS) lesion detection models encounter challenges, including uncertain noise and difficulties in locating small lesions. This paper proposes SUA-MT, a multi-video object detection network using uncertainty-aware transformers and temporal-spatial mamba. To address the performance degradation caused by low-quality frames and noise interference in input data, we propose the uncertainty-gated transformer decoder (UGTD), which dynamically adjusts attention weights to focus on high-confidence regions while suppressing attention to redundant areas. The spatialtemporal(ST) mamba module is designed to model the long-term dependencies and 3D spatial features of different plane video frames, making full use of temporal and spatial information. In general, SUA-MT not only supports dynamic video length input, but also combines multidimensional video modeling of lesions (transverse, sagittal and coronal plane), making full use of the complementarity of multi-view information and spatiotemporal information to enhance the ability to locate small lesions. Experimental results on a combined set of internal and publicly datasets demonstrate that the SUA-MT method achieves state-of-the-art performance compared to existing video detection approaches. Specifically, SUA-MT attains a mean precision of 80.1 % and mean recall of 84.2 %, providing an efficient and robust solution for lesion detection in ABUS. Xin Qian 0001, Luoxi Zhu, Zhang Xiong 0001, Wei Ke 0001, Lingyun Bao, Tao Tan 0002 |
BIBM | 4 |
| 2025 | GSHOI Denoiser: Denoising Gaussian Hand-Object Interaction for Photorealistic RenderingabstractMany VR/AR applications require the photorealistic rendering of hand-object interactions. Virtual hands are driven by users' hand poses captured via motion tracking to interact with virtual objects. The driven pose can be very noisy due to the constraints of tracking hardware and computation accuracy. This noise may lead to distorted hand poses and penetration artifacts during rendering. In this paper, we introduce the Gaussian Hand-Object Interaction Denoiser, the Gaussian splatting-based hand-object interaction denoising method, which effectively denoises the input twisted and penetrated hand poses to produce photorealistic results. We first propose the innovative joint-to-Gaussian surface representation, which accurately models the spatial relationships between hand skeleton joints and object Gaussians while highlighting hand-object penetrations and generalizing well to new hand poses and objects. Then, we propose a geometry-aware de-penetration algorithm that eliminates penetrations by detecting intersections between skeleton bones and object Gaussians and reposing any penetrated fingers onto the estimated underlying surface of the object. Experiments demonstrate that our method not only effectively reduces hand-object penetration depth but also produces more realistic rendering quality compared to the state-of-the-art methods MANUS+GEARS, MANUS+GeneOH, and$2 \text{DGS}+\text{Gene} \text{OH}$. The user study results show that our method significantly improves the users' visual perceptual experience regarding penetration and stability metrics. Project page: https://github.com/ZhaoLizz/GSHOIDenoiser Lizhi Zhao, Xuequan Lu, Wei Ke 0001, Lili Wang 0006 |
ISMAR | 4 |
| 2025 | FDF-VQVAE: A Frequency Disentanglement and Fusion Learning Framework for Multi-sequence MRI Enhancement
Xinghe Xie, Luyi Han, Yue Sun 0001, Chi Kin Lam, Jian Zheng 0001, Tong Tong 0001, Wei Ke 0001, Chan-Tong Lam, Tao Tan 0002 |
MICCAI (3) | 7 |
| 2025 | Cold-Start Microservice Workload Prediction by Dynamically Annealed Graph-Regularized Matrix Factorization
Xiaoxuan Luo, Hong Shen 0001, Wei Ke 0001 |
PDCAT | 3 |
| 2025 | Pyramid Mamba with Multi-directional Scanning for Light Field Image Super-Resolution
Wenqi Lyu, Wei Ke 0001, Hao Sheng 0001 |
PDCAT | 2 |
| 2025 | Robust Aggregation on Federated Distillation in Adversarial Environments via Entropy and Semantic-Aware Logit Filtering
Hong Shen 0001, Wei Ke 0001, Wenqi Lyu |
PDCAT | 3 |
| 2025 | Enhance Privacy Protection by Reducing Information Exchanges in Multi-agent System Consensus
Hong Shen 0001, Wei Ke 0001 |
PDCAT | 3 |
| 2025 | Stereo matching on epipolar plane image for light field depth estimation via oriented structure
Rongshan Chen, Hao Sheng 0001, Ruixuan Cong, Da Yang 0001, Zhenglong Cui, Wei Ke 0001 |
Eng. Appl. Artif. Intell. | 7 |
| 2025 | Spatio-temporal attention based collaborative local-global learning for traffic flow prediction
Haiyang Chi, Yuhuan Lu 0001, Can Xie, Wei Ke 0001, Bidong Chen |
Eng. Appl. Artif. Intell. | 4 |
| 2025 | Latent temporal smoothness-induced Schatten-p norm factorization for sequential subspace clustering
Zhen-Zhen Zhao, Tong-Wei Lu, Wei Ke 0001, Yang Zhang 0032, Ming-Qing Zhang |
Eng. Appl. Artif. Intell. | 4 |
| 2025 | Regression loss-assisted conditional style generative adversarial network for virtual sample generation with small data in soft sensing
Xue-Yu Zhang, Wei Ke 0001, Ming-Qing Zhang, Yuan Xu 0026 |
Eng. Appl. Artif. Intell. | 3 |
| 2025 | An algorithm for improving lower bounds in dynamic time warping
Yuqi Luo, Xinyi Fang, Wei Ke 0001, Chan-Tong Lam, Sio Kei Im, Luís Paquete |
Expert Syst. Appl. | 3 |
| 2025 | Category-wise Fine-Tuning: Resisting incorrect pseudo-labels in multi-label image classification with partial labels
Chak Fong Chong, Xinyi Fang, Jielong Guo, Pedro H. Abreu, Yapeng Wang 0001, Xu Yang 0010, Wei Ke 0001, Sio Kei Im |
Neurocomputing | 7 |
| 2025 | Bi-branch bidirectional coupled interaction fusion network for multi-retinal diseases diagnosis
Shaobin Chen, Tao Tan 0002, Wei Ke 0001, Xiayu Xu, Yanwu Xu 0004, Chan-Tong Lam, Yue Sun 0001 |
Knowl. Based Syst. | 4 |
| 2025 | Dynamic spatiotemporal graph convolutional network collaborative pre-training learning for traffic flow prediction
Haiyang Chi, Yuhuan Lu 0001, Yirong Zhu, Wei Ke 0001, Hanbin Mao |
Knowl. Based Syst. | 4 |
| 2025 | A GPU-Enabled Framework for Light Field Efficient Compression and Real-Time RenderingabstractReal-time rendering offers instantaneous visual feedback, making it crucial for mixed-reality applications. The light field captures both light intensity and direction in a 3D environment, serving as a data-rich medium to enhance mixed-reality experiences. However, two major challenges remain: 1) current light field rendering techniques are unsuitable for real-time computation, and 2) existing real-time methods cannot efficiently process high-dimensional light field data on GPU platforms. To overcome these challenges, we propose an framework utilizing a compact neural representation of light field data, implemented on a GPU platform for real-time rendering. This framework provides both compact storage and high-fidelity real-time computation. Specifically, we introduce a ray global alignment strategy to simplify the framework and improve practicality. This strategy enables the learning of an optimal embedding for all local rays in a globally consistent way, removing the need for camera pose calculations. To achieve effective compression, the neural light field is employed to map each embedded ray to its corresponding color. To enable real-time rendering, we design a novel super-resolution network to enhance rendering speed. Extensive experiments demonstrate that our framework significantly enhances compression efficiency and real-time rendering performance, achieving nearly 50$\mathbf{\times}$compression ratio and 100 FPS rendering. Mingyuan Zhao 0001, Hao Sheng 0001, Rongshan Chen, Ruixuan Cong, Tun Wang, Zhenglong Cui, Da Yang 0001, Shuai Wang 0027, Wei Ke 0001 |
IEEE Trans. Computers | 9 |
| 2025 | PFL-ALP: Personalized Federated Learning Against Backdoor Attacks via Attention-Based Local PurificationabstractFederated learning (FL) enables collaborative model training with local data privacy preserving, but is vulnerable to backdoor attacks from malicious clients. These attacks can manipulate the global model to produce malicious output when encountering specific triggers. Existing defenses, categorized as server-side and client-side approaches, have limitations such as reliance on auxiliary data availability, susceptibility to inference attacks, and instability under non-independent and identically distributed (Non-IID) data. In response to these challenges, we propose a Personalized Federated Learning via Attention-based Local Purification (PFL-ALP) algorithm, a hybrid defense mechanism integrating server-side dynamic clustering and client-side purification enhanced with personalized model knowledge. This approach effectively mitigates bias introduced by Non-IID data on the server side and further purifies the backdoored model on the client side. Specifically, we employ neural attention distillation (NAD) for model purification and enhance it with personalized model knowledge, extending the effectiveness of NAD in Non-IID FL settings. This design makes PFL-ALP compatible with privacy protocols to mitigate inference attacks. Moreover, we establish a convergence guarantee for PFL-ALP and experimentally validate its superior performance in defending against various backdoor attacks compared to multiple state-of-the-art (SOTA) defenses across three datasets. The results show that even with malicious rates ranging from 30% to 90%, PFL-ALP can reduce the attack success rate by more than 69.4 percentage points, with the reduction in main task accuracy less than 12.4 percentage points. Yifeng Jiang 0008, Xiaochen Yuan, Wei Ke 0001, Chan-Tong Lam, Sio Kei Im |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2025 | A Symmetric Self-Embedding Mechanism for High-Fidelity Image Recovery Against TamperingabstractDigital images are inherently fragile and vulnerable to malicious tampering, significantly compromising their authenticity and integrity. Image recovery is crucial for restoring altered content and preserving the reliability of digital images. Traditional fragile watermarking methods achieve high-quality recovery but fail under post-processing attacks, while existing deep learning-based approaches offer some robustness, yet often produce lower-quality recovered images, typically with a PSNR of around 28 dB. To address these challenges, we propose a novel Symmetric Self-embedding Mechanism for High-Fidelity Image Recovery against tampering (SSEM-HIR), which is capable of restoring tampered images with high quality while maintaining some robustness against common attacks. Unlike existing methods that use the fragility of watermarking solely for tampering localization, SSEM-HIR is the first work to integrate fragility with spatial symmetry, enabling high-quality tampering recovery. Specifically, our SSEM-HIR employs a hierarchical watermark embedding module to embed an inverted version of the original image, utilizing spatial symmetry to retrieve lost information from the extracted watermark. To further improve recovery quality, we design a Dual-branch Region-based Self-Recovery module, where a Spatial-based Watermark Extraction block restores tampered regions using embedded watermark information, while a Frequency-assisted Image Repair block compensates for quality degradation in the untampered area. Extensive experiments show that our method achieves an average PSNR of 34.14 dB under common attack scenarios, including noise addition, image scaling, Gaussian blurring, and no post-processing. This represents an improvement of over 5 dB and 18% in recovered image quality compared to state-of-the-art approaches. Tong Liu 0021, Xiaochen Yuan, Wei Ke 0001, Chan-Tong Lam, Sio Kei Im, Pedro Martins 0003 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2024 | A Parameterized Generative Adversarial Network Using Cyclic Projection for Explainable Medical Image ClassificationsabstractAlthough current data augmentation methods are successful to alleviate the data insufficiency, conventional augmentation are primarily intra-domain while advanced generative adversarial networks (GANs) generate images remaining uncertain, particularly in small-scale datasets. In this paper, we propose a parameterized GAN (ParaGAN) that effectively controls the changes of synthetic samples among domains and highlights the attention regions for downstream classification. Specifically, ParaGAN incorporates projection distance parameters in cyclic projection and projects the source images to the decision boundary to obtain the class-difference maps. Our experiments show that ParaGAN can consistently outperform the existing augmentation methods with explainable classification on two small-scale medical datasets. Xiangyu Xiong, Yue Sun 0001, Xiaohong Liu 0001, Chan-Tong Lam, Tong Tong 0001, Hao Chen 0037, Qinquan Gao, Wei Ke 0001, Tao Tan 0002 |
ICASSP | 8 |
| 2024 | MSFGNet: Multi-Scale Features Gathering Network for Change Detection of Remote Sensing ImagesabstractChange detection is an important research area in remote sensing. To achieve accurate results, it is essential to extract multi-scale spatial information from images while filtering out noise. However, existing models lack this capability. Therefore, Multi-Scale Feature Gathering Network (MSFGNet) is proposed. Within MSFGNet, Bi-Temporal Image Multi-Level Fusion Module (BMF) is utilized to fuse bi-temporal remote sensing images. Additionally, Multi-Receptive Field Features Extraction Module (MRFE) is utilized to extract deep features. Within MRFE, Large Receptive Field Features Extraction Module (LRFE) and Multi-Scale Information Fusion Module (MSIF) are designed, which use large kernel convolution and dilated convolution respectively to capture spatial information with large receptive fields. Furthermore, Cross-Dimension Feature Sifting Fusion Module (CDFSF) is designed to sift noise from various dimensions, fusing valuable information. Across multiple public datasets, MSFGNet consistently achieves the best experimental results. The code can be accessed at https://github.com/juncyan/msfgnet.git. Junqing Huang, Xiaochen Yuan, Chan-Tong Lam, Wei Ke 0001 |
ICME | 4 |
| 2024 | Dual Hypergraph Convolution Networks for Image Forgery Localization
Xiaochen Yuan, Wei Ke 0001, Chan-Tong Lam |
ICPR (22) | 3 |
| 2024 | Enhancing Federated Learning Robustness in Non-IID Data Environments via MMD-Based Distribution Alignment
Hong Shen 0001, Wenqi Lyu, Wei Ke 0001 |
PDCAT | 4 |
| 2024 | Advancing Evasion: Distributed Backdoor Attacks in Federated Learning
Hong Shen 0001, Wei Ke 0001 |
PDCAT | 3 |
| 2024 | Online 3D behavioral tracking of aquatic model organism with a dual-camera system
Zewei Wu, Wei Zhang 0245, Guodong Sun 0001, Wei Ke 0001, Zhang Xiong 0001 |
Adv. Eng. Informatics | 5 |
| 2024 | VVIR-OM: Efficient Object Manipulation in VR with Variable Virtual Interaction RegionabstractManipulating with virtual objects is a fundamental requirement in virtual environments. A good manipulation method needs to consider efficiency, accuracy, and comfort. This paper proposes VVIR-OM, an object manipulation method in virtual reality (VR) based on a variable virtual interaction region (VVIR). A hand interaction hemisphere region (HIHR) is introduced and constructed in real space, where the user is more comfortable manipulating the objects. Then a VVIR is introduced, and an interaction heat volume (IHV) based method is proposed to update VVIR during the process of the manipulation. At last, a mapping algorithm is proposed to map the user hand position in HIHR to the position in VVIR. Two user studies are designed to evaluate the performance of VVIR-OM. Compared to the state-of-the-art methods, VVIR-OM achieves significant improvements in task completion time, manipulation precision, and a significant reduction in fatigue. Moreover, VVIR-OM outperforms other methods in terms of task load and usability without the cost of cybersickness. Qinwen Zheng, Lili Wang 0006, Wei Ke 0001, Sio Kei Im |
Int. J. Hum. Comput. Interact. | 3 |
| 2024 | Object manipulation based on the head manipulation space in VR
Lili Wang 0006, Wei Ke 0001, Sio Kei Im |
Int. J. Hum. Comput. Stud. | 3 |
| 2024 | A simple transformer-based baseline for crowd tracking with Sequential Feature Aggregation and Hybrid Group Training
Zewei Wu, Wei Ke 0001, Zhang Xiong 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2024 | An accurate slicing method for dynamic time warping algorithm and the segment-level early abandoning optimization
Yuqi Luo, Wei Ke 0001, Chan-Tong Lam, Sio Kei Im |
Knowl. Based Syst. | 2 |
| 2024 | Light Field Depth Estimation for Non-Lambertian Objects via Adaptive Cross OperatorabstractLight field (LF) depth estimation is a crucial basis for LF-related applications. Most existing methods are based on the Lambertian assumption and cannot deal with non-Lambertian surfaces represented by transparent objects and mirrors. In this paper, we propose a novel Adaptive-Cross-Operator-based(ACO) depth estimation algorithm for non-Lambertian LF. By analyzing the imaging characteristics of non-Lambertian regions, it is found that the difficulty of depth estimation lies in the photo inconsistency of the center view. Combining with the two-branch structure, we propose ACO with an inter-branch cooperation strategy to adaptively separate depth information with different reflectance coefficients. We discover that the bimodal distribution feature of the operator filtering results can assist in the separation of multi-layer scene information. The first detection branch filters the EPI and implicitly records the severity of multi-layer scene aliasing. According to the identification of bimodal distribution features, the non-Lambertian regions are marked out and the depth of the foreground is estimated. The second branch receives guidance from the first to dynamically adjust the inner weight and infer the background’s depth after weakening the interference from the foreground. Finally, the depth information separation of multi-layer scenes is achieved by extracting the unique X-shaped linear structure. Without the reflection coefficients of the non-Lambertian object, the proposed method can produce high-quality depth estimation under the transparency of 90% to 20%. Experimental results show that the proposed ACO outperforms state-of-the-art LF depth estimation methods in terms of accuracy and robustness. Zhenglong Cui, Hao Sheng 0001, Da Yang 0001, Rongshan Chen, Wei Ke 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Discriminative Feature Learning With Co-Occurrence Attention Network for Vehicle ReIDabstractVehicle Re-Identification (ReID) aims to find images of the same vehicle from different videos. It remains a challenging task in the video analysis field due to the huge appearance discrepancy of the same vehicle in cross-view matching and the subtle difference of different similar vehicles in same-view matching. In this paper, we propose a Co-occurrence Attention Net (CAN) to deal with these two challenges. Specifically, CAN consists of two branches, a main branch and an aware branch. The main branch is in charge of extracting global features that are consistent in most views. This feature encodes holistic information such as color and pose, however, it can not handle cross/same-view hard cases, as shown in Fig.1. Therefore, the aware branch is designed to focus on the local details and viewpoint information, which can become an important complement for those hard cases. Considering that the positions of local areas such as wheels and logos change with the viewpoint, Aware Attention Module is introduced to find the hidden relationship among local areas and seamlessly combine the viewpoint information simultaneously. Then, CAN is trained by a partition-and-reunion-based loss, which can narrow the intra-class distance and increase the inter-class distance. Further, an adaptive co-occurrence view emphasize strategy is adopted to fully utilize the learned features. Experimental results on three widely used datasets including VeRi-776, VehicleID and VERI-Wild demonstrate the effectiveness of our method and competitive performance with other state-of-the-art methods. Hao Sheng 0001, Shuai Wang 0027, Haobo Chen, Da Yang 0001, Wei Ke 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2024 | Distributed Collaborative Object Retrieval With Blockchain-Based Edge ComputingabstractIn the current industrial informatics society, the numerous cameras deployed in the modern city promote the development of various video services, such as security monitoring and object retrieval. However, traditional methods encounter data leakage risks. Some camera owners are reluctant to share their data since the video contains confidential information. Meanwhile, domain diversities between cameras bring obstacles to practical object retrieval applications. To deal with these dilemmas, we propose a blockchain-based collaborative object retrieval (BCOR) system that can protect privacy as much as possible. BCOR includes two core components: multicamera reidentification framework (MC-ReF) and multicamera collaborative chain (M2C-Chain). Specifically, MC-ReF leverages visual relevance attention net (VRANet) to distinguish object identities in edge nodes. Through domain adaptation gradient optimization, VRANet can adapt to different cameras without the need for private camera data. M2C-Chain is responsible for maintaining the security and trust of the system. Through M2C-Chain, the collaboration among different nodes is transferred into a transaction-based manner, which is validated by a deeply integrated consensus. Finally, we implement a prototype system and deploy it into a real-world outdoor scene. The experiments indicate that BCOR achieves 30%–35% average improvement in domain adaptation on mean average precision and Rank-1 indicators. The performance analysis and security experiments also prove the efficiency and stability of BCOR. Shuai Wang 0027, Hao Sheng 0001, Dazhi Yang 0003, Da Yang 0001, Yang Zhang 0032, Wei Ke 0001 |
IEEE Trans. Ind. Informatics | 7 |
| 2024 | Triple Consistency for Transparent Cheating Problem in Light Field Depth EstimationabstractDepth estimation extracting scenes' structural information is a key step in various light field(LF) applications. However, most existing depth estimation methods are based on the Lambertian assumption, which limits the application in non-Lambertian scenes. In this paper, we discover a unique transparent cheating problem for non-Lambertian scenes which can effectively spoof depth estimation algorithms based on photo consistency. It arises because the spatial consistency and the linear structure superimposed on the epipolar plane image form new spurious lines. Therefore, we propose centrifugal consistency and centripetal consistency for separating the depth information of multi-layer scenes and correcting the error due to the transparent cheating problem, respectively. By comparing the distributional characteristics and the number of minimal values of photo consistency and centrifugal consistency, non-Lambertian regions can be efficiently identified and initial depth estimates obtained. Then centripetal consistency is exploited to reject the projection from different layers and to address transparent cheating. By assigning decreasing weights radiating outward from the central view, pixels with a concentration of colors close to the central viewpoint are considered more significant. The problem of underestimating the depth of background caused by transparent cheating is effectively solved and corrected. Experiments on synthetic and real-world data show that our method can produce high-quality depth estimation under the transparency and the reflectivity of 90% to 20%. The proposed triple-consistency-based algorithm outperforms state-of-the-art LF depth estimation methods in terms of accuracy and robustness. Zhenglong Cui, Da Yang 0001, Hao Sheng 0001, Rongshan Chen, Ruixuan Cong, Wei Ke 0001 |
IEEE Trans. Multim. | 7 |
| 2024 | VPRF: Visual Perceptual Radiance Fields for Foveated Image SynthesisabstractNeural radiance fields (NeRF) has achieved revolutionary breakthrough in the novel view synthesis task for complex 3D scenes. However, this new paradigm struggles to meet the requirements for real-time rendering and high perceptual quality in virtual reality. In this paper, we propose VPRF, a novel visual perceptual based radiance fields representation method, which for the first time integrates the visual acuity and contrast sensitivity models of human visual system (HVS) into the radiance field rendering framework. Initially, we encode both the appearance and visual sensitivity information of the scene into our radiance field representation. Then, we propose a visual perceptual sampling strategy, allocating computational resources according to the HVS sensitivity of different regions. Finally, we propose a sampling weight-constrained training scheme to ensure the effectiveness of our sampling strategy and improve the representation of the radiance field based on the scene content. Experimental results demonstrate that our method renders more efficiently, with higher PSNR and SSIM in the foveal and salient regions compared to the state-of-the-art FoV-NeRF. The results of the user study confirm that our rendering results exhibit high-fidelity visual perception. Jian Wu 0033, Runze Fan, Wei Ke 0001, Lili Wang 0006 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2024 | SMigraPH: a perceptually retained method for passive haptics-based migration of MR indoor scenes
Qixiang Ma, Lili Wang 0006, Wei Ke 0001, Sio Kei Im |
Vis. Comput. | 3 |
| 2024 | Joint attribute soft-sharing and contextual local: a multi-level features learning network for person re-identification
Wangmeng Wang, Yanbing Chen, Dengwen Wang, Zhixin Tie, Linbing Tao, Wei Ke 0001 |
Vis. Comput. | 6 |
| 2023 | Category-Wise Fine-Tuning for Image Multi-label Classification with Partial Labels
Chak Fong Chong, Xu Yang 0010, Tenglong Wang, Wei Ke 0001, Yapeng Wang 0001 |
ICONIP (11) | 4 |
| 2023 | Light Field Image Super-Resolution via Global-View Information Adaption and Angular Attention Fusion
Wei Zhang 0245, Wei Ke 0001, Hao Sheng 0001 |
ICONIP (11) | 2 |
| 2023 | Locomotion-aware Foveated RenderingabstractOptimizing rendering performance improves the user's immersion in virtual scene exploration. Foveated rendering uses the features of the human visual system (HVS) to improve rendering performance without sacrificing perceptual visual quality. We collect and analyze the viewing motion of different locomotion methods, and describe the effects of these viewing motions on HVS's sensitivity, as well as the advantages of these effects that may bring to foveated rendering. Then we propose the locomotion-aware foveated rendering method (LaFR) to further accelerate foveated rendering by leveraging the advantages. In LaFR, we first introduce the framework of LaFR. Secondly, we propose an eccentricity-based shading rate controller that provides the shading rate control of the given region in foveated rendering. Thirdly, we propose a locomotion-aware log-polar mapping method, which controls the foveal average shading rate, the peripheral shading rate decrease speed, and the overall shading quantity with the locomotion-aware coefficients based on the eccentricity-based shading rate controller. LaFR achieves similar perceptual visual quality as the conventional foveated rendering while achieving up to 1.6× speedup. Compared with the full resolution rendering, LaFR achieves up to 3.8× speedup. Xuehuai Shi, Lili Wang 0006, Jian Wu 0033, Wei Ke 0001, Chan-Tong Lam |
VR | 4 |
| 2023 | Wearable Real-time Air-writing System Employing KNN and Constrained Dynamic Time WarpingabstractIn the digital world, gesture recognition plays a crucial role in human-computer interaction (HCI). In this paper, we propose an innovative wearable air-writing system that allows users to write the English alphabet and Arabic numerals in free space without using any predefined gestures or rules. Based on an Inertial Measurement Unit (IMU), the proposed air-writing wearable device uses the constrained dynamic time warping (cDTW) algorithm for the distance measure and K-nearest neighbors (KNN) as the classifier. In addition, to increase the recognition accuracy and meet HCI requirements, we develop a novel method that allows users to rapidly switch to correct recognition results when the initial results are erroneous. In the experiment, the accuracy rate is 88.9% for the alphabet and 10 decimal digits in the user-dependent condition, and the recognition is in real-time, consuming only 0.427s for each character, which is superior to many other approaches that employ classic DTW or FastDTW. With the proposed HCI design, character input accuracy of over 95% can be obtained in about 1 second. We also simulated the application scenarios of Parkinson’s disease patients and obtained a high accuracy rate of 85.4%. Besides, we explored the variety of K values in KNN and w values in cDTW, and propose a multi-template system that gives new optimization directions for the KNN-cDTW algorithm. Yuqi Luo, Wei Ke 0001, Chan-Tong Lam |
WCNC | 2 |
| 2023 | MFSRNet: spatial-angular correlation retaining for light field super-resolution
Hao Sheng 0001, Da Yang 0001, Zhenglong Cui, Ruixuan Cong, Wei Ke 0001 |
Appl. Intell. | 6 |
| 2023 | Light field super-resolution using complementary-view feature attentionabstractLight field (LF) cameras record multiple perspectives by a sparse sampling of real scenes, and these perspectives provide complementary information. This information is beneficial to LF super-resolution (LFSR). Compared with traditional single-image super-resolution, LF can exploit parallax structure and perspective correlation among different LF views. Furthermore, the performance of existing methods are limited as they fail to deeply explore the complementary information across LF views. In this paper, we propose a novel network, called the light field complementary-view feature attention network (LF-CFANet), to improve LFSR by dynamically learning the complementary information in LF views. Specifically, we design a residual complementary-view spatial and channel attention module (RCSCAM) to effectively interact with complementary information between complementary views. Moreover, RCSCAM captures the relationships between different channels, and it is able to generate informative features for reconstructing LF images while ignoring redundant information. Then, a maximum-difference information supplementary branch (MDISB) is used to supplement information from the maximum-difference angular positions based on the geometric structure of LF images. This branch also can guide the process of reconstruction. Experimental results on both synthetic and real-world datasets demonstrate the superiority of our method. The proposed LF-CFANet has a more advanced reconstruction performance that displays faithful details with higher SR accuracy than state-of-the-art methods. Wei Zhang 0245, Wei Ke 0001, Da Yang 0001, Hao Sheng 0001, Zhang Xiong 0001 |
Comput. Vis. Media | 2 |
| 2023 | Hybrid Motion Model for Multiple Object Tracking in Mobile DevicesabstractFor an intelligent transportation system, multiple object tracking (MOT) is more challenging from the traditional static surveillance camera to mobile devices of the Internet of Things (IoT). To cope with this problem, previous works always rely on additional information from multivision, various sensors, or precalibration. Only based on a monocular camera, we propose a hybrid motion model to improve the tracking accuracy in mobile devices. First, the model evaluates camera motion hypotheses by measuring optical flow similarity and transition smoothness to perform robust camera trajectory estimation. Second, along the camera trajectory, smooth dynamic projection is used to map objects from image to world coordinate. Third, to deal with trajectory motion inconsistency, which is caused by occlusion and interaction of long time interval, tracklet motion is described by the multimode motion filter for adaptive modeling. Fourth, in tracklets association, we propose a spatiotemporal evaluation mechanism, which achieves higher discriminability in motion measurement. Experiments on MOT15, MOT17, and KITTI benchmarks show that our proposed method improves the trajectory accuracy, especially in mobile devices and our method achieves competitive results over other state-of-the-art methods. Yubin Wu, Hao Sheng 0001, Yang Zhang 0032, Shuai Wang 0027, Zhang Xiong 0001, Wei Ke 0001 |
IEEE Internet Things J. | 6 |
| 2023 | Uncertainty-guided joint attention and contextual relation network for person re-identification
Dengwen Wang, Yanbing Chen, Wangmeng Wang, Zhixin Tie, Xian Fang, Wei Ke 0001 |
J. Vis. Commun. Image Represent. | 6 |
| 2023 | AEA-Net:Affinity-supervised entanglement attentive network for person re-identification
Dengwen Wang, Yanbing Chen, Lingbing Tao, Chentao Hu, Zhixin Tie, Wei Ke 0001 |
Pattern Recognit. Lett. | 6 |
| 2022 | Group Guided Data Association for Multiple Object Tracking
Yubin Wu, Hao Sheng 0001, Shuai Wang 0027, Yang Liu 0088, Zhang Xiong 0001, Wei Ke 0001 |
ACCV (7) | 6 |
| 2022 | Online 3D Reconstruction of Zebrafish Behavioral Trajectories within A Holistic PerspectiveabstractRecording activities of zebrafish is a fundamental task in biological research that aims to accurately track individuals and recover their real-world movement trajectories from multiple viewpoint videos. In this paper, we propose a novel online tracking solution based on a holistic perspective that leverages the correlation of appearance and location across views. It first reconstructs the 3D coordinates of targets frame by frame and then tracks them directly in 3D space instead of a 2D image plane. However, it is not trivial to implement such a solution which requires the association of targets across views and neighboring frames under occlusion and parallax distortion. To cope with that, we propose the view-invariant feature representation and the Kalman filter-based 3D state estimation, and combine the advantages of both to generate robust 3D trajectories. Extensive experiments on public datasets verify the efficiency and effectiveness of the approach. Zewei Wu, Wei Ke 0001, Wei Zhang 0245, Zhang Xiong 0001 |
BIBM | 2 |
| 2022 | Transformation from MVC Applications to Smart ContractsabstractA smart contract is a program running on a blockchain platform. Smart devices send data to smart contracts or change their own status based on smart contracts. Businesses want to use smart devices and smart contracts to streamline workflow because smart contracts reduce the need in trusted intermediators and cut down enforcement costs. However, developing smart contract applications is challenging due to different memory models, different interaction models, and a dearth of supporting tools and libraries. When developers are asked to reimplement conventional applications to smart contracts, it is thus ideal to automatically transform them, avoiding manual labor, as well as ensuring reliability and security. This paper contributes a set of rules to transform conventional applications in the Model-View-Controller (MVC) pattern into smart contracts running on Hyperledger Fabric, a blockchain platform preferred by businesses. Major transformations are performed in the model, while in the controller, model calls are replaced by smart contract calls. The source application and the target smart contract are all written in Java. Our rules add read-your-writes consistency that Hyperledger Fabric does not natively support. Runtime pre- and post-condition checking in the original application is supported in the transformed smart contract. We evaluated our rules on CoCoME and other MVC applications, and all code ran correctly and passed unit tests. Qiqi Gu 0001, Wei Ke 0001, Yilong Yang 0001 |
EUC | 2 |
| 2022 | Partition and Reunion: A Viewpoint-Aware Loss for Vehicle Re-IdentificationabstractVehicle Re-Identification (ReID) aims to retrieve images of vehicles with the same identity from different scenarios. It is a challenging task due to the large intra-identity discrepancy caused by viewpoint variations and the subtle inter-identity difference produced by similar appearances. In this paper, we propose a Viewpoint-Aware Loss (VAL) function to deal with these challenges. Specifically, we propose partition and reunion operations in VAL, which significantly shrinks the intra-identity distance and acquires viewpoint-invariant representations. In addition, we embed a multi-decision boundary mechanism in VAL. It contributes to enlarging the inter-identity distance. A comprehensive evaluation on two benchmarks shows the superiority of our method in contrast to a series of existing state-of-the-arts. Haobo Chen, Yang Liu 0088, Wei Ke 0001, Hao Sheng 0001 |
ICIP | 4 |
| 2022 | Mask Guided Spatial-Temporal Fusion Network for Multiple Object TrackingabstractMulti-object trackers make the association almost perfectly when no occlusion occurred between two or more targets. However, it is hard to extract reliable features on account of partial occlusion caused by a nearby object, which often leads to tracking failure. In this paper, we utilize mask to guide attention of the neural network in order to focus on the visible part of the target and design a tracklet-level feature extraction method. Then, a tracking framework is proposed based on a mask guided fusion network and multi-hypothesis tracking algorithm. Comprehensive evaluation on the MOT17 dataset shows that our approach achieves competitive results. Shuangye Zhao, Yubin Wu, Shuai Wang 0027, Wei Ke 0001, Hao Sheng 0001 |
ICIP | 4 |
| 2022 | Data Association with Graph Network for Multi-Object Tracking
Yubin Wu, Hao Sheng 0001, Shuai Wang 0027, Yang Liu 0088, Wei Ke 0001, Zhang Xiong 0001 |
KSEM (1) | 5 |
| 2022 | A Local Rotation Transformation Model for Vehicle Re-Identification
Yanbing Chen, Wei Ke 0001, Hao Sheng 0001, Zhang Xiong 0001 |
WASA (1) | 2 |
| 2022 | Enhancing Efficiency and Quality of Image Caption Generation with CARU
Xuefei Huang, Wei Ke 0001, Hao Sheng 0001 |
WASA (2) | 2 |
| 2022 | Local perspective based synthesis for vehicle re-identification: A transformation state adversarial method
Yanbing Chen, Wei Ke 0001, Hong Lin 0006, Chan-Tong Lam, Kai Lv 0002, Hao Sheng 0001, Zhang Xiong 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2022 | Multiple classifier for concatenate-designed neural network
Ka-Hou Chan, Sio Kei Im, Wei Ke 0001 |
Neural Comput. Appl. | 3 |
| 2021 | Near-Realtime Face Mask Wearing Recognition Based on Deep LearningabstractCOVID-19 pandemic has led to serious economic and life losses. Face Masks serve as first infection barrier when used in public spaces. In this paper, we propose a new near-realtime method to automatically recognize face mask wearing that combines human posture recognition with convolutional neural network (CNN). We use the power of human posture recognition to perform background filtering and spatial reduction in the original images. The outcome is then used by a trained CNN model to identify if the subject is wearing a mask. We exploit Openpose to identify the skeleton of human body and locate the facial region thus spatially reducing the area to be processed by the CNN framework. We then adopt supervised learning approach to detect if a face mask is present. The CNN is trained using images, cropped to the supposed face mask covered region. This approach led to a substantial reduction in neural network complexity yet improving the recognition accuracy. The system has been evaluated in a multitude of scenarios using images taken in public places at different time of day and with different angles. Overall, our system achieves a recognition accuracy of 95.8% and 94.6% in daytime and nighttime respectively. Hong Lin 0006, Rita Tse, Su-Kit Tang, Yanbing Chen, Wei Ke 0001, Giovanni Pau 0001 |
CCNC | 5 |
| 2021 | A Neural Architecture for Detecting Identifier Renaming from Diff
Qiqi Gu 0001, Wei Ke 0001 |
IDEAL | 2 |
| 2021 | Combining Pose Invariant and Discriminative Features for Vehicle ReidentificationabstractVehicle reidentification, aiming at identifying vehicles across images, has drawn a lot of attention and has made significant achievements in recent years. However, vehicle reidentification remains a challenging task caused by severe appearance changes due to different orientations. In practice, the result of reidentification is greatly influenced by the pose of vehicles, and we call this influence as a pose barrier problem. One way to address the pose barrier problem is to train a feature representation that is invariant for various vehicle poses. To this end, we present pose robust features (PRFs) that contains two components: 1) pose-invariant features (PIFs) and 2) pose discriminative features (PDFs). On the one hand, PIF is the expert in exploring the overall characteristic of vehicles. When training PIF, we adopt an identity classifier as well as an orientation classifier. In addition, an adversarial loss is deployed in the PIF network. On the other hand, we design a PDF network, which has a similar architecture to the PIF network but can distinguish the difference between local details. The difference between PDF and PIF is that the network of training PDF does not apply the adversarial loss. Finally, by combining PIF and PDF, PRF has the advantages of the two features and can alleviate the influence of the pose barrier problem. Experiments are conducted on the VeRi-776 and VehicleID data sets. We show that PIF and PDF are complementary and that PRF produces competitive performance compared with state-of-the-art approaches. Hao Sheng 0001, Kai Lv 0002, Yang Liu 0088, Wei Ke 0001, Weifeng Lyu, Zhang Xiong 0001, Wei Li 0022 |
IEEE Internet Things J. | 4 |
| 2020 | Variable-Depth Convolutional Neural Network for Text Classification
Ka-Hou Chan, Sio Kei Im, Wei Ke 0001 |
ICONIP (5) | 3 |
| 2020 | CARU: A Content-Adaptive Recurrent Unit for the Transition of Hidden State in NLP
Ka-Hou Chan, Wei Ke 0001, Sio Kei Im |
ICONIP (1) | 2 |
| 2020 | Camera Style Guided Feature Generation for Person Re-identification
Hantao Hu, Yang Liu 0088, Kai Lv 0002, Yanwei Zheng, Wei Zhang 0245, Wei Ke 0001, Hao Sheng 0001 |
WASA (1) | 6 |
| 2020 | A Dual Scale Matching Model for Long-Term Association
Yubin Wu, Shuai Wang 0027, Yang Zhang 0032, Yanbing Chen, Wei Ke 0001, Hao Sheng 0001 |
WASA (1) | 6 |
| 2020 | Mining Hard Samples Globally and Efficiently for Person ReidentificationabstractPerson reidentification (ReID) is an important application of Internet of Things (IoT). ReID recognizes pedestrians across camera views at different locations and time, which is usually treated as a ranking task. An essential part of this task is the hard sample mining. Technically, two strategies could be employed, i.e., global hard mining and local hard mining. For the former, hard samples are mined within the entire training set, while for the latter, it is done in mini-batches. In literature, most existing methods operate locally. Examples include batch-hard sample mining and semihard sample mining. The reason for the rare use of global hard mining is the high computational complexity. In this article, we argue that global mining helps to find harder samples that benefit model training. To this end, this article introduces a new system to: 1) efficiently mine hard samples (positive and negative) from the entire training set and 2) effectively use them in training. Specifically, a ranking list network coupled with a multiplet loss is proposed. On the one hand, the multiplet loss makes the ranking list progressively created to avoid the time-consuming initialization. On the other hand, the multiplet loss aims to make effective use of the hard and easy samples during training. In addition, the ranking list makes it possible to globally and effectively mine hard positive and negative samples. In the experiments, we explore the performance of the global and local sample mining methods, and the effects of the semihard, the hardest, and the randomly selected samples. Finally, we demonstrate the validity of our theories using various public data sets and achieve competitive results via a quantitative evaluation. Hao Sheng 0001, Yanwei Zheng, Wei Ke 0001, Dongxiao Yu, Xiuzhen Cheng, Weifeng Lyu, Zhang Xiong 0001 |
IEEE Internet Things J. | 3 |
| 2020 | Multiplex Labeling Graph for Near-Online Tracking in Crowded ScenesabstractIn recent years, the demand for intelligent devices related to the Internet of Things (IoT) is rapidly increasing. In the field of computer vision, many algorithms have been preinstalled in IoT devices to achieve higher efficiency, such as face recognition, area detection, target tracking, etc. Tracking is an important but complex task that needs high efficiency solutions in real applications. There is a common assumption that detection can only represent one pedestrian to describe nonoverlapping in physical space. In fact, the pixels of the image do not exactly correspond to the positions in the real world. In order to overcome the limitation of this assumption, we remove this unreasonable assumption and present a novel idea that each detector response can have multiple labels to describe different targets at the same time. Therefore, we propose a graph-based method for near-online tracking in this article. We introduce a detection multiplexing method for tracking in the monocular image and propose a multiplex labeling graph (MLG) model. Each node in MLG has the ability to represent multiple targets. In addition, we improve the shortage of graph-based trackers in using temporal features. We construct long short-term memory networks to model motion and appearance features for MLG optimization. On the public multiobject tracking challenge benchmark, our near-online method gains satisfactory efficiency and achieves state-of-the-art results without additional private detection as well. Yang Zhang 0032, Hao Sheng 0001, Yubin Wu, Shuai Wang 0027, Wei Ke 0001, Zhang Xiong 0001 |
IEEE Internet Things J. | 5 |
| 2020 | Hypothesis Testing Based Tracking With Spatio-Temporal Joint Interaction ModelingabstractData association is one of the key research in tracking-by-detection framework. Due to frequent interactions among targets, there are various relationships among trajectories in crowded scenes which leads to problems in data association, such as association ambiguity, association omission, etc. To handle these problems, we propose hypothesis-testing based tracking (HTBT) framework to build potential associations between target by constructing and testing hypotheses. In addition, a spatio-temporal interaction graph (STIG) model is introduced to describe the basic interaction patterns of trajectories and test the potential hypotheses. Based on network flow optimization, we formulate offline tracking as a MAP problem. Experimental results show that our tracking framework improves the robustness of tracklet association when detection failure occurs during tracking. On the public MOT16, MOT17 and MOT20 benchmark, our method achieves competitive results compared with other state-of-the-art methods. Hao Sheng 0001, Yang Zhang 0032, Yubin Wu, Shuai Wang 0027, Weifeng Lyu, Wei Ke 0001, Zhang Xiong 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2020 | Long-Term Tracking With Deep Tracklet AssociationabstractRecently, most multiple object tracking (MOT) algorithms adopt the idea of tracking-by-detection. Relevant research shows that the performance of the detector obviously affects the tracker, while the improvement of detector is gradually slowing down in recent years. Therefore, trackers using tracklet (short trajectory) are proposed to generate more complete trajectories. Although there are various tracklet generation algorithms, the fragmentation problem still often occurs in crowded scenes. In this paper, we introduce an iterative clustering method that generates more tracklets while maintaining high confidence. Our method shows robust performance on avoiding internal identity switch. Then we propose a deep association method for tracklet association. In terms of motion and appearance, we construct motion evaluation network (MEN) and appearance evaluation network (AEN) to learn long-term features of tracklets for association. In order to explore more robust features of tracklets, a tracklet-based training mechanism is also introduced. Tracklet groups are used as the input of the networks instead of discrete detections. Experimental results show that our training method enhances the performance of the networks. In addition, our tracking framework generates more complete trajectories while maintaining the unique identity of each target as the same time. On the latest MOT 2017 benchmark, we achieve state-of-the-art results. Yang Zhang 0032, Hao Sheng 0001, Yubin Wu, Shuai Wang 0027, Weifeng Lyu, Wei Ke 0001, Zhang Xiong 0001 |
IEEE Trans. Image Process. | 6 |
| 2020 | Automated Prototype Generation From Formal Requirements ModelabstractPrototyping is an effective and efficient way of requirements validation to avoid introducing errors in the early stage of software development. However, manually developing a prototype of a software system requires additional efforts, which would increase the overall cost of software development. In this article, we present an approach with a developed tool RM2PT to automated prototype generation from formal requirements models for requirements validation. A requirements model consists of a use case diagram, a conceptual class diagram, use case definitions specified by system sequence diagrams, and the contracts of their system operations. A system operation contract is formally specified by a pair of pre and postconditions in object constraint language. We propose a method with a set of transformation rules to decompose a contract into executable parts and nonexecutable parts. An executable part can be automatically transformed into a sequence of primitive operations by applying their corresponding rules, and a nonexecutable part is not transformable with the rules. The tool RM2PT provides a mechanism for developers to develop a piece of program for each nonexecutable part manually, which can be plugged into the generated prototype source code automatically. We have conducted four case studies with over 50 use cases. The experimental result shows that the 93.65% system operations are executable, and only 6.35% are nonexecutable, which can be implemented by developers manually or invoking the third-party application programming interface (APIs). Overall, the result is satisfactory. Each 1 s generated prototype of four case studies requires approximate one day's manual implementation by a skilled programmer. The proposed approach with the developed computer-aided software engineering tool can be applied to the software industry for requirements engineering. Yilong Yang 0001, Wei Ke 0001, Zhiming Liu 0001 |
IEEE Trans. Reliab. | 3 |
| 2019 | Image resizing enhancement with DCT coefficientsabstractAccording to the principle of scale transformation, signal expansion in the time domain corresponds to compression in the frequency domain, so that information energy is concentrated in the low frequency part. In this paper, an image enlargement algorithm based on the Discrete Cosine Transform (DCT) is proposed, which preserves the low frequency of the image and combines the corresponding enhancement coefficients to realize the resizing operation in the DCT domain. This paper also reach to proving of the enhancement value determinant. Then comparing the experiment on the scaled image with other interpolation algorithms, the result shows that our algorithm performs better than other methods. This method can also be carried out during DCT transformation, and is easier to implement than other methods. Ka-Hou Chan, Sio Kei Im, Wei Ke 0001 |
ICMV | 3 |
| 2019 | Pedestrian Similarity Extraction to Improve People Counting AccuracyabstractCurrent state-of-the-art single shot object detection pipelines, composed by an object detector such as Yolo, generate multiple detections for each object, requiring a post-processing Non-Maxima Suppression (NMS) algorithm to remove redundant detections. However, this pipeline struggles to achieve high accuracy, particularly in object counting applications, due to a trade-off between precision and recall rates. A higher NMS threshold results in fewer detections suppressed and, consequently, in a higher recall rate, as well as lower precision and accuracy. In this paper, we have explored a new pedestrian detection pipeline which is more flexible, able to adapt to different scenarios and with improved precision and accuracy. A higher NMS threshold is used to retain all true detections and achieve a high recall rate for different scenarios, and a Pedestrian Similarity Extraction (PSE) algorithm is used to remove redundant detentions, consequently improving counting accuracy. The PSE algorithm significantly reduces the detection accuracy volatility and its dependency on NMS thresholds, improving the mean detection accuracy for different input datasets. Xu Yang 0010, José Gaspar, Wei Ke 0001, Chan-Tong Lam, Yanwei Zheng, Weng Hong Lou, Yapeng Wang 0001 |
ICPRAM | 3 |
| 2019 | Deep Learning for Web Services ClassificationabstractAutomated service classification plays a crucial role in service discovery, selection, and composition. Machine learning has been used for service classification in recent years. However, the performance of conventional machine learning methods highly depends on the quality of manual feature engineering. In this paper, we present a deep neural network to automatically abstract low-level representation of service description to high-level features without feature engineering and then predict service classification on 50 service categories. To demonstrate the effectiveness of our approach, we conduct a comprehensive experimental study by comparing 10 machine learning methods on 10,000 real-world web services. The result shows that the proposed deep neural network can achieve higher accuracy than other machine learning methods. Yilong Yang 0001, Wei Ke 0001, Weiru Wang 0001 |
ICWS | 2 |
| 2019 | Spatio-Temporal Correlation Graph for Association Enhancement in Multi-object Tracking
Hao Sheng 0001, Yang Zhang 0032, Yubin Wu, Jiahui Chen 0001, Wei Ke 0001 |
KSEM (1) | 6 |
| 2019 | RM2PT: Requirements Validation through Automatic PrototypingabstractPrototyping is an effective and efficient way of requirements validation to avoid introducing errors in the early stage of software development. Our previous work presents a tool RM2PT to automatically generate prototypes from requirements models. The stakeholders can easily check whether the requirements reflect their real needs by investigating the executions of use cases in the generated prototypes. However, the conflict and contradictory of the requirements are hard to be discovered. In this paper, we enhance RM2PT by introducing consistency checking and state observations in the generated prototypes. Requirements inconsistency can be automatically detected and further fixed through carefully analyzing the contracts of system operations and system state observations. We have conducted four case studies with over 50 use cases. The experimental result shows that 107 requirements inconsistency are founded in requirements validations. Overall, the result is satisfiable, and the enhanced RM2PT can be further applied to the software industry for requirements validation. The tool can be downloaded at http://rm2pt.mydreamy.net and a demo video casting its features is at https://youtu.be/Y7GNa57WGfA. Yilong Yang 0001, Wei Ke 0001 |
RE | 2 |
| 2019 | An easy-to-hard learning strategy for within-image co-saliency detection
Shaoyue Song, Hongkai Yu, Zhenjiang Miao, Dazhou Guo, Wei Ke 0001, Cong Ma 0004, Song Wang 0002 |
Neurocomputing | 5 |
| 2019 | Iterative Multiple Hypothesis Tracking With Tracklet-Level AssociationabstractThis paper proposes a novel iterative maximum weighted independent set (MWIS) algorithm for multiple hypothesis tracking (MHT) in a tracking-by-detection framework. MHT converts the tracking problem into a series of MWIS problems across the tracking time. Previous works solve these NP-hard MWIS problems independently without the use of any prior information from each frame, and they ignore the relevance between adjacent frames. In this paper, we iteratively solve the MWIS problems by using the MWIS solution from the previous frame rather than solving the problem from scratch each time. First, we define five hypothesis categories and a hypothesis transfer model, which explicitly describes the hypothesis relationship between adjacent frames. We also propose a polynomial-time approximation algorithm for the MWIS problem in MHT. In addition to that, we present a confident short tracklet generation method and incorporate tracklet-level association into MHT, which further improves the computational efficiency. Our experiments on both MOT16 and MOT17 benchmarks show that our tracker outperforms all the previously published tracking algorithms on both MOT16 and MOT17 benchmarks. Finally, we demonstrate that the polynomial-time approximate tracker reaches nearly the same tracking performance. Hao Sheng 0001, Jiahui Chen 0001, Yang Zhang 0032, Wei Ke 0001, Zhang Xiong 0001, Jingyi Yu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2018 | Student motivation towards learning to programabstractThis Research to Practice Full Paper presents a study on student's motivation towards learning to program. Motivation is a key factor in learning. Hence, stimulating student motivation strategies should be present in any pedagogical approach. This is particularly true in courses where a very active student attitude is fundamental. Introductory programming courses in higher education are a good example, which are known to be difficult for many students. To be successful students need to be motivated, as effort and commitment are necessary to overcome the difficulties many of them experience. In our study we analyzed several motivational aspects separately and then we correlated that information with the marks students obtained in introductory programming courses. We used two questionnaires. The Course Interest Survey (CIS) and the Instructional Materials Motivation Survey (IMMS). We could find some interesting correlations that confirm the importance of different motivational aspects to learning. We found other issues that demand more investigation, in order to create the best context to promote student motivation and learning. Anabela Jesus Gomes, Wei Ke 0001, Chan-Tong Lam, Maria José Marcelino, António J. Mendes |
FIE | 2 |
| 2018 | Community Evolution Model for Network Flow Based Multiple Object TrackingabstractMultiple object tracking is a research hotspot in the artificial intelligent field, and tracking-by-detection is one of the most popular paradigms in recent years. Among these methods, the network flow based tracker is quite popular due to its computational efficiency and optimality, but it still has one main drawback: Object detection is the processing unit, so high-order information is hard to be taken into consideration directly, and it is usually processed hierarchically, which leads to error propagation. To address this problem, we propose community evolution model for network flow based trackers. We introduce a novel community, which maintains detections and tracklets dynamically. The community allows modeling the connectivities of detections and tracklets jointly, which adaptively incorporates all-level correlations among detections and tracklets, including low-level optical flow, mid-level color histogram, and high-level ranking model. We demonstrate the validity of our method on PETS09 dataset and the MOT17 benchmark, and our method achieves competitive results. Our results on the MOT17 benchmark are available on the website. Jiahui Chen 0001, Hao Sheng 0001, Yang Zhang 0032, Wei Ke 0001, Zhang Xiong 0001 |
ICTAI | 4 |
| 2018 | Combine Coarse and Fine Cues: Multi-grained Fusion Network for Video-Based Person Re-identification
Chao Li 0001, Lei Liu 0016, Kai Lv 0002, Hao Sheng 0001, Wei Ke 0001 |
KSEM (1) | 5 |
| 2018 | Enhancing Network Flow for Multi-target Tracking with Detection Group Analysis
Chao Li 0001, Kun Qian 0007, Jiahui Chen 0001, Guangtao Xue, Hao Sheng 0001, Wei Ke 0001 |
KSEM (1) | 6 |
| 2018 | W-Shaped Selection for Light Field Super-Resolution
Bing Su 0004, Hao Sheng 0001, Shuo Zhang 0003, Da Yang 0001, Nengcheng Chen, Wei Ke 0001 |
KSEM (1) | 6 |
| 2018 | Particle-mesh coupling in the interaction of fluid and deformable bodies with screen space refraction renderingabstractAbstract On the basis of the smoothed particle hydrodynamics and finite element method (FEM) model, we propose a method integrating several improvements for the real‐time simulation of fluid interacting with deformable bodies. We improve the particle neighbor search in smoothed particle hydrodynamics, so that the predefined scene containers are no longer needed. This improvement can also be applied to the simulation of fluid interacting with other materials, such as rigid and soft bodies. We also propose a two‐way coupling method for fluid and deformable bodies, where the particle–mesh interaction is obtained by the ray‐traced collision detection method instead of the proxy/ghost particle generation. By using the forward ray‐tracing method for both velocity and position, we are able to calculate the coupling forces based on the conservation of momentum and kinetic energy in the particle–mesh interaction. We use the screen space fluid rendering for fluid, and on the basis of that, we introduce a screen space refraction rendering method to improve the refraction effect. We implement our method in NVIDIA CUDA and OptiX to make use of the full computational power of a graphics processing unit. The simulation results are analyzed and discussed to show the efficiency of our method. Ka-Hou Chan, Wei Ke 0001, Sio Kei Im |
Comput. Animat. Virtual Worlds | 2 |
| 2017 | Fast Binarisation with Chebyshev InequalityabstractIn order to enhance the binarization result of degraded document images with smudged and bleed-through background, we present a fast binarization technique that applies the Chebyshev theory in the image preprocessing. We introduce the Chebyshev filter which uses the Chebyshev inequality in the segmentation of objects and background. Our result shows that the Chebyshev filter is not only effective, but also simple, robust and easy to implement. Because of its simplicity, our method is sufficiently efficient to process live image sequences in real-time. We have implemented and compared with the Document Image Binarization Contest datasets (H-DIBCO 2014) for testing and evaluation. The experimental outcomes have demonstrated that this method achieved good result in this literature. Ka-Hou Chan, Sio Kei Im, Wei Ke 0001 |
DocEng | 3 |
| 2017 | A teacher's view about introductory programming teaching and learning - Portuguese and Macanese perspectivesabstractThe difficulties faced by students and teachers in learning and teaching introductory programming has been a research issue over the years. Demotivation is common in many novice programming students, who are not able to cope with the natural difficulties associated to programming learning. It is up to the teacher to find strategies to help students and keep them motivated during the course. The objective of our research was to know more about the pedagogical and motivational strategies used by teachers in the author's institutions to promote programming student's motivation and learning. Some time ago we interviewed a few Portuguese teachers with diversified experiences in programming teaching. Recently we had the opportunity to do the same kind of research with Professors from Macao (China). The problems identified, the teachers' motivation to teach programming, the educational and motivational strategies and the student-teacher relationship were specifically addressed. This paper describes the research done, stressing the Macanese teacher's views and relating them with the views previously expressed by their Portuguese colleagues. Anabela Jesus Gomes, Wei Ke 0001, Sio Kei Im, Andrew Siu, António J. Mendes, Maria José Marcelino |
FIE | 2 |
| 2015 | Simulation of Interaction Between Fluid and Deformable Bodies
Ka-Hou Chan, Wei Ke 0001 |
ICIG (3) | 3 |
| 2014 | GEARS: A General and Efficient Algorithm for Rendering ShadowsabstractAbstract We present a soft shadow rendering algorithm that is general, efficient and accurate. The algorithm supports fully dynamic scenes, with moving and deforming blockers and receivers, and with changing area light source parameters. For each output image pixel, the algorithm computes a tight but conservative approximation of the set of triangles that block the light source as seen from the pixel sample. The set of potentially blocking triangles allows estimating visibility between light points and pixel samples accurately and efficiently. As the light source size decreases to a point, our algorithm converges to rendering pixel accurate hard shadows. Lili Wang 0006, Shiheng Zhou, Wei Ke 0001, Voicu Popescu |
Comput. Graph. Forum | 3 |
| 2014 | Fast and compact dynamic data compression based on composite rigid body construction
Lili Wang 0006, Wei Ke 0001, Qinping Zhao |
Sci. China Inf. Sci. | 4 |
| 2014 | Interactive texture design and synthesis from mesh sketches
Lili Wang 0006, Qinglin Qi, Wei Ke 0001, Aimin Hao |
Frontiers Comput. Sci. | 4 |
| 2014 | Second-Order Feed-Forward Renderingfor Specular and Glossy ReflectionsabstractThe feed-forward pipeline based on projection followed by rasterization handles the rays that leave the eye efficiently: these first-order rays are modeled with a simple camera that projects geometry to screen. Second-order rays however, as, for example, those resulting from specular reflections, are challenging for the feed-forward approach. We propose an extension of the feed-forward pipeline to handle second-order rays resulting from specular and glossy reflections. The coherence of second-order rays is leveraged through clustering, the geometry reflected by a cluster is approximated with a depth image, and the color samples captured by the second-order rays of a cluster are computed by intersection with the depth image. We achieve quality specular and glossy reflections at interactive rates in fully dynamic scenes. Lili Wang 0006, Naiwen Xie, Wei Ke 0001, Voicu Popescu |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2013 | Composite Rigid Body Construction for Fast and Compact Dynamic Data CompressionabstractCompression of 3D dynamic datasets in remote visualization still remains two challenges. One is low time performance due to the grown data and complex computation of compression algorithm. Another is small compression factor because of dynamic scenes without known equations of their motions. In this paper, we propose a fast and compact compression for 3D dynamic datasets. It accelerates compression with KD-tree construction and node-grid mapping for the dynamic data, which allow parallel rigid body decomposition and merging with disjoint union method. To increase the compression factor, composite rigid body is introduced with consideration of temporary motion consistency among rigid bodies. The results of the experiments show that our algorithm can compress dynamic datasets quickly and obtain high compression factor to reduce limitation of bandwidth. Lili Wang 0006, Xinwe Zhang, Wei Ke 0001, Qinping Zhao |
CAD/Graphics | 4 |
| 2013 | A graph-based generic type system for object-oriented programs
Wei Ke 0001, Zhiming Liu 0001, Shuling Wang 0003, Liang Zhao 0022 |
Frontiers Comput. Sci. | 1 |
| 2012 | rCOS: a formal model-driven engineering method for component-based software
Wei Ke 0001, Zhiming Liu 0001, Volker Stolz |
Frontiers Comput. Sci. China | 1 |
| 2009 | A Graph-Based Operational Semantics of OO Programs
Wei Ke 0001, Zhiming Liu 0001, Shuling Wang 0003, Liang Zhao 0022 |
ICFEM | 1 |