VLDB 2026 Research / reviewers in the wild / expert
Wai Keung Wong
dblp:122/1539 · also Wai-Keung Wong, Waikeung Wong
· DBLP profile ↗
155ranked-venue papers
20as first author
75since 2021 · last 2026
0000-0002-5214-7114ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 91 · 13 first-author · 43 since 2021Graphics, computer vision, multimedia, augmented reality and games · 58 · 5 first-author · 34 since 2021Databases, data management, data science and information retrieval · 12 · 7 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 first-author · 2 since 2021Systems, architecture and hardware · 2Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Quality-aware and Soft Consistency Driven Representation Fusion for Incomplete Multi-view Multi-label ClassificationabstractMulti-view multi-label classification aims to utilize the rich information contained in multiple views for accurate classification. However, in real-world applications, its performance is often severely constrained by the concurrent missingness of both views and labels. To address this problem, this paper first targets the drawback of representation degradation in traditional feature disentanglement methods caused by strong consistency constraints and proposes a soft consistency constraint. This constraint not only effectively aligns the shared information and maximally avoids the compression of information beneficial to the classification task, but it also enhances the aggregation effect of high-quality representations on other representations. Furthermore, to address the coarse-grained problem of traditional fusion strategies, we designed a quality assessment network that achieves instance-level dynamic weighted fusion in a data-driven manner. Extensive experiments on multiple benchmark datasets demonstrate that our method achieves state-of-the-art performance in both incomplete and complete data scenarios, showcasing its robustness and generality. Wai Keung Wong, Jie Wen 0001 |
AAAI | 2 |
| 2026 | Towards intelligent online cross-selling
Kaicheng Pang, Xingxing Zou, Zowie Broach, Wai Keung Wong |
Expert Syst. Appl. | 4 |
| 2026 | Attention-based multi-scale dual flow reverse distillation for unsupervised anomaly detection
Jiayin Zhou, Wai Keung Wong, Fangjian Liao |
Expert Syst. Appl. | 2 |
| 2026 | Revisiting kernel complexity in Gaussian process regression: An empirical study on text-to-visual embedding mapping in fashion design
Dongmei Mo, Daniel Augusto R. M. A. de Souza, Xingxing Zou, Wai Keung Wong |
Neurocomputing | 4 |
| 2026 | EUAD: End-to-end unsupervised anomaly detection based on few normal data
Jiayin Zhou, Wai Keung Wong, Kaihang Jiang, Pengjie Tan |
Neurocomputing | 2 |
| 2026 | An Efficient Regenerated Cross-Modal Hashing: Improving Existing Hash Codes With the Arbitrary LengthabstractIn recent years, numerous hashing techniques have been developed to boost efficient cross-modal retrieval. Once a retrieval model is deployed, the hash code length is fixed to achieve optimal performance. To address different retrieval scenarios while maintaining retrieval accuracy, a common approach is to redesign and retrain the original model with different hash code length. However, this retraining process can increase the training load and may lead to worse results. To tackle these challenges, we present Regenerated Cross-Modal Hashing (RCMH), a novel cross-modal hashing framework designed to improve the quality of existing hash codes and convert them to arbitrary lengths with high efficiency. First, we clip or pad the existing hash codes to initialize them with the target length, under the supervision of the similarity matrix generated by the augmented label information. Second, we introduce a linear-nonlinear competitive reconstruction approach to reduce the semantic gaps and further capture the deeper relationships from linear image features and nonlinear text features. In this way, each pair of samples is compared and selected to obtain reconstructed binary codes that can preserve the modality-specific properties. Finally, to reduce the training costs caused by iterations of variables, the regenerate hashing term is utilized to regenerate final hash codes with the reconstructed binary codes while preserving the information from the existing hash codes without iterative optimization. Notably, RCMH can be integrated with existing state-of-the-art (SOTA) methods with robustness, helping them to adjust the hash code length and achieve better retrieval performance. Kaihang Jiang, Wai Keung Wong, Xiaozhao Fang, Weijun Sun, Guoxu Zhou, Shengli Xie 0001, Xiaochun Cao |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | AdaSAM-AD: Boosting SAM2 for fine-grained pixel-level anomaly detection via spatial-channel calibration and deformable cascades
Xianbing Zhao, Xinyang Yang, Wai Keung Wong, Chengliang Liu 0003 |
Pattern Recognit. | 4 |
| 2026 | One-shot unsupervised industrial anomaly detection: Enhanced performance under extreme data scarcity
Jiayin Zhou, Wai Keung Wong, Fangjian Liao |
Pattern Recognit. | 2 |
| 2026 | LayerDiffusion: High-Fidelity Layered Virtual Try-On via Semantic-Aware DiffusionabstractImage-based Virtual Try-On (VTON) aims to generate virtual try-on results by transferring input garment images onto target person images. Current algorithms have achieved high generalization and quality on public datasets including VITON and DressCode. However, current algorithms primarily process images at the pixel level while neglecting the semantic information of garments. Moreover, layered outerwear, as an essential aspect of try-on, has been largely overlooked. Layering increases the difficulty of precise garment control due to the need to consider inner garments. Furthermore, the task-specific nature of current methods prevents them from capturing layering characteristics in try-on. To address these limitations, we propose LayerDiffusion, a novel framework that enhances both generalization capability and semantic controllability for virtual try-on generation. Our approach makes three key contributions. First, we construct LayerDataset, a specialized dataset focusing on diverse outerwear garments with balanced gender representation and multi-view captures. Second, we integrate a pretrained image encoder to capture rich semantic garment information, combined with multi-conditional masking and adapter fine-tuning to enable flexible layering control. Third, we introduce a joint training strategy on both try-on and try-off tasks, which substantially improves model generalization by learning bidirectional garment-body correlations. Extensive experiments demonstrate that LayerDiffusion significantly outperforms existing state-of-the-art methods on public benchmarks (VITON-HD and DressCode) across multiple metrics including LPIPS, SSIM, FID, KID, and CLIP-based scores. Our method excels in both single-garment and layered try-on scenarios while naturally supporting try-off capabilities, offering superior semantic preservation and detail fidelity compared to existing approaches. Fangjian Liao, Xingxing Zou, Wai Keung Wong |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | A Two-Stage Conditional Diffusion Model With Differential Attention for Hyperspectral and Multispectral Image FusionabstractDiffusion models have been used extensively for hyperspectral and multispectral image fusion; however, their intrinsic hallucination phenomenon frequently results in a loss of high-frequency details in the fused images. To address this limitation, this paper proposes a fusion method that integrates a differential attention mechanism with a two-stage conditional diffusion model. The proposed method leverages differential attention to compute the difference between two independent attention maps, thereby producing sparser and more focused attention representations. In addition, the two-stage conditional injection strategy is implemented to realize precise control over the image generation process. In the first stage, feature-level linear modulation via affine transformation is applied within the encoder to maintain global structural consistency. Then, in the second stage, wavelet features extracted from the conditioning images are injected into the decoder to facilitate the restoration of fine-grained details. Extensive validation experiments on the CAVE, Harvard, Pavia Center and Chikusei datasets verify the effectiveness of the proposed method. Compared with numerous state-of-the-art approaches, our method consistently achieves superior performance across key evaluation metrics. On the CAVE dataset, the ERGAS and RMSE metrics improved by 2.16% and 5.01%, and these metrics increased by 1.78% and 1.38% on the Pavia Center dataset, respectively. The code will be available at https://github.com/Ruijie2580/DifferentialDiff. Yingxia Chen, Wai Keung Wong, Jie Wen 0001 |
IEEE Trans. Image Process. | 3 |
| 2026 | Deep Multi-View Clustering via Cluster-Semantic GuidanceabstractDeep multi-view clustering aims to exploit the rich semantic information contained in heterogeneous multi-view data to uncover the underlying relationships among samples. However, existing deep multi-view clustering models often overlook inter-cluster separability and the effective integration of semantic information across views, resulting in insufficient feature discriminability and consequently limited clustering performance. To address the above issues, this paper proposes a novel deep multi-view clustering method via cluster-semantic guidance. We separate clusters to enhance inter-cluster discriminability, while incorporating a knowledge distillation mechanism to ensure cluster stability and facilitate the learning of clustering-friendly representations. Furthermore, by aggregating sample-level semantic information, the model is guided to follow a cluster-oriented learning strategy that promotes the extraction of discriminative features, thereby strengthening the sample representation capability. Our method effectively learns discriminative and clustering-friendly representations, guiding the model to acquire distinctive feature embeddings from a cluster-oriented perspective. Our comprehensive experiments across datasets of varying scales confirm the model's effectiveness, showing superior clustering performance over existing state-of-the-art methods. Jinrong Cui, Xiaohuang Wu, Wai Keung Wong, Linlin Tang, Jie Wen 0001 |
IEEE Trans. Image Process. | 3 |
| 2026 | Expert-Guided Cross-View Fusion With Self-Derived Lesion Proposals for Multi-View Diabetic Retinopathy GradingabstractRecent advances in multi-view fundus imaging show great promise for automated diabetic retinopathy (DR) grading. However, mainstream end-to-end CNN/Transformer pipelines rely on striding or tokenization that compresses spatial detail, causing small, low-contrast lesions (e.g., microaneurysms) to be under-represented and creating performance ceilings. Prior efforts have mitigated this by incorporating external lesion- or vessel-level annotations into models. However, such labels are costly to acquire, break the end-to-end training, and make performance over-reliant on the annotation quality. To reduce dependence on expensive annotations, we propose an end-to-end framework that generates lesion proposals on the fly during training and inference, providing self-derived cues for grading. First, we introduce a Grade-Activated Lesion Proposal (GALP) module that derives grade-conditioned evidence maps (GEMs) from stage-wise auxiliary classifiers and selects the top-K high-evidence regions per view as lesion proposals. Second, we propose a Cross-View Lesion Expert Guided Regional Fusion (LGRF) module, which selectively activates experts for a view's lesion proposals based on contextual guidance from other views, ensuring that only the most relevant feature extractors contribute to fusion. Experimental results on two multi-view DR datasets show that our method matches or surpasses strong baselines without external annotations, confirming that self-generated proposals can substantially reduce annotation needs. Wai Keung Wong, Xueling Zhou, Junlin Hou, Jie Wen 0001 |
IEEE Trans. Image Process. | 1 |
| 2026 | Evidential Reliable Fusion for Partial Multi-View Incomplete Multi-Label Classification
Jiaying Zhou, Wai Keung Wong, Xiaohuan Lu, Youliang Tian, Jie Wen 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2026 | Multi-View Hilbert Curve-Based Hierarchical Information Aggregation for Incomplete Multimodal Alzheimer's Disease DiagnosisabstractTimely identification of Alzheimer's disease (AD) benefits from combining neuroimaging, fluid biomarkers, and cognitive assessments, yet in practice one or more modalities are often unavailable due to various factors such as cost, patient compliance, and procedural risks. Furthermore, conventional convolutional neural network (CNN) architectures and even Transformer-based models struggle to efficiently capture both local and global dependencies, especially when dealing with high-dimensional and highly heterogeneous medical data. In this study, we introduce a novel hierarchical information aggregation and dynamic fusion (HI-AD) framework for incomplete multimodal AD diagnosis. Our method couples a multi-view Hilbert curve-guided Mamba block with hierarchical spatial feature extraction to retain spatial continuity, model long-range dependencies, and integrate local context in neuroimaging data. To balance semantic alignment and modality-specific information, we propose a unified mutual information-driven learning objective with an active confidence evaluation mechanism, thereby preventing modality collapse and promoting robust representation learning. Extensive experiments on real-world datasets validate that our HI-AD framework consistently outperforms existing state-of-the-art methods across a diverse range of modality-missing scenarios, establishing an effective and generalizable solution for early-stage AD screening in heterogeneous clinical data environments. Chengliang Liu 0003, Yuanxi Que, Wai Keung Wong, Xiaoling Luo 0001 |
IEEE Trans. Medical Imaging | 3 |
| 2026 | Energy-Driven Explicit Alignment Network: A Blended-Target Domain Adaptation ApproachabstractAs a specialized paradigm of domain adaptation, blended-target domain adaptation (BTDA) transfers knowledge from a source domain to a blended target domain. In this paper, we propose an Energy-Driven Explicit Alignment Network (EDEAN) framework that innovatively applies energy-based models (EBMs) to address BTDA problems. We observe that EBMs display free energy biases when the source domain and the target domain data originate from different distributions. Therefore, we use these biases as a measure of the discrepancies between the source domain and the target domain and align them by minimizing these biases via the free energy alignment (FEA) module. We further propose the balanced weight distribution (BWD) module, which comprehensively considers the complementary information between the linear and semantic pseudo-labels and obtains the corresponding complementary information by mixing both label types. Moreover, we propose the normalized free energy (NFE) module, which assigns higher weights to high free energy samples and dynamically corrects the pseudo-labels by continuously updating these weights. We also conducted experiments on four widely used BTDA databases and achieved substantial improvements over the latest BTDA methods. Yuwu Lu, Wai Keung Wong, Anne Toomey, Zhihui Lai 0001, Xuelong Li 0001 |
IEEE Trans. Multim. | 3 |
| 2025 | Multi-view Evidential Learning-based Medical Image SegmentationabstractMedical image segmentation provides useful information about the shape and size of organs, which is beneficial for improving diagnosis, analysis, and treatment. Despite traditional deep learning-based models can extract domain-specific knowledge, they face a generalization bottleneck due to the limited embedded knowledge scope. Vision foundation models have been demonstrated to be effective in extracting generalizable knowledge, but they cannot extract domain-specific knowledge without fine-tuning. In this work, we propose a novel multi-view evidential learning-based framework, which can extract both domain-specific and generalizable knowledge from multi-view features by combining the advantages of traditional and vision foundation models. Specifically, a novel multi-view state space model (MV-SSM) is designed to extract task-related knowledge while removing redundant information within multi-view features. The proposed MV-SSM utilizes Mamba, a state space model, to model cross-view contextual dependencies between domain-specific and generalizable features. Additionally, evidential learning is adopted to quantify the segmentation uncertainty of the model for boundary. In special, variational Dirichlet is introduced to characterize the distribution of the result probabilities, parameterized with collected evidence to quantify uncertainty. As a result, the model can reduce the segmentation uncertainties of boundaries by optimizing the parameters of the Dirichlet distribution. Experimental results on three datasets show that our method obtains superior segmentation performance. Chao Huang 0008, Yushu Shi, Wai Keung Wong, Chengliang Liu 0003, Wei Wang 0169, Zhihua Wang 0002, Jie Wen 0001 |
AAAI | 3 |
| 2025 | DiffusionREC: Diffusion Model with Adaptive Condition for Referring Expression ComprehensionabstractThe objective of referring expression comprehension (REC) is to accurately identify the object in an image described by a given expression. Existing REC methods, including transformer-based and graph-based approaches among others, have shown robust performance in REC tasks. In this study, we present a groundbreaking framework named DiffusionREC for REC task. This framework reimagines the REC task as a text guided bounding box denoising diffusion process, through which noisy bounding boxes are refined and distilled to pinpoint the target box. Throughout the training process, the bounding box of the target object diffuses from its ground-truth position towards a random distribution. Simultaneously, a filtering-based object decoder is introduced to reverse this diffusion of noise, conditional on the provided expression, the result from previous denoised step and the interaction between the expression and the image. At the inference stage, we begin by randomly generating a collection of boxes. Subsequently, the filtering-based object decoder is iteratively employed to refine and prune these bounding boxes, taking into account the conditions on the given expression, the results from the previous denoised step, and the interaction between the expression and the image. Extensive experiments conducted on six datasets demonstrate that DiffusionREC outperforms previous REC methods, yielding superior performances. Jingcheng Ke, Wai Keung Wong, Jia Wang 0020, Mu Li 0005, Lunke Fei, Jie Wen 0001 |
AAAI | 2 |
| 2025 | Collaborative Semantic Consistency Alignment for Blended-Target Domain AdaptationabstractBlended-target domain adaptation (BTDA) leverages learned source knowledge to adapt the model to a blended-target domain that is composed of multiple unlabeled sub-target domains with distinct statistical characteristics. The existing BTDA methods usually overlook semantic correlation information across multiple domains and domain shifts among sub-target domains, resulting in suboptimal adaptation performance. To fully harness semantic knowledge and alleviate domain shifts in hybrid data distribution, we propose a collaborative semantic consistency alignment (CSCA) method for BTDA. Specifically, we achieve distribution alignment by minimizing the sliced Wasserstein distance between the source and target feature distributions. To alleviate complex domain shifts among all sub-target domains in the hybrid feature space, we design graph networks to propagate and share semantic knowledge across domains, which reduces semantic discrepancies among multiple domains. Additionally, we propose a double consistency regularization method to reduce the susceptibility of the model to domain-specific information, further facilitating semantic alignment and alleviating domain shifts. Extensive experiments on several datasets show that CSCA achieves promising classification performance. Yuwu Lu, Wai Keung Wong |
AAAI | 3 |
| 2025 | Invertible Projection and Conditional Alignment for Multi-Source Blended-Target Domain AdaptationabstractMulti-source domain adaptation (MSDA), which utilizes multiple source domains to align the distribution of a single target domain, is a popular and challenging setting in domain adaptation (DA). However, existing MSDA approaches are difficult to obtain sufficient target domain knowledge, which serve as the transfer object. Furthermore, the target distributions are confused in the real world, i.e., the model cannot obtain the domain labels of target domains. To tackle these problems, we consider a more realistic DA setting Multi-Source Blended-Target Domain Adaptation (MBDA) and propose an Invertible Projection and Conditional Alignment (IPCA) method. Specifically, to reduce the impact of the distribution discrepancy, we construct an invertible projection for the source and blended-target domains. Then, we adopt a projection consistency regularization to our model, which makes the model more robust on the domain-specific parts. In addition, because the labels of the blended-target domain are unseen, we introduce conditional discrepancy to obtain the domain-level discriminative information and guide the classifier to serve as the discriminator, which is suitable for MBDA settings. Extensive experiment results on the ImageCLEF-DA, Office-Home, and DomainNet datasets validate the effectiveness of our method. Yuwu Lu, Wai Keung Wong |
AAAI | 3 |
| 2025 | MVREC: A General Few-shot Defect Classification Model Using Multi-View Region-ContextabstractFew-shot defect multi-classification (FSDMC) is an emerging trend in quality control within industrial manufacturing. However, current FSDMC research often lacks generalizability due to its focus on specific datasets. Additionally, defect classification heavily relies on contextual information within images, and existing methods fall short of effectively extracting this information. To address these challenges, we propose a general FSDMC framework called MVREC, which offers two primary advantages: (1) MVREC extracts general features for defect instances by incorporating the pre-trained AlphaCLIP model. (2) It utilizes a region-context framework to enhance defect features by leveraging mask region input and multi-view context augmentation. Furthermore, Few-shot Zip-Adapter(-F) classifiers within the model are introduced to cache the visual features of the support set and perform few-shot classification. We also introduce MVTec-FS, a new FSDMC benchmark based on MVTec AD, which includes 1228 defect images with instance-level mask annotations and 46 defect types. Extensive experiments conducted on MVTec-FS and four additional datasets demonstrate its effectiveness in general defect classification and its ability to incorporate contextual information to improve classification performance. Shuai Lyu, Rongchen Zhang, Zeqi Ma, Fangjian Liao, Dongmei Mo, Wai Keung Wong |
AAAI | 6 |
| 2025 | Palm-vein images reconstruction against adversarial attacksabstractPalm-vein has received widespread attention for reliable biometric recognition due to its robust resistance to replicate and forge. However, the rise of adversarial attacks poses a high risk of vulnerability for palm-vein recognition, leaving most existing methods vulnerable to small and human-imperceptible adversarial perturbations. In this paper, we propose a palm-vein image reconstruction network for palm-vein image protection, which mainly consists of palm-vein-specific exploration, feature refinement, and image reconstruction sub-networks. Specifically, we first specially learn the noise-insensitive palm-vein-specific feature by decoupling non-vein noise information via cascaded noise-injected and Canny-based convolution layers, and then refine palm-vein-specific features via multiple stacked basic convolution and transposed convolution pairs. Lastly, we convert the fine-grained palm-vein features into the latent sharp palm-vein images via two transposed convolution layers. Moreover, we develop both identity-aware and visual-aware loss functions to ensure the high-quality of the reconstructed palm-vein images. Experimental results on the widely used PolyU palm-vein dataset demonstrate the promising effectiveness of the proposed palm-vein image reconstruction network. Lunke Fei, Wai Keung Wong, Shuping Zhao, Anne Toomey, Jiehang Deng |
ICASSP | 3 |
| 2025 | Coarse-To-Fine Graph Reasoning for 3D Hand Mesh Reconstructionabstract3D hand mesh reconstruction from 2D images is crucial for various computer visual tasks such as virtual reality and human-computer interaction, while it remains a challenging problem due to changed hand poses and diverse self-/cross-hand occlusions. In this paper, we propose a graph-based reasoning network for 3D hand mesh reconstruction from a single 2D RGB image by recovering the fine-grained 3D hand mesh keypoints in a coarse-to-fine manner. First, we extract the hand-joint-specific features to capture the overall hand structure via a CNN backbone with a linear projection sampling operation. Based on the hand-joint locations, we further progressively learn more hand mesh keypoints to construct the coarse hand shape via hierarchical attention-embedded graph learning layers. Finally, we leverage the 2D shallow semantic features to refine the coarse hand mesh keypoints into fine-grained 3D hand mesh vertices coordinates via cascaded graph learning layer with linear mapping. Experimental results on the widely used hand databases show that our method achieves outstanding performance in both single-hand and two-interactive-hand 3D mesh reconstruction. Dan Fu, Wai Keung Wong, Lunke Fei, Tingting Chai, Yuzhu Ji, Qinghua Zhu 0001 |
ICME | 2 |
| 2025 | Mutual Learning for SAM Adaptation: A Dual Collaborative Network Framework for Source-Free Domain TransferabstractSegment Anything Model (SAM) has demonstrated remarkable zero-shot segmentation capabilities across various visual tasks. However, its performance degrades significantly when deployed in new target domains with substantial distribution shifts. While existing self-training methods based on fixed teacher-student architectures have shown improvements, they struggle to ensure that the teacher network consistently outperforms the student under severe domain shifts. To address this limitation, we propose a novel Collaborative Mutual Learning Framework for source-free SAM adaptation, leveraging dual-networks in a dynamic and cooperative manner. Unlike fixed teacher-student paradigms, our method dynamically assigns the teacher and student roles by evaluating the reliability of each collaborative network in each training iteration. Our framework incorporates a dynamic mutual learning mechanism with three key components: a direct alignment loss for knowledge transfer, a reverse distillation loss to encourage diversity, and a triplet relationship loss to refine feature representations. These components enhance the adaptation capabilities of the collaborative networks, enabling them to generalize effectively to target domains while preserving their pre-trained knowledge. Extensive experiments on diverse target domains demonstrate that our proposed framework achieves state-of-the-art adaptation performance. Wai Keung Wong, Chengliang Liu 0003, Xiaoling Luo 0001, Yong Xu 0001 |
ICML | 2 |
| 2025 | Label Prediction Inherited Hashing for Cross-Modal Retrieval: Applying Supervised Hashing to Unsupervised TasksabstractSupervised cross-modal hashing has achieved remarkable progress in retrieving related items across different modalities. However, in practical applications, a significant portion of data remains unlabeled, such as online data on websites, which must be included for effective retrieval. To address this challenge, while maintaining the high accuracy and efficiency of supervised methods, few works have attempted to adapt existing supervised techniques to handle unsupervised tasks through a general modular approach. To this end, we introduce a novel cross-modal hashing method, termed Label Prediction Inherited Hashing (LPIH). Initially, LPIH leverages labeled data to learn high-quality general label functions using supervised methods. Subsequently, it inherits the existing hash codes from existing supervised methods to further refine the pseudo-label information. Finally, LPIH integrates the refined pseudo-label information with the existing hash functions to learn new hash functions specifically tailored for unsupervised tasks. Extensive experimental results on three public datasets demonstrate the superior performance of LPIH compared to state-of-the-art (SOTA) cross-modal hashing methods. Specifically, LPIH achieves an average precision improvement of 5% over SOTA methods, highlighting its effectiveness in bridging the gap between supervised and unsupervised learning in the context of cross-modal retrieval. Kaihang Jiang, Wai Keung Wong, Jianyang Qin, Xiaozhao Fang, Jie Wen 0001, Bingzhi Chen, Hongbo Gao 0001 |
ACM Multimedia | 2 |
| 2025 | Patch distance based auto-encoder for industrial anomaly detection
Zeqi Ma, Jiaxing Li 0009, Wai Keung Wong |
Expert Syst. Appl. | 3 |
| 2025 | LoopNet for fine-grained fashion attributes editing
Xingxing Zou, Shumin Zhu, Wai Keung Wong |
Expert Syst. Appl. | 3 |
| 2025 | Self-inferring incomplete multi-view clusteringabstractAbstract With the advantage of exploiting complementary and consensus information across multiple views, techniques for Multi‐view Clustering have attracted increasing attention in recent years. However, it is common that data on some views is not completed in real‐world applications, which brings the challenge of partial mapping between the views. To explore the information hidden in the local geometric structure and recover missing instances through mining the information hidden in existing instances, a self‐inferring incomplete multi‐view clustering algorithm is proposed. Firstly, the incomplete multi‐view data is replenished directly and exploited as variables for inferring the missing instances. And then, a feature graph constraint is united in consensus learning. Besides, a similarity graph learning method is imposed to preserve the local manifold structure. At last, the inferred instances are filled in the missing instances for learning better consensus representation in the iterative process. Extensive experiment results show that this method can improve the clustering performance compared with the state‐of‐the‐art methods. Junjun Fan, Zeqi Ma, Jiajun Wen 0001, Zhihui Lai 0001, Weicheng Xie 0001, Wai Keung Wong |
IET Comput. Vis. | 6 |
| 2025 | Learning to estimate 3D interactive two-hand poses with attention perception
Wai Keung Wong, Hongkun Sun, Weijun Sun, Shuping Zhao, Lunke Fei |
Image Vis. Comput. | 1 |
| 2025 | Integrating local and global correlations with Mamba-Transformer for multi-class anomaly detection
Zeqi Ma, Jiaxing Li 0009, Kaihang Jiang, Wai Keung Wong |
Knowl. Based Syst. | 4 |
| 2025 | MLDF: Multi-scale local descriptors fusion for lace fabric image retrieval
Rongchen Zhang, Wai Keung Wong, Shuai Lyu |
Knowl. Based Syst. | 2 |
| 2025 | Task-augmented cross-view imputation network for partial multi-view incomplete multi-label classification
Lian Zhao, Jie Wen 0001, Xiaohuan Lu, Wai Keung Wong, Wulin Xie |
Neural Networks | 4 |
| 2025 | Any Fashion Attribute Editing: Dataset and Pretrained ModelsabstractFashion attribute editing is essential for combining the expertise of fashion designers with the potential of generative artificial intelligence. In this work, we focus on 'any' fashion attribute editing: 1) the ability to edit 78 fine-grained design attributes commonly observed in daily life; 2) the capability to modify desired attributes while keeping the rest components still; and 3) the flexibility to continuously edit on the edited image. To this end, we present the Any Fashion Attribute Editing (AFED) dataset, which includes 830 K high-quality fashion images from sketch and product domains, filling the gap for a large-scale, openly accessible fine-grained dataset. We also propose Twin-Net, a twin encoder-decoder GAN inversion method that offers diverse and precise information for high-fidelity image reconstruction. This inversion model, trained on the new dataset, serves as a robust foundation for attribute editing. Additionally, we introduce PairsPCA to identify semantic directions in latent space, enabling accurate editing without manual supervision. Comprehensive experiments, including comparisons with ten state-of-the-art image inversion methods and four editing algorithms, demonstrate the effectiveness of our Twin-Net and editing algorithm. All data and models are available at https://github.com/ArtmeScienceLab/AnyFashionAttributeEditing. Shumin Zhu, Xingxing Zou, Wenhan Yang, Wai Keung Wong |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Semantic decomposition and enhancement hashing for deep cross-modal retrieval
Lunke Fei, Wai Keung Wong, Qi Zhu 0001, Shuping Zhao, Jie Wen 0001 |
Pattern Recognit. | 3 |
| 2025 | Separation of Unknown Features and Samples for Unbiased Source-free Open Set Domain Adaptation
Yifan Lan, Yuwu Lu, Wai Keung Wong, Ming Zhao 0011, Zhihui Lai 0001, Xuelong Li 0001 |
Pattern Recognit. | 4 |
| 2025 | Dual structure-aware consensus graph learning for incomplete multi-view clustering
Lilei Sun, Wai Keung Wong, Yusen Fu, Jie Wen 0001, Mu Li 0005, Yuwu Lu, Lunke Fei |
Pattern Recognit. | 2 |
| 2025 | Approximate geometric structure transfer for cross-domain image classification
Wai Keung Wong, Yuwu Lu, Zhihui Lai 0001, Xuelong Li 0001 |
Pattern Recognit. | 1 |
| 2025 | Graph-Based Group Division Network for Referring Expression ComprehensionabstractReferring expression comprehension (REC) aims at locating the target object described by an expression. We observe that most of the graph-based REC methods only focus on establishing relations between all objects in an image and the given expression during the graph construction while ignoring the relationships between objects in the same category. As a result, these methods are sub-optimal in locating the target object described by the expression, particularly when the target object is surrounded by objects of similar categories. Meanwhile, during reasoning, numerous irrelevant objects are considered for expression, which will introduce significant harmful noise. To address these issues, this paper proposes a new graph-based group division network (GBGDN). Different from the existing works, our work partitions the constructed graphs into several sub-graphs based on the categories of objects and expressions. In each sub-graph, the common visual features of objects will be strengthened through a feature enhancement strategy. Subsequently, the enhanced sub-graphs and expressions undergo joint processing via a filtering-based reasoning module designed to reduce the influence of unrelated nodes in each sub-graph, facilitating more accurate reasoning and matching. Experimental results across various datasets, including RefCOCO /+/g, Flickr30K Entities, RefClef, and Ref-reasoning, showcase the superiority of our proposed method over existing approaches. Most importantly, our method does not need pre-training. Jingcheng Ke, Jia Wang 0020, Wai Keung Wong, Anne Toomey, Jie Wen 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Adaptive Dispersal and Collaborative Clustering for Few-Shot Unsupervised Domain AdaptationabstractUnsupervised domain adaptation is mainly focused on the tasks of transferring knowledge from a fully-labeled source domain to an unlabeled target domain. However, in some scenarios, the labeled data are expensive to collect, which cause an insufficient label issue in the source domain. To tackle this issue, some works have focused on few-shot unsupervised domain adaptation (FUDA), which transfers predictive models to an unlabeled target domain through a source domain that only contains a few labeled samples. Yet the relationship between labeled and unlabeled source domains are not well exploited in generating pseudo-labels. Additionally, the few-shot setting further prevents the transfer tasks as an excessive domain gap is introduced between the source and target domains. To address these issues, we newly proposed an adaptive dispersal and collaborative clustering (ADCC) method for FUDA. Specifically, for the shortage of the labeled source data, a collaborative clustering algorithm is constructed that expands the labeled source data to obtain more distribution information. Furthermore, to alleviate the negative impact of domain-irrelevant information, we construct an adaptive dispersal strategy that introduces an intermediate domain and pushes both the source and target domains to this intermediate domain. Extensive experiments on the Office31, Office-Home, miniDomainNet, and VisDA-2017 datasets showcase the superior performance of ADCC compared to the state-of-the-art FUDA methods. Yuwu Lu, Wai Keung Wong, Zhihui Lai 0001, Xuelong Li 0001 |
IEEE Trans. Image Process. | 3 |
| 2025 | Generative AI in Fashion: OverviewabstractGenerative Artificial Intelligence (GenAI) has recently gained immense popularity by offering various applications for generating high-quality and aesthetically pleasing content of image, 3D, and video data format. The innovative GenAI solutions have shifted paradigms across various design-related industries, particularly fashion. In this paper, we explore the incorporation of GenAI into fashion-related tasks and applications. Our examination encompasses a thorough review of more than 470 research papers and an in-depth analysis of over 300 applications, focusing on their contributions to the field. These contributions are identified as 13 tasks within four categories: multi-modal fashion understanding, and fashion synthesis of image, 3D, and dynamic (video and animatable 3D) formats We delve into these methods, recognizing their potential to propel future endeavours toward achieving state-of-the-art (SOTA) performance. Furthermore, we present a comprehensive overview of 53 publicly available datasets suitable for training and benchmarking fashion-centric models, accompanied by the relevant evaluation metrics. Finally, we review real-world applications, unveiling existing challenges and future directions. With comprehensive investigation and in-depth analysis, this paper is targeted to serve as a useful resource for understanding the current landscape of GenAI in fashion, paving the way for future innovations in this dynamic field. Papers discussed in this paper, along with public code and datasets links are available at: https://github.com/wendashi/Cool-GenAI-Fashion-Papers/ . Wenda Shi, Wai Keung Wong, Xingxing Zou |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2025 | Collaboratively Semantic Alignment and Metric Learning for Cross-Modal HashingabstractCross-modal retrieval is a promising technique nowadays to find semantically similar instances in other modalities while a query instance is given from one modality. However, there still exists many challenges for reducing heterogeneous modality gap by embedding label information to discrete hash codes effectively, solving the binary optimization when generating unified hash codes and reducing the discrepancy of data distribution efficiently during common space learning. In order to overcome the above-mentioned challenges, we propose a Collaboratively Semantic alignment and Metric learning for cross-modal Hashing (CSMH) in this paper. Specifically, by a kernelization operation, CSMH first extracts the non-linear data features for each modality, which are projected into a latent subspace to align both marginal and conditional distributions simultaneously. Then, a maximum mean discrepancy-based metric strategy is customized to mitigate the distribution discrepancies among features from different modalities. Finally, semantic information obtained from the label similarity matrix, is further incorporated to embed the latent semantic structure into the discriminant subspace. Experimental results of CSMH and baseline methods on four widely-used datasets show that CSMH outperforms some state-of-the-art hashing baseline methods for cross-modal retrieval on efficiency and precision. Jiaxing Li 0009, Wai Keung Wong, Kaihang Jiang, Xiaozhao Fang, Shengli Xie 0001, Jie Wen 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2025 | Heterogeneous Pairwise-Semantic Enhancement Hashing for Large-Scale Cross-Modal RetrievalabstractCross-modal hash learning has drawn widespread attention for large-scale multimodal retrieval because of its stability and efficiency in approximate similarity searches. However, most existing cross-modal hashing approaches employ discrete label-guided information to coarsely reflect intra- and intermodality correlations, making them less effective to measuring the semantic similarity of data with multiple modalities. In this paper, we propose a new heterogeneous pairwise-semantic enhancement hashing (HPsEH) for large-scale cross-modal retrieval by distilling higher-level pairwise-semantic similarity from supervision information. First, we adopt a supervised self-expression to learn a data-specific quantified semantic matrix, which uses real values to measure both the similarity and dissimilarity ranks of paired instances, such that the intrinsic semantics of the data can be well captured. Then, we fuse the label-based information and quantified semantic similarity to collaboratively learn the hash codes of multimodal data, such that both the intermodality consistency and modality-specific features can be simultaneously obtained during hash code learning. Moreover, we employ effective iterative optimization to address the discrete binary solution and massive pairwise matrix calculation, making the HPsEH scalable to large-scale datasets. Extensive experimental results on three widely used datasets demonstrate the superiority of our proposed HPsEH method over most state-of-the art approaches. Wai Keung Wong, Lunke Fei, Jianyang Qin, Shuping Zhao, Jie Wen 0001 |
IEEE Trans. Multim. | 1 |
| 2025 | Random Online Hashing for Cross-Modal RetrievalabstractIn the past decades, supervised cross-modal hashing methods have attracted considerable attentions due to their high searching efficiency on large-scale multimedia databases. Many of these methods leverage semantic correlations among heterogeneous modalities by constructing a similarity matrix or building a common semantic space with the collective matrix factorization method. However, the similarity matrix may sacrifice the scalability and cannot preserve more semantic information into hash codes in the existing methods. Meanwhile, the matrix factorization methods cannot embed the main modality-specific information into hash codes. To address these issues, we propose a novel supervised cross-modal hashing method called random online hashing (ROH) in this article. ROH proposes a linear bridging strategy to simplify the pair-wise similarities factorization problem into a linear optimization one. Specifically, a bridging matrix is introduced to establish a bidirectional linear relation between hash codes and labels, which preserves more semantic similarities into hash codes and significantly reduces the semantic distances between hash codes of samples with similar labels. Additionally, a novel maximum eigenvalue direction (MED) embedding method is proposed to identify the direction of maximum eigenvalue for the original features and preserve critical information into modality-specific hash codes. Eventually, to handle real-time data dynamically, an online structure is adopted to solve the problem of dealing with new arrival data chunks without considering pairwise constraints. Extensive experimental results on three benchmark datasets demonstrate that the proposed ROH outperforms several state-of-the-art cross-modal hashing methods. Kaihang Jiang, Wai Keung Wong, Xiaozhao Fang, Jiaxing Li 0009, Jianyang Qin, Shengli Xie 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | Confident Local Structure-Aware Incomplete Multiview Spectral ClusteringabstractExploring the structure information is crucial for data clustering task, particularly for the sceneries of incomplete multiview clustering (IMVC) when some views are missing. However, almost all of the existing graph-based IMVC methods either introduce the Laplacian constraint with fixed graphs or simply fuse the graphs of all views, which are vulnerable to the quality of the constructed graphs. To address this issue, we propose a new graph-based method, called confident local structure-aware incomplete multiview spectral clustering. Different from existing works, our method seeks to adaptively uncover the inherent similarity structure among the available instances in each view and learn the optimal consensus graph within a unified learning framework. Moreover, to mitigate the adverse effects of imbalance information across incomplete views and improve the quality of consensus graph, we further impose some adaptive weights on the consensus graph learning model w.r.t. each view and introduce some confident structure graphs to explore the most confident similarity information in the model. In contrast to existing works, our approach simultaneously takes into account the pairwise similarity information and neighbor group-based confident structure information. This dual consideration makes our method more effective in achieving the optimal consensus graph and delivering superior IMVC performance. Experimental results on several datasets demonstrate that our method effectively learns a high-quality and clustering-friendly graph from incomplete multiview data, and it outperforms many state-of-the-art IMVC methods in terms of clustering performance. Wai Keung Wong, Lusi Li, Lunke Fei, Bob Zhang 0001, Anne Toomey, Jie Wen 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2024 | Denoising High-Order Graph ClusteringabstractHigh-Order Graph (HOG) clustering has received much attention for its advantage of exploiting the rich intrinsic structure of data. However, the construction of HOG involves the generation of a large number of redundant walks, which dilutes the useful walks and thus leads to untrustworthy high-order similarity and, consequently, suboptimal clustering results may be obtained. We formalize this issue as the Weight Explosion (WE) problem. Furthermore, current works rarely focus on exploiting the correlation between multi-order graphs that can capture high-order relations at various levels. In this paper, we first analyze the pattern of redundant walks, also termed as noise, and subsequently propose a novel$h$-length Simple Path Search ($h$-SPS) algorithm to solve the WE problem.$h$-SPS aims to find valid walks to denoise HOG and thus avoids enumerating walks to report the similarity. Regarding the second problem, we propose a multi-order graphs fusion method, which adaptively integrates graphs of varying orders by solving a convex problem. This allows us to capture information across different order levels effectively. Extensive experiments on benchmark datasets demonstrate that our method11https://github.com/YonghaoChen511/DenoHOG can effectively solve the proposed WE problem, while also well exploiting the correlation of multi-order graphs. Yonghao Chen, Ruibing Chen, Qiaoyun Li, Xiaozhao Fang, Jiaxing Li 0009, Wai Keung Wong |
ICDE | 6 |
| 2024 | Diffusion-based Missing-view Generation With the Application on Incomplete Multi-view ClusteringabstractAs a branch of clustering, multi-view clustering has received much attention in recent years. In practical applications, a common phenomenon is that partial views of some samples may be missing in the collected multi-view data, which poses a severe challenge to design the multi-view learning model and explore complementary and consistent information. Currently, most of the incomplete multi-view clustering methods only focus on exploring the information of available views while few works study the missing view recovery for incomplete multi-view learning. To this end, we propose an innovative diffusion-based missing view generation (DMVG) network. Moreover, for the scenarios with high missing rates, we further propose an incomplete multi-view data augmentation strategy to enhance the recovery quality for the missing views. Extensive experimental results show that the proposed DMVG can not only accurately predict missing views, but also further enhance the subsequent clustering performance in comparison with several state-of-the-art incomplete multi-view clustering methods. Jie Wen 0001, Wai Keung Wong, Guoqing Chao, Chao Huang 0008, Lunke Fei, Yong Xu 0001 |
ICML | 3 |
| 2024 | Uni-DlLoRA: Style Fine-Tuning for Fashion Image TranslationabstractImage-to-image (i2i) translation has achieved notable success, yet remains challenging in scenarios like real-to-illustrative style transfer of fashion. Existing methods focus on enhancing the generative model with diversity while lacking ID-preserved domain translation. This paper introduces a novel model named Uni-DlLoRA to release this constraint. The proposed model combines the original images within a pretrained diffusion-based model using the proposed Uni-adapter extractors, while adopting the proposed Dual-LoRA module to provide distinct style guidance. This approach optimizes generative capabilities and reduces the number of additional parameters required. In addition, a new multimodal dataset featuring higher-quality images with captions built upon an existing real-to-illustration dataset is proposed. Experimentation validates the effectiveness of our proposed method. Fangjian Liao, Xingxing Zou, Wai Keung Wong |
ACM Multimedia | 3 |
| 2024 | Learning Visual Body-shape-Aware Embeddings for Fashion CompatibilityabstractBody shape is a crucial factor in outfit recommendation. Previous studies that directly used body measurement data to investigate the relationship between body shape and outfit have achieved limited performance due to oversimplified body shape representations. This paper proposes a Visual Body-shape-Aware Network (ViBA-Net) to improve the fashion compatibility model’s awareness of human body shape through visual-level information. Specifically, ViBA-Net consists of three modules: a body-shape embedding module, which extracts visual and anthropometric features of body shape from a newly introduced large-scale body shape dataset; an outfit embedding module, which learns the outfit representation based on visual features extracted from a try-on image and textual features extracted from fashion attributes; and a joint embedding module, which jointly models the relationship between the representations of body shape and outfit. ViBA-Net is designed to generate attribute-level explanations for the evaluation results based on the computed attention weights. The effectiveness of ViBA-Net is evaluated on two mainstream datasets through qualitative and quantitative analysis. Data and code are released1. Kaicheng Pang, Xingxing Zou, Wai Keung Wong |
WACV | 3 |
| 2024 | Attentional pixel-wise deformation for pose-based human image generation
Fangjian Liao, Xingxing Zou, Wai Keung Wong |
Expert Syst. Appl. | 3 |
| 2024 | REB: Reducing biases in representation for industrial anomaly detection
Shuai Lyu, Dongmei Mo, Wai Keung Wong |
Knowl. Based Syst. | 3 |
| 2024 | Unsupervised anomaly detection and localization with one model for all category
Pengjie Tan, Wai Keung Wong |
Knowl. Based Syst. | 2 |
| 2024 | Decoupling visual and identity features for adversarial palm-vein image attack
Wai Keung Wong, Lunke Fei, Shuping Zhao, Jie Wen 0001, Shaohua Teng |
Neural Networks | 2 |
| 2024 | Graph correlated discriminant embedding for multi-source domain adaptationabstractAs a main branch of domain adaptation (DA), multi-source DA (MSDA) has attracted increasing attention for exploiting information from multi-source domain data. However, how to effectively explore useful information from each source domain for target tasks is still a key problem. In this paper, to fully explore multiple information of different domain data, we propose a graph correlated discriminant embedding (GCDE) method for MSDA. In GCDE, the category-discriminative information, manifold structure, and correlation learning are fully considered. Specifically, GCDE encodes the within- and between- class information of each domain data, preserves the local and global structure information of the data, and extracts the maximization correlative features from different domains by designing a novel correlative learning scheme. We also extend GCDE to a nonlinear case and obtain kernel GCDE (KGCDE). We have conducted extensive experiments on four public data benchmarks to verify the performance of GCDE and KGCDE. The promising performance on the databases prove the efficiency of our methods with the comparison of the advanced approaches. Wai Keung Wong, Yuwu Lu, Zhihui Lai 0001, Xuelong Li 0001 |
Pattern Recognit. | 1 |
| 2024 | CKDH: CLIP-Based Knowledge Distillation Hashing for Cross-Modal RetrievalabstractRecently, deep hashing-based cross-modal retrieval has attracted much attention of researchers, due to its advantages of fast retrieval efficiency and low storage overhead, etc. However, the existing deep hashing-based cross-modal retrieval methods typically 1) suffer from inadequately capturing the semantic relevance and coexistent information for cross-modal data, which may result in sub-optimal retrieval performance, 2) require a more comprehensive similarity measurement for cross-modal features to ensure high retrieval accuracy, 3) lack of scalability for lightweight deployment framework. To handle the issues mentioned above, we propose a CLIP-based knowledge distillation hashing (CKDH) for cross-modal retrieval, by referring the research trend of combining traditional methods and modern neural architecture to design lightweight networks based on large language models. Specifically, to effectively help capture the semantic relevance and coexistent information, CLIP is fine-tuned to extract visual features, while a graph attention network is used to enhance textual features extracted by bag-of-words model in the teacher model. Then, for better supervising the training of student model, a more comprehensive similarity measurement is introduced to represent distilled knowledge by jointly preserving the log-likelihood, intra and inter modality similarities. Finally, the student model extracts deep features by a lightweight networks, and generates the hash codes under the supervision of the similarity matrix produced by the teacher model. Experimental results on three widely used datasets demonstrate that CKDH can outperform some state-of-the-art methods, by delivering the best result consistently. Jiaxing Li 0009, Wai Keung Wong, Xiaozhao Fang, Shengli Xie 0001, Yong Xu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Low-Rank Correlation Learning for Unsupervised Domain AdaptationabstractIn unsupervised domain adaptation (UDA), negative transfer is one of the most challenging problems. Due to complex environments, the used domain data are always corrupted by noise or outliers in many applications. If the noisy data are directly used for domain adaptation, the disturbances and negative influence of the noise are also shifted for the target tasks. Thus, preventing disturbances and negative effects caused by noise are key problems in UDA that need to be addressed. In this article, a low-rank correlation learning (LRCL) method is proposed for UDA. In LRCL, the noisy domain data are recovered by low-rank learning; then both domain data are cleaned. Hence, the disturbances and negative effects of the noise are prevented. The maximized correlated features of the clean data from the source and target domains are learned by a novel correlation regularization term in a latent common space. LRCL also reduces the distribution difference of the learned clean source and target data by constructing a reconstruction term, in which the clean target data are linearly represented by the clean source data. To explore the temporal and structural information of the data, we further extend LRCL into a graph case and propose graph LRCL (GLRCL). Extensive experiments have been conducted on several public data benchmarks, and the experimental results demonstrate that our methods can effectively prevent negative transfer and obtain better classification outcomes than other compared approaches. Yuwu Lu, Wai Keung Wong, Chun Yuan 0003, Zhihui Lai 0001, Xuelong Li 0001 |
IEEE Trans. Multim. | 2 |
| 2024 | Correlation-Guided Distribution and Geometry Alignments for Heterogeneous Domain AdaptationabstractIn this paper, we present a novel approach named correlation-guided distribution and geometry alignments (CDGA) for heterogeneous domain adaptation. Unlike existing methods that typically combine feature alignment and domain alignment into a single objective function, our proposed CDGA separates the two alignments into distinct steps. The two adaptation steps are: paired canonical correlation analysis (PCCA) and distribution and geometry alignments (DGA). In the PCCA step, CDGA focuses on maximizing the within-category correlation between source and target samples to produce the dimension-aligned feature representations for the next adaptation step. In the DGA step, CDGA is responsible for learning a classifier that incorporates both distribution and geometry alignments. Furthermore, during this step, the highly confident pseudo labeled samples are carefully selected for the next iteration of PCCA, establishing a beneficial coupling between PCCA and DGA to improve the adaptation performance in an iterative manner. Experimental results on various visual cross-domain benchmarks demonstrate that CDGA achieves remarkable performance compared to the existing shallow heterogeneous domain adaptation methods and even exhibits superiority over the state-of-the-art neural network-based approaches. Wai Keung Wong, Dewei Lin, Yuwu Lu, Jiajun Wen 0001, Zhihui Lai 0001, Xuelong Li 0001 |
IEEE Trans. Multim. | 1 |
| 2024 | Learning Structured Relation Embeddings for Fine-Grained Fashion Attribute RecognitionabstractFashion attribute recognition is a not-new topic, but rather a core task in understanding fashion from the perspective of computer vision. This article proposes a structured relation-aware network (sRA-Net), which exploits multiple hidden relations in fashion images to enrich and achieve accurate attribute representations to boost the performance of fashion attribute recognition. Specifically, it deconstructs the features of a clothing fashion item into three levels, including low-level attribute-related image region information, mid-level attribute dependency information, and high-level clothing look information. To learn these multi-relational embeddings, we present three relation-aware attention mechanisms. The attribute attention mechanism describes the relationship among different attribute vectors through self-attention and uses the attention map to update the attribute embedding. Then, the spatial attention mechanism associates the attribute with the image features and enhances the attribute embedding by leveraging the attribute-related image region. Finally, the channel attention mechanism selects attribute-related image feature channels to obtain a more fine-grained attribute embedding. Furthermore, we introduce structure-aware embedding to constrain attribute recognition in images from a global perspective by identifying the inner structure of the clothing. Without bells and whistles, sRA-Net outperforms all state-of-the-art attribute recognition methods on two mainstream fashion attribute datasets, namely the DeepFashion-C dataset and iFashion-Attribute dataset, with over 1%-3% improvement. Shumin Zhu, Xingxing Zou, Jianjun Qian, Wai Keung Wong |
IEEE Trans. Multim. | 4 |
| 2023 | MVCINN: Multi-View Diabetic Retinopathy Detection Using a Deep Cross-Interaction Neural NetworkabstractDiabetic retinopathy (DR) is the main cause of irreversible blindness for working-age adults. The previous models for DR detection have difficulties in clinical application. The main reason is that most of the previous methods only use single-view data, and the single field of view (FOV) only accounts for about 13% of the FOV of the retina, resulting in the loss of most lesion features. To alleviate this problem, we propose a multi-view model for DR detection, which takes full advantage of multi-view images covering almost all of the retinal field. To be specific, we design a Cross-Interaction Self-Attention based Module (CISAM) that interfuses local features extracted from convolutional blocks with long-range global features learned from transformer blocks. Furthermore, considering the pathological association in different views, we use the feature jigsaw to assemble and learn the features of multiple views. Extensive experiments on the latest public multi-view MFIDDR dataset with 34,452 images demonstrate the superiority of our method, which performs favorably against state-of-the-art models. To the best of our knowledge, this work is the first study on the public large-scale multi-view fundus images dataset for DR detection. Xiaoling Luo 0001, Chengliang Liu 0003, Wai Keung Wong, Jie Wen 0001, Xiaopeng Jin, Yong Xu 0001 |
AAAI | 3 |
| 2023 | Personalized Fashion Recommendation via Deep Personality Learning
Dongmei Mo, Xingxing Zou, Wai Keung Wong |
BMVC | 3 |
| 2023 | CLOTH4D: A Dataset for Clothed Human ReconstructionabstractClothed human reconstruction is the cornerstone for creating the virtual world. To a great extent, the quality of recovered avatars decides whether the Metaverse is a passing fad. In this work, we introduce CLOTH4D, a clothed human dataset containing 1,000 subjects with varied appearances, 1,000 3D outfits, and over 100,000 clothed meshes with paired unclothed humans, to fill the gap in large-scale and high-quality 4D clothing data. It enjoys appealing characteristics: 1) Accurate and detailed clothing textured meshes-all clothing items are manually created and then simulated in professional software, strictly following the general standard in fashion design. 2) Separated textured clothing and under-clothing body meshes, closer to the physical world than single-layer raw scans. 3) Clothed human motion sequences simulated given a set of 289 actions, covering fundamental and complicated dynamics. Upon CLOTH4D, we novelly designed a series of temporally-aware metries to evaluate the temporal stability of the generated 3D human meshes, which has been over-looked previously. Moreover, by assessing and retraining current state-of-the-art clothed human reconstruction methods, we reveal insights, present improved performance, and propose potential future research directions, confirming our dataset's advancement. The dataset is available at. Xingxing Zou, Xintong Han, Wai Keung Wong |
CVPR | 3 |
| 2023 | Towards private stylists via personalized compatibility learning
Dongmei Mo, Xingxing Zou, Kaicheng Pang, Wai Keung Wong |
Expert Syst. Appl. | 4 |
| 2023 | Binary multi-view clustering with spectral embedding
Zeqi Ma, Wai Keung Wong |
Neurocomputing | 2 |
| 2023 | Guided Discrimination and Correlation Subspace Learning for Domain AdaptationabstractAs a branch of transfer learning, domain adaptation leverages useful knowledge from a source domain to a target domain for solving target tasks. Most of the existing domain adaptation methods focus on how to diminish the conditional distribution shift and learn invariant features between different domains. However, two important factors are overlooked by most existing methods: 1) the transferred features should be not only domain invariant but also discriminative and correlated, and 2) negative transfer should be avoided as much as possible for the target tasks. To fully consider these factors in domain adaptation, we propose a guided discrimination and correlation subspace learning (GDCSL) method for cross-domain image classification. GDCSL considers the domain-invariant, category-discriminative, and correlation learning of data. Specifically, GDCSL introduces the discriminative information associated with the source and target data by minimizing the intraclass scatter and maximizing the interclass distance. By designing a new correlation term, GDCSL extracts the most correlated features from the source and target domains for image classification. The global structure of the data can be preserved in GDCSL because the target samples are represented by the source samples. To avoid negative transfer issues, we use a sample reweighting method to detect target samples with different confidence levels. A semi-supervised extension of GDCSL (Semi-GDCSL) is also proposed, and a novel label selection scheme is introduced to ensure the correction of the target pseudo-labels. Comprehensive and extensive experiments are conducted on several cross-domain data benchmarks. The experimental results verify the effectiveness of the proposed methods over state-of-the-art domain adaptation methods. Yuwu Lu, Wai Keung Wong, Zhihui Lai 0001, Xuelong Li 0001 |
IEEE Trans. Image Process. | 2 |
| 2023 | OTFace: Hard Samples Guided Optimal Transport Loss for Deep Face RepresentationabstractFace representation in the wild is extremely hard due to the large scale face variations. Some deep convolutional neural networks (CNNs) have been developed to learn discriminative feature by designing properly margin-based losses, which perform well on easy samples but fail on hard samples. Although some methods mainly adjust the weights of hard samples in training stage to improve the feature discrimination, they overlook the distribution property of feature. It is worth noting that the miss-classified hard samples may be corrected from the feature distribution view. To overcome this problem, this paper proposes the hard samples guided optimal transport (OT) loss for deep face representation, OTFace in short. OTFace aims to enhance the performance of hard samples by introducing the feature distribution discrepancy while maintaining the performance on easy samples. Specifically, we embrace triplet scheme to indicate hard sample groups in one mini-batch during training. OT is then used to characterize the distribution differences of features from the high level convolutional layer. Finally, we integrate the margin-based-softmax (e.g. ArcFace or AM-Softmax) and OT together to guide deep CNN learning. Extensive experiments were conducted on several benchmark databases. The quantitative results demonstrate the advantages of the proposed OTFace over state-of-the-art methods. Jianjun Qian, Shumin Zhu, Chaoyu Zhao, Jian Yang 0003, Wai Keung Wong |
IEEE Trans. Multim. | 5 |
| 2023 | Learning Spectrum-Invariance Representation for Cross-Spectral Palmprint RecognitionabstractPalmprint recognition provides a potential solution for noninvasive personal authentication due to its excellent contactless property and user-security, and it has attracted tremendous research interest in recent years. However, most existing methods focus on intraspectral palmprint recognition, which requires gallery and probe images to be captured under similar illumination, and thus significantly limit its practical applications in open environments with variant illuminations. In this study, we present a spectrum-invariant feature learning method for cross-spectral palmprint recognition to address the problem that gallery and probe samples are captured under different spectra. First, the blockwise direction-based ordinal measure vectors are formed to represent the intrinsic information of palmprint images. Then, a unified feature projection is jointly learned to map two different spectra of palmprint images into a common feature space, in which the different spectral features have enhanced discriminative power by enlarging their variances while the intraclass features learned from different spectral images are similar. The proposed method can be easily extended to seek the unified spectrum-invariant representation of multiple spectral palmprint images, making it feasible to perform palmprint recognition crossing one spectrum to multiple spectra. Experimental results on two multispectral palmprint image databases demonstrate the promising effectiveness of the proposed method on cross-spectral palmprint recognition. Lunke Fei, Wai Keung Wong, Shuping Zhao, Jie Wen 0001, Jian Zhu 0001, Yong Xu 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2022 | Dress Well via Fashion Cognitive Learning
Kaicheng Pang, Xingxing Zou, Wai Keung Wong |
BMVC | 3 |
| 2022 | How Good Is Aesthetic Ability of a Fashion Model?abstractWe introduce A100 (Aesthetic 100) to assess the aesthetic ability of the fashion compatibility models. To date, it is the first work to address the AI model's aesthetic ability with detailed characterization based on the professional fashion domain knowledge. A100 has several desirable characteristics: 1. Completeness. It covers all types of standards in the fashion aesthetic system through two tests, namely LAT (Liberalism Aesthetic Test) and AAT (Academicism Aesthetic Test); 2. Reliability. It is training data agnostic and consistent with major indicators. It provides a fair and objective judgment for model comparison. 3. Explainability. Better than all previous indicators, the A100 further identifies essential characteristics of fashion aesthetics, thus showing the model's performance on more fine-grained dimensions, such as Color, Balance, Material, etc. Experimental results prove the advance of the A100 in the aforementioned aspects. All data can be found at https://github.com/AemikaChow/AiDLab-fAshIon-Data. Xingxing Zou, Kaicheng Pang, Wen Zhang 0010, Wai Keung Wong |
CVPR | 4 |
| 2022 | Neural stylist: Towards online styling service
Dongmei Mo, Xingxing Zou, Wai Keung Wong |
Expert Syst. Appl. | 3 |
| 2022 | Joint Optimal Transport With Convex Regularization for Robust Image ClassificationabstractThe critical step of learning the robust regression model from high-dimensional visual data is how to characterize the error term. The existing methods mainly employ the nuclear norm to describe the error term, which are robust against structure noises (e.g., illumination changes and occlusions). Although the nuclear norm can describe the structure property of the error term, global distribution information is ignored in most of these methods. It is known that optimal transport (OT) is a robust distribution metric scheme due to that it can handle correspondences between different elements in the two distributions. Leveraging this property, this article presents a novel robust regression scheme by integrating OT with convex regularization. The OT-based regression with$L_{2} $norm regularization (OTR) is first proposed to perform image classification. The alternating direction method of multipliers is developed to handle the model. To further address the occlusion problem in image classification, the extended OTR (EOTR) model is then presented by integrating the nuclear norm error term with an OTR model. In addition, we apply the alternating direction method of multipliers with Gaussian back substitution to solve EOTR and also provide the complexity and convergence analysis of our algorithms. Experiments were conducted on five benchmark datasets, including illumination changes and various occlusions. The experimental results demonstrate the performance of our robust regression model on biometric image classification against several state-of-the-art regression-based classification methods. Jianjun Qian, Wai Keung Wong, Hengmin Zhang, Jin Xie 0001, Jian Yang 0003 |
IEEE Trans. Cybern. | 2 |
| 2022 | Leveraging Multiple Relations for Fashion Trend Forecasting Based on Social MediaabstractFashion trend forecasting is of great research significance in providing useful suggestions for both fashion companies and fashion lovers. Although various studies have been devoted to tackling this challenging task, they only studied limited fashion elements with highly seasonal or simple patterns, which could hardly reveal the real complex fashion trends. Moreover, the mainstream solutions for this task are still statistical-based and solely focus on time-series data modeling, which limit the forecast accuracy. Towards insightful fashion trend forecasting, previous work[1]proposed to analyze more fine-grained fashion elements which can informatively reveal fashion trends. Specifically, it focused on detailed fashion element trend forecasting for specific user groups based on social media data. In addition, it proposed a neural network-based method, namely KERN, to address the problem of fashion trend modeling and forecasting. In this work, to extend the previous work[1], we propose an improved model named Relation Enhanced Attention Recurrent (REAR) network. Compared to KERN, the REAR model leverages not only the relations among fashion elements, but also those among user groups, thus capturing more types of correlations among various fashion trends. To further improve the performance of long-range trend forecasting, the REAR method devises a sliding temporal attention mechanism, which is able to capture temporal patterns on future horizons better. Extensive experiments and more analysis have been conducted on the FIT[1]and GeoStyle[2]datasets to evaluate the performance of REAR. Experimental and analytical results demonstrate the effectiveness of the proposed REAR model in fashion trend forecasting, which also show the improvement of REAR compared to the KERN. Yujuan Ding, Yunshan Ma 0002, Lizi Liao, Wai Keung Wong, Tat-Seng Chua |
IEEE Trans. Multim. | 4 |
| 2022 | Modeling Instant User Intent and Content-Level Transition for Sequential Fashion RecommendationabstractFashion recommendation, aiming to explore specific user preference in fashion, has become an important research topic for its practical significance to the fashion business sector. However, little work has been done on an important sub-task called sequential fashion recommendation, which aims to capture additional short-term fashion interest of users by modeling the item-to-item transitions. In this paper, we propose a novel Attentional Content-level Translation-based Recommender (ACTR) framework, which simultaneously models the instant user intent of each transition and the intent-specific transition probability. Specifically, we define instant intent with the relationships between adjacent items that the users interacted, which are the three fundamental domain-specific relationships of:match,substituteandothers. To further exploit the characteristics of fashion domain and alleviate the item transition sparsity problem, we augment the item-level transition modeling with multiple sub-transitions using various content-level attributes. An attention mechanism is further devised to effectively aggregate multiple content-level transitions. To the best of our knowledge, this is the first work that specifies the implicit user actions in online fashion shopping with explicit instant intent, which enhances the connectivity of fashion items and boosts the recommendation performance. Extensive experiments on two real-world fashion E-commerce datasets demonstrate the effectiveness of the proposed method in sequential fashion recommendation. Yujuan Ding, Yunshan Ma 0002, Wai Keung Wong, Tat-Seng Chua |
IEEE Trans. Multim. | 3 |
| 2021 | Leveraging Two Types of Global Graph for Sequential Fashion RecommendationabstractSequential fashion recommendation is of great significance in online fashion shopping, which accounts for an increasing portion of either fashion retailing or online e-commerce. The key to building an effective sequential fashion recommendation model lies in capturing two types of patterns: the personal fashion preference of users and the transitional relationships between adjacent items. The two types of patterns are usually related to user-item interaction and item-item transition modeling respectively. However, due to the large sets of users and items as well as the sparse historical interactions, it is difficult to train an effective and efficient sequential fashion recommendation model. To tackle these problems, we propose to leverage two types of global graph, i.e., the user-item interaction graph and item-item transition graph, to obtain enhanced user and item representations by incorporating higher-order connections over the graphs. In addition, we adopt the graph kernel of LightGCN [9] for the information propagation in both graphs and propose a new design for item-item transition graph. Extensive experiments on two established sequential fashion recommendation datasets validate the effectiveness and efficiency of our approach. Yujuan Ding, Yunshan Ma 0002, Wai Keung Wong, Tat-Seng Chua |
ICMR | 3 |
| 2021 | Reproducibility Companion Paper: Knowledge Enhanced Neural Fashion Trend ForecastingabstractThis companion paper supports the replication of the fashion trend forecasting experiments with the KERN (Knowledge Enhanced Recurrent Network) method that we presented in the ICMR 2020. We provide an artifact that allows the replication of the experiments using a Python implementation. The artifact is easy to deploy with simple installation, training and evaluation. We reproduce the experiments conducted in the original paper and obtain similar performance as previously reported. The replication results of the experiments support the main claims in the original paper. Yunshan Ma 0002, Yujuan Ding, Xun Yang 0001, Lizi Liao, Wai Keung Wong, Tat-Seng Chua, Jinyoung Moon, Hong-Han Shuai |
ICMR | 5 |
| 2021 | Dual robust regression for pattern classification
Jianjun Qian, Shumin Zhu, Wai Keung Wong, Hengmin Zhang, Zhihui Lai 0001, Jian Yang 0003 |
Inf. Sci. | 3 |
| 2021 | MVDRNet: Multi-view diabetic retinopathy detection by combining DCNNs and attention mechanisms
Xiaoling Luo 0001, Zuhui Pu, Yong Xu 0001, Wai Keung Wong, Jingyong Su, Xiaoyan Dou, Baikang Ye, Jiying Hu, Lisha Mou |
Pattern Recognit. | 4 |
| 2021 | Weighted Double-Low-Rank Decomposition With Application to Fabric Defect DetectionabstractRecently, many methods based on low-rank representation have been proposed for fabric defect detection. Most of them relax the low-rank decomposition problem to a nuclear norm minimization (NNM) problem to pursue the convexity of the objective function. When solving the standard NNM problem, matrix singular values have to be treated equally. This, however, would be impractical in the scenario of fabric defect detection as the matrix singular values have clear physical meanings, and thus, they should be treated differently. In this article, we propose a weighted double-low-rank decomposition method (WDLRD) to treat the matrix singular values differently by assigning different weights. Thus, the most important/distinguishing characteristics of a fabric image can be preserved. Another difference between WDLRD and the other existing low-rank-based methods is that WDLRD considers a defective fabric image being decomposed to two low-rank matrices, i.e., low-rank defect-free matrix and low-rank defect matrix, as the defect-free and defective regions are usually composed of homogeneous objects that have a high correlation. Besides, WDLRD is more robust for defect detection in various situations by adding a noise term to avoid noise or other interference on the fabric surface. In addition, a defect prior is incorporated into the objective function of WDLRD to guide locating the defective regions. The proposed optimization problem can be easily solved by an iterative algorithm based on augmented Lagrange multipliers. Experimental results on TILDA, periodically patterned fabric, and Textile & Apparel Artificial Intelligence databases show that the proposed WDLRD obtains better performance than state-of-the-art methods in locating the defective regions on fabric images.Note to Practitioners—This article is motivated by the problem that the performance of fabric defect detection in the textile industry is poor. It is necessary to develop an effective method to improve the defect detection accuracy and reduce overall manufacturing cost. Existing automatic defect detection approaches usually contain two stages: first, capture fabric images from the weaving machine and then use a defect detection algorithm in a host computer to conduct a real-time inspection and give an alarm if defects occur. This article focuses on locating defects for given defective images after the procedure of binary classification (which determines an image as defective or defect free). The article proposes an objective function to mathematically interpret the optimization problem between fabric images and predictive defects. The optimal solution can be obtained by employing the alternating direction method of multipliers (ADMMs). The proposed method is described as a new defect detection algorithm. Extensive experiments were conducted to evaluate the algorithm, and the experimental results indicate that the proposed method is superior to many existing fabric defect detection methods. Preliminary experiments suggest that this method is feasible but has not yet been really used in production. In future research, we will collect more fabric images from the textile industry and develop large-scale databases for verifying the proposed method for real-life applications. Dongmei Mo, Wai Keung Wong, Zhihui Lai 0001, Jie Zhou 0009 |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2020 | Knowledge Enhanced Neural Fashion Trend ForecastingabstractFashion trend forecasting is a crucial task for both academia andindustry. Although some efforts have been devoted to tackling this challenging task, they only studied limited fashion elements with highly seasonal or simple patterns, which could hardly reveal thereal fashion trends. Towards insightful fashion trend forecasting,this work focuses on investigating fine-grained fashion element trends for specific user groups. We first contribute a large-scale fashion trend dataset (FIT) collected from Instagram with extracted time series fashion element records and user information. Furthermore, to effectively model the time series data of fashion elements with rather complex patterns, we propose a Knowledge Enhanced Recurrent Network model (KERN) which takes advantage of the capability of deep recurrent neural networks in modeling time series data. Moreover, it leverages internal and external knowledgein fashion domain that affects the time-series patterns of fashion element trends. Such incorporation of domain knowledge further enhances the deep learning model in capturing the patterns of specific fashion elements and predicting the future trends. Extensive experiments demonstrate that the proposed KERN model can effectively capture the complicated patterns of objective fashion elements, therefore making preferable fashion trend forecast. Yunshan Ma 0002, Yujuan Ding, Xun Yang 0001, Lizi Liao, Wai Keung Wong, Tat-Seng Chua |
ICMR | 5 |
| 2020 | The Devil is in the Detail: Deep Feature Based Disguised Face Recognition Method
Shumin Zhu, Jianjun Qian, Yangwei Dong, Wai Keung Wong |
PRCV (2) | 4 |
| 2020 | Discriminative dual-stream deep hashing for large-scale image retrieval
Yujuan Ding, Wai Keung Wong, Zhihui Lai 0001, Zheng Zhang 0006 |
Inf. Process. Manag. | 2 |
| 2020 | Discriminative deep multi-task learning for facial expression recognition
Ruili Wang 0001, Wanting Ji, Ming Zong, Wai Keung Wong, Zhihui Lai 0001, Hexin Lv |
Inf. Sci. | 5 |
| 2020 | Low-rank discriminative regression learning for image classification
Yuwu Lu, Zhihui Lai 0001, Wai Keung Wong, Xuelong Li 0001 |
Neural Networks | 3 |
| 2020 | Bilinear Supervised Hashing Based on 2D Image FeaturesabstractHashing has been recognized as an efficient representation learning method to effectively handle big data due to its low computational complexity and memory cost. Most of the existing hashing methods focus on learning the low-dimensional vectorized binary features based on the high-dimensional raw vectorized features. However, the studies on how to obtain preferable binary codes from the original 2D image features for retrieval is very limited. This paper proposes a bilinear supervised discrete hashing (BSDH) method based on 2D image features which utilizes bilinear projections to binarize the image matrix features such that the intrinsic characteristics in the 2D image space are preserved in the learned binary codes. Meanwhile, the bilinear projection approximation and vectorization binary codes regression are seamlessly integrated together to formulate the final robust learning framework. Furthermore, a discrete optimization strategy is developed to alternatively update each variable for obtaining the high-quality binary codes. In addition, two 2D image features, traditional SURF-based FVLAD feature, and CNN-based AlexConv5 feature are designed for further improving the performance of the proposed BSDH method. The results of extensive experiments conducted on four benchmark datasets show that the proposed BSDH method almost outperforms all competing hashing methods with different input features by different evaluation protocols. Yujuan Ding, Wai Keung Wong, Zhihui Lai 0001, Zheng Zhang 0006 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | Double Relaxed Regression for Image ClassificationabstractThis paper addresses two fundamental problems: 1) learning discriminative model parameters and 2) avoiding over-fitting, which often occurs in regression-based classification tasks. We formulate these two problems in terms of relaxing both the strict binary label matrix and graph regularization term into more flexible forms so that the margins between different classes are enlarged as much as possible and the problem of over-fitting is avoided to some extent. This task is accomplished by the proposed double relaxed regression (DRR) method. The convex problem of DRR is solved efficiently with an iterative procedure. Extensive experiments on synthetic and real world image data sets demonstrate the effectiveness of the proposed method in terms of both classification accuracy and running time. Na Han, Jigang Wu, Xiaozhao Fang, Wai Keung Wong, Yong Xu 0001, Jian Yang 0003, Xuelong Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | Clustering Structure-Induced Robust Multi-View Graph RecoveryabstractGraph based classification methods have been widely applied in the fields of computer vision and machine learning. The quality of the graph highly affects the performance of these methods. The same object is commonly represented by different features, i.e., multi-view features, which leads to multiple graphs corresponding to different features in multi-view learning. However, what kind of graph is important for the task is unknown in advance. Moreover, existing multi-view learning methods become weak in dealing with noisy graphs when the data is corrupted by the noise. In this paper, we address this problem by observing that the noise of each graph has specific structure. Then, based on this observation we propose a robust multi-view graph recovery (RMGR) method in which the specific structure is used to clean the multiple input noisy graphs and these cleaned graphs are simultaneously aggregated into a consensus graph by adaptively assigning great weighted coefficients for important graphs. To make the consensus graph suit classification, the clustering structure is introduced to restrain the rank of Laplacian matrix of the consensus graph such that the number of its connected components is equal to that of clustering. In doing so, the graph is adaptively adjusted during optimization to more accurately partition data. The optimization problem is solved by proposed the iterative update algorithm. Extensive experiments on synthetic and several benchmark data sets show the effectiveness of the proposed method. Wai Keung Wong, Na Han, Xiaozhao Fang, Shanhua Zhan, Jie Wen 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2020 | Robust Flexible Preserving EmbeddingabstractNeighborhood preserving embedding (NPE) has been proposed to encode overall geometry manifold embedding information. However, the class-special structure of the data is destroyed by noise or outliers existing in the data. To address this problem, in this article, we propose a novel embedding approach called robust flexible preserving embedding (RFPE). First, RFPE recovers the noisy data by low-rank learning and obtains clean data. Then, the clean data are used to learn the projection matrix. In this way, the projective learning is totally unaffected by noise or outliers. By encoding a flexible regularization term, RFPE can keep the property of the data points with a nonlinear manifold and be more flexible. RFPE searches the optimal projective subspace for feature extraction. In addition, we also extend the proposed RFPE to a kernel case and propose kernel RFPE (KRFPE). Extensive experiments on six public image databases show the superiority of the proposed methods over other state-of-the-art methods. Yuwu Lu, Wai Keung Wong, Zhihui Lai 0001, Xuelong Li 0001 |
IEEE Trans. Cybern. | 2 |
| 2020 | Study on 2D Feature-Based Hash LearningabstractHashing is an important topic in image processing, as it can help save a considerable amount of storage and computational cost. Recently, inspired by 2D strategies employed in other areas of image processing, such as feature extraction, some 2D-based hashing methods were proposed. Related papers have shown that these methods may have better image retrieval performance in terms of both effectiveness and efficiency. However, the difference in the retrieval performances of hashing methods resulting from different forms of input (1D or 2D) has not been previously studied. Whether the widely used bilinear strategy in 2D-based hashing can truly help improve the retrieval precision has not been investigated in existing research. In this paper, we conduct a comparison study on 1D and 2D feature-based hashing methods and attempt to theoretically and experimentally analyse the differences in using 1D and 2D features in hashing. Furthermore, two new hashing methods are proposed for conducting the comparison experiments. Through a comprehensive study, we obtain three main conclusions in this paper: 1) Linear projection on 1D features and bilinear projection on 2D features are essentially the same. 2) 2D-based hashing methods are obviously more efficient than 1D-based methods for analysing high-dimensional input features. 3) 2D-based hashing methods show generally better performance for solving small sample size problems. Yujuan Ding, Wai Keung Wong, Zhihui Lai 0001, Yudong Chen 0002 |
IEEE Trans. Multim. | 2 |
| 2020 | Jointly Sparse Locality Regression for Image Feature ExtractionabstractThis paper proposes a novel method called Jointly Sparse Locality Regression (JSLR) for feature extraction and selection. JSLR utilizes joint L2,1-norm minimization on regularization term, and also introduces the locality to characterize the local geometric structure of the data. There are three main contributions in JSLR for face recognition. Firstly, it eliminates the drawback in ridge regression and Linear Discriminant Analysis (LDA) that when the number of the classes is too small, not enough projections can be obtained for feature extraction. Secondly, by using the local geometric structure as the regularization term, JSLR is able to preserve local information and find an embedding subspace which can detect the most essential data manifold structure. Moreover, since the L2,1-norm based loss function is robust to outliers in data points, JSLR provides the joint sparsity for robust feature selection. The theoretical connections of the proposed method and the previous regression methods are explored and the convergence of the proposed algorithm is also proved. Experimental evaluation on several well-known data sets shows the merits of the proposed method on feature selection and classification. Dongmei Mo, Zhihui Lai 0001, Xizhao Wang, Wai Keung Wong |
IEEE Trans. Multim. | 4 |
| 2019 | Deep Supervised Hashing With Anchor GraphabstractRecently, a series of deep supervised hashing methods were proposed for binary code learning. However, due to the high computation cost and the limited hardware's memory, these methods will first select a subset from the training set, and then form a mini-batch data to update the network in each iteration. Therefore, the remaining labeled data cannot be fully utilized and the model cannot directly obtain the binary codes of the entire training set for retrieval. To address these problems, this paper proposes an interesting regularized deep model to seamlessly integrate the advantages of deep hashing and efficient binary code learning by using the anchor graph. As such, the deep features and label matrix can be jointly used to optimize the binary codes, and the network can obtain more discriminative feedback from the linear combinations of the learned bits. Moreover, we also reveal the algorithm mechanism and its computation essence. Experiments on three large-scale datasets indicate that the proposed method achieves better retrieval performance with less training time compared to previous deep hashing methods. Yudong Chen 0002, Zhihui Lai 0001, Yujuan Ding, Kaiyi Lin, Wai Keung Wong |
ICCV | 5 |
| 2019 | Granular maximum decision entropy-based monotonic uncertainty measure for attribute reduction
Can Gao, Zhihui Lai 0001, Jie Zhou 0009, Jiajun Wen 0001, Wai Keung Wong |
Int. J. Approx. Reason. | 5 |
| 2019 | Binary sparse signal recovery algorithms based on logic observation
Xiao-Li Hu, Jiajun Wen 0001, Zhihui Lai 0001, Wai Keung Wong, LinLin Shen |
Pattern Recognit. | 4 |
| 2019 | Generalized Robust Regression for Jointly Sparse Subspace LearningabstractRidge regression is widely used in multiple variable data analysis. However, in very high-dimensional cases such as image feature extraction and recognition, conventional ridge regression or its extensions have the small-class problem, that is, the number of the projections obtained by ridge regression is limited by the number of the classes. In this paper, we proposed a novel method called generalized robust regression (GRR) for jointly sparse subspace learning which can address the problem. GRR not only imposes L2,1-norm penalty on both loss function and regularization term to guarantee the joint sparsity and the robustness to outliers for effective feature selection, but also utilizes L2,1-norm as the measurement to take the intrinsic local geometric structure of the data into consideration to improve the performance. Moreover, by incorporating the elastic factor on the loss function, GRR can enhance the robustness to obtain more projections for feature selection or classification. To obtain the optimal solution of GRR, an iterative algorithm was proposed and the convergence was also proved. Experiments on six wellknown data sets demonstrate the merits of the proposed method. The result indicates that GRR is a robust and efficient regression method for face recognition. Zhihui Lai 0001, Dongmei Mo, Jiajun Wen 0001, LinLin Shen, Wai Keung Wong |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2019 | Horizontal and Vertical Nuclear Norm-Based 2DLDA for Image Representationabstract2-D linear discriminant analysis (2DLDA) has been widely used in pattern recognition and image classification. 2DLDA selects discriminative features from the up and left corner of images. However, 2DLDA uses the Frobenius norm (F-norm), which is sensitive to noise or outliers in data, as a metric. In this paper, we propose a novel framework, called horizontal and vertical nuclear norm-based 2DLDA (HVNN-2DLDA) for image representation. In the proposed framework, HVNN-2DLDA methods (i.e., HNN-2DLDA and VNN-2DLDA) are proposed, and both use the nuclear norm as a criterion. The nuclear norm can provide more structure and global information for the reconstruction of noisy images. HNN-2DLDA and VNN-2DLDA represent images in the row and column directions, respectively. In addition, by combining the row and column directions, we propose a bilateral nuclear norm-based 2DLDA method called BNN-2DLDA. The advantage of BNN-2DLDA over HNN-2DLDA and VNN-2DLDA is that an image sample can be represented by both the row and the column directions instead of only the row or column direction. HVNN-2DLDA learns a set of local optimal projection vectors by maximizing the ratio of the nuclear norm of the between-class scatter matrix and the nuclear norm of the within-class scatter matrix. To verify the robustness and recognition performance in image classification of HVNN-2DLDA, six public image databases are used for experiments. The experimental results demonstrate the effectiveness and the feasibility of the proposed framework. Yuwu Lu, Chun Yuan 0003, Zhihui Lai 0001, Xuelong Li 0001, David Zhang 0001, Wai Keung Wong |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2019 | Low-Rank 2-D Neighborhood Preserving Projection for Enhanced Robust Image Representationabstract2-D neighborhood preserving projection (2DNPP) uses 2-D images as feature input instead of 1-D vectors used by neighborhood preserving projection (NPP). 2DNPP requires less computation time than NPP. However, both NPP and 2DNPP use the L2norm as a metric, which is sensitive to noise in data. In this paper, we proposed a novel NPP method called low-rank 2DNPP (LR-2DNPP). This method divided the input data into a component part that encoded low-rank features, and an error part that ensured the noise was sparse. Then, a nearest neighbor graph was learned from the clean data using the same procedure as 2DNPP. To ensure that the features learned by LR-2DNPP were optimal for classification, we combined the structurally incoherent learning and low-rank learning with NPP to form a unified model called discriminative LR-2DNPP (DLR2DNPP). By encoding the structural incoherence of the learned clean data, DLR-2DNPP could enhance the discriminative ability for feature extraction. Theoretical analyses on the convergence and computational complexity of LR-2DNPP and DLR-2DNPP were presented in details. We used seven public image databases to verify the performance of the proposed methods. The experimental results showed the effectiveness of our methods for robust image representation. Yuwu Lu, Zhihui Lai 0001, Xuelong Li 0001, Wai Keung Wong, Chun Yuan 0003, David Zhang 0001 |
IEEE Trans. Cybern. | 4 |
| 2019 | Scalable Supervised Asymmetric Hashing With Semantic and Latent Factor EmbeddingabstractCompact hash code learning has been widely applied to fast similarity search owing to its significantly reduced storage and highly efficient query speed. However, it is still a challenging task to learn discriminative binary codes for perfectly preserving the full pairwise similarities embedded in the high-dimensional real-valued features, such that the promising performance can be guaranteed. To overcome this difficulty, in this paper, we propose a novel scalable supervised asymmetric hashing (SSAH) method, which can skillfully approximate the full-pairwise similarity matrix based on maximum asymmetric inner product of two different non-binary embeddings. In particular, to comprehensively explore the semantic information of data, the supervised label information and the refined latent feature embedding are simultaneously considered to construct the high-quality hashing function and boost the discriminant of the learned binary codes. Specifically, SSAH learns two distinctive hashing functions in conjunction of minimizing the regression loss on the semantic label alignment and the encoding loss on the refined latent features. More importantly, instead of using only part of similarity correlations of data, the full-pairwise similarity matrix is directly utilized to avoid information loss and performance degeneration, and its cumbersome computation complexity on n ×n matrix can be dexterously manipulated during the optimization phase. Furthermore, an efficient alternating optimization scheme with guaranteed convergence is designed to address the resulting discrete optimization problem. The encouraging experimental results on diverse benchmark datasets demonstrate the superiority of the proposed SSAH method in comparison with many recently proposed hashing algorithms. Zheng Zhang 0006, Zhihui Lai 0001, Zi Huang, Wai Keung Wong, Guosen Xie, Li Liu 0004, Ling Shao 0001 |
IEEE Trans. Image Process. | 4 |
| 2019 | Locally Joint Sparse Marginal Embedding for Feature ExtractionabstractClassical linear discriminant analysis (LDA) has the limitation that it requires the within-class scatter matrix to be nonsingular so that it can perform eigen-decomposition to obtain optimal solutions. To break through this limitation, many methods based on LDA have been proposed. However, these methods are either sensitive to outliers or lack joint sparsity for effective feature extraction. To release these problems, this paper proposes a locally joint sparse marginal embedding (LJSME) method. LJSME reconstructs the scatter matrices and utilizes the locality graph to weigh each pair of data, such that it is robust to outliers and able to preserve the neighborhood relationship of the data. Moreover, LJSME can easily avoid the small sample-size problem by a maximum margin criterion and obtain joint sparsity for effective feature extraction by using joint sparse regularization. The comprehensive analysis between the proposed LJSME and the related methods is presented, which indicates the advantages of the proposed method. A series of experiments was conducted to evaluate the performance of LJSME when compared with the state-of-the-art methods. The MATLAB code of LJSME can be downloaded fromhttps://github.com/TungmeeMo/LJSME.git. Dongmei Mo, Zhihui Lai 0001, Wai Keung Wong |
IEEE Trans. Multim. | 3 |
| 2019 | Flexible Affinity Matrix Learning for Unsupervised and Semisupervised ClassificationabstractIn this paper, we propose a unified model called flexible affinity matrix learning (FAML) for unsupervised and semisupervised classification by exploiting both the relationship among data and the clustering structure simultaneously. To capture the relationship among data, we exploit the self-expressiveness property of data to learn a structured matrix in which the structures are induced by different norms. A rank constraint is imposed on the Laplacian matrix of the desired affinity matrix, so that the connected components of data are exactly equal to the cluster number. Thus, the clustering structure is explicit in the learned affinity matrix. By making the estimated affinity matrix approximate the structured matrix during the learning procedure, FAML allows the affinity matrix itself to be adaptively adjusted such that the learned affinity matrix can well capture both the relationship among data and the clustering structure. Thus, FAML has the potential to perform better than other related methods. We derive optimization algorithms to solve the corresponding problems. Extensive unsupervised and semisupervised classification experiments on both synthetic data and real-world benchmark data sets show that the proposed FAML consistently outperforms the state-of-the-art methods. Xiaozhao Fang, Na Han, Wai Keung Wong, Shaohua Teng, Jigang Wu, Shengli Xie 0001, Xuelong Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2018 | An integrated optimisation algorithm for feature extraction, dictionary learning and classification
Yan Cui 0007, Jielin Jiang, Zhihui Lai 0001, Wai Keung Wong, Zuojin Hu |
Neurocomputing | 4 |
| 2018 | Kernelized random KISS metric learning for person re-identification
Cairong Zhao, Yipeng Chen, Xuekuan Wang, Wai Keung Wong, Duoqian Miao 0001, Jingsheng Lei |
Neurocomputing | 4 |
| 2018 | Low-rank and sparse embedding for dimensionality reduction
Na Han, Jigang Wu, Yingyi Liang, Xiaozhao Fang, Wai Keung Wong, Shaohua Teng |
Neural Networks | 5 |
| 2018 | Supervised discrete discriminant hashing for image retrieval
Yan Cui 0007, Jielin Jiang, Zhihui Lai 0001, Zuojin Hu, Wai Keung Wong |
Pattern Recognit. | 5 |
| 2018 | On uniqueness of sparse signal recovery
Xiao-Li Hu, Jiajun Wen 0001, Wai Keung Wong, Le Tong, Jinrong Cui |
Signal Process. | 3 |
| 2018 | New semi-supervised classification using a multi-modal feature joint L21-norm based sparse representation
Yan Cui 0007, Jielin Jiang, Zhihui Lai 0001, Zuojin Hu, Yuquan Jiang, Wai Keung Wong |
Signal Process. Image Commun. | 6 |
| 2018 | Learning Parts-Based and Global Representation for Image ClassificationabstractNonnegative matrix factorization (NMF), known as a famous matrix factorization technique, has been widely used in pattern recognition and computer vision. NMF represents the input data matrix as a product of two nonnegative factors. As NMF is based on the Euclidean distance, which is sensitive to noise or errors in the data, some robust NMF methods are proposed. Mainly focusing on parts-based representation, these robust NMF methods often neglect global representation of data. In fact, the global geometry information of data is more robust than the local information about the noisy data in terms of image classification. In order to effectively improve the robustness of NMF and learn part-based and global representation of the data, a novel method low-rank nonnegative factorization (LRNF) is proposed in this paper. First, we assume that the data are grossly corrupted, and the$L_{1} $norm is used as a sparse constraint on the assumed noise matrix. Then, LRNF learns a low-rank matrix with the global representation ability. Finally, we make a nonnegative factorization of the learned low-rank matrix. We can obtain a base matrix, which preserves locality and globality properties of the data in the meantime. Extensive experiments have been conducted on nine real-world image databases to verify the performance of the proposed LRNF method by comparing with the state-of-the-art algorithms on robust dimensionality reduction. Yuwu Lu, Zhihui Lai 0001, Xuelong Li 0001, David Zhang 0001, Wai Keung Wong, Chun Yuan 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2018 | Robust Discriminant Regression for Feature ExtractionabstractRidge regression (RR) and its extended versions are widely used as an effective feature extraction method in pattern recognition. However, the RR-based methods are sensitive to the variations of data and can learn only limited number of projections for feature extraction and recognition. To address these problems, we propose a new method called robust discriminant regression (RDR) for feature extraction. In order to enhance the robustness, the L2,1-norm is used as the basic metric in the proposed RDR. The designed robust objective function in regression form can be solved by an iterative algorithm containing an eigenfunction, through which the optimal orthogonal projections of RDR can be obtained by eigen decomposition. The convergence analysis and computational complexity are presented. In addition, we also explore the intrinsic connections and differences between the RDR and some previous methods. Experiments on some well-known databases show that RDR is superior to the classical and very recent proposed methods reported in the literature, no matter the L2-norm or the L2,1-norm-based regression methods. The code of this paper can be downloaded from http://www.scholat.com/laizhihui. Zhihui Lai 0001, Dongmei Mo, Wai Keung Wong, Yong Xu 0001, Duoqian Miao 0001, David Zhang 0001 |
IEEE Trans. Cybern. | 3 |
| 2018 | Jointly Sparse Hashing for Image RetrievalabstractRecently, hash learning attracts great attentions since it can obtain fast image retrieval on large-scale datasets by using a series of discriminative binary codes. The popular methods include manifold-based hashing methods, which aim to learn the binary codes by embedding the original high-dimensional data into low-dimensional intrinsic subspace. However, most of these methods tend to relax the discrete constraint to compute the final binary codes in an easier way. Therefore, the information loss will increase. In this paper, we propose a novel jointly sparse regression model to minimize the locality information loss and obtain jointly sparse hashing method. The proposed model integrates locality, joint sparsity and rotation operation together with a seamless formulation. Thus, the drawback in previous methods using two separated and independent stages such as PCA-ITQ and the similar methods can be addressed. Moreover, since we introduce the joint sparsity, the feature extraction and jointly sparse feature selection can also be realized in a single projection operation, which has the potentials to select more discriminant features. The convergence of the proposed algorithm is proved, and the essences of the iterative procedures are also revealed. The experimental results on large-scale datasets demonstrate the performance of the proposed method. Zhihui Lai 0001, Yudong Chen 0002, Wai Keung Wong, Fumin Shen |
IEEE Trans. Image Process. | 4 |
| 2018 | Low-Rank Linear Embedding for Image RecognitionabstractLocality preserving projections (LPP) has been widely studied and extended in recent years, because of its promising performance in feature extraction. In this paper, we propose a modified version of the LPP by constructing a novel regression model. To improve the performance of the model, we impose a low-rank constraint on the regression matrix to discover the latent relations between different neighbors. By using the L2,1-norm as a metric for the loss function, we can further minimize the reconstruction error and derive a robust model. Furthermore, the L2,1-norm regularization term is added to obtain a jointly sparse regression matrix for feature selection. An iterative algorithm with guaranteed convergence is designed to solve the optimization problem. To validate the recognition efficiency, we apply the algorithm to a series of benchmark datasets containing face and character images for feature extraction. The experimental results show that the proposed method is better than some existing methods. The code of this paper can be downloaded from http://www.scholat.com/laizhihui. Yudong Chen 0002, Zhihui Lai 0001, Wai Keung Wong, LinLin Shen, Qinghua Hu |
IEEE Trans. Multim. | 3 |
| 2018 | Approximate Low-Rank Projection Learning for Feature ExtractionabstractFeature extraction plays a significant role in pattern recognition. Recently, many representation-based feature extraction methods have been proposed and achieved successes in many applications. As an excellent unsupervised feature extraction method, latent low-rank representation (LatLRR) has shown its power in extracting salient features. However, LatLRR has the following three disadvantages: 1) the dimension of features obtained using LatLRR cannot be reduced, which is not preferred in feature extraction; 2) two low-rank matrices are separately learned so that the overall optimality may not be guaranteed; and 3) LatLRR is an unsupervised method, which by far has not been extended to the supervised scenario. To this end, in this paper, we first propose to use two different matrices to approximate the low-rank projection in LatLRR so that the dimension of obtained features can be reduced, which is more flexible than original LatLRR. Then, we treat the two low-rank matrices in LatLRR as a whole in the process of learning. In this way, they can be boosted mutually so that the obtained projection can extract more discriminative features. Finally, we extend LatLRR to the supervised scenario by integrating feature extraction with the ridge regression. Thus, the process of feature extraction is closely related to the classification so that the extracted features are discriminative. Extensive experiments are conducted on different databases for unsupervised and supervised feature extraction, and very encouraging results are achieved in comparison with many state-of-the-arts methods. Xiaozhao Fang, Na Han, Jigang Wu, Yong Xu 0001, Jian Yang 0003, Wai Keung Wong, Xuelong Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2018 | Robust Latent Subspace Learning for Image ClassificationabstractThis paper proposes a novel method, called robust latent subspace learning (RLSL), for image classification. We formulate an RLSL problem as a joint optimization problem over both the latent SL and classification model parameter predication, which simultaneously minimizes: 1) the regression loss between the learned data representation and objective outputs and 2) the reconstruction error between the learned data representation and original inputs. The latent subspace can be used as a bridge that is expected to seamlessly connect the origin visual features and their class labels and hence improve the overall prediction performance. RLSL combines feature learning with classification so that the learned data representation in the latent subspace is more discriminative for classification. To learn a robust latent subspace, we use a sparse item to compensate error, which helps suppress the interference of noise via weakening its response during regression. An efficient optimization algorithm is designed to solve the proposed optimization problem. To validate the effectiveness of the proposed RLSL method, we conduct experiments on diverse databases and encouraging recognition results are achieved compared with many state-of-the-arts methods. Xiaozhao Fang, Shaohua Teng, Zhihui Lai 0001, Zhaoshui He, Shengli Xie 0001, Wai Keung Wong |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2018 | Regularized Label Relaxation Linear RegressionabstractLinear regression (LR) and some of its variants have been widely used for classification problems. Most of these methods assume that during the learning phase, the training samples can be exactly transformed into a strict binary label matrix, which has too little freedom to fit the labels adequately. To address this problem, in this paper, we propose a novel regularized label relaxation LR method, which has the following notable characteristics. First, the proposed method relaxes the strict binary label matrix into a slack variable matrix by introducing a nonnegative label relaxation matrix into LR, which provides more freedom to fit the labels and simultaneously enlarges the margins between different classes as much as possible. Second, the proposed method constructs the class compactness graph based on manifold learning and uses it as the regularization item to avoid the problem of overfitting. The class compactness graph is used to ensure that the samples sharing the same labels can be kept close after they are transformed. Two different algorithms, which are, respectively, based on -norm and -norm loss functions are devised. These two algorithms have compact closed-form solutions in each iteration so that they are easily implemented. Extensive experiments show that these two algorithms outperform the state-of-the-art algorithms in terms of the classification accuracy and running time. Xiaozhao Fang, Yong Xu 0001, Xuelong Li 0001, Zhihui Lai 0001, Wai Keung Wong, Bingwu Fang |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2017 | Multiple metric learning based on bar-shape descriptor for person re-identification
Cairong Zhao, Xuekuan Wang, Wai Keung Wong, Wei-Shi Zheng 0001, Jian Yang 0003, Duoqian Miao 0001 |
Pattern Recognit. | 3 |
| 2017 | Directional Gaussian Model for Automatic Speeding Event DetectionabstractThis paper proposes a velocity learning method based on the directional Gaussian model to detect speeding events in surveillance scenarios. The proposed method is of an uncalibrated type, but yet has considered the influences of projective transformation on estimating the motion velocities in an image plane, which is convenient and feasible to use in real application. We have analyzed the theory that the velocity in the image plane varies due to the changes of the direction of the moving object with constant velocity in a real world plane. With the support of this theory, we propose to learn the velocities calculated on a certain position in different direction bins to tolerate the effect of projective transformation on velocity modeling. To facilitate the whole framework, two key issues have to be addressed. First, we have designed an improved Fisher model to optimize the direction bins, which reflect the major moving directions in a scenario. Second, we have adopted a weighted sampling strategy and surface fitting to solve the lack of sample problem during the learning process. Experiments conducted on real surveillance videos show that the proposed method can obtain competitive results compared with the state-of-the-art methods. Jiajun Wen 0001, Zhihui Lai 0001, Zhong Ming 0001, Wai Keung Wong, Zuofeng Zhong |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2017 | Low-Rank Embedding for Robust Image Feature ExtractionabstractRobustness to noises, outliers, and corruptions is an important issue in linear dimensionality reduction. Since the sample-specific corruptions and outliers exist, the class-special structure or the local geometric structure is destroyed, and thus, many existing methods, including the popular manifold learning- based linear dimensionality methods, fail to achieve good performance in recognition tasks. In this paper, we focus on the unsupervised robust linear dimensionality reduction on corrupted data by introducing the robust low-rank representation (LRR). Thus, a robust linear dimensionality reduction technique termed low-rank embedding (LRE) is proposed in this paper, which provides a robust image representation to uncover the potential relationship among the images to reduce the negative influence from the occlusion and corruption so as to enhance the algorithm's robustness in image feature extraction. LRE searches the optimal LRR and optimal subspace simultaneously. The model of LRE can be solved by alternatively iterating the argument Lagrangian multiplier method and the eigendecomposition. The theoretical analysis, including convergence analysis and computational complexity, of the algorithms is presented. Experiments on some well-known databases with different corruptions show that LRE is superior to the previous methods of feature extraction, and therefore, it indicates the robustness of the proposed method. The code of this paper can be downloaded from http://www.scholat.com/laizhihui. Wai Keung Wong, Zhihui Lai 0001, Jiajun Wen 0001, Xiaozhao Fang, Yuwu Lu |
IEEE Trans. Image Process. | 1 |
| 2017 | Nuclear Norm-Based 2DLPP for Image ClassificationabstractTwo-dimensional locality preserving projections (2DLPP) that use 2D image representation in preserving projection learning can preserve the intrinsic manifold structure and local information of data. However, 2DLPP is based on the Euclidean distance, which is sensitive to noise and outliers in data. In this paper, we propose a novel locality preserving projection method called nuclear norm-based two-dimensional locality preserving projections (NN-2DLPP). First, NN-2DLPP recovers the noisy data matrix through low-rank learning. Second, noise in data is removed and the learned clean data points are projected on a new subspace. Without the disturbance of noise, data points belonging to the same class are kept as close to each other as possible in the new projective subspace. Experimental results on six public image databases with face recognition, object classification, and handwritten digit recognition tasks demonstrated the effectiveness of the proposed method. Yuwu Lu, Chun Yuan 0003, Zhihui Lai 0001, Xuelong Li 0001, Wai Keung Wong, David Zhang 0001 |
IEEE Trans. Multim. | 5 |
| 2016 | Differential evolution-based optimal Gabor filter model for fabric inspection
Le Tong, Wai Keung Wong, C. K. Kwong 0001 |
Neurocomputing | 2 |
| 2016 | Robust Semi-Supervised Subspace Clustering via Non-Negative Low-Rank RepresentationabstractLow-rank representation (LRR) has been successfully applied in exploring the subspace structures of data. However, in previous LRR-based semi-supervised subspace clustering methods, the label information is not used to guide the affinity matrix construction so that the affinity matrix cannot deliver strong discriminant information. Moreover, these methods cannot guarantee an overall optimum since the affinity matrix construction and subspace clustering are often independent steps. In this paper, we propose a robust semi-supervised subspace clustering method based on non-negative LRR (NNLRR) to address these problems. By combining the LRR framework and the Gaussian fields and harmonic functions method in a single optimization problem, the supervision information is explicitly incorporated to guide the affinity matrix construction and the affinity matrix construction and subspace clustering are accomplished in one step to guarantee the overall optimum. The affinity matrix is obtained by seeking a non-negative low-rank matrix that represents each sample as a linear combination of others. We also explicitly impose the sparse constraint on the affinity matrix such that the affinity matrix obtained by NNLRR is non-negative low-rank and sparse. We introduce an efficient linearized alternating direction method with adaptive penalty to solve the corresponding optimization problem. Extensive experimental results demonstrate that NNLRR is effective in semi-supervised subspace clustering and robust to different types of noise than other state-of-the-art methods. Xiaozhao Fang, Yong Xu 0001, Xuelong Li 0001, Zhihui Lai 0001, Wai Keung Wong |
IEEE Trans. Cybern. | 5 |
| 2016 | Approximate Orthogonal Sparse Embedding for Dimensionality ReductionabstractLocally linear embedding (LLE) is one of the most well-known manifold learning methods. As the representative linear extension of LLE, orthogonal neighborhood preserving projection (ONPP) has attracted widespread attention in the field of dimensionality reduction. In this paper, a unified sparse learning framework is proposed by introducing the sparsity or L1-norm learning, which further extends the LLE-based methods to sparse cases. Theoretical connections between the ONPP and the proposed sparse linear embedding are discovered. The optimal sparse embeddings derived from the proposed framework can be computed by iterating the modified elastic net and singular value decomposition. We also show that the proposed model can be viewed as a general model for sparse linear and nonlinear (kernel) subspace learning. Based on this general model, sparse kernel embedding is also proposed for nonlinear sparse feature extraction. Extensive experiments on five databases demonstrate that the proposed sparse learning framework performs better than the existing subspace learning algorithm, particularly in the cases of small sample sizes. Zhihui Lai 0001, Wai Keung Wong, Yong Xu 0001, Jian Yang 0003, David Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2015 | Sparse nonlocal priors based two-phase approach for mixed noise removal
Jielin Jiang, Jian Yang 0003, Yan Cui 0007, Wai Keung Wong, Zhihui Lai 0001 |
Signal Process. | 4 |
| 2015 | Joint Tensor Feature Analysis For Visual Object RecognitionabstractTensor-based object recognition has been widely studied in the past several years. This paper focuses on the issue of joint feature selection from the tensor data and proposes a novel method called joint tensor feature analysis (JTFA) for tensor feature extraction and recognition. In order to obtain a set of jointly sparse projections for tensor feature extraction, we define the modified within-class tensor scatter value and the modified between-class tensor scatter value for regression. The k-mode optimization technique and the L(2,1)-norm jointly sparse regression are combined together to compute the optimal solutions. The convergent analysis, computational complexity analysis and the essence of the proposed method/model are also presented. It is interesting to show that the proposed method is very similar to singular value decomposition on the scatter matrix but with sparsity constraint on the right singular value matrix or eigen-decomposition on the scatter matrix with sparse manner. Experimental results on some tensor datasets indicate that JTFA outperforms some well-known tensor feature extraction and selection algorithms. Wai Keung Wong, Zhihui Lai 0001, Yong Xu 0001, Jiajun Wen 0001, Chu Po Ho |
IEEE Trans. Cybern. | 1 |
| 2015 | Learning a Nonnegative Sparse Graph for Linear RegressionabstractPrevious graph-based semisupervised learning (G-SSL) methods have the following drawbacks: 1) they usually predefine the graph structure and then use it to perform label prediction, which cannot guarantee an overall optimum and 2) they only focus on the label prediction or the graph structure construction but are not competent in handling new samples. To this end, a novel nonnegative sparse graph (NNSG) learning method was first proposed. Then, both the label prediction and projection learning were integrated into linear regression. Finally, the linear regression and graph structure learning were unified within the same framework to overcome these two drawbacks. Therefore, a novel method, named learning a NNSG for linear regression was presented, in which the linear regression and graph learning were simultaneously performed to guarantee an overall optimum. In the learning process, the label information can be accurately propagated via the graph structure so that the linear regression can learn a discriminative projection to better fit sample labels and accurately classify new samples. An effective algorithm was designed to solve the corresponding optimization problem with fast convergence. Furthermore, NNSG provides a unified perceptiveness for a number of graph-based learning methods and linear regression methods. The experimental results showed that NNSG can obtain very high classification accuracy and greatly outperforms conventional G-SSL methods, especially some conventional graph construction methods. Xiaozhao Fang, Yong Xu 0001, Xuelong Li 0001, Zhihui Lai 0001, Wai Keung Wong |
IEEE Trans. Image Process. | 5 |
| 2015 | Stochastic Stability of Delayed Neural Networks With Local Impulsive EffectsabstractIn this paper, the stability problem is studied for a class of stochastic neural networks (NNs) with local impulsive effects. The impulsive effects considered can be not only nonidentical in different dimensions of the system state but also various at distinct impulsive instants. Hence, the impulses here can encompass several typical impulses in NNs. The aim of this paper is to derive stability criteria such that stochastic NNs with local impulsive effects are exponentially stable in mean square. By means of the mathematical induction method, several easy-to-check conditions are obtained to ensure the mean square stability of NNs. Three examples are given to show the effectiveness of the proposed stability criterion. Wenbing Zhang, Yang Tang 0001, Wai Keung Wong, Qingying Miao |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2014 | Optimal Feature Selection for Robust Classification via l2, 1-Norms RegularizationabstractThis paper aims to explore the optimal feature selection with dimensionality reduction and jointly sparse representation scheme for classification. The proposed method is called Optimal Feature Selection Classification (OFSC). Our model simultaneously learns an orthogonal subspace for jointly sparse feature selection and representation via l2,1-norms regularization. To solve the proposed model, an alternately iterative algorithm is proposed to optimize both the jointly sparse projection matrix and representation matrix. Experimental results on three public face datasets and one action dataset validate the quick convergence of our algorithm and show that the proposed method is more competitive than the state-of-the-art methods. Jiajun Wen 0001, Zhihui Lai 0001, Wai Keung Wong, Jinrong Cui, Minghua Wan |
ICPR | 3 |
| 2014 | Robust H∞ control for switched systems with input delays: A sojourn-probability-dependent method
Engang Tian, Wai Keung Wong, Dong Yue 0001 |
Inf. Sci. | 2 |
| 2014 | A seasonal discrete grey forecasting model for fashion retailing
Min Xia 0002, Wai Keung Wong |
Knowl. Based Syst. | 2 |
| 2014 | Sequence Memory Based on Coherent Spin-Interaction Neural NetworksabstractSequence information processing, for instance, the sequence memory, plays an important role on many functions of brain. In the workings of the human brain, the steady-state period is alterable. However, in the existing sequence memory models using heteroassociations, the steady-state period cannot be changed in the sequence recall. In this work, a novel neural network model for sequence memory with controllable steady-state period based on coherent spininteraction is proposed. In the proposed model, neurons fire collectively in a phase-coherent manner, which lets a neuron group respond differently to different patterns and also lets different neuron groups respond differently to one pattern. The simulation results demonstrating the performance of the sequence memory are presented. By introducing a new coherent spin-interaction sequence memory model, the steady-state period can be controlled by dimension parameters and the overlap between the input pattern and the stored patterns. The sequence storage capacity is enlarged by coherent spin interaction compared with the existing sequence memory models. Furthermore, the sequence storage capacity has an exponential relationship to the dimension of the neural network. Min Xia 0002, Wai Keung Wong, Zhijie Wang 0001 |
Neural Comput. | 2 |
| 2014 | 3D soft-tissue tracking using spatial-color joint probability distribution and thin-plate spline model
Bo Yang 0022, Wai Keung Wong, Chao Liu 0003, Philippe Poignet |
Pattern Recognit. | 2 |
| 2014 | Regularized discriminant entropy analysis
Wai Keung Wong |
Pattern Recognit. | 2 |
| 2014 | Sparse Alignment for Robust Tensor LearningabstractMultilinear/tensor extensions of manifold learning based algorithms have been widely used in computer vision and pattern recognition. This paper first provides a systematic analysis of the multilinear extensions for the most popular methods by using alignment techniques, thereby obtaining a general tensor alignment framework. From this framework, it is easy to show that the manifold learning based tensor learning methods are intrinsically different from the alignment techniques. Based on the alignment framework, a robust tensor learning method called sparse tensor alignment (STA) is then proposed for unsupervised tensor feature extraction. Different from the existing tensor learning methods, L1- and L2-norms are introduced to enhance the robustness in the alignment step of the STA. The advantage of the proposed technique is that the difficulty in selecting the size of the local neighborhood can be avoided in the manifold learning based tensor feature extraction algorithms. Although STA is an unsupervised learning method, the sparsity encodes the discriminative information in the alignment step and provides the robustness of STA. Extensive experiments on the well-known image databases as well as action and hand gesture databases by encoding object images as tensors demonstrate that the proposed STA algorithm gives the most competitive performance when compared with the tensor-based unsupervised learning methods. Zhihui Lai 0001, Wai Keung Wong, Yong Xu 0001, Cairong Zhao, Mingming Sun 0006 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2013 | A multivariate intelligent decision-making model for retail sales forecasting
Zhaoxia Guo, Wai Keung Wong |
Decis. Support Syst. | 2 |
| 2013 | Key role of voltage-dependent properties of synaptic currents in robust network synchronization
Wai Keung Wong |
Neural Networks | 2 |
| 2013 | Distributed Synchronization of Coupled Neural Networks via Randomly Occurring ControlabstractIn this paper, we study the distributed synchronization and pinning distributed synchronization of stochastic coupled neural networks via randomly occurring control. Two Bernoulli stochastic variables are used to describe the occurrences of distributed adaptive control and updating law according to certain probabilities. Both distributed adaptive control and updating law for each vertex in a network depend on state information on each vertex's neighborhood. By constructing appropriate Lyapunov functions and employing stochastic analysis techniques, we prove that the distributed synchronization and the distributed pinning synchronization of stochastic complex networks can be achieved in mean square. Additionally, randomly occurring distributed control is compared with periodically intermittent control. It is revealed that, although randomly occurring control is an intermediate method among the three types of control in terms of control costs and convergence rates, it has fewer restrictions to implement and can be more easily applied in practice than periodically intermittent control. Yang Tang 0001, Wai Keung Wong |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2012 | Intelligent multivariate sales forecasting using wrapper approach and neural networksabstractThis research investigated a retail sales forecasting problem based on early sales. An effective multivariate intelligent decision-making (MID) model is developed to handle this problem by integrating a data preparation and preprocessing module, a harmony search-wrapper-based variable selection (HWVS) module and a multivariate intelligent forecaster (MIF) module. The HWVS module selects out the optimal input variable subset from given candidate inputs as the inputs of MIF. The MIF is proposed to model the relationship between the selected input variables and the sales volumes of retail products, and then employed to forecast the sales volumes of retail products. Experiments were conducted to evaluate the effectiveness of the proposed model. Results show that it is statistically significant that the proposed MID model can provide superior forecasts to ELM-based model and generalized linear model. Zhaoxia Guo, Wai Keung Wong |
INDIN | 3 |
| 2012 | A hybrid particle swarm optimization and its application in neural networks
Sunney Yung-Sun Leung, Yang Tang 0001, Wai Keung Wong |
Expert Syst. Appl. | 3 |
| 2012 | Feedback controlled particle swarm optimization and its application in time-series prediction
Wai Keung Wong, Sunney Yung-Sun Leung, Zhaoxia Guo |
Expert Syst. Appl. | 1 |
| 2012 | Relationship between Applicability of Current-Based Synapses and Uniformity of Firing PatternsabstractThe purpose of this paper is to identify situations in neural network modeling where current-based synapses are applicable. The applicability of current-based synapse model for studying post-transient behavior of neural networks is discussed in terms of average synaptic current strength induced by per spike during one firing cycle of a neuron (or briefly per spike synaptic current strength). It was found that current-based synapse models are applicable in both situations where both the interspike intervals of the neurons and the distribution of firing times of the neurons are uniform, and where the firing of all neurons is synchronized. If neither the interspike intervals nor the distribution of firing times of the neurons is uniform or the reversal potential is between the rest and threshold potentials, current-based synapse models may be oversimplified. Wai Keung Wong, B. Zhen, Sunney Yung-Sun Leung |
Int. J. Neural Syst. | 1 |
| 2012 | Sparsely connected neural network-based time series forecasting
Zhaoxia Guo, Wai Keung Wong |
Inf. Sci. | 2 |
| 2012 | Discover latent discriminant information for dimensionality reduction: Non-negative Sparseness Preserving Embedding
Wai Keung Wong |
Pattern Recognit. | 1 |
| 2012 | Supervised optimal locality preserving projection
Wai Keung Wong |
Pattern Recognit. | 1 |
| 2012 | Sparse Approximation to the Eigensubspace for DiscriminationabstractTwo-dimensional (2-D) image-matrix-based projection methods for feature extraction are widely used in many fields of computer vision and pattern recognition. In this paper, we propose a novel framework called sparse 2-D projections (S2DP) for image feature extraction. Different from the existing 2-D feature extraction methods, S2DP iteratively learns the sparse projection matrix by using elastic net regression and singular value decomposition. Theoretical analysis shows that the optimal sparse subspace approximates the eigensubspace obtained by solving the corresponding generalized eigenequation. With the S2DP framework, many 2-D projection methods can be easily extended to sparse cases. Moreover, when each row/column of the image matrix is regarded as an independent high-dimensional vector (1-D vector), it is proven that the vector-based eigensubspace is also approximated by the sparse subspace obtained by the same method used in this paper. Theoretical analysis shows that, when compared with the vector-based sparse projection learning methods, S2DP greatly saves both computation and memory costs. This property makes S2DP more tractable for real-world applications. Experiments on well-known face databases indicate the competitive performance of the proposed S2DP over some 2-D projection methods when facial expressions, lighting conditions, and time vary. Zhihui Lai 0001, Wai Keung Wong, Zhong Jin, Jian Yang 0003, Yong Xu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2011 | A heuristic time-invariant model for fuzzy time series forecasting
Enjian Bai, Wai Keung Wong, W. C. Chu, Min Xia 0002 |
Expert Syst. Appl. | 2 |
| 2011 | Fuzzy logic control of a novel robotic hanger for garment inspection: Modeling, simulation and experimental implementation
E. H. K. Fung, X. Z. Zhang, C. W. M. Yuen, Wai Keung Wong |
Expert Syst. Appl. | 6 |
| 2011 | Deep Learning Regularized Fisher MappingsabstractFor classification tasks, it is always desirable to extract features that are most effective for preserving class separability. In this brief, we propose a new feature extraction method called regularized deep Fisher mapping (RDFM), which learns an explicit mapping from the sample space to the feature space using a deep neural network to enhance the separability of features according to the Fisher criterion. Compared to kernel methods, the deep neural network is a deep and nonlocal learning architecture, and therefore exhibits more powerful ability to learn the nature of highly variable datasets from fewer samples. To eliminate the side effects of overfitting brought about by the large capacity of powerful learners, regularizers are applied in the learning procedure of RDFM. RDFM is evaluated in various types of datasets, and the results reveal that it is necessary to apply unsupervised regularization in the fine-tuning phase of deep learning. Thus, for very flexible models, the optimal Fisher feature extractor may be a balance between discriminative ability and descriptive ability. Wai Keung Wong |
IEEE Trans. Neural Networks | 1 |
| 2010 | Tangent space discriminant analysis for feature extractionabstractIn this paper, a novel method called tangent space discriminant analysis is proposed for dimensionality reduction and feature extraction. Differing from the recently proposed manifold learning methods completely operating on raw feature space, TSDA completely uses the local tangent space to represent the local within-class geometry and local between-class geometry. Assume that the face images of different people reside on different intrinsically low-dimensional sub-manifolds, TSDA is developed to preserve the locality of each sub-manifold and simultaneously maximize the local separability of different sub-manifolds by using local tangent space alignment. Experimental results show that TSDA achieves higher recognition rates than a few the state-of-the-art techniques. Zhihui Lai 0001, Zhong Jin, Wai Keung Wong |
ICIP | 3 |
| 2010 | Sparse Local Discriminant Projections for Feature ExtractionabstractOne of the major disadvantages of the linear dimensionality reduction algorithms, such as Principle Component Analysis (PCA) and Linear Discriminant Analysis (LDA), are that the projections are linear combination of all the original features or variables and all weights in the linear combination known as loadings are typically non-zero. Thus, they lack physical interpretation in many applications. In this paper, we propose a novel supervised learning method called Sparse Local Discriminant Projections (SLDP) for linear dimensionality reduction. SLDP introduces a sparse constraint into the objective function and obtains a set of sparse projective axes with directly physical interpretation. The sparse projections can be efficiently computed by the Elastic Net combining with spectral analysis. The experimental results show that SLDP give the explicit interpretation on its projections and achieves competitive performance compared with some dimensionality reduction techniques. Zhihui Lai 0001, Zhong Jin, Jian Yang 0003, Wai Keung Wong |
ICPR | 4 |
| 2010 | Impulsive pinning synchronization of stochastic discrete-time networks
Yang Tang 0001, Sunney Yung-Sun Leung, Wai Keung Wong |
Neurocomputing | 3 |
| 2010 | Adaptive Time-Variant Models for Fuzzy-Time-Series ForecastingabstractA fuzzy time series has been applied to the prediction of enrollment, temperature, stock indices, and other domains. Related studies mainly focus on three factors, namely, the partition of discourse, the content of forecasting rules, and the methods of defuzzification, all of which greatly influence the prediction accuracy of forecasting models. These studies use fixed analysis window sizes for forecasting. In this paper, an adaptive time-variant fuzzy-time-series forecasting model (ATVF) is proposed to improve forecasting accuracy. The proposed model automatically adapts the analysis window size of fuzzy time series based on the prediction accuracy in the training phase and uses heuristic rules to generate forecasting values in the testing phase. The performance of the ATVF model is tested using both simulated and actual time series including the enrollments at the University of Alabama, Tuscaloosa, and the Taiwan Stock Exchange Capitalization Weighted Stock Index (TAIEX). The experiment results show that the proposed ATVF model achieves a significant improvement in forecasting accuracy as compared to other fuzzy-time-series forecasting models. Wai Keung Wong, Enjian Bai, Alice Wai-Ching Chu |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2009 | A Suboptimal Modified Code Tracking Loop for Synchronous DS-CDMA SystemsabstractThis paper presents a suboptimal modified code tracking loop (SO-MCTL) that mitigates the effect of multiple- access interference (MAI) for synchronous direct-sequence code division multiple access (DS-CDMA) systems. Weighted linear combinations of the interfering users' signals are introduced into the local signals to facilitate the optimization of the tracking performance. Instead of search for the optimal weighting vectors that yields the minimum mean square tracking error, the mean square tracking error is minimized with respect to the weighting parameters given the constraints that ensure a stable bias-free S-curve. Simulations based on Gold codes compare the S-curve and root-mean-square (rms) tracking error of the SO-MCTL and traditional modified code tracking loop (T-MCTL). It is shown that SO-MCTL outperforms T-MCTL greatly in that it provides unbiased code tracking with reduced mean square tracking error and near-far resistance. Yun T. Wu, Shu Hung Leung, Wai Keung Wong, Y. S. Zhu |
VTC Spring | 3 |
| 2009 | Intelligent production control decision support system for flexible assembly lines
Zhaoxia Guo, Wai Keung Wong, Sunney Yung-Sun Leung, J. T. Fan |
Expert Syst. Appl. | 2 |
| 2009 | Solving the two-dimensional irregular objects allocation problems by using a two-stage packing approach
Wai Keung Wong, X. X. Wang, P. Y. Mok 0001, Sunney Yung-Sun Leung, C. K. Kwong 0001 |
Expert Syst. Appl. | 1 |
| 2009 | Stitching defect detection and classification using wavelet transform and BP neural network
Wai Keung Wong, C. W. M. Yuen, D. D. Fan, L. K. Chan, E. H. K. Fung |
Expert Syst. Appl. | 1 |
| 2009 | A decision support tool for apparel coordination through integrating the knowledge-based attribute evaluation expert system and the T-S fuzzy neural network
Wai Keung Wong, X. H. Zeng, W. M. R. Au |
Expert Syst. Appl. | 1 |
| 2009 | A fashion mix-and-match expert system for fashion retailers using fuzzy screening approach
Wai Keung Wong, X. H. Zeng, W. M. R. Au, P. Y. Mok 0001, Sunney Yung-Sun Leung |
Expert Syst. Appl. | 1 |
| 2009 | A hybrid model using genetic algorithm and neural network for classifying garment defects
C. W. M. Yuen, Wai Keung Wong, S. Q. Qian, L. K. Chan, E. H. K. Fung |
Expert Syst. Appl. | 2 |
| 2008 | Genetic optimization of order scheduling with multiple uncertainties
Zhaoxia Guo, Wai Keung Wong, Sunney Yung-Sun Leung, J. T. Fan, S. F. Chan |
Expert Syst. Appl. | 2 |
| 2008 | A Genetic-Algorithm-Based Optimization Model for Solving the Flexible Assembly Line Balancing Problem With Work Sharing and Workstation RevisitingabstractThis paper investigates a flexible assembly line balancing (FALB) problem with work sharing and workstation revisiting. The mathematical model of the problem is presented, and its objective is to meet the desired cycle time of each order and minimize the total idle time of the assembly line. An optimization model is developed to tackle the addressed problem, which involves two parts. A bilevel genetic algorithm with multiparent crossover is proposed to determine the operation assignment to workstations and the task proportion of each shared operation being processed on different workstations. A heuristic operation routing rule is then presented to route the shared operation of each product to an appropriate workstation when it should be processed. Experiments based on industrial data are conducted to validate the proposed optimization model. The experimental results demonstrate the effectiveness of the proposed model to solve the FALB problem. Zhaoxia Guo, Wai Keung Wong, Sunney Yung-Sun Leung, J. T. Fan, S. F. Chan |
IEEE Trans. Syst. Man Cybern. Part C | 2 |
| 2007 | A Modified De-Correlated Delay Lock Loop with Better Static Response for Synchronous DS-CDMA SystemsabstractIn this paper we present a modified de-correlated delay lock loop (MD-DLL) for synchronous direct-sequence code division multiple access (DS-CDMA) systems. With appropriate choice of weighting and offset parameters of the local signals, the proposed MD-DLL scheme solves the tracking bias problems in de-correlated delay lock loop (D-DLL) under low SNR condition. By diminishing multiple access interference (MAI) at the on-time code position instant and minimizing the weighting imbalance between the two local signals, it is shown that MD-DLL achieves a better S-curve with negligible tracking bias and good near far resistance compared with the traditional delay lock loop (T-DLL) and D-DLL. Computer simulations based on Gold code sequences are used to verify the findings. Wai Keung Wong, Shu Hung Leung, Y. S. Zhu |
VTC Spring | 2 |
| 2000 | FPGA Implementation of a Microcoded Elliptic Curve Cryptographic ProcessorabstractElliptic curve cryptography (ECC) has been the focus of much recent attention since it offers the highest security per bit of any known public key cryptosystem. This benefit of smaller key sizes makes ECC particularly attractive for embedded applications since its implementation requires less memory and processing power. In this paper a microcoded Xilinx Virtex based elliptic curve processor is described. In contrast to previous implementations, it implements curve operations as well as optimal normal basis field operations in F(2/sup n/); the design is parameterized for arbitrary n; and it is microcoded to allow for rapid development of the control part of the processor. The design was successfully tested on a Xilinx Virtex XCV300-4 and, for n=113 bits, utilized 1290 slices at a maximum frequency of 45 MHz and achieved a thirty-fold speedup over an optimized software implementation. Ka Hei Leung, K. W. Ma, Wai Keung Wong, Philip H. W. Leong |
FCCM | 3 |