Shaozi Li

dblp:51/2064 · also Shao-Zi Li · DBLP profile ↗
← Back
199ranked-venue papers
10as first author
101since 2021 · last 2026
0000-0001-5403-9945ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 82 · 3 first-author · 44 since 2021Graphics, computer vision, multimedia, augmented reality and games · 82 · 1 first-author · 39 since 2021Systems, architecture and hardware · 30 · 22 since 2021Human-computer interaction and ubiquitous computing · 19 · 6 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 1 first-author · 10 since 2021Security and privacy · 5 · 5 since 2021Computer networks · 3 · 2 since 2021Databases, data management, data science and information retrieval · 2
YearPublicationVenuePosition
2026 OneFont: A Unified Agent for End-to-End Font Creation
abstract
Despite recent advancements in font generation, practitioners still grapple with a laborious trial-and-error workflow. To streamline this, we propose OneFont, an end-to-end framework that interprets user intents via free-form dialogue, seamlessly integrating both glyph synthesis and refinement modules. We introduce the Font with Thought (FwT) paradigm, reframing font design as a reasoning task where the model plans actions and articulates design rationales. OneFont’s core planner is trained via a two-stage regimen to master this paradigm. First, we instill reasoning abilities via Supervised Fine-Tuning (SFT) on a new, comprehensive benchmark of 1,500 font families we built. Second, we refine the model's policy with a novel reinforcement learning algorithm, Group Relative Policy Optimization (GRPO), guided by a hybrid reward that assesses visual fidelity, rationale coherence, and transformation correctness. Extensive experiments show OneFont significantly surpasses existing methods in design quality and stroke precision across diverse scripts, validated on our new benchmark. We will release our dataset, code, and models.
Yingxin Lai, Yufei Liu 0003, Jiaxing Chai, Zhiming Luo, Shaozi Li
AAAI6
2026 SCF-Net: Spatial-Channel Fusion and Feature Refinement for Vessel Re-Identification
abstract
ABSTRACT Vessel re‐identification (ReID) plays a critical role in maritime surveillance by matching vessels across different camera views. Compared with person or vehicle ReID, vessel ReID faces unique challenges due to subtle interclass differences and large intraclass variations caused by viewpoint changes. These issues are further exacerbated by the highly similar appearances of vessels and the lack of fine‐grained identity cues commonly found in other ReID tasks. To address these challenges, we propose a spatial‐channel fusion network (SCF‐Net), a dual‐branch deep framework that integrates a spatial‐channel fusion (SCF) module and a feature refinement and alignment (FRA) module. The SCF module captures interdependent relationships between spatial and channel dimensions, enabling the network to emphasize discriminative regions while suppressing irrelevant background information. The FRA module refines high‐dimensional embeddings into a compact representation and enforces intraclass similarity via a learnable multilayer perceptron (MLP) and a supervised mean squared error (MSE) loss. By jointly optimizing the two branches and the FRA output, SCF‐Net effectively learns both interclass discrimination and intraclass compactness. Extensive experiments demonstrate that SCF‐Net achieves competitive performance on public vessel ReID benchmarks, highlighting its effectiveness in handling subtle interclass differences and large intraclass variations.
Gangzhu Lin, Yongguo Ling, Wenhao Shao, Shaozi Li, Hongfeng Xu
Concurr. Comput. Pract. Exp.5
2026 Disentangled self-supervised video camouflaged object detection and salient object detection
Haoke Xiao, Lv Tang, Bo Li 0115, Zhiming Luo, Shaozi Li
Neural Networks5
2026 Mitigating low-frequency bias: Feature recalibration and frequency attention regularization for adversarial robustness
Kejia Zhang 0003, Juanjuan Weng, Yuanzheng Cai, Shaozi Li, Zhiming Luo
Neural Networks4
2026 MTSCL-Net: Multi-level temporal spatial contrastive learning for robust breast tumor segmentation in DCE-MRI
Jiezhou He, Zhiming Luo, Songzhi Su, Shaozi Li
Pattern Recognit.7
2026 FADMB: Fully attention-based dual memory bank network for weakly supervised video anomaly detection
Zhiming Luo, Shuheng Huang, Jianzhe Gao, Shaozi Li
Pattern Recognit.5
2026 Exploring Frequencies via Feature Mixing and Meta-Learning for Improving Adversarial Transferability
abstract
Recent studies have shown that Deep Neural Networks (DNNs) are susceptible to adversarial attacks, with frequency-domain analysis underscoring the significance of high-frequency components in influencing model predictions. Conversely, targeting low-frequency components has been effective in enhancing attack transferability on black-box models. In this study, we introduce a frequency decomposition-based feature mixing method to exploit these frequency characteristics in both clean and adversarial samples. Our findings suggest that incorporating features of clean samples into adversarial features extracted from adversarial examples is more effective in attacking normally-trained models, while combining clean features with the adversarial features extracted from low-frequency parts decomposed from the adversarial samples yields better results in attacking defense models. However, a conflict issue arises when these two mixing approaches are employed simultaneously. To tackle the issue, we propose a cross-frequency meta-optimization approach comprising the meta-train step, meta-test step, and final update. In the meta-train step, we leverage the low-frequency components of adversarial samples to boost the transferability of attacks against defense models. Meanwhile, in the meta-test step, we utilize adversarial samples to stabilize gradients, thereby enhancing the attack's transferability against normally trained models. For the final update, we update the adversarial sample based on the gradients obtained from both meta-train and meta-test steps. Our proposed method is evaluated through extensive experiments on the ImageNet-Compatible dataset, affirming its effectiveness in improving the transferability of attacks on both normally-trained CNNs and defense models. The source code is available at https://github.com/WJJLL/MetaSSA.
Juanjuan Weng, Zhiming Luo, Shaozi Li
IEEE Trans. Image Process.3
2025 Long-Tailed Out-of-Distribution Detection: Prioritizing Attention to Tail
abstract
Current out-of-distribution (OOD) detection methods typically assume balanced in-distribution (ID) data, while most real-world data follow a long-tailed distribution. Previous approaches to long-tailed OOD detection often involve balancing the ID data by reducing the semantics of head classes. However, this reduction can severely affect the classification accuracy of ID data. The main challenge of this task lies in the severe lack of features for tail classes, leading to confusion with OOD data. To tackle this issue, we introduce a novel Prioritizing Attention to Tail (PATT) method using augmentation instead of reduction. Our main intuition involves using a mixture of von Mises-Fisher (vMF) distributions to model the ID data and a temperature scaling module to boost the confidence of ID data. This enables us to generate infinite contrastive pairs, implicitly enhancing the semantics of ID classes while promoting differentiation between ID and OOD data. To further strengthen the detection of OOD data without compromising the classification performance of ID data, we propose feature calibration during the inference phase. By extracting an attention weight from the training set that prioritizes the tail classes and reduces the confidence in OOD data, we improve the OOD detection capability. Extensive experiments verified that our method outperforms the current state-of-the-art methods on various benchmarks.
Yina He, Yongcun Zhang, Juanjuan Weng, Shaozi Li, Zhiming Luo
AAAI5
2025 Font-Agent: Enhancing Font Understanding with Large Language Models
abstract
The rapid development of generative models has significantly advanced font generation. However, limited exploration has been devoted to the evaluation and interpretability of graphical fonts. Existing quality assessment models can only provide basic visual analyses, such as recognizing clarity and brightness, without in-depth explanations. To address these limitations, we first constructed a large-scale multimodal dataset named the Diversity Font Dataset (DFD), comprising 135,000 font-text pairs. This dataset encompasses a wide range of generated font types and annotations, including language descriptions and quality assessments, thus providing a robust foundation for training and evaluating font analysis models. Based on this dataset, we developed a font agent built upon a Vision-Language Model (VLM) aiming to enhance font quality assessment and offer interpretable question-answering capabilities. Alongside the original visual encoder in VLM, we integrated an Edge-Aware Traces (EAT) module to capture detailed edge information of font strokes and components. Furthermore, we introduced a Dynamic Direct Preference Optimization (D-DPO) strategy to facilitate efficient model fine-tuning. Experimental results demonstrate that Font-Agent achieves state-of-the-art performance on the established dataset. To further evaluate the generalization ability of our algorithm, we conducted additional experiments on several public datasets. The results highlight the notable advantage of Font-Agent in both assessing the quality of generated fonts and comprehending their content.
Yingxin Lai, Cuijie Xu, Haitian Shi, Zhiming Luo, Shaozi Li
CVPR7
2025 Towards Adversarial Robustness via Debiased High-Confidence Logit Alignment
abstract
Despite the remarkable progress of deep neural networks (DNNs) in various visual tasks, their vulnerability to adversarial examples raises significant security concerns. Recent adversarial training methods leverage inverse adversarial attacks to generate high-confidence examples, aiming to align adversarial distributions with high-confidence class regions. However, our investigation reveals that under inverse adversarial attacks, high-confidence outputs are influenced by biased feature activations, causing models to rely on background features that lack a causal relationship with the labels. This spurious correlation bias leads to overfitting irrelevant background features during adversarial training, thereby degrading the model's robust performance and generalization capabilities. To address this issue, we propose Debiased High-Confidence Adversarial Training (DHAT), a novel approach that aligns adversarial logits with debiased high-confidence logits and restores proper attention by enhancing foreground logit orthogonality. Extensive experiments demonstrate that DHAT achieves state-of-the-art robustness on both CIFAR and ImageNet-1K benchmarks, while significantly improving generalization by mitigating the feature bias inherent in inverse adversarial training approaches. Code is available at https://github.com/KejiaZhang-Robust/DHAT.
Kejia Zhang 0003, Juanjuan Weng, Shaozi Li, Zhiming Luo
ICCV3
2025 HCCM: Hierarchical Cross-Granularity Contrastive and Matching Learning for Natural Language-Guided Drones
Hao Ruan, Jinliang Lin, Yingxin Lai, Zhiming Luo, Shaozi Li
ACM Multimedia5
2025 PETformer: Prototype-Enhanced Transformer with Multi-Scale Convolution Attention for Multivariate Time Series Anomaly Detection
abstract
Unsupervised anomaly detection in multivariate time series is of significant practical importance in industrial monitoring and IoT device management. However, existing methods still face significant challenges in modeling complex temporal patterns. Although Transformer models have demonstrated notable potential in point-wise representation learning, their attention mechanisms primarily focus on direct dependencies between time steps, making it difficult to capture structural features at the segment level. To address these issues, we propose PETformer, a prototype-enhanced multi-scale unsupervised Transformer architecture. PETformer introduces the Multi-Head Scale Convolutional Attention (MHSCA) module, which employs parallel convolutional kernels to extract multi-scale features. This design enhances the model’s ability to capture both local and global dependencies. Additionally, it incorporates Multi-Scale Cosine Prototypes (MSCP) as inductive biases that interact with the input sequence, strengthening prior modeling of normal patterns across sequences and improving the model’s sensitivity to potential pattern deviations. Experimental results show that PETformer consistently outperforms existing state-of-the-art unsupervised anomaly detection methods on three widely used benchmark datasets.
Qianchen Ren, Yuliang Tang, Shaozi Li
SMC4
2025 A Two-Stage GNN for Joint UAV Positioning and Relay Routing
abstract
In modern warfare, bionic robots are increasingly deployed to execute high-risk missions. To ensure robust remote control and command in complex urban environments, unmanned aerial vehicles (UAVs) are utilized as aerial relays to facilitate reliable data transmission. This paper investigates the joint optimization of UAV positioning and multi-hop relay path selection in UAV-assisted wireless networks. We formulate the problem as a graph-based optimization task and propose a two-stage Graph Neural Network (GNN) framework. In the first stage, a reinforcement learning (RL) enhanced Relay Path GNN (RPG) is developed to enable low-latency and efficient routing. In the second stage, a UAV Position GNN (UPG) determines near-optimal UAV deployment strategies. Both modules are designed to train without labeled data, relying instead on unsupervised and RL techniques tailored to the graph-structured problem, making the approach data-efficient and robust. Simulation results show that the proposed framework UPG-RPG achieves near-optimal performance with substantially reduced computational complexity and demonstrates superior scalability and adaptability compared to conventional heuristic or rule-based methods.
Qianchen Ren, Yuliang Tang, Shaozi Li
SMC4
2025 HRCUNet: Hierarchical Region Contrastive Learning for Segmentation of Breast Tumors in DCE-MRI
abstract
ABSTRACT Segmenting breast tumors from dynamic contrast‐enhanced magnetic resonance images is a critical step in the early detection and diagnosis of breast cancer. However, this task becomes significantly more challenging due to the diverse shapes and sizes of tumors, which make it difficult to establish a unified perception field for modeling them. Moreover, tumor regions are often subtle or imperceptible during early detection, exacerbating the issue of extreme class imbalance. This imbalance can lead to biased training and challenge accurately segmenting tumor regions from the predominant normal tissues. To address these issues, we propose a hierarchical region contrastive learning approach for breast tumor segmentation. Our approach introduces a novel hierarchical region contrastive learning loss function that addresses the class imbalance problem. This loss function encourages the model to create a clear separation between feature embeddings by maximizing the inter‐class margin and minimizing the intra‐class distance across different levels of the feature space. In addition, we design a novel Attention‐based 3D Multi‐scale Feature Fusion Residual Module to explore more granular multi‐scale representations to improve the feature learning ability of tumors. Extensive experiments on two breast DCE‐MRI datasets demonstrate that the proposed algorithm is more competitive against several state‐of‐the‐art approaches under different segmentation metrics.
Jiezhou He, Zhiming Luo, Wei Peng 0009, Songzhi Su, Shaozi Li
Concurr. Comput. Pract. Exp.7
2025 An Efficient Lattice-Based Heterogeneous Signcryption Scheme for VANETs
abstract
ABSTRACT Nowadays, vehicular ad‐hoc networks (VANETs) offer increased convenience to drivers and enable intelligent traffic management. However, the public wireless transmission channel in VANETs brings challenges related to security vulnerabilities and privacy leakage, in addition, vehicles produced by different manufacturers may use different cryptosystems such as certificateless cryptosystems (CLCs) and identity‐based cryptosystems (IBC). To address privacy leakage during cross‐cryptosystem communication in VANETs, we propose a lattice‐based heterogeneous signcryption scheme named LHS‐C2I. The scheme facilitates secure multi‐cryptosystem bidirectional communication as CLC‐based vehicles to IBC‐based vehicles and IBC‐based vehicles to CLC‐based vehicles. The confidentiality and authenticity of LHS‐C2I help to prevent the users from privacy leakage during cross‐cryptosystem communication and to authenticate the message integrity and the sender's identity legitimacy. The proposed scheme is proven to achieve Indistinguishability under Chosen Ciphertext Attack (IND‐CCA2) and Existential Unforgeability against Adaptive Chosen Messages Attack (EUF‐CMA) within the random oracle model. Performance analysis demonstrates that LHS‐C2I outperforms existing schemes in terms of computational overhead, communication overhead, and overall security features. It is particularly well‐suited for scenarios requiring secure communication across different cryptosystems in VANETs.
Jintao Jiao, Lei Guo 0020, Wensen Yu, Shaozi Li
Concurr. Comput. Pract. Exp.5
2025 Advancing Continuous Sign Language Recognition Through Denoising Diffusion Transformer-Based Spatial-Temporal Enhancement
abstract
ABSTRACT The intricate spatial‐temporal dynamics and variability of sign language gestures pose significant challenges for Continuous Sign Language Recognition (CSLR) systems. Existing models often fall short in accurately capturing these complexities, leading to performance issues and frequent misalignments. To address these shortcomings, we introduce a new approach that leverages Denoising Diffusion Models (DDMs) to improve feature representation in the visual‐sequential module of CSLR systems. Originally intended for generative tasks, DDMs have shown strong potential in representation learning through a denoising process akin to Denoising Autoencoders. Our method incorporates a denoising diffusion transformer into the CSLR framework to refine spatial‐temporal features, capitalizing on the ability of diffusion models to enhance representation quality. By conditionally denoising visual feature sequences, our approach increases the discriminative capability of the system. Additionally, we introduce an additional classifier, trained with Connectionist Temporal Classification (CTC) loss, to provide complementary supervision and further boost performance. Extensive experiments demonstrate that our method significantly improves CSLR accuracy by effectively capturing the subtle details of continuous sign language gestures and overcoming the representation limitations of current models.
Suhail Muhammad Kamal, Yidong Chen 0001, Shaozi Li
Concurr. Comput. Pract. Exp.3
2025 Enabling Interactive Education With Low-Latency Large Language Models
abstract
ABSTRACT Large language models (LLMs) have transformed educational applications through personalized learning and intelligent tutoring systems. However, educational LLMs (EduLLMs) face significant deployment challenges due to their massive computational demands and autoregressive nature, particularly in resource‐constrained environments. This paper presents a novel approach for inference acceleration aimed at facilitating the deployment of EduLLMs on a commodity GPU. This technique substantially diminishes the memory footprint and the volume of data transfers between the CPU and GPU by strategically preloading critical neurons directly onto the GPU, thereby enabling rapid access. Concurrently, computations pertaining to non‐critical neurons are processed on the CPU. Implementation of our optimized approach, AccEduLLM, on a single NVIDIA RTX 3090 GPU using FP16 type, achieves 60.09 tokens/s, which is 8.32 faster than the previous work, while preserving the model's accuracy. The approach demonstrates particular effectiveness for variable‐length educational content like essays and textbooks.
Wei Wu 0072, Lei Guo 0020, Wensen Yu, Tzong-Jer Chen, Xing Ruan, Shaozi Li
Concurr. Comput. Pract. Exp.6
2025 Niching-Based Two-Stage Differential Evolution for Feature Selection
abstract
ABSTRACT Feature selection is a vital preprocessing step aimed at identifying a subset of the most relevant features from high‐dimensional data to enhance model performance and reduce computational complexity. Differential Evolution (DE) algorithms have been extensively applied to this task by iteratively optimizing the selection probabilities or weights of features. However, many existing DE‐based approaches suffer from premature convergence and local optima entrapment due to an insufficient balance between global exploration and local exploitation. To address these challenges, we propose a niching‐based two‐stage mutation DE algorithm for feature selection. Firstly, the mutual information is utilized to initialize the population and reduce the number of features. Then, an improved two‐stage mutation operator is employed to balance the algorithm's exploitation and exploration. Additionally, duplicate individuals generated during the evolutionary process are replaced using a one‐bit evolutionary mutation to aid population evolution. Experimental results on 16 benchmark datasets demonstrate that the proposed method achieves superior classification accuracy compared to several state‐of‐the‐art approaches, validating its effectiveness and robustness in diverse feature selection scenarios.
Deyong Wu, Zhiming Luo, Jiezhou He, Shaozi Li
Concurr. Comput. Pract. Exp.4
2025 Research on Transient Stability Model of Power System Based on Periodic Enhanced Informer
abstract
ABSTRACT In view of the nonlinearity and complexity of power system connection, it is difficult to directly identify the key characteristics of power system stability, while there are characteristics unrelated to system stability. For this reason, this paper proposes a method to distinguish the transient stability margin of power grid based on long short memory recurrent network. This method uses unbalanced data clustering algorithm to perform clustering analysis on the historical data set, extract a variety of typical instability characteristics, and consider the periodic characteristics of the unstable sequence. The unstable value in the same period of the input sequence is combined with the output of the long and short term memory network (LSTM) network to output the instability prediction results. At the same time, the key value of the periodic instability in the input sequence is connected in series with the output of the Informer model to build a full connection layer, and combined with the convolutional neural network to build a power system transient stability model based on periodically enhanced Informer. Experimental results show that the model effectively solves the problem of periodic mode attenuation of non‐stationary load series and alleviates the redundancy problem of key features of power grid stability by building a multi‐dimensional periodic feature extraction channel.
Zhiming Luo, Shaozi Li
Concurr. Comput. Pract. Exp.3
2025 Pick and mix reliable pseudo labels for scribble-supervised medical image segmentation
Jiawei Su, Zhiming Luo, Dazhen Lin, Lihui Lin, Shaozi Li
Neurocomputing5
2025 Text-Guided Multiround Learning for Unsupervised Visible-Infrared Person Re-Identification
abstract
Unsupervised Visible-Infrared Person Re-Identification (USVI-ReID) aims to match images of the same individual across visible and infrared modalities without relying on identity annotations. This task is critical in Visual Internet of Things (VIoT) systems. Existing methods typically generate pseudo-labels through clustering and cross-modality associations. However, the quality of these pseudo-labels is often compromised by unstable clustering performance, modality discrepancies, and unreliable matching strategies, resulting in suboptimal accuracy. Therefore, obtaining more reliable and robust pseudo-labels remains challenging in this domain. To address these challenges, we propose a Text-Guided Multi-Round Learning (TGMRL) framework that enhances the reliability and robustness of cross-modality pseudo-label associations. TGMRL comprises two core components: (1) the Text Dual-Cross Similarity Matching (TDSM) module, which facilitates cross-modality cluster alignment by constructing both visual and text-based cluster centers and integrating their similarity matrices for more accurate correspondence; and (2) the Multi-Round Pseudo-Label Guidance (MRPG) module, which enhances label consistency by imposing temporal regularization across clustering iterations through both similarity and Intersection-Over-Union (IoU) based measures. Extensive experiments on two benchmark datasets demonstrate that TGMRL significantly improves the reliability of pseudo-labels and achieves state-of-the-art performance on the USVI-ReID task.
Yongguo Ling, Zihao Hu, Wenhao Shao, Shaozi Li, Thomas Wu 0001
IEEE Internet Things J.5
2025 OTMA: Optimal transfer modality alignment for visible-thermal person re-identification
Yongguo Ling, Zihao Hu, Gangzhu Lin, Shaozi Li, Min Jiang 0005
Knowl. Based Syst.4
2025 Cross-modality average precision optimization for visible thermal person re-identification
Yongguo Ling, Zhiming Luo, Dazhen Lin, Shaozi Li, Min Jiang 0005, Nicu Sebe, Zhun Zhong
Pattern Recognit.4
2025 Improving Transferable Targeted Adversarial Attack via Normalized Logit Calibration and Truncated Feature Mixing
abstract
This paper aims to enhance the transferability of adversarial samples in targeted attacks, where attack success rates remain comparatively low. To achieve this objective, we propose two distinct techniques for improving the targeted transferability from the loss and feature aspects. First, in previous approaches, logit calibrations used in targeted attacks primarily focus on the logit margin between the targeted class and the untargeted classes among samples, neglecting the standard deviation of the logit. In contrast, we introduce a new normalized logit calibration method that jointly considers the logit margin and the standard deviation of logits. This approach effectively calibrates the logits, enhancing the targeted transferability. Second, previous studies have demonstrated that mixing the features of clean samples during optimization can significantly increase transferability. Building upon this, we further investigate a truncated feature mixing method to reduce the impact of the source training model, resulting in additional improvements. The truncated feature is determined by removing the Rank-1 feature associated with the largest singular value decomposed from the high-level convolutional layers of the clean sample. Extensive experiments conducted on the ImageNet-Compatible, CIFAR-10 and ImageNet-1k datasets demonstrate the individual and mutual benefits of our proposed two components, which outperform the state-of-the-art methods by a large margin in black-box targeted attacks.
Juanjuan Weng, Zhiming Luo, Shaozi Li
IEEE Trans. Inf. Forensics Secur.3
2025 A Self-Adaptive Feature Extraction Method for Aerial-View Geo-Localization
abstract
Cross-view geo-localization aims to match the same geographic location from different view images, e.g., drone-view images and geo-referenced satellite-view images. Due to UAV cameras' different shooting angles and heights, the scale of the same captured target building in the drone-view images varies greatly. Meanwhile, there is a difference in size and floor area for different geographic locations in the real world, such as towers and stadiums, which also leads to scale variants of geographic targets in the images. However, existing methods mainly focus on extracting the fine-grained information of the geographic targets or the contextual information of the surrounding area, which overlook the robust feature for scale changes and the importance of feature alignment. In this study, we argue that the key underpinning of this task is to train a network to mine a discriminative representation against scale variants. To this end, we design an effective and novel end-to-end network called Self-Adaptive Feature Extraction Network (Safe-Net) to extract powerful scale-invariant features in a self-adaptive manner. Safe-Net includes a global representation-guided feature alignment module and a saliency-guided feature partition module. The former applies an affine transformation guided by the global feature for adaptive feature alignment. Without extra region annotations, the latter computes saliency distribution for different regions of the image and adopts the saliency information to guide a self-adaptive feature partition on the feature map to learn a visual representation against scale variants. Experiments on two prevailing large-scale aerial-view geo-localization benchmarks, i.e., University-1652 and SUES-200, show that the proposed method achieves state-of-the-art results. In addition, our proposed Safe-Net has a significant scale adaptive capability and can extract robust feature representations for those query images with small target buildings. The source code of this study is available at: https://github.com/AggMan96/Safe-Net.
Jinliang Lin, Zhiming Luo, Dazhen Lin, Shaozi Li, Zhun Zhong
IEEE Trans. Image Process.4
2025 Dual-Modality-Shared Learning and Label Refinement for Unsupervised Visible-Infrared Person ReID
abstract
Unsupervised visible-infrared person re-identification (USVI-ReID) aims to match a person across two modalities without annotations. Current research primarily addresses the modality gap by establishing cross-modality correspondences through matching algorithms and utilizing memory banks for contrastive learning. However, the inherent noise in pseudo labels and neglect of hard samples often limit the efficacy of cross-modality learning. In this article, we propose a dual-modality-shared learning and label refinement (DLLR) algorithm for USVI-ReID. First, we leverage a cluster similarity matching (CSM) module and a cluster relationship-based label refinement (CRLR) algorithm to create and refine pseudo labels. Then, we adopt a weighted modality-shared memory (WMM) to construct memory banks by jointly considering sample distribution and feature differences, thereby enhancing the effectiveness of cross-modality learning. Extensive experiments on three publicly available datasets validate the effectiveness of our proposed method, which outperforms state-of-the-art methods. The code is available at https://github.com/CharRic/DLLR .
Licun Dai, Zhiming Luo, Yongguo Ling, Jiaxing Chai, Shaozi Li
ACM Trans. Multim. Comput. Commun. Appl.5
2024 Diversity-Authenticity Co-constrained Stylization for Federated Domain Generalization in Person Re-identification
abstract
This paper tackles the problem of federated domain generalization in person re-identification (FedDG re-ID), aiming to learn a model generalizable to unseen domains with decentralized source domains. Previous methods mainly focus on preventing local overfitting. However, the direction of diversifying local data through stylization for model training is largely overlooked. This direction is popular in domain generalization but will encounter two issues under federated scenario: (1) Most stylization methods require the centralization of multiple domains to generate novel styles but this is not applicable under decentralized constraint. (2) The authenticity of generated data cannot be ensured especially given limited local data, which may impair the model optimization. To solve these two problems, we propose the Diversity-Authenticity Co-constrained Stylization (DACS), which can generate diverse and authentic data for learning robust local model. Specifically, we deploy a style transformation model on each domain to generate novel data with two constraints: (1) A diversity constraint is designed to increase data diversity, which enlarges the Wasserstein distance between the original and transformed data; (2) An authenticity constraint is proposed to ensure data authenticity, which enforces the transformed data to be easily/hardly recognized by the local-side global/local model. Extensive experiments demonstrate the effectiveness of the proposed DACS and show that DACS achieves state-of-the-art performance for FedDG re-ID.
Fengxiang Yang, Zhun Zhong, Zhiming Luo, Yifan He 0002, Shaozi Li, Nicu Sebe
AAAI5
2024 Learning to Distinguish Samples for Generalized Category Discovery
Fengxiang Yang, Nan Pu, Wenjing Li 0005, Zhiming Luo, Shaozi Li, Nicu Sebe, Zhun Zhong
ECCV (65)5
2024 CC-DA: Cross-Domain Consistency Data Augmentation for 3D Tumor Segmentation
abstract
Deep learning-based tumor segmentation in 3D medical images faces the challenges of limited annotated data and class imbalance. In this paper, we proposed a novel Cross-domain Consistency Data Augmentation (CC-DA) for 3D tumor segmentation. Specifically, we copy the tumor from source data and apply random transformations to enhance its diversity. Then, we paste the enhanced tumor into the organ area of target data to generate a new sample. This process can alleviate class imbalance by regulating the merged tumor pixel ratio. To further enhance the generated data credibility, we proposed a domain consistency constraint that aligns the source data distribution with the target data distribution. We conduct extensive experiments on KiTS19 and LiTS17 datasets. The promising results clearly show that our CC-DA method can effectively improve the existing state-of-the-art 3D tumor segmentation performance.
Jiezhou He, Zhiming Luo, Wei Peng 0009, Songzhi Su, Shaozi Li
ICASSP5
2024 Selective Domain-Invariant Feature for Generalizable Deepfake Detection
abstract
With diverse presentation forgery methods emerging continually, detecting the authenticity of images has drawn growing attention. Although existing methods have achieved impressive accuracy in training dataset detection, they still perform poorly in the unseen domain and suffer from forgery of irrelevant information such as background and identity, affecting generalizability. To solve this problem, we proposed a novel framework Selective Domain-Invariant Feature (SDIF), which reduces the sensitivity to face forgery by fusing content features and styles. Specifically, we first use a Farthest-Point Sampling (FPS) training strategy to construct a task-relevant style sample representation space for fusing with content features. Then, we propose a dynamic feature extraction module to generate features with diverse styles to improve the performance and effectiveness of the feature extractor. Finally, a domain separation strategy is used to retain domain-related features to help distinguish between real and fake faces. Both qualitative and quantitative results in existing benchmarks and proposals demonstrate the effectiveness of our approach.
Yingxin Lai, Yifan He 0002, Zhiming Luo, Shaozi Li
ICASSP5
2024 Zero-Shot Co-Salient Object Detection Framework
abstract
Co-salient Object Detection (CoSOD) endeavors to replicate the human visual system’s capacity to recognize common and salient objects within a collection of images. Despite recent advancements in deep learning models, these models still rely on training with well-annotated CoSOD datasets. The exploration of training-free zero-shot CoSOD frameworks has been limited. In this paper, taking inspiration from the zero-shot transfer capabilities of foundational computer vision models, we introduce the first zero-shot CoSOD framework that harnesses these models without any training process. To achieve this, we introduce two novel components in our proposed framework: the group prompt generation (GPG) module and the co-saliency map generation (CMP) module. We evaluate the framework’s performance on widely-used datasets and observe impressive results. Our approach surpasses existing unsupervised methods and even outperforms fully supervised methods developed before 2020, while remaining competitive with some fully supervised methods developed before 2022.
Haoke Xiao, Lv Tang, Bo Li 0115, Zhiming Luo, Shaozi Li
ICASSP5
2024 Omni-Granularity Embedding Network for Text-to-Image Person Retrieval
abstract
Text-to-image person retrieval aims to identify the desired individual based on a textual description. As an instance-level retrieval problem, it has a large intra-class variance and a small inter-class variance. Although significant progress has been made, the omni-granularity matching issue remains unaddressed. Omni-granularity matching involves aligning words with multi-granularity image regions, challenging models to learn in an omni-granularity embedding space. In this paper, we introduce a novel Omni-Granularity Embedding Network (OGEN) for person representation learning. It addresses the omni-granularity matching issue by developing a Cross-Granularity Aggregation Module (CGAM). This module dynamically consolidates diverse granularity features for learning granularity-dependent and omni-granularity person representations. Additionally, a teacher-student knowledge transfer framework is introduced to minimize the inter-modality discrepancy, allowing CGAM to focus on modality-shared semantics. Due to the effectiveness of CGAM and the knowledge transfer framework, our OGEN enhances the Rank-1 accuracy of the Baseline by 8.54%, 9.89%, and 11.09% on three public datasets, respectively.
Chengji Wang, Zhiming Luo, Shaozi Li
ICME3
2024 Mask Matching Network for Self-supervised Few-shot Medical Image Segmentation
abstract
Existing few-shot segmentation methods have achieved remarkable progress in medical image segmentation. However, many existing methods yield incomplete and discontinuous boundary predictions. In contrast, the Segment Anything Model (SAM) consistently produces clear, continuous, and comprehensive segmentation boundaries. Building on this observation, we propose a new two-step network called Mask Matching Network (MMNet) to introduce extra knowledge learned by SAM in natural images for few-shot medical image segmentation. Firstly, Q-Net has been utilized to locate some Regions of Interest (RoI) as prompts for SAM, allowing for the automatic generation of masks without relying on manual prompts. Secondly, we propose a novel Mask Matching Module (MMM), which considers both feature similarity and volume similarity as guidance to collaboratively mine the final segmentation from proposal masks. MMNet achieves state-of-the-art performance with remarkable improvements on two widely used datasets, abdominal MR (ABD) and cardiac MR (CMR), under two different settings.
Zeyun Zhao, Jianzhe Gao, Zhiming Luo, Shaozi Li
ICME5
2024 TSESNet: Temporal-Spatial Enhanced Breast Tumor Segmentation in DCE-MRI Using Feature Perception and Separability
Jiezhou He, Zhiming Luo, Songzhi Su, Shaozi Li
IJCAI5
2024 A Collaborative Framework Using Multimodal Data and Adaptive Noise for Human Behavior Anomaly Detection
abstract
Human behavior anomaly detection in video aims to identify unusual behaviors that are crucial for public safety. Recently, there has been an increase in reconstruction or prediction-based methods that integrate diverse modal features to enhance anomaly detection. However, they use methods that independently or directly fusion multimodal features without fully considering the collaborative potential between multimodal features, which are susceptible to interference from semantic differences, thereby impacting detection performance. In contrast, we design a collaborative framework using multimodal data and adaptive noise for behavior anomaly detection. Our framework detects anomalies by analyzing the contrastive differences between two modalities alongside single-frame reconstruction errors. Specifically, we first learn the correlation between RGB and skeletal modalities for normal behavior through contrastive learning and use inter-modal contrast difference to detect motion anomalies. Additionally, we propose a single-frame reconstruction network that adaptively adds noise based on the importance of foreground features to detect appearance anomalies. Anomalies often occur in the motion foreground, and increasing noise in this area can make it more difficult to reconstruct anomalies. Extensive experiments validate the state-of-the-art performance of our method on three public datasets.
Jianzhe Gao, Kejia Zhang 0003, Yifan He 0002, Zhiming Luo, Shaozi Li
IJCNN6
2024 QueryNet: A Unified Framework for Accurate Polyp Segmentation and Detection
Jiaxing Chai, Zhiming Luo, Jianzhe Gao, Licun Dai, Yingxin Lai, Shaozi Li
MICCAI (8)6
2024 VCLIPSeg: Voxel-Wise CLIP-Enhanced Model for Semi-supervised Medical Image Segmentation
Lei Li 0048, Sheng Lian, Zhiming Luo, Beizhan Wang, Shaozi Li
MICCAI (9)5
2024 A Multilevel Guidance-Exploration Network and Behavior-Scene Matching Method for Human Behavior Anomaly Detection
abstract
Human behavior anomaly detection aims to identify unusual human actions, playing a crucial role in intelligent surveillance and other areas. The current mainstream methods still adopt reconstruction or future frame prediction techniques. However, reconstructing or predicting low-level pixel features easily enables the network to achieve overly strong generalization ability, allowing anomalies to be reconstructed or predicted as effectively as normal data. Different from their methods, inspired by the Student-Teacher Network, we propose a novel framework called the Multilevel Guidance-Exploration Network (MGENet), which detects anomalies through the difference in high-level representation between the Guidance and Exploration network. Specifically, we first utilize the Normalizing Flow that takes skeletal keypoints as input to guide an RGB encoder, which takes unmasked RGB frames as input, to explore latent motion features. Then, the RGB encoder guides the mask encoder, which takes masked RGB frames as input, to explore the latent appearance feature. Additionally, we design a Behavior-Scene Matching Module to detect scene-related behavioral anomalies. Extensive experiments demonstrate that our proposed method achieves state-of-the-art performance on ShanghaiTech and UBnormal datasets, with AUC of 86.9% and 74.3%, respectively. The code is available at https://github.com/molu-ggg/GENet.
Zhiming Luo, Jianzhe Gao, Yingxin Lai, Yifan He 0002, Shaozi Li
ACM Multimedia7
2024 CPNet: Cross Prototype Network for Few-Shot Medical Image Segmentation
Zeyun Zhao, Jianzhe Gao, Zhiming Luo, Shaozi Li
PRCV (15)4
2024 Comparative evaluation of recent universal adversarial perturbations in image classification
Juanjuan Weng, Zhiming Luo, Dazhen Lin, Shaozi Li
Comput. Secur.4
2024 Learning transferable targeted universal adversarial perturbations by sequential meta-learning
Juanjuan Weng, Zhiming Luo, Dazhen Lin, Shaozi Li
Comput. Secur.4
2024 Adversarial autoencoder for continuous sign language recognition
abstract
Summary Sign language serves as a vital communication medium for the deaf community, encompassing a diverse array of signs conveyed through distinct hand shapes along with non‐manual gestures like facial expressions and body movements. Accurate recognition of sign language is crucial for bridging the communication gap between deaf and hearing individuals, yet the scarcity of large‐scale datasets poses a significant challenge in developing robust recognition technologies. Existing works address this challenge by employing various strategies, such as enhancing visual modules, incorporating pretrained visual models, and leveraging multiple modalities to improve performance and mitigate overfitting. However, the exploration of the contextual module, responsible for modeling long‐term dependencies, remains limited. This work introduces an Adversarial Autoencoder for Continuous Sign Language Recognition, AA‐CSLR, to address the constraints imposed by limited data availability, leveraging the capabilities of generative models. The integration of pretrained knowledge, coupled with cross‐modal alignment, enhances the representation of sign language by effectively aligning visual and textual features. Through extensive experiments on publicly available datasets (PHOENIX‐2014, PHOENIX‐2014T, and CSL‐Daily), we demonstrate the effectiveness of our proposed method in achieving competitive performance in continuous sign language recognition.
Suhail Muhammad Kamal, Yidong Chen 0001, Shaozi Li
Concurr. Comput. Pract. Exp.3
2024 Learning multi-organ and tumor segmentation from partially labeled datasets by a conditional dynamic attention network
abstract
Summary Multi‐organ segmentation is a critical prerequisite for many clinical applications. Deep learning‐based approaches have recently achieved promising results on this task. However, they heavily rely on massive data with multi‐organ annotated, which is labor‐ and expert‐intensive and thus difficult to obtain. In contrast, single‐organ datasets are easier to acquire, and many well‐annotated ones are publicly available. It leads to the partially labeled issue: How to learn a unified multi‐organ segmentation model from several single‐organ datasets? Pseudo‐label‐based methods and conditional information‐based methods make up the majority of existing solutions, where the former largely depends on the accuracy of pseudo‐labels, and the latter has a limited capacity for task‐related features. In this paper, we propose the Conditional Dynamic Attention Network (CDANet). Our approach is designed with two key components: (1) multisource parameter generator, fusing the conditional and multiscale information to better distinguish among different tasks, and (2) dynamic attention module, promoting more attention to task‐related features. We have conducted extensive experiments on seven partially labeled challenging datasets. The results show that our method achieved competitive results compared with the advanced approaches, with an average Dice score of 75.08%. Additionally, the Hausdorff Distance is 26.31, which is a competitive result.
Lei Li 0048, Sheng Lian, Dazhen Lin, Zhiming Luo, Beizhan Wang, Shaozi Li
Concurr. Comput. Pract. Exp.6
2024 Mutual learning with reliable pseudo label for semi-supervised medical image segmentation
Jiawei Su, Zhiming Luo, Sheng Lian, Dazhen Lin, Shaozi Li
Medical Image Anal.5
2024 Reconstruct incomplete relation for incomplete modality brain tumor segmentation
Jiawei Su, Zhiming Luo, Chengji Wang, Sheng Lian, Xuejuan Lin, Shaozi Li
Neural Networks6
2024 Semi-supervised imbalanced multi-label classification with label propagation
Guodong Du 0002, Jia Zhang 0019, Hanrui Wu, Peiliang Wu, Shaozi Li
Pattern Recognit.6
2024 Bridge Gap in Pixel and Feature Level for Cross-Modality Person Re-Identification
abstract
Visible thermal person re-identification (VT-ReID) plays a vital role in intelligent surveillance systems, particularly in weak lighting environments. VT-ReID faces substantial challenges, including the cross-modality gap and intra-class variations. Existing methods address these challenges through either pixel-level image translation techniques or feature-level metric learning techniques. However, the former approaches require additional computational costs and often generate noisy images, making model training challenging. The latter methods focus on constraining the relations between individual instances or class centers, while often ignoring joint consideration of the relationship between the two aspects. In addition, these works do not fully investigate the mutual benefits at both pixel-level and feature-level. To address these limitations, we propose a unified Dual-level Smooth Gap (DSG) learning framework that simultaneously smooths the cross-modality gap at the pixel and feature levels. Specifically, on the one hand, we develop a parameter-free Class-aware Modality Mix (CMM) to smooth the cross-modality gap at the pixel level. CMM can capture and explore internal information between the two modalities by mixing images from different modalities belonging to the same class. On the other hand, we devise an efficient Center-guided Metric Learning (CML) to reduce the inter-modality discrepancy and intra-class variations at the feature level. CML enhances model discrimination and generalization by enforcing constraints on both class centers and instances. Experiments on two benchmark datasets demonstrate the mutual benefits of our proposed and show the superior performance of our method over state-of-the-art methods.
Yongguo Ling, Zhun Zhong, Zhiming Luo, Shaozi Li, Nicu Sebe
IEEE Trans. Circuits Syst. Video Technol.4
2024 Boosting Adversarial Transferability via Logits Mixup With Dominant Decomposed Feature
abstract
Recent research has shown that adversarial samples are highly transferable and can be used to attack other unknown black-box Deep Neural Networks (DNNs). To improve the transferability of adversarial samples, several feature-based adversarial attack methods have been proposed to disrupt neuron activation in the middle layers. However, current state-of-the-art feature-based attack methods typically require additional computation costs for estimating the importance of neurons. To address this challenge, we propose a Singular Value Decomposition (SVD)-based feature-level attack method. Our approach is inspired by the discovery that eigenvectors associated with the larger singular values decomposed from the middle layer features exhibit superior generalization and attention properties. Specifically, we conduct the attack by retaining the dominant decomposed feature that corresponds to the largest singular value (i.e., Rank-1 decomposed feature) for computing the output logits before the final softmax. These logits are later integrated with the original logits to optimize adversarial examples. Our extensive experimental results verify the effectiveness of our proposed method, which can be easily integrated into various baselines to significantly enhance the transferability of adversarial samples for disturbing normally trained CNNs and advanced defense strategies. The source code is available at Link.
Juanjuan Weng, Zhiming Luo, Shaozi Li, Dazhen Lin, Zhun Zhong
IEEE Trans. Inf. Forensics Secur.3
2024 Fast Multilabel Feature Selection via Global Relevance and Redundancy Optimization
abstract
Information theoretical-based methods have attracted a great attention in recent years and gained promising results for multilabel feature selection (MLFS). Nevertheless, most of the existing methods consider a heuristic way to the grid search of important features, and they may also suffer from the issue of fully utilizing labeling information. Thus, they are probable to deliver a suboptimal result with heavy computational burden. In this article, we propose a general optimization framework global relevance and redundancy optimization (GRRO) to solve the learning problem. The main technical contribution in GRRO is a formulation for MLFS while feature relevance, label relevance (i.e., label correlation), and feature redundancy are taken into account, which can avoid repetitive entropy calculations to obtain a global optimal solution efficiently. To further improve the efficiency, we extend GRRO to filter out inessential labels and features, thus facilitating fast MLFS. We call the extension as GRROfast, in which the key insights are twofold: 1) promising labels and related relevant features are investigated to reduce ineffective calculations in terms of features, even labels and 2) the framework of GRRO is reconstructed to generate the optimal result with an ensemble. Moreover, our proposed algorithms have an excellent mechanism for exploiting the inherent properties of multilabel data; specifically, we provide a formulation to enhance the proposal with label-specific features. Extensive experiments clearly reveal the effectiveness and efficiency of our proposed algorithms.
Jia Zhang 0019, Yidong Lin, Min Jiang 0005, Shaozi Li, Yong Tang 0001, Jinyi Long, Jian Weng 0001, Kay Chen Tan
IEEE Trans. Neural Networks Learn. Syst.4
2023 Cross-Modality Earth Mover's Distance for Visible Thermal Person Re-identification
abstract
Visible thermal person re-identification (VT-ReID) suffers from inter-modality discrepancy and intra-identity variations. Distribution alignment is a popular solution for VT-ReID, however, it is usually restricted to the influence of the intra-identity variations. In this paper, we propose the Cross-Modality Earth Mover's Distance (CM-EMD) that can alleviate the impact of the intra-identity variations during modality alignment. CM-EMD selects an optimal transport strategy and assigns high weights to pairs that have a smaller intra-identity variation. In this manner, the model will focus on reducing the inter-modality discrepancy while paying less attention to intra-identity variations, leading to a more effective modality alignment. Moreover, we introduce two techniques to improve the advantage of CM-EMD. First, Cross-Modality Discrimination Learning (CM-DL) is designed to overcome the discrimination degradation problem caused by modality alignment. By reducing the ratio between intra-identity and inter-identity variances, CM-DL leads the model to learn more discriminative representations. Second, we construct the Multi-Granularity Structure (MGS), enabling us to align modalities from both coarse- and fine-grained levels with the proposed CM-EMD. Extensive experiments show the benefits of the proposed CM-EMD and its auxiliary techniques (CM-DL and MGS). Our method achieves state-of-the-art performance on two VT-ReID benchmarks.
Yongguo Ling, Zhun Zhong, Zhiming Luo, Fengxiang Yang, Donglin Cao, Yaojin Lin, Shaozi Li, Nicu Sebe
AAAI7
2023 Exploring Non-target Knowledge for Improving Ensemble Universal Adversarial Attacks
abstract
The ensemble attack with average weights can be leveraged for increasing the transferability of universal adversarial perturbation (UAP) by training with multiple Convolutional Neural Networks (CNNs). However, after analyzing the Pearson Correlation Coefficients (PCCs) between the ensemble logits and individual logits of the crafted UAP trained by the ensemble attack, we find that one CNN plays a dominant role during the optimization. Consequently, this average weighted strategy will weaken the contributions of other CNNs and thus limit the transferability for other black-box CNNs. To deal with this bias issue, the primary attempt is to leverage the Kullback–Leibler (KL) divergence loss to encourage the joint contribution from different CNNs, which is still insufficient. After decoupling the KL loss into a target-class part and a non-target-class part, the main issue lies in that the non-target knowledge will be significantly suppressed due to the increasing logit of the target class. In this study, we simply adopt a KL loss that only considers the non-target classes for addressing the dominant bias issue. Besides, to further boost the transferability, we incorporate the min-max learning framework to self-adjust the ensemble weights for each CNN. Experiments results validate that considering the non-target KL loss can achieve superior transferability than the original KL loss by a large margin, and the min-max training can provide a mutual benefit in adversarial ensemble attacks. The source code is available at: https://github.com/WJJLL/ND-MM.
Juanjuan Weng, Zhiming Luo, Zhun Zhong, Dazhen Lin, Shaozi Li
AAAI5
2023 Frequency-Aware Attentional Feature Fusion for Deepfake Detection
abstract
Various face manipulation techniques develop rapidly and can easily generate high-quality fake images or videos, posing significant ethical concerns when used for malicious purposes. Although recent works achieve significant performance in deepfake detection, they still suffer from overfitting issues. To deal with this problem, we propose a novel framework to aggregate diverse information for deepfake detection from both RGB and frequency. Specially, we first introduce a channel attention module to assemble local and global contexts to overcome the potential semantic inconsistency on local artifacts and global features. Then we design a spatial-frequency feature fusion module to fuse the RGB-frequency information comprehensively. Moreover, a variant attention module is further proposed to improve feature discrimination. Extensive experiments demonstrate that our method maintains comparable performance in intra-dataset and cross-dataset evaluation.
Zhiming Luo, Guimin Shi, Shaozi Li
ICASSP4
2023 Boundary Difference over Union Loss for Medical Image Segmentation
Zhiming Luo, Shaozi Li
MICCAI (4)3
2023 TPNet: Enhancing Weakly Supervised Polyp Frame Detection with Temporal Encoder and Prototype-Based Memory Bank
Jianzhe Gao, Zhiming Luo, Shaozi Li
PRCV (12)4
2023 Label distribution learning with high-order label correlations
abstract
Summary Label distribution learning (LDL) is an emerging learning paradigm, which can be used to solve the label ambiguity problem. In spite of the recent great progress in LDL algorithms considering label correlations, the majority of existing methods only measure pairwise label correlations through the commonly used similarity metric, which is incapable of accurately reflecting the complex relationship between labels. To solve this problem, a novel label distribution learning method—based on high‐order label correlations (LDL‐HLC) is proposed. By virtue of the ‐regularization sparse reconstruction of the label space, the high‐order label correlations matrix is firstly obtained. Then, a new regular term can be constructed to fit the final prediction label distribution via the correction matrix. Furthermore, efficient classification performance and complete feature selection are guaranteed by common features learning via ‐regularization. Finally, the performance and effectiveness of the proposed algorithm are well illustrated through extensive experiments on 14 label distribution datasets and comparisons with some existing algorithms.
Yulin Li 0002, Yaojin Lin, Xiehua Yu, Lei Guo 0020, Shaozi Li
Concurr. Comput. Pract. Exp.5
2023 Multi-label feature selection based on relative entropy and fuzzy neighborhood mutual discrimination index
abstract
Abstract Multi‐label feature selection eliminates irrelevant and redundant features, and then improves the performance of multi‐label classification models. Most multi‐label feature selection algorithms assume that the training set contains logical labels, which means that labels are equally important for instances. However, in practical applications, there are different importances with respect to labels. To solve the problem, a multi‐label feature selection method based on relative entropy and fuzzy neighborhood mutual discriminant index is proposed. Firstly, logical labels are converted to label distribution through label enhancement. Secondly, the neighborhood and relative entropy are introduced into the label distribution, the label neighborhood similarity matrix is constructed to describe the similarity of samples under label space. Finally, the fuzzy neighborhood mutual discrimination index is used to combine the candidate features with the label neighborhood similarity matrix, which is used to judge the distinguishing ability of the candidate features. Comprehensive experiment of eight multi‐label datasets shows that the proposed algorithm has better classification performance than other compared algorithms.
Chenxi Wang 0002, Chen E, Mengli Ren, Lei Guo 0020, Xiehua Yu, Yaojin Lin, Shaozi Li
Concurr. Comput. Pract. Exp.7
2023 Online feature selection for hierarchical classification learning based on improved ReliefF
abstract
Abstract In hierarchical classification learning, the feature space of data has high dimensionality and is unknown with emergent features. To solve the above problems, we propose an online hierarchical feature selection algorithm based on adaptive ReliefF. Firstly, ReliefF is adaptively improved via using the density information of instances around the target sample, making it unnecessary to prespecify parameters. Secondly, the hierarchical relationship between classes is used, and a new method for calculating the feature weight of hierarchical data is defined. Then, an online correlation analysis method based on feature interaction is designed. Finally, the adaptive ReliefF algorithm is improved based on feature redundancy, and the feature weight is scaled by the correlation between features in order to achieve the dynamic updating of feature redundancy. A large number of experiments verify the effectiveness of the proposed algorithm.
Chenxi Wang 0002, Mengli Ren, Chen E, Lei Guo 0020, Xiehua Yu, Yaojin Lin, Shaozi Li
Concurr. Comput. Pract. Exp.7
2023 Group-preserving label-specific feature selection for multi-label learning
Jia Zhang 0019, Hanrui Wu, Min Jiang 0005, Shaozi Li, Yong Tang 0001, Jinyi Long
Expert Syst. Appl.5
2023 Toward embedding-based multi-label feature selection with label and feature collaboration
Jia Zhang 0019, Guodong Du 0002, Candong Li, Rong Wei, Shaozi Li
Neural Comput. Appl.6
2023 Towards Robust Person Re-Identification by Defending Against Universal Attackers
abstract
Recent studies show that deep person re-identification (re-ID) models are vulnerable to adversarial examples, so it is critical to improving the robustness of re-ID models against attacks. To achieve this goal, we explore the strengths and weaknesses of existing re-ID models, i.e., designing learning-based attacks and training robust models by defending against the learned attacks. The contributions of this paper are three-fold: First, we build a holistic attack-defense framework to study the relationship between the attack and defense for person re-ID. Second, we introduce a combinatorial adversarial attack that is adaptive to unseen domains and unseen model types. It consists of distortions in pixel and color space (i.e., mimicking camera shifts). Third, we propose a novel virtual-guided meta-learning algorithm for our attack-defense system. We leverage a virtual dataset to conduct experiments under our meta-learning framework, which can explore the cross-domain constraints for enhancing the generalization of the attack and the robustness of the re-ID model. Comprehensive experiments on three large-scale re-ID benchmarks demonstrate that: 1) Our combinatorial attack is effective and highly universal in cross-model and cross-dataset scenarios; 2) Our meta-learning algorithm can be readily applied to different attack and defense approaches, which can reach consistent improvement; 3) The defense model trained on the learning-to-learn framework is robust to recent SOTA attacks that are not even used during training.
Fengxiang Yang, Juanjuan Weng, Zhun Zhong, Hong Liu 0009, Zheng Wang 0007, Zhiming Luo, Donglin Cao, Shaozi Li, Shin'ichi Satoh 0001, Nicu Sebe
IEEE Trans. Pattern Anal. Mach. Intell.8
2023 Dual-Stream Transformer With Distribution Alignment for Visible-Infrared Person Re-Identification
abstract
Visible-infrared person re-identification(VI-ReID) aims to match the person images captured by visible and infrared cameras and suffers from severe cross-modality discrepancy and intra-modality variations. Existing approaches mainly use convolution neural network (CNN)-based architectures to extract pedestrian features, which fail to capture the long-range dependencies within an image. In addition, previous works usually attempt to bridge the modality gap by using adversarial learning to generate style-consistent images or designing different feature-level metric learning constraints. However, few works consider the cross-modality disparity from the perspective of assessing overall distance distribution discrepancy. To address these problems, we design a pure Transformer-based Visible-Infrared (TransVI) network with a conventional two-stream structure, which can explicitly capture modality-specific representations and learn multi-modality sharable knowledge. TransVI can efficiently address the lack of global dependency in CNN-based architectures due to the multi-head self-attention modules in the transformer, which allows us to capture the long-range dependencies of pedestrian images. Furthermore, we introduce the Cross-Modality Dissimilarity-based Maximum Mean Discrepancy (CMD-MMD) constraint to handle the cross-modality discrepancy at the distance distribution level. Specifically, CMD-MMD leverages intra-modality distribution separability to guide inter-modality distribution separability learning, aligning pair-wise distance distributions of intra- and inter-modality for within-class and between-class, respectively. In this way, the distance distributions of intra- and inter-modality become more similar, significantly mitigating the cross-modality discrepancy and learning more modality invariant representations. Extensive experimental results on two public VI-ReID datasets confirm that our proposed framework can achieve state-of-the-art performance.
Zehua Chai, Yongguo Ling, Zhiming Luo, Dazhen Lin, Min Jiang 0005, Shaozi Li
IEEE Trans. Circuits Syst. Video Technol.6
2023 Identity-Aware Contrastive Knowledge Distillation for Facial Attribute Recognition
abstract
Facial attribute recognition (FAR) is an important and yet challenging multi-label learning task in computer vision. Existing FAR methods have achieved promising performance with the development of deep learning. However, they usually suffer from prohibitive computational and memory costs. In this paper, we propose an identity-aware contrastive knowledge distillation method, termed ICKD, to compress the FAR model. A nonlinear weight-sharing mapping (NWSM) mechanism is firstly designed to avoid the difficulty of directly matching features of the teacher and student networks due to the lower representation ability of the student network. Furthermore, an identity-aware contrastive distillation (ICD) loss is employed to guide the student network to effectively learn the mutual relations between samples with multiple attributes. In addition, an adjustable ladder distillation (ALD) loss is developed to automatically adjust the importance of different distillation points with the progress of training. Extensive experiments demonstrate that our method can significantly improve the performance of student networks and outperforms the existing FAR methods on the public challenging datasets.
Si Chen 0002, Xueyan Zhu, Yan Yan 0001, Shunzhi Zhu, Shaozi Li, Dahan Wang
IEEE Trans. Circuits Syst. Video Technol.5
2023 Logit Margin Matters: Improving Transferable Targeted Adversarial Attack by Logit Calibration
abstract
Previous works have extensively studied the transferability of adversarial samples in untargeted black-box scenarios. However, it still remains challenging to craft targeted adversarial examples with higher transferability than non-targeted ones. Recent studies reveal that the traditional Cross-Entropy (CE) loss function is insufficient to learn transferable targeted adversarial examples due to the issue of vanishing gradient. In this work, we provide a comprehensive investigation of the CE loss function and find that the logit margin between the targeted and untargeted classes will quickly obtain saturation in CE, which largely limits the transferability. Therefore, in this paper, we devote to the goal of continually increasing the logit margin along the optimization to deal with the saturation issue and propose two simple and effective logit calibration methods, which are achieved by downscaling the logits with a temperature factor and an adaptive margin, respectively. Both of them can effectively encourage optimization to produce a larger logit margin and lead to higher transferability. Besides, we show that minimizing the cosine distance between the adversarial examples and the classifier weights of the target class can further improve the transferability, which is benefited from downscaling logits via L2-normalization. Experiments conducted on the ImageNet dataset validate the effectiveness of the proposed methods, which outperform the state-of-the-art methods in black-box targeted attacks. The source code is available at Link.
Juanjuan Weng, Zhiming Luo, Shaozi Li, Nicu Sebe, Zhun Zhong
IEEE Trans. Inf. Forensics Secur.3
2023 Graph-Based Class-Imbalance Learning With Label Enhancement
abstract
Class imbalance is a common issue in the community of machine learning and data mining. The class-imbalance distribution can make most classical classification algorithms neglect the significance of the minority class and tend toward the majority class. In this article, we propose a label enhancement method to solve the class-imbalance problem in a graph manner, which estimates the numerical label and trains the inductive model simultaneously. It gives a new perspective on the class-imbalance learning based on the numerical label rather than the original logical label. We also present an iterative optimization algorithm and analyze the computation complexity and its convergence. To demonstrate the superiority of the proposed method, several single-label and multilabel datasets are applied in the experiments. The experimental results show that the proposed method achieves a promising performance and outperforms some state-of-the-art single-label and multilabel class-imbalance learning methods.
Guodong Du 0002, Jia Zhang 0019, Min Jiang 0005, Jinyi Long, Yaojin Lin, Shaozi Li, Kay Chen Tan
IEEE Trans. Neural Networks Learn. Syst.6
2023 SFANet: A Spectrum-Aware Feature Augmentation Network for Visible-Infrared Person Reidentification
abstract
Visible-Infrared person reidentification (VI-ReID) is a challenging matching problem due to large modality variations between visible and infrared images. Existing approaches usually bridge the modality gap with only feature-level constraints, ignoring pixel-level variations. Some methods employ a generative adversarial network (GAN) to generate style-consistent images, but it destroys the structure information and incurs a considerable level of noise. In this article, we explicitly consider these challenges and formulate a novel spectrum-aware feature augmentation network named SFANet for cross-modality matching problem. Specifically, we put forward to employ grayscale-spectrum images to fully replace RGB images for feature learning. Learning with the grayscale-spectrum images, our model can apparently reduce modality discrepancy and detect inner structure relations across the different modalities, making it robust to color variations. At feature level, we improve the conventional two-stream network by balancing the number of specific and sharable convolutional blocks, which preserve the spatial structure information of features. Additionally, a bidirectional tri-constrained top-push ranking loss (BTTR) is embedded in the proposed network to improve the discriminability, which efficiently further boosts the matching accuracy. Meanwhile, we further introduce an effective dual-linear with batch normalization identification (ID) embedding method to model the identity-specific information and assist BTTR loss in magnitude stabilizing. On SYSU-MM01 and RegDB datasets, we conducted extensively experiments to demonstrate that our proposed framework contributes indispensably and achieves a very competitive VI-ReID performance.
Shun Ma, Daoxun Xia, Shaozi Li
IEEE Trans. Neural Networks Learn. Syst.4
2022 Generalized Person Re-identification by Locating and Eliminating Domain-Sensitive Features
Fengxiang Yang, Zhiming Luo, Shaozi Li
ACCV (6)4
2022 Symmetrical Supervision with Transformer for Few-shot Medical Image Segmentation
abstract
Few-shot learning can potentially learn the target knowledge in extremely few data regimes. Existing few-shot medical image segmentation methods fail to consider the global anatomy correlation between the support and query sets. They generally adopt a weak one-way information transmission that can not fully explore the knowledge to segment query data. To address this problem, we propose a novel Symmetrical Supervision network based on traditional two-branch methods. We raise two main contributions: (1) The Symmetrical Supervision Mechanism is leveraged to strengthen the supervision of network training; (2) A transformer-based Global Feature Alignment module is introduced to increase the global consistency between the two branches. Experimental results on two challenging datasets (abdominal segmentation dataset CHAOS and cardiac segmentation dataset MS-CMRSeg) show a remarkable performance compared to other comparing methods.
Yao Niu, Zhiming Luo, Sheng Lian, Lei Li 0048, Shaozi Li, Haixin Song
BIBM5
2022 Dynamic Selection Network For Rgb-D Salient Object Detection
abstract
Existing RGB-D salient object detection (SOD) methods usually use elaborate fusion modules for exploring cross-modal information, which is computationally expensive and ignores the noise depth information. To deal with this issue, we propose a dynamic selection network (DSNet) for RGB-D salient object detection. Specifically, a cross-modal combination module (CCM) is proposed to fuse two modalities with a light computation. Then a dynamic selection module (DSM) adaptively learns the model parameter for the decoding based on the fused features. Furthermore, skip connection is used for hierarchical features combination between encoder and decoder. Experiments on four popular datasets demonstrate our model outperforms other state-of-the-art methods.
Jinlin Zhou, Zhiming Luo, Shaozi Li
ICIP3
2022 Hierarchical community-discovery algorithm combining core nodes and three-order structure model
abstract
Abstract A community structure in a complex network often exhibits hierarchical characteristics. Current hierarchical community‐discovery algorithms generally consider a single node as a community during the initial stage. This approach leads to over‐fine clustering granularity, too‐deep clustering levels, and other issues. Therefore, this article proposes a hierarchical community‐discovery algorithm that combines the core nodes and the three‐order structure model. Between neighboring nodes, there is a first‐order structure. The core node is identified based on its influence, and the similarity between the core node and its neighboring nodes is defined as the second‐order structure. The nodes satisfying the second‐order structure are then formed into a friend circle. The similarity between friend circles is defined as the third‐order structure. According to this structure, the friend circles are construed as a hierarchical clustering tree (HCT) where one HCT represents a community. The HCT built by this algorithm has relatively fewer levels and exhibits a flat feature. Experimental results on both artificial and real networks show that the algorithm performs well on various indicators. Additionally, the algorithm exhibits near‐linear time complexity.
Lei Guo 0020, Shaozi Li, Qingshou Wu
Concurr. Comput. Pract. Exp.3
2022 Consistent response for automated multilabel thoracic disease classification
abstract
Summary While recent studies on automated multilabel chest X‐ray (CXR) images classification have shown remarkable progress in leveraging complicated network and attention mechanisms, the automated detection on chest radiographs is still challenging because the pathological patterns are usually highly diverse in their sizes and locations. The CNN model will suffer from the complicated background and high diversity of diseases, which reduce the generalization and performance of the model. To solve these problems, we propose a dual‐distribution consistency (DDC) model, which increases the consistency from two aspects, that is, feature‐level and label‐level. This model integrates two novel loss functions: multilabel response consistency (MRC) loss and distribution consistency (DC) loss. Specifically, we use the original image and its transformed image as inputs to imitate different views of CXR images. The MRC loss encourages the multilabel‐wise attention maps to be consistent between the original CXR image and its transformed counterpart. And the DC loss can force their output probability distributions to be uniform. In this manner, we can make sure that the model can learn discriminative features by using a different view of CXR images. Experiments conducted on the ChestX‐ray14 dataset show the effectiveness of the proposed method.
Jiawei Su, Zhiming Luo, Shaozi Li
Concurr. Comput. Pract. Exp.3
2022 Multilabel causal variable discovery in multisource
abstract
Abstract Multilabel causal feature selection, as a well‐known and effective approach in dealing with high‐dimensional multilabel data, is a popular topic. Amount of causal feature selection algorithms have achieved a great deal of success in classification and prediction tasks. However, the descriptive information of data is collected from different data sources in many practical applications. While few researches focus on the causal variable discovery in multisource environments due to the complex causal relationships. To address these problems, we propose a causal feature selection framework in multisource environments to solve the above problems. Firstly, we mine the causal mechanism with respect to the class attribute under the assumption that only a single data source is included. Secondly, by utilizing the concept of causal invariance in causal inference, we formulate the problem of causal feature selection with multiple data sources as a search problem for an invariant set across data sources. In addition, we give the upper and lower bounds of the causal invariant set. Finally, we design a novel multisource multilabel causal feature selection (MMCFS) algorithm. To verify the effectiveness of the proposed algorithm, we compare it with 12 feature selection methods on synthetic datasets. Experiment results show that the classification performance of MMCFS achieves highly competitive performance against other comparing algorithms.
Yun-an Wang, Yaojin Lin, Xiehua Yu, Zhisen Wei, Shaozi Li
Concurr. Comput. Pract. Exp.7
2022 Online streaming feature selection for multigranularity hierarchical classification learning
abstract
Abstract Hierarchical classification learning is a hot research topic in machine learning and data mining domains, and many feature selection algorithms with category hierarchy have been proposed. However, existing algorithms assume that the feature space of data is completely obtained in advance, and ignore its uncertainty and dynamicity. To address these problems, we propose an online streaming feature selection framework with a hierarchical structure to solve the above two problems simultaneously. First, we apply the hierarchical relationship between nodes in a hierarchical structure to the Relief algorithm, so that it can be used to compute the weights of dynamic features. Second, we dynamically select important features for each internal node via comparing the magnitude of the weights of features on these nodes with their parent and sibling nodes. In addition, we perform redundancy analysis of features by calculating the covariance between features to obtain a superior online feature subset for each internal node. Finally, the proposed algorithm is compared with six online streaming feature selection methods on six hierarchical data sets, and experimental results shows that the classification performance of the proposed algorithm is effective.
Chenxi Wang 0002, Xiaoqing Zhang 0019, Liqin Ye, Shaozi Li, Yaojin Lin
Concurr. Comput. Pract. Exp.5
2022 Neighborhood rough set based multi-label feature selection with label correlation
abstract
Summary Neighborhood rough set (NRS) is considered as an effective tool for feature selection and has been widely used in processing high‐dimensional data. However, most of the existing methods are difficult to deal with multi‐label data and are lack of considering label correlation (LC), which is an important issue in multi‐label learning. Therefore, in this article, we introduce a new NRS model with considering LC. First, we explore LC by calculating the similarity relation between labels and divide the related labels into several label subsets. Then, a new neighborhood relation is proposed, which can solve the problem of neighborhood granularity selection by using the nearest neighbor information distribution of instances under the related labels. On this basis, the NRS model is reconstructed by embedding LC information, and the related properties of the model are discussed. Moreover, we design a new feature significance function to evaluate the quality of features, which can well capture the specific relationship between features and labels. Finally, a greedy forward feature selection algorithm is designed. Extensive experiments which are conducted on different types of datasets verify the effectiveness of the proposed algorithm.
Yilin Wu 0001, Xiehua Yu, Yaojin Lin, Shaozi Li
Concurr. Comput. Pract. Exp.5
2022 Stock movement prediction via gated recurrent unit network based on reinforcement learning with incorporated attention mechanisms
Hongfeng Xu, Lei Chai, Zhiming Luo, Shaozi Li
Neurocomputing4
2022 From the whole to detail: Progressively sampling discriminative parts for fine-grained recognition
Yaojin Lin, Shengyu Chen, Zhichun Zeng, Ming-Wen Shao, Shaozi Li
Knowl. Based Syst.6
2022 A self-regulated generative adversarial network for stock price movement prediction based on the historical price and tweets
Hongfeng Xu, Donglin Cao, Shaozi Li
Knowl. Based Syst.3
2022 Improving embedding learning by virtual attribute decoupling for text-based person search
Chengji Wang, Zhiming Luo, Yaojin Lin, Shaozi Li
Neural Comput. Appl.4
2022 Learning From Weakly Labeled Data Based on Manifold Regularized Sparse Model
abstract
In multilabel learning, each training example is represented by a single instance, which is relevant to multiple class labels simultaneously. Generally, all relevant labels are considered to be available for labeled data. However, instances with a full label set are difficult to obtain in real-world applications, thus leading to the weakly multilabel learning problem, that is, relevant labels of training data are partially known and many relevant labels are missing, and even abundant training data are associated with an empty label set. To address the problem, we propose a new multilabel method to learn from weakly labeled data. To be specific, an optimization framework is constructed based on the manifold regularized sparse model, in which the correlations among labels and feature structure are considered to model global and local label correlations, thereby achieving discriminative feature analysis for mapping training data to ground-truth label space. Moreover, the proposed method has an excellent mechanism to conduct semisupervised multilabel learning by exploiting training data with the predicted label set of the unlabeled. Experiments on various real-world tasks reveal that the proposed method outperforms some state-of-the-art methods.
Jia Zhang 0019, Shaozi Li, Min Jiang 0005, Kay Chen Tan
IEEE Trans. Cybern.2
2022 An Online Prediction Approach Based on Incremental Support Vector Machine for Dynamic Multiobjective Optimization
abstract
Real-world multiobjective optimization problems usually involve conflicting objectives that change over time, which requires the optimization algorithms to quickly track the Pareto-optimal front (POF) when the environment changes. In recent years, evolutionary algorithms based on prediction models have been considered promising. However, most existing approaches only make predictions based on the linear correlation between a finite number of optimal solutions in two or three previous environments. These incomplete information extraction strategies may lead to low prediction accuracy in some instances. In this article, an incremental support vector machine (ISVM)-based dynamic multiobjective evolutionary algorithm, in short called ISVM-DMOEA, is proposed. We treat the solving of dynamic multiobjective optimization problems (DMOPs) as an online learning process, using the continuously obtained optimal solution to update an ISVM without discarding the solution information at earlier time. ISVM is then used to filter random solutions and generate an initial population for the next moment. To overcome the obstacle of insufficient training samples, a synthetic minority oversampling strategy is implemented before the training of ISVM. The advantage of this approach is that the nonlinear correlation between solutions can be explored online by ISVM, and the information contained in all historical optimal solutions can be exploited to a greater extent. The experimental results and comparison with the chosen state-of-the-art algorithms demonstrate that the proposed algorithm can effectively tackle DMOPs.
Dejun Xu, Min Jiang 0005, Weizhen Hu, Shaozi Li, Renhu Pan, Gary G. Yen
IEEE Trans. Evol. Comput.4
2022 A Spectral Feature Selection Approach With Kernelized Fuzzy Rough Sets
abstract
Feature evaluation is an important issue in constructing a feature selection algorithm in kernelized fuzzy rough sets, which has been proven to be an effective approach to deal with nonlinear classification tasks and uncertainty in learning problems. However, the feature evaluation function developed with kernelized fuzzy rough sets cannot better reflect the affinity relationship of samples and is time-consuming. To overcome these drawbacks, in this article, the problem of feature selection with kernelized fuzzy rough sets is studied based on the spectral graph theory. First, the within-class and between-class sample similarity matrices by using kernelized fuzzy approximation operators are constructed. Two operators, which can capture the affinity relationship of samples, are then introduced based on the sample similarity matrices. The proposed operator can be regarded as the sum of the weighted kernelized fuzzy approximation operators. Second, based on the ratio criterion, a feature evaluation function and its corresponding feature selection algorithm FRKF are presented, which can effectively evaluate the importance of features. Third, to illustrate the performance of the proposed algorithm, extensive experiments have been carried out to compare FRKF and other well-known feature selection methods, including the feature ranking methods and feature subset selection methods on various classification tasks. The experimental results on real-world datasets demonstrate that FRKF achieves the high performances in terms of the robustness, efficiency, and effectiveness.
Jinkun Chen, Yaojin Lin, Ju-Sheng Mi, Shaozi Li, Weiping Ding 0001
IEEE Trans. Fuzzy Syst.4
2022 Joint Representation Learning and Keypoint Detection for Cross-View Geo-Localization
abstract
In this paper, we study the cross-view geo-localization problem to match images from different viewpoints. The key motivation underpinning this task is to learn a discriminative viewpoint-invariant visual representation. Inspired by the human visual system for mining local patterns, we propose a new framework called RK-Net to jointly learn the discriminative Representation and detect salient Keypoints with a single Network. Specifically, we introduce a Unit Subtraction Attention Module (USAM) that can automatically discover representative keypoints from feature maps and draw attention to the salient regions. USAM contains very few learning parameters but yields significant performance improvement and can be easily plugged into different networks. We demonstrate through extensive experiments that (1) by incorporating USAM, RK-Net facilitates end-to-end joint learning without the prerequisite of extra annotations. Representation learning and keypoint detection are two highly-related tasks. Representation learning aids keypoint detection. Keypoint detection, in turn, enriches the model capability against large appearance changes caused by viewpoint variants. (2) USAM is easy to implement and can be integrated with existing methods, further improving the state-of-the-art performance. We achieve competitive geo-localization accuracy on three challenging datasets, i. e., University-1652, CVUSA and CVACT. Our code is available at https://github.com/AggMan96/RK-Net.
Jinliang Lin, Zhedong Zheng, Zhun Zhong, Zhiming Luo, Shaozi Li, Yi Yang 0001, Nicu Sebe
IEEE Trans. Image Process.5
2021 Learning to Attack Real-World Models for Person Re-identification via Virtual-Guided Meta-Learning
abstract
Recent advances in person re-identification (re-ID) have led to impressive retrieval accuracy. However, existing re-ID models are challenged by the adversarial examples crafted by adding quasi-imperceptible perturbations. Moreover, re-ID systems face the domain shift issue that training and testing domains are not consistent. In this study, we argue that learning powerful attackers with high universality that works well on unseen domains is an important step in promoting the robustness of re-ID systems. Therefore, we introduce a novel universal attack algorithm called ``MetaAttack'' for person re-ID. MetaAttack can mislead re-ID models on unseen domains by a universal adversarial perturbation. Specifically, to capture common patterns across different domains, we propose a meta-learning scheme to seek the universal perturbation via the gradient interaction between meta-train and meta-test formed by two datasets. We also take advantage of a virtual dataset (PersonX), instead of real ones, to conduct meta-test. This scheme not only enables us to learn with more comprehensive variation factors but also mitigates the negative effects caused by biased factors of real datasets. Experiments on three large-scale re-ID datasets demonstrate the effectiveness of our method in attacking re-ID models on unseen domains. Our final visualization results reveal some new properties of existing re-ID systems, which can guide us in designing a more robust re-ID model. Code and supplemental material are available at \url{https://github.com/FlyingRoastDuck/MetaAttack_AAAI21}.
Fengxiang Yang, Zhun Zhong, Hong Liu 0009, Zheng Wang 0007, Zhiming Luo, Shaozi Li, Nicu Sebe, Shin'ichi Satoh 0001
AAAI6
2021 Joint Noise-Tolerant Learning and Meta Camera Shift Adaptation for Unsupervised Person Re-Identification
abstract
This paper considers the problem of unsupervised person re-identification (re-ID), which aims to learn discriminative models with unlabeled data. One popular method is to obtain pseudo-label by clustering and use them to optimize the model. Although this kind of approach has shown promising accuracy, it is hampered by 1) noisy labels produced by clustering and 2) feature variations caused by camera shift. The former will lead to incorrect optimization and thus hinders the model accuracy. The latter will result in assigning the intra-class samples of different cameras to different pseudo-label, making the model sensitive to camera variations. In this paper, we propose a unified framework to solve both problems. Concretely, we propose a Dynamic and Symmetric Cross-Entropy loss (DSCE) to deal with noisy samples and a camera-aware meta-learning algorithm (MetaCam) to adapt camera shift. DSCE can alleviate the negative effects of noisy samples and accommodate the change of clusters after each clustering step. MetaCam simulates cross-camera constraint by splitting the training data into meta-train and meta-test based on camera IDs. With the interacted gradient from meta-train and meta-test, the model is enforced to learn camera-invariant features. Extensive experiments on three re-ID benchmarks show the effectiveness and the complementary of the proposed DSCE and MetaCam. Our method outperforms the state-of-the-art methods on both fully unsupervised re-ID and unsupervised domain adaptive re-ID.
Fengxiang Yang, Zhun Zhong, Zhiming Luo, Yuanzheng Cai, Yaojin Lin, Shaozi Li, Nicu Sebe
CVPR6
2021 Learning to Generalize Unseen Domains via Memory-based Multi-Source Meta-Learning for Person Re-Identification
abstract
Recent advances in person re-identification (ReID) obtain impressive accuracy in the supervised and unsupervised learning settings. However, most of the existing methods need to train a new model for a new domain by accessing data. Due to public privacy, the new domain data are not always accessible, leading to a limited applicability of these methods. In this paper, we study the problem of multi-source domain generalization in ReID, which aims to learn a model that can perform well on unseen domains with only several labeled source domains. To address this problem, we propose the Memory-based Multi-Source Meta-Learning (M3L) framework to train a generalizable model for unseen domains. Specifically, a meta-learning strategy is introduced to simulate the train-test process of domain generalization for learning more generalizable models. To overcome the unstable meta-optimization caused by the parametric classifier, we propose a memory-based identification loss that is non-parametric and harmonizes with meta-learning. We also present a meta batch normalization layer (MetaBN) to diversify meta-test features, further establishing the advantage of meta-learning. Experiments demonstrate that our M3L can effectively enhance the generalization ability of the model for unseen domains and can outperform the state-of-the-art methods on four large-scale ReID datasets.
Zhun Zhong, Fengxiang Yang, Zhiming Luo, Yaojin Lin, Shaozi Li, Nicu Sebe
CVPR6
2021 OpenMix: Reviving Known Knowledge for Discovering Novel Visual Categories in an Open World
abstract
In this paper, we tackle the problem of discovering new classes in unlabeled visual data given labeled data from disjoint classes. Existing methods typically first pre-train a model with labeled data, and then identify new classes in unlabeled data via unsupervised clustering. However, the labeled data that provide essential knowledge are often underexplored in the second step. The challenge is that the labeled and unlabeled examples are from non-overlapping classes, which makes it difficult to build a learning relationship between them. In this work, we introduce Open-Mix to mix the unlabeled examples from an open set and the labeled examples from known classes, where their non-overlapping labels and pseudo-labels are simultaneously mixed into a joint label distribution. OpenMix dynamically compounds examples in two ways. First, we produce mixed training images by incorporating labeled examples with unlabeled examples. With the benefit of unique prior knowledge in novel class discovery, the generated pseudo-labels will be more credible than the original unlabeled predictions. As a result, OpenMix helps preventing the model from overfitting on unlabeled samples that may be assigned with wrong pseudo-labels. Second, the first way encourages the unlabeled examples with high class-probabilities to have considerable accuracy. We introduce these examples as reliable anchors and further integrate them with un-labeled samples. This enables us to generate more combinations in unlabeled examples and exploit finer object relations among the new classes. Experiments on three classification datasets demonstrate the effectiveness of the proposed OpenMix, which is superior to state-of-the-art methods in novel class discovery.
Zhun Zhong, Linchao Zhu, Zhiming Luo, Shaozi Li, Yi Yang 0001, Nicu Sebe
CVPR4
2021 A Multi-Constraint Similarity Learning with Adaptive Weighting for Visible-Thermal Person Re-Identification
abstract
The challenges of visible-thermal person re-identification (VT-ReID) lies in the inter-modality discrepancy and the intra-modality variations. An appropriate metric learning plays a crucial role in optimizing the feature similarity between the two modalities. However, most existing metric learning-based methods mainly constrain the similarity between individual instances or class centers, which are inadequate to explore the rich data relationships in the cross-modality data. Besides, most of these methods fail to consider the importance of different pairs, incurring an inefficiency and ineffectiveness of optimization. To address these issues, we propose a Multi-Constraint (MC) similarity learning method that jointly considers the cross-modality relationships from three different aspects, i.e., Instance-to-Instance (I2I), Center-to-Instance (C2I), and Center-to-Center (C2C). Moreover, we devise an Adaptive Weighting Loss (AWL) function to implement the MC efficiently. In the AWL, we first use an adaptive margin pair mining to select informative pairs and then adaptively adjust weights of mined pairs based on their similarity. Finally, the mined and weighted pairs are used for the metric learning. Extensive experiments on two benchmark datasets demonstrate the superior performance of the proposed over the state-of-the-art methods.
Yongguo Ling, Zhiming Luo, Yaojin Lin, Shaozi Li
IJCAI4
2021 Text-based Person Search via Multi-Granularity Embedding Learning
abstract
Most existing text-based person search methods highly depend on exploring the corresponding relations between the regions of the image and the words in the sentence. However, these methods correlated image regions and words in the same semantic granularity. It 1) results in irrelevant corresponding relations between image and text, 2) causes an ambiguity embedding problem. In this study, we propose a novel multi-granularity embedding learning model for text-based person search. It generates multi-granularity embeddings of partial person bodies in a coarse-to-fine manner by revisiting the person image at different spatial scales. Specifically, we distill the partial knowledge from image scrips to guide the model to select the semantically relevant words from the text description. It can learn discriminative and modality-invariant visual-textual embeddings. In addition, we integrate the partial embeddings at each granularity and perform multi-granularity image-text matching. Extensive experiments validate the effectiveness of our method, which can achieve new state-of-the-art performance by the learned discriminative partial embeddings.
Chengji Wang, Zhiming Luo, Yaojin Lin, Shaozi Li
IJCAI4
2021 Learning Consistency- and Discrepancy-Context for 2D Organ Segmentation
Lei Li 0048, Sheng Lian, Zhiming Luo, Shaozi Li, Beizhan Wang, Shuo Li 0001
MICCAI (1)4
2021 Recommendation algorithm based on community structure and user trust
abstract
Abstract While contemporary community‐based recommendation algorithms based on a single community structure are more capable of processing large datasets than ever, they lack recommendation precision. This article proposes a collaborative filtering recommendation algorithm that integrates community structure and user implicit trust. The algorithm first applies a method based on the Gaussian function to fill the matrix of item ratings of users to alleviate data sparsity. It then uses the trust matrix to obtain the asymmetric trust relationship of the trustor and trustee, based on which the degree of users' implicit trust is calculated. The users are divided into communities based on the implicit trust degree to determine the influence among users more accurately. The algorithm then predicts the target user's rating using the ratings of users in the community to generate recommendations. To verify the performance of the proposed algorithm, we compared the proposed algorithm with three contemporary algorithms under the same conditions using FilmTrust datasets. The recommendation accuracy as well as the mean absolute error and root mean square error values of the proposed algorithm were better than those of the other four algorithms by approximately 14% and 4%, respectively. The experimental results demonstrate that the proposed algorithm can achieve better recommendation efficiency than existing algorithms.
Lei Guo 0020, Shaozi Li, Qingshou Wu, Wensen Yu
Concurr. Comput. Pract. Exp.3
2021 Causality-based online streaming feature selection
abstract
Abstract Online streaming feature selection, as a well‐known and effective preprocessing approach in machine learning, is an eternal topic. Amount of online streaming feature selection algorithms have achieved a great deal of success in classification and prediction tasks. However, most of these existing algorithms only concentrate on the relevance between features and labels, and neglect the causal relationships between them. Discovering the potential causal relationships between features and labels, that is, the Markov blanket (MB) of class label, which can build a more interpretable and robust classification model. In this paper, we put forward a causality‐based online streaming feature selection algorithm with neighborhood conditional mutual information. First, we apply neighborhood symmetrical uncertainty to discover a candidate Markov blanket (CMB) with causal information. Then, neighborhood conditional mutual information instead of conditional independence test is used to delete the false positives in CMB, which can significantly alleviate the computational cost. Moreover, we utilize the updated CMB to choose the true spouses, which may be mistakenly deleted during the process of removing false positives, and then acquire an optimal MB as the online selected feature subset. Finally, causality‐based online streaming feature selection with neighborhood conditional mutual information is compared with four well‐established online streaming feature selection methods on 13 real‐world datasets. Experiment results show that the proposed algorithm outperforms these online streaming feature selection algorithms.
Longzhu Li, Yaojin Lin, Hong Zhao 0002, Jinkun Chen, Shaozi Li
Concurr. Comput. Pract. Exp.5
2021 Feature interaction based online streaming feature selection via buffer mechanism
abstract
Abstract Feature selection is a nontrivial preprocessing technique in many practical application domains. There are three key challenges with respect to real‐world data. Firstly, the dimensionality of data keeps growing and will achieve hundreds of millions. Secondly, the data has the characteristic of high‐dimensional and small‐size. Thirdly, practical applications need to process each feature in an online manner. However, most of the previous methods only pay much attention to solving the challenges of high dimensionality and online stream. To address all issues above, we propose OFSI, in this article, an nline streaming eature election based on feature nteraction method for feature selection. OFSI can effectively select the streaming features that are strongly related to each other in high‐dimensional and small‐size data, via using feature interaction. Furthermore, to address upcoming features that arrive by groups, we present a new group‐OFSI algorithm for online group feature selection. An extensive experiment using a series of benchmark data sets shows that the proposed two algorithms, OFSI and group‐OFSI, outperform six state‐of‐the‐art online streaming feature selection methods.
Yaojin Lin, Xiangyan Chen, Chenxi Wang 0002, Shaozi Li
Concurr. Comput. Pract. Exp.5
2021 A weakly supervised tooth-mark and crack detection method in tongue image
abstract
Abstract Tongue diagnosis is one of the primary clinical diagnostic methods in Traditional Chinese Medicine. Recognizing the tooth‐marked tongue and the crackled tongue plays an essential role in evaluating the status of patients. Previous methods mainly focus on identifying whether a tongue image is a tooth‐marked tongue (cracked tongue) or not, while cannot provide more details. In this study, we propose a weakly supervised method for training the tooth‐mark and crack detection model by leveraging fully bounding‐box level annotated and coarse image‐level annotated tongue images. The proposed model is extended from the YOLO object detection model, and we add several classification branches for recognizing the tooth‐marked tongue and cracked tongue. The classification branch aims to predict the coarse label for both coarse‐labeled data and fully annotated data. The detection branch is used to locate the position of tooth marks and cracks from the fully annotated data. Finally, we utilize a multitask loss function for training the model. Experimental results on a challenging tongue image dataset demonstrate the effectiveness of our proposed weakly supervised method.
Hui Weng, Lei Li 0048, Huangwei Lei, Zhiming Luo, Candong Li, Shaozi Li
Concurr. Comput. Pract. Exp.6
2021 Grammar guided embedding based Chinese long text sentiment classification
abstract
Abstract Although the state‐of‐the‐art sentiment classification approaches, such as LSTM and TextCNN, have achieved a good performance on Chinese short text sentiment analysis, the Chinese long text sentiment classification is still a challenge because of the sentiment change problem and the long text structure problem. Therefore, we propose a grammar guided embedding model (GGE) and a novel Chinese long text sentiment classification framework. First, the part‐of‐speech (POS) tags are introduced as the Chinese long text grammar guided information which can help classification approaches to model the Chinese long text structure and the important structure of sentiment change. Second, we proposed a simple GGE training method which considers the combination representation of word sequence and POS sequence. Finally, we proposed a unified framework which combines our novel GGE with TextCNN. Experiment results show that after using GGE, the model outperforms the state‐of‐the‐art approaches. At the same time, we also found that the GGE achieves the model converge faster, that is, it can achieve better results than without GGE when there is only a small amount of training data. Thus, we believe that the GGE can help machines better understand human language sentiment expression structure.
Dazhen Lin, Donglin Cao, Shaozi Li
Concurr. Comput. Pract. Exp.4
2021 Towards graph-based class-imbalance learning for hospital readmission
Guodong Du 0002, Jia Zhang 0019, Fenglong Ma, Yaojin Lin, Shaozi Li
Expert Syst. Appl.6
2021 Learning from class-imbalance and heterogeneous data for 30-day hospital readmission
Guodong Du 0002, Jia Zhang 0019, Shaozi Li, Candong Li
Neurocomputing3
2021 High-Order-Interaction for weakly supervised Fine-Grained Visual Categorization
Nanyu Li, Zhiming Luo, Zhun Zhong, Shaozi Li
Neurocomputing5
2021 Divide-and-Merge the embedding space for cross-modality person search
Chengji Wang, Zhiming Luo, Zhun Zhong, Shaozi Li
Neurocomputing4
2021 APRIL: Anatomical prior-guided reinforcement learning for accurate carotid lumen diameter and intima-media thickness measurement
Sheng Lian, Zhiming Luo, Shaozi Li, Shuo Li 0001
Medical Image Anal.4
2021 SAFD: single shot anchor free face detector
Chengji Wang, Zhiming Luo, Zhun Zhong, Shaozi Li
Multim. Tools Appl.4
2021 Learning to Adapt Invariance in Memory for Person Re-Identification
abstract
This work considers the problem of unsupervised domain adaptation in person re-identification (re-ID), which aims to transfer knowledge from the source domain to the target domain. Existing methods are primary to reduce the inter-domain shift between the domains, which however usually overlook the relations among target samples. This paper investigates into the intra-domain variations of the target domain and proposes a novel adaptation framework w.r.t three types of underlying invariance, i.e., Exemplar-Invariance, Camera-Invariance, and Neighborhood-Invariance. Specifically, an exemplar memory is introduced to store features of samples, which can effectively and efficiently enforce the invariance constraints over the global dataset. We further present the Graph-based Positive Prediction (GPP) method to explore reliable neighbors for the target domain, which is built upon the memory and is trained on the source samples. Experiments demonstrate that 1) the three invariance properties are complementary and indispensable for effective domain adaptation, 2) the memory plays a key role in implementing invariance learning and improves the performance with limited extra computation cost, 3) GPP can facilitate the invariance learning and thus significantly improves the results, and 4) our approach produces new state-of-the-art adaptation accuracy on three re-ID large-scale benchmarks.
Zhun Zhong, Liang Zheng 0001, Zhiming Luo, Shaozi Li, Yi Yang 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2021 A Global and Local Enhanced Residual U-Net for Accurate Retinal Vessel Segmentation
abstract
Retinal vessel segmentation is a critical procedure towards the accurate visualization, diagnosis, early treatment, and surgery planning of ocular diseases. Recent deep learning-based approaches have achieved impressive performance in retinal vessel segmentation. However, they usually apply global image pre-processing and take the whole retinal images as input during network training, which have two drawbacks for accurate retinal vessel segmentation. First, these methods lack the utilization of the local patch information. Second, they overlook the geometric constraint that retina only occurs in a specific area within the whole image or the extracted patch. As a consequence, these global-based methods suffer in handling details, such as recognizing the small thin vessels, discriminating the optic disk, etc. To address these drawbacks, this study proposes a Global and Local enhanced residual U-nEt (GLUE) for accurate retinal vessel segmentation, which benefits from both the globally and locally enhanced information inside the retinal region. Experimental results on two benchmark datasets demonstrate the effectiveness of the proposed method, which consistently improves the segmentation accuracy over a conventional U-Net and achieves competitive performance compared to the state-of-the-art.
Sheng Lian, Lei Li 0048, Guiren Lian, Zhiming Luo, Shaozi Li
IEEE ACM Trans. Comput. Biol. Bioinform.6
2020 Asymmetric Co-Teaching for Unsupervised Cross-Domain Person Re-Identification
abstract
Person re-identification (re-ID), is a challenging task due to the high variance within identity samples and imaging conditions. Although recent advances in deep learning have achieved remarkable accuracy in settled scenes, i.e., source domain, few works can generalize well on the unseen target domain. One popular solution is assigning unlabeled target images with pseudo labels by clustering, and then retraining the model. However, clustering methods tend to introduce noisy labels and discard low confidence samples as outliers, which may hinder the retraining process and thus limit the generalization ability. In this study, we argue that by explicitly adding a sample filtering procedure after the clustering, the mined examples can be much more efficiently used. To this end, we design an asymmetric co-teaching framework, which resists noisy labels by cooperating two models to select data with possibly clean labels for each other. Meanwhile, one of the models receives samples as pure as possible, while the other takes in samples as diverse as possible. This procedure encourages that the selected training samples can be both clean and miscellaneous, and that the two models can promote each other iteratively. Extensive experiments show that the proposed framework can consistently benefit most clustering based methods, and boost the state-of-the-art adaptation accuracy. Our code is available at https://github.com/FlyingRoastDuck/ACT_AAAI20.
Fengxiang Yang, Ke Li 0015, Zhun Zhong, Zhiming Luo, Xing Sun 0001, Hao Cheng 0012, Feiyue Huang, Rongrong Ji, Shaozi Li
AAAI10
2020 Random Erasing Data Augmentation
abstract
In this paper, we introduce Random Erasing, a new data augmentation method for training the convolutional neural network (CNN). In training, Random Erasing randomly selects a rectangle region in an image and erases its pixels with random values. In this process, training images with various levels of occlusion are generated, which reduces the risk of over-fitting and makes the model robust to occlusion. Random Erasing is parameter learning free, easy to implement, and can be integrated with most of the CNN-based recognition models. Albeit simple, Random Erasing is complementary to commonly used data augmentation techniques such as random cropping and flipping, and yields consistent improvement over strong baselines in image classification, object detection and person re-identification. Code is available at: https://github.com/zhunzhong07/Random-Erasing.
Zhun Zhong, Liang Zheng 0001, Guoliang Kang, Shaozi Li, Yi Yang 0001
AAAI4
2020 Multi-label Feature Selection via Global Relevance and Redundancy Optimization
abstract
Information theoretical based methods have attracted a great attention in recent years, and gained promising results to deal with multi-label data with high dimensionality. However, most of the existing methods are either directly transformed from heuristic single-label feature selection methods or inefficient in exploiting labeling information. Thus, they may not be able to get an optimal feature selection result shared by multiple labels. In this paper, we propose a general global optimization framework, in which feature relevance, label relevance (i.e., label correlation), and feature redundancy are taken into account, thus facilitating multi-label feature selection. Moreover, the proposed method has an excellent mechanism for utilizing inherent properties of multi-label learning. Specially, we provide a formulation to extend the proposed method with label-specific features. Empirical studies on twenty multi-label data sets reveal the effectiveness and efficiency of the proposed method. Our implementation of the proposed method is available online at: https://jiazhang-ml.pub/GRRO-master.zip.
Jia Zhang 0019, Yidong Lin, Min Jiang 0005, Shaozi Li, Yong Tang 0001, Kay Chen Tan
IJCAI4
2020 Class-Aware Modality Mix and Center-Guided Metric Learning for Visible-Thermal Person Re-Identification
abstract
Visible thermal person re-identification (VT-REID) is an important and challenging task in that 1) weak lighting environments are inevitably encountered in real-world settings and 2) the inter-modality discrepancy is serious. Most existing methods either aim at reducing the cross-modality gap in pixel- and feature-level or optimizing cross-modality network by metric learning techniques. However, few works have jointly considered these two aspects and studied their mutual benefits. In this paper, we design a novel framework to jointly bridge the modality gap in pixel- and feature-level without additional parameters, as well as reduce the inter- and intra-modalities variations by a center-guided metric learning constraint. Specifically, we introduce the Class-aware Modality Mix (CMM) to generate internal information of the two modalities for reducing the modality gap in pixel-level. In addition, we exploit the KL-divergence to further align modality distributions on feature-level. On the other hand, we propose an efficient Center-guided Metric Learning (CML) method for decreasing the discrepancy within the inter- and intra-modalities, by enforcing constraints on class centers and instances. Extensive experiments on two datasets show the mutual advantage of the proposed components and demonstrate the superiority of our method over the state of the art.
Yongguo Ling, Zhun Zhong, Zhiming Luo, Paolo Rota, Shaozi Li, Nicu Sebe
ACM Multimedia5
2020 Joint multilabel classification and feature selection based on deep canonical correlation analysis
abstract
Summary In recent years, multilabel learning has been applied to a lot of application areas and is yet a challenging task. In multilabel learning, an instance often belongs to multiple class labels simultaneously. The labels usually have correlations with others, and mining label correlations is helpful to enhance the multilabel classification performance. Aiming at increasing the accuracy of prediction, Label embedding (LE) is an important technique, and conducive to extracting label information for multilabel learning. In this paper, we present a novel multilabel learning approach via exploiting label correlations, which can be naturally extended to tackle feature selection problem. First, to obtain the discriminative features shared by all labels, the proposed algorithm learns a latent space by employing deep canonical correlation analysis. Then we exploit label correlations by enforcing predictions on similar labels to be similar, thereby improving the prediction performance. Results on several multiple datasets illustrate that the proposed algorithm has the advantages on multilabel classification and feature selection.
Guodong Du 0002, Jia Zhang 0019, Candong Li, Rong Wei, Shaozi Li
Concurr. Comput. Pract. Exp.6
2020 An iterative transfer learning framework for cross-domain tongue segmentation
abstract
Summary Tongue diagnosis is an important clinical examination in Traditional Chinese Medicine. As the first step of the diagnosis, the accuracy of tongue image segmentation directly affects the subsequent diagnosis. Recently, deep learning‐based methods have been applied for tongue image segmentation and achieve promising results. However, these methods usually work well on one dataset and degenerate significantly on different distributed datasets. To deal with this issue, we propose a framework named Iterative cross‐domain tongue segmentation in the study. First, we train a tongue image segmentation U‐Net model on the source dataset. Then, we propose a tongue assessment filter to select satisfying samples based on predictions of the U‐Net model from the target dataset. Following, we fine‐tune the model on the selected samples along with the source domain. Finally, we iterate between the filtering and the fine‐tuning steps until the model is converged. Experimental results on two tongue datasets show that our proposed method can improve the dice score on the target domain from 70.11% to 98.26%, as well as outperform state‐of‐the‐art comparing methods.
Lei Li 0048, Zhiming Luo, Yuanzheng Cai, Candong Li, Shaozi Li
Concurr. Comput. Pract. Exp.6
2020 Research on line overload identification of power system based on improved neural network algorithm
abstract
Summary Due to the continuous appearance of safety fault accidents in the practice process, operation safety has become the central task of various operation and management tasks of the power grid. Therefore, to establish a line overload identification and data control model for the power system, we first defined the vulnerability of complex power systems based on the analysis of each line and node. For finding the optimal parameters of this model, we proposed an improved optimization strategy by combining the genetic algorithm and BP neural network. To verified the effectiveness of our proposed method, we conducted experiments on a simulation on the IEEE 30‐node power system environment. Experimental results demonstrate that the proposed algorithms can establish an optimized overload identification model with better performance. This study can help to conduct reasonable adjustment when overload happens to the power system, and then reduce similar failure as well as enhance the operation safety.
Zhiming Luo, Wangqing Lin, Shaozi Li
Concurr. Comput. Pract. Exp.4
2020 Hand gesture recognition based on attentive feature fusion
abstract
Summary Video‐based hand gesture recognition plays an important role in human‐computer interaction (HCI). Recent advanced methods usually add 3D convolutional neural networks to capture the information from both spatial and temporal dimensions. However, these methods suffer the issue of requiring large‐scale training data and high computational complexity. To address this issue, we proposed an attentive feature fusion framework for efficient hand‐gesture recognition. In our proposed model, we utilize a shallow two‐stream CNNs to capture the low‐level features from the original video frame and its corresponding optical flow. Following, we designed an attentive feature fusion module to selectively combine useful information from the previous two streams based on the attention mechanism. Finally, we obtain a compact embedding of a video by concatenating features from several short segments. To evaluate the effectiveness of our proposed framework, we train and test our method on a large‐scale video‐based hand gesture recognition dataset, Jester. Experimental results demonstrate that our approach obtains very competitive performance on the Jester dataset with a classification accuracy of 95.77%.
Zhiming Luo, Huangbin Wu, Shaozi Li
Concurr. Comput. Pract. Exp.4
2020 A multi-source heterogeneous data analytic method for future price fluctuation prediction
Lei Chai, Hongfeng Xu, Zhiming Luo, Shaozi Li
Neurocomputing4
2020 Stock movement predictive network via incorporative attention mechanisms based on tweet and historical prices
Hongfeng Xu, Lei Chai, Zhiming Luo, Shaozi Li
Neurocomputing4
2020 Towards Chinese clinical named entity recognition by dynamic embedding using domain-specific knowledge
Yuan Li 0024, Guodong Du 0002, Shaozi Li, Lei Ma 0010, Xiongbin Wang
J. Biomed. Informatics4
2020 Joint imbalanced classification and feature selection for hospital readmissions
Guodong Du 0002, Jia Zhang 0019, Zhiming Luo, Fenglong Ma, Lei Ma 0010, Shaozi Li
Knowl. Based Syst.6
2020 Leveraging Virtual and Real Person for Unsupervised Person Re-Identification
abstract
Person re-identification (re-ID) is a challenging instance retrieval problem, especially when identity annotations are not available for training. Although modern deep re-ID approaches have achieved great improvement, it is still difficult to optimize the deep re-ID model and learn discriminative person representation without annotations in training data. To address this challenge, this study considers the problem of unsupervised person re-ID and introduces a novel approach to solve this problem by leveraging virtual and real data. Our approach includes two components: virtual person generation and training of the deep re-ID model. For virtual person generation, we learn a person generation model and a camera style transfer model using unlabeled real data to generate virtual persons with different poses and camera styles. The virtual data is formed as labeled training data, enabling subsequent training deep re-ID model in supervision. For training of the deep re-ID model, we divide it into three steps: 1) pre-training a coarse re-ID model by using virtual data; 2) collaborative filtering based positive pair mining from the real data; and 3) fine-tuning of the coarse re-ID model by leveraging the mined positive pairs and virtual data. The final re-ID model is achieved by iterating between step 2 and step 3 until convergence. Extensive experiments demonstrate the effectiveness of our method. Experimental results on two large-scale datasets, Market-1501 and DukeMTMC-reID, show the advantages of our method over state-of-the-art approaches in unsupervised person re-ID. Our code is now available online1.
Fengxiang Yang, Zhun Zhong, Zhiming Luo, Sheng Lian, Shaozi Li
IEEE Trans. Multim.5
2019 Invariance Matters: Exemplar Memory for Domain Adaptive Person Re-Identification
abstract
This paper considers the domain adaptive person re-identification (re-ID) problem: learning a re-ID model from a labeled source domain and an unlabeled target domain. Conventional methods are mainly to reduce feature distribution gap between the source and target domains. However, these studies largely neglect the intra-domain variations in the target domain, which contain critical factors influencing the testing performance on the target domain. In this work, we comprehensively investigate into the intra-domain variations of the target domain and propose to generalize the re-ID model w.r.t three types of the underlying invariance, i.e., exemplar-invariance, camera-invariance and neighborhood-invariance. To achieve this goal, an exemplar memory is introduced to store features of the target domain and accommodate the three invariance properties. The memory allows us to enforce the invariance constraints over global training batch without significantly increasing computation cost. Experiment demonstrates that the three invariance properties and the proposed memory are indispensable towards an effective domain adaptation system. Results on three re-ID domains show that our domain adaptation accuracy outperforms the state of the art by a large margin. Code is available at: https://github.com/zhunzhong07/ECN.
Zhun Zhong, Liang Zheng 0001, Zhiming Luo, Shaozi Li, Yi Yang 0001
CVPR4
2019 Cover patches: A general feature extraction strategy for spoofing detection
abstract
Summary Face anti‐spoofing has attracted many attentions in security applications, such as mobile payment and entrance guard. Until now, face anti‐spoofing technique is still a challenging task. Mainstream image‐based spoofing algorithms usually use global motion or texture information to distinguish whether an input face is live or fake. However, the performance of these methods are sensitive in light changes, or images acquired from different sensors. The main reason is that spoofed face image always has slight different texture in local areas, such as landmark or salient region of face. To this end, this paper proposes a novel multi‐patches feature extraction strategy to detect spoofing. First, a set of patches with specific combination scheme is selected to cover the face image. Second, features such as hand‐crafted Gray Level Co‐occurrence Matrix (GLCM), Local Binary Patterns (LBP), or deep features are extracted from these patches. Third, all features are combined as the global descriptor of the face image, then fed into an SVM classifier to verify the anti‐spoofing detection. Experimental results show that the proposed strategy can effectively enhance the performance, concerning with the accuracy of spoofed face detection in four widely used anti‐spoofing databases.
Guo-Rong Cai, Songzhi Su, Chengcai Leng, Jipeng Wu, Yun-Dong Wu, Shaozi Li
Concurr. Comput. Pract. Exp.6
2019 Multi-label feature selection with application to TCM state identification
abstract
Summary The goal of TCM state identification is to identify the patient's syndromes and locations and natures of diseases according to symptoms. Generally, symptoms of a patient are associated with several syndromes and multiple locations and natures of diseases; hence, the TCM state identification is a typical multi‐label problem. In this paper, a new method is proposed to predict syndromes and locations and natures of diseases according to the diagnostic information of TCM. In detail, the correlation between features and the correlation between class labels are combined into a new uniform feature space. After that, the MDMR algorithm is used to select the most discriminatory features from the new uniform feature space, which is helpful to reduce the data dimensionality. Lastly, a KNN‐like algorithm is modified to calculate the label similarity of test data, and the finite set of labels of test data is predicted by ML‐KNN. In this paper, the test data is collected by Fujian University of Traditional Chinese Medicine according to the theory of TCM and medical ethics. The experiments show that the performance of the proposed method is superior to some other popular methods and is helpful in the identification of health state in TCM.
Jia Zhang 0019, Candong Li, Changen Zhou, Shaozi Li
Concurr. Comput. Pract. Exp.5
2019 Automated classification of Wuyi rock tealeaves based on support vector machine
abstract
Summary This paper describes a new automated classification method for Wuyi rock tealeaves based on the best penalty parameter selection for the support vector machine with RBF (Radial Basis Function) kernel. A total of 3590 fresh tealeaf images of the representative Rou Gui and Shui Hsien varieties of Wuyi rock tea are collected in their natural habitat. Fourteen image features are extracted in terms of the leaf shape and texture. The automatic selection method is used to find the optimum RBF kernel parameter sigma, which is then applied to design an automatic parameter selection method to screen the best penalty parameter C for the classification of Wuyi rock tealeaves. In this study, the SVM classifier is used for the automated classification and recognition of the 14 image features. The contribution of the various features to the recognition rate of fresh tealeaves is evaluated to identify the key features for the classification and recognition of fresh Wuyi rock tealeaf images. The experimental results show that the proposed method improves the recognition rate of fresh tealeaves to 91.00%.
Li-Hui Lin, Cheng-Hsuan Li, Shaozi Li
Concurr. Comput. Pract. Exp.4
2019 Chinese microblog rumor detection based on deep sequence context
abstract
Summary Rumor is one of the main problems in social media, which often shows deeply and rapidly undesirable affection on the society. Although many rumor detection models consider content features and social features, all of them are based on the word independence assumption, which lacks the sequence context. Thus, if we use some words that often appear in rumors, our posts will be recognized as a rumor. To solve this problem, we propose a deep sequence context model (DSCM) for Chinese microblog rumor detection. This model considers two important factors of rumors: falsity and influence. Firstly, to learn falsity, we abolish the word independence assumption and use long short‐term memory (LSTM) units to capture bi‐direction sequence context information in content. Secondly, to learn influence, we combine the deep sequence context information with social features to learn the connection between content and social features. In our experiment, our results show that our approach outperforms several state‐of‐the‐art machine learning approaches, including term frequency and inverse document frequency (TFIDF), LSTM, and gated recurrent unit (GRU) in rumor detection.
Dazhen Lin, Donglin Cao, Shaozi Li
Concurr. Comput. Pract. Exp.4
2019 Mutual information based multi-label feature selection via constrained convex optimization
Zhenqiang Sun, Jia Zhang 0019, Candong Li, Changen Zhou, Jiliang Xin, Shaozi Li
Neurocomputing7
2019 Manifold regularized discriminative feature selection for multi-label learning
Jia Zhang 0019, Zhiming Luo, Candong Li, Changen Zhou, Shaozi Li
Pattern Recognit.5
2019 CamStyle: A Novel Data Augmentation Method for Person Re-Identification
abstract
Person re-identification (re-ID) is a cross-camera retrieval task that suffers from image style variations caused by different cameras. The art implicitly addresses this problem by learning a camera-invariant descriptor subspace. In this paper, we explicitly consider this challenge by introducing camera style (CamStyle). CamStyle can serve as a data augmentation approach that reduces the risk of deep network overfitting and that smooths the CamStyle disparities. Specifically, with a style transfer model, labeled training images can be style transferred to each camera, and along with the original training samples, form the augmented training set. This method, while increasing data diversity against overfitting, also incurs a considerable level of noise. In the effort to alleviate the impact of noise, the label smooth regularization (LSR) is adopted. The vanilla version of our method (without LSR) performs reasonably well on few camera systems in which overfitting often occurs. With LSR, we demonstrate consistent improvement in all systems regardless of the extent of overfitting. We also report competitive accuracy compared with the state of the art on Market-1501 and DukeMTMC-re-ID. Importantly, CamStyle can be employed to the challenging problems of one view learning and unsupervised domain adaptation (UDA) in person re-identification (re-ID), both of which have critical research and application significance. The former only has labeled data in one camera view and the latter only has labeled data in the source domain. Experimental results show that CamStyle significantly improves the performance of the baseline in the two problems. Specially, for UDA, CamStyle achieves state-of-the-art accuracy based on a baseline deep re-ID model on Market-1501 and DukeMTMC-reID. Our code is available at: https://github.com/zhunzhong07/CamStyle .
Zhun Zhong, Liang Zheng 0001, Zhedong Zheng, Shaozi Li, Yi Yang 0001
IEEE Trans. Image Process.4
2018 Camera Style Adaptation for Person Re-Identification
abstract
Being a cross-camera retrieval task, person re-identification suffers from image style variations caused by different cameras. The art implicitly addresses this problem by learning a camera-invariant descriptor subspace. In this paper, we explicitly consider this challenge by introducing camera style (CamStyle) adaptation. CamStyle can serve as a data augmentation approach that smooths the camera style disparities. Specifically, with CycleGAN, labeled training images can be style-transferred to each camera, and, along with the original training samples, form the augmented training set. This method, while increasing data diversity against over-fitting, also incurs a considerable level of noise. In the effort to alleviate the impact of noise, the label smooth regularization (LSR) is adopted. The vanilla version of our method (without LSR) performs reasonably well on few-camera systems in which over-fitting often occurs. With LSR, we demonstrate consistent improvement in all systems regardless of the extent of over-fitting. We also report competitive accuracy compared with the state of the art. Code is available at: https://github.com/zhunzhong07/CamStyle.
Zhun Zhong, Liang Zheng 0001, Zhedong Zheng, Shaozi Li, Yi Yang 0001
CVPR4
2018 Generalizing a Person Retrieval Model Hetero- and Homogeneously
Zhun Zhong, Liang Zheng 0001, Shaozi Li, Yi Yang 0001
ECCV (13)3
2018 Anchor Free Network for Multi-Scale Face Detection
abstract
Anchor-based deep methods are the most widely used methods for face detection and have reached the state-of-the-art result. Compared with anchor-based methods that estimates the bounding-box rely on some pre-defined anchor boxes, anchor-free methods perform the localization by predicting the offsets of a pixel inside a face to its outside boundaries whose accuracies are much more precise. However, anchor-free methods suffer the drawback of low recall-rate mainly because 1) only using single scale features lead to miss detection of small faces, 2) the highly intra-class imbalance problem among different size faces. In this paper, to address these problems, we propose a unified anchor-free network for detecting multi-scale faces by leveraging the local and global contextual information of multi-layer features. We also utilize a scale aware sampling strategy to mitigate the intra-class imbalance issue which can adaptivity select the positive samples. Furthermore, a revised focal loss function is adopted to deal with the foreground/background imbalance issue. Experimental results on two benchmark datasets demonstrate the effective of our proposed method.
Chengji Wang, Zhiming Luo, Sheng Lian, Shaozi Li
ICPR4
2018 Combining 2D and 3D features to improve road detection based on stereo cameras
abstract
Road detection is a fundamental component of autonomous driving systems since it provides validspace and candidate regions of objects for driving decision. The core of roaddetection methods is extracting effective and discriminative features. Sincetwo‐dimensional (2D) and 3D features are complementary, the authors propose arobust multi‐feature combination and optimisation framework for stereo imagepairs, called Feature++. First, several 2D and 3D features such as Gabor andplane are, respectively, extracted after the generation of 2D super‐pixel and a3D depth image from stereo matching. Second, the combined features are fed intoa three‐layer shallow neural network classifier to decide whether a super‐pixelis road region or not. Finally, the classified results are further refined usingfully connected conditional random field (CRF), taking the content informationinto consideration. We extensively evaluate the performance of four 2D features,four 3D features, and their combinations. Experiments conducted on the KITTIROAD benchmark show that (i) the combinations of 2D and 3D features greatlyimprove the road detection performance and (ii) using CRF as a refinement stepis necessary. Overall, their proposed ‘Feature + +’ method outperforms mostmanually designed features, and is comparable with state‐of‐the‐art methods thatare based on deep learning methods.
Guo-Rong Cai, Songzhi Su, Wenli He, Yun-Dong Wu, Shaozi Li
IET Comput. Vis.5
2018 Discriminative parts learning for 3D human action recognition
Min Huang 0004, Guo-Rong Cai, Hongbo Zhang 0002, Sheng Yu 0007, Dong-Ying Gong, Donglin Cao, Shaozi Li, Songzhi Su
Neurocomputing7
2018 Attention guided U-Net for accurate iris segmentation
Sheng Lian, Zhiming Luo, Zhun Zhong, Songzhi Su, Shaozi Li
J. Vis. Commun. Image Represent.6
2018 Electroencephalogram-based brain-computer interface for the Chinese spelling system: a survey
abstract
Electroencephalogram (EEG) based brain-computer interfaces allow users to communicate with the external environment by means of their EEG signals, without relying on the brain’s usual output pathways such as muscles. A popular application for EEGs is the EEG-based speller, which translates EEG signals into intentions to spell particular words, thus benefiting those suffering from severe disabilities, such as amyotrophic lateral sclerosis. Although the EEG-based English speller (EEGES) has been widely studied in recent years, few studies have focused on the EEG-based Chinese speller (EEGCS). The EEGCS is more difficult to develop than the EEGES, because the English alphabet contains only 26 letters. By contrast, Chinese contains more than 11 000 logographic characters. The goal of this paper is to survey the literature on EEGCS systems. First, the taxonomy of current EEGCS systems is discussed to get the gist of the paper. Then, a common framework unifying the current EEGCS and EEGES systems is proposed, in which the concept of EEG-based choice acts as a core component. In addition, a variety of current EEGCS systems are investigated and discussed to highlight the advances, current problems, and future directions for EEGCS.
Minghui Shi, Changle Zhou, Jun Xie 0002, Shaozi Li, Qingyang Hong, Min Jiang 0005, Fei Chao 0001, Weifeng Ren, Xiangqian Liu, Dajun Zhou
Frontiers Inf. Technol. Electron. Eng.4
2018 Improving deep ensemble vehicle classification by using selected adversarial samples
Wei Liu 0052, Zhiming Luo, Shaozi Li
Knowl. Based Syst.3
2018 Multi-label learning with label-specific features by resolving label correlations
Jia Zhang 0019, Candong Li, Donglin Cao, Yaojin Lin, Songzhi Su, Shaozi Li
Knowl. Based Syst.7
2018 Traffic Analytics With Low-Frame-Rate Videos
abstract
In this paper, we investigate the possibility of monitoring highway traffic based on videos whose frame rate is too low to accurately estimate motion features. The goal of the proposed method is to recognize traffic conditions instead of measuring them, as is usually the case. The main advantage of our approach comes from its ability to process low-frame-rate videos for which motion features cannot be estimated. Our method takes advantage of the highly redundant nature of traffic scenes that are pictured from a top-down perspective showing vehicles on a predominant asphalted road surrounded by background objects. Due to the limited variety of objects pictured in traffic scenes, our method gets to learn features that are specific to such images. With these features, our method is able to segment traffic images, classify traffic scenes, and estimate traffic density without requiring motion features. Different convolutional neural network models are proposed to segment traffic images in three different classes (Road, Car, and Background), classify traffic images into different categories (Empty, Fluid, Heavy, and Jam), and predict traffic density. We also propose a procedure to perform transfer learning of any of these models to new traffic scenes.
Zhiming Luo, Pierre-Marc Jodoin, Songzhi Su, Shaozi Li, Hugo Larochelle
IEEE Trans. Circuits Syst. Video Technol.4
2018 MIO-TCD: A New Benchmark Dataset for Vehicle Classification and Localization
abstract
The ability to train on a large dataset of labeled samples is critical to the success of deep learning in many domains. In this paper, we focus on motor vehicle classification and localization from a single video frame and introduce the "MIOvision Traffic Camera Dataset" (MIO-TCD) in this context. MIO-TCD is the largest dataset for motorized traffic analysis to date. It includes 11 traffic object classes such as cars, trucks, buses, motorcycles, bicycles, pedestrians. It contains 786,702 annotated images acquired at different times of the day and different periods of the year by hundreds of traffic surveillance cameras deployed across Canada and the United States. The dataset consists of two parts: a "localization dataset", containing 137,743 full video frames with bounding boxes around traffic objects, and a "classification dataset", containing 648,959 crops of traffic objects from the 11 classes. We also report results from the 2017 CVPR MIO-TCD Challenge, that leveraged this dataset, and compare them with results for state-of-the-art deep learning architectures. These results demonstrate the viability of deep learning methods for vehicle localization and classification from a single video frame in real-life traffic scenarios. The topperforming methods achieve both accuracy and Kappa score above 96% on the classification dataset and mean-average precision of 77% on the localization dataset. We also identify scenarios in which state-of-the-art methods still fail and we suggest avenues to address these challenges. Both the dataset and detailed results are publicly available on-line [1].
Zhiming Luo, Frederic Branchaud-Charron, Carl Lemaire, Janusz Konrad, Shaozi Li, Akshaya Mishra, Andrew Achkar, Justin A. Eichel, Pierre-Marc Jodoin
IEEE Trans. Image Process.5
2018 Multifeature Selection for 3D Human Action Recognition
abstract
In mainstream approaches for 3D human action recognition, depth and skeleton features are combined to improve recognition accuracy. However, this strategy results in high feature dimensions and low discrimination due to redundant feature vectors. To solve this drawback, a multi-feature selection approach for 3D human action recognition is proposed in this paper. First, three novel single-modal features are proposed to describe depth appearance, depth motion, and skeleton motion. Second, a classification entropy of random forest is used to evaluate the discrimination of the depth appearance based features. Finally, one of the three features is selected to recognize the sample according to the discrimination evaluation. Experimental results show that the proposed multi-feature selection approach significantly outperforms other approaches based on single-modal feature and feature fusion.
Min Huang 0004, Songzhi Su, Hongbo Zhang 0002, Guo-Rong Cai, Dong-Ying Gong, Donglin Cao, Shaozi Li
ACM Trans. Multim. Comput. Commun. Appl.7
2017 Non-local Deep Features for Salient Object Detection
abstract
Saliency detection aims to highlight the most relevant objects in an image. Methods using conventional models struggle whenever salient objects are pictured on top of a cluttered background while deep neural nets suffer from excess complexity and slow evaluation speeds. In this paper, we propose a simplified convolutional neural network which combines local and global information through a multi-resolution 4×5 grid structure. Instead of enforcing spacial coherence with a CRF or superpixels as is usually the case, we implemented a loss function inspired by the Mumford-Shah functional which penalizes errors on the boundary. We trained our model on the MSRA-B dataset, and tested it on six different saliency benchmark datasets. Results show that our method is on par with the state-of-the-art while reducing computation time by a factor of 18 to 100 times, enabling near real-time, high performance saliency detection.
Zhiming Luo, Akshaya Kumar Mishra, Andrew Achkar, Justin A. Eichel, Shaozi Li, Pierre-Marc Jodoin
CVPR5
2017 Re-ranking Person Re-identification with k-Reciprocal Encoding
abstract
When considering person re-identification (re-ID) as a retrieval process, re-ranking is a critical step to improve its accuracy. Yet in the re-ID community, limited effort has been devoted to re-ranking, especially those fully automatic, unsupervised solutions. In this paper, we propose a k-reciprocal encoding method to re-rank the re-ID results. Our hypothesis is that if a gallery image is similar to the probe in the k-reciprocal nearest neighbors, it is more likely to be a true match. Specifically, given an image, a k-reciprocal feature is calculated by encoding its k-reciprocal nearest neighbors into a single vector, which is used for re-ranking under the Jaccard distance. The final distance is computed as the combination of the original distance and the Jaccard distance. Our re-ranking method does not require any human interaction or any labeled data, so it is applicable to large-scale datasets. Experiments on the large-scale Market-1501, CUHK03, MARS, and PRW datasets confirm the effectiveness of our method.
Zhun Zhong, Liang Zheng 0001, Donglin Cao, Shaozi Li
CVPR4
2017 Computational drug repositioning using collaborative filtering via multi-source fusion
Jia Zhang 0019, Candong Li, Yaojin Lin, Youwei Shao, Shaozi Li
Expert Syst. Appl.5
2017 Meta-action descriptor for action recognition in RGBD video
abstract
Action recognition is one of the hottest research topics in computer vision. Recent methods represent actions based on global or local video features. These approaches, however, lack semantic structure and may not provide a deep insight into the essence of an action. In this work, the authors argue that semantic clues, such as joint positions and part‐level motion clustering, help verify actions. To this end, a meta‐action descriptor for action recognition in RGBD video is proposed in this study. Specifically, two discrimination‐based strategies – dynamic and discriminative part clustering – are introduced to improve accuracy. Experiments conducted on the MSR Action 3D dataset show that the proposed method significantly outperforms the methods without joint position semantic.
Min Huang 0004, Songzhi Su, Guo-Rong Cai, Hongbo Zhang 0002, Donglin Cao, Shaozi Li
IET Comput. Vis.6
2017 Fully convolutional networks for action recognition
abstract
Human action recognition is an important and challenging topic in computer vision. Recently, convolutional neural networks (CNNs) have established impressive results for many image recognition tasks. The CNNs usually contain million parameters which prone to overfit when training on small datasets. Therefore, the CNNs do not produce superior performance over traditional methods for action recognition. In this study, the authors design a novel two‐stream fully convolutional networks architecture for action recognition which can significantly reduce parameters while keeping performance. To utilise the advantage of spatial‐temporal features, a linear weighted fusion method is used to fuse two‐stream networks’ feature maps and a video pooling method is adopted to construct the video‐level features. At the meantime, the authors also demonstrate that the improved dense trajectories has significant impact for action recognition. The authors’ method can achieve the state‐of‐the‐art performance on two challenging datasets UCF101 (93.0%) and HMDB51 (70.2%).
Sheng Yu 0007, Shaozi Li
IET Comput. Vis.4
2017 Class-specific object proposals re-ranking for object detection in automatic driving
Zhun Zhong, Mingyi Lei, Donglin Cao, Jianping Fan 0001, Shaozi Li
Neurocomputing5
2017 A novel recurrent hybrid network for feature fusion in action recognition
Sheng Yu 0007, Zhiming Luo, Min Huang 0004, Shaozi Li
J. Vis. Commun. Image Represent.6
2017 Learning rich features from objectness estimation for human lying-pose detection
Daoxun Xia, Songzhi Su, Li-Chuan Geng, Guoxi Wu, Shaozi Li
Multim. Syst.5
2017 Stratified pooling based deep convolutional neural networks for human action recognition
Sheng Yu 0007, Songzhi Su, Guo-Rong Cai, Shaozi Li
Multim. Tools Appl.5
2017 Detecting ground control points via convolutional neural network for stereo matching
Zhun Zhong, Songzhi Su, Donglin Cao, Shaozi Li, Zhihan Lyu
Multim. Tools Appl.4
2016 Dynamic programming based optimized product quantization for approximate nearest neighbor search
Yuanzheng Cai, Rongrong Ji, Shaozi Li
Neurocomputing3
2016 CBDF: Compressed Binary Discriminative Feature
Li-Chuan Geng, Pierre-Marc Jodoin, Songzhi Su, Shaozi Li
Neurocomputing4
2016 Detection based object labeling of 3D point cloud for indoor scenes
Wei Liu 0005, Shaozi Li, Donglin Cao, Songzhi Su, Rongrong Ji
Neurocomputing2
2016 Local consistent hierarchical Hough Match for image re-ranking
Yuanzheng Cai, Shaozi Li, Rongrong Ji
J. Vis. Commun. Image Represent.2
2016 Discriminative local collaborative representation for online object tracking
Si Chen 0002, Shaozi Li, Rongrong Ji, Yan Yan 0001, Shunzhi Zhu
Knowl. Based Syst.2
2016 A cross-media public sentiment analysis system for microblog
Donglin Cao, Rongrong Ji, Dazhen Lin, Shaozi Li
Multim. Syst.4
2016 Decomposed human localization from social photo album
Shaozi Li, Songzhi Su, Bing Shuai, Rongrong Ji
Multim. Syst.1
2016 Spectral-spatial co-clustering of hyperspectral image data based on bipartite graph
Wei Liu 0005, Shaozi Li, Xianming Lin, Yun-Dong Wu, Rongrong Ji
Multim. Syst.2
2016 Fast verification via statistical geometric for mobile visual search
Shaozi Li, Xianming Lin, Songzhi Su, Rongrong Ji
Multim. Syst.2
2016 Visual sentiment topic model based microblog image sentiment analysis
Donglin Cao, Rongrong Ji, Dazhen Lin, Shaozi Li
Multim. Tools Appl.4
2016 Multi-view fall detection based on spatio-temporal interest points
Songzhi Su, Sin-Sian Wu, Shu-Yuan Chen, Der-Jyh Duh, Shaozi Li
Multim. Tools Appl.5
2015 Towards 3D object detection with bimodal deep Boltzmann machines over RGBD imagery
abstract
Nowadays, detecting objects in 3D scenes like point clouds has become an emerging challenge with various applications. However, it retains as an open problem due to the deficiency of labeling 3D training data. To deploy an accurate detection algorithm typically resorts to investigating both RGB and depth modalities, which have distinct statistics while correlated with each other. Previous research mainly focus on detecting objects using only one modality, which ignores exploiting the cross-modality cues. In this work, we propose a cross-modality deep learning framework based on deep Boltzmann Machines for 3D Scenes object detection. In particular, we demonstrate that by learning cross-modality feature from RGBD data, it is possible to capture their joint information to reinforce detector trainings in individual modalities. In particular, we slide a 3D detection window in the 3D point cloud to match the exemplar shape, which the lack of training data in 3D domain is conquered via (1) We collect 3D CAD models and 2D positive samples from Internet. (2) adopt pretrained R-CNNs [2] to extract raw feature from both RGB and Depth domains. Experiments on RMRC dataset demonstrate that the bimodal based deep feature learning framework helps 3D scene object detection.
Wei Liu 0005, Rongrong Ji, Shaozi Li
CVPR3
2015 Sentiment analysis of Chinese micro-blog based on multi-modal correlation model
abstract
Text, emoticons and images, various modalities have been used to express users' feelings on social media, which significantly challenges traditional text-based sentiment analysis approaches. In this paper, we propose a Multi-modal Correlation Model (MCM) for multi-modal sentiment analysis. Compared with other multi-modal methods, MCM models hierarchical correlations among modalities, as well as between modalities and sentiments. Specifically, a probabilistic graphical model (PGM) is subsequently built upon the proposed MCM model, which considers the hierarchical correlations and preserves the classification ability of each modality. In order to compute the posterior probabilities of sentiments in PGM, we optimize the model by Maximum Likelihood Estimation. Experimental results demonstrate: 1) the hierarchical correlations among different modalities and sentiment; 2) the importance of hierarchical correlations to sentiment analysis.
Donglin Cao, Shaozi Li, Rongrong Ji
ICIP3
2015 Traffic analysis without motion features
abstract
In this paper, we investigate the possibility of monitoring traffic without using any motion features. The goal of our system is to process videos with ultra-low frame rate, i.e. videos for which reliable motion features cannot be computed. In this work, we investigate how 2D spatial features combined with a machine learning method can assess traffic conditions such as fluid traffic, dense traffic, and traffic jam. The underlying hypothesis that we ought to validate is that traffic images are heavily characterized by their 2D spatial textures. In that perspective, we tested different 2D texture features and machine learning methods to see how accurate such an approach can be. We also performed a regression on the image descriptor in order to estimate traffic density. Experimental results obtained on the UCSD traffic dataset reveal that our approach generalizes well to various weather and lighting conditions. It even outperforms state-of-the-art traffic analysis methods relying on spatio-temporal features.
Zhiming Luo, Pierre-Marc Jodoin, Shaozi Li, Songzhi Su
ICIP3
2015 Feature learning based on SAE-PCA network for human gesture recognition in RGBD images
Shaozi Li, Wei Wu 0072, Songzhi Su, Rongrong Ji
Neurocomputing1
2015 Sparse auto-encoder based feature learning for human body detection in depth image
Songzhi Su, Zhi-Hui Liu, Suping Xu, Shaozi Li, Rongrong Ji
Signal Process.4
2014 Lying-pose detection with training dataset expansion
abstract
We propose a rotation and scale invariant method to locate people lying on the ground. Unlike conventional human-shape detection methods which assume that all human shapes are in upright position, a person lying on the ground can have arbitrary orientation and pose. Accounting for every possible body configuration would thus require a huge training dataset that would be challenging to gather. In this paper, we propose a method which increases the size of a small training dataset and allows to detect multiple body poses. To do so, our method increases the size of the dataset with a geometric distortion method followed by a rejection sampling method. Then, it automatically identifies K body configurations in the training set, realign it in upright position and trains K SVM classifiers, one for each body configuration. Lying pose detection is then performed by considering a max pooling strategy across all K SVM classifiers.
Daoxun Xia, Songzhi Su, Shaozi Li, Pierre-Marc Jodoin
ICIP3
2014 Pursuing Detector Efficiency for Simple Scene Pedestrian Detection
De-Dong Yuan, Songzhi Su, Shaozi Li, Rongrong Ji
MMM (2)4
2014 Kinship classification based on discriminative facial patches
abstract
Recently there has been a large explosive growth of image data on social networks and how to use computer vision and machine learning technology to verify people relationships on these huge amount of human-centered image data remains a challenging issue. Remarkably, there have been few research attempts to analyze the possible human relationships on images, especially kin relationships. In this paper, we tackle a challenging and relatively new issue in kinship classification: determining the family that a query face image belongs to. To address this challenge, we propose a kinship classification method in three steps: (l)Discriminative patches are detected automatically in the facial landmark regions. (2) Appearance features, Histogram of Gradient (HOG), Scale-Invariant Feature Transform (SIFT) and Four-Patch Local Binary Pattern (FPLBP) are extracted from these patches respectively, and then we concatenate the features to create a high-dimensional feature vector. (3) Linear Support Vector Machine (SVM) with polynomial kernel is adopted to accomplish kinship classification task. Experimental evaluation results on Cornell Family 101 dataset demonstrate that our proposed method significantly outperforms the state-of-the-art kinship classification approaches.
Songzhi Su, Shaozi Li
VCIP4
2014 Online MIL tracking with instance-level semi-supervised learning
Si Chen 0002, Shaozi Li, Songzhi Su, Qi Tian 0001, Rongrong Ji
Neurocomputing2
2014 Perspective-Invariant Image Matching Framework with Binary Feature Descriptor and APSO
abstract
A novel perspective invariant image matching framework is proposed in this paper, noted as Perspective-Invariant Binary Robust Independent Elementary Features (PBRIEF). First, we use the homographic transformation to simulate the distortion between two corresponding patches around the feature points. Then, binary descriptors are constructed by comparing the intensity of sample points surrounding the feature location. We transform the location of the sample points with simulated homographic matrices. This operation is to ensure that the intensities which we compared are the realistic corresponding pixels between two image patches. Since the exact perspective transform matrix is unknown, an Adaptive Particle Swarm Optimization (APSO) algorithm-based iterative procedure is proposed to estimate the real transformation angles. Experimental results obtained on five different datasets show that PBRIEF outperforms significantly the existing methods on images with large viewpoint difference. Moreover, the efficiency of our framework is also improved comparing with Affine-Scale Invariant Feature Transform (ASIFT).
Li-Chuan Geng, Songzhi Su, Donglin Cao, Shaozi Li
Int. J. Pattern Recognit. Artif. Intell.4
2014 Online semi-supervised compressive coding for robust visual tracking
Si Chen 0002, Shaozi Li, Songzhi Su, Donglin Cao, Rongrong Ji
J. Vis. Commun. Image Represent.2
2014 Logo detection with extendibility and discrimination
Kuo-Wei Li, Shu-Yuan Chen, Songzhi Su, Der-Jyh Duh, Hongbo Zhang 0002, Shaozi Li
Multim. Tools Appl.6
2014 Adaptive photograph retrieval method
Hongbo Zhang 0002, Shang-An Li, Shu-Yuan Chen, Songzhi Su, Der-Jyh Duh, Shaozi Li
Multim. Tools Appl.6
2013 Saliency detection by adaptive clustering
abstract
Saliency detection plays an important role in image segmentation, content-aware resizing and object recognition. Most approaches obtain promising performance recently, which is useful for the postprocessing. We propose a clustering-based method to detect refined regions with comparative performance. For coarse-grained classification with unknown clusters number, an adaptive algorithm called f-means is developed in this paper. Pixels are clustered by f-means based on color and spatial features, and then the centroids are used to compute their saliency values. Experiments show that our algorithm generates more fine maps, which outperform the state-of-the-art approaches on MSRA dataset. Relying on the saliency map, we also get superior results in foreground extracting, image resizing and thumbnails generation.
Hai Cao, Shaozi Li, Songzhi Su, Rongrong Ji
VCIP2
2013 A new camera self-calibration method based on CSA
abstract
A large number of computer vision applications rely on camera calibration. Camera self-calibration which only depends on the relationship between corresponding points of a pair of images draws much attention for its simplicity. Almost all the camera self-calibration methods rely on the solution of Kruppa equations which are difficult to be directly solved. The state-of-the-art self-calibration algorithms usually convert the solution of these equations to non-linear optimization problem, traditional optimization methods usually have the drawback of convergent to local extreme. Artificial immune system (AIS) has the ability to fast convergent to global extreme. To address this problem, we proposed an artificial immune system based method which can fast convergent to the global optimization solutions. We demonstrate the performance of the proposed method with synthetic and real data.
Li-Chuan Geng, Shaozi Li, Songzhi Su, Donglin Cao, Rongrong Ji
VCIP2
2013 Decomposed human localization in personal photo albums
abstract
Recent years have seen tremendous progress in human detection, whereas only upright poses are usually considered. In this paper, we relax this constraint to localizing highly deformable persons, as commonly exhibited in personal photo albums. Human localization based on arbitrary pose is extremely challenging, due to the large pose variances, disabling the traditional part based template detectors. To tackle this issue, we propose a decomposition-based human localization model dealing with this issue in three-step: a stable upper-body is firstly detected, then a set of bigger bounding boxes are extended, from which the most appropriate instance is distinguished by a discriminative Whole Person Model. The experiment results demonstrated that our decomposition-based model worked very well at localizing deformable persons, which boosted the average precision by 10% compared to state-of-the-art person detectors. On the other hand, Similar Pose Feature(SPF) provides the feasibility of projecting persons with similar poses into same clusters, facilitating a novel pose-based photo album browsing functionality.
Bing Shuai, Songzhi Su, Shaozi Li, Rongrong Ji
VCIP3
2013 Seeing actions through scene context
abstract
Recognizing human actions is not alone, as hinted by the scene herein. In this paper, we investigate the possibility to boost the action recognition performance by exploiting their scene context associated. To this end, we model the scene as a mid-level “hidden layer” to bridge action descriptors and action categories. This is achieved via a scene topic model, in which hybrid visual descriptors including spatiotemporal action features and scene descriptors are first extracted from the video sequence. Then, we learn a joint probability distribution between scene and action by a Naive Bayesian N-earest Neighbor algorithm, which is adopted to jointly infer the action categories online by combining off-the-shelf action recognition algorithms. We demonstrate our merits by comparing to state-of-the-arts in several action recognition benchmarks.
Hongbo Zhang 0002, Songzhi Su, Shaozi Li, Duansheng Chen, Bineng Zhong 0001, Rongrong Ji
VCIP3
2013 Perspective-SIFT: An efficient tool for low-altitude remote sensing image registration
Guo-Rong Cai, Pierre-Marc Jodoin, Shaozi Li, Yun-Dong Wu, Songzhi Su, Zhenkun Huang
Signal Process.3
2012 A two-level model for automatic image annotation
Xiao Ke, Shaozi Li, Donglin Cao
Multim. Tools Appl.2
2011 Sensitivity-based data selection for predicting individual's sub-health on TCM doctors' diagnosis
abstract
In this paper we propose an approach of predicting individual's sub-health based on the principle of TCM (Traditional Chinese medicine) as a preventive medicine. The object's vision features like features of tongue, eye and face are extracted for modeling a process of TCM doctor's diagnosis. As a consequence of the diversity and uncertainty of TCM doctors' diagnosis, the sensitivity is defined as a criterion to select the appropriate features from the derived features, and integrate the diagnosis data given by multiple doctors as training data for constructing the sub-health inference model. The experiment results show that the sensitivity-based feature selection and diagnosis data integration improve the model's inference performance on the accuracy, correlation and residual variance.
Yi Wang 0046, Ying Dai 0001, Feng Guo 0005, Shaozi Li
SMC4
2011 Particle swarm optimizer for variable weighting in clustering high-dimensional data
Yanping Lv, Shengrui Wang, Shaozi Li, Changle Zhou
Mach. Learn.3
2010 Inferring individuals' sub-health and their TCM syndrome based on the diagnosis of TCM doctors
abstract
This paper represents an approach automatically inferring users' sub-healthy state and the corresponding Traditional Chinese Medicine (TCM) syndrome based on TCM doctors' knowledge, so as to construct a system which helps users recuperating from the sub-health by means of showing them the relevant regimen. The system's framework involves utilizing objects' facial features, feelings, and physical status to infer their sub-healthy state and TCM syndrome, and providing the relevant mental care plan and diet plan according to their TCM syndrome. This paper also introduces the data collection for training the inference model, and describes the algorithm of association rule mining while the individuality of TCM doctor's criteria in diagnosing is considered. Experiment results demonstrate the effectiveness and efficiency of the proposed method, and show the potential important role of the TCM in diagnosing people's sub-healthy state, and helping them to recuperate from it.
Feng Guo 0005, Ying Dai 0001, Shaozi Li, Kenzo Ito
SMC3
2010 Making intelligent business decisions by mining the implicit relation from bloggers' posts
Dazhen Lin, Shaozi Li, Donglin Cao
Soft Comput.2
2009 Particle swarm optimizer for variable weighting in clustering high-dimensional data
abstract
This paper proposes a particle swarm optimizer to solve the variable weighting problem in subspace clustering of high-dimensional data. Many subspace clustering algorithms fail to yield good cluster quality because they do not employ an efficient search strategy. In this paper, we are interested in soft subspace clustering and design a suitable weighting k-means objective function, on which a change of variable weights is exponentially reflected. We transform the original constrained variable weighting problem into a problem with bound constraints using a potential solution coding method and we develop a particle swarm optimizer to minimize the objective function in order to obtain global optima to the variable weighting problem in clustering. Our experimental results on synthetic datasets show that the proposed algorithm greatly improves cluster quality. In addition, the result of the new algorithm is much less dependent on the initial cluster centroids.
Yanping Lv, Shengrui Wang, Shaozi Li, Changle Zhou
SIS3
2009 Text Clustering via Particle Swarm Optimization
abstract
This paper presents an approach which extends a particle swarm optimizer for variable weighting (PSOVW) to handle the problem of text clustering, called Text Clustering via Particle Swarm Optimization (TCPSO). PSOVW has been exploited for evolving optimal feature weights for clusters and has demonstrated to improve the clustering quality of high-dimensional data. However, when applying it for text clustering, there exist some modifications such as the similarity measure, parameter selection and the criterion function. Our experimental results on both four structured text datasets built from 20 newsgroups as well as four large-scale text datasets selected from CLUTO show that the proposed algorithm is able to greatly improve the quality of text clustering compared to four typical clustering algorithms and one competitive subspace clustering method.
Yanping Lv, Shengrui Wang, Shaozi Li, Changle Zhou
SIS3
2008 A tuplespace-based coordination architecture for service composition in pervasive computing environments
abstract
With increasing number and wireless connected computing elements, pervasive computing environments become more and more complex. This environment need an improved coordination mechanisms and composted services to provide the users with a sufficient level of quality. This paper introduces a tuplesapce-based coordinate service composition architecture that addresses these two requirements. Our proposal supports a high level coordination model that is based on the use of tuplespace to achieve service composition among the nodes. We use DAG-based (directed acyclic graph) algorithm to compose services efficiently. In addition, a case study is implemented and some future work is described.
Qingcong Lv, Qiying Cao, Shaozi Li
CSCWD4
2008 A Novel Language Model Based on Cognition Attention Attenuation in Web Retrieval
abstract
Language model is widely used in many retrieval systems. Its document representation is based on the bag of words assumption. Hence, each term in document is treated as an equal object and only the term frequency is considered as the evidence of the importance of term. In this paper, we study the problem of cognition attention attenuation in processing documents and present a cognition attention attenuation based language model. This model estimates the document model by attenuation process of term in document. Compared with the classical language model, the advantage of this model is considering about the document structure which is often used in text summarization. From the experiments results, our novel cognition attention attenuation based language model outperformed the classical language model with Dirichlet smoothing in blog page and Web page.
Donglin Cao, Shuo Bai, Xueqi Cheng 0001, Shaozi Li
Web Intelligence5
2007 A Document Recommendation System Based on Clustering P2P Networks
Feng Guo 0005, Shaozi Li
CDVE2
2007 A VoIP Client System Framework and its Implementation
abstract
This paper describes the design of the whole framework of a VoIP Client System. In this design, we use the design idea of automaton to build the entire dialing and hang up processes, use multithread to implement event synchronous mechanism. Then we use a strategy to couple the Jitter-Buffer design with the automatic adjustment of the redundant packets. This strategy gets a good trade-off at the client between the methods of solving delay and those of solving dithering. The experiment result indicates that this framework keeps good QoS performance even in bad network conditions.
Fengmei Zou, Yuhui Chen, Shaozi Li, Tangqiu Li, Xiaodong Shi
CSCWD4
2006 Improved Genetic Algorithm for Multiple Sequence Alignment Using Segment Profiles (GASP)
Yanping Lv, Shaozi Li, Changle Zhou, Wenzhong Guo, Zhengming Xu
ADMA2
2006 A Collaborative Multimedia Editing System Based on Shallow Nature Language Parsing
Donglin Cao, Dazhen Lin, Shaozi Li
CDVE3
2006 A New Migration Algorithm of Mobile Agent Based on Ant Colony Algorithm in P2P Network
Shaozi Li, Huowang Chen
CDVE1
2006 Research on P2P Hybrid Information Retrieval Based on Ant Colony Algorithm
abstract
Peer-to-peer (P2P) architectures make each network node capable of being both a client and a server. Such a network is potentially powerful for developing large-scale distributed information sharing. This paper presents our recent research work on the P2P hybrid information retrieval based on ant colony algorithm (ACA). From various perspectives, our work focuses on how to improve retrieval efficiency in a P2P file-sharing network. A new model is proposed to implement P2P information retrieval. In order to evaluate and validate the model, we built a simulated P2P application consist of a network of peer nodes; mobile agent travel through the network, making peer nodes communicate with each other. The results show some advantages of the proposed approach for the P2P retrieval based on ant colony algorithm using information recommendation services
Shaozi Li
CSCWD2
2006 A Cooperative Authoring System for Teaching Chinese as a Foreign Language
abstract
Shortage of a collaborative and automated platform has put textbook authoring for teaching Chinese as a foreign language into trouble. Furthermore, for lots of published textbooks, there are lots of defects which make them unacceptable. In this paper, a novel system based on computer supported cooperative work (CSCW) was developed to solve these problems. In order to improve the speed and accuracy of automatic phonetic notation, a method that combines string matching with statistics based on hidden Markov model was proposed, and a method of cross-text search and computing is adopted to control the amount and recurrence rate of new words. To get suitable examples for words, an algorithm for corpus indexing and retrieval is presented. Result indicates that textbooks authored on this system are scientific
Shaoyong Yu, Shaozi Li
CSCWD2
2005 Research on Mobile Agent Based Information Content-Sharing in Peer to Peer System
Shaozi Li, Changle Zhou, Huowang Chen
CDVE1
2005 Research on general adaptive management mechanism of active network
abstract
This paper studies a general QoS adaptive management mechanism in view of video server and active node. For the video server, we put forward an idea of video period task set as well as defining some video services with this set. This paper also designs a kind of mixed real time task scheduling arithmetic which taking QoS critical level as a character of task priority vector to build video period task set and accomplish dynamic resource management, based on QoS conveniently. Therefore, it realizes the service quality control of small granularity. Based on it, the paper further puts forward active node adaptive QoS control component. Through designing dynamic resource trading active queue and using mixed priority of video active packet, the new idea not only makes system support video application as much as possible, but also enables it to satisfy resource trading requirements of user. At last, a simulation experiment is carried out to illustrate the advantage of active queues after amelioration.
Zhongpan Qiu, Shaozi Li, Tangqiu Li
CSCWD (2)3
2005 Web document retrieval based on multi-agent
abstract
A mass of distributed and dynamic information on the Web has resulted in "information overload". With the flood of information, it has become an important research issue to search the Web based on traditional information retrieval technology. However, various systems and ambiguous terminology of information retrieval on the Web bring much trouble to users in application and researchers in development as well. This paper proposes the same interface of Web document retrieval to users, it is the model based on multi-agent. Each document in the documents base or from Web is represented as a vector in the vector space of classifiable sememes. The query from user is also represented as a vector. The relevance between them can be measured by using the cosine angle between the query and its k nearest neighbors in the vector space. Experiments have been done and their results shown that this scheme yields good results.
Shaozi Li, Changle Zhou, Huowang Chen
CSCWD (1)1
2005 An improved shifting bottleneck algorithm for job shop scheduling problem
abstract
The shifting bottleneck procedures are mainly based on one-machine scheduling algorithm, so the results of one-machine scheduling algorithm significantly affect the performance of the shifting bottleneck procedures. In this paper, inspired by two facts that the function of the release time and the delivery time of each job is uniform, an efficient algorithm for one-machine scheduling problem is presented. Based on this algorithm, an improved shifting bottleneck procedure is developed. The simulation results to the well-known benchmark problems show that the performance of the improved procedure can compete with that of the shifting bottleneck procedure and the modified ones.
Tangqiu Li, Shaozi Li
CSCWD (2)3
2005 Research on Chinese Ellipsis Recovering Based on Discourse Representation Theory
abstract
This paper presents our research work on Chinese ellipsis recovering based on discourse representation theory (DRT). From various perspectives, our work focuses on how to construct the Chinese DRT and the semantic model. Then it proposes an ellipsis recovering model based on discourse, which selects the most suitable one and fill in the syntactic ellipsis and possible semantic ellipsis. The simulated results show some advantages of the proposed approach based on DRT and potential applications in Chinese oral system and discourse machine understanding
Shaozi Li, Changle Zhou
ICTAI1
2003 Cross-Lingual Text Filtering Based on Text Concepts and kNN
Shaozi Li, Weifeng Su, Tangqiu Li, Huowang Chen
PACLIC1
2002 IP Multicasting Technique and its Application on Multimedia Network Teaching System
abstract
This paper presents the basic principles of IP multicasting. The design and application in a multimedia network teaching system are then given.
Shaozi Li, Tangqiu Li, Yidong Chen 0001
CSCWD1
2001 Cooperative Work in Multimedia Network Teaching System
abstract
This paper presents the basic architecture of CSCW (computer supported cooperative work). Then some issues of CSCW systems such as concurrency control and synchronization are discussed. Lastly, its application in a multimedia network teaching system such as a whiteboard, chatting system, etc. is given.
Shaozi Li, Tangqiu Li
CSCWD1
2001 A cross lingual texts filtering module in classifiable sememes vector space
abstract
The WWW is increasingly being used as a source of information. The volume of this information is accessed by users using direct manipulation tools. The paper describes a module that sifts through a large number of texts retrieved by the user. We describe a system that learns a model of the user's preferences, filters the information, and notifies the user when relevant information becomes available. The user's model is represented as a vector in the vector space of classifiable sememes. The document is also represented as a vector. The relevance of the text to the user's interest can be measured by using the cosine angle between the two vectors. Experiments are given to demonstrate it to be a good idea.
Shaozi Li, Weifeng Su, Tangqiu Li
SMC1
2001 A hybrid method for syntactic and semantic structure disambiguation for Chinese
abstract
This paper presents a method of syntactic and semantic structure disambiguation for Chinese. The method finds the most plausible interpretation of a phrase or a sentence in Chinese by evaluation of the similarity between the structure and examples in related word entries in a knowledge base, Hownet. First we put forward the general idea of the method and briefly introduce its semantic knowledge resource-the Hownet Dictionary. The main algorithm of the method is then proposed with detail. The experimental result shows that the method is effective.
Tangqiu Li, Qingyang Hong, Shaozi Li
SMC4