EDBT 2026 Demo / reviewers in the wild / expert
Jie Gui
dblp:45/794
· DBLP profile ↗
110ranked-venue papers
29as first author
76since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 52 · 11 first-author · 41 since 2021Artificial intelligence and machine learning · 44 · 16 first-author · 27 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 1 first-author · 5 since 2021Security and privacy · 10 · 3 first-author · 9 since 2021Computer networks · 6 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 6 · 4 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Diversifying Counterattacks: Orthogonal Exploration for Robust CLlP InferenceabstractVision-language pre-training models (VLPs) demonstrate strong multimodal understanding and zero-shot generalization, yet remain vulnerable to adversarial examples, raising concerns about their reliability. Recent work, Test-Time Counterattack (TTC), improves robustness by generating perturbations that maximize the embedding deviation of adversarial inputs using PGD, pushing them away from their adversarial representations. However, due to the fundamental difference in optimization objectives between adversarial attacks and counterattacks, generating counterattacks solely based on gradients with respect to the adversarial input confines the search to a narrow space. As a result, the counterattacks could overfit limited adversarial patterns and lack the diversity to fully neutralize a broad range of perturbations. In this work, we argue that enhancing the diversity and coverage of counterattacks is crucial to improving adversarial robustness in test-time defense. Accordingly, we propose Directional Orthogonal Counterattack (DOC), which augments counterattack optimization by incorporating orthogonal gradient directions and momentum-based updates. This design expands the exploration of the counterattack space and increases the diversity of perturbations, which facilitates the discovery of more generalizable counterattacks and ultimately improves the ability to neutralize adversarial perturbations. Meanwhile, we present a directional sensitivity score based on averaged cosine similarity to boost DOC by improving example discrimination and adaptively modulating the counterattack strength. Extensive experiments on 16 datasets demonstrate that DOC improves adversarial robustness under various attacks while maintaining competitive clean accuracy. Chengze Jiang, Minjing Dong, Xinli Shi, Jie Gui |
AAAI | 4 |
| 2026 | KaeTE: Towards Practical Neural Traffic Engineering with Lagrangian Duality and Learning-to-optimizeabstractTraffic engineering (TE) is becoming increasingly important in modern networks, as it can improve network performance by splitting traffic across paths. However, traditional TE solvers can be too slow for rapid changes, while recent machine learning (ML) solvers are fast but often fail to support dynamic network conditions, such as topology changes or link capacity changes. Moreover, they usually support only simple TE objectives that do not account for potential link overload, which makes them less practical. In this paper, we present KaeTE, an ML-based TE solver that supports dynamic network conditions and the throughput objective. The design of KaeTE leverages the convexity of the throughput objective. Specifically, KaeTE first makes the throughput objective strongly convex through regularization. KaeTE then adopts a learning-to-optimize (L2O)-inspired model to iteratively refine the dual variables. To ensure that the final TE solutions do not overload any link, KaeTE generates the final TE solutions and the corresponding throughput objective value through a constraint-aware loss function. Evaluations on both dynamic and static network conditions show that KaeTE consistently outperforms baselines in our evaluated settings. Zirui Ou, Yanghao Zhang, Jie Gui, Qun Huang 0001 |
APNet | 3 |
| 2026 | Frequent Checkpointing through Mergeable Delta Compression and Semi-Reliable Transmission
Jie Gui, Xiandong Lu, Qun Huang 0001 |
INFOCOM | 1 |
| 2026 | Deep Learning-Based Point Cloud Registration: A Comprehensive Survey and Taxonomy
Yu-Xin Zhang 0004, Jie Gui, Baosheng Yu, Xiaofeng Cong, Xin Gong 0001, Wenbing Tao, Dacheng Tao |
Int. J. Comput. Vis. | 2 |
| 2026 | Brightness-Aware Synthetic-to-Real Learning for Nighttime Hazy Image EnhancementabstractNighttime hazy vision is severely limited by the presence of haze and multi-colored light sources. Different from the daytime image dehazing task which has been widely studied, less progress has been made in nighttime image dehazing. In this paper, through extensive analysis and experimentation, we find that game engine simulations offer strong real-world generalization but suffer from unrealistic brightness. To tackle this, we introduce a three-step, brightness-aware synthetic-to-real learning approach. First, we use supervised learning to train a spatial-frequency network (SFN) on synthetic data to produce pseudo-labels. With these pseudo-labels, we develop a semi-supervised dehazing model (SFN+) that minimizes domain discrepancy through a brightness consistency loss applied to local windows. Building on SFN+, we fine-tune the model for better vision using a relative brightness improvement strategy that accounts for color shifts from lighting and brightness shifts during enhancement (SFN++). Experiments on popular benchmark datasets confirm our method's superiority over state-of-the-art approaches. Jie Gui, Xiaofeng Cong, Yu-Xin Zhang 0004, Junming Hou, Dacheng Tao |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | Zero-shot skeleton-based action recognition with dual visual-text alignment
Jidong Kuang, Hongsong Wang 0001, Chaolei Han 0001, Yang Zhang 0002, Jie Gui |
Pattern Recognit. | 5 |
| 2026 | Efficient Diffusion-Based 3D Human Pose Estimation With Hierarchical Temporal PruningabstractDiffusion models have demonstrated strong capabilities in generating high-fidelity 3D human poses, yet their iterative nature and multi-hypothesis requirements incur substantial computational cost. In this paper, we propose an efficient diffusion-based 3D human pose estimation framework with a Hierarchical Temporal Pruning (HTP) strategy, which dynamically prunes redundant pose tokens across both frame and semantic levels while preserving critical motion dynamics. HTP operates in a staged, top-down manner: (1) Temporal Correlation-Enhanced Pruning (TCEP) identifies essential frames by analyzing inter-frame motion correlations through adaptive temporal graph construction; (2) Sparse-Focused Temporal MHSA (SFT MHSA) leverages the resulting frame-level sparsity to reduce attention computation, focusing on motion-relevant tokens; and (3) Mask-Guided Pose Token Pruner (MGPTP) performs fine-grained semantic pruning via clustering, retaining only the most informative pose tokens. Experiments on Human3.6M and MPI-INF-3DHP show that HTP reduces training MACs by 38.5%, inference MACs by 56.8%, and improves inference speed by an average of 81.1% compared to prior diffusion-based methods, while achieving state-of-the-art performance. Yuquan Bi, Hongsong Wang 0001, Xinli Shi, Zhipeng Gui, Jie Gui, Yuan Yan Tang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | Gradient Perturbation Guidance for Boosting Sparse Adversarial Attack TransferabilityabstractSparse adversarial attacks perturb only a few pixels to achieve an attack, making them harder to detect and more dangerous. Recently, generative sparse attacks decouple the generation of sparse adversarial examples (AEs) into dense perturbations and sparse masks. By modeling the data distribution from clean examples to sparse AEs, generative sparse attacks mitigate the poor transferability that arises from over-reliance on gradients. These methods put effort into deriving optimal sparse masks on the generated perturbation. However, the quality of perturbation generation has always been overlooked, which limits the transferability of sparse AEs. To explore the influence of perturbation quality, we conduct empirical analyses of sparse gradient-based perturbations. The results show that directly applying sparsity to gradient-based perturbations disrupts their holistic adversarial information, leading to degraded attack performance. Therefore, it is critical to extract key adversarial knowledge from gradient-based perturbations while preserving their overall integrity to guide sparse adversarial attacks. Motivated by this observation, we propose to extract essential adversarial information from gradient-based AEs to guide the generator to produce higher-quality dense perturbations and stronger transferable sparse AEs. Specifically, we introduce the Gradient Perturbation Guidance (GPG) sparse adversarial attack, which integrates gradient adversarial feature guidance and gradient perturbation guidance regularization. The former guides the generator to capture gradient-based adversarial features during encoding, while the latter refines adversarial knowledge from gradient-based perturbations during decoding. Extensive experiments on ImageNet-1K show that our GPG significantly boosts transferability compared to state-of-the-art methods under consistent sparsity constraints. Our code is available at Github. Chengze Jiang, Minjing Dong, Jie Gui, Lu Dong 0002, Yuan Yan Tang, James T. Kwok |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Improving Fast Adversarial Training Paradigm: An Example Taxonomy PerspectiveabstractWhile adversarial training is an effective defense method against adversarial attacks, it notably increases the training cost. To this end, fast adversarial training (FAT) is presented for efficient training and has become a hot research topic. However, FAT suffers from catastrophic overfitting, which leads to a performance drop compared with multi-step adversarial training. However, the cause of catastrophic overfitting remains unclear and lacks exploration. In this paper, we present an example taxonomy in FAT, which suggests that catastrophic overfitting is correlated with the imbalance between the inner and outer optimization in FAT. Furthermore, we investigated the impact of varying degrees of training loss, revealing a correlation between training loss and catastrophic overfitting. Based on these observations, we redesign the loss function in FAT with the proposed dynamic label relaxation to concentrate the loss range and reduce the impact of misclassified examples. Meanwhile, we introduce batch momentum initialization to enhance diversity and prevent catastrophic overfitting in an efficient manner. Furthermore, we also propose Catastrophic Overfitting aware Loss Adaptation (COLA), which employs a separate training strategy for examples based on their loss degree. Our proposed method, named example taxonomy aware FAT (ETA), establishes an improved paradigm for FAT. Experiment results demonstrate that our ETA achieves higher robust accuracy than all other evaluated methods. Comprehensive experiments on four standard datasets demonstrate the competitiveness of our method. The source code and model checkpoints will be publicly released. Jie Gui, Chengze Jiang, Minjing Dong, Kun Tong, Xinli Shi, Yuan Yan Tang, Dacheng Tao |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2026 | SMInject: Specious Malignant Injection Attacks With Semantically-Enhanced Tokens in Cross-Modal RetrievalabstractThe pre-training multimodal models have achieved remarkable success with powerful cross-modal understanding capabilities, while easily being affected by deliberate injection attacks. Although the deceptive injection attacks are harmful, they are valuable in revealing the vulnerability and improving the robustness for multimodal models. Unfortunately, the existing multimodal injection attacks pay less attention to the complicated roles of different modality-related causal correlation, which results in such attacks being susceptible to detection and defense. To alleviate this issue, we propose a novel specious malignant injection attack framework, calledSMInject, which exploits both the irrationality and causal correlation across diverse modalities to stealthily manipulate the space of output. To enhance the stealthiness, we generate deceptive injections to assemble the concepts by analyzing causal correlation under four types of attacks. To further boost the effectiveness, the malignant injections are guided to penetrate in the encoded embedding space by designing the premise-hypothesis consensus alignment. Extensive experiments on representative multimodal models demonstrate that ourSMInjectachieves over 14% higher attack success rate and 6% higher Hit@5 metric than state-of-the-art methods while preserving the overall utility of models. Moreover, we highlight that theSMInjectalso exhibits the desired transferability by investigating the impact of contextual factors, such as similar attack profiles, imperceptible noise perturbations,etc. Our code is available athttps://anonymous.4open.science/r/SMInject-0DBC. Ju Jia, Jiabao Guo, Xiaojun Jia, Siqi Ma 0001, Jie Gui, Robert H. Deng |
IEEE Trans. Dependable Secur. Comput. | 6 |
| 2026 | MeanCut: Greedy Graph Clustering by Fast Maximum Spanning Tree and Degree Descent CriterionabstractAs the most typical graph clustering method, spectral clustering is popular and attractive due to its remarkable performance, easy implementation, and strong adaptability. Classical spectral clustering measures the edge weights using pairwise Euclidean similarity and resolves the optimal graph partitioning by relaxing the constraints of indicator matrix and decomposing the Laplacian matrix. However, Euclidean similarity might cause skew graph cuts when handling non-spherical clusters, and the relaxation strategy introduces information loss. Meanwhile, spectral clustering requires specifying the number of clusters, which is difficult to determine without enough prior knowledge. In this work, we propose a greedy-optimized scheme for resolving the indicator matrix using path-based similarity and degree descent criterion. It yields an indicator matrix with strictly binary entries without destructive relaxation and discretization steps. Path-based similarity can enhance the intra-cluster associations of arbitrary-shaped clusters, while degree descending is theoretically proven to be the best order to minimize our proposed objective function MeanCut. Moreover, we define a density gradient factor to separate clusters with fuzzy boundaries, and develop a fast maximum spanning tree algorithm to improve the scalability of similarity calculation. The effectiveness of MeanCut has been demonstrated on synthetic datasets and real-world benchmarks. By fusing multi-view image features, MeanCut outperforms cutting-edge subspace clustering methods in face recognition. The code is available at:https://github.com/ZPGuiGroupWhu/MeanCut. Dehua Peng, Zhipeng Gui, Jie Gui, Huayi Wu |
IEEE Trans. Fuzzy Syst. | 5 |
| 2026 | Axial-View-Oriented Contrastive Adversarial Training for Robust Point Cloud RecognitionabstractContrastive adversarial training emerges as an effective approach to enhancing model robustness in safety-critical applications, particularly point cloud recognition for autonomous driving and medical imaging. However, existing point cloud adversarial training methods mainly emphasize global contrastive learning while overlooking local geometric variations induced by adversarial perturbations. Motivated by the spatial and intensity variations of perturbations across axial views, we propose AVOC, a novel local-global adversarial training framework that utilizes axial-view-oriented contrastive learning. This framework leverages the smallest axial view for local contrastive learning, as it exhibits the highest perturbation differences, and utilizes the largest axial view for global contrastive learning, as it preserves global structural consistency. We conduct comprehensive experiments across four representative architectures, demonstrating significant robustness improvements on widely-adopted recognition benchmarks, including ModelNet40, ShapeNetPart, ModelNet40-C, and ScanObjectNN-C, and further validate its effectiveness on the large-scale KITTI benchmark for 3D object detection. Our results across diverse perturbation scenarios, encompassing white-box attacks, black-box attacks, and natural perturbations, demonstrate the consistent and significant model robustness enhancement of our proposed method. Jie Gui, Yu-Xin Zhang 0004, Xiaofeng Cong, Baosheng Yu, Zhipeng Gui, Yuan Yan Tang, James T. Kwok |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2026 | Rethinking Frequency Modeling: Tail-Aware Dynamic Adversarial Training for Long-Tailed RobustnessabstractAdversarial training (AT) is among the most effective defenses against adversarial attacks on deep neural networks. However, in real-world scenarios where data often follow long-tailed distributions, conventional AT methods struggle to handle such imbalance, resulting in severe robustness disparities across classes and limited overall robustness. Although recent efforts attempt to improve robustness through class frequency-aware weighting or distribution adjustments, our empirical analysis reveals that class frequency alone is an insufficient indicator of adversarial vulnerability, as robust accuracy does not correlate with the number of examples per class. Furthermore, AT under long-tailed distributions exhibits optimization instability, particularly for tail classes with limited data. To address these challenges, we present Tail-Aware Dynamic Adversarial Training (TAD-AT), which integrates three complementary components targeting the training loss, attack strategy, and weight average. TAD-AT captures data imbalance and performance disparity, improving adversarial robustness under long-tailed distributions. First, our training loss incorporates frequency- and accuracy-aware regularization to emphasize learning for vulnerable classes. Second, our attack adjusts perturbations based on class-wise vulnerability, encouraging robust feature learning around vulnerable regions, thereby mitigating robustness overfitting and improving clean accuracy. Third, our weight average improves robust generalization and training stability by adaptively controlling the decay rate across classes. Experiments on long-tailed benchmarks demonstrate that our TAD-AT significantly improves adversarial robustness, offering a systematic and practical solution to robustness challenges under long-tail distributions. Our code is publicly available on https://github.com/bookman233/TADAT. Chengze Jiang, Minjing Dong, Jie Gui, Ju Jia, Yuan Yan Tang, James T. Kwok |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2026 | PANDA: Diffusion-Guided Purification and Adaptation for Robust Point Cloud Classification Against Adversarial Attack
Yu-Xin Zhang 0004, Xiaofeng Cong, Minjing Dong, Zhipeng Gui, Jie Gui, Yuan Yan Tang, James T. Kwok |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2026 | Long-Tailed Approaching Cross-Modal Hashing With Multi-Expert Collaborative LearningabstractCross-modal hashing enables efficient retrieval across different modalities by mapping heterogeneous data into compact binary codes within a shared Hamming space. However, most existing methods assume that data from each class are evenly distributed, which contradicts the long-tailed nature of real-world data. Consequently, these approaches often exhibit suboptimal performance when handling imbalanced datasets. The only existing cross-modal hashing method that considers long-tailed data attempts to mine both the individuality and commonality across modalities, yet it relies on a negative log-likelihood pairwise loss that tends to bias the model toward head categories. To address this issue, we propose a novel Long-tailed Approaching Cross-modal Hashing (LACH) framework based on multi-expert collaborative learning. Specifically, LACH constructs a multi-expert architecture with a Graph Convolutional Network (GCN) backbone to facilitate knowledge transfer. Unlike conventional multi-expert models that either share identical data copies or employ entirely distinct data subsets, we introduce a partial data replication strategy that ensures each expert receives a balanced yet overlapping training set. Furthermore, we design a proxy-based pointwise loss to treat head and tail categories equitably, along with an inter-modal approaching loss to enhance the alignment of hash codes across modalities within each class. Extensive experiments demonstrate that LACH achieves accuracy improvements of up to 4.2% and 6.3% over state-of-the-art baselines on balanced and long-tailed datasets, respectively. Our code is available at https://github.com/caoyuan618/LACH. Yuan Cao 0005, Zifan Liu, Weikang Gao, Jie Gui, Yanwei Yu |
IEEE Trans. Image Process. | 4 |
| 2026 | Focus on Finding Deepfakes: A Robust Proactive Detection Method Based on Orthogonal Moment WatermarkingabstractDeepfake detection remains a challenging research topic, especially when the quality of forged images degrades, leading to unreliable detection results. In this paper, we propose a watermarking-based proactive method for robust proactive deepfake detection. First, we embed a watermark into the Fractional-order Quaternion Exponent Moments (FrQEMs) space of the host face image, achieving a balance between imperceptibility and robustness of the watermarking algorithm. Then, we introduce the Frequency Mamba (FreMamba) block to enhance feature extraction by leveraging correlations between frequency-domain subbands, thereby enabling the extraction of more discriminative feature representations. Finally, at the detection stage, we construct a dual-branch framework comprising a watermark extractor and a forgery discriminator. Through knowledge distillation, the watermark extractor guides the forgery discriminator to perceive forgery traces. Specifically, the integrity of the extracted watermark is compromised only when the host image is subjected to a deepfake attack, while conventional attacks do not affect the integrity. Experimental results on benchmark datasets demonstrate that the proposed method achieves superior deepfake detection accuracy. In particular, when images are subjected to conventional attacks, our method surpasses state-of-the-art approaches by more than 5.3% in terms of ACC. Chunpeng Wang 0001, Shanshan Zhang 0001, Jie Gui, Qi Li 0029, Yunan Liu 0001 |
IEEE Trans. Image Process. | 4 |
| 2026 | Data-Free Class-Incremental Gesture Recognition With Prototype-Guided Pseudo-Feature ReplayabstractGesture recognition is an important research area in the field of computer vision. Most existing efforts focus on close-set scenarios, thereby limiting the capacity to effectively handle unseen or novel gestures. We aim to address class-incremental gesture recognition, which entails the ability to accommodate new and previously unseen gestures over time. Specifically, we introduce a Prototype-Guided Pseudo Feature Replay framework for data-free class-incremental learning. This framework comprises four components: Pseudo Feature Generation with Batch Prototypes (PFGBP), Variational Prototype Replay for old classes, Truncated Cross-Entropy for new classes, and Continual Classifier Re-Training. To tackle the issue of catastrophic forgetting, the PFGBP dynamically generates a diversity of pseudo features in an online manner, leveraging class prototypes of old classes along with batch class prototypes of new classes. Furthermore, the Variational Prototype Replay enforces consistency between the classifier's weights and the prototypes of old classes, leveraging class prototypes and covariance matrices to enhance robustness and generalization capabilities. The Truncated Cross-Entropy mitigates the impact of domain differences of the classifier caused by pseudo features. Finally, the Continual Classifier Re-Training training strategy is designed to prevent overfitting to new classes and ensure the stability of features extracted from old classes. Extensive experiments conducted on two widely used gesture recognition datasets, namely SHREC 2017 3D and EgoGesture 3D, demonstrate that our approach outperforms existing state-of-the-art methods by 11.8% and 12.8% in terms of mean global accuracy, respectively. The code is available on https://github.com/sunao-101/PGPFR-3/. Hongsong Wang 0001, Jie Gui, Liang Wang 0001 |
IEEE Trans. Image Process. | 3 |
| 2026 | Enabling General and Efficient Window Mechanism for In-Network TelemetryabstractRecent network telemetry solutions typically target programmable switches to achieve high performance and in-network visibility. They partition the packet stream into windows and then apply various stream processing techniques to summarize flow-level statistics. However, existing studies focus on the measurement within each window. Window management is still a missing piece due to the resource limitation of programmable switches. In this paper, we propose OmniWindow, a general and efficient window mechanism framework. OmniWindow splits the original window into fine-grained sub-windows such that the sub-windows can be merged into various window types. To deal with the resource restriction, OmniWindow carefully designs its data plane memory layout and proposes a window synchronization method. It also employs a collaborative architecture that can collect and reset stateful data in sub-windows within a limited time. We prototype OmniWindow on Tofino. We incorporate OmniWindow into a SOTA query-driven telemetry system and eight sketch-based telemetry algorithms. Our experiments demonstrate that OmniWindow enables these telemetry solutions to achieve higher accuracy than conventional window mechanism. Haifeng Sun 0004, Jintao He, Jie Gui, Qun Huang 0001 |
IEEE Trans. Netw. | 4 |
| 2025 | Deep Graph Online Hashing for Multi-Label Image RetrievalabstractOnline hashing has attracted much research attention for large-scale image retrieval in a streaming way. The main challenge lies in keeping balance between high retrieval accuracy and low training time. Existing online hashing methods almost rely on shallow models rather than deep networks due to high training costs, because it is unacceptable to update hash functions on an order of hours. In addition, the multi-label supervision information is not fully utilized to guide the hash learning process and the affinity matrix is always fixed once constructed. In this paper, we propose a novel Deep Graph Online Hashing (DGOH) method, which for the first time introduces inductive graph neural networks (GNNs) to realize deep online hashing with acceptable training costs on an order of seconds. Furthermore, we mine the multi-label information of the images by constructing a label network and learn label-wise weights dynamically to help to update the affinity matrix. In addition, we provide a strategy to obtain examples from the old data to solve the catastrophic forgetting problem. An integrated objective function is designed to train the entire architecture. Extensive experiments on two common benchmarks demonstrate that the proposed method achieves up to 13.3% accuracy gains over state-of-the-art baselines and shows competitive performance on training time. Yuan Cao 0005, Xiangru Chen 0001, Zifan Liu, Wenzhe Jia, Fanlei Meng, Jie Gui |
AAAI | 6 |
| 2025 | External Reliable Information-enhanced Multimodal Contrastive Learning for Fake News DetectionabstractWith the rapid development of the Internet, the information dissemination paradigm has changed and the efficiency has been improved greatly. While this also brings the quick spread of fake news and leads to negative impacts on cyberspace. Currently, the information presentation formats have evolved gradually, with the news formats shifting from texts to multimodal contents. As a result, detecting multimodal fake news has become one of the research hotspots. However, multimodal fake news detection research field still faces two main challenges: the inability to fully and effectively utilize multimodal information for detection, and the low credibility or static nature of the introduced external information, which limits dynamic updates. To bridge the gaps, we propose ERIC-FND, an external reliable information-enhanced multimodal contrastive learning framework for fake news detection. ERIC-FND strengthens the representation of news contents by entity-enriched external information enhancement method. It also enriches the multimodal news information via multimodal semantic interaction method where the multimodal constrative learning is employed to make different modality representations learn from each other. Moreover, an adaptive fusion method is taken to integrate the news representations from different dimensions for the eventual classification. Experiments are done on two commonly used datasets in different languages, X (Twitter) and Weibo. Experiment results demonstrate that our proposed model ERIC-FND outperforms existing state-of-the-art fake news detection methods under the same settings. Biwei Cao, Qihang Wu, Jiuxin Cao, Bo Liu 0004, Jie Gui |
AAAI | 5 |
| 2025 | Dual Conditioned Motion Diffusion for Pose-Based Video Anomaly DetectionabstractVideo Anomaly Detection (VAD) is essential for computer vision and multimedia research. Existing VAD methods utilize either reconstruction-based or prediction-based frameworks. The former excels at detecting irregular patterns or structures, whereas the latter is capable of spotting abnormal deviations or trends. We address pose-based video anomaly detection and introduce a novel framework called Dual Conditioned Motion Diffusion (DCMD), which enjoys the advantages of both approaches. The DCMD integrates conditioned motion and conditioned embedding to comprehensively utilize the pose characteristics and latent semantics of observed movements, respectively. In the reverse diffusion process, a motion transformer is proposed to capture potential correlations from multi-layered characteristics within the spectrum space of human motion. To enhance the discriminability between normal and abnormal instances, we design a novel United Association Discrepancy (UAD) regularization that primarily relies on a Gaussian kernel-based time association and a self-attention-based global association. Finally, a mask completion strategy is introduced during the inference stage of the reverse diffusion process to enhance the utilization of conditioned motion for the prediction branch of anomaly detection. Extensive experiments conducted on four datasets demonstrate that our method dramatically outperforms state-of-the-art methods and exhibits superior generalization performance. Hongsong Wang 0001, Andi Xu, Pinle Ding, Jie Gui |
AAAI | 4 |
| 2025 | Heterogeneous Skeleton-Based Action Representation Learning
Hongsong Wang 0001, Jidong Kuang, Jie Gui |
CVPR | 4 |
| 2025 | Backdooring Self-Supervised Contrastive Learning by Noisy AlignmentabstractSelf-supervised contrastive learning (CL) effectively learns transferable representations from unlabeled data containing images or image-text pairs but suffers vulnerability to data poisoning backdoor attacks (DPCLs). An adversary can inject poisoned images into pretraining datasets, causing compromised CL encoders to exhibit targeted misbehavior in downstream tasks. Existing DPCLs, however, achieve limited efficacy due to their dependence on fragile implicit co-occurrence between backdoor and target object and inadequate suppression of discriminative features in backdoored images. We propose Noisy Alignment (NA), a DPCL method that explicitly suppresses noise components in poisoned images. Inspired by powerful training-controllable CL attacks, we identify and extract the critical objective of noisy alignment, adapting it effectively into data-poisoning scenarios. Our method implements noisy alignment by strategically manipulating contrastive learning's random cropping mechanism, formulating this process as an image layout optimization problem with theoretically derived optimal parameters. The resulting method is simple yet effective, achieving state-of-the-art performance compared to existing DPCLs, while maintaining clean-data accuracy. Furthermore, Noisy Alignment demonstrates robustness against common backdoor defenses. Codes can be found at https://github.com/jsrdcht/Noisy-Alignment. Tuo Chen, Jie Gui, Minjing Dong, Ju Jia, Lanting Fang |
ICCV | 2 |
| 2025 | LOTA: Bit-Planes Guided AI-Generated Image Detection
Hongsong Wang 0001, Renxi Cheng, Yang Zhang 0002, Chaolei Han 0001, Jie Gui |
ICCV | 5 |
| 2025 | FD-Filter: A Compact Data Structure for Fine-Grained Intra-Flow Packet Delay Monitoring
Jintao He, Jie Gui, Tian Lv, Qun Huang 0001 |
INFOCOM | 2 |
| 2025 | FingerVeinSyn-5M: A Million-Scale Dataset and Benchmark for Finger Vein RecognitionabstractA major challenge in finger vein recognition is the lack of large-scale public datasets. Existing datasets contain few identities and limited samples per finger, restricting the advancement of deep learning-based methods. To address this, we introduce FVeinSyn, a synthetic generator capable of producing diverse finger vein patterns with rich intra-class variations. Using FVeinSyn, we created FingerVeinSyn-5M -- the largest available finger vein dataset -- containing 5 million samples from 50,000 unique fingers, each with 100 variations including shift, rotation, scale, roll, varying exposure levels, skin scattering blur, optical blur, and motion blur. FingerVeinSyn-5M is also the first to offer fully annotated finger vein images, supporting deep learning applications in this field. Models pretrained on FingerVeinSyn-5M and fine-tuned with minimal real data achieve an average 53.91% performance gain across multiple benchmarks. The dataset is publicly available at: https://github.com/EvanWang98/FingerVeinSyn-5M. Yifan Wang 0036, Jie Gui, Baosheng Yu, Qi Li 0005, Zhenan Sun, Juho Kannala, Guoying Zhao 0001 |
ACM Multimedia | 2 |
| 2025 | Resilient and Efficient Multirobot Pickup and Delivery Against Strategic Attacks: A Three-Layered FrameworkabstractThe widespread application of intelligent storage systems has urged more requirements for system security. In this work, we consider a resilient pickup and delivery task of the multi-robot systems against strategic attacks. The strategic attacks terminate the normal motion of a fraction of key operating robots with a well-designed strategy, and thus jeopardize the cooperation among robots. A three-layered decision framework is proposed to suppress the above strategic attacks: The first layer introduces defensive strategies, which incorporate reputation mechanisms and counterfactual rescue mode allocations. The counterfactual rescue-mode allocation mechanism dynamically assesses the benefit differences between rescuing others for resilience and conducting self-tasks for efficiency. Based on the reputation mechanisms, the second layer employs a bi-level programming method for pickup/delivery mode allocations and task assignments. Then, the third layer calculates the collision-free path for the swarm using a mixed integer linear programming model. The practicality and resilience of this algorithm against strategic attacks have been demonstrated through numerical simulations involving various robot scales and task burdens. Xin Gong 0001, Jie Gui, Zhan Shu 0001, Tingwen Huang |
IEEE Internet Things J. | 3 |
| 2025 | Resilient Human-in-the-Loop Formation-Tracking of Multi-UAV Systems Against Byzantine AttacksabstractThis study addresses resilient human-in-the-loop (HiTL) formation-tracking of multi-UAV systems against$f$-local Byzantine attacks. In the HiTL settings, a human operator plays a key role in detecting any physical hazard, monitoring the whole UAV swarm, and sending secure execution signals to a non-autonomous leader UAV. Moreover, there exists a fraction of Byzantine UAVs in the multi-UAV systems, which propagate incorrect information to their neighbors (called Byzantine edge attacks (BEAs)) and adopt false input signals (called Byzantine node attacks (BNAs)) when swarming. In order to suppress the above aggressive Byzantine attacks, this paper proposes a Byzantine-resilient hierarchical control scheme, including a virtual Digital Twin Layer (DTL) apart from a Cyber-Physical Layer (CPL). First, a distributed resilient estimation scheme is proposed on the DTL, which can realize resilient estimation on the state of the non-autonomous leader UAV against BEAs on the premise that the DTL topology is strongly$(2f+1)$-robust. Second, a series of decentralized and chattering-free controllers is formulated on the CPL, which is resilient to both BNAs and inter-layered faults. The asymptotical control performance of the above controllers is strictly proven based on Cromwell-Bellman Lemma. To demonstrate the practicality of the theoretical results, a resilient HiTL multi-UAV systems experiment has been further conducted. The experimental results verify the effectiveness and practicality of the designed two-layered controllers against$f$-local Byzantine attacks.Note to Practitioners—Owing to the wide application of multi-UAV systems, the resilience of the whole swarm against malicious attacks has grasped the great attention of both academia and industry. This work considers a rather aggressive kind of attacks, named Byzantine attacks, where a fraction of unidentified UAVs act as traitors. Inspired by the digital twin technology, a two-layered control architecture for multi-UAV systems is formatted, including a Digital Twin Layer (DTL) and a Cyber-Physical Layer (CPL). Here are the highlights: 1) Control Architecture: The DTL handles Byzantine edge attacks (BEAs), while the CPL addresses Byzantine node attacks (BNAs), ensuring reliable human-swarm cooperation in adversarial environments. 2) Resilient Estimation against BEAs: A novel resilient estimation scheme on the DTL is designed, using edge-based feedback, which can estimate the states of the leader UAV manipulated by human operators. 3) Adaptive Controller against BNAs and Inter-layered Faults: On the CPL, a decentralized adaptive controller with adjustable and exponential convergence is proposed, enhancing its precision and flexibility. 4) Practical Application: A UAV swarm formation-tracking experiment validates the control architecture’s effectiveness in human-in-the-loop scenarios, demonstrating its practicality in the realm of swarm robotics and human-swarm interaction. Xin Gong 0001, Jie Gui, Yong Chen 0006, Wenwu Yu, Tingwen Huang |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2025 | A Privacy-Preserving Large-Scale Image Retrieval Framework With Vision GNN HashingabstractWith the growing popularity of cloud services, companies and individuals outsource images to cloud servers to reduce storage and computing burdens. The images are encrypted before outsourcing for privacy protection. It has become urgent to solve the privacy-preserving image retrieval problem on the cloud. There are three main challenges in this area. First, how can we achieve high retrieval accuracy on the encryption domain? Second, how can we improve efficiency in large-scale encrypted image retrieval? Third, how can we ensure the reliability of the retrieval results? The existing schemes only consider some of these characteristics and the retrieval accuracy is insufficient. In this paper, we propose a privacy-preserving large-scale image retrieval framework with vision graph convolutional neural network hashing (ViGH). To the best of our knowledge, this is the first framework that is able to address all the above challenges with more advanced accuracy performance. To be specific, cycle-consistent adversarial networks and vision graph convolutional networks (ViG) are utilized to increase retrieval accuracy. By embedding encrypted images into hash codes, we can obtain high retrieval efficiency by Hamming distances. Cloud servers store the hash codes on the blockchain (Ethereum). The retrieval algorithm on the smart contracts and the consensus mechanism of blockchain ensure reliability of the retrieval results. The experimental results on three common datasets verify the effectiveness and efficiency of the proposed privacy-preserving image retrieval framework. The reliability of the retrieval results is ensured by the consensus mechanism of blockchain with no need for verification. Yuan Cao 0005, Fanlei Meng, Xinzheng Shang, Jie Gui, Yuan Yan Tang |
IEEE Trans. Big Data | 4 |
| 2025 | No Place to Hide: Dual Deep Interaction Channel Network for Fake News Detection With Data AugmentationabstractOnline social network has emerged as a prominent place for the propagation of fake news due to its low cost of information dissemination. Although the existing methods have made many attempts in news content and propagation structure, the detection of fake news is still facing two challenges: one is how to mine the unique key features and evolution patterns, and the other is how to tackle the problem of small samples to build the high-performance model. Different from popular methods, which take full advantage of the propagation topology structure, in this article, we propose a novel framework for fake news detection from perspectives of semantics, emotion and data enhancement. The semantic and emotional features of news and comments, the inconsistent emotion between news and news participants as well as the emotion evolution features in comments are fused by the designed dual deep interaction channel network to obtain a more comprehensive and fine-grained news representation. Meanwhile, with the construction of large language model (LLM) prompt, a LLM-based data enhancement module is used to obtain more diverse labeled data of high quality filtered by confidence, further improving the performance of the classification model. Experiments show that the proposed approach outperforms the state-of-the-art methods. Biwei Cao, Jiuxin Cao, Lulu Hua, Bo Liu 0004, Jie Gui, James T. Kwok |
IEEE Trans. Comput. Soc. Syst. | 7 |
| 2025 | A Robust and Efficient Boundary Point Detection Method by Measuring Local Direction DispersionabstractBoundary point detection aims to outline the external contour structure of clusters and enhance the inter-cluster discrimination, thus bolstering the performance of the downstream classification and clustering tasks. However, existing boundary point detectors are sensitive to density heterogeneity or cannot identify boundary points in concave structures and high-dimensional manifolds. In this work, we propose a robust and efficient boundary point detection method based on Local Direction Dispersion (LoDD). The core of boundary point detection lies in measuring the difference between boundary points and internal points. It is a common observation that an internal point is surrounded by its neighbors in all directions, while the neighbors of a boundary point tend to be distributed only in a certain directional range. By considering this observation, we adopt density-independent K-Nearest Neighbors (KNN) method to determine neighboring points and design a centrality metric LoDD using the eigenvalues of the covariance matrix to depict the distribution uniformity of KNN. We also develop a grid-structure assumption of data distribution to determine the parameters adaptively. The effectiveness of LoDD is demonstrated on synthetic datasets, real-world benchmarks, and application of training set split for deep learning model and hole detection on point cloud data. The datasets and toolkit are available at:https://github.com/ZPGuiGroupWhu/lodd. Dehua Peng, Zhipeng Gui, Jie Gui, Huayi Wu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Improving Fast Adversarial Training via Self-Knowledge GuidanceabstractAdversarial training has achieved remarkable advancements in defending against adversarial attacks. Among them, fast adversarial training (FAT) is gaining attention for its ability to achieve competitive robustness with fewer computing resources. Existing FAT methods typically employ a uniform strategy that optimizes all training data equally without considering the influence of different examples, which leads to an imbalanced optimization. However, this imbalance remains unexplored in the field of FAT. In this paper, we conduct a comprehensive study of the imbalance issue in FAT and observe an obvious class disparity regarding their performances. This disparity could be embodied from a perspective of alignment between clean and robust accuracy. Based on the analysis, we mainly attribute the observed misalignment and disparity to the imbalanced optimization in FAT, which motivates us to optimize different training data adaptively to enhance robustness. Specifically, we take disparity and misalignment into consideration. First, we introduce self-knowledge guided regularization, which assigns differentiated regularization weights to each class based on its training state, alleviating class disparity. Additionally, we propose self-knowledge guided label relaxation, which adjusts label relaxation according to the training accuracy, alleviating the misalignment and improving robustness. By combining these methods, we formulate the Self-Knowledge Guided FAT (SKG-FAT), leveraging naturally generated knowledge during training to enhance the adversarial robustness without compromising training efficiency. Extensive experiments on four standard datasets demonstrate that the SKG-FAT improves the robustness and preserves competitive clean accuracy, outperforming the state-of-the-art methods. Code and checkpoints are available at SFG-FAT Code Implementation. Chengze Jiang, Minjing Dong, Jie Gui, Xinli Shi, Yuan Cao 0005, Yuan Yan Tang, James T. Kwok |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2025 | ColorVein: Colorful Cancelable Vein BiometricsabstractVein recognition technologies have become one of the primary solutions for high-security identification systems. However, the issue of biometric information leakage can still pose a serious threat to user privacy and anonymity. Currently, there is no cancelable biometric template generation scheme specifically designed for vein biometrics. Therefore, this paper proposes an innovative cancelable vein biometric generation scheme: ColorVein. Unlike previous cancelable template generation schemes, ColorVein does not destroy the original biometric features and introduces additional color information to grayscale vein images. This method significantly enhances the information density of vein images by transforming static grayscale information into dynamically controllable color representations through interactive colorization. ColorVein allows users/administrators to define a controllable pseudo-random color space for grayscale vein images by editing the position, number, and color of hint points, thereby generating protected cancelable templates. Additionally, we propose a new secure center loss to optimize the training process of the protected feature extraction model, effectively increasing the feature distance between enrolled users and any potential impostors. Finally, we evaluate ColorVein’s performance on all types of vein biometrics, including recognition performance, unlinkability, irreversibility, and revocability, and conduct security and privacy analyses. ColorVein achieves competitive performance compared with state-of-the-art methods. Yifan Wang 0036, Jie Gui, Xinli Shi, Linqing Gui, Yuan Yan Tang, James T. Kwok |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2025 | Divide and Conquer: Frequency-Aware Contrastive Adversarial Training for Robust Point Cloud ClassificationabstractContrastive adversarial training has shown great potential in enhancing model robustness and has been adopted in point cloud classification. There are varying spatial distributions and densities across different regions in point cloud data, which makes adversarial perturbations always exhibit non-uniform patterns of attack intensity and distribution in different regions. However, existing approaches always rely on uniform feature contrast without considering the granularity in the context of point cloud data, limiting their capacities to counter adversarial perturbations effectively. To address this issue, we propose a novel frequency-aware contrastive adversarial training framework, which considers feature contrast via a “divide-and-conquer” method. Specifically, we systematically “divide” point clouds into distinct frequency components and “conquer” feature contrast within each frequency band, which fosters fine-grained feature consistency learning and leads to more informative as well as robust representations. Besides, existing methods typically apply group-level contrastive learning, which emphasizes category-wise similarity but often overlooks the nuanced structural variations among instances. To remedy this, we incorporate instance-level contrastive learning to capture per-instance geometric variations. Moreover, a frequency-specific hard-masked sample generation module is designed to construct challenging sample pairs by masking keypoint features in each frequency band, thereby promoting the model to learn more robust feature representations. Extensive experiments on multiple benchmark datasets demonstrate that our proposed method significantly outperforms existing state-of-the-art approaches in adversarial robustness for point cloud classification. The code is available on DiCon-FAT. Yu-Xin Zhang 0004, Jie Gui, Minjing Dong, Xiaofeng Cong, Yuan Cao 0005, Xin Gong 0001, Yuan Yan Tang, James T. Kwok |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2025 | Exploring the Coordination of Frequency and Attention in Masked Image ModelingabstractRecently, masked image modeling (MIM), which learns visual representations by reconstructing the masked patches of an image, has become a popular self-supervised paradigm. However, the pre-training of MIM always takes massive time due to the large-scale data and large-size backbones. We mainly attribute it to the random patch masking in previous MIM works, which fails to leverage the crucial semantic information for effective visual representation learning. To tackle this issue, we propose the Frequency & Attention-driven Masking and Throwing Strategy (FAMT), which can detect semantic patches and reduce the number of training patches to boost model performance and training efficiency simultaneously. Specifically, FAMT utilizes the self-attention mechanism to extract semantic information from the image for masking during training in an unsupervised manner. However, attention alone could sometimes focus on inappropriate areas regarding the semantic information. Thus, we are motivated to incorporate the information from the frequency domain into the self-attention mechanism to derive the sampling weights for masking, which captures semantic patches for visual representation learning. Furthermore, we introduce a patch throwing strategy based on the derived sampling weights to reduce the training cost. FAMT can be seamlessly integrated as a plug-and-play module and surpasses previous works, e.g. reducing the training phase time by nearly 50% and improving the linear probing accuracy of MAE by $1.8$ % ~ $ 6.3$ % across various datasets, including CIFAR-10/100, Tiny ImageNet, and ImageNet-1K. FAMT also demonstrates superior performance in downstream detection and segmentation tasks. Jie Gui, Tuo Chen, Minjing Dong, Zhengqi Liu, Hao Luo 0004, James T. Kwok, Yuan Yan Tang |
IEEE Trans. Image Process. | 1 |
| 2025 | Unrevealed Threats: Adversarial Robustness Analysis of Underwater Image Enhancement ModelsabstractLearning-based methods for underwater image enhancement (UWIE) have undergone extensive exploration. However, learning-based models are usually vulnerable to adversarial examples so as the UWIE models. To the best of our knowledge, there is no comprehensive study on the adversarial robustness of UWIE models, which indicates that UWIE models are potentially under the threat of adversarial attacks. In this paper, we propose a general adversarial attack protocol. We make a first attempt to conduct adversarial attacks on five well-designed UWIE models on three common underwater image benchmark datasets. Considering the scattering and absorption of light in the underwater environment, there exists a strong correlation between color correction and underwater image enhancement. On the basis of that, we also design two effective UWIE-oriented adversarial attack methods, Pixel Attack and Color Shift Attack targeting different color spaces. The results show that five models exhibit varying degrees of vulnerability to adversarial attacks and well-designed small perturbations on degraded images are capable of preventing UWIE models from generating enhanced results. In addition, we conduct adversarial training on these models and successfully mitigated the effectiveness of adversarial attacks. In summary, we reveal the adversarial vulnerability of UWIE models and propose a new evaluation dimension of UWIE models. Siyu Zhai, Zhibo He, Xiaofeng Cong, Junming Hou, Jie Gui, Jian Wei You, Xin Gong 0001, James T. Kwok, Yuan Yan Tang |
IEEE Trans. Multim. | 5 |
| 2024 | Underwater Organism Color Fine-Tuning via Decomposition and GuidanceabstractDue to the wavelength dependent light attenuation and scattering, the color of the underwater organism usually appears distorted. The existing underwater image enhancement methods mainly focus on designing networks capable of generating enhanced underwater organisms with fixed color. Due to the complexity of the underwater environment, ground truth labels are difficult to obtain, which results in the non-existence of perfect enhancement effects. Different from the existing methods, this paper proposes an algorithm with color enhancement and color fine-tuning (CECF) capabilities. The color enhancement behavior of CECF is the same as that of existing methods, aiming to restore the color of the distorted underwater organism. Beyond this general purpose, the color fine-tuning behavior of CECF can adjust the color of organisms in a controlled manner, which can generate enhanced organisms with diverse colors. To achieve this purpose, four processes are used in CECF. A supervised enhancement process learns the mapping from a distorted image to an enhanced image by the decomposition of color code. A self reconstruction process and a cross-reconstruction process are used for content-invariant learning. A color fine-tuning process is designed based on the guidance for obtaining various enhanced results with different colors. Experimental results have proven the enhancement ability and color fine-tuning ability of the proposed CECF. The source code is provided in https://github.com/Xiaofeng-life/CECF. Xiaofeng Cong, Jie Gui, Junming Hou |
AAAI | 2 |
| 2024 | Taxonomy Driven Fast Adversarial TrainingabstractAdversarial training (AT) is an effective defense method against gradient-based attacks to enhance the robustness of neural networks. Among them, single-step AT has emerged as a hotspot topic due to its simplicity and efficiency, requiring only one gradient propagation in generating adversarial examples. Nonetheless, the problem of catastrophic overfitting (CO) that causes training collapse remains poorly understood, and there exists a gap between the robust accuracy achieved through single- and multi-step AT. In this paper, we present a surprising finding that the taxonomy of adversarial examples reveals the truth of CO. Based on this conclusion, we propose taxonomy driven fast adversarial training (TDAT) which jointly optimizes learning objective, loss function, and initialization method, thereby can be regarded as a new paradigm of single-step AT. Compared with other fast AT methods, TDAT can boost the robustness of neural networks, alleviate the influence of misclassified examples, and prevent CO during the training process while requiring almost no additional computational and memory resources. Our method achieves robust accuracy improvement of 1.59%, 1.62%, 0.71%, and 1.26% on CIFAR-10, CIFAR-100, Tiny ImageNet, and ImageNet-100 datasets, when against projected gradient descent PGD10 attack with perturbation budget 8/255. Furthermore, our proposed method also achieves state-of-the-art robust accuracy against other attacks. Code is available at https://github.com/bookman233/TDAT. Kun Tong, Chengze Jiang, Jie Gui, Yuan Cao 0005 |
AAAI | 3 |
| 2024 | A Semi-Supervised Nighttime Dehazing Baseline with Spatial-Frequency Aware and Realistic Brightness ConstraintabstractExisting research based on deep learning has extensively explored the problem of daytime image dehazing. However, few studies have considered the characteristics of nighttime hazy scenes. There are two distinctions between nighttime and daytime haze. First, there may be multiple active col-ored light sources with lower illumination intensity in night-time scenes, which may cause haze, glow and noise with localized, coupled and frequency inconsistent characteris-tics. Second, due to the domain discrepancy between simulated and real-world data, unrealistic brightness may occur when applying a dehazing model trained on simulated data to real-world data. To address the above two issues, we propose a semi-supervised model for real-world nighttime dehazing. First, the spatial attention and frequency spectrum filtering are implemented as a spatial-frequency do-main information interaction module to handle the first is-sue. Second, a pseudo-label-based retraining strategy and a local window-based brightness loss for semi-supervised training process is designed to suppress haze and glow while achieving realistic brightness. Experiments on public benchmarks validate the effectiveness of the proposed method and its superiority over state-of-the-art methods. The source code and Supplementary Materials are placed in the https://github.com/Xiaofeng-life/SFSNiD. Xiaofeng Cong, Jie Gui, Jing Zhang 0037, Junming Hou, Hao Shen 0006 |
CVPR | 2 |
| 2024 | A Comprehensive Survey and Taxonomy on Point Cloud Registration Based on Deep Learning
Yu-Xin Zhang 0004, Jie Gui, Xiaofeng Cong, Xin Gong 0001, Wenbing Tao |
IJCAI | 2 |
| 2024 | MDNet: Morphology-Driven Weakly Supervised Polyp Detection
Jiajia Chen 0006, Jie Gui, Xiuquan Du, Wen Sha |
PRCV (15) | 3 |
| 2024 | MFIS-Net: A Deep Learning Framework for Left Atrial Segmentation
Jie Gui, Wen Sha, Xiuquan Du |
PRCV (15) | 1 |
| 2024 | GCNet: Global Context-Guided Uncertainty Boundary for Polyp Segmentation
Jiajia Chen 0006, Jie Gui, Xiuquan Du, Wen Sha |
PRCV (15) | 3 |
| 2024 | Region-aware image-based human action retrieval with transformers
Hongsong Wang 0001, Jie Gui |
Comput. Vis. Image Underst. | 3 |
| 2024 | HyperComm: Hypergraph-based communication in multi-agent reinforcement learning
Xinli Shi, Xiangping Xu, Jie Gui, Jinde Cao |
Neural Networks | 4 |
| 2024 | A Survey on Self-Supervised Learning: Algorithms, Applications, and Future TrendsabstractDeep supervised learning algorithms typically require a large volume of labeled data to achieve satisfactory performance. However, the process of collecting and labeling such data can be expensive and time-consuming. Self-supervised learning (SSL), a subset of unsupervised learning, aims to learn discriminative features from unlabeled data without relying on human-annotated labels. SSL has garnered significant attention recently, leading to the development of numerous related algorithms. However, there is a dearth of comprehensive studies that elucidate the connections and evolution of different SSL variants. This paper presents a review of diverse SSL methods, encompassing algorithmic aspects, application domains, three key trends, and open research questions. First, we provide a detailed introduction to the motivations behind most SSL algorithms and compare their commonalities and differences. Second, we explore representative applications of SSL in domains such as image processing, computer vision, and natural language processing. Lastly, we discuss the three primary trends observed in SSL research and highlight the open questions that remain. Jie Gui, Tuo Chen, Jing Zhang 0037, Qiong Cao, Zhenan Sun, Hao Luo 0004, Dacheng Tao |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | From Simple to Complex Scenes: Learning Robust Feature Representations for Accurate Human ParsingabstractHuman parsing has attracted considerable research interest due to its broad potential applications in the computer vision community. In this paper, we explore several useful properties, including high-resolution representation, auxiliary guidance, and model robustness, which collectively contribute to a novel method for accurate human parsing in both simple and complex scenes. Starting from simple scenes: we propose the boundary-aware hybrid resolution network (BHRN), an advanced human parsing network. BHRN utilizes deconvolutional layers and multi-scale supervision to generate rich high-resolution representations. Additionally, it includes an edge perceiving branch designed to enhance the fineness of part boundaries. Building on BHRN, we construct a dual-task mutual learning (DTML) framework. It not only provides implicit guidance to assist the parser by incorporating boundary features, but also explicitly maintains the high-order consistency between the parsing prediction and the ground truth. Toward complex scenes: we develop a domain transform method to enhance the model robustness. By transforming the input space from the spatial domain to the polar harmonic Fourier moment domain, the mapping relationship to the output semantic space is highly stable. This transformation yields robust representations for both clean and corrupted data. When evaluated on standard benchmark datasets, our method achieves superior performance compared to state-of-the-art human parsing methods. Furthermore, our domain transform strategy significantly improves the robustness of DTML dramatically in most complex scenes. Yunan Liu 0001, Chunpeng Wang 0001, Mingyu Lu, Jian Yang 0003, Jie Gui, Shanshan Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | Response Generation in Social Network With Topic and Emotion ConstraintsabstractResponse generation is the task of automatically generating human-like content based on the provided context. One of its prominent applications is to simulate realistic response content for social network posts. In the digital age, social network platforms play a vital role in information exchange and social interaction. This study focuses on response generation techniques for the platform of public opinion evolution simulation that simulate realistic response content, enabling a deeper understanding of the emotional expressions of network users. Recent advancements in deep learning techniques, particularly the sequence-to-sequence (Seq2Seq) model, have shown promise in the response generation field. However, we still face two challenges: content variety, topic and emotion relevancy. To this end, we propose the EmoTG-ETRS model which comprises three parts. The first is a response generation module based on Transformer architecture. Then, an auxiliary emotion improvement module is incorporated to enhance the emotional expressiveness of the response candidates. Finally, a reverse selection module, which combines maximum mutual information (MMI) evaluation, emotional expression evaluation, and topic consistency evaluation, is devised to select the highest-scoring response. Extensive experiments have been conducted to evaluate the effectiveness of the proposed model and the results demonstrate that the EmoTG-ETRS model improves the quality of produced replies in terms of topic consistency and emotional accuracy rate when compared with the SOTA research works. Biwei Cao, Jiuxin Cao, Bo Liu 0004, Jie Gui, Jun Zhou 0027, Yuan Yan Tang, James T. Kwok |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2024 | Fooling the Image Dehazing Models by First Order GradientabstractThe research on the single image dehazing task has been widely explored. However, as far as we know, no comprehensive study has been conducted on the robustness of the well-trained dehazing models. Therefore, there is no evidence that the dehazing networks can resist malicious attacks. In this paper, we focus on designing a group of attack methods based on first order gradient to verify the robustness of the existing dehazing algorithms. By analyzing the general purpose of image dehazing task, four attack methods are proposed, which are predicted dehazed image attack, hazy layer mask attack, haze-free image attack and haze-preserved attack. The corresponding experiments are conducted on six datasets with different scales. Further, the defense strategy based on adversarial training is adopted for reducing the negative effects caused by malicious attacks. In summary, this paper defines a new challenging problem for the image dehazing area, which can be called as adversarial attack on dehazing networks (AADN). Code is available at https://github.com/Xiaofeng-life/AADN_Dehazing. Jie Gui, Xiaofeng Cong, Chengwei Peng, Yuan Yan Tang, James T. Kwok |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | CFVNet: An End-to-End Cancelable Finger Vein Network for RecognitionabstractFinger vein recognition technology has become one of the primary solutions for high-security identification systems. However, it still has information leakage problems, which seriously jeopardizes user’s privacy and anonymity and cause great security risks. In addition, there is no work to consider a fully integrated secure finger vein recognition system. So, different from the previous systems, we integrate preprocessing and template protection into an integrated deep learning model. We propose an end-to-end cancelable finger vein network (CFVNet), which can be used to design an secure finger vein recognition system. It includes a plug-and-play BWR-ROIAlign unit, which consists of three sub-modules: Localization, Compression and Transformation. The localization module achieves automated localization of stable and unique finger vein ROI. The compression module losslessly removes spatial and channel redundancies. The transformation module uses the proposed BWR method to introduce unlinkability, irreversibility and revocability to the system. BWR-ROIAlign can directly plug into the model to introduce the above features for DCNN-based finger vein recognition systems. We perform extensive experiments on four public datasets to study the performance and cancelable biometric attributes of the CFVNet-based recognition system. The average accuracy, EERs and$D_{\leftrightarrow } ^{sys}$on the four datasets are 99.82%, 0.01% and 0.025, respectively, and achieves competitive performance compared with the state-of-the-arts. Yifan Wang 0036, Jie Gui, Yuan Yan Tang, James T. Kwok |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2024 | Constructing Diverse Inlier Consistency for Partial Point Cloud RegistrationabstractPartial point cloud registration aims to align partial scans into a shared coordinate system. While learning-based partial point cloud registration methods have achieved remarkable progress, they often fail to take full advantage of the relative positional relationships both within (intra-) and between (inter-) point clouds. This oversight hampers their ability to accurately identify overlapping regions and search for reliable correspondences. To address these limitations, a diverse inlier consistency (DIC) method has been proposed that adaptively embeds the positional information of a reliable correspondence in the intra- and inter-point cloud. Firstly, a diverse inlier consistency-driven region perception (DICdRP) module is devised, which encodes the positional information of the selected correspondence within the intra-point cloud. This module enhances the sensitivity of all points to overlapping regions by recognizing the position of the selected correspondence. Secondly, a diverse inlier consistency-aware correspondence search (DICaCS) module is developed, which leverages relative positions in the inter-point cloud. This module studies an inter-point cloud DIC weight to supervise correspondence compatibility, allowing for precise identification of correspondences and effective outlier filtration. Thirdly, diverse information is integrated throughout our framework to achieve a more holistic and detailed registration process. Extensive experiments on object-level and scene-level datasets demonstrate the superior performance of the proposed algorithm. The code is available at https://github.com/yxzhang15/DIC. Yu-Xin Zhang 0004, Jie Gui, James T. Kwok |
IEEE Trans. Image Process. | 2 |
| 2024 | Illumination Controllable Dehazing Network based on Unsupervised Retinex EmbeddingabstractOn the one hand, the dehazing task is an ill-posedness problem, which means that no unique solution exists. On the other hand, the dehazing task should take into account the subjective factor, which is to give the user selectable dehazed images rather than a single result. Therefore, this paper proposes a multi-output dehazing network by introducing illumination controllable ability, called IC-Dehazing. The proposed IC-Dehazing can change the illumination intensity by adjusting the factor of the illumination controllable module, which is realized based on the interpretable Retinex model. Moreover, the backbone dehazing network of IC-Dehazing consists of a Transformer with double decoders for high-quality image restoration. Further, the prior-based loss function and unsupervised training strategy enable IC-Dehazing to complete the parameter learning process without the need for paired data. To demonstrate the effectiveness of the proposed IC-Dehazing, quantitative and qualitative experiments are conducted. Code is available athttps://github.com/Xiaofeng-life/ICDehazing. Jie Gui, Xiaofeng Cong, Yuan Yan Tang, James T. Kwok |
IEEE Trans. Multim. | 1 |
| 2023 | Fast Online Hashing with Multi-Label ProjectionabstractHashing has been widely researched to solve the large-scale approximate nearest neighbor search problem owing to its time and storage superiority. In recent years, a number of online hashing methods have emerged, which can update the hash functions to adapt to the new stream data and realize dynamic retrieval. However, existing online hashing methods are required to update the whole database with the latest hash functions when a query arrives, which leads to low retrieval efficiency with the continuous increase of the stream data. On the other hand, these methods ignore the supervision relationship among the examples, especially in the multi-label case. In this paper, we propose a novel Fast Online Hashing (FOH) method which only updates the binary codes of a small part of the database. To be specific, we first build a query pool in which the nearest neighbors of each central point are recorded. When a new query arrives, only the binary codes of the corresponding potential neighbors are updated. In addition, we create a similarity matrix which takes the multi-label supervision information into account and bring in the multi-label projection loss to further preserve the similarity among the multi-label data. The experimental results on two common benchmarks show that the proposed FOH can achieve dramatic superiority on query time up to 6.28 seconds less than state-of-the-art baselines with competitive retrieval accuracy. Wenzhe Jia, Yuan Cao 0005, Jie Gui |
AAAI | 4 |
| 2023 | Good Helper Is around You: Attention-Driven Masked Image ModelingabstractIt has been witnessed that masked image modeling (MIM) has shown a huge potential in self-supervised learning in the past year. Benefiting from the universal backbone vision transformer, MIM learns self-supervised visual representations through masking a part of patches of the image while attempting to recover the missing pixels. Most previous works mask patches of the image randomly, which underutilizes the semantic information that is beneficial to visual representation learning. On the other hand, due to the large size of the backbone, most previous works have to spend much time on pre-training. In this paper, we propose Attention-driven Masking and Throwing Strategy (AMT), which could solve both problems above. We first leverage the self-attention mechanism to obtain the semantic information of the image during the training process automatically without using any supervised methods. Masking strategy can be guided by that information to mask areas selectively, which is helpful for representation learning. Moreover, a redundant patch throwing strategy is proposed, which makes learning more efficient. As a plug-and-play module for masked image modeling, AMT improves the linear probing accuracy of MAE by 2.9% ~ 5.9% on CIFAR-10/100, STL-10, Tiny ImageNet, and ImageNet-1K, and obtains an improved performance with respect to fine-tuning accuracy of MAE and SimMIM. Moreover, this design also achieves superior performance on downstream detection and segmentation tasks. Zhengqi Liu, Jie Gui, Hao Luo 0004 |
AAAI | 2 |
| 2023 | SK-Gradient: Efficient Communication for Distributed Machine Learning with Data SketchabstractWith the explosive growth of data volume, distributed machine learning has become the mainstream approach for training deep neural networks. However, distributed machine learning incurs non-trivial communication overhead. To this end, various compression schemes are proposed to alleviate the communication volume among nodes. Nevertheless, existing compression schemes, such as gradient quantization or gradient sparsification, suffer from low compression ratios and/or high computational overheads. Recent studies advocate leveraging sketch techniques to assist these schemes. However, the limitations of gradient quantization and gradient sparsification remain. In this paper, we propose SK-Gradient, a novel gradient compression scheme that solely builds on sketch. The core component of SK-Gradient is a novel sketch namely FGC Sketch that is tailored for gradient compression. FGC Sketch precomputes the costly hash functions to alleviate computational overheads. Its simplified design makes it convenient for GPU acceleration. In addition, SK-Gradient leverages various techniques including selective gradient compression and periodic synchronization strategy to improve computational efficiency and compression accuracy. Compared with the state-of-the-art schemes, SK-Gradient achieves up to 92.9% reduction in computational overhead and up to 95.2% improvement in training speedups at the same compression ratio. Jie Gui, Zezhou Wang, Chenhong He, Qun Huang 0001 |
ICDE | 1 |
| 2023 | Occluded Skeleton-Based Human Action Recognition with Dual Inhibition TrainingabstractRecently, skeleton-based human action recognition has received widespread attention in computer vision community. However, most existing research focuses on improving the recognition accuracy on complete skeleton data, while ignoring the performance on the incomplete skeleton data with occlusion or noise. This paper addresses occluded and noise-robust skeleton-based action recognition and presents a novel Dual Inhibition Training strategy. Specifically, we propose Part-aware and Dual-inhibition Graph Convolutional Network (PDGCN), which comprises of three parts: Input Skeleton Inhibition (ISI), Part-Aware Representation Learning (PARL) and Predicted Score Inhibition (PSI). The ISI and PSI are plug and play modules which could encourage the model to learn discriminative features from diversified body joints by effectively simulating key body part occlusions and random occlusions. The PARL module learns both the global and local representations from the whole body and body parts, respectively, and progressively fuses them during representation learning to enhance the model robustness under occlusions. Finally, we design different settings for occluded skeleton-based human action recognition to deep study this problem and better evaluate different approaches. Our approach achieves state-of-the-art results on different benchmarks and dramatically outperforms the recent skeleton-based action recognition approaches, especially under large-scale temporal occlusion. Zhenjie Chen, Hongsong Wang 0001, Jie Gui |
ACM Multimedia | 3 |
| 2023 | OmniWindow: A General and Efficient Window Mechanism Framework for Network TelemetryabstractRecent network telemetry solutions typically target programmable switches to achieve high performance and in-network visibility. They partition the packet stream into windows and then apply various stream processing techniques to summarize flow-level statistics. However, existing studies focus on the measurement within each window. Window management is still a missing piece due to the resource limitation of programmable switches. In this paper, we propose OmniWindow, a general and efficient window mechanism framework. OmniWindow splits the original window into fine-grained sub-windows such that the sub-windows can be merged into various window types. To deal with the resource restriction, OmniWindow carefully designs its data plane memory layout and proposes a window synchronization method. It also employs a collaborative architecture that can collect and reset stateful data in sub-windows within a limited time. We prototype OmniWindow on Tofino. We incorporate OmniWindow into a SOTA query-driven telemetry system and eight sketch-based telemetry algorithms. Our experiments demonstrate that OmniWindow enables these telemetry solutions to achieve higher accuracy than conventional window mechanism. Haifeng Sun 0004, Jintao He, Jie Gui, Qun Huang 0001 |
SIGCOMM | 4 |
| 2023 | Anonymous Edge Representation for Inductive Anomaly Detection in Dynamic Bipartite GraphsabstractThe activities in many real-world applications, such as e-commerce and online education, are usually modeled as a dynamic bipartite graph that evolves over time. It is a critical task to detect anomalies inductively in a dynamic bipartite graph. Previous approaches either focus on detecting pre-defined types of anomalies or cannot handle nodes that are unseen during the training stage. To address this challenge, we propose an effective method to learn anonymous edge representation (AER) that captures the characteristics of an edge without using identity information. We further propose a model named AER-AD to utilize AER to detect anomalies in dynamic bipartite graphs in an inductive setting. Extensive experiments on both real-life and synthetic datasets are conducted to illustrate that AER-AD outperforms state-of-the-art baselines. In terms of AUC and F1, AER-AD is able to achieve 8.38% and 14.98% higher results than the best inductive representation baselines, and 6.99% and 19.59% than the best anomaly detection baselines. Lanting Fang, Kaiyu Feng, Jie Gui, Shanshan Feng 0001, Aiqun Hu |
Proc. VLDB Endow. | 3 |
| 2023 | Feedback Pyramid Attention Networks for Single Image Super-ResolutionabstractRecently, convolutional neural network (CNN) based image super-resolution (SR) methods have achieved significant performance improvement. However, most CNN-based methods mainly focus on feed-forward architecture design and neglect to explore the feedback mechanism, which usually exists in the human visual system. In this paper, we propose feedback pyramid attention networks (FPAN) to fully exploit the mutual dependencies of features. Specifically, a novel feedback connection structure is developed to enhance low-level feature expression with high-level information. In our method, the output of each layer in the first stage is also used as the input of the corresponding layer in the next state to re-update the previous low-level filters. Moreover, we introduce a pyramid non-local structure to model global contextual information in different scales and improve the discriminative representation of the network. Extensive experimental results on various datasets demonstrate the superiority of our FPAN in comparison with the state-of-the-art SR methods. Huapeng Wu, Jie Gui, Jun Zhang 0024, James T. Kwok, Zhihui Wei |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Learning the Relation Between Similarity Loss and Clustering Loss in Self-Supervised LearningabstractSelf-supervised learning enables networks to learn discriminative features from massive data itself. Most state-of-the-art methods maximize the similarity between two augmentations of one image based on contrastive learning. By utilizing the consistency of two augmentations, the burden of manual annotations can be freed. Contrastive learning exploits instance-level information to learn robust features. However, the learned information is probably confined to different views of the same instance. In this paper, we attempt to leverage the similarity between two distinct images to boost representation in self-supervised learning. In contrast to instance-level information, the similarity between two distinct images may provide more useful information. Besides, we analyze the relation between similarity loss and feature-level cross-entropy loss. These two losses are essential for most deep learning methods. However, the relation between these two losses is not clear. Similarity loss helps obtain instance-level representation, while feature-level cross-entropy loss helps mine the similarity between two distinct images. We provide theoretical analyses and experiments to show that a suitable combination of these two losses can get state-of-the-art results. Code is available at https://github.com/guijiejie/ICCL. Jidong Ge, Jie Gui, Lanting Fang, Ming Lin 0002, James T. Kwok, LiGuo Huang, Bin Luo 0003 |
IEEE Trans. Image Process. | 3 |
| 2023 | A Review on Generative Adversarial Networks: Algorithms, Theory, and ApplicationsabstractGenerative adversarial networks (GANs) have recently become a hot research topic; however, they have been studied since 2014, and a large number of algorithms have been proposed. Nevertheless, few comprehensive studies explain the connections among different GAN variants and how they have evolved. In this paper, we attempt to provide a review of the various GAN methods from the perspectives of algorithms, theory, and applications. First, the motivations, mathematical representations, and structures of most GAN algorithms are introduced in detail, and we compare their commonalities and differences. Second, theoretical issues related to GANs are investigated. Finally, typical applications of GANs in image processing and computer vision, natural language processing, music, speech and audio, the medical field, and data science are discussed. Jie Gui, Zhenan Sun, Yonggang Wen 0001, Dacheng Tao, Jieping Ye |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2023 | AlignVE: Visual Entailment Recognition Based on Alignment RelationsabstractVisual entailment (VE) is to recognize whether the semantics of a hypothesis text can be inferred from the given premise image, which is one special task among recent emerged vision and language understanding tasks. Currently, most of the existing VE approaches are derived from the methods of visual question answering. They recognize visual entailment by quantifying the similarity between the hypothesis and premise in the content semantic features from multi modalities. Such approaches, however, ignore the VE's unique nature of relation inference between the premise and hypothesis. Therefore, in this paper, a new architecture called AlignVE is proposed to solve the visual entailment problem with a relation interaction method. It models the relation between the premise and hypothesis as an alignment matrix. Then it introduces a pooling operation to get feature vectors with a fixed size. Finally, it goes through the fully-connected layer and normalization layer to complete the classification. Experiments show that our alignment-based architecture reaches 72.45% accuracy on SNLI-VE dataset, outperforming previous content-based models under the same settings. Biwei Cao, Jiuxin Cao, Jie Gui, Jiayun Shen, Bo Liu 0004, Yuan Yan Tang, James T. Kwok |
IEEE Trans. Multim. | 3 |
| 2023 | Noah: Reinforcement-Learning-Based Rate Limiter for Microservices in Large-Scale E-Commerce ServicesabstractModern large-scale online service providers typically deploy microservices into containers to achieve flexible service management. One critical problem in such container-based microservice architectures is to control the arrival rate of requests in the containers to avoid containers from being overloaded. In this article, we present our experience of rate limit for the containers in Alibaba, one of the largest e-commerce services in the world. Given the highly diverse characteristics of containers in Alibaba, we point out that the existing rate limit mechanisms cannot meet our demand. Thus, we design Noah, a dynamic rate limiter that can automatically adapt to the specific characteristic of each container without human efforts. The key idea of Noah is to use deep reinforcement learning (DRL) that automatically infers the most suitable configuration for each container. To fully embrace the advantages of DRL in our context, Noah addresses two technical challenges. First, Noah uses a lightweight system monitoring mechanism to collect container status. In this way, it minimizes the monitoring overhead while ensuring a timely reaction to system load changes. Second, Noah injects synthetic extreme data when training its models. Thus, its model gains knowledge on unseen special events and hence remains highly available in extreme scenarios. To guarantee model convergence with the injected training data, Noah adopts task-specific curriculum learning to train the model from normal data to extreme data gradually. Noah has been deployed in the production of Alibaba for two years, serving more than 50000 containers and around 300 types of microservice applications. Experimental results show that Noah can well adapt to three common scenarios in the production environment. It effectively achieves better system availability and shorter request response time compared with four state-of-the-art rate limiters. Zhao Li 0007, Haifeng Sun 0004, Zheng Xiong, Qun Huang 0001, Zehong Hu, Shasha Ruan, Hai Hong, Jie Gui, Jintao He, Zebin Xu |
IEEE Trans. Neural Networks Learn. Syst. | 9 |
| 2022 | Pyramidal dense attention networks for single image super-resolutionabstractAbstract Recently, residual and dense networks have effectively promoted the development of image super‐resolution (SR). However, most dense networks based SR methods do not make full use of dense feature information. To solve this problem, a pyramidal dense attention network for single image super‐resolution is proposed in this paper. In this method, the proposed pyramidal dense learning can gradually increase the width of the densely connected layer inside a pyramidal dense block to extract deep features efficiently. Meanwhile, the adaptive group convolution that the number of groups grows linearly with dense convolutional layers is introduced to relieve the parameter explosion. Besides, a novel joint attention to capture cross‐dimension interaction between the spatial dimensions and channel dimension in an efficient way for providing rich discriminative feature representations is also proposed. Extensive experimental results show that the method achieves comparable performance in comparison with the state‐of‐the‐art SR methods. Huapeng Wu, Jie Gui, Jun Zhang 0024, James T. Kwok, Zhihui Wei |
IET Image Process. | 2 |
| 2022 | Deep human answer understanding for natural reverse QA
Rujing Yao, Linlin Hou, Jie Gui, Ou Wu 0001 |
Knowl. Based Syst. | 4 |
| 2022 | Hash Learning With Variable Quantization for Large-Scale RetrievalabstractApproximate Nearest Neighbor(ANN) search is the core problem in many large-scale machine learning and computer vision applications such as multimodal retrieval. Hashing is becoming increasingly popular, since it can provide efficient similarity search and compact data representations suitable for handling such large-scale ANN search problems. Most hashing algorithms concentrate on learning more effective projection functions. However, the accuracy loss in the quantization step has been ignored and barely studied. In this paper, we analyse the importance of various projected dimensions, distribute them into several groups and quantize them with two types of values which can both better preserve the neighborhood structure among data. One is Variable Integer-based Quantization (VIQ) that quantizes each projected dimension with integer values. The other is Variable Codebook-based Quantization (VCQ) that quantizes each projected dimension with corresponding codebook values. We conduct experiments on five common public data sets containing up to one million vectors. The results show that the proposed VCQ and VIQ algorithms can both achieve much higher accuracy than state-of-the-art quantization methods. Furthermore, although VCQ performs better than VIQ, ANN search with VIQ provides much higher search efficiency. Yuan Cao 0005, Sheng Chen 0015, Jie Gui, Heng Qi, Zhiyang Li 0001, Chao Liu 0008 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | An Ensemble Framework for Improving the Prediction of Deleterious Synonymous MutationabstractIn recent years, the association between synonymous mutations (SMs) and human diseases has been uncovered in many studies. It is a challenge for identifying deleterious SMs in the field of medical genomics. Although there are several computational methods proposed in the past years, the precise prediction of deleterious SMs is still challenging. In this work, we proposed a predictor named as EnDSM, which is an accurate method based on the ensemble framework. We explored multimodal features across four groups including functional score, conservation, splicing, and sequence features, and we then trained eight conceptually different machine learning classifiers for each of them, resulting in 32 base classification models. We further selected four base models referring to their prediction performance and the predictive probabilities of these base classification models were subsequently used as the input feature vectors of logistic regression classifier to construct the ensemble learning model. The results suggested that EnDSM achieved better performance comparing with other state-of-the-art predictors on the training and independent test datasets. We anticipate that our ensemble predictor EnDSM will become a valuable tool for deleterious SM prediction.The EnDSM server interface along with the benchmarking data sets are freely available athttp://bioinfo.ahu.edu.cn/EnDSM. Jie Gui, Chun-Hou Zheng 0001, Junfeng Xia |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | An Efficient Cross-Modality Self-Calibrated Network for Hyperspectral and Multispectral Image FusionabstractRecently, deep convolutional neural network based hyperspectral and multispectral image fusion methods have shown significant performance. Nevertheless, the rich spatial and spectral details of hyperspectral images (HSIs) have not been fully explored, leaving room for further improve the representation ability of the model. In this paper, we propose an efficient cross-modality self-calibrated network (CMSCN) for hyperspectral and multispectral image fusion. Specifically, we use a cross-modality non-local module to fuse a high-resolution multispectral image (HR-MSI) and a low-resolution hyperspectral image (LR-HSI) to get an enhanced LR-HSI. In addition, a novel cross-scale self-calibrated convolution structure is proposed to explore and exploit multi-scale and hierarchical spatial-spectral features, which can improve the learning ability of the model. The introduced efficient spatial-spectral attention mechanism can calibrate the feature representation at different dimensions, thereby providing more efficient and accurate information for hyperspectral image reconstruction. Extensive experimental results on various hyperspectral images demonstrate the superiority of our method in comparison with the state-of-the-art image fusion methods. Huapeng Wu, Jie Gui, Yang Xu 0006, Zebin Wu 0001, Yuan Yan Tang, Zhihui Wei |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Scalable Distributed Hashing for Approximate Nearest Neighbor SearchabstractHashing has been widely applied to the large-scale approximate nearest neighbor search problem owing to its high efficiency and low storage requirement. Most investigations concentrate on learning hashing methods in a centralized setting. However, in existing big data systems, data is often stored across different nodes. In some situations, data is even collected in a distributed manner. A straightforward way to solve this problem is to aggregate all the data into the fusion center to obtain the search result (aggregating method). However, this strategy is not feasible because of the prohibitive communication cost. Although a few distributed hashing methods have been proposed to reduce this cost, they only focus on designing a distributed algorithm for a specific global optimization objective without considering scalability. Moreover, existing distributed hashing methods aim at finding a distributed solution to hashing, meanwhile avoiding accuracy loss, rather than improving accuracy. To address these challenges, we propose a Scalable Distributed Hashing (SDisH) model in which most existing hashing methods can be extended to process distributed data with no changes. Furthermore, to improve accuracy, we utilize the search radius as a global variable across different nodes to achieve a global optimum search result for every iteration. In addition, a voting algorithm is presented based on the results produced by multiple iterations to further reduce search errors. Theoretical analyses of communication, computation, and accuracy demonstrate the superiority of the proposed model. Numerical simulations on three large-scale and two relatively small benchmark datasets also show that the SDisH model achieves up to 44.75% and 10.23% accuracy gains compared to the aggregating method and state-of-the-art distributed hashing methods, respectively. Yuan Cao 0005, Heng Qi, Jie Gui, Keqiu Li, Jieping Ye, Chao Liu 0008 |
IEEE Trans. Image Process. | 4 |
| 2021 | Delving into Variance Transmission and Normalization: Shift of Average Gradient Makes the Network Collapse
Jidong Ge, Chuanyi Li, Jie Gui |
AAAI | 4 |
| 2021 | A Comprehensive Survey on Image Dehazing Based on Deep LearningabstractThe presence of haze significantly reduces the quality of images. Researchers have designed a variety of algorithms for image dehazing (ID) to restore the quality of hazy images. However, there are few studies that summarize the deep learning (DL) based dehazing technologies. In this paper, we conduct a comprehensive survey on the recent proposed dehazing methods. Firstly, we conclude the commonly used datasets, loss functions and evaluation metrics. Secondly, we group the existing researches of ID into two major categories: supervised ID and unsupervised ID. The core ideas of various influential dehazing models are introduced. Finally, the open issues for future research on ID are pointed out. Jie Gui, Xiaofeng Cong, Yuan Cao 0005, Wenqi Ren, Jun Zhang 0011, Jing Zhang 0037, Dacheng Tao |
IJCAI | 1 |
| 2021 | MapEmbed: Perfect Hashing with High Load Factor and Fast UpdateabstractPerfect hashing is a hash function that maps a set of distinct keys to a set of continuous integers without collision. However,most existing perfect hash schemes are static, which means that they cannot support incremental updates, while most datasets in practice are dynamic. To address this issue, we propose a novel hashing scheme, namely MapEmbed Hashing. Inspired by divide-and-conquer and map-and-reduce, our key idea is named map-and-embed and includes two phases: 1) Map all keys into many small virtual tables; 2) Embed all small tables into a large table by circular move. Our experimental results show that under the same experimental setting, the state-of-the-art perfect hashing (dynamic perfect hashing) can achieve around 15% load factor, around 0.3 Mops update speed, while our MapEmbed achieves around 90% ~ 95% load factor, and around 8.0 Mops update speed per thread. All codes of ours and other algorithms are open-sourced at GitHub. Yuhan Wu 0001, Zirui Liu 0002, Jie Gui, Haochen Gan, Yuhao Han, Tao Li 0008, Ori Rottenstreich, Tong Yang 0003 |
KDD | 4 |
| 2021 | Learning rates for multi-task regularization networks
Jie Gui, Haizhang Zhang |
Neurocomputing | 1 |
| 2021 | Multi-Grained Attention Networks for Single Image Super-ResolutionabstractDeep Convolutional Neural Networks (CNN) have drawn great attention in image super-resolution (SR). Recently, visual attention mechanism, which exploits both of the feature importance and contextual cues, has been introduced to image SR and proves to be effective to improve CNN-based SR performance. In this paper, we make a thorough investigation on the attention mechanisms in a SR model and shed light on how simple and effective improvements on these ideas improve the state-of-the-arts. We further propose a unified approach called “multi-grained attention networks (MGAN)” which fully exploits the advantages of multi-scale and attention mechanisms in SR tasks. In our method, the importance of each neuron is computed according to its surrounding regions in a multi-grained fashion and then is used to adaptively re-scale the feature responses. More importantly, the “channel attention” and “spatial attention” strategies in previous methods can be essentially considered as two special cases of our method. We also introduce multi-scale dense connections to extract the image features at multiple scales and capture the features of different layers through dense skip connections. Ablation studies on benchmark datasets demonstrate the effectiveness of our method. In comparison with other state-of-the-art SR methods, our method shows the superiority in terms of both accuracy and model size. Huapeng Wu, Zhengxia Zou, Jie Gui, Wen-Jun Zeng, Jieping Ye, Jun Zhang 0024, Hongyi Liu 0001, Zhihui Wei |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Fast kNN Search in Weighted Hamming Space With Multiple TablesabstractHashing methods have been widely used in Approximate Nearest Neighbor (ANN) search for big data due to low storage requirements and high search efficiency. These methods usually map the ANN search for big data into the k -Nearest Neighbor ( k NN) search problem in Hamming space. However, Hamming distance calculation ignores the bit-level distinction, leading to confusing ranking. In order to further increase search accuracy, various bit-level weights have been proposed to rank hash codes in weighted Hamming space. Nevertheless, existing ranking methods in weighted Hamming space are almost based on exhaustive linear scan, which is time consuming and not suitable for large datasets. Although Multi-Index hashing that is a sub-linear search method has been proposed, it relies on Hamming distance rather than weighted Hamming distance. To address this issue, we propose an exact k NN search approach with Multiple Tables in Weighted Hamming space named WHMT, in which the distribution of bit-level weights is incorporated into the multi-index building. By WHMT, we can get the optimal candidate set for exact k NN search in weighted Hamming space without exhaustive linear scan. Experimental results show that WHMT can achieve dramatic speedup up to 69.8 times over linear scan baseline without losing accuracy in weighted Hamming space. Jie Gui, Yuan Cao 0005, Heng Qi, Keqiu Li, Jieping Ye, Chao Liu 0008, Xiaowei Xu 0005 |
IEEE Trans. Image Process. | 1 |
| 2021 | Learning to Hash With Dimension Analysis Based Quantizer for Image RetrievalabstractThe last few years have witnessed the rise of the big data era in which approximate nearest neighbor search is a fundamental problem in many applications, such as large-scale image retrieval. Recently, many research results have demonstrated that hashing can achieve promising performance due to its appealing storage and search efficiency. Since complex optimization problems for loss functions are difficult to solve, most hashing methods decompose the hash code learning problem into two steps: projection and quantization. In the quantization step, binary codes are widely used because ranking them by the Hamming distance is very efficient. However, the massive information loss produced by the quantization step should be reduced in applications where high search accuracy is required, such as in image retrieval. Since many two-step hashing methods produce uneven projected dimensions in the projection step, in this paper, we propose a novel dimension analysis-based quantization (DAQ) on two-step hashing methods for image retrieval. We first perform an importance analysis of the projected dimensions and select a subset of them that are more informative than others, and then we divide the selected projected dimensions into several regions with our quantizer. Every region is quantized with its corresponding codebook. Finally, the similarity between two hash codes is estimated by the Manhattan distance between their corresponding codebooks, which is also efficient. We conduct experiments on three public benchmarks containing up to one million descriptors and show that the proposed DAQ method consistently leads to significant accuracy improvements over state-of-the-art quantization methods. Yuan Cao 0005, Heng Qi, Jie Gui, Keqiu Li, Yuan Yan Tang, James T. Kwok |
IEEE Trans. Multim. | 3 |
| 2020 | Randomized Kernel Multi-View Discriminant AnalysisabstractIn many artificial intelligence and computer vision systems, the same object can be observed at distinct viewpoints or by diverse sensors, which raises the challenges for recognizing objects from different, even heterogeneous views. Multi-view discriminant analysis (MvDA) is an effective multi-view subspace learning method, which finds a discriminant common subspace by jointly learning multiple view-specific linear projections for object recognition from multiple views, in a non-pairwise way. In this paper, we propose the kernel version of multi-view discriminant analysis, called kernel multi-view discriminant analysis (KMvDA). To overcome the well-known computational bottleneck of kernel methods, we also study the performance of using random Fourier features (RFF) to approximate Gaussian kernels in KMvDA, for large scale learning. Theoretical analysis on stability of this approximation is developed. We also conduct experiments on several popular multi-view datasets to illustrate the effectiveness of our proposed strategy. Jie Gui, Ping Li 0001 |
ECAI | 2 |
| 2020 | Discrete Haze Level Dehazing NetworkabstractIn contrast to traditional dehazing methods, deep learning based single image dehazing (SID) algorithms have achieved better performances by creating a mapping function from haze to haze-free images. Usually, the images taken from the natural scenes have different haze levels, but deep SID algorithms only process the hazy images as one group. It makes the deep SID algorithms difficult to deal with the image set with some images having specific haze density. In this paper, a Discrete Haze Level Dehazing network (DHL-Dehaze), a very effective method to dehaze multiple different haze level images, is proposed. The proposed approach considers a single image dehazing problem as a multi-domain image-to-image translation, instead of grouping all hazy images into the same domain. DHL-Dehaze provides computational derivation to describe the role of different haze levels for image translation. To verify the proposed approach, we synthesize two largescale datasets with multiple haze level images based on the NYU-Depth and DIML/CVL datasets. The experiments show that DHL-Dehaze can obtain excellent quantitative and qualitative dehazing results, especially when the haze concentration is high. Xiaofeng Cong, Jie Gui, Kai-Chao Miao, Jun Zhang 0011, Bing Wang 0004, Peng Chen 0001 |
ACM Multimedia | 2 |
| 2020 | Local Attention Networks for Occluded Airplane Detection in Remote Sensing ImagesabstractDespite the great progress of deep learning and target detection in recent years, the accurate detection of the occluded targets in remote sensing images still remains a challenge. In this letter, we propose a new detection method called local attention networks to improve the detection of occluded airplanes. Following the idea of “divide and conquer,” the proposed method is designed by first dividing an airplane target into four visual parts: head, left/right wings, body, and tail, and then considering the detection as the prediction of the individual key points in each of the visual parts. We further introduce an additional attention branch in the standard detection pipeline to enhance the features and make the model focus on individual parts of a target even if it is only partially visible in the image. Detection results and ablation studies on three remote sensing target detection data sets (including two publicly available ones) demonstrate the effectiveness of our method, especially for occluded airplane targets. In addition, our method outperforms the other state-of-the-art detection methods on these data sets. Zhengxia Zou, Zhenwei Shi 0001, Wen-Jun Zeng, Jie Gui |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2019 | General Distributed Hash Learning on Image Descriptors for $k$-Nearest Neighbor SearchabstractHashing methods have attracted much attention due to their superior time and storage properties for image retrieval. To learn similarity-preserving hash function, most existing methods are designed for the centralized setting. However, the current data storage systems are distributed to increase scalability. Obviously, it is infeasible to aggregate all the data into a fusion center because of the prohibitively expensive communication and computation overhead. Motivated by this, some methods are proposed to achieve hashing for distributed data. However, these methods mostly focus on extending one specific hashing to a distributed model without considering the generality. In this letter, we propose a novel general distributed hash learning model, which can be viewed as an effective distributed model of most hashing methods. The proposed model can achieve up to 15.2% accuracy gains over state-of-the-art distributed hashing methods, while the communication cost is independent on the data size. Yuan Cao 0005, Heng Qi, Jie Gui, Shuai Li 0002, Keqiu Li |
IEEE Signal Process. Lett. | 3 |
| 2018 | Multi-view Feature Selection for Heterogeneous Face RecognitionabstractWhile the task of feature selection has been studied for many years, the topic of multi-view feature selection for heterogeneous face recognition (HFR) such as visible (VIS) image versus near infrared (NIR) image recognition, photo versus sketch recognition, and face recognition across pose, is rarely studied. In this paper, we propose a multi-view feature selection method (MvFS) for HFR. To the best of our knowledge, MvFS is the first algorithm to address the problem of multiview feature selection for HFR, in which the dimensionalities of different views are the same and the number of selected features of different views are the same. The proposed algorithm is simple and computationally efficient. Our experiments confirm the effectiveness of MvFS. Jie Gui, Ping Li 0001 |
ICDM | 1 |
| 2018 | R 2 SDH: Robust Rotated Supervised Discrete HashingabstractLearning-based hashing has recently received considerable attentions due to its capability of supporting efficient storage and retrieval of high-dimensional data such as images, videos, and documents. In this paper, we propose a learning-based hashing algorithm called "Robust Rotated Supervised Discrete Hashing" (R 2 SDH), by extending the previous work on "Supervised Discrete Hashing" (SDH). In R 2 SDH, correntropy is adopted to replace the least square regression (LSR) model in SDH for achieving better robustness. Furthermore, considering the commonly used distance metrics such as cosine and Euclidean distance are invariant to rotational transformation, rotation is integrated into the original zero-one label matrix used in SDH, as additional freedom to promote flexibility without sacrificing accuracy. The rotation matrix is learned through an optimization procedure. Experimental results on three image datasets (MNIST, CIFAR-10, and NUS-WIDE) confirm that R 2 SDH generally outperforms SDH. Jie Gui, Ping Li 0001 |
KDD | 1 |
| 2018 | miRBaseConverter: an R/Bioconductor package for converting and retrieving miRNA name, accession, sequence and family information in different versions of miRBaseabstractBACKGROUND: miRBase is the primary repository for published miRNA sequence and annotation data, and serves as the "go-to" place for miRNA research. However, the definition and annotation of miRNAs have been changed significantly across different versions of miRBase. The changes cause inconsistency in miRNA related data between different databases and articles published at different times. Several tools have been developed for different purposes of querying and converting the information of miRNAs between different miRBase versions, but none of them individually can provide the comprehensive information about miRNAs in miRBase and users will need to use a number of different tools in their analyses. RESULTS: We introduce miRBaseConverter, an R package integrating the latest miRBase version 22 available in Bioconductor to provide a suite of functions for converting and retrieving miRNA name (ID), accession, sequence, species, version and family information in different versions of miRBase. The package is implemented in R and available under the GPL-2 license from the Bioconductor website ( http://bioconductor.org/packages/miRBaseConverter/ ). A Shiny-based GUI suitable for non-R users is also available as a standalone application from the package and also as a web application at http://nugget.unisa.edu.au:3838/miRBaseConverter . miRBaseConverter has a built-in database for querying miRNA information in all species and for both pre-mature and mature miRNAs defined by miRBase. In addition, it is the first tool for batch querying the miRNA family information. The package aims to provide a comprehensive and easy-to-use tool for miRNA research community where researchers often utilize published miRNA data from different sources. CONCLUSIONS: The Bioconductor package miRBaseConverter and the Shiny-based web application are presented to provide a suite of functions for converting and retrieving miRNA name, accession, sequence, species, version and family information in different versions of miRBase. The package will serve a wide range of applications in miRNA research and could provide a full view of the miRNAs of interest. Taosheng Xu, Lin Liu 0003, Junpeng Zhang 0001, Weijia Zhang 0001, Jie Gui, Kui Yu, Jiuyong Li, Thuc Duy Le |
BMC Bioinform. | 7 |
| 2018 | Fast Supervised Discrete HashingabstractLearning-based hashing algorithms are "hot topics" because they can greatly increase the scale at which existing methods operate. In this paper, we propose a new learning-based hashing method called "fast supervised discrete hashing" (FSDH) based on "supervised discrete hashing" (SDH). Regressing the training examples (or hash code) to the corresponding class labels is widely used in ordinary least squares regression. Rather than adopting this method, FSDH uses a very simple yet effective regression of the class labels of training examples to the corresponding hash code to accelerate the algorithm. To the best of our knowledge, this strategy has not previously been used for hashing. Traditional SDH decomposes the optimization into three sub-problems, with the most critical sub-problem - discrete optimization for binary hash codes - solved using iterative discrete cyclic coordinate descent (DCC), which is time-consuming. However, FSDH has a closed-form solution and only requires a single rather than iterative hash code-solving step, which is highly efficient. Furthermore, FSDH is usually faster than SDH for solving the projection matrix for least squares regression, making FSDH generally faster than SDH. For example, our results show that FSDH is about 12-times faster than SDH when the number of hashing bits is 128 on the CIFAR-10 data base, and FSDH is about 151-times faster than FastHash when the number of hashing bits is 64 on the MNIST data-base. Our experimental results show that FSDH is not only fast, but also outperforms other comparative methods. Jie Gui, Tongliang Liu, Zhenan Sun, Dacheng Tao, Tieniu Tan |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2018 | Velocity-Level Control With Compliance to Acceleration-Level Constraints: A Novel Scheme for Manipulator Redundancy ResolutionabstractManipulators are subject to physical constraints at different levels, i.e., joint angle limits, joint velocity limits, and acceleration limits. Effective resolution of redundant manipulators with compliance to the physical constraints is a fundamental issue for safe operation. Existing results generally resolve the manipulator redundancy either at the velocity level or the acceleration level. On the one hand, the velocity-level redundancy resolution scheme is able to deal with the joint angle and joint velocity limits successfully but cannot address the joint acceleration limit. On the other hand, although the existing acceleration-level redundancy resolution scheme is able to overcome the failure of the velocity-level one in complying with acceleration constraints, it is at the cost of making the system equation more complicated, e.g., the dependence on the time derivative of the Jacobian matrix. Whether it is possible to conduct redundancy resolution at the velocity level but with the compliance to joint angle constraints, joint velocity constraints, and joint acceleration constraints remains an open problem in past decades. This paper gives a positive answer to this pending problem by providing a novel scheme. In the proposed scheme, the redundancy resolution problem is formulated as a quadratic program subject to joint angle, velocity, and acceleration constraints with the joint velocity being the decision variable and joint velocity norm as the performance index, which is widely adopted and closely related to the energy consumption. Then, a projection neural network is designed and proposed to online solve the problem with the joint acceleration constraint handled. Theoretical analysis is performed to guarantee the global convergence of the proposed projection neural network to the optimal solution to the redundancy resolution problem. Besides, simulation results based on a PUMA 560 industrial manipulator are presented and compared to verify the theoretical result and substantiate the efficacy and superiority of the proposed scheme. Yinyan Zhang, Shuai Li 0002, Jie Gui, Xin Luo 0001 |
IEEE Trans. Ind. Informatics | 3 |
| 2018 | Supervised Discrete Hashing With RelaxationabstractData-dependent hashing has recently attracted attention due to being able to support efficient retrieval and storage of high-dimensional data, such as documents, images, and videos. In this paper, we propose a novel learning-based hashing method called "supervised discrete hashing with relaxation" (SDHR) based on "supervised discrete hashing" (SDH). SDH uses ordinary least squares regression and traditional zero-one matrix encoding of class label information as the regression target (code words), thus fixing the regression target. In SDHR, the regression target is instead optimized. The optimized regression target matrix satisfies a large margin constraint for correct classification of each example. Compared with SDH, which uses the traditional zero-one matrix, SDHR utilizes the learned regression target matrix and, therefore, more accurately measures the classification error of the regression model and is more flexible. As expected, SDHR generally outperforms SDH. Experimental results on two large-scale image data sets (CIFAR-10 and MNIST) and a large-scale and challenging face data set (FRGC) demonstrate the effectiveness and efficiency of SDHR. Jie Gui, Tongliang Liu, Zhenan Sun, Dacheng Tao, Tieniu Tan |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2017 | Feature Selection Based on Structured Sparsity: A Comprehensive StudyabstractFeature selection (FS) is an important component of many pattern recognition tasks. In these tasks, one is often confronted with very high-dimensional data. FS algorithms are designed to identify the relevant feature subset from the original features, which can facilitate subsequent analysis, such as clustering and classification. Structured sparsity-inducing feature selection (SSFS) methods have been widely studied in the last few years, and a number of algorithms have been proposed. However, there is no comprehensive study concerning the connections between different SSFS methods, and how they have evolved. In this paper, we attempt to provide a survey on various SSFS methods, including their motivations and mathematical representations. We then explore the relationship among different formulations and propose a taxonomy to elucidate their evolution. We group the existing SSFS methods into two categories, i.e., vector-based feature selection (feature selection based on lasso) and matrix-based feature selection (feature selection based on lr,p-norm). Furthermore, FS has been combined with other machine learning algorithms for specific applications, such as multitask learning, multilabel learning, multiview learning, classification, and clustering. This paper not only compares the differences and commonalities of these methods based on regression and regularization strategies, but also provides useful guidelines to practitioners working in related fields to guide them how to do feature selection. Jie Gui, Zhenan Sun, Shuiwang Ji, Dacheng Tao, Tieniu Tan |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2016 | Representative Vector Machines: A Unified Framework for Classical ClassifiersabstractClassifier design is a fundamental problem in pattern recognition. A variety of pattern classification methods such as the nearest neighbor (NN) classifier, support vector machine (SVM), and sparse representation-based classification (SRC) have been proposed in the literature. These typical and widely used classifiers were originally developed from different theory or application motivations and they are conventionally treated as independent and specific solutions for pattern classification. This paper proposes a novel pattern classification framework, namely, representative vector machines (or RVMs for short). The basic idea of RVMs is to assign the class label of a test example according to its nearest representative vector. The contributions of RVMs are twofold. On one hand, the proposed RVMs establish a unified framework of classical classifiers because NN, SVM, and SRC can be interpreted as the special cases of RVMs with different definitions of representative vectors. Thus, the underlying relationship among a number of classical classifiers is revealed for better understanding of pattern classification. On the other hand, novel and advanced classifiers are inspired in the framework of RVMs. For example, a robust pattern classification method called discriminant vector machine (DVM) is motivated from RVMs. Given a test example, DVM first finds its k -NNs and then performs classification based on the robust M-estimator and manifold regularization. Extensive experimental evaluations on a variety of visual recognition tasks such as face recognition (Yale and face recognition grand challenge databases), object categorization (Caltech-101 dataset), and action recognition (Action Similarity LAbeliNg) demonstrate the advantages of DVM over other classifiers. Jie Gui, Tongliang Liu, Dacheng Tao, Zhenan Sun, Tieniu Tan |
IEEE Trans. Cybern. | 1 |
| 2016 | Large Margin Multi-Modal Multi-Task Feature Extraction for Image ClassificationabstractThe features used in many image analysis-based applications are frequently of very high dimension. Feature extraction offers several advantages in high-dimensional cases, and many recent studies have used multi-task feature extraction approaches, which often outperform single-task feature extraction approaches. However, most of these methods are limited in that they only consider data represented by a single type of feature, even though features usually represent images from multiple modalities. We, therefore, propose a novel large margin multi-modal multi-task feature extraction (LM3FE) framework for handling multi-modal features for image classification. In particular, LM3FE simultaneously learns the feature extraction matrix for each modality and the modality combination coefficients. In this way, LM3FE not only handles correlated and noisy features, but also utilizes the complementarity of different modalities to further help reduce feature redundancy in each modality. The large margin principle employed also helps to extract strongly predictive features, so that they are more suitable for prediction (e.g., classification). An alternating algorithm is developed for problem optimization, and each subproblem can be efficiently solved. Experiments on two challenging real-world image data sets demonstrate the effectiveness and superiority of the proposed method. Yong Luo 0002, Yonggang Wen 0001, Dacheng Tao, Jie Gui, Chao Xu 0006 |
IEEE Trans. Image Process. | 4 |
| 2014 | An optimal set of code words and correntropy for rotated least squares regressionabstractThis paper presents a robust feature extraction method for face recognition based on least squares regression (LSR). Our focus is to enhance the robustness and discriminability of the LSR. First, an optimal set of code words is introduced in LSR. Compared to the traditional set of code words, this new set uses less number of code words. Furthermore, it can make the distance of the regression targets of different classes as large as possible. Then, correntropy is integrated into the LSR model for better robustness. Furthermore, considering the commonly used distance metrics such as Euclidean distance and Cosine distance in the subspace are invariant to rotation transformation, rotation is introduced as additional freedom to promote flexibility without sacrificing accuracy. Our objective function is optimized using half-quadratic (HQ) optimization, which facilitates algorithm development and convergence study. Experimental results show that our method outperforms several subspace methods for face recognition, which indicates the validity of the proposed method. Jie Gui, Zhenan Sun, Guangqi Hou, Tieniu Tan |
IJCB | 1 |
| 2014 | How to Estimate the Regularization Parameter for Spectral Regression Discriminant Analysis and its Kernel Version?abstractSpectral regression discriminant analysis (SRDA) has recently been proposed as an efficient solution to large-scale subspace learning problems. There is a tunable regularization parameter in SRDA, which is critical to algorithm performance. However, how to automatically set this parameter has not been well solved until now. So this regularization parameter was only set to be a constant in SRDA, which is obviously suboptimal. This paper proposes to automatically estimate the optimal regularization parameter of SRDA based on the perturbation linear discriminant analysis (PLDA). In addition, two parameter estimation methods for the kernel version of SRDA are also developed. One is derived from the method of optimal regularization parameter estimation for SRDA. The other is to utilize the kernel version of PLDA. Experiments on a number of publicly available databases demonstrate the effectiveness of the proposed methods for face recognition, spoken letter recognition, handwritten digit recognition, and text categorization. Jie Gui, Zhenan Sun, Shuiwang Ji, Xindong Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2014 | Group Sparse Multiview Patch Alignment Framework With View Consistency for Image ClassificationabstractNo single feature can satisfactorily characterize the semantic concepts of an image. Multiview learning aims to unify different kinds of features to produce a consensual and efficient representation. This paper redefines part optimization in the patch alignment framework (PAF) and develops a group sparse multiview patch alignment framework (GSM-PAF). The new part optimization considers not only the complementary properties of different views, but also view consistency. In particular, view consistency models the correlations between all possible combinations of any two kinds of view. In contrast to conventional dimensionality reduction algorithms that perform feature extraction and feature selection independently, GSM-PAF enjoys joint feature extraction and feature selection by exploiting l(2,1)-norm on the projection matrix to achieve row sparsity, which leads to the simultaneous selection of relevant features and learning transformation, and thus makes the algorithm more discriminative. Experiments on two real-world image data sets demonstrate the effectiveness of GSM-PAF for image classification. Jie Gui, Dacheng Tao, Zhenan Sun, Yong Luo 0002, Xinge You, Yuan Yan Tang |
IEEE Trans. Image Process. | 1 |
| 2014 | Angular Pattern and Binary Angular Pattern for Shape RetrievalabstractIn this paper, we propose two novel shape descriptors, angular pattern (AP) and binary angular pattern (BAP), and a multiscale integration of them for shape retrieval. Both AP and BAP are intrinsically invariant to scale and rotation. More importantly, being global shape descriptors, the proposed shape descriptors are computationally very efficient, while possessing similar discriminability as state-of-the-art local descriptors. As a result, the proposed approach is attractive for real world shape retrieval applications. The experiments on the widely used MPEG-7 and TARI-1000 data sets demonstrate the effectiveness of the proposed method in comparison with existing methods. Rong-Xiang Hu, Wei Jia 0001, Haibin Ling, Yang Zhao 0002, Jie Gui |
IEEE Trans. Image Process. | 5 |
| 2014 | Histogram of Oriented Lines for Palmprint RecognitionabstractSubspace learning methods are very sensitive to the illumination, translation, and rotation variances in image recognition. Thus, they have not obtained promising performance for palmprint recognition so far. In this paper, we propose a new descriptor of palmprint named histogram of oriented lines (HOL), which is a variant of histogram of oriented gradients (HOG). HOL is not very sensitive to changes of illumination, and has the robustness against small transformations because slight translations and rotations make small histogram value changes. Based on HOL, even some simple subspace learning methods can achieve high recognition rates. Wei Jia 0001, Rong-Xiang Hu, Ying-Ke Lei, Yang Zhao 0002, Jie Gui |
IEEE Trans. Syst. Man Cybern. Syst. | 5 |
| 2013 | Face recognition via Weighted Sparse Representation
Canyi Lu, Hai Min, Jie Gui, Lin Zhu 0008, Ying-Ke Lei |
J. Vis. Commun. Image Represent. | 3 |
| 2012 | Regularization parameter estimation for spectral regression discriminant analysis based on perturbation theory
Jie Gui, Zhenan Sun, Tieniu Tan |
ICPR | 1 |
| 2012 | Newborn footprint recognition using orientation feature
Wei Jia 0001, Hai-Yang Cai, Jie Gui, Rong-Xiang Hu, Ying-Ke Lei |
Neural Comput. Appl. | 3 |
| 2012 | Feature extraction using orthogonal discriminant local tangent space alignment
Ying-Ke Lei, Yangming Xu, Junan Yang, Zhiguo Ding 0004, Jie Gui |
Pattern Anal. Appl. | 5 |
| 2012 | Discriminant sparse neighborhood preserving embedding for face recognition
Jie Gui, Zhenan Sun, Wei Jia 0001, Rong-Xiang Hu, Ying-Ke Lei, Shuiwang Ji |
Pattern Recognit. | 1 |
| 2012 | Perceptually motivated morphological strategies for shape retrieval
Rong-Xiang Hu, Wei Jia 0001, Yang Zhao 0002, Jie Gui |
Pattern Recognit. | 4 |
| 2012 | Hand shape recognition based on coherent distance shape contexts
Rong-Xiang Hu, Wei Jia 0001, David Zhang 0001, Jie Gui, Liang-Tu Song |
Pattern Recognit. | 4 |
| 2010 | Newborn Footprint Recognition Using Subspace Learning Methods
Wei Jia 0001, Jie Gui, Rong-Xiang Hu, Ying-Ke Lei, Xue-Yang Xiao |
ICIC (1) | 2 |
| 2010 | A Method for ICA with Reference Signals
Jian-Xun Mi, Jie Gui |
ICIC (2) | 2 |
| 2010 | Dual Unsupervised Discriminant Projection for Face Recognition
Jie Gui |
ICIC (1) | 2 |
| 2010 | Multi-step dimensionality reduction and semi-supervised graph-based tumor classification using gene expression data
Jie Gui, Shu-Lin Wang, Ying-Ke Lei |
Artif. Intell. Medicine | 1 |
| 2010 | Using manifold embedding for assessing and predicting protein interactions from high-throughput experimental dataabstractMOTIVATION: High-throughput protein interaction data, with ever-increasing volume, are becoming the foundation of many biological discoveries, and thus high-quality protein-protein interaction (PPI) maps are critical for a deeper understanding of cellular processes. However, the unreliability and paucity of current available PPI data are key obstacles to the subsequent quantitative studies. It is therefore highly desirable to develop an approach to deal with these issues from the computational perspective. Most previous works for assessing and predicting protein interactions either need supporting evidences from multiple information resources or are severely impacted by the sparseness of PPI networks. RESULTS: We developed a robust manifold embedding technique for assessing the reliability of interactions and predicting new interactions, which purely utilizes the topological information of PPI networks and can work on a sparse input protein interactome without requiring additional information types. After transforming a given PPI network into a low-dimensional metric space using manifold embedding based on isometric feature mapping (ISOMAP), the problem of assessing and predicting protein interactions is recasted into the form of measuring similarity between points of its metric space. Then a reliability index, a likelihood indicating the interaction of two proteins, is assigned to each protein pair in the PPI networks based on the similarity between the points in the embedded space. Validation of the proposed method is performed with extensive experiments on densely connected and sparse PPI network of yeast, respectively. Results demonstrate that the interactions ranked top by our method have high-functional homogeneity and localization coherence, especially our method is very efficient for large sparse PPI network with which the traditional algorithms fail. Therefore, the proposed algorithm is a much more promising method to detect both false positive and false negative interactions in PPI networks. AVAILABILITY: MATLAB code implementing the algorithm is available from the web site http://home.ustc.edu.cn/∼yzh33108/Manifold.htm. Zhu-Hong You, Ying-Ke Lei, Jie Gui, De-Shuang Huang, Xiaobo Zhou 0001 |
Bioinform. | 3 |
| 2010 | Locality preserving discriminant projections for face and palmprint recognition
Jie Gui, Wei Jia 0001, Shu-Ling Wang, De-Shuang Huang |
Neurocomputing | 1 |
| 2009 | Locality Preserving Discriminant Projections
Jie Gui |
ICIC (2) | 1 |
| 2008 | A Novel Hybrid Method of Gene Selection and Its Application on Tumor Classification
Zhu-Hong You, Shulin Wang, Jie Gui, Shanwen Zhang |
ICIC (2) | 3 |
| 2008 | An improvement on learning with local and global consistencyabstractA modified version for semi-supervised learning algorithm with local and global consistency was proposed in this paper. The new method adds the label information, and adopts the geodesic distance rather than Euclidean distance as the measure of the difference between two data points when conducting calculation. In addition we add class prior knowledge. It was found that the effect of class prior knowledge was different between under high label rate and low label rate. The experimental results show that the changes attain the satisfying classification performance better than the original algorithms. Jie Gui, De-Shuang Huang, Zhu-Hong You |
ICPR | 1 |