VLDB 2026 Research / reviewers in the wild / expert
Xuequan Lu
dblp:137/2585
· DBLP profile ↗
125ranked-venue papers
8as first author
109since 2021 · last 2026
0000-0003-0959-408XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 84 · 7 first-author · 74 since 2021Artificial intelligence and machine learning · 54 · 2 first-author · 47 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Security and privacy · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DAPointMamba: Domain Adaptive Point Mamba for Point Cloud CompletionabstractDomain adaptive point cloud completion (DA PCC) aims to narrow the geometric and semantic discrepancies between the labeled source and unlabeled target domains. Existing methods either suffer from limited receptive fields or quadratic complexity due to using CNNs or vision Transformers. In this paper, we present the first work that studies the adaptability of state space models (SSMs) in DA PCC and find that directly applying SSMs to DA PCC will encounter several challenges: directly serializing 3D point clouds into 1D sequences often disrupts the spatial topology and local geometric features of the target domain. Besides, the overlook of designs in the learning domain-agnostic representations hinders the adaptation performance. To address these issues, we propose a novel framework, DAPointMamba for DA PCC, that exhibits strong adaptability across domains and has the advantages of global receptive fields and efficient linear complexity. It has three novel modules. In particular, Cross-Domain Patch-Level Scanning introduces patch-level geometric correspondences, enabling effective local alignment. Cross-Domain Spatial SSM Alignment further strengthens spatial consistency by modulating patch features based on cross-domain similarity, effectively mitigating fine-grained structural discrepancies. Cross-Domain Channel SSM Alignment actively addresses global semantic gaps by interleaving and aligning feature channels. Extensive experiments on both synthetic and real-world benchmarks demonstrate that our DAPointMamba outperforms state-of-the-art methods with less computational complexity and inference latency. Qianyu Zhou 0001, Di Shao, Ye Zhu 0002, Richard Dazeley, Xuequan Lu |
AAAI | 7 |
| 2026 | PointDGRWKV: Generalizing RWKV-like Architecture to Unseen Domains for Point Cloud ClassificationabstractDomain Generalization (DG) has been recently explored to enhance the generalizability of Point Cloud Classification (PCC) models toward unseen domains. Prior works are based on convolutional networks, Transformer or Mamba architectures, either suffering from limited receptive fields or high computational cost, or insufficient long-range dependency modeling. RWKV, as an emerging architecture, possesses superior linear complexity, global receptive fields, and long-range dependency. In this paper, we present the first work that studies the generalizability of RWKV models in DG PCC. We find that directly applying RWKV to DG PCC encounters two significant challenges: RWKV's fixed direction token shift methods, like Q-Shift, introduce spatial distortions when applied to unstructured point clouds, weakening local geometric modeling and reducing robustness. In addition, the Bi-WKV attention in RWKV amplifies slight cross-domain differences in key distributions through exponential weighting, leading to attention shifts and degraded generalization. To this end, we propose PointDGRWKV, the first RWKV-based framework tailored for DG PCC. It introduces two core modules to enhance spatial modeling and cross-domain robustness, while maintaining RWKV's linear efficiency. In particular, we present Adaptive Geometric Token Shift to model local neighborhood structures to improve geometric context awareness. In addition, Cross-Domain key feature Distribution Alignment is designed to mitigate attention drift by aligning key feature distributions across domains. Extensive experiments on multiple benchmarks demonstrate that PointDGRWKV achieves state-of-the-art performance on DG PCC. Qianyu Zhou 0001, Haijia Sun, Xiangtai Li, Xuequan Lu, Lizhuang Ma, Shuicheng Yan |
AAAI | 5 |
| 2025 | DAPoinTr: Domain Adaptive Point Transformer for Point Cloud CompletionabstractPoint Transformers (PoinTr) have shown great potential in point cloud completion recently. Nevertheless, effective domain adaptation that improves transferability toward target domains remains unexplored. In this paper, we delve into this topic and empirically discover that direct feature alignment on point Transformer’s CNN backbone only brings limited improvements since it cannot guarantee sequence-wise domain-invariant features in the Transformer. To this end, we propose a pioneering Domain Adaptive Point Transformer (DAPoinTr) framework for point cloud completion. DAPoinTr consists of three novel components: Domain Query-based Feature Alignment (DQFA), Point Token-wise Feature alignment (PTFA), and Voted Prediction Consistency (VPC). In particular, DQFA is presented to narrow the global domain gaps from the sequence via the presented domain proxy and domain query at the Transformer encoder and decoder, respectively. PTFA is proposed to close the local domain shifts by aligning the tokens, i.e., point proxy and dynamic query, at the Transformer encoder and decoder, respectively. VPC is designed to consider different Transformer decoders as multiple of experts (MoE) for ensembled prediction voting and pseudo-label generation. Extensive experiments with visualization on several challenging domain adaptation benchmarks demonstrate the effectiveness and superiority of our DAPoinTr compared with other state-of-the-art methods. Qianyu Zhou 0001, Jingyu Gong, Ye Zhu 0002, Richard Dazeley, Xinkui Zhao, Xuequan Lu |
AAAI | 7 |
| 2025 | RI-MAE: Rotation-Invariant Masked AutoEncoders for Self-Supervised Point Cloud Representation LearningabstractMasked point modeling methods have recently achieved great success in self-supervised learning for point cloud data. However, these methods are sensitive to rotations and often exhibit sharp performance drops when encountering rotational variations. In this paper, we propose a novel Rotation-Invariant Masked AutoEncoders (RI-MAE) to address two major challenges: 1) achieving rotation-invariant latent representations, and 2) facilitating self-supervised reconstruction in a rotation-invariant manner. For the first challenge, we introduce RI-Transformer, which features disentangled geometry content, rotation-invariant relative orientation and position embedding mechanisms for constructing rotation-invariant point cloud latent space. For the second challenge, a novel dual-branch student-teacher architecture is devised. It enables the self-supervised learning via the reconstruction of masked patches within the learned rotation-invariant latent space. Each branch is based on an RI-Transformer, and they are connected with an additional RI-Transformer predictor. The teacher encodes all point patches, while the student solely encodes unmasked ones. Finally, the predictor predicts the latent features of the masked patches using the output latent embeddings from the student, supervised by the outputs from the teacher. Extensive experiments demonstrate that our method is robust to rotations, achieving the state-of-the-art performance on various downstream tasks. Kunming Su, Qiuxia Wu, Panpan Cai, Xiaogang Zhu 0001, Xuequan Lu, Zhiyong Wang 0001, Kun Hu 0008 |
AAAI | 5 |
| 2025 | PointDGMamba: Domain Generalization of Point Cloud Classification via Generalized State Space ModelabstractDomain Generalization (DG) has been recently explored to improve the generalizability of point cloud classification (PCC) models toward unseen domains. However, they often suffer from limited receptive fields or quadratic complexity due to the use of convolution neural networks or vision Transformers. In this paper, we present the first work that studies the generalizability of state space models (SSMs) in DG PCC and find that directly applying SSMs into DG PCC will encounter several challenges: the inherent topology of the point cloud tends to be disrupted and leads to noise accumulation during the serialization stage. Besides, the lack of designs in domain-agnostic feature learning and data scanning will introduce unanticipated domain-specific information into the 3D sequence data. To this end, we propose a novel framework, PointDGMamba, that excels in strong generalizability toward unseen domains and has the advantages of global receptive fields and efficient linear complexity. PointDGMamba consists of three innovative components: Masked Sequence Denoising (MSD), Sequence-wise Cross-domain Feature Aggregation (SCFA), and Dual-level Domain Scanning (DDS). In particular, MSD selectively masks out the noised point tokens of the point cloud sequences, SCFA introduces cross-domain but same-class point cloud features to encourage the model to learn how to extract more generalized features. DDS includes intra-domain scanning and cross-domain scanning to facilitate information exchange between features. In addition, we propose a new and more challenging benchmark PointDG-3to1 for multi-domain generalization. Extensive experiments demonstrate the effectiveness and state-of-the-art performance of PointDGMamba. Qianyu Zhou 0001, Haijia Sun, Xiangtai Li, Fengqi Liu, Xuequan Lu, Lizhuang Ma, Shuicheng Yan |
AAAI | 6 |
| 2025 | Cross-Rejective Open-Set SAR Image RegistrationabstractSynthetic Aperture Radar (SAR) image registration is an essential upstream task in geoscience applications, in which pre-detected keypoints from two images are employed as observed objects to seek matched-point pairs. In general, the registration is regarded as a typical closed-set classification, which forces each keypoint to be classified into the given classes, but ignoring an essential issue that numerous redundant keypoints are beyond the given classes, which unavoidably results in capturing incorrect matched-point pairs. Based on this, we propose a Cross-Rejective Open-set SAR Image Registration (CroR-OSIR) method. In this work, these redundant keypoints are regarded as out-of-distribution (OOD) samples, and we formulate the registration as a special open-set task with two modules: supervised contrastive feature-tuning and cross-rejective open-set recognition (CroR-OSR). Unlike traditional open-set recognition, all samples, including OOD samples, are available in the CroR-OSR module. CroR-OSR conducts the closed-set classifications in individual open-set domains from two images, meanwhile employing the cross-domain rejection during training, to exclude these OOD samples based on confidence and consistency. Moreover, a new supervised contrastive tuning strategy is incorporated for feature-tuning. Especially, the cross-domain estimation labels obtained by CroR-OSR are fed back to the feature-tuning module for feature-tuning, to enhance feature discriminability. The experimental results illustrate that the proposed method achieves more precise registration than the state-of-the-art methods. The code is released at https://github.com/XDyaoshi/CroR-OSIR-main. Shasha Mao, Shiming Lu, Zhaolong Du, Licheng Jiao, Shuiping Gou, Luntian Mou, Xuequan Lu |
CVPR | 7 |
| 2025 | Rethinking Multiple-Instance Learning From Feature Space to Probability SpaceabstractMultiple-instance learning (MIL) was initially proposed to identify key instances within a set (bag) of instances when only one bag-level label is provided. Current deep MIL models mostly solve multi-instance problem in feature space. Nevertheless, with the increasing complexity of data, we found this paradigm faces significant risks in representation learning stage, which could lead to algorithm degradation in deep MIL models. We speculate that the degradation issue stems from the persistent drift of instances in feature space during learning. In this paper, we propose a novel Probability-Space MIL network (PSMIL) as a countermeasure. In PSMIL, a self-training alignment strategy is introduced in probability space to cope with the drift problem in feature space, and the alignment target objective is proven mathematically optimal. Furthermore, we reveal that the widely-used attention-based pooling mechanism in current deep MIL models is easily affected by the perturbation in feature space and further introduce an alternative called probability-space attention pooling. It effectively captures the key instance in each bag from feature space to probability space, and further eliminates the impact of selection drift in the pooling stage. To summarize, PSMIL seeks to solve a MIL problem in probability space rather than feature space. Experimental results illustrate that PSMIL could potentially achieve performance close to supervised learning level in complex tasks (gap within 5\%), with the incremental alignment in propability space bring more than 19\% accuracy improvements for current existing mainstream models in simulated CIFAR datasets. For existing publicly available MIL benchmarks/datasets, attention in probability space also achieves competitive performance to the state-of-the-art deep MIL models. Codes are available at \url{https://github.com/LMBDA-design/PSAMIL}. Zhaolong Du, Shasha Mao, Xuequan Lu, Mengnan Qi, Licheng Jiao |
ICLR | 3 |
| 2025 | SU-SAM: A Simple Unified Framework for Adapting SAM in Underperformed SceneabstractSegment Anything Model (SAM) excels in common vision tasks but struggles with specialized data. Recent methods fine-tune SAM using parameter-efficient techniques and task-specific designs, but they rely heavily on handcrafting and pre/post-processing, limiting the generalizability. In this paper, we propose SU-SAM, a simple and unified framework that adapts SAM efficiently without task-specific designs, improving its adaptability to underperforming scenes. SU-SAM abstracts parameter-efficient modules into basic design elements, offering four variants: series, parallel, mixed, and LoRA structures. Experiments across nine datasets and six tasks, including medical and defect segmentation, demonstrate SU-SAM’s superior performance. We analyze the effectiveness of different parameter-efficient designs and present a generalized model and benchmark, highlighting SU-SAM’s adaptability across diverse datasets. Yiran Song, Qianyu Zhou 0001, Xuequan Lu, Zhiwen Shao, Lizhuang Ma |
ICME | 3 |
| 2025 | Adversarial Attacks on Both Face Recognition and Face Anti-spoofing ModelsabstractAdversarial attacks on Face Recognition (FR) systems have demonstrated significant effectiveness against standalone FR models. However, their practicality diminishes in complete FR systems that incorporate Face Anti-Spoofing (FAS) models, as these models can detect and mitigate a substantial number of adversarial examples. To address this critical yet under-explored challenge, we introduce a novel attack setting that targets both FR and FAS models simultaneously, thereby enhancing the practicability of adversarial attacks on integrated FR systems. Specifically, we propose a new attack method, termed Reference-free Multi-level Alignment (RMA), designed to improve the capacity of black-box attacks on both FR and FAS models. The RMA framework is built upon three key components. Firstly, we propose an Adaptive Gradient Maintenance module to address the imbalances in gradient contributions between FR and FAS models. Secondly, we develop a Reference-free Intermediate Biasing module to improve the transferability of adversarial examples against FAS models. In addition, we introduce a Multi-level Feature Alignment module to reduce feature discrepancies at various levels of representation. Extensive experiments showcase the superiority of our proposed attack method to state-of-the-art adversarial attacks. Fengfan Zhou, Qianyu Zhou 0001, Heifei Ling, Xuequan Lu |
IJCAI | 4 |
| 2025 | GSHOI Denoiser: Denoising Gaussian Hand-Object Interaction for Photorealistic RenderingabstractMany VR/AR applications require the photorealistic rendering of hand-object interactions. Virtual hands are driven by users' hand poses captured via motion tracking to interact with virtual objects. The driven pose can be very noisy due to the constraints of tracking hardware and computation accuracy. This noise may lead to distorted hand poses and penetration artifacts during rendering. In this paper, we introduce the Gaussian Hand-Object Interaction Denoiser, the Gaussian splatting-based hand-object interaction denoising method, which effectively denoises the input twisted and penetrated hand poses to produce photorealistic results. We first propose the innovative joint-to-Gaussian surface representation, which accurately models the spatial relationships between hand skeleton joints and object Gaussians while highlighting hand-object penetrations and generalizing well to new hand poses and objects. Then, we propose a geometry-aware de-penetration algorithm that eliminates penetrations by detecting intersections between skeleton bones and object Gaussians and reposing any penetrated fingers onto the estimated underlying surface of the object. Experiments demonstrate that our method not only effectively reduces hand-object penetration depth but also produces more realistic rendering quality compared to the state-of-the-art methods MANUS+GEARS, MANUS+GeneOH, and$2 \text{DGS}+\text{Gene} \text{OH}$. The user study results show that our method significantly improves the users' visual perceptual experience regarding penetration and stability metrics. Project page: https://github.com/ZhaoLizz/GSHOIDenoiser Lizhi Zhao, Xuequan Lu, Wei Ke 0001, Lili Wang 0006 |
ISMAR | 2 |
| 2025 | Learning Adaptive Node Selection with External Attention for Human Interaction RecognitionabstractMost GCN-based methods model interacting individuals as independent graphs, neglecting their inherent inter-dependencies. Although recent approaches utilize predefined interaction adjacency matrices to integrate participants, these matrices fail to adaptively capture the dynamic and context-specific joint interactions across different actions. In this paper, we propose the Active Node Selection with External Attention Network (ASEA), an innovative approach that dynamically captures interaction relationships without predefined assumptions. Our method models each participant individually using a GCN to capture intra-personal relationships, facilitating a detailed representation of their actions. To identify the most relevant nodes for interaction modeling, we introduce the Adaptive Temporal Node Amplitude Calculation (AT-NAC) module, which estimates global node activity by combining spatial motion magnitude with adaptive temporal weighting, thereby highlighting salient motion patterns while reducing irrelevant or redundant information. A learnable threshold, regularized to prevent extreme variations, is defined to selectively identify the most informative nodes for interaction modeling. To capture interactions, we design the External Attention (EA) module to operate on active nodes, effectively modeling the interaction dynamics and semantic relationships between individuals. Extensive evaluations show that our method captures interaction relationships more effectively and flexibly, achieving state-of-the-art performance. Chen Pang 0001, Xuequan Lu, Qianyu Zhou 0001, Lei Lyu 0001 |
ACM Multimedia | 2 |
| 2025 | Walking the Schrödinger Bridge: A Direct Trajectory for Text-to-3D GenerationabstractRecent advancements in optimization-based text-to-3D generation heavily rely on distilling knowledge from pre-trained text-to-image diffusion models using techniques like Score Distillation Sampling (SDS), which often introduce artifacts such as over-saturation and over-smoothing into the generated 3D assets. In this paper, we address this essential problem by formulating the generation process as learning an optimal, direct transport trajectory between the distribution of the current rendering and the desired target distribution, thereby enabling high-quality generation with smaller Classifier-free Guidance (CFG) values. At first, we theoretically establish SDS as a simplified instance of the Schrödinger Bridge framework. We prove that SDS employs the reverse process of an Schrödinger Bridge, which, under specific conditions (e.g., a Gaussian noise as one end), collapses to SDS's score function of the pre-trained diffusion model. Based upon this, we introduce Trajectory-Centric Distillation (TraCe), a novel text-to-3D generation framework, which reformulates the mathematically trackable framework of Schrödinger Bridge to explicitly construct a diffusion bridge from the current rendering to its text-conditioned, denoised target, and trains a LoRA-adapted model on this trajectory's score dynamics for robust 3D optimization. Comprehensive experiments demonstrate that TraCe consistently achieves superior quality and fidelity to state-of-the-art techniques. Our code will be released to the community. Ziying Li, Xuequan Lu, Xinkui Zhao, Guanjie Cheng, Shuiguang Deng, Jianwei Yin |
NeurIPS | 2 |
| 2025 | DS-MAE: Dual-Siamese Masked Autoencoders for Point Cloud AnalysisabstractMasked autoencoders (MAEs) have emerged as a powerful self-supervised approach for point cloud analysis. Nevertheless, existing methods often separately focus on global structures or multi-scale features, ignoring their complementary potential. In this paper, we propose a novel dual-Siamese masked autoencoder (DS-MAE) framework that explores integrating global and hierarchical feature learning in a unified architecture for point cloud analysis. In particular, we introduce a consistent dual-branch patch embedding strategy to partition the point cloud into patches using shared group centers, ensuring both global and hierarchical branches process point patches centered at the same spatial locations. Each branch employs dual-branch Siamese encoders to process original and augmented point patches, learning representations that capture both local details and global context. In addition, we have designed cross-attention Siamese decoders to reconstruct masked point patches and align features both within and between branches with cross-attention mechanisms. Comprehensive experiments demonstrate our method consistently achieves superior results to prior methods. Code is available at https://github.com/shaoandy1211/DS-MAE.git. Di Shao, Yaping Jing, Xinkui Zhao, Shasha Mao, Lei Lyu 0001, Xiao Liu 0004, Xuequan Lu |
Comput. Vis. Media | 7 |
| 2025 | Multi-Scale Adaptive Large Kernel Graph Convolutional Network for Skeleton-Based Action Recognition
Yu-Qing Zhang, Chen Pang 0001, Pei Geng, Xuequan Lu, Lei Lyu 0001 |
J. Comput. Sci. Technol. | 4 |
| 2025 | Skeleton-based action recognition through attention guided heterogeneous graph neural network
Tianchen Li, Pei Geng, Xuequan Lu, Wanqing Li 0001, Lei Lyu 0001 |
Knowl. Based Syst. | 3 |
| 2025 | I-Nema: a large-scale microscopic image dataset for nematode recognition
Shenglin Lu, Sheldon Fung, Xuequan Lu, Wanli Ouyang, Xue Qing |
Neural Comput. Appl. | 4 |
| 2025 | MOL: Joint Estimation of Micro-Expression, Optical Flow, and Landmark via Transformer-Graph-Style ConvolutionabstractFacial micro-expression recognition (MER) is a challenging problem, due to transient and subtle micro-expression (ME) actions. Most existing methods depend on hand-crafted features, key frames like onset, apex, and offset frames, or deep networks limited by small-scale and low-diversity datasets. In this paper, we propose an end-to-end micro-action-aware deep learning framework with advantages from transformer, graph convolution, and vanilla convolution. In particular, we propose a novel F5C block composed of fully-connected convolution and channel correspondence convolution to directly extract local-global features from a sequence of raw frames, without the prior knowledge of key frames. The transformer-style fully-connected convolution is proposed to extract local features while maintaining global receptive fields, and the graph-style channel correspondence convolution is introduced to model the correlations among feature patterns. Moreover, MER, optical flow estimation, and facial landmark detection are jointly trained by sharing the local-global features. The two latter tasks contribute to capturing facial subtle action information for MER, which can alleviate the impact of insufficient training data. Extensive experiments demonstrate that our framework (i) outperforms the state-of-the-art MER methods on CASME II, SAMM, and SMIC benchmarks, (ii) works well for optical flow estimation and facial landmark detection, and (iii) can capture facial subtle muscle actions in local regions associated with MEs. Zhiwen Shao, Feiran Li, Yong Zhou 0003, Xuequan Lu, Yuan Xie 0006, Lizhuang Ma |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | Non-Rigid Point Cloud Registration via Anisotropic Hybrid Field HarmonizationabstractCurrent point cloud registration algorithms struggle to effectively handle both deformations and occlusions simultaneously. Our manifold analysis reveals this limitation arises from the inaccurate modeling of the shape's underlying manifold and the lack of an effective optimization strategy for fragmented manifold structures. In this paper, we present AniSym-Net, a novel non-rigid registration framework designed to address near-isometric deformation registration in the presence of occlusions. To encode object's coarse topological properties and local geometric information, AniSym-Net introduces a novel anisotropic hybrid shape-motion deformation field. The effectiveness of the anisotropic hybrid shape-motion fields relies on both the holonomic constraints from the symplectic structure modeling in AniSym-Net and the motion-conditional cross-attention during fusion, which calibrates geometric features using velocity-boundary constrained point motion patterns. The harmonization of correspondences derived from anisotropic hybrid fields and those from motion-shape fields significantly mitigates registration errors and occlusions. This is achieved through the optimization of loop closures of cotangent bundles within the symplectic manifold framework. We conduct comprehensive evaluation across five popular benchmarks, namely CAPE, DT4D, SAPIEN, FAUST, and DeepDeform, to demonstrate our AniSym-Net's superior performance compared to the state-of-the-art methods. Code will be publicly available. Xuequan Lu, Mohammed Bennamoun, Bin Sheng 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Masked Autoencoders in 3D Point Cloud Representation LearningabstractTransformer-based Self-supervised Representation Learning methods learn generic features from unlabeled datasets for providing useful network initialization parameters for downstream tasks. Recently, methods based upon masking Autoencoders have been explored in the fields. The input can be intuitively masked due to regular content, like sequence words and 2D pixels. However, the extension to 3D point cloud is challenging due to irregularity. In this paper, we propose masked Autoencoders in 3D point cloud representation learning (abbreviated as MAE3D), a novel autoencoding paradigm for self-supervised learning. We first split the input point cloud into patches and mask a portion of them, then use our Patch Embedding Module to extract the features of unmasked patches. Secondly, we employ patch-wise MAE3D Transformers to learn both local features of point cloud patches and high-level contextual relationships between patches, then complete the latent representations of masked patches. We use our Point Cloud Reconstruction Module with multi-task loss to complete the incomplete point cloud as a result. We conduct self-supervised pre-training on ShapeNet55 with the point cloud completion pre-text task and fine-tune the pre-trained model on ModelNet40 and ScanObjectNN (PB_T50_RS, the hardest variant). Comprehensive experiments demonstrate that the local features extracted by our MAE3D from point cloud patches are beneficial for downstream classification tasks, soundly outperforming state-of-the-art methods (93.4% and 86.2% classification accuracy, respectively).Our source codes are available at:https://github.com/Jinec98/MAE3D. Jincen Jiang, Xuequan Lu, Lizhi Zhao, Richard Dazeley, Meili Wang 0001 |
IEEE Trans. Multim. | 2 |
| 2025 | A Divide-and-Conquer Approach for Global Orientation of Non-Watertight Scene-Level Point Clouds Using 0-1 Integer OptimizationabstractOrienting point clouds is a fundamental problem in computer graphics and 3D vision, with applications in reconstruction, segmentation, and analysis. While significant progress has been made, existing approaches mainly focus on watertight, object-level 3D models. The orientation of large-scale, non-watertight 3D scenes remains an underexplored challenge. To address this gap, we propose DACPO (Divide-And-Conquer Point Orientation), a novel framework that leverages a divide-and-conquer strategy for scalable and robust point cloud orientation. Rather than attempting to orient an unbounded scene at once, DACPO segments the input point cloud into smaller, manageable blocks, processes each block independently, and integrates the results through a global optimization stage. For each block, we introduce a two-step process: estimating initial normal orientations by a randomized greedy method and refining them by an adapted iterative Poisson surface reconstruction. To achieve consistency across blocks, we model inter-block relationships using an an undirected graph, where nodes represent blocks and edges connect spatially adjacent blocks. To reliably evaluate orientation consistency between adjacent blocks, we introduce the concept of the visible connected region , which defines the region over which visibility-based assessments are performed. The global integration is then formulated as a 0-1 integer-constrained optimization problem, with block flip states as binary variables. Despite the combinatorial nature of the problem, DACPO remains scalable by limiting the number of blocks (typically a few hundred for 3D scenes) involved in the optimization. Experiments on benchmark datasets demonstrate DACPO's strong performance, particularly in challenging large-scale, non-watertight scenarios where existing methods often fail. The source code is available at https://github.com/zd-lee/DACPO. Zhuodong Li, Fei Hou 0001, Wencheng Wang 0001, Xuequan Lu, Ying He 0001 |
ACM Trans. Graph. | 4 |
| 2025 | DeSC: Learning Deep Semantic Descriptor for NeRF RegistrationabstractNeRF registration has gained increasing attention recently. While existing research demonstrates considerable potential for this task, most methods primarily focus on either global geometric or rendering photometric information during feature learning, overlooking the rich cross-modal information inherent in the NeRF embedding feature space. In this paper, we propose DeSC, a novel NeRF registration approach that leverages the rich cross-modal features from NeRF to learn robust semantic descriptors. In particular, we propose a Deep Semantic Aggregation module, which employs a weighted graph convolution network to capture high-frequency texture details in NeRF patches. This approach reveals the underlying semantics shared across different NeRFs of the same scene, thereby yielding more robust global feature descriptors that lead to better alignment accuracy and robustness. In addition, we design a density-aware photometric consistency loss that facilitates the learning of robust features. Extensive experimental results on Objaverse datasets demonstrate that our approach produces superior registration performance to state-of-the-art techniques. Sheldon Fung, Wei Pan 0010, Kui Su, Hui Cui 0002, Xinkui Zhao, Xuequan Lu |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2025 | CloudMix: Dual Mixup Consistency for Unpaired Point Cloud CompletionabstractDue to the unsatisfactory performance of supervised methods on unpaired real-world scans, point cloud completion via cross-domain adaptation has recently drawn growing attention. Nevertheless, previous approaches only focus on alleviating the distribution shift through domain alignment, resulting in massive information loss of real-world domain data. To tackle this issue, we propose a dual mixup-induced consistency regularization to integrate both source and target domain to improve robustness and generalization capability. Specifically, we mix up virtual and real-world shapes in the input and latent feature space respectively, and then regularize the completion network by forcing two kinds of mixed completion predictions to be consistent. To further adapt to each instance within the real-world domain, we design a novel density-aware refiner to utilize local context information to preserve the fine-grained details and remove noise or outliers for coarse completion. Extensive experiments on real-world scans and our synthetic unpaired datasets demonstrate the superiority of our method over existing state-of-the-art approaches. Fengqi Liu, Jingyu Gong, Qianyu Zhou 0001, Xuequan Lu, Ran Yi 0002, Yuan Xie 0006, Lizhuang Ma |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2025 | TriCI: Triple Cross-Intra Branch Contrastive Learning for Point Cloud AnalysisabstractWhereas contrastive learning eliminates the need for labeled data, existing methods may suffer from inadequate features due to the conventional single shared encoder structure and struggle to fully harness the rich spectrum of 3D augmentations. In this paper, we propose TriCI, a self-supervised method that designs a triple-branch contrastive learning architecture. During contrastive pre-training, we generate three augmented versions of each input point cloud sample and pair each augmented sample with the original one, resulting in three unique positive pairs. We subsequently feed the pairs into three distinct encoders, each of which extracts features from its corresponding input positive pair. We design a novel cross-branch contrastive loss and use it along with the intra-branch contrastive loss to jointly train our network. The proposed cross-branch loss effectively aligns the output features from different perspectives for pre-training and facilitates their integration for downstream tasks, particularly in object-level scenarios. The intra-branch loss helps maximize the feature correspondences within positive pairs. Extensive experiments demonstrate the superiority of our TriCI in self-supervised learning, and show its strong ability in enhancing the performance of downstream object classification and part segmentation tasks. Interestingly, our TriCI achieves a 92.9% accuracy for linear SVM evaluation on ModelNet40, exceeding its closest competitor by 1.7% and even exceeding some supervised methods. Di Shao, Xuequan Lu, Xiao Liu 0004, Ajmal Mian |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2025 | Learning Implicit Fields for Point Cloud FilteringabstractSince point clouds acquired by scanners inevitably contain noise, recovering a clean version from a noisy point cloud is essential for further 3D geometry processing applications. Several data-driven approaches have been recently introduced to overcome the drawbacks of traditional filtering algorithms, such as less robust preservation of sharp features and tedious tuning for multiple parameters. Most of these methods achieve filtering by directly regressing the position/displacement of each point, which may blur detailed features and is prone to uneven distribution. In this article, we propose a novel data-driven method that explores the implicit fields. Our assumption is that the given noisy points implicitly define a surface, and we attempt to obtain a point's movement direction and distance separately based on the predicted signed distance fields (SDFs). Taking a noisy point cloud as input, we first obtain a consistent alignment by incorporating the global points into local patches. We then feed them into an encoder-decoder structure and predict a 7D vector consisting of SDFs. Subsequently, the distance can be obtained directly from the first element in the vector, and the movement direction can be obtained by computing the gradient descent from the last six elements (i.e., six surrounding SDFs). We finally obtain the filtered results by moving each point with its predicted distance along its movement direction. Our method can produce feature-preserving results without requiring explicit normals. Experiments demonstrate that our method visually outperforms state-of-the-art methods and generally produces better quantitative results than position-based methods (both learning and non-learning). Jinxi Wang, Xuequan Lu, Meili Wang 0001, Fei Hou 0001, Ying He 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2025 | GaussianHand: Real-Time 3D Gaussian Rendering for Hand Avatar AnimationabstractRendering animatable and realistic hand avatars is pivotal for enhancing user experiences in human-centered AR/VR applications. While recent initiatives have utilized neural radiance fields to forge hand avatars with lifelike appearances, these methods are often hindered by high computational demands and the necessity for extensive training views. In this paper, we introduce GaussianHand, the first Gaussian-based real-time 3D rendering approach that enables efficient free-view and free-pose hand avatar animation from sparse view images. Our approach encompasses two key innovations. We first propose Hand Gaussian Blend Shapes that effectively models hand surface geometry while ensuring consistent appearance across various poses. Second, we introduce the Neural Residual Skeleton, equipped with Residual Skinning Weights, designed to rectify inaccuracies involved in Linear Blend Skinning deformations due to geometry offsets. Experiments demonstrate that our method not only achieves far more realistic rendering quality with as few as 5 or 20 training views, compared to the 139 views required by existing methods, but also excels in efficiency, achieving up to 125 frames per second for real-time rendering and remarkably surpassing recent methods. Lizhi Zhao, Xuequan Lu, Runze Fan, Sio Kei Im, Lili Wang 0006 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2025 | Enhanced ping pong training assessment via VR: integrating time-spatial alignment and multi-modal fusion
Xuequan Lu, Xiaojun Huang |
Vis. Comput. | 2 |
| 2025 | $\hbox {D}^3$M-GS: Dynamic endoscopy reconstruction via dual-domain deformation model
Xuequan Lu, Xiaojun Huang |
Vis. Comput. | 2 |
| 2025 | VR table tennis techniques learning for humanoid robot via AMP-Former
Xuequan Lu, Xiaojun Huang |
Vis. Comput. | 2 |
| 2024 | DHGCN: Dynamic Hop Graph Convolution Network for Self-Supervised Point Cloud LearningabstractRecent works attempt to extend Graph Convolution Networks (GCNs) to point clouds for classification and segmentation tasks. These works tend to sample and group points to create smaller point sets locally and mainly focus on extracting local features through GCNs, while ignoring the relationship between point sets. In this paper, we propose the Dynamic Hop Graph Convolution Network (DHGCN) for explicitly learning the contextual relationships between the voxelized point parts, which are treated as graph nodes. Motivated by the intuition that the contextual information between point parts lies in the pairwise adjacent relationship, which can be depicted by the hop distance of the graph quantitatively, we devise a novel self-supervised part-level hop distance reconstruction task and design a novel loss function accordingly to facilitate training. In addition, we propose the Hop Graph Attention (HGA), which takes the learned hop distance as input for producing attention weights to allow edge features to contribute distinctively in aggregation. Eventually, the proposed DHGCN is a plug-and-play module that is compatible with point-based backbone networks. Comprehensive experiments on different backbones and tasks demonstrate that our self-supervised method achieves state-of-the-art performance. Our source codes are available at: https://github.com/Jinec98/DHGCN. Jincen Jiang, Lizhi Zhao, Xuequan Lu, Muhammad Imran Razzak, Meili Wang 0001 |
AAAI | 3 |
| 2024 | 3D Face Recognition with Contrastive Learning Network on Low-Quality Data
Yaping Jing, Ajmal Mian, Leo Zhang, Shang Gao 0003, Xuequan Lu |
CGI (1) | 5 |
| 2024 | TopFormer: Topology-Aware Transformer for Point Cloud Registration
Sheldon Fung, Wei Pan 0010, Xiao Liu 0004, John Yearwood, Richard Dazeley, Xuequan Lu |
CVM (1) | 6 |
| 2024 | Test-Time Domain Generalization for Face Anti-SpoofingabstractFace Anti-Spoofing (FAS) is pivotal in safeguarding facial recognition systems against presentation attacks. While domain generalization (DG) methods have been developed to enhance FAS performance, they predominantly focus on learning domain-invariant features during training, which may not guarantee generalizability to unseen data that dif-fers largely from the source distributions. Our insight is that testing data can serve as a valuable resource to enhance the generalizability beyond mere evaluation for DG FAS. In this paper, we introduce a novel Test-Time Domain Generalization (TTDG) framework for FAS, which leverages the testing data to boost the model's generalizability. Our method, consisting of Test-Time Style Projection (TTSP) and Diverse Style Shifts Simulation (DSSS), effectively projects the unseen data to the seen domain space. In particular, we first introduce the innovative TTSP to project the styles of the arbitrarily unseen samples of the testing distribution to the known source space of the training distributions. We then design the efficient DSSS to synthesize diverse style shifts via learnable style bases with two specifically designed losses in a hyperspherical feature space. Our method elimi-nates the need for model updates at the test time and can be seamlessly integrated into not only the CNN but also ViT backbones. Comprehensive experiments on widely used cross-domain FAS benchmarks demonstrate our method's state-of-the-art performance and effectiveness. Qianyu Zhou 0001, Ke-Yue Zhang, Taiping Yao, Xuequan Lu, Shouhong Ding, Lizhuang Ma |
CVPR | 4 |
| 2024 | StraightPCF: Straight Point Cloud FilteringabstractPoint cloud filtering is a fundamental 3D vision task, which aims to remove noise while recovering the underlying clean surfaces. State-of-the-art methods remove noise by moving noisy points along stochastic trajectories to the clean surfaces. These methods often require regularization within the training objective and/or during post-processing, to ensure fidelity. In this paper, we introduce StraightPCF, a new deep learning based method for point cloud filtering. It works by moving noisy points along straight paths, thus reducing discretization errors while ensuring faster convergence to the clean surfaces. We model noisy patches as intermediate states between high noise patch variants and their clean counterparts, and design the VelocityModule to infer a constant flow velocity from the former to the latter. This constant flow leads to straight filtering trajectories. In addition, we introduce a DistanceModule that scales the straight trajectory using an estimated distance scalar to attain convergence near the clean surface. Our network is lightweight and only has ~530K parameters, being 17% of IterativePFn (a most recent point cloud filtering network). Extensive experiments on both synthetic and real-world data show our method achieves state-of-the-art results. Our method also demonstrates nice distributions of filtered points without the need for regularization. The implementation code can be found at: https://github.com/ddsediri/StraightPCF. Dasith de Silva Edirimuni, Xuequan Lu, Gang Li 0009, Lei Wei 0002, Antonio Robles-Kelly, Hongdong Li |
CVPR | 2 |
| 2024 | BA-SAM: Scalable Bias-Mode Attention Mask for Segment Anything ModelabstractIn this paper, we address the challenge of image resolution variation for the Segment Anything Model (SAM). SAM, known for its zero-shot generalizability, exhibits a performance degradation when faced with datasets with varying image sizes. Previous approaches tend to resize the image to a fixed size or adopt structure modifications, hindering the preservation of SAM's rich prior knowledge. Besides, such task-specific tuning necessitates a complete retraining of the model, which is cost-expensive and unacceptable for deployment in the downstream tasks. In this paper, we reformulate this challenge as a length extrapolation problem, where token sequence length varies while maintaining a consistent patch size for images with different sizes. To this end, we propose a Scalable Bias-Mode Attention Mask (BA-SAM) to enhance SAM's adaptability to varying image resolutions while eliminating the need for structure modifications. Firstly, we introduce a new scaling factor to ensure consistent magnitude in the attention layer's dot product values when the token sequence length changes. Secondly, we present a bias-mode attention mask that allows each token to prioritize neighboring information, mitigating the impact of untrained distant information. Our BA-SAM demonstrates efficacy in two scenarios: zero-shot and finetuning. Extensive evaluation of diverse datasets, including DIS5K, DUTS, ISIC, COD10K, and COCO, reveals its ability to significantly mitigate performance degradation in the zero-shot setting and achieve state-of-the-art performance with minimal fine-tuning. Furthermore, we propose a generalized model and benchmark, showcasing BA-SAM's generalizability across all four datasets simultaneously. Yiran Song, Qianyu Zhou 0001, Xiangtai Li, Deng-Ping Fan, Xuequan Lu, Lizhuang Ma |
CVPR | 5 |
| 2024 | SemReg: Semantics Constrained Point Cloud Registration
Sheldon Fung, Xuequan Lu, Dasith de Silva Edirimuni, Wei Pan 0010, Xiao Liu 0004, Hongdong Li |
ECCV (41) | 2 |
| 2024 | DG-PIC: Domain Generalized Point-In-Context Learning for Point Cloud Understanding
Jincen Jiang, Qianyu Zhou 0001, Yuhang Li 0011, Xuequan Lu, Meili Wang 0001, Lizhuang Ma, Jian Chang 0001, Jian J. Zhang 0001 |
ECCV (6) | 4 |
| 2024 | Fine-Granularity Face Sketch SynthesisabstractGenerative Adversarial Networks (GANs) are often used in face sketch synthesis due to their powerful ability in image generation. However, most GAN based synthesis methods took the entire face as the minimum unit. Differently, we propose a novel fine-granularity face sketch synthesis framework in this paper. The core idea is to first capture local information at a fine granularity (i.e., facial component), and then generate a complete face sketch based on the fine-grained information. Specifically, we partition the face sketch into multiple components, and then train a parallel network for each component. A condition enhanced detail repair network is further designed to correct the mismatches and deformations produced during parallel generation. Extensive experiments show that our approach outperforms state-of-the-art methods from both the qualitative and quantitative perspectives. Yangdong Chen, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Xuequan Lu, Shang Gao 0003 |
ICASSP | 6 |
| 2024 | Emphasizing Semantic Consistency of Salient Posture for Speech-Driven Gesture GenerationabstractSpeech-driven gesture generation aims at synthesizing a gesture sequence synchronized with the input speech signal. Previous methods leverage neural networks to directly map a compact audio representation to the gesture sequence, ignoring the semantic association of different modalities and failing to deal with salient gestures. In this paper, we propose a novel speech-driven gesture generation method by emphasizing the semantic consistency of salient posture. Specifically, we first learn a joint manifold space for the individual representation of audio and body pose to exploit the inherent semantic association between two modalities, and propose to enforce semantic consistency via a consistency loss. Furthermore, we emphasize the semantic consistency of salient postures by introducing a weakly-supervised detector to identify salient postures, and reweighting the consistency loss to focus more on learning the correspondence between salient postures and the high-level semantics of speech content. In addition, we propose to extract audio features dedicated to facial expression and body gesture separately, and design separate branches for face and body gesture synthesis. Extensive experimental results demonstrate the superiority of our method over the state-of-the-art approaches. Fengqi Liu, Jingyu Gong, Ran Yi 0002, Qianyu Zhou 0001, Xuequan Lu, Jiangbo Lu, Lizhuang Ma |
ACM Multimedia | 6 |
| 2024 | DGMamba: Domain Generalization via Generalized State Space ModelabstractDomain generalization (DG) aims at solving distribution shift problems in various scenes. Existing approaches are based on Convolution Neural Networks (CNNs) or Vision Transformers (ViTs), which suffer from limited receptive fields or quadratic complexity issues. Mamba, as an emerging state space model (SSM), possesses superior linear complexity and global receptive fields. Despite this, it can hardly be applied to DG to address distribution shifts, due to the hidden state issues and inappropriate scan mechanisms. In this paper, we propose a novel framework for DG, named DGMamba, that excels in strong generalizability toward unseen domains and meanwhile has the advantages of global receptive fields, and efficient linear complexity. Our DGMamba compromises two core components: Hidden State Suppressing (HSS) and Semantic-aware Patch Refining (SPR). In particular, HSS is introduced to mitigate the influence of hidden states associated with domain-specific features during output prediction. SPR strives to encourage the model to concentrate more on objects rather than context, consisting of two designs: Prior-Free Scanning (PFS), and Domain Context Interchange (DCI). Concretely, PFS aims to shuffle the non-semantic patches within images, creating more flexible and effective sequences from images, and DCI is designed to regularize Mamba with the combination of mismatched non-semantic and semantic information by fusing patches among domains. Extensive experiments on four commonly used DG benchmarks demonstrate that the proposed DGMamba achieves remarkably superior results to state-of-the-art models. The code will be made publicly available at https://github.com/longshaocong/DGMamba. Shaocong Long, Qianyu Zhou 0001, Xiangtai Li, Xuequan Lu, Chenhao Ying 0001, Yuan Luo 0003, Lizhuang Ma, Shuicheng Yan |
ACM Multimedia | 4 |
| 2024 | Rethinking Impersonation and Dodging Attacks on Face Recognition SystemsabstractFace Recognition (FR) systems can be easily deceived by adversarial examples that manipulate benign face images through imperceptible perturbations. Adversarial attacks on FR encompass two types: impersonation (targeted) attacks and dodging (untargeted) attacks. Previous methods often achieve a successful impersonation attack on FR, however, it does not necessarily guarantee a successful dodging attack on FR in the black-box setting. In this paper, our key insight is that the generation of adversarial examples should perform both impersonation and dodging attacks simultaneously. To this end, we propose a novel attack method termed as Adversarial Pruning (Adv-Pruning), to fine-tune existing adversarial examples to enhance their dodging capabilities while preserving their impersonation capabilities. Adv-Pruning consists of Priming, Pruning, and Restoration stages. Concretely, we propose Adversarial Priority Quantification to measure the region-wise priority of original adversarial perturbations, identifying and releasing those with minimal impact on absolute model output variances. Then, Biased Gradient Adaptation is presented to adapt the adversarial examples to traverse the decision boundaries of both the attacker and victim by adding perturbations favoring dodging attacks on the vacated regions, preserving the prioritized features of the original perturbations while boosting dodging performance. As a result, we can maintain the impersonation capabilities of original adversarial examples while effectively enhancing dodging capabilities. Comprehensive experiments demonstrate the superiority of our method compared with state-of-the-art adversarial attack methods. Fengfan Zhou, Qianyu Zhou 0001, Bangjie Yin, Xuequan Lu, Lizhuang Ma |
ACM Multimedia | 5 |
| 2024 | SITransformer: Shared Information-Guided Transformer for Extreme Multimodal Summarization
Lintao Wang 0002, Xiaogang Zhu 0001, Xuequan Lu, Zhiyong Wang 0001, Kun Hu 0008 |
MMAsia | 4 |
| 2024 | CFRL: Coarse-Fine Decoupled Representation Learning For Long-Tailed RecognitionabstractData often faces a severe class imbalance issue in the real world, meaning that the number of instances within classes varies greatly, following a long-tailed distribution.In this case, the direct application of supervised learning yields poor performance.Existing long-tailed recognition (LTR) methods often heavily rely on the label information to enhance tail classes' accuracy at the expense of head class by an image-level end-to-end resampling strategy to address data distribution imbalance.Nevertheless, they neglect label bias, which can severely affect the LTR model's accuracy.In this paper, we propose a novel approach, namely Coarse-Fine Decoupled Representation Learning (CFRL) for LTR.Our core idea is to decouple data representations from the classifier and decompose representation learning into two stages: image-level and patch-level.Specifically, in the image-level stage, we leverage unsupervised learning on image-level information to reduce the impact of label bias caused by imbalanced datasets.In the patch-level stage, we introduce patch-level rotation augmentation as negative samples, forcing the model to acquire more comprehensive information.Our theoretical and empirical analyses demonstrate that the approach does not sacrifice the accuracy of head classes while significantly reducing the overfitting of tail classes, improving both of them.We showcase state-of-the-art results on CIFAR, ImageNet, and iNaturalist datasets.Furthermore, we illustrate that this training methodology can be combined with various existing Long-Tailed Recognition (LTR) methods, further enhancing their performance. Yiran Song, Qianyu Zhou 0001, Kun Hu 0008, Lizhuang Ma, Xuequan Lu |
MMAsia | 5 |
| 2024 | Point Cloud Normal Estimation via Representation Learning on Height MapsabstractPoint Cloud Normal Estimation via Representation Learning on Height Maps Dasith de Silva Edirimuni, Ye Zhu 0002, Shang Gao 0003, Zhiyong Wang 0001, Antonio Robles-Kelly, Xuequan Lu |
MMAsia | 7 |
| 2024 | PCoTTA: Continual Test-Time Adaptation for Multi-Task Point Cloud UnderstandingabstractIn this paper, we present PCoTTA, an innovative, pioneering framework for Continual Test-Time Adaptation (CoTTA) in multi-task point cloud understanding, enhancing the model's transferability towards the continually changing target domain. We introduce a multi-task setting for PCoTTA, which is practical and realistic, handling multiple tasks within one unified model during the continual adaptation. Our PCoTTA involves three key components: automatic prototype mixture (APM), Gaussian Splatted feature shifting (GSFS), and contrastive prototype repulsion (CPR). Firstly, APM is designed to automatically mix the source prototypes with the learnable prototypes with a similarity balancing factor, avoiding catastrophic forgetting. Then, GSFS dynamically shifts the testing sample toward the source domain, mitigating error accumulation in an online manner. In addition, CPR is proposed to pull the nearest learnable prototype close to the testing feature and push it away from other prototypes, making each prototype distinguishable during the adaptation. Experimental comparisons lead to a new benchmark, demonstrating PCoTTA's superiority in boosting the model's transferability towards the continually changing target domain. Our source code is available at: https://github.com/Jinec98/PCoTTA. Jincen Jiang, Qianyu Zhou 0001, Yuhang Li 0011, Xinkui Zhao, Meili Wang 0001, Lizhuang Ma, Jian Chang 0001, Jian J. Zhang 0001, Xuequan Lu |
NeurIPS | 9 |
| 2024 | Noise4Denoise: Leveraging noise for unsupervised point cloud denoisingabstractExisting deep learning-based point cloud denoising methods are generally trained in a supervised manner that requires clean data as ground-truth labels. However, in practice, it is not always feasible to obtain clean point clouds. In this paper, we introduce a novel unsupervised point cloud denoising method that eliminates the need to use clean point clouds as groundtruth labels during training. We demonstrate that it is feasible for neural networks to only take noisy point clouds as input, and learn to approximate and restore their clean versions. In particular, we generate two noise levels for the original point clouds, requiring the second noise level to be twice the amount of the first noise level. With this, we can deduce the relationship between the displacement information that recovers the clean surfaces across the two levels of noise, and thus learn the displacement of each noisy point in order to recover the corresponding clean point. Comprehensive experiments demonstrate that our method achieves outstanding denoising results across various datasets with synthetic and real-world noise, obtaining better performance than previous unsupervised methods and competitive performance to current supervised methods. Xiao Liu 0004, Hailing Zhou, Lei Wei 0002, Zhigang Deng 0001, M. Manzur Murshed, Xuequan Lu |
Comput. Vis. Media | 7 |
| 2024 | Counting in congested crowd scenes with hierarchical scale-aware encoder-decoder network
Run Han, Ran Qi, Xuequan Lu, Lei Huang 0010, Lei Lyu 0001 |
Expert Syst. Appl. | 3 |
| 2024 | JRC: Deepfake detection via joint reconstruction and classificationabstractDeep learning has enabled realistic face manipulation for malicious purposes (e.g., deepfakes), which poses significant concerns over the integrity of the media in circulation. Most existing deep learning techniques for deepfake detection can achieve promising performance in the intra-dataset evaluation setting, but are unable to perform satisfactorily in the inter-dataset evaluation setting. Most previous methods use a backbone network to extract global features for making predictions and only employ binary supervision to train the network. Classification merely based on the learning of global features often leads to weak generalizability to deepfakes of unseen manipulation methods. In this paper, we design a two-branch Convolutional AutoEncoder (CAE), which considers the reconstruction and classification tasks simultaneously for deepfake detection. This Joint Reconstruction and Classification (JRC) method shares the information learned by one task with the other, each focusing on different aspects, and hence boosts the overall performance. JRC is end-to-end, and experiments demonstrate that it achieves state-of-the-art performance on three commonly-used datasets, particularly in the cross-dataset evaluation setting. Bosheng Yan, Chang-Tsun Li, Xuequan Lu |
Neurocomputing | 3 |
| 2024 | PainterAR: A Self-Painting AR Interface for Mobile DevicesabstractABSTRACT Painting is a complex and creative process that involves the use of various drawing skills to create artworks. The concept of training artificial intelligence models to imitate this process is referred to as neural painting. To enable ordinary people to engage in the process of painting, we propose PainterAR, a novel interface that renders any paintings stroke‐by‐stroke in an immersive and realistic augmented reality (AR) environment. PainterAR is composed of two components: the neural painting model and the AR interface. Regarding the neural painting model, unlike previous models, we introduce the Kullback–Leibler divergence to replace the original Wasserstein distance existed in the baseline paint transformer model, which solves an important problem of encountering different scales of strokes (big or small) during painting. We then design an interactive AR interface, which allows users to upload an image and display the creation process of the neural painting model on the virtual drawing board. Experiments demonstrate that the paintings generated by our improved neural painting model are more realistic and vivid than previous neural painting models. The user study demonstrates that users prefer to control the painting process interactively in our AR environment. Yinghan Shi, Lizhi Zhao, Xuequan Lu, Henry Been-Lirn Duh, Meili Wang 0001 |
Comput. Animat. Virtual Worlds | 4 |
| 2024 | Variation-aware directed graph convolutional networks for skeleton-based action recognition
Tianchen Li, Pei Geng, Guohui Cai, Xinran Hou, Xuequan Lu, Lei Lyu 0001 |
Knowl. Based Syst. | 5 |
| 2024 | Single-stage object detector with attention mechanism for squamous cell carcinoma feature detection using histopathological imagesabstractAbstract Squamous cell carcinoma is the most common type of cancer that occurs in squamous cells of epithelial tissue. Histopathological evaluation of tissue samples is the gold standard approach used for carcinoma diagnosis. SCC detection based on various histopathological features often employs traditional machine learning approaches or pixel-based deep CNN models. This study aims to detect keratin pearl, the most prominent SCC feature, by implementing RetinaNet one-stage object detector. Further, we enhance the model performance by incorporating an attention module. The proposed method is more efficient in detection of small keratin pearls. This is the first work detecting keratin pearl resorting to the object detection technique to the extent of our knowledge. We conducted a comprehensive assessment of the model both quantitatively and qualitatively. The experimental results demonstrate that the proposed approach enhanced the mAP by about 4% compared to default RetinaNet model. Swathi Prabhu, Keerthana Prasad, Xuequan Lu, Antonio Robles-Kelly, Thuong N. Hoang |
Multim. Tools Appl. | 3 |
| 2024 | IIAM: Intra and Inter Attention With Mutual Consistency Learning Network for Medical Image SegmentationabstractMedical image segmentation provides a reliable basis for diagnosis analysis and disease treatment by capturing the global and local features of the target region. To learn global features, convolutional neural networks are replaced with pure transformers, or transformer layers are stacked at the deepest layers of convolutional neural networks. Nevertheless, they are deficient in exploring local-global cues at each scale and the interaction among consensual regions in multiple scales, hindering the learning about the changes in size, shape, and position of target objects. To cope with these defects, we propose a novel Intra and Inter Attention with Mutual Consistency Learning Network (IIAM). Concretely, we design an intra attention module to aggregate the CNN-based local features and transformer-based global information on each scale. In addition, to capture the interaction among consensual regions in multiple scales, we devise an inter attention module to explore the cross-scale dependency of the object and its surroundings. Moreover, to reduce the impact of blurred regions in medical images on the final segmentation results, we introduce multiple decoders to estimate the model uncertainty, where we adopt a mutual consistency learning strategy to minimize the output discrepancy during the end-to-end training and weight the outputs of the three decoders as the final segmentation result. Extensive experiments on three benchmark datasets verify the efficacy of our method and demonstrate superior performance of our model to state-of-the-art techniques. Chen Pang 0001, Xuequan Lu, Renfeng Zhang, Lei Lyu 0001 |
IEEE J. Biomed. Health Informatics | 2 |
| 2024 | ARFL: Adaptive and Robust Federated LearningabstractFederated Learning (FL) is a machine learning technique that enables multiple local clients holding individual datasets to collaboratively train a model, without exchanging the clients' datasets. Conventional FL approaches often assign a fixed workload (local epoch) and step size (learning rate) to the clients during the client-side local model training and utilize all collaborating trained models' parameters evenly during the server-side global model aggregation. Consequently, they frequently experience problems with data heterogeneity and high communication costs. In this paper, we propose a novel FL approach to mitigate the above problems. On the client side, we propose an adaptive model update approach that optimally allocates a needful number of local epochs and dynamically adjusts the learning rate to train the local model and regularizes the conventional objective function by adding a proximal term to it. On the server side, we propose a robust model aggregation strategy that potentially supplants the local outlier updates (models' weights) prior to the aggregation. We provide the theoretical convergence results and perform extensive experiments on different data setups over the MNIST, CIFAR-10, and Shakespeare datasets, which manifest that our FL scheme surpasses the baselines in terms of communication speedup, test-set performance, and global convergence. Md Palash Uddin, Yong Xiang 0001, Borui Cai, Xuequan Lu, John Yearwood, Longxiang Gao |
IEEE Trans. Mob. Comput. | 4 |
| 2024 | Hierarchical Aggregated Graph Neural Network for Skeleton-Based Action RecognitionabstractSupervised human action recognition methods based on skeleton data have achieved impressive performance recently. However, many current works emphasize the design of different contrastive strategies to gain stronger supervised signals, ignoring the crucial role of the model's encoder in encoding fine-grained action representations. Our key insight is that a superior skeleton encoder can effectively exploit the fine-grained dependencies between different skeleton information (e.g., joint, bone, angle) in mining more discriminative fine-grained features. In this paper, we devise an innovative hierarchical aggregated graph neural network (HA-GNN) that involves several core components. In particular, the proposed hierarchical graph convolution (HGC) module learns the complementary semantic information among joint, bone, and angle in a hierarchical manner. The designed pyramid attention fusion mechanism (PAFM) fuses the skeleton features successively to compensate for the action representations obtained by the HGC. We use the multi-scale temporal convolution (MSTC) module to enrich the expression capability of temporal features. In addition, to learn more comprehensive semantic representations of the skeleton, we construct a multi-task learning framework with simple contrastive learning and design the learnable data-enhanced strategy to acquire different data representations. Extensive experiments on NTU RGB+D 60/120, NW-UCLA, Kinetics-400, UAV-Human, and PKUMMD datasets prove that the proposed HA-GNN without contrastive learning achieves state-of-the-art performance in skeleton-based action recognition, and it achieves even better results with contrastive learning. Pei Geng, Xuequan Lu, Wanqing Li 0001, Lei Lyu 0001 |
IEEE Trans. Multim. | 2 |
| 2024 | GRA: Graph Representation Alignment for Semi-Supervised Action RecognitionabstractGraph convolutional networks (GCNs) have emerged as a powerful tool for action recognition, leveraging skeletal graphs to encapsulate human motion. Despite their efficacy, a significant challenge remains the dependency on huge labeled datasets. Acquiring such datasets is often prohibitive, and the frequent occurrence of incomplete skeleton data, typified by absent joints and frames, complicates the testing phase. To tackle these issues, we present graph representation alignment (GRA), a novel approach with two main contributions: 1) a self-training (ST) paradigm that substantially reduces the need for labeled data by generating high-quality pseudo-labels, ensuring model stability even with minimal labeled inputs and 2) a representation alignment (RA) technique that utilizes consistency regularization to effectively reduce the impact of missing data components. Our extensive evaluations on the NTU RGB+D and Northwestern-UCLA (N-UCLA) benchmarks demonstrate that GRA not only improves GCN performance in data-constrained environments but also retains impressive performance in the face of data incompleteness. Kuan-Hung Huang, Yao-Bang Huang, Yong-Xiang Lin, Kai-Lung Hua, Muhammad Tanveer 0001, Xuequan Lu, Muhammad Imran Razzak |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | Contrastive Learning for Joint Normal Estimation and Point Cloud FilteringabstractPoint cloud filtering and normal estimation are two fundamental research problems in the 3D field. Existing methods usually perform normal estimation and filtering separately and often show sensitivity to noise and/or inability to preserve sharp geometric features such as corners and edges. In this article, we propose a novel deep learning method to jointly estimate normals and filter point clouds. We first introduce a 3D patch based contrastive learning framework, with noise corruption as an augmentation, to train a feature encoder capable of generating faithful representations of point cloud patches while remaining robust to noise. These representations are consumed by a simple regression network and supervised by a novel joint loss, simultaneously estimating point normals and displacements that are used to filter the patch centers. Experimental results show that our method well supports the two tasks simultaneously and preserves sharp features and fine details. It generally outperforms state-of-the-art techniques on both tasks. Dasith de Silva Edirimuni, Xuequan Lu, Gang Li 0009, Antonio Robles-Kelly |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2024 | Unsupervised contrastive learning with simple transformation for 3D point cloud data
Jincen Jiang, Xuequan Lu, Wanli Ouyang, Meili Wang 0001 |
Vis. Comput. | 2 |
| 2024 | Veintr: robust end-to-end full-hand vein identification with transformer
Shenglin Lu, Sheldon Fung, Wei Pan 0010, Nilmini Wickramasinghe, Xuequan Lu |
Vis. Comput. | 5 |
| 2024 | Segmentation-driven feature-preserving mesh denoising
Wei Pan 0010, Chaofan Dai, Richard Dazeley, Lei Wei 0002, Bernard Rolfe, Xuequan Lu |
Vis. Comput. | 7 |
| 2023 | Instance-Aware Domain Generalization for Face Anti-SpoofingabstractFace anti-spoofing (FAS) based on domain generalization (DG) has been recently studied to improve the generalization on unseen scenarios. Previous methods typically rely on domain labels to align the distribution of each domain for learning domain-invariant representations. However, artificial domain labels are coarse-grained and subjective, which cannot reflect real domain distributions accurately. Besides, such domain-aware methods focus on domain-level alignment, which is not fine-grained enough to ensure that learned representations are insensitive to domain styles. To address these issues, we propose a novel perspective for DG FAS that aligns features on the instance level without the need for domain labels. Specifically, Instance-Aware Domain Generalization framework is proposed to learn the generalizable feature by weakening the features' sensitivity to instance-specific styles. Concretely, we propose Asymmetric Instance Adaptive Whitening to adaptively eliminate the style-sensitive feature correlation, boosting the generalization. Moreover, Dynamic Kernel Generator and Categorical Style Assembly are proposed to first extract the instance-specific features and then generate the style-diversified features with large style shifts, respectively, further facilitating the learning of style-insensitive features. Extensive experiments and analysis demonstrate the superiority of our method over state-of-the-art competitors. Code will be publicly available at this link. Qianyu Zhou 0001, Ke-Yue Zhang, Taiping Yao, Xuequan Lu, Ran Yi 0002, Shouhong Ding, Lizhuang Ma |
CVPR | 4 |
| 2023 | Boosting Semi-Supervised Learning by Exploiting All Unlabeled DataabstractSemi-supervised learning (SSL) has attracted enormous attention due to its vast potential of mitigating the dependence on large labeled datasets. The latest methods (e.g., FixMatch) use a combination of consistency regularization and pseudo-labeling to achieve remarkable successes. However, these methods all suffer from the waste of complicated examples since all pseudo-labels have to be selected by a high threshold to filter out noisy ones. Hence, the examples with ambiguous predictions will not contribute to the training phase. For better leveraging all unlabeled examples, we propose two novel techniques: Entropy Meaning Loss (EML) and Adaptive Negative Learning (ANL). EML incorporates the prediction distribution of non-target classes into the optimization objective to avoid competition with target class, and thus generating more high-confidence predictions for selecting pseudo-label. ANL introduces the additional negative pseudo-label for all unlabeled data to leverage low-confidence examples. It adaptively allocates this label by dynamically evaluating the top-k performance of the model. EML and ANL do not introduce any additional parameter and hyperparameter. We integrate these techniques with FixMatch, and develop a simple yet powerful framework called FullMatch. Extensive experiments on several common SSL benchmarks (CIFAR-10/100, SVHN, STL-10 and ImageNet) demonstrate that FullMatch exceeds FixMatch by a large margin. Integrated with FlexMatch (an advanced FixMatch-based framework), we achieve state-of-the-art performance. Source code is available at https://github.com/megvii-research/FullMatch. Xin Tan 0002, Borui Zhao, Zhaowei Chen, Renjie Song, Jiajun Liang, Xuequan Lu |
CVPR | 7 |
| 2023 | IterativePFN: True Iterative Point Cloud FilteringabstractThe quality of point clouds is often limited by noise introduced during their capture process. Consequently, a fundamental 3D vision task is the removal of noise, known as point cloud filtering or denoising. State-of-the-art learning based methods focus on training neural networks to infer filtered displacements and directly shift noisy points onto the underlying clean surfaces. In high noise conditions, they iterate the filtering process. However, this iterative filtering is only done at test time and is less effective at ensuring points converge quickly onto the clean surfaces. We propose IterativePFN (iterative point cloud filtering network), which consists of multiple IterationModules that model the true iterative filtering process internally, within a single network. We train our IterativePFn network using a novel loss function that utilizes an adaptive ground truth target at each iteration to capture the relationship between intermediate filtering results during training. This ensures that the filtered results converge faster to the clean surfaces. Our method is able to obtain better performance compared to state-of-the-art methods. The source code can be found at: https://github.com/ddsediri/IterativePFN Dasith de Silva Edirimuni, Xuequan Lu, Zhiwen Shao, Gang Li 0009, Antonio Robles-Kelly, Ying He 0001 |
CVPR | 2 |
| 2023 | Motion-Aware Video Paragraph Captioning via Exploring Object-Centered Internal KnowledgeabstractVideo paragraph captioning task aims at generating a fine-grained, coherent and relevant paragraph for a video. Different from the images where objects are static, the temporal states of objects are changing in videos. The dynamic information could be contributed to understanding the whole video content. Existing works rarely put focus on modeling the dynamic changing state of the objects in the videos, causing the activities occurred in videos are poorly or wrongly depicted in paragraphs. To address this problem, we propose a novel Object State Tracking Network, which can capture the temporal state change of objects. However, due to the similarity of the consecutive frames in the videos, the information of the video is redundant and noisy. We further propose a semantic alignment mechanism, and enable the sentence information to refine the visual information. Extensive experiments on ActivityNet Captions demonstrate the effectiveness of our method. Yimin Hu, Guorui Yu, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Xuequan Lu, Shang Gao 0003 |
ICASSP | 6 |
| 2023 | Snow Removal in Video: A New Dataset and A Novel MethodabstractSnowfall is a common weather phenomenon that can severely affect computer vision tasks by obscuring objects and scenes. However, existing deep learning-based snow removal methods are designed for single images only. In this paper, we target a more complex task - video snow removal, which aims to restore the clear video from the snowy video. To facilitate this task, we propose the first high-quality video dataset, which simulates realistic physical characteristics of snow and haze using a rendering engine and augmentation techniques. We also develop a deep learning framework for video snow removal. Specifically, we propose a snow-query temporal aggregation module and a snow-aware contrastive learning loss function. The module aggregates features between video frames and removes snow effectively, while the loss function helps identify and eliminate snow features. We conduct extensive experiments and demonstrate that our proposed dataset is more realistic than previous datasets, and the models trained on it achieve better performance in real-world snowing images. Our proposed method outperforms state-of-the-art video and image-based methods on both synthetic and real snowy videos. Haoyu Chen 0003, Jinjin Gu, Xuequan Lu, Haoming Cai, Lei Zhu 0003 |
ICCV | 5 |
| 2023 | Weighted Point Cloud Normal EstimationabstractExisting normal estimation methods for point clouds are often less robust to severe noise and complex geometric structures. Also, they usually ignore the contributions of different neighbouring points during normal estimation, which leads to less accurate results. In this paper, we introduce a weighted normal estimation method for 3D point cloud data. We innovate in two key points: 1) we develop a novel weighted normal regression technique that predicts point-wise weights from local point patches and use them for robust, feature-preserving normal regression; 2) we propose to conduct contrastive learning between point patches and the corresponding ground-truth normals of the patches’ central points as a pre-training process to facilitate normal regression. Comprehensive experiments demonstrate that our method can robustly handle noisy and complex point clouds, achieving state-of-the-art performance on both synthetic and real-world datasets. Xuequan Lu, Di Shao, Xiao Liu 0004, Richard Dazeley, Antonio Robles-Kelly, Wei Pan 0010 |
ICME | 2 |
| 2023 | CAMG: Context-Aware Moment Graph Network for Multimodal Temporal Activity Localization via Language
Yuelin Hu, Yuanwu Xu, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Xuequan Lu, Shang Gao 0003 |
NLPCC (1) | 6 |
| 2023 | Random screening-based feature aggregation for point cloud denoisingabstractRaw point clouds captured by sensing devices are often contaminated with noise, which perturbs the fidelity of the original geometric information. Point cloud denoising is therefore an inseparable post-processing step, aiming to remove the noise in the point clouds. Existing point cloud denoising approaches are typically trained on datasets that have uniform point distributions and densities, making them unsuitable for effectively denoising point clouds with severe noise or irregular point distributions. In this paper, we introduce a novel random screening-based feature aggregation method for point cloud denoising. Our key insight is that merging features of dense and sparse points assists with enhancing the quality of point cloud denoising results. In specific, our approach involves randomly screening the features of local point patches and fusing richer geometric information of denser points into sparser point representations. Comprehensive experiments demonstrate that our method achieves state-of-the-art performance in the point cloud denoising task on both synthetic and real-world datasets. Wei Pan 0010, Xiao Liu 0004, Kui Su, Bernard Rolfe, Xuequan Lu |
Comput. Graph. | 6 |
| 2023 | Towards uniform point distribution in feature-preserving point cloud filteringabstractWhile a popular representation of 3D data, point clouds may contain noise and need filtering before use. Existing point cloud filtering methods either cannot preserve sharp features or result in uneven point distributions in the filtered output. To address this problem, this paper introduces a point cloud filtering method that considers both point distribution and feature preservation during filtering. The key idea is to incorporate a repulsion term with a data term in energy minimization. The repulsion term is responsible for the point distribution, while the data term aims to approximate the noisy surfaces while preserving geometric features. This method is capable of handling models with fine-scale features and sharp features. Extensive experiments show that our method quickly yields good results with relatively uniform point distribution. Shuaijun Chen, Jinxi Wang, Wei Pan 0010, Shang Gao 0003, Meili Wang 0001, Xuequan Lu |
Comput. Vis. Media | 6 |
| 2023 | 3D face recognition: A comprehensive survey in 2022abstractIn the past ten years, research on face recognition has shifted to using 3D facial surfaces, as 3D geometric information provides more discriminative features. This comprehensive survey reviews 3D face recognition techniques developed in the past decade, both conventional methods and deep learning methods. These methods are evaluated with detailed descriptions of selected representative works. Their advantages and disadvantages are summarized in terms of accuracy, complexity, and robustness to facial variations (expression, pose, occlusion, etc.). A review of 3D face databases is also provided, and a discussion of future research challenges and directions of the topic. Yaping Jing, Xuequan Lu, Shang Gao 0003 |
Comput. Vis. Media | 2 |
| 2023 | Self-Adversarial Disentangling for Specific Domain AdaptationabstractDomain adaptation aims to bridge the domain shifts between the source and the target domain. These shifts may span different dimensions such as fog, rainfall, etc. However, recent methods typically do not consider explicit prior knowledge about the domain shifts on a specific dimension, thus leading to less desired adaptation performance. In this article, we study a practical setting called Specific Domain Adaptation (SDA) that aligns the source and target domains in a demanded-specific dimension. Within this setting, we observe the intra-domain gap induced by different domainness (i.e., numerical magnitudes of domain shifts in this dimension) is crucial when adapting to a specific domain. To address the problem, we propose a novel Self-Adversarial Disentangling (SAD) framework. In particular, given a specific dimension, we first enrich the source domain by introducing a domainness creator with providing additional supervisory signals. Guided by the created domainness, we design a self-adversarial regularizer and two loss functions to jointly disentangle the latent representations into domainness-specific and domainness-invariant features, thus mitigating the intra-domain gap. Our method can be easily taken as a plug-and-play framework and does not introduce any extra costs in the inference time. We achieve consistent improvements over state-of-the-art methods in both object detection and semantic segmentation. Qianyu Zhou 0001, Jiangmiao Pang, Xuequan Lu, Lizhuang Ma |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Robust image clustering via context-aware contrastive graph learning
Uno Fang, Jianxin Li 0001, Xuequan Lu, Ajmal Mian, Zhaoquan Gu |
Pattern Recognit. | 3 |
| 2023 | Graph classification via discriminative edge feature learningabstractSpectral graph convolutional neural networks (GCNNs) have been producing encouraging results in graph classification tasks. However, most spectral GCNNs utilize fixed graphs when aggregating node features, while omitting edge feature learning and failing to get an optimal graph structure. Moreover, many existing graph datasets do not provide initialized edge features, further restraining the ability of learning edge features via spectral GCNNs. In this paper, we try to address this issue by designing an edge feature scheme and an add-on layer between every two stacked graph convolution layers in spectral GCNN. Both are lightweight while effective in filling the gap between edge feature learning and performance enhancement of graph classification. The edge feature scheme makes edge features adapt to node representations at different spectral graph convolution layers. The add-on layers help adjust the edge features to an optimal graph structure. To test the effectiveness of our method, we take Euclidean positions as initial node features and extract graphs with semantic information from point cloud objects. The node features of our extracted graphs are more scalable for edge feature learning than most existing graph datasets (in one-hot encoded label format). Three new graph datasets are constructed based on ModelNet40, ModelNet10 and ShapeNet Part datasets. Experimental results show that our method outperforms state-of-the-art graph classification methods on the new datasets. Our code and the constructed graph datasets will be released to the community. Xuequan Lu, Shang Gao 0003, Antonio Robles-Kelly, Yuejie Zhang |
Pattern Recognit. | 2 |
| 2023 | Focusing Fine-Grained Action by Self-Attention-Enhanced Graph Neural Networks With Contrastive LearningabstractWith the aid of graph convolution neural network and transformer model, human action recognition has achieved significant performance based on skeleton data. However, the majority of existing works rarely focus on identifying fine-grained motion information (i.e., “read”, “write”, etc.). Furthermore, they tend to explore correlations between joints and bones ignoring the angular information. Consequently, the recognition accuracy for fine-grained actions with most models is still less desired. To address this issue, we first attempt to bring angular information as a complement to familiar joint and bone information, while learning the potential dependencies of the three kinds of information using graph neural networks. Based on this, we propose a self-attention-enhanced graph neural network (SAE-GNN), which consists of a kernel-unified graph convolution (KUGC) module and an enhanced attention graph convolution (EAGC) module. The KUGC module is devised to effectively extract rich features in the skeleton information. The EAGC consisting of a multi-scale enhanced graph convolution block and a multi-headed self-attention block is designed to learn the potential high-level semantic information in the features. Besides, we introduce contrastive learning in the two blocks to enhance feature representation by maximizing their mutual information. We conduct extensive experiments on four publicly available datasets, and results show that our model outperforms state-of-the-art methods in recognizing fine-grained actions. Pei Geng, Xuequan Lu, Chunyu Hu 0001, Hong Liu 0013, Lei Lyu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Context-Aware Mixup for Domain Adaptive Semantic SegmentationabstractUnsupervised domain adaptation (UDA) aims to adapt a model of the labeled source domain to an unlabeled target domain. Existing UDA-based semantic segmentation approaches always reduce the domain shifts in pixel level, feature level, and output level. However, almost all of them largely neglect the contextual dependency, which is generally shared across different domains, leading to less-desired performance. In this paper, we propose a novel Context-Aware Mixup (CAMix) framework for domain adaptive semantic segmentation, which exploits this important clue of context-dependency as explicit prior knowledge in a fully end-to-end trainable manner for enhancing the adaptability toward the target domain. Firstly, we present a contextual mask generation strategy by leveraging the accumulated spatial distributions and prior contextual relationships. The generated contextual mask is critical in this work and will guide the context-aware domain mixup on three different levels. Besides, provided the context knowledge, we introduce a significance-reweighted consistency loss to penalize the inconsistency between the mixed student prediction and the mixed teacher prediction, which alleviates the negative transfer of the adaptation, e.g., early performance degradation. Extensive experiments and analysis demonstrate the effectiveness of our method against the state-of-the-art approaches on widely-used UDA benchmarks. Qianyu Zhou 0001, Zhengyang Feng, Jiangmiao Pang, Xuequan Lu, Jianping Shi, Lizhuang Ma |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2023 | 3D Intracranial Aneurysm Classification and Segmentation via Unsupervised Dual-Branch LearningabstractIntracranial aneurysms are common nowadays and how to detect them intelligently is of great significance in digital health. Whereas most existing deep learning research focused on medical images in a supervised way, we introduce an unsupervised method for the detection of intracranial aneurysms based on 3D point cloud data. In particular, our method consists of two stages: unsupervised pre-training and downstream tasks. As for the former, the main idea is to pair each point cloud with its jittering counterpart and maximise their correspondence. Then we design a dual-branch contrastive network with an encoder for each branch and a subsequent common projection head. As for the latter, we design simple networks for supervised classification and segmentation training. Experiments on the public dataset (IntrA) show that our unsupervised method achieves comparable or even better performance than some state-of-the-art supervised techniques, and it is most prominent in the detection of aneurysmal vessels. Experiments on the ModelNet-40 also show that our method achieves the accuracy of 90.79% which outperforms existing state-of-the-art unsupervised models. Di Shao, Xuequan Lu, Xiao Liu 0004 |
IEEE J. Biomed. Health Informatics | 2 |
| 2023 | Skeleton-Based Action Recognition Through Contrasting Two-Stream Spatial-Temporal NetworksabstractFor pursuing accurate skeleton-based action recognition, most prior methods use the strategy of combining Graph Convolution Networks (GCNs) with attention-based methods in a serial way. However, they regard the human skeleton as a complete graph, resulting in less variations between different actions (e.g., the connection between the elbow and head in action “clapping hands”). For this, we propose a novel Contrastive GCN-Transformer Network (ConGT) which fuses the spatial and temporal modules in a parallel way. The ConGT involves two parallel streams: Spatial-Temporal Graph Convolution stream (STG) and Spatial-Temporal Transformer stream (STT). The STG is designed to obtain action representations maintaining the natural topology structure of the human skeleton. The STT is devised to acquire action representations containing the global relationships among joints. Since the action representations produced from these two streams contain different characteristics, and each of them knows little information of the other, we introduce the contrastive learning paradigm to guide their output representations of the same sample to be as close as possible in a self-supervised manner. Through the contrastive learning, they can learn information from each other to enrich the action features by maximizing the mutual information between the two types of action representations. To further improve action recognition accuracy, we introduce the Cyclical Focal Loss (CFL) which can focus on confident training samples in early training epochs, with an increasing focus on hard samples during the middle epochs. We conduct experiments on three benchmark datasets, which demonstrate that our model achieves state-of-the-art performance in action recognition. Chen Pang 0001, Xuequan Lu, Lei Lyu 0001 |
IEEE Trans. Multim. | 2 |
| 2023 | Federated Learning via Disentangled Information BottleneckabstractExisting Federated Learning (FL) algorithms generally suffer from high communication costs and data heterogeneity due to the use of conventional loss function for local model update and the equal consideration of each local model for global model aggregation. In this paper, we propose a novel FL approach to address the above issues. For local model update, we propose a disentangled Information Bottleneck (IB) principle-based loss function. For global model aggregation, we suggest a model selection strategy based on Mutual Information (MI). Particularly, we design a Lagrangian-based loss function using the IB principle and “disentanglement” for maximizing MI between the ground truth and model prediction and minimizing MI between the intermediate representations. We calculate MI ratio between the ground truth and model prediction, and between the original input and ground truth to select the effective models for aggregation. We analyze the theoretical optimal cost of the loss function and manifest optimal convergence rate, and quantify the outlier robustness of the aggregation scheme. Experiments demonstrate the superiority of the proposed FL approach, in terms of testing performance and communication speedup (i.e., 3.00-14.88 times for IID MNIST, 2.5-50.75 times for non-IID MNIST, 1.87-18.40 times for IID CIFAR-10, and 1.24-2.10 times for non-IID MIMIC-III). Md Palash Uddin, Yong Xiang 0001, Xuequan Lu, John Yearwood, Longxiang Gao |
IEEE Trans. Serv. Comput. | 3 |
| 2022 | CREAM: Weakly Supervised Object Localization via Class RE-Activation MappingabstractWeakly Supervised Object Localization (WSOL) aims to localize objects with image-level supervision. Existing works mainly rely on Class Activation Mapping (CAM) de-rived from a classification model. However, CAM-based methods usually focus on the most discriminative parts of an object (i.e., incomplete localization problem). In this paper, we empirically prove that this problem is associated with the mixup of the activation values between less discrimi-native foreground regions and the background. To address it, we propose Class RE-Activation Mapping (CREAM), a novel clustering-based approach to boost the activation values of the integral object regions. To this end, we in-troduce class-specific foreground and background context embeddings as cluster centroids. A CAM-guided momen-tum preservation strategy is developed to learn the context embeddings during training. At the inference stage, the re-activation mapping is formulated as a parameter es-timation problem under Gaussian Mixture Model, which can be solved by deriving an unsupervised Expectation- Maximization based soft-clustering algorithm. By simply integrating CREAM into various WSOL approaches, our method significantly improves their performance. CREAM achieves the state-of-the-art performance on CUB, ILSVRC and OpenImages benchmark datasets. Code will be avail-able at https://github.com/lazzcharles/CREAM. Jilan Xu, Junlin Hou, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Xuequan Lu, Shang Gao 0003 |
CVPR | 7 |
| 2022 | Domain Adaptive Semantic Segmentation via Regional Contrastive Consistency RegularizationabstractUnsupervised domain adaptation (UDA) for semantic seg-mentation has been well-studied in recent years. However, most existing works largely neglect the local regional consis-tency across different domains, and are less robust to changes in outdoor environments. In this paper, we propose a novel and fully end-to-end trainable approach, called regional contrastive consistency regularization (RCCR) for domain adaptive semantic segmentation. Our core idea is to pull the sim-ilar regional features extracted from the same location of dif-ferent images, i.e., the original image and augmented image, to be closer, and meanwhile push the features from the dif-ferent locations of the two images to be separated. We pro-pose a region-wise contrastive loss with two sampling strate-gies to realize effective regional consistency. Besides, we present momentum projection heads, where the teacher pro-jection head is the exponential moving average of the student. Finally, a memory bank mechanism is designed to learn more robust and stable region-wise features under varying environ-ments. Extensive experiments demonstrate that our approach outperforms the state-of-the-art methods. Qianyu Zhou 0001, Chuyun Zhuang, Ran Yi 0002, Xuequan Lu, Lizhuang Ma |
ICME | 4 |
| 2022 | Semantic-Driven Saliency-Context Separation for Video CaptioningabstractVideo captioning aims at generating a natural language de-scription for a given video clip including not only salient sce-narios but also contextual scenarios. The former reveal the highlight of a video and are usually the focus of most existing captioning methods. The latter, however, are not well ex-plored and even ignored easily, though they may provide cer-tain detailed and latent information that can help with a better understanding of the video. To effectively exploit the infor-mation contained in both, a novel video captioning network is proposed. It has two key modules: Cross-Modality Selection (CMS) and Saliency-Context Adaptive Decoder (SCAD). Specifically, CMS mainly focuses on utilizing the semantic information to distinguish saliency and context. Meanwhile, SCAD adaptively identifies both the saliency and context to generate more detailed and precise captions. Experiments on two benchmark datasets, i.e., MSVD and MSR-VTT, demon-strate the effectiveness of our model through the comparison with state-of-the-art methods. Heming Jing, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Xuequan Lu, Shang Gao 0003 |
ICME | 6 |
| 2022 | Deep Point Cloud Normal Estimation Via Triplet LearningabstractCurrent normal estimation methods for 3D point clouds often show limited accuracy in predicting normals at sharp features (e.g., edges and corners) and less robustness to noise. In this paper, we propose a novel normal estimation method for point clouds which consists of two phases: (a) feature encoding to learn representations of local patches, and (b) normal estimation that takes the learned representation as input and regresses the normal vector. We are motivated that local patches on isotropic and anisotropic surfaces respectively have similar and distinct normals, and these separable features or representations can be learned to facilitate normal estimation. To realise this, we design a triplet learning network for feature encoding and a normal estimation network to regress normals. Despite having a smaller network size compared with most other methods, experiments show that our method preserves sharp features and achieves better normal estimation results especially on computer-aided design (CAD) shapes. Xuequan Lu, Dasith de Silva Edirimuni, Xiao Liu 0004, Antonio Robles-Kelly |
ICME | 2 |
| 2022 | STDNet: Spatio-Temporal Decomposed Network for Video GroundingabstractPrevious methods for video grounding treated either the query or the video as a whole, while neglecting their respective semantics in the orthogonal space and time dimensions. Since spatial semantics appears frequently in a video, temporal semantics is more discriminative and deserves more attention. Based on such considerations, we propose a novel Spatio-Temporal Decomposed Network (STDNet) which decomposes the query and the video into their spatial and temporal semantics, respectively. Specifically, spatial and temporal words are selected from the query, and the video is split into two pathways. Spatial cross-modal attention is computed first and serves as prior knowledge for temporal attention. A new localization strategy is also devised which regresses the segment's start conditioned on the end and essentially breaks the independence assumption made in previous methods. Experimental results on three public benchmark datasets show that our STDNet outperforms the state-of-the-art methods. Yuanwu Xu, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Xuequan Lu, Shang Gao 0003 |
ICME | 6 |
| 2022 | Anatomical Landmarks Localization for 3D Foot Point Clouds
Sheldon Fung, Xuequan Lu, Mantas Mykolaitis, Muhammad Imran Razzak, Gediminas Kostkevicius, Domantas Ozerenskis |
ICONIP (3) | 2 |
| 2022 | TCCNet: Temporally Consistent Context-Free Network for Semi-supervised Video Polyp SegmentationabstractAutomatic video polyp segmentation (VPS) is highly valued for the early diagnosis of colorectal cancer. However, existing methods are limited in three respects: 1) most of them work on static images, while ignoring the temporal information in consecutive video frames; 2) all of them are fully supervised and easily overfit in presence of limited annotations; 3) the context of polyp (i.e., lumen, specularity and mucosa tissue) varies in an endoscopic clip, which may affect the predictions of adjacent frames. To resolve these challenges, we propose a novel Temporally Consistent Context-Free Network (TCCNet) for semi-supervised VPS. It contains a segmentation branch and a propagation branch with a co-training scheme to supervise the predictions of unlabeled image. To maintain the temporal consistency of predictions, we design a Sequence-Corrected Reverse Attention module and a Propagation-Corrected Reverse Attention module. A Context-Free Loss is also proposed to mitigate the impact of varying contexts. Extensive experiments show that even trained under 1/15 label ratio, TCCNet is comparable to the state-of-the-art fully supervised methods for VPS. Also, TCCNet surpasses existing semi-supervised methods for natural image and other medical image segmentation tasks. Jilan Xu, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Xuequan Lu, Shang Gao 0003 |
IJCAI | 7 |
| 2022 | Spoof Face Detection Via Semi-Supervised Adversarial TrainingabstractFace spoofing causes severe security threats in face recognition systems. The previous anti-spoofing mainly focused on supervised techniques, typically with either binary or auxiliary supervision. Most of them have to ‘see’ both spoofing face data and live face data during training to realize the task of face anti-spoofing. In this paper, we propose a semi-supervised adversarial learning framework for spoof face detection, which largely relaxes the supervision condition. To capture the underlying structure of live face data in latent representation space, we propose to train the live face data only, with a convolutional Encoder-Decoder network acting as a Generator, and a second convolutional network serving as a Discriminator. The generator and discriminator are trained by competing with each other while collaborating to understand the live faces. Since the spoof face detection is video-based (i.e., temporal information), we intuitively take the optical flow maps converted from consecutive video frames as input. Our approach is free of the spoof faces, thus being robust and general to different types of face spoofing (even unknown spoofing). Experiments on cross-dataset tests show that our semi-supervised method achieves better or comparable results to state-of-the-art supervised techniques. We also conduct ablation studies for the proposed method. Chengwei Chen, Yaping Jing, Xuequan Lu, Wang Yuan, Lizhuang Ma |
IJCNN | 3 |
| 2022 | In-Place Gestures Classification via Long-term Memory Augmented NetworkabstractIn-place gesture-based virtual locomotion techniques enable users to control their viewpoint and intuitively move in the 3D virtual environment. A key research problem is to accurately and quickly recognize in-place gestures, since they can trigger specific movements of virtual viewpoints and enhance user experience. However, to achieve real-time experience, only short-term sensor sequence data (up to about 300ms, 6 to 10 frames) can be taken as input, which actually affects the classification performance due to limited spatiotemporal information. In this paper, we propose a novel long-term memory augmented network for in-place gestures classification. It takes as input both short-term gesture sequence samples and their corresponding long-term sequence samples that provide extra relevant spatio-temporal information in the training phase. We store long-term sequence features with an external memory queue. In addition, we design a memory augmented loss to help cluster features of the same class and push apart features from different classes, thus enabling our memory queue to memorize more relevant long-term sequence features. In the inference phase, we input only short-term sequence samples to recall the stored features accordingly, and fuse them together to predict the gesture class. We create a large-scale in-place gestures dataset from 25 participants with 11 gestures. Our method achieves a promising accuracy of 95.1% with a latency of 192ms, and an accuracy of 97.3% with a latency of 312ms, and is demonstrated to be superior to recent in-place gesture classification techniques. User study also validates our approach. Our source code and dataset will be made available to the community. Lizhi Zhao, Xuequan Lu, Qianyue Bao, Meili Wang 0001 |
ISMAR | 2 |
| 2022 | A Differential Privacy Mechanism for Deceiving Cyber Attacks in IoT Networks
Guizhen Yang, Mengmeng Ge 0001, Shang Gao 0003, Xuequan Lu, Leo Yu Zhang, Robin Doss |
NSS | 4 |
| 2022 | Rethinking Point Cloud Filtering: A Non-Local Position Based Approach
Jinxi Wang, Jincen Jiang, Xuequan Lu, Meili Wang 0001 |
Comput. Aided Des. | 3 |
| 2022 | SPCNet: Stepwise Point Cloud Completion NetworkabstractAbstract How will you repair a physical object with large missings? You may first recover its global yet coarse shape and stepwise increase its local details. We are motivated to imitate the above physical repair procedure to address the point cloud completion task. We propose a novel stepwise point cloud completion network (SPCNet) for various 3D models with large missings. SPCNet has a hierarchical bottom‐to‐up network architecture. It fulfills shape completion in an iterative manner, which 1) first infers the global feature of the coarse result; 2) then infers the local feature with the aid of global feature; and 3) finally infers the detailed result with the help of local feature and coarse result. Beyond the wisdom of simulating the physical repair, we newly design a cycle loss to enhance the generalization and robustness of SPCNet. Extensive experiments clearly show the superiority of our SPCNet over the state‐of‐the‐art methods on 3D point clouds with large missings. Code is available at https://github.com/1127368546/SPCNet . Honghua Chen, Xuequan Lu, Zhe Zhu, Jun Wang 0039, Weiming Wang 0002, Fu Lee Wang, Mingqiang Wei |
Comput. Graph. Forum | 3 |
| 2022 | SO(3)-Pose: SO(3)-Equivariance Learning for 6D Object Pose EstimationabstractAbstract 6D pose estimation of rigid objects from RGB‐D images is crucial for object grasping and manipulation in robotics. Although RGB channels and the depth (D) channel are often complementary, providing respectively the appearance and geometry information, it is still non‐trivial on how to fully benefit from the two cross‐modal data. From the simple yet new observation, when an object rotates, its semantic label is invariant to the pose while its keypoint offset direction is variant to the pose. To this end, we present SO(3)‐Pose, a new representation learning network to explore SO(3)‐equivariant and SO(3)‐invariant features from the depth channel for pose estimation. The SO(3)‐invariant features facilitate to learn more distinctive representations for segmenting objects with similar appearance from RGB channels. The SO(3)‐equivariant features communicate with RGB features to deduce the (missed) geometry for detecting keypoints of an object with the reflective surface from the depth channel. Unlike most of existing pose estimation methods, our SO(3)‐Pose not only implements the information communication between the RGB and depth channels, but also naturally absorbs the SO(3)‐equivariance geometry knowledge from depth images, leading to better appearance and geometry representation learning. Comprehensive experiments show that our method achieves the state‐of‐the‐art performance on three benchmarks. Code is available at https://github.com/phaoran9999/SO3-Pose . Haoran Pan, Jun Zhou 0007, Xuequan Lu, Weiming Wang 0002, Xuefeng Yan 0001, Mingqiang Wei |
Comput. Graph. Forum | 4 |
| 2022 | Uncertainty-aware consistency regularization for cross-domain semantic segmentation
Qianyu Zhou 0001, Zhengyang Feng, Xuequan Lu, Jianping Shi, Lizhuang Ma |
Comput. Vis. Image Underst. | 5 |
| 2022 | A robust scheme for copy detection of 3D object point clouds
Xuequan Lu, Wenzhi Chen |
Neurocomputing | 2 |
| 2022 | CGSNet: Contrastive Graph Self-Attention Network for Session-based Recommendation
Fuyun Wang, Xuequan Lu, Lei Lyu 0001 |
Knowl. Based Syst. | 2 |
| 2022 | DMT: Dynamic mutual training for semi-supervised learning
Zhengyang Feng, Qianyu Zhou 0001, Xin Tan 0002, Xuequan Lu, Jianping Shi, Lizhuang Ma |
Pattern Recognit. | 6 |
| 2022 | Example-based color transfer with Gaussian mixture modeling
Chunzhi Gu, Xuequan Lu, Chao Zhang 0030 |
Pattern Recognit. | 2 |
| 2022 | Unconstrained Facial Action Unit Detection via Latent Feature DomainabstractFacial action unit (AU) detection in the wild is a challenging problem, due to the unconstrained variability in facial appearances and the lack of accurate annotations. Most existing methods depend on either impractical labor-intensive labeling or inaccurate pseudo labels. In this paper, we propose an end-to-end unconstrained facial AU detection framework based on domain adaptation, which transfers accurate AU labels from a constrained source domain to an unconstrained target domain by exploiting labels of AU-related facial landmarks. Specifically, we map a source image with label and a target image without label into a latent feature domain by combining source landmark-related feature with target landmark-free feature. Due to the combination of source AU-related information and target AU-free information, the latent feature domain with transferred source label can be learned by maximizing the target-domain AU detection performance. Moreover, we introduce a novel landmark adversarial loss to disentangle the landmark-free feature from the landmark-related feature by treating the adversarial learning as a multi-player minimax game. Our framework can also be naturally extended for use with target-domain pseudo AU labels. Extensive experiments show that our method soundly outperforms lower-bounds and upper-bounds of the basic model, as well as state-of-the-art approaches on the challenging in-the-wild benchmarks. The code is available athttps://github.com/ZhiwenShao/ADLD. Zhiwen Shao, Jianfei Cai 0001, Tat-Jen Cham, Xuequan Lu, Lizhuang Ma |
IEEE Trans. Affect. Comput. | 4 |
| 2022 | Low Rank Matrix Approximation for 3D Geometry FilteringabstractWe propose a robust normal estimation method for both point clouds and meshes using a low rank matrix approximation algorithm. First, we compute a local isotropic structure for each point and find its similar, non-local structures that we organize into a matrix. We then show that a low rank matrix approximation algorithm can robustly estimate normals for both point clouds and meshes. Furthermore, we provide a new filtering method for point cloud data to smooth the position data to fit the estimated normals. We show the applications of our method to point cloud filtering, point set upsampling, surface reconstruction, mesh denoising, and geometric texture removal. Our experiments show that our method generally achieves better results than existing methods. Xuequan Lu, Scott Schaefer, Jun Luo 0001, Lizhuang Ma, Ying He 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2021 | PIT: Position-Invariant Transform for Cross-FoV Domain AdaptationabstractCross-domain object detection and semantic segmentation have witnessed impressive progress recently. Existing approaches mainly consider the domain shift resulting from external environments including the changes of background, illumination or weather, while distinct camera intrinsic parameters appear commonly in different domains and their influence for domain adaptation has been very rarely explored. In this paper, we observe that the Field of View (FoV) gap induces noticeable instance appearance differences between the source and target domains. We further discover that the FoV gap between two domains impairs domain adaptation performance under both the FoV-increasing (source FoV < target FoV) and FoV-decreasing cases. Motivated by the observations, we propose the Position-Invariant Transform (PIT) to better align images in different domains. We also introduce a reverse PIT for mapping the transformed/aligned images back to the original image space, and design a loss reweighting strategy to accelerate the training process. Our method can be easily plugged into existing cross-domain detection/segmentation frameworks, while bringing about negligible computational overhead. Extensive experiments demonstrate that our method can soundly boost the performance on both cross-domain object detection and segmentation for state-of-the-art techniques. Our code is available at https://github.com/sheepooo/PIT-Position-Invariant-Transform. Qianyu Zhou 0001, Zhengyang Feng, Xuequan Lu, Jianping Shi, Lizhuang Ma |
ICCV | 6 |
| 2021 | DeepfakeUCL: Deepfake Detection via Unsupervised Contrastive LearningabstractFace deepfake detection has seen impressive results recently. Nearly all existing deep learning techniques for face deepfake detection are fully supervised and require labels during training. In this paper, we design a novel deepfake detection method via unsupervised contrastive learning. We first generate two different transformed versions of an image and feed them into two sequential sub-networks, i.e., an encoder and a projection head. The unsupervised training is achieved by maximizing the correspondence degree of the outputs of the projection head. To evaluate the detection performance of our unsupervised method, we further use the unsupervised features to train an efficient linear classification network. Extensive experiments show that our unsupervised learning method enables comparable detection performance to state-of-the-art supervised techniques, in both the intra- and inter-dataset settings. We also conduct ablation studies for our method. Sheldon Fung, Xuequan Lu, Chao Zhang 0030, Chang-Tsun Li |
IJCNN | 2 |
| 2021 | Classifying In-Place Gestures with End-to-End Point Cloud LearningabstractWalking in place for moving through virtual environments has attracted noticeable attention recently. Recent attempts focused on training a classifier to recognize certain patterns of gestures (e.g., standing, walking, etc) with the use of neural networks like CNN or LSTM. Nevertheless, they often consider very few types of gestures and/or induce less desired latency in virtual environments. In this paper, we propose a novel framework for accurate and efficient classification of in-place gestures. Our key idea is to treat several consecutive frames as a “point cloud”. The HMD and two VIVE trackers provide three points in each frame, with each point consisting of 12-dimensional features (i.e., three-dimensional position coordinates, velocity, rotation, angular velocity). We create a dataset consisting of 9 gesture classes for virtual in-place locomotion. In addition to the supervised point-based network, we also take unsupervised domain adaptation into account due to inter-person variations. To this end, we develop an end-to-end joint framework involving both a supervised loss for supervised point learning and an unsupervised loss for unsupervised domain adaptation. Experiments demonstrate that our approach generates very promising outcomes, in terms of high overall classification accuracy (95.0%) and real-time performance (192ms latency). We will release our dataset and source code to the community. Lizhi Zhao, Xuequan Lu, Meili Wang 0001 |
ISMAR | 2 |
| 2021 | Automated Security Assessment for the Internet of ThingsabstractInternet of Things (IoT) based applications face an increasing number of potential security risks, which need to be systematically assessed and addressed. Expert-based manual assessment of IoT security is a predominant approach, which is usually inefficient. To address this problem, we propose an automated security assessment framework for IoT networks. Our framework first leverages machine learning and natural language processing to analyze vulnerability descriptions for predicting vulnerability metrics. The predicted metrics are then input into a two-layered graphical security model, which consists of an attack graph at the upper layer to present the network connectivity and an attack tree for each node in the network at the bottom layer to depict the vulnerability information. This security model automatically assesses the security of the IoT network by capturing potential attack paths. We evaluate the viability of our approach using a proof-of-concept smart building system model which contains a variety of real-world IoT devices and poten-tial vulnerabilities. Our evaluation of the proposed framework demonstrates its effectiveness in terms of automatically predicting the vulnerability metrics of new vulnerabilities with more than 90% accuracy, on average, and identifying the most vulnerable attack paths within an IoT network. The produced assessment results can serve as a guideline for cybersecurity professionals to take further actions and mitigate risks in a timely manner. Xuanyu Duan, Mengmeng Ge 0001, Triet Huynh Minh Le, Faheem Ullah, Shang Gao 0003, Xuequan Lu, Muhammad Ali Babar 0001 |
PRDC | 6 |
| 2021 | Scene image representation by foreground, background and hybrid features
Chiranjibi Sitaula, Yong Xiang 0001, Sunil Aryal, Xuequan Lu |
Expert Syst. Appl. | 4 |
| 2021 | Self-supervised cross-iterative clustering for unlabeled plant disease images
Uno Fang, Jianxin Li 0001, Xuequan Lu, Longxiang Gao, Mumtaz Ali 0003, Yong Xiang 0001 |
Neurocomputing | 3 |
| 2021 | Bas-relief layout arrangement via automatic method optimizationabstractAbstract It is significant to achieve automatic arrangement for bas‐relief layout which can be noticeably more efficient than the time‐consuming manual process. In fact, nearly none work has been reported in terms of bas‐relief layout arrangement. In this paper, we propose a novel approach to tackle this problem. Specifically, we first identify the evaluation indicators to account for different aesthetic factors, and model the goodness of each indicator. We then cast the bas‐relief layout as a combinatorial optimization problem based on those evaluation indicators and a geometric mean model. The contribution of this paper is to propose an objective function for bas‐relief layout and apply simulated annealing algorithm for optimization. Experiments show that our method is effective, in terms of layout arrangement for bas‐relief generation. In addition, this method can synthesize a few models arrangement and investigate which evaluation indicators will affect the aesthetic perception of the bas‐relief. Jiahui Mao, Meili Wang 0001, Jian Chang 0001, Xuequan Lu |
Comput. Animat. Virtual Worlds | 6 |
| 2021 | KeyFrame extraction for human motion capture data via multiple binomial fittingabstractAbstract In this paper, we make two contributions. The first is to propose a new keyframe extraction algorithm, which reduces the keyframe redundancy and reduces the motion sequence reconstruction error. Secondly, a new motion sequence reconstruction method is proposed, which further reduces the error of motion sequence reconstruction. Specifically, we treated the input motion sequence as curves, then the binomial fitting was extended to obtain the points where the slope changes dramatically in the vicinity. Then we took these points as inputs to obtain keyframes by density clustering. Finally, the motion curves were segmented by keyframes and the segmented curves were fitted by binomial formula again to obtain the binomial parameters for motion reconstruction. Experiments show that our methods outperform existing techniques, in terms of reconstruction error. Chenxu Xu, Yanran Li, Xuequan Lu, Meili Wang 0001, Xiaosong Yang |
Comput. Animat. Virtual Worlds | 4 |
| 2021 | Content and context features for scene image representation
Chiranjibi Sitaula, Sunil Aryal, Yong Xiang 0001, Anish Basnet, Xuequan Lu |
Knowl. Based Syst. | 5 |
| 2021 | Blur Removal Via Blurred-Noisy Image PairabstractComplex blur such as the mixup of space-variant and space-invariant blur, which is hard to model mathematically, widely exists in real images. In this article, we propose a novel image deblurring method that does not need to estimate blur kernels. We utilize a pair of images that can be easily acquired in low-light situations: (1) a blurred image taken with low shutter speed and low ISO noise; and (2) a noisy image captured with high shutter speed and high ISO noise. Slicing the blurred image into patches, we extend the Gaussian mixture model (GMM) to model the underlying intensity distribution of each patch using the corresponding patches in the noisy image. We compute patch correspondences by analyzing the optical flow between the two images. The Expectation Maximization (EM) algorithm is utilized to estimate the parameters of GMM. To preserve sharp features, we add an additional bilateral term to the objective function in the M-step. We eventually add a detail layer to the deblurred image for refinement. Extensive experiments on both synthetic and real-world data demonstrate that our method outperforms state-of-the-art techniques, in terms of robustness, visual quality, and quantitative metrics. Chunzhi Gu, Xuequan Lu, Ying He 0001, Chao Zhang 0030 |
IEEE Trans. Image Process. | 2 |
| 2021 | Explicit Facial Expression Transfer via Fine-Grained RepresentationsabstractFacial expression transfer between two unpaired images is a challenging problem, as fine-grained expression is typically tangled with other facial attributes. Most existing methods treat expression transfer as an application of expression manipulation, and use predicted global expression, landmarks or action units (AUs) as a guidance. However, the prediction may be inaccurate, which limits the performance of transferring fine-grained expression. Instead of using an intermediate estimated guidance, we propose to explicitly transfer facial expression by directly mapping two unpaired input images to two synthesized images with swapped expressions. Specifically, considering AUs semantically describe fine-grained expression details, we propose a novel multi-class adversarial training method to disentangle input images into two types of fine-grained representations: AU-related feature and AU-free feature. Then, we can synthesize new images with preserved identities and swapped expressions by combining AU-free features with swapped AU-related features. Moreover, to obtain reliable expression transfer results of the unpaired input, we introduce a swap consistency loss to make the synthesized images and self-reconstructed images indistinguishable. Extensive experiments show that our approach outperforms the state-of-the-art expression manipulation methods for transferring fine-grained expressions while preserving other attributes including identity and pose. Zhiwen Shao, Hengliang Zhu, Junshu Tang, Xuequan Lu, Lizhuang Ma |
IEEE Trans. Image Process. | 4 |
| 2021 | Mutual Information Driven Federated LearningabstractFederated Learning (FL) is an emerging research field that yields a global trained model from different local clients without violating data privacy. Existing FL techniques often ignore the effective distinction between local models and the aggregated global model when doing the client-side weight update, as well as the distinction of local models for the server-side aggregation. In this article, we propose a novel FL approach with resorting to mutual information (MI). Specifically, in client-side, the weight update is reformulated through minimizing the MI between local and aggregated models and employing Negative Correlation Learning (NCL) strategy. In server-side, we select top effective models for aggregation based on the MI between an individual local model and its previous aggregated model. We also theoretically prove the convergence of our algorithm. Experiments conducted on MNIST, CIFAR-10, ImageNet, and the clinical MIMIC-III datasets manifest that our method outperforms the state-of-the-art techniques in terms of both communication and testing performance. Md Palash Uddin, Yong Xiang 0001, Xuequan Lu, John Yearwood, Longxiang Gao |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2021 | Pointfilter: Point Cloud Filtering via Encoder-Decoder ModelingabstractPoint cloud filtering is a fundamental problem in geometry modeling and processing. Despite of significant advancement in recent years, the existing methods still suffer from two issues: 1) they are either designed without preserving sharp features or less robust in feature preservation; and 2) they usually have many parameters and require tedious parameter tuning. In this article, we propose a novel deep learning approach that automatically and robustly filters point clouds by removing noise and preserving their sharp features. Our point-wise learning architecture consists of an encoder and a decoder. The encoder directly takes points (a point and its neighbors) as input, and learns a latent representation vector which goes through the decoder to relate the ground-truth position with a displacement vector. The trained neural network can automatically generate a set of clean points from a noisy input. Extensive experiments show that our approach outperforms the state-of-the-art deep learning techniques in terms of both visual quality and quantitative error metrics. The source code and dataset can be found at https://github.com/dongbo-BUAA-VR/Pointfilter. Dongbo Zhang 0004, Xuequan Lu, Hong Qin 0001, Ying He 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2020 | Deep Detection for Face Manipulation
Disheng Feng, Xuequan Lu, Xufeng Lin |
ICONIP (5) | 2 |
| 2020 | Deep Patch-Based Human Segmentation
Dongbo Zhang 0004, Zheng Fang 0008, Xuequan Lu, Hong Qin 0001, Antonio Robles-Kelly, Chao Zhang 0030, Ying He 0001 |
ICONIP (1) | 3 |
| 2020 | HDF: Hybrid Deep Features for Scene Image RepresentationabstractNowadays it is prevalent to take features extracted from pre-trained deep learning models as image representations which have achieved promising classification performance. Existing methods usually consider either object-based features or scene-based features only. However, both types of features are important for complex images like scene images, as they can complement each other. In this paper, we propose a novel type of features - hybrid deep features, for scene images. Specifically, we exploit both object-based and scene-based features at two levels: part image level (i.e., parts of an image) and whole image level (i.e., a whole image), which produces a total number of four types of deep features. Regarding the part image level, we also propose two new slicing techniques to extract part based features. Finally, we aggregate these four types of deep features via the concatenation operator. We demonstrate the effectiveness of our hybrid deep features on three commonly used scene datasets (MIT-67, Scene-15, and Event-8), in terms of the scene image classification task. Extensive comparisons show that our introduced features can produce state-of-the-art classification accuracies which are more consistent and stable than the results of existing features across all datasets. Chiranjibi Sitaula, Yong Xiang 0001, Anish Basnet, Sunil Aryal, Xuequan Lu |
IJCNN | 5 |
| 2020 | Deep feature-preserving normal estimation for point cloud filtering
Dening Lu, Xuequan Lu, Yangxing Sun, Jun Wang 0039 |
Comput. Aided Des. | 2 |
| 2020 | HLO: Half-kernel Laplacian operator for surface smoothing
Wei Pan 0010, Xuequan Lu, Yuanhao Gong, Wenming Tang, Ying He 0001, Guoping Qiu |
Comput. Aided Des. | 2 |
| 2020 | G2MF-WA: Geometric multi-model fitting with weakly annotated dataabstractIn this paper we address the problem of geometric multi-model fitting using a few weakly annotated data points, which has been little studied so far. In weak annotating (WA), most manual annotations are supposed to be correct yet inevitably mixed with incorrect ones. Such WA data can naturally arise through interaction in various tasks. For example, in the case of homography estimation, one can easily annotate points on the same plane or object with a single label by observing the image. Motivated by this, we propose a novel method to make full use of WA data to boost multi-model fitting performance. Specifically, a graph for model proposal sampling is first constructed using the WA data, given the prior that WA data annotated with the same weak label has a high probability of belonging to the same model. By incorporating this prior knowledge into the calculation of edge probabilities, vertices (i.e., data points) lying on or near the latent model are likely to be associated and further form a subset or cluster for effective proposal generation. Having generated proposals, o-expansion is used for labeling, and our method in return updates the proposals. This procedure works in an iterative way. Extensive experiments validate our method and show that it produces noticeably better results than state-of-the-art techniques in most cases. Chao Zhang 0030, Xuequan Lu, Katsuya Hotta, Xi Yang 0017 |
Comput. Vis. Media | 2 |
| 2020 | Multi-label zero-shot learning with graph convolutional networks
Guangjin Ou, Guoxian Yu, Carlotta Domeniconi, Xuequan Lu, Xiangliang Zhang 0001 |
Neural Networks | 4 |
| 2019 | Tag-Based Semantic Features for Scene Image Classification
Chiranjibi Sitaula, Yong Xiang 0001, Anish Basnet, Sunil Aryal, Xuequan Lu |
ICONIP (3) | 5 |
| 2019 | Unsupervised Deep Features for Privacy Image Classification
Chiranjibi Sitaula, Yong Xiang 0001, Sunil Aryal, Xuequan Lu |
PSIVT | 4 |
| 2019 | 3D articulated skeleton extraction using a single consumer-grade depth camera
Xuequan Lu, Zhigang Deng 0001, Jun Luo 0001, Wenzhi Chen, Sai-Kit Yeung, Ying He 0001 |
Comput. Vis. Image Underst. | 1 |
| 2018 | Unsupervised Articulated Skeleton Extraction From Point Set Sequences Captured by a Single Depth CameraabstractHow to robustly and accurately extract articulated skeletons from point set sequences captured by a single consumer-grade depth camera still remains to be an unresolved challenge to date. To address this issue, we propose a novel, unsupervised approach consisting of three contributions (steps): (i) a non-rigid point set registration algorithm to first build one-to-one point correspondences among the frames of a sequence; (ii) a skeletal structure extraction algorithm to generate a skeleton with reasonable numbers of joints and bones; (iii) a skeleton joints estimation algorithm to achieve accurate joints. At the end, our method can produce a quality articulated skeleton from a single 3D point sequence corrupted with noise and outliers. The experimental results show that our approach soundly outperforms state of the art techniques, in terms of both visual quality and accuracy. Xuequan Lu, Honghua Chen, Sai-Kit Yeung, Zhigang Deng 0001, Wenzhi Chen |
AAAI | 1 |
| 2018 | GPF: GMM-Inspired Feature-Preserving Point Set FilteringabstractPoint set filtering, which aims at reconstructing noise-free point sets from their corresponding noisy inputs, is a fundamental problem in 3D geometry processing. The main challenge of point set filtering is to preserve geometric features of the underlying geometry while at the same time removing the noise. State-of-the-art point set filtering methods still struggle with this issue: some are not designed to recover sharp features, and others cannot well preserve geometric features, especially fine-scale features. In this paper, we propose a novel approach for robust feature-preserving point set filtering, inspired by the Gaussian Mixture Model (GMM). Taking a noisy point set and its filtered normals as input, our method can robustly reconstruct a high-quality point set which is both noise-free and feature-preserving. Various experiments show that our approach can soundly outperform the selected state-of-the-art methods, in terms of both filtering quality and reconstruction accuracy. Xuequan Lu, Honghua Chen, Sai-Kit Yeung, Wenzhi Chen, Matthias Zwicker |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2017 | Robust mesh denoising via vertex pre-filtering and L1-median normal filtering
Xuequan Lu, Wenzhi Chen, Scott Schaefer |
Comput. Aided Geom. Des. | 1 |
| 2016 | A Robust Scheme for Feature-Preserving Mesh DenoisingabstractIn recent years researchers have made noticeable progresses in mesh denoising, that is, recovering high-quality 3D models from meshes corrupted with noise (raw or synthetic). Nevertheless, these state of the art approaches still fall short for robustly handling various noisy 3D models. The main technical challenge of robust mesh denoising is to remove noise while maximally preserving geometric features. In particular, this issue becomes more difficult for models with considerable amount of noise. In this paper we present a novel scheme for robust feature-preserving mesh denoising. Given a noisy mesh input, our method first estimates an initial mesh, then performs feature detection, identification and connection, and finally, iteratively updates vertex positions based on the constructed feature edges. Through many experiments, we show that our approach can robustly and effectively denoise various input mesh models with synthetic noise or raw scanned noise. The qualitative and quantitative comparisons between our method and the selected state of the art methods also show that our approach can noticeably outperform them in terms of both quality and robustness. Xuequan Lu, Zhigang Deng 0001, Wenzhi Chen |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2014 | AA-FVDM: An accident-avoidance full velocity difference model for animating realistic street-level traffic in rural scenesabstractABSTRACT Most of existing traffic simulation efforts focus on urban regions with a coarse two‐dimensional representation; relatively few studies have been conducted to simulate realistic three‐dimensional traffic flows on a large, complex road web in rural scenes. In this paper, we present a novel agent‐based approach called accident‐avoidance full velocity difference model (abbreviated as AA‐FVDM) to simulate realistic street‐level rural traffics, on top of the existing FVDM. The main distinction between FVDM and AA‐FVDM is that FVDM cannot handle a critical real‐world traffic problem while AA‐FVDM settles this problem and retains the essence of FVDM. We also design a novel scheme to animate the lane‐changing maneuvering process (in particular, the execution course). Through numerous simulations, we demonstrate that besides addressing a previously unaddressed real‐world traffic problem, our AA‐FVDM method efficiently (in real time) simulates large‐scale traffic flows (tens of thousands of vehicles) with realistic, smooth effects. Furthermore, we validate our method using real‐world traffic data, and the validation results show that our method measurably outperforms state‐of‐the‐art traffic simulation methods.Copyright © 2013 John Wiley & Sons, Ltd. Xuequan Lu, Wenzhi Chen, Zonghui Wang, Zhigang Deng 0001, Yangdong Ye |
Comput. Animat. Virtual Worlds | 1 |
| 2014 | A personality model for animating heterogeneous traffic behaviorsabstractABSTRACT How to automatically generate realistic and heterogeneous traffic behaviors has been a much needed yet challenging problem for numerous traffic simulation and urban planning applications. In this paper, we propose a novel approach to model heterogeneous traffic behaviors by adapting a well‐established personality trait model (i.e., Eysenck's PEN (psychoticism, extraversion and neuroticism) model) into widely used traffic simulation approaches. First, we collected a large amount of user feedback while users watch a variety of computer‐generated traffic simulation video clips. Then, we trained regression models to bridge low‐level traffic simulation parameters and high‐level perceived traffic behaviors (i.e., adjectives according to the PEN model and the three PEN traits). We also conducted an additional user study to validate the effectiveness and usefulness of our approach, in particular, high correlation coefficients and the Pearson values between users’ feedback and our model predictions prove the effectiveness of our approach. Furthermore, our approach can also produce interesting emergent traffic patterns including faster‐is‐slower effect and sticking‐in‐a‐pin‐wherever‐there‐is‐room effect. Copyright © 2014 John Wiley & Sons, Ltd. Xuequan Lu, Zonghui Wang, Wenzhi Chen, Zhigang Deng 0001 |
Comput. Animat. Virtual Worlds | 1 |