EDBT 2026 Demo / reviewers in the wild / expert
Xian Wei
dblp:139/0725
· DBLP profile ↗
76ranked-venue papers
9as first author
60since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 46 · 7 first-author · 35 since 2021Graphics, computer vision, multimedia, augmented reality and games · 25 · 5 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 11 since 2021Systems, architecture and hardware · 9 · 9 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Automatic Paper Reviewing with Heterogeneous Graph Reasoning over LLM-Simulated Reviewer-Author DebatesabstractExisting paper review methods often rely on superficial manuscript features or directly on large language models (LLMs), which are prone to hallucinations, biased scoring, and limited reasoning capabilities. Moreover, these methods often fail to capture the complex argumentative reasoning and negotiation dynamics inherent in reviewer-author interactions. To address these limitations, we propose ReViewGraph (Reviewer-Author Debates Graph Reasoner), a novel framework that performs heterogeneous graph reasoning over LLM-simulated multi-round reviewer-author debates. In our approach, reviewer-author exchanges are simulated through LLM-based multi-agent collaboration. Diverse opinion relations (e.g., acceptance, rejection, clarification, and compromise) are then explicitly extracted and encoded as typed edges within a heterogeneous interaction graph. By applying graph neural networks to reason over these structured debate graphs, ReViewGraph captures fine-grained argumentative dynamics and enables more informed review decisions. Extensive experiments on three datasets demonstrate that ReViewGraph outperforms strong baselines with an average relative improvement of 15.73%, underscoring the value of modeling detailed reviewer–author debate structures. Shuaimin Li, Liyang Fan, Yufang Lin, Xian Wei, Shiwen Ni, Hamid Alinejad-Rokny, Min Yang 0007 |
AAAI | 5 |
| 2026 | Resolving Natural Language Ambiguity for Verified Code Generation: A Two-Stage LLM Framework
Yanhong Huang, Jianqi Shi, Haibin Cai, Xian Wei |
ICIC (23) | 6 |
| 2026 | Robust optoelectronic dual-mode memristor enabled by ZnO/MoS2 heterojunction for synaptic bionics and in-memory computing
Youshan Gui, Junyuan Chen, Tenglong Guo, Suo Zhang, Xian Wei, Xiaojuan Lian, Er-Tao Hu, Haifeng Ling |
Sci. China Inf. Sci. | 9 |
| 2026 | Zero-Shot Visual Grounding via Cascade Semantic Prototype LearningabstractCompared to standard visual grounding (VG), zero-shot visual grounding (ZSVG) requires robust localization of novel image–query pairs unseen during training, creating a clear gap from supervised settings. Existing two-stage approaches rely too heavily on external knowledge from pre-trained vision-language models and underexplore in-context semantics, while current one-stage models fail to leverage region-level cues to enhance discriminative representation learning. In this paper, inspired by the information processing procedure of human beings, we propose a novel one-stage network, dubbed Cascade Semantic Prototype Learning (CaSePro), to tackle these limitations. Specifically, CaSePro builds two types of prototypes that capture prototypical multimodal representations for global contexts and local semantics, respectively. With global context prototypes, CaSePro explores context-related cues to enhance the coarsely fused global multimodal features, enabling the model to learn more comprehensive representations. With local semantic prototypes, CaSePro searches for region-level multimodal correspondences to augment the prototypical global patterns, thereby strengthening the cross-modal relationship between the query and visual representations. Moreover, to enhance the generalization ability of the proposed approach on sample-imbalanced benchmark datasets, CaSePro further incorporates an Adaptive Focal Loss that dynamically reweights hard and easy samples by accounting for the model’s grounding uncertainty. Extensive experiments on ten benchmarks demonstrate that CaSePro outperforms all existing one-stage approaches and achieves performance comparable to state-of-the-art two-stage methods. Jinpeng Mi, Jinbo Yang, Xian Wei, Jianwei Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Relaxed Rotational Equivariance via G-Biases in VisionabstractGroup Equivariant Convolution (GConv) can capture rotational equivariance from original data. It assumes uniform and strict rotational equivariance across all features as the transformations under the specific group. However, the presentation or distribution of real-world data rarely conforms to strict rotational equivariance, commonly referred to as Rotational Symmetry-Breaking (RSB) in the system or dataset, making GConv unable to adapt effectively to this phenomenon. Motivated by this, we propose a simple but highly effective method to address this problem, which utilizes a set of learnable biases called G-Biases under the group order to break strict group constraints and then achieve a Relaxed Rotational Equivariant Convolution (RREConv). To validate the efficiency of RREConv, we conduct extensive ablation experiments on the discrete rotational group Cn. Experiments demonstrate that the proposed RREConv-based methods achieve excellent performance compared to existing GConv-based methods in both classification and 2D object detection tasks on the natural image datasets. Licheng Sun, Jian Yang 0034, Shing-Ho J. Lin, Jinpeng Mi, Xian Wei |
AAAI | 10 |
| 2025 | CE-FFT: Communication-Efficient Federated Fine-Tuning for Large Language Models via Quantization and In-Context LearningabstractAlthough Federated Fine-Tuning (FFT) facilitates the fine-tuning of Large Language Models (LLMs) across data owners without compromising their privacy, it suffers from severe communication overheads caused by numerous parameters of LLMs even with Parameter-Efficient Fine-Tuning (PEFT) methods. To address this, we propose a novel communication-efficient FFT framework by reducing per-round transmission costs and simultaneously diminishing the required convergence rounds. Specifically, we freeze the parameter-intensive backbone of LLMs, activate low-rank adapters for fine-tuning, and design a parameter quantization mechanism to reduce per-round communication costs. Additionally, inspired by in-context learning, we align the distributions of local data by identifying and rewriting unimportant samples using the global model, retaining important local information, and guiding local fine-tuning towards a uniform trajectory, thereby accelerating convergence and reducing transmission burdens. Experiments on two instruction datasets show that our method can reduce communication costs by up to 7.76× without compromising performance. Yingbo Zhou 0001, Ming Hu 0003, Xian Wei, Mingsong Chen 0001 |
ICASSP | 5 |
| 2025 | EqGAN: Reformation-based Feature Equalization Fusion for Few-shot Image GenerationabstractDue to the absence or mismatch of semantic information, existing few-shot image generation methods suffer from unsatisfactory generation quality and diversity, which have minimal benefits as data augmentation for downstream classification tasks. Reformatting the contextual and textural information of features at different scales, we propose a novel Feature Equalization Fusion Generative Adversarial Network (EqGAN) for few-shot image generation. Specifically, we first decompose the encoded features into textual and structural components to mitigate the influence of irrelevant and redundant information. Based on feature correlation learning and attention mechanism, we then obtain fused features by refining different contents (i.e., textures and structures) with a more fine-grained semantic alignment. Moreover, an attention-based reconstruction loss and a consistency-based equalization loss are devised to provide better training stability and generation performance. Comprehensive experiments on three public datasets demonstrate that EqGAN not only significantly improves the FID scores (by up to 14.10%) and LPIPS scores (by up to 3.17%) of generated images, but also outperforms the state-of-the-art in terms of accuracy (by up to 3.89%) for downstream classification. Yingbo Zhou 0001, Zhihao Yue, Yutong Ye 0001, Xian Wei, Mingsong Chen 0001 |
ICASSP | 5 |
| 2025 | FiTGAN: Content Fusion with Style Transformation for Few-shot Image GenerationabstractDue to the semantic entanglement in fusion strategies or unstable training in complicated image transformations, existing few-shot image generation methods still suffer from low generation quality and diversity. To tackle the above problems, we propose a novel fusion- and transformation-based framework named content Fusion with style Transformation Generative Adversarial Network (FiTGAN) for few-shot image generation. The basic assumption is that any image consists of a collection of content-related and style-related features. FiTGAN disentangles internal representations with two independent encoders and combines the fused contents and transformed styles to generate new images. Specifically, we design a multi-scale content fusion strategy and a reparameterized style transformation mechanism to learn more fine-grained semantics without changing category-relevant attributes. Furthermore, we formulate a content reconstruction loss and a style divergence loss to provide better training stability and generation performance. Comprehensive experiments on three well-known datasets demonstrate that FiTGAN can not only produce more realistic and diverse images for few-shot image generation but also achieve better classification accuracy for downstream visual applications with limited data. Yingbo Zhou 0001, Yutong Ye 0001, Zhihao Yue, Xian Wei, Mingsong Chen 0001 |
ICASSP | 5 |
| 2025 | R2Det: Exploring Relaxed Rotation Equivariance in 2D Object DetectionabstractGroup Equivariant Convolution (GConv) empowers models to explore underlying symmetry in data, improving performance. However, real-world scenarios often deviate from ideal symmetric systems caused by physical permutation, characterized by non-trivial actions of a symmetry group, resulting in asymmetries that affect the outputs, a phenomenon known as Symmetry Breaking. Traditional GConv-based methods are constrained by rigid operational rules within group space, assuming data remains strictly symmetry after limited group transformations. This limitation makes it difficult to adapt to Symmetry-Breaking and non-rigid transformations. Motivated by this, we mainly focus on a common scenario: Rotational Symmetry-Breaking. By relaxing strict group transformations within Strict Rotation-Equivariant group $\mathbf{C}_n$, we redefine a Relaxed Rotation-Equivariant group $\mathbf{R}_n$ and introduce a novel Relaxed Rotation-Equivariant GConv (R2GConv) with only a minimal increase of $4n$ parameters compared to GConv. Based on R2GConv, we propose a Relaxed Rotation-Equivariant Network (R2Net) as the backbone and develop a Relaxed Rotation-Equivariant Object Detector (R2Det) for 2D object detection. Experimental results demonstrate the effectiveness of the proposed R2GConv in natural image classification, and R2Det achieves excellent performance in 2D object detection with improved generalization capabilities and robustness. The code is available in \texttt{https://github.com/wuer5/r2det}. Jian Yang 0034, Mingsong Chen 0001, Xian Wei |
ICLR | 8 |
| 2025 | Supplementary Material for STTODE: Spatio-Temporal Transformer Ordinary Differential Equation Networks for Pedestrian Trajectory ForecastingabstractDataset 1: ETH-UCY [2] [3] dataset: ETH-UCY consists of the following sub-datasets: ETH, HOTEL, UNIV, ZARA1, ZARA2. In accordance with previous research experimental setups, we split each trajectory sample into 8-second segments, using a time interval of 0.4 seconds. We use the first 3.2 seconds (8 steps) of trajectory data to predict the next 4.8 seconds (12 steps). For training, we use a leave-one-out strategy with 4 subsets and the remaining subset for testing. YingJie Liu, Jian Yang 0034, Mingsong Chen 0001, Xian Wei |
ICME | 6 |
| 2025 | 3D Dense Captioning via Prototypical Momentum Distillationabstract3D dense captioning aims to describe the crucial regions in 3D visual scenes in the form of natural language. Recent prevailing approaches achieve promising results by leveraging complicated structures incorporated with large-scale models, which necessitate abundant parameters and pose challenges regarding its practical applications. Besides, with limited training data, 3D dense captioners are often susceptible to overfitting, directly degrading caption generation performance. Drawing inspiration from the recent advancements in knowledge distillation, we propose a novel approach termed Prototypical Momentum Distillation (PMD) to prompt the model to generate more detailed captions. PMD incorporates Momentum Distillation (MD) with an Uncertainty-aware Prototype-anchored Clustering (UPC) strategy to transfer knowledge by considering the uncertainty of the teacher knowledge. Specifically, we employ the original captioner as the student model and maintain an Exponential Moving Average (EMA) copy of the captioner as the teacher model to impart knowledge as the auxiliary supervision of the student. To abate the misleading caused by uncertain knowledge, we present an Uncertainty-aware Prototype-anchored Clustering (UPC) strategy to cluster the distilled knowledge according to its confidence. We then transfer the rearranged knowledge from the teacher to guide the training route of the student. We conduct extensive experiments and ablation studies on two widely used benchmark datasets, ScanRefer and Nr3D. Experimental results demonstrate that PMD outperforms all state-of-the-art approaches on the benchmarks with MLE training, highlighting its effectiveness. Jinpeng Mi, Shaofei Jin, Xian Wei, Jianwei Zhang 0001 |
ICRA | 5 |
| 2025 | Dual-BEV Nav: Dual-Layer BEV-Based Heuristic Path Planning for Robotic Navigation in Unstructured Outdoor EnvironmentsabstractPath planning with strong environmental adaptability plays a crucial role in robotic navigation in unstructured outdoor environments, especially in the case of low-quality location and map information. The path planning ability of a robot depends on the identification of the traversability of global and local ground areas. In real-world scenarios, the complexity of outdoor open environments makes it difficult for robots to identify the traversability of ground areas that lack a clearly defined structure. Moreover, most existing methods have rarely analyzed the integration of local and global traversability identifications in unstructured outdoor scenarios. To address this problem, we propose a novel method, Dual-BEV Nav, first introducing Bird's Eye View (BEV) representations into local planning to generate high-quality traversable paths. Then, these paths are projected into the global traversability probability map generated by the global BEV planning model to obtain the optimal path. By integrating the traversability from both local and global BEV, we establish a dual-layer BEV heuristic planning paradigm, enabling long-distance navigation in unstructured outdoor environments. We test our approach through both public dataset evaluations and real-world robot deployments, yielding promising results. Compared to baselines, the Dual-BEV Nav improved temporal distance prediction accuracy by up to 18.26%. In the real-world deployment, under conditions significantly different from the training set and with notable occlusions in the global BEV, the Dual-BEV Nav successfully achieved a 65-meter-long outdoor navigation. Further analysis demonstrates that the local BEV representation significantly enhances the rationality of the planning, while the global BEV probability map ensures the robustness of the overall planning. Jian Yang 0034, Shibo Huang, Ke Li 0005, Xian Wei, Xiong You |
ICRA | 8 |
| 2025 | PDDFormer: Pairwise Distance Distribution Graph Transformer for Crystal Material Property PredictionabstractCrystal structures can be simplified as a periodic point set that repeats across three-dimensional space along an underlying lattice. Traditionally, crystal representation methods rely on descriptors such as lattice parameters, symmetry, and space groups to characterize the structure. However, in reality, atoms in materials always vibrate above absolute zero, causing their positions to fluctuate continuously. This dynamic behavior disrupts the fundamental periodicity of the lattice, making crystal graphs based on static lattice parameters and conventional descriptors discontinuous under slight perturbations. Chemists proposed the pairwise distance distribution (PDD) method to address this. However, the completeness of PDD requires defining a large number of neighboring atoms, leading to high computational costs. Additionally, PDD does not account for atomic information, making it challenging to apply it directly to crystal material property prediction tasks. To tackle these challenges, we introduce the atom-weighted Pairwise Distance Distribution (WPDD) and Unit cell Pairwise Distance Distribution (UPDD) for the first time, applying them to the construction of multi-edge crystal graphs. We demonstrate the continuity and general completeness of crystal graphs under slight atomic position perturbations. Moreover, by modeling PDD as global information and integrating it into matrix-based message passing, we significantly reduce computational costs. Comprehensive evaluation results show that WPDDFormer achieves state-of-the-art predictive accuracy across tasks on benchmark datasets such as the Materials Project and JARVIS-DFT. Xiangxiang Shen, Lingfeng Wen, Licheng Sun, Jian Yang 0034, Shing-Ho J. Lin, Xiao He 0004, Mingsong Chen 0001, Xian Wei |
IJCAI | 10 |
| 2025 | Generative Multi-Agent Collaboration in Embodied AI: A Systematic ReviewabstractEmbodied multi-agent systems (EMAS) have attracted growing attention for their potential to address complex, real-world challenges in areas such as logistics and robotics. Recent advances in foundation models pave the way for generative agents capable of richer communication and adaptive problem-solving. This survey provides a systematic examination of how EMAS can benefit from these generative capabilities. We propose a taxonomy that categorizes EMAS by system architectures and embodiment modalities, emphasizing how collaboration spans both physical and virtual contexts. Central building blocks, perception, planning, communication, and feedback, are then analyzed to illustrate how generative techniques bolster system robustness and flexibility. Through concrete examples, we demonstrate the transformative effects of integrating foundation models into embodied, multi-agent frameworks. Finally, we discuss challenges and future directions, underlining the significant promise of EMAS to reshape the landscape of AI-driven collaboration. Xian Wei, Guang Chen 0001, Hao Shen 0002, Bo Jin 0003 |
IJCAI | 2 |
| 2025 | KiteRunner: Language-Driven Cooperative Local-Global Navigation Policy with UAV Mapping in Outdoor EnvironmentsabstractAutonomous navigation in open-world outdoor environments faces challenges in integrating dynamic conditions, long-distance spatial reasoning, and semantic understanding. Traditional methods struggle to balance local planning, global planning, and semantic task execution, while existing large language models (LLMs) enhance semantic comprehension but lack spatial reasoning capabilities. Although diffusion models excel in local optimization, they fall short in large-scale long-distance navigation. To address these gaps, this paper proposes KiteRunner, a language-driven cooperative local-global navigation strategy that combines UAV orthophoto-based global planning with diffusion model-driven local path generation for long-distance navigation in open-world scenarios. Our method innovatively leverages real-time UAV orthophotography to construct a global probability map, providing traversability guidance for the local planner, while integrating large models like CLIP and GPT to interpret natural language instructions. Experiments demonstrate that KiteRunner achieves 5.6% and 12.8% improvements in path efficiency over state-of-the-art methods in structured and unstructured environments, respectively, with significant reductions in human interventions and execution time. Shibo Huang, Chenfan Shi, Jian Yang 0034, Jinpeng Mi, Ke Li 0005, Miao Ding, Peidong Liang, Xiong You, Xian Wei |
IROS | 11 |
| 2025 | NeuroLoc: Encoding Navigation Cells for 6-DOF Camera LocalizationabstractRecently, camera localization has been widely adopted in autonomous robotic navigation due to its efficiency and convenience. However, autonomous navigation in unknown environments often suffers from scene ambiguity, environmental disturbances, and dynamic object transformation in camera localization. To address this problem, inspired by the brain cognitive navigation mechanism (such as grid cells, place cells, and head direction cells), we propose a novel neurobiological camera location method, namely NeuroLoc. Firstly, we designed a Hebbian learning module driven by place cells to save and replay historical information, aiming to restore the details of historical representations and solve the issue of scene fuzziness. Secondly, we utilized the head direction cell-inspired internal direction learning as multi-head attention embedding to help restore the true orientation in similar scenes. Finally, we added a 3D grid center prediction in the pose regression module to reduce the final wrong prediction. We evaluate the proposed NeuroLoc on commonly used benchmark indoor and outdoor datasets. The experimental results show that our NeuroLoc can enhance the robustness in complex environments and improve the performance of pose regression by using only a single image. Jian Yang 0034, Fenli Jia, Muyu Wang, Jinpeng Mi, Jilin Hu, Peidong Liang, Ke Li 0005, Xiong You, Xian Wei |
IROS | 12 |
| 2025 | MuseCNN: Embedding-Guided Polyphonic Music Accompaniment GenerationabstractAlthough various methods are proposed to generate music accompaniment tracks according to the main music track, they suffer from the problems of modeling music dependencies and representing music data, therefore, generating high-quality music accompaniment remains a challenging task.To address this issue, we propose a multi-track sequential convolutional neural network (MuseCNN) to generate accompaniment tracks corresponding to the main music track.Inspired by the similarity between pianoroll representation and pictures, we transform the pianoroll data into a two-channel music representation matrix and feed it into the convolutional neural network (CNN).Using a hierarchical loss function, we integrate all music tracks to maintain the coherence and harmony of the music.Based on the three-level loss design and multiple CNNs, the problem of modeling music dependencies can be solved.The experimental results indicate that our multi-CNN model can effectively learn complex music dependencies and generate harmonious longsequence polyphonic music. Yutong Ye 0001, Yingbo Zhou 0001, Qi Wen 0002, Xiang Lian 0001, Xian Wei, Mingsong Chen 0001 |
SEKE | 6 |
| 2025 | When foresight pruning meets zeroth-order optimization: Efficient federated learning for low-memory devices
Yingbo Zhou 0001, Xian Wei, Mingsong Chen 0001 |
J. Syst. Archit. | 4 |
| 2025 | Hyperbolic prototype rectification for few-shot 3D point cloud classification
Yuanzhi Feng, Shing-Ho J. Lin, Mu-Yu Wang, Jianzhang Zheng, Ziyao He, Zi-Yi Pang, Jian Yang 0034, Mingsong Chen 0001, Xian Wei |
Pattern Recognit. | 10 |
| 2025 | Efficient Unsupervised Deep Learning for Simultaneous Seismic Noise Attenuation and InterpolationabstractWe have explored an unsupervised deep learning (DL)-based approach for the efficient and effective reconstruction of noisy and incomplete (N-I) seismic data. This method does not require clean and complete (C-C) seismic data as labeled data. In each iteration, the seismic data input to the network is first subjected to patching techniques, dividing the 2-D or 3-D data into many 1-D signals. Then, to boost the efficiency, a reconstruction error-based patch selection (REBPS) is employed to choose patches that contain more complex structures, which are then fed into the network. The network adopts an encoder and corresponding decoder architecture to compress and reconstruct data features, attenuating noise within the seismic data and performing an initial reconstruction of the missing parts. To improve the reconstruction accuracy, we employ the projection onto convex sets (POCS) algorithm, ultimately obtaining reconstructed data from one iteration. In this process, the output results of each POCS iteration serve as the input for the next round. Through experimental verification using both synthetic and field seismic data, the results show that our proposed method surpasses other comparative methods in the quality of seismic data reconstruction. Anyu Li, Wei Chen 0031, Xian Wei, Yangkang Chen |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | FuseFormer: A Manifold Metric Fusing Attention for Pedestrian Trajectory PredictionabstractAccurate pedestrian trajectory prediction is critical for ensuring the safety of autonomous vehicles and advancing higher levels of driving automation. However, the complex interpersonal interactions and highly dynamic trajectory patterns in real-world scenarios pose significant challenges to achieving precise predictions. Recently, Transformers have shown remarkable success in pedestrian trajectory prediction, primarily due to their effective modeling of temporal and spatial dependencies via Multi-Head Self-Attention (MHA) mechanisms. Despite these advancements, existing self-attention methods often rely on Euclidean distance-based metrics and dot-product operations, which are inadequate for capturing interaction-induced trajectory curvatures. To address this limitation, we propose a novel hybrid Transformer architecture, FuseFormer, that incorporates Geodesic Self-Attention (GSA) mechanisms. GSA utilizes geodesic distances to characterize interaction features effectively, complementing MHA, which excels in capturing local features and maintaining temporal correlations. FuseFormer employs a gating network to adaptively combine GSA and MHA embeddings, leveraging their complementary strengths. Additionally, FuseFormer integrates a Transformer-based Neural Ordinary Differential Equation (ODE) decoder to model trajectory temporal dynamics. This design enables the generation of future trajectories that align closely with motion trends while adapting the network depth to input sequence lengths. Experimental results demonstrate that FuseFormer achieves state-of-the-art performance across widely used pedestrian trajectory prediction datasets, including ETH/UCY, SDD, and NBA. These results underscore the model’s effectiveness and generalization capability in capturing complex interaction patterns and handling diverse scenarios. Kohsin Ko, Jian Yang 0034, Ke Li 0005, Xiong You, Jinpeng Mi, Mingsong Chen 0001, Xian Wei |
IEEE Trans. Intell. Transp. Syst. | 10 |
| 2025 | Integrating Convolution and Sparse Coding for Learning Low-Dimensional Discriminative Image RepresentationsabstractThis work investigates the problem of efficiently learning discriminative low-dimensional (LD) representations of multiclass image objects. We propose a generic end-to-end approach that jointly optimizes sparse dictionary and convolutions for learning LOW-dimensional discriminative image representations, named SparConvLow, taking advantage of convolutional neural networks (CNNs), dictionary learning, and orthogonal projections. The whole learning process can be summarized as follows. First, a CNN module is employed to extract high-dimensional (HD) preliminary convolutional features. Second, to avoid the high computational cost of direct sparse coding on HD CNN features, we learn sparse representation (SR) over a task-driven dictionary in the space with the feature being orthogonally projected. We then exploit the discriminative projection on SR. The whole learning process is consistently treated as an end-to-end joint optimization problem of trace quotient maximization. The cost function is well-defined on the product of the CNN parameters space, the Stiefel manifold, the Oblique manifold, and the Grassmann manifold. By using the explicit gradient delivery, the cost function is optimized via a geometrical stochastic gradient descent (SGD) algorithm along with the chain rule and the backpropagation. The experimental results show that the proposed method can achieve a highly competitive performance with the state-of-the-art (SOTA) image classification, object categorization, and face recognition methods, under both supervised and semi-supervised settings. The code is available at https://github.com/MVPR-Group/SparConvLow. Xian Wei, Shui Yu 0001, Mingsong Chen 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Hyperbolic Graph Diffusion ModelabstractDiffusion generative models (DMs) have achieved promising results in image and graph generation. However, real-world graphs, such as social networks, molecular graphs, and traffic graphs, generally share non-Euclidean topologies and hidden hierarchies. For example, the degree distributions of graphs are mostly power-law distributions. The current latent diffusion model embeds the hierarchical data in a Euclidean space, which leads to distortions and interferes with modeling the distribution. Instead, hyperbolic space has been found to be more suitable for capturing complex hierarchical structures due to its exponential growth property. In order to simultaneously utilize the data generation capabilities of diffusion models and the ability of hyperbolic embeddings to extract latent hierarchical distributions, we propose a novel graph generation method called, Hyperbolic Graph Diffusion Model (HGDM), which consists of an auto-encoder to encode nodes into successive hyperbolic embeddings, and a DM that operates in the hyperbolic latent space. HGDM captures the crucial graph structure distributions by constructing a hyperbolic potential node space that incorporates edge information. Extensive experiments show that HGDM achieves better performance in generic graph and molecule generation benchmarks, with a 48% improvement in the quality of graph generation with highly hierarchical structures. Lingfeng Wen, Mingjie Ouyang, Xiangxiang Shen, Jian Yang 0034, Daxin Zhu, Mingsong Chen 0001, Xian Wei |
AAAI | 8 |
| 2024 | Exact Fusion via Feature Distribution Matching for Few-Shot Image GenerationabstractFew-shot image generation, as an important yet challenging visual task, still suffers from the trade-off between generation quality and diversity. According to the principle of feature-matching learning, existing fusion-based methods usually fuse different features by using similarity measurements or attention mechanisms, which may match features inaccurately and lead to artifacts in the texture and structure of generated images. In this paper, we propose an exact Fusion via Feature Distribution matching Generative Adversarial Network (F2DGAN) for few-shot image generation. The rationale behind this is that feature distribution matching is much more reliable than feature matching to explore the statistical characters in image feature space for limited real-world data. To model feature distributions from only a few examples for feature fusion, we design a novel variational feature distribution matching fusion module to perform exact fusion by empirical cumulative distribution functions. Specifically, we employ a variational autoencoder to transform deep image features into distributions and fuse different features exactly by applying histogram matching. Additionally, we formulate two effective losses to guide the matching process for better fitting our fusion strategy. Extensive experiments compared with state-of-the-art methods on three public datasets demonstrate the superiority of F2DGAN for few-shot image generation in terms of generation quality and diversity, and the effectiveness of data augmentation in downstream classification tasks11Code is available at: https:/github.com/ZYBOBO/F2DGAN. Yingbo Zhou 0001, Yutong Ye 0001, Xian Wei, Mingsong Chen 0001 |
CVPR | 4 |
| 2024 | Social Lode: Human Trajectory Prediction with Latent OdesabstractHuman trajectory prediction is crucial in human-computer interaction and even in the safety of autonomous driving. In this work, A new method, called Social Latent Ordinary Differential Equation (Social LODE), is introduced for predicting human trajectories. The backbone of Social LODE consists of a conditional Variational Autoencoder (VAE) architecture based on Recurrent Neural Network (RNN). The hidden state updated by RNN is often discrete, but the human trajectory is continuous and uncertain. Thus, we use Latent ODEs as the decoder of VAE to overcome the limitation of RNN. Finally, we demonstrate that Social LODE achieves state-of-the-art compared to other methods, such as those involving the ETH/UCY and SDD datasets. Kexin Ke, Jian Yang 0034, Mingsong Chen 0001, Xian Wei |
ICASSP | 5 |
| 2024 | ProEqBEV: Product Group Equivariant BEV Network for 3D Object Detection in Road Scenes of Autonomous DrivingabstractWith the rapid development of autonomous driving systems, 3D object detection based on Bird’s Eye View (BEV) in road scenes has witnessed great progress over the past few years. As a road scene exhibits a part-whole hierarchy between the within objects and the scene itself, simple parts (e.g., roads, lane lines, vehicles and pedestrians) can be assembled into progressively more complex shapes to form a BEV representation of the whole road scene. Therefore, a BEV often has multiple levels of freedom on motion, i.e., the rotation and the moving shift of the whole BEV, and the random movements of objects (e.g., pedestrians and vehicles) inside the BEV. However, most of the current single-sensor or multi-sensor fusion-based BEV object detection methods have not yet taken into account capturing such multi-level motion in a BEV. To address this problem, we propose a product group equivariant object detection network framework that is equivariant with respect to multiple levels of symmetry groups based on multi-sensor fusion. The proposed framework extracts local equivariant features of objects in point clouds, while global equivariant features are extracted in both point clouds and images. Furthermore, the network learns diverse rotation-equivariant features and mitigates a significant amount of detection errors caused by rotations of BEV and objects inside a BEV, thereby further enhancing the performance of object detection. The experiment results show that the network architecture significantly improves object detection on mAP and NDS, respectively. In addition, in order to demonstrate the effectiveness of the proposed local-multi-global equivariant components, we conduct sufficient ablation experiments. The results show that the individual components are indispensable for the object detection performance improvement of the overall network architecture. Jian Yang 0034, Ke Li 0005, Jianzhang Zheng, Xihao Wang, Mingsong Chen 0001, Xiong You, Xian Wei |
ICRA | 10 |
| 2024 | Continual Learning of 3D Point Cloud with Hyperbolic Manifold ReplayabstractAs an irregular and sparse form of data, 3D point cloud data is widely used in the fields of computer vision and machine learning. Its effective classification and recognition are of great significance. In many cases, task data does not arrive all at once but is acquired in batches over time. However, in deep learning, continual learning faces the problem of catastrophic forgetting, where the model loses memory of old tasks while learning new ones. How to retain old knowledge while learning new knowledge is a significant challenge. To address these challenges, existing replay-based strategies alleviate this issue by storing small portions of old samples. However, this often leads to an imbalance between new and old data, affecting performance. Additionally, point cloud data typically has complex non-Euclidean structures, with potential hierarchical relationships within and between point cloud objects. Current deep learning models based on Euclidean space struggle to capture the hierarchical prior features of point clouds. Therefore, in this paper, we introduce a continual learning method based on a replay mechanism and explore how to incorporate hyperbolic space into continual learning tasks to enhance feature representation capabilities. We propose a manifold replay strategy in hyperbolic space, termed HyMR. Specifically, this paper presents a knowledge distillation strategy that combines global and local information, and utilizes manifold spherical projection to select representative old data for replay. Experiments demonstrate that this method achieves good results in continual learning tasks on point cloud datasets such as ShapeNet and ModelNet. Zihao Xu 0002, Mingsong Chen 0001, Xian Wei |
ICTAI | 6 |
| 2024 | Continuous Geodesic Self-Attention Models with Gated Fusion for Trajectory PredictionabstractDriven by the rapid development of intelligent driving vehicles, predicting the trajectories of pedestrians on the road is crucial for decision-making during driving and even road safety. In this paper, we propose a novel method for trajectory prediction, namely, Continuous Geodesic Self-Attention Models with Gated Fusion (CGSAG). We use geodesic attention to measure the similarity between trajectory points, and utilize a gating mechanism to fuse the geodesic features extracted by multi-layer graph convolution. We then use Neural Ordinary Differential Equations (Neural ODE) to model the continuous-time dynamics of the trajectory. We show that CGSAG improves state-of-the-art performances on several human trajectory prediction datasets, including ETH/UCY, SDD, and Ind. At the same time, we conduct ablation studies to prove the effectiveness and efficiency of our proposed method. Kexin Ke, Huining Chen, Xian Wei, Jian Yang 0034 |
IJCNN | 5 |
| 2024 | Is Aggregation the Only Choice? Federated Learning via Layer-wise Model RecombinationabstractAlthough Federated Learning (FL) enables global model training across clients without compromising their raw data, due to the un- evenly distributed data among clients, existing Federated Averaging (FedAvg)-based methods suffer from the problem of low inference performance. Specifically, different data distributions among clients lead to various optimization directions of local models. Aggregat- ing local models usually results in a low-generalized global model, which performs worse on most of the clients. To address the above issue, inspired by the observation from a geometric perspective that a well-generalized solution is located in a flat area rather than a sharp area, we propose a novel and heuristic FL paradigm named FedMR (Federated Model Recombination). The goal of FedMR is to guide the recombined models to be trained towards a flat area. Unlike conventional FedAvg-based methods, in FedMR, the cloud server recombines collected local models by shuffling each layer of them to generate multiple recombined models for local training on clients rather than an aggregated global model. Since the area of the flat area is larger than the sharp area, when local models are located in different areas, recombined models have a higher probability of locating in a flat area. When all recombined models are located in the same flat area, they are optimized towards the same direction. We theoretically analyze the convergence of model recombination. Experimental results show that, compared with state-of-the-art FL methods, FedMR can significantly improve the inference accuracy without exposing the privacy of each client. Ming Hu 0003, Zhihao Yue, Xiaofei Xie, Cheng Chen 0015, Yihao Huang 0001, Xian Wei, Xiang Lian 0001, Yang Liu 0003, Mingsong Chen 0001 |
KDD | 6 |
| 2024 | Approximate Matrix Multiplication over Sliding WindowsabstractLarge-scale streaming matrix multiplication is very common in various applications, sparking significant interest in develop efficient algorithms for approximate matrix multiplication (AMM) over streams. In addition, many practical scenarios require to process time-sensitive data and aim to compute matrix multiplication for most recent columns of the data matrices rather than the entire matrices, which motivated us to study efficient AMM algorithms over sliding windows. In this paper, we present two novel deterministic algorithms for this problem and provide corresponding error guarantees. We further reduce the space and time costs of our methods for sparse matrices by performing an approximate singular value decomposition which can utilize the sparsity of matrices. Extensive experimental results on both synthetic and real-world datasets validate our theoretical analysis and highlight the efficiency of our methods. Ziqi Yao, Lianzhi Li, Mingsong Chen 0001, Xian Wei, Cheng Chen 0015 |
KDD | 4 |
| 2024 | WaveAttack: Asymmetric Frequency Obfuscation-based Backdoor Attacks Against Deep Neural NetworksabstractDue to the increasing popularity of Artificial Intelligence (AI), more and more backdoor attacks are designed to mislead Deep Neural Network (DNN) predictions by manipulating training samples or processes. Although backdoor attacks have been investigated in various scenarios, they still suffer from the problems of both low fidelity of poisoned samples and non-negligible transfer in latent space, which make them easily identified by existing backdoor detection algorithms. To overcome this weakness, this paper proposes a novel frequency-based backdoor attack method named WaveAttack, which obtains high-frequency image features through Discrete Wavelet Transform (DWT) to generate highly stealthy backdoor triggers. By introducing an asymmetric frequency obfuscation method, our approach adds an adaptive residual to the training and inference stages to improve the impact of triggers, thus further enhancing the effectiveness of WaveAttack. Comprehensive experimental results show that, WaveAttack can not only achieve higher effectiveness than state-of-the-art backdoor attack methods, but also outperform them in the fidelity of images (i.e., by up to 28.27\% improvement in PSNR, 1.61\% improvement in SSIM, and 70.59\% reduction in IS). Our code is available at https://github.com/BililiCode/WaveAttack. Jun Xia 0003, Zhihao Yue, Yingbo Zhou 0001, Zhiwei Ling, Yiyu Shi 0001, Xian Wei, Mingsong Chen 0001 |
NeurIPS | 6 |
| 2024 | SampDetox: Black-box Backdoor Defense via Perturbation-based Sample DetoxificationabstractThe advancement of Machine Learning has enabled the widespread deployment of Machine Learning as a Service (MLaaS) applications. However, the untrustworthy nature of third-party ML services poses backdoor threats. Existing defenses in MLaaS are limited by their reliance on training samples or white-box model analysis, highlighting the need for a black-box backdoor purification method. In our paper, we attempt to use diffusion models for purification by introducing noise in a forward diffusion process to destroy backdoors and recover clean samples through a reverse generative process. However, since a higher noise also destroys the semantics of the original samples, it still results in a low restoration performance. To investigate the effectiveness of noise in eliminating different types of backdoors, we conducted a preliminary study, which demonstrates that backdoors with low visibility can be easily destroyed by lightweight noise and those with high visibility need to be destroyed by high noise but can be easily detected. Based on the study, we propose SampDetox, which strategically combines lightweight and high noise. SampDetox applies weak noise to eliminate low-visibility backdoors and compares the structural similarity between the recovered and original samples to localize high-visibility backdoors. Intensive noise is then applied to these localized areas, destroying the high-visibility backdoors while preserving global semantic information. As a result, detoxified samples can be used for inference, even by poisoned models. Comprehensive experiments demonstrate the effectiveness of SampDetox in defending against various state-of-the-art backdoor attacks. Yanxin Yang, Chentao Jia, Dengke Yan, Ming Hu 0003, Tianlin Li, Xiaofei Xie, Xian Wei, Mingsong Chen 0001 |
NeurIPS | 7 |
| 2024 | Zero-shot visual grounding via coarse-to-fine representation learning
Jinpeng Mi, Shaofei Jin, Zhiqian Chen, Xian Wei, Jianwei Zhang 0001 |
Neurocomputing | 5 |
| 2024 | rpcPRF: Generalizable MPI Neural Radiance Field for Satellite Camera With Single and Sparse ViewsabstractMost advances in neural radiance fields (NeRF) assume sufficient input views from pinhole cameras. This paper proposes rpcPRF, a Multiplane Images (MPI) based Planar neural Radiance Field for Rational Polynomial Camera (RPC), exhibits robust performance even with single or sparse inputs. Unlike coordinate-based NeRFs that require sufficient views of one scene, our model can be applied to new scenes and has shown positive results for single or sparse test images. This allows for generalization across various scenes. To achieve generalization across scenes, we use reprojection supervision to ensure that the predicted MPI (multiplane image) accurately captures the geometry between the 3D coordinates and the images. Additionally, we have obviated the requisite for dense depth supervision in multiview-stereo-based methods by introducing rendering techniques of radiance fields. rpcPRF combines the superiority of implicit representations and the advantages of the RPC model, to capture the continuous altitude space while learning the 3D structure. On the DFC2019 dataset with sparse input of the same scene, rpcPRF achieves the best results. On the TLC dataset and the SatMVS3D dataset with changing scenes in every batch, rpcPRF outperforms state-of-the-art NeRF-based methods by a significant margin in terms of image fidelity, reconstruction accuracy, and efficiency, for both single-view and multiview task. Tongtong Zhang, Xian Wei |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | SatensoRF: Fast Satellite Tensorial Radiance Field for Multidate Satellite Imagery of Large SizeabstractExisting NeRF models for satellite imagery have limitations in processing large images and require solar input, leading to slow speeds. As a response, we introduce SatensoRF, which speeds up the entire process significantly while using fewer parameters for large satellite imagery. We have noticed that the common assumption of Lambertian surfaces in satellite neural radiance fields is not sufficient for vegetative and aquatic elements. In contrast to the traditional hierarchical MLP-based scene representation, we have chosen a multiscale tensor decomposition approach for color, volume density, and auxiliary variables to model the light field with specular color. Additionally, to rectify inconsistencies in multi-date imagery, we incorporate total variation denoising to restore the density tensor field, thus mitigating the negative impact of transient objects. To validate our approach, we conducted assessments of SatensoRF using subsets from the spacenet multi-view dataset, which includes both multi-date and single-date multi-view RGB images. Our results demonstrate that SatensoRF surpasses the state-of-the-art Sat-NeRF series regarding novel view synthesis performance. Significantly, SatensoRF requires fewer parameters for training, resulting in faster training and inference speeds and reduced computational demands. Tongtong Zhang, Xian Wei |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | DuEqNet: Dual-Equivariance Network in Outdoor 3D Object Detection for Autonomous DrivingabstractOutdoor 3D object detection has played an essential role in the environment perception of autonomous driving. In complicated traffic situations, precise object recognition provides indispensable information for prediction and planning in the dynamic system, improving self-driving safety and reliability. However, with the vehicle's veering, the constant rotation of the surrounding scenario makes a challenge for the perception systems. Yet most existing methods have not focused on alleviating the detection accuracy impairment brought by the vehicle's rotation, especially in outdoor 3D detection. In this paper, we propose DuEqNet, which first introduces the concept of equivariance into 3D object detection network by leveraging a hierarchical embedded framework. The dual-equivariance of our model can extract the equivariant features at both local and global levels, respectively. For the local feature, we utilize the graph-based strategy to guarantee the equivariance of the feature in point cloud pillars. In terms of the global feature, the group equivariant convolution layers are adopted to aggregate the local feature to achieve the global equivariance. In the experiment part, we evaluate our approach with different baselines in 3D object detection tasks and obtain State-Of-The-Art performance. According to the results, our model presents higher accuracy on orientation and better prediction efficiency. Moreover, our dual-equivariance strategy exhibits the satisfied plug-and-play ability on various popular object detection frameworks to improve their performance. Xihao Wang, JiaMing Lei, Arafat Al-Jawari, Xian Wei |
ICRA | 5 |
| 2023 | SparGE: Sparse Coding-based Patient Similarity Learning via Low-rank Constraints and Graph EmbeddingabstractPatient similarity assessment (PSA) is pivotal to evidence-based and personalized medicine, enabled by analyzing the increasingly available electronic health records (EHRs). However, machine learning approaches for PSA have to deal with inherent data deficiencies of EHRs, namely missing values, noise, and small sample sizes. In this work, an end-to-end discriminative learning framework, called SparGE, is proposed to address these data challenges of EHR for PSA. SparGE measures similarity by jointly sparse coding and graph embedding. First, we use low-rank constrained sparse coding to identify and calculate weight for similar patients, while denoising against missing values. Then, graph embedding on sparse representations is adopted to measure the similarity between patient pairs via preserving local relationships defined by distances. Finally, a global cost function is constructed to optimize related parameters. Experimental results on two private and public real-world healthcare datasets, namely SingHEART and MIMIC-III, show that the proposed SparGE significantly outperforms other machine learning patient similarity methods. Xian Wei, See-Kiong Ng, Tongtong Zhang, Mingsong Chen 0001 |
IJCNN | 2 |
| 2023 | Model-Contrastive Learning for Backdoor EliminationabstractDue to the popularity of Artificial Intelligence (AI) techniques, we are witnessing an increasing number of backdoor injection attacks that are designed to maliciously threaten Deep Neural Networks (DNNs) causing misclassification. Although there exist various defense methods that can effectively erase backdoors from DNNs, they greatly suffer from both high Attack Success Rate (ASR) and a non-negligible loss in Benign Accuracy (BA). Inspired by the observation that a backdoored DNN tends to form a new cluster in its feature spaces for poisoned data, in this paper, we propose a novel two-stage backdoor defense method, named MCLDef, based on Model-Contrastive Learning (MCL). MCLDef can purify the backdoored model by pulling the feature representations of poisoned data towards those of their clean data counterparts. Due to the shrunken cluster of poisoned data, the backdoor formed by end-to-end supervised learning can be effectively eliminated. Comprehensive experimental results show that, with only 5% of clean data, MCLDef significantly outperforms state-of-the-art defense methods by up to 95.79% reduction in ASR, while in most cases, the BA degradation can be controlled within less than 2%. Our code is available at https://github.com/Zhihao151/MCL. Zhihao Yue, Jun Xia 0003, Zhiwei Ling, Ming Hu 0003, Ting Wang 0001, Xian Wei, Mingsong Chen 0001 |
ACM Multimedia | 6 |
| 2023 | Learning Dictionary for Visual AttentionabstractRecently, the attention mechanism has shown outstanding competence in capturing global structure information and long-range relationships within data, thus enhancing the performance of deep vision models on various computer vision tasks. In this work, we propose a novel dictionary learning-based attention (\textit{Dic-Attn}) module, which models this issue as a decomposition and reconstruction problem with the sparsity prior, inspired by sparse coding in the human visual perception system. The proposed \textit{Dic-Attn} module decomposes the input into a dictionary and corresponding sparse representations, allowing for the disentanglement of underlying nonlinear structural information in visual data and the reconstruction of an attention embedding. By applying transformation operations in the spatial and channel domains, the module dynamically selects the dictionary's atoms and sparse representations. Finally, the updated dictionary and sparse representations capture the global contextual information and reconstruct the attention maps. The proposed \textit{Dic-Attn} module is designed with plug-and-play compatibility, allowing for integration into deep attention encoders. Our approach offers an intuitive and elegant means to exploit the discriminative information from data, promoting visual attention construction. Extensive experimental results on various computer vision tasks, e.g., image and point cloud classification, validate that our method achieves promising performance, and shows a strong competitive comparison with state-of-the-art attention methods. Xian Wei |
NeurIPS | 5 |
| 2023 | RECO: Rotation Equivariant COnvolutional Neural Network for Human Trajectory Forecasting
Jijun Cheng, Dongheng Shao, Jian Yang 0034, Mingsong Chen 0001, Xian Wei |
PRCV (3) | 6 |
| 2023 | Brief Industry Paper: Towards Efficient Task Scheduling for AUTOSAR using Parallel PruningabstractAs a standardized software framework and open E/E system architecture, the AUTomotive Open System ARchitecture (AUTOSAR) has been widely applied to autonomous driving systems to enable real-time control. However, due to the increasing design complexity and the lack of efficient algorithms and design automation tools, it is difficult to quickly figure out an optimal task scheduling scheme for an AUTOSAR-based system. To address this problem, we introduce a novel task scheduling method that can parallelly search for an optimal solution with the help of our proposed pruning strategy. Experimental results on a real-world AUTOSAR-based autonomous driving system demonstrate that our approach can achieve much better task scheduling solutions than the ones obtained manually and significantly reduce the overall task scheduling time. Yanxing Yang, Nan Zhang 0019, Dengke Yan, Xian Wei, Junlong Zhou, Mingsong Chen 0001 |
RTSS | 4 |
| 2023 | Brief Industry Paper: RTLight: Digital Twin-Based Real-Time Federated Traffic Signal ControlabstractAlthough Reinforcement Learning (RL)-based methods have been widely researched in Traffic Signal Control (TSC), they still suffer from the problems of poor adaptation to real-world traffic scenarios and slow convergence to optimized solutions. This is because RL-based TSC methods have a high dependency on accurate modeling of the environment. With transportation infrastructure constraints, some vehicle dynamic information in the road network is difficult to obtain in real-time, which strongly limits the capability of RL agents. To address this problem, we propose a novel real-time federated traffic signal control system named RTLight, which can efficiently control traffic lights in real-time for multi-intersection scenarios. Based on the digital twin, the RL agent can obtain sufficient traffic information and interact with the environment in real time. Inspired by federated learning, our system supports knowledge sharing among intersections, which improves the overall convergence rate and control performance. Note that, we have deployed our RTLight system for large-scale application validation in Xishan district, Wuxi, China. Experimental results obtained from various real-world traffic scenarios demonstrate that RTLight can significantly improve the control performance. Yutong Ye 0001, Zhiwei Ling, Yaning Yang, Xian Wei, Cheng Chen 0027, Mingsong Chen 0001 |
RTSS | 4 |
| 2023 | Continual Learning via Manifold Expansion ReplayabstractIn continual learning, the learner learns multiple tasks in sequence, with data being acquired only once for each task. Catastrophic forgetting is a major challenge to continual learning. To reduce forgetting, some existing rehearsal-based methods use episodic memory to replay samples of previous tasks. However, in the process of knowledge integration when learning a new task, this strategy also suffers from catastrophic forgetting due to an imbalance between old and new knowledge. To address this problem, we propose a novel replay strategy called Manifold Expansion Replay (MaER). We argue that expanding the implicit manifold of the knowledge representation in the episodic memory helps to improve the robustness and expressiveness of the model. To this end, we propose a greedy strategy to keep increasing the diameter of the implicit manifold represented by the knowledge in the buffer during memory management. In addition, we introduce Wasserstein distance instead of cross entropy as distillation loss to preserve previous knowledge. With extensive experimental validation on MNIST, CIFAR10, CIFAR100, and TinyImageNet, we show that the proposed method significantly improves the accuracy in continual learning setup, outperforming the state of the arts. Zihao Xu 0002, Jian Yang 0034, Mingsong Chen 0001, Xian Wei |
SMC | 7 |
| 2023 | Couplformer: Rethinking Vision Transformer with Coupling AttentionabstractWith the development of the self-attention mechanism, the Transformer model has demonstrated its outstanding performance in the computer vision domain. However, the massive computation brought from the full attention mechanism became a heavy burden for memory consumption. Sequentially, the limitation of memory consumption hinders the deployment of the Transformer model on the embedded system where the computing resources are limited. To remedy this problem, we propose a novel memory economy attention mechanism named Couplformer, which decouples the attention map into two sub-matrices and generates the alignment scores from spatial information. Our method enables the Transformer model to improve time and memory efficiency while maintaining expressive power. A series of different scale image classification tasks are applied to evaluate the effectiveness of our model. The result of experiments shows that on the ImageNet-1K classification task, the Couplformer can significantly decrease 42% memory consumption compared with the regular Transformer. Meanwhile, it accesses sufficient accuracy requirements, which outperforms 0.56% on Top-1 accuracy and occupies the same memory footprint. Besides, the Couplformer achieves state-of-art performance in MS COCO 2017 object detection and instance segmentation tasks. As a result, the Couplformer can serve as an efficient backbone in visual tasks and provide a novel perspective on deploying attention mechanisms for researchers. Xihao Wang, Hao Shen 0002, Peidong Liang, Xian Wei |
WACV | 5 |
| 2023 | RGBT tracking using randomly projected CNN features
Yong Wang 0032, Xian Wei, Keping Yu, Lingkun Luo |
Expert Syst. Appl. | 2 |
| 2023 | Missing-edge aware knowledge graph inductive inference through dual graph learning and traversing
Yini Zhang, Yongshen Yang, Xian Wei, Jianhua Luo |
Expert Syst. Appl. | 6 |
| 2023 | Graph matching for knowledge graph alignment using edge-coloring propagation
Xian Wei, Yi Lu Murphey |
Pattern Recognit. | 3 |
| 2023 | FairLight: Fairness-Aware Autonomous Traffic Signal Control With Hierarchical Action SpaceabstractAlthough reinforcement learning (RL) approaches are promising in autonomous traffic signal control (TSC), they often suffer from the unfairness problem that causes extremely long waiting time at intersections for partial vehicles. This is mainly because the traditional RL methods focus on optimizing the overall traffic performance, while the fairness of individual vehicles is neglected. To address this problem, we propose a novel RL-based method named FairLight for the fair and efficient control of traffic with variable phase duration. Inspired by the concept of user satisfaction index (USI) proposed in the transportation field, we introduce a fairness index in the design of key RL elements, which specially considers the travel quality (e.g., fairness). Based on our proposed hierarchical action space method, FairLight can accurately allocate the duration of traffic lights for selected phases. Experimental results obtained from various well-known traffic benchmarks show that, compared with the state-of-the-art RL-based TSC methods, FairLight can not only achieve better fairness performance but also improve the control quality from the perspectives of the average travel time of vehicles and RL convergence speed. Yutong Ye 0001, Jiepin Ding, Ting Wang 0001, Junlong Zhou, Xian Wei, Mingsong Chen 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2023 | Energy and Reliability-Aware Task Scheduling for Cost Optimization of DVFS-Enabled Cloud WorkflowsabstractDue to the increasing complexity, the execution of workflow applications on cloud typically involves a large number of virtual machines (VMs), which makes the cost as well as energy consumption a great concern. To alleviate this issue, more and more cloud service providers introduce new pricing policies considering Dynamic Voltage and Frequency Scaling (DVFS), where users are charged on the basis of allocated CPU frequencies together with various combinations of VM configurations and prices. However, the customizable CPU frequencies make resource provisioning and scheduling harder to achieve a cost-optimal solution. The things become even worse, since lowering CPU voltages of VMs will increase their chance of suffering soft errors, which results in a high rate of completion time failures of workflow applications. To address the above problem, this paper proposes a novel task scheduling method for the purpose of cost optimization based on the genetic algorithm. By introducing new genetic operators and frequency scaling scheme for DVFS-enabled cloud workflows, our approach can quickly figure out cost-optimal resource provisioning and task scheduling solutions by allocating tasks to appropriate VMs with specific operating frequencies under energy, reliability, makespan and memory constraints. Extensive experiments on various well-known scientific workflow benchmarks validate the effectiveness of the proposed method. Comparing with state-of-the-art methods, our approach can significantly reduce the overall cost and energy consumption without violating the given constraints. E. Cao, Saira Musa, Mingsong Chen 0001, Tongquan Wei, Xian Wei, Xin Fu 0001, Meikang Qiu |
IEEE Trans. Cloud Comput. | 5 |
| 2023 | Adaptive Spectral Rotation via Joint Cluster and Pairwise StructureabstractDensity structure and pairwise structure serve as two different but complementary perspectives for clustering. Either side of road is frequently visited and explored by multiple clustering methods. However, there are seldom approaches, which could mutually exploit both structures for clustering. To address this problem, in this paper, we develop a novel adaptive joint clustering algorithm, which combines unsupervised discrete orthogonal least squares discriminant analysis (DOLSDA) and discrete spectral clustering (DSC) with adaptive neighbors and side information into a unified model. Firstly, we extend supervised OLSDA to a discrete kernel clustering problem. To further achieve a clear pairwise structure, a new similarity with adaptive neighbors is then derived to establish sparse Laplacian matrix. In addition, side information could be incorporated to formulate clearer graph by modifying the proposed similarity. Based on the constructed graph, DSC is embedded with the discrete kernel OLSDA (DKOLSDA) clustering to exploit both cluster and pairwise data structures. Equipped with the proposed framework regarding quadratic weighted optimization, adaptive weight can be obtained automatically to leverage both unsupervised DKOLSDA and DSC. Since the unified problem is still discrete, we develop an increment scheme to achieve the optimal spectral rotation for the approximate solution to the predicted indicator. Tong Wu 0006, Rui Zhang 0017, Ziheng Jiao, Xian Wei, Xuelong Li 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2022 | Learning Extremely Lightweight and Robust Model with Differentiable Constraints on Sparsity and Condition Number
Xian Wei, Yangyu Xu, Yanhui Huang, Hairong Lv, Mingsong Chen 0001 |
ECCV (4) | 1 |
| 2022 | Spatial-Frequency Domain Information Integration for Pan-Sharpening
Man Zhou 0003, Jie Huang 0017, Hu Yu 0001, Xueyang Fu, Aiping Liu, Xian Wei, Feng Zhao 0004 |
ECCV (18) | 7 |
| 2022 | Eliminating Backdoor Triggers for Deep Neural Networks Using Attention Relation Graph DistillationabstractDue to the prosperity of Artificial Intelligence (AI) techniques, more and more backdoors are designed by adversaries to attack Deep Neural Networks (DNNs). Although the state-of-the-art method Neural Attention Distillation (NAD) can effectively erase backdoor triggers from DNNs, it still suffers from non-negligible Attack Success Rate (ASR) together with lowered classification ACCuracy (ACC), since NAD focuses on backdoor defense using attention features (i.e., attention maps) of the same order. In this paper, we introduce a novel backdoor defense framework named Attention Relation Graph Distillation (ARGD), which fully explores the correlation among attention features with different orders using our proposed Attention Relation Graphs (ARGs). Based on the alignment of ARGs between teacher and student models during knowledge distillation, ARGD can more effectively eradicate backdoors than NAD. Comprehensive experimental results show that, against six latest backdoor attacks, ARGD outperforms NAD by up to 94.85% reduction in ASR, while ACC can be improved by up to 3.23%. Jun Xia 0003, Ting Wang 0001, Jiepin Ding, Xian Wei, Mingsong Chen 0001 |
IJCAI | 4 |
| 2022 | Geodesic Self-Attention for 3D Point CloudsabstractDue to the outstanding competence in capturing long-range relationships, self-attention mechanism has achieved remarkable progress in point cloud tasks. Nevertheless, point cloud object often has complex non-Euclidean spatial structures, with the behavior changing dynamically and unpredictably. Most current self-attention modules highly rely on the dot product multiplication in Euclidean space, which cannot capture internal non-Euclidean structures of point cloud objects, especially the long-range relationships along the curve of the implicit manifold surface represented by point cloud objects. To address this problem, in this paper, we introduce a novel metric on the Riemannian manifold to capture the long-range geometrical dependencies of point cloud objects to replace traditional self-attention modules, namely, the Geodesic Self-Attention (GSA) module. Our approach achieves state-of-the-art performance compared to point cloud Transformers on object classification, few-shot classification and part segmentation benchmarks. Zihao Xu 0002, Xihao Wang, Mingsong Chen 0001, Xian Wei |
NeurIPS | 7 |
| 2022 | An adaptive social-aware device-to-device communication mechanism for wireless networks
Feras ZenAlden, Suhaidi Hassan, Adib Habbal, Xian Wei |
Ad Hoc Networks | 4 |
| 2022 | Response map evaluation for RGBT tracking
Yong Wang 0032, Xian Wei, Jiangxiong Fang |
Neural Comput. Appl. | 2 |
| 2022 | Adaptive Fusion CNN Features for RGBT Object TrackingabstractThermal sensors play an important role in intelligent transportation system. This paper studies the problem of RGB and thermal (RGBT) tracking in challenging situations by leveraging multimodal data. A RGBT object tracking method is proposed in correlation filter tracking framework based on short term historical information. Given the initial object bounding box, hierarchical convolutional neural network (CNN) is employed to extract features. The target is tracked for RGB and thermal modalities separately. Then the backward tracking is implemented in the two modalities. The difference between each pair is computed, which is an indicator of the tracking quality in each modality. Considering the temporal continuity of sequence frames, we also incorporate the history data into the weights computation to achieve a robust fusion of different source data. Experiments on three RGBT datasets show the proposed method achieves comparable results to state-of-the-art methods. Yong Wang 0032, Xian Wei, Hao Shen 0002, Huanlong Zhang |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2021 | Boost Neural Networks by CheckpointsabstractTraining multiple deep neural networks (DNNs) and averaging their outputs is a simple way to improve the predictive performance. Nevertheless, the multiplied training cost prevents this ensemble method to be practical and efficient. Several recent works attempt to save and ensemble the checkpoints of DNNs, which only requires the same computational cost as training a single network. However, these methods suffer from either marginal accuracy improvements due to the low diversity of checkpoints or high risk of divergence due to the cyclical learning rates they adopted. In this paper, we propose a novel method to ensemble the checkpoints, where a boosting scheme is utilized to accelerate model convergence and maximize the checkpoint diversity. We theoretically prove that it converges by reducing exponential loss. The empirical evaluation also indicates our proposed ensemble outperforms single model and existing ensembles in terms of accuracy and efficiency. With the same training budget, our method achieves 4.16% lower error on Cifar-100 and 6.96% on Tiny-ImageNet with ResNet-110 architecture. Moreover, the adaptive sample weights in our method make it an effective solution to address the imbalanced class distribution. In the experiments, it yields up to 5.02% higher accuracy over single EfficientNet-B0 on the imbalanced datasets. Feng Wang 0047, Guoyizhe Wei, Qiao Liu 0008, Jinxiang Ou, Xian Wei, Hairong Lv |
NeurIPS | 5 |
| 2021 | Performance evaluation of low resolution visual tracking for unmanned aerial vehicles
Yong Wang 0032, Xian Wei, Hao Shen 0002, Jilin Hu, Lingkun Luo |
Neural Comput. Appl. | 2 |
| 2021 | Deep Clustering With Intraclass Distance Constraint for Hyperspectral ImagesabstractThe high dimensionality of hyperspectral images often results in the degradation of clustering performance. Due to the powerful ability of potential feature extraction and nonlinear representation, deep clustering algorithms have become a hot topic in hyperspectral remote sensing. Different tasks often need different features. However, the current deep clustering algorithms generally separate feature extraction from clustering, which results in the extracted features that are not constrained by clustering tasks. Therefore, the features extracted by these algorithms may not be suitable for clustering. To address this issue, we adopt intraclass distance as a constraint condition and proposed an intraclass distance constrained deep clustering algorithm for hyperspectral images. The proposed algorithm propagates the clustering error back to the feature mapping process of the autoencoder network, so as to realize the constraint of clustering objective on feature extraction and make the extracted features more suitable for clustering tasks. In addition, the proposed algorithm simultaneously completes network optimization and clustering, which is more efficient. Experimental results demonstrate the intense competitiveness of the proposed algorithm in comparison with state-of-the-art clustering methods for hyperspectral images. Jinguang Sun, Xian Wei, Xiaoliang Tang, Yusheng Xu, Wei Yao 0008 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2020 | Adaptive Spatio-Temporal Graph Convolutional Neural Network for Remaining Useful Life EstimationabstractAccurate remaining useful life (RUL) estimation is of crucial importance to numerous industrial applications where safety and reliability are among primary concerns. Recently, deep learning based prognostics methods have been emerging as an effective method to improve RUL prediction results. However, these methods, e.g. recurrent neural networks (RNNs), convolutional neural networks (CNNs), only capture temporal information of the sensory data while ignoring intrinsic spatial relations between sensors. To solve this problem, in this work, we propose a framework, namely, adaptive spatio-temporal graph convolutional neural network (ASTGCNN). The proposed framework consists of two parts. In the spatial domain, since the intrinsic graph structure of sensors is not provided in most situations, a dynamic graph neural network is proposed to learn the sensors' spatial relation. In the temporal domain, a stacked dilated ID CNN is utilized to capture long range dependency of input sensor signals. These two parts are integrated in a unified framework and can be trained in an end-to-end manner. The performance of ASTGCNN is investigated on the turbofan engine dataset Experimental results show that the proposed framework can improve the RUL prediction performance of the current deep learning methods, and learn the intrinsic spatial information of sensors. Xian Wei |
IJCNN | 3 |
| 2020 | Robust RGB-D tracking via compact CNN features
Yong Wang 0032, Xian Wei, Lingkun Luo |
Eng. Appl. Artif. Intell. | 2 |
| 2020 | CNN tracking based on data augmentation
Yong Wang 0032, Xian Wei, Hao Shen 0002 |
Knowl. Based Syst. | 2 |
| 2020 | Trace Quotient with Sparsity Priors for Learning Low Dimensional Image RepresentationsabstractThis work studies the problem of learning appropriate low dimensional image representations. We propose a generic algorithmic framework, which leverages two classic representation learning paradigms, i.e., sparse representation and the trace quotient criterion, to disentangle underlying factors of variation in high dimensional images. Specifically, we aim to learn simple representations of low dimensional, discriminant factors by applying the trace quotient criterion to well-engineered sparse representations. We construct a unified cost function, coined as the SPARse LOW dimensional representation (SparLow) function, for jointly learning both a sparsifying dictionary and a dimensionality reduction transformation. The SparLow function is widely applicable for developing various algorithms in three classic machine learning scenarios, namely, unsupervised, supervised, and semi-supervised learning. In order to develop efficient joint learning algorithms for maximizing the SparLow function, we deploy a framework of sparse coding with appropriate convex priors to ensure the sparse representations to be locally differentiable. Moreover, we develop an efficient geometric conjugate gradient algorithm to maximize the SparLow function on its underlying Riemannian manifold. Performance of the proposed SparLow algorithmic framework is investigated on several image processing tasks, such as 3D data visualization, face/digit recognition, and object/scene categorization. Xian Wei, Hao Shen 0002, Martin Kleinsteuber |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2020 | UAV tracking based on saliency detection
Yong Wang 0032, Xinbin Luo, Lingkun Luo, Huanlong Zhang, Xian Wei |
Soft Comput. | 5 |
| 2020 | Adaptive model updating for robust object tracking
Yong Wang 0032, Xian Wei, Hao Shen 0002 |
Signal Process. Image Commun. | 2 |
| 2020 | A robust visual tracking method via local feature extraction and saliency detection
Yong Wang 0032, Xian Wei, Xiaoliang Tang, Huanlong Zhang |
Vis. Comput. | 2 |
| 2019 | Robust visual tracking based on response stability
Yong Wang 0032, Xinbin Luo, Shan Fu, Xian Wei |
Eng. Appl. Artif. Intell. | 5 |
| 2019 | Detection based visual tracking with convolutional neural network
Yong Wang 0032, Xinbin Luo, Shan Fu, Xian Wei |
Knowl. Based Syst. | 5 |
| 2019 | Joint Learning of Fuzzy k-Means and Nonnegative Spectral Clustering With Side InformationabstractAs one of the most widely used clustering techniques, the fuzzy k-means (FKM) assigns every data point to each cluster with a certain degree of membership. However, conventional FKM approach relies on the square data fitting term, which is sensitive to the outliers with ignoring the prior information. In this paper, we develop a novel and robust fuzzy k-means clustering algorithm, namely, joint learning of fuzzy k-means and nonnegative spectral clustering with side information. The proposed method combines fuzzy k-means and nonnegative spectral clustering into a unified model, which can further exploit the prior knowledge of data pairs such that both the quality of affinity graph and the clustering performance can be improved. In addition, for the purpose of enhancing the robustness, the adaptive loss function is adopted in the objective function, since it smoothly interpolates between 11-norm and 12-norm. Finally, experimental results on benchmark datasets verify the effectiveness and the superiority of our clustering method. Rui Zhang 0017, Feiping Nie 0001, Muhan Guo, Xian Wei, Xuelong Li 0001 |
IEEE Trans. Image Process. | 4 |
| 2019 | Reconstructible Nonlinear Dimensionality Reduction via Joint Dictionary LearningabstractThis paper presents a parametric low-dimensional (LD) representation learning method that allows to reconstruct high-dimensional (HD) input vectors in an unsupervised manner. Under the assumption that the HD data and its LD representation share the same or similar local sparse structure, the proposed method achieves reconstructible dimensionality reduction via jointly learning dictionaries in both the original HD data space and its LD representation space. By regarding the sparse representation as a smooth function with respect to a specific dictionary, we construct an encoding-decoding block for learning LD representations from sparse coefficients of HD data. It is expected that this learning process preserves the desirable structure of HD data in the LD representation space, and simultaneously allows a reliable reconstruction from the LD space back to the original HD space. In addition, the proposed single layer encoding-decoding block can be easily extended to deep learning structures. Numerical experiments on both synthetic data sets and real images show that the proposed method achieves strongly competitive and robust performance in data DR, reconstruction, and synthesis, even on heavily corrupted data. The proposed method can be used as an alternative approach to compressive sensing (CS); however, it can outperform the traditional CS methods in: 1) task-driven learning problems, such as 2-D/3-D data visualization, and 2) data reconstruction at a lower dimensional space. Xian Wei, Hao Shen 0002, Martin Kleinsteuber, Yi Lu Murphey |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2017 | Joint learning sparsifying linear transformation for low-resolution image synthesis and recognition
Xian Wei, Hao Shen 0002, Weidong Xiang, Yi Lu Murphey |
Pattern Recognit. | 1 |
| 2017 | Dynamical Textures Modeling via Joint Video Dictionary LearningabstractVideo representation is an important and challenging task in the computer vision community. In this paper, we consider the problem of modeling and classifying video sequences of dynamic scenes which could be modeled in a dynamic textures (DTs) framework. At first, we assume that image frames of a moving scene can be modeled as a Markov random process. We propose a sparse coding framework, named joint video dictionary learning (JVDL), to model a video adaptively. By treating the sparse coefficients of image frames over a learned dictionary as the underlying "states", we learn an efficient and robust linear transition matrix between two adjacent frames of sparse events in time series. Hence, a dynamic scene sequence is represented by an appropriate transition matrix associated with a dictionary. In order to ensure the stability of JVDL, we impose several constraints on such transition matrix and dictionary. The developed framework is able to capture the dynamics of a moving scene by exploring both the sparse properties and the temporal correlations of consecutive video frames. Moreover, such learned JVDL parameters can be used for various DT applications, such as DT synthesis and recognition. Experimental results demonstrate the strong competitiveness of the proposed JVDL approach in comparison with the state-of-the-art video representation methods. Especially, it performs significantly better in dealing with DT synthesis and recognition on heavily corrupted data. Xian Wei, Hao Shen 0002, Martin Kleinsteuber, Zhongfeng Wang 0001 |
IEEE Trans. Image Process. | 1 |
| 2016 | Trace Quotient Meets Sparsity: A Method for Learning Low Dimensional Image RepresentationsabstractThis paper presents an algorithm that allows to learn low dimensional representations of images in an unsupervised manner. The core idea is to combine two criteria that play important roles in unsupervised representation learning, namely sparsity and trace quotient. The former is known to be a convenient tool to identify underlying factors, and the latter is known as a disentanglement of underlying discriminative factors. In this work, we develop a generic cost function for learning jointly a sparsifying dictionary and a dimensionality reduction transformation. It leads to several counterparts of classic low dimensional representation methods, such as Principal Component Analysis, Local Linear Embedding, and Laplacian Eigenmap. Our proposed optimisation algorithm leverages the efficiency of geometric optimisation on Riemannian manifolds and a closed form solution to the elastic net problem. Xian Wei, Hao Shen 0002, Martin Kleinsteuber |
CVPR | 1 |
| 2016 | Joint learning dictionary and discriminative features for high dimensional dataabstractRecently, sparse representation (SR) over a redundant dictionary has become a popular way of representing the data. It has been verified as an efficient and useful tool to promote the discrimination between signals. This work develops a joint learning approach to find the low dimensional discriminative features for high dimensional data. To avoid the high computational cost of direct sparse coding on large scale input data, we first learn SR in an orthogonal projected space over a task-driven sparsifying dictionary. We then exploit the discriminative projection on SR. The whole learning process is treated as an optimization problem of trace quotient maximization, which involves an orthogonal projection on original data space, a dictionary and a discriminative projection on sparse codes. The related cost function is well defined on a product manifold of the Stiefel manifold, the Oblique manifold and the Grassmann manifold. Finally, we employ a stochastic gradient descent algorithm on the smooth product manifold to maximize the cost function. Our numerical experiments on visual recognition demonstrate the effectiveness of the proposed algorithm, in comparison with the state of the arts. Xian Wei, Hao Shen 0002, Martin Kleinsteuber, Yi Lu Murphey |
ICPR | 1 |
| 2014 | An adaptive dictionary learning approach for modeling dynamical texturesabstractVideo representation is an important and challenging task in the computer vision community. In this paper, we assume that image frames of a moving scene can be modeled as a Markov random process. We propose a sparse coding framework, named adaptive video dictionary learning (AVDL), to model a video adaptively. The developed framework is able to capture the dynamics of a moving scene by exploring both sparse properties and the temporal correlations of consecutive video frames. The proposed method is compared with state of the art video processing methods on several benchmark data sequences, which exhibit appearance changes and heavy occlusions. Xian Wei, Hao Shen 0002, Martin Kleinsteuber |
ICASSP | 1 |