VLDB 2026 Research / reviewers in the wild / expert
Kai-Xuan Chen 0001
dblp:220/5629 · also Kaixuan Chen 0004
· DBLP profile ↗
33ranked-venue papers
7as first author
29since 2021 · last 2026
0000-0002-2492-5230ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 5 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 2 first-author · 11 since 2021Databases, data management, data science and information retrieval · 5 · 3 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Evolutionary Negative Module Pruning for Better LoRA MergingabstractAnda Cao, Zhuo Gou, Yi Wang, Kaixuan Chen, Yu Wang, Can Wang, Mingli Song, Jie Song. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Anda Cao, Zhuo Gou, Yi Wang 0068, Kai-Xuan Chen 0001, Yu Wang 0176, Mingli Song, Jie Song 0011 |
ACL (1) | 4 |
| 2026 | Unsupervised Action Segmentation via Multi-Scale Temporal-Interaction EnhancementabstractUnsupervised action segmentation (UAS) aims to identify action boundaries in long, untrimmed videos without the use of annotations. This involves learning discriminative frame features and applying a segmentation mechanism to organize frames into coherent action segments. However, most of the common approaches ignore the importance of multi-scale temporal interactions within the video sequence, resulting in a limited frame representation capability and inaccurate action boundary detection. In this paper, we propose MulSclTE, a novel UAS framework that incorporates multi-scale temporal interactions across global, clip, and frame levels to enhance the overall performance. To address the limited representation capability, we first present global-level interaction enhancement by implementing a bi-directional temporal encoding mechanism, designed to capture comprehensive information across the entire sequence. Then, we devise a hierarchical self-supervised loss function equipped with a clip-level interaction constraint that aims to bring temporally adjacent clips closer while separating non-adjacent ones. To precisely identify action boundaries, we provide comprehensive information by integrating frame-level prediction errors and similarity scores to alleviate the under-segmentation issue, and present a refinement mechanism to mitigate the over-segmentation issue. Extensive experiments on Breakfast, YouTube Instructions, 50Salads, and EPIC-KITCHENS show that MulSclTE attains leading or second-best performance across all datasets, and even exceeds some supervised methods in MoF and F1 metrics, underscoring its robustness and effectiveness. Zhiying Song, Kai-Xuan Chen 0001, Mingli Song, Nenggan Zheng |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Agent-Aware Training for Agent-Agnostic Action Advising in Deep Reinforcement LearningabstractAction advising endeavors to leverage supplementary guidance from expert teachers to alleviate the issue of sampling inefficiency in Deep Reinforcement Learning (DRL). Previous agent-specific action advising methods are hindered by imperfections in the agent itself, while agent-agnostic approaches exhibit limited adaptability to the learning agent. In this study, we propose a novel framework called Agent-Aware trAining yet Agent-Agnostic Action Advising (A7) to strike a balance between the two. The underlying concept of A7 revolves around utilizing the similarity of state features as an indicator for soliciting advice. However, unlike prior methodologies, the measurement of state feature similarity is performed by neither the error-prone learning agent nor the agent-agnostic advisor. Instead, we employ a proxy model to extract state features that are both discriminative (adaptive to the agent) and generally applicable (robust to agent noise). Furthermore, we utilize behavior cloning to train a model for reusing advice and introduce an intrinsic reward for the advised samples to incentivize the utilization of expert guidance. Experiments are conducted on the GridWorld, LunarLander, and six prominent scenarios from Atari games. The results demonstrate that A7 significantly accelerates the learning process and surpasses existing methods (both agent- specific and agent-agnostic) by a substantial margin. Our code will be made publicly available. Yaoquan Wei, Shunyu Liu 0001, Jie Song 0011, Tongya Zheng, Kai-Xuan Chen 0001, Mingli Song |
AAAI | 5 |
| 2025 | Disentangled Table-Graph Representation for Interpretable Transmission Line Fault LocationabstractThe fault location task in power grids is crucial for maintaining social order and ensuring public safety. However, existing methods that rely on tabular state records often neglect the intrinsic topological influences of transmission lines, resulting in a segmented approach to fault location that consists of multiple stages. In this paper, we propose an Disentangled Table-Graph representation framework, termed DTG, which integrates fault location tasks at coarse-grained line levels and fine-grained point levels within an end-to-end learning paradigm. Our innovative disentanglement strategy produces interpretable attribution coefficients that connect tabular records and transmission line topology, thereby facilitating fault location at both line- and point-levels. The joint prediction tasks designed around our disentangled tabular graph representation promote mutual information exchange between features and topology of transmission lines in an interpretable manner. Experimental results on the 7-bus system, 36-bus system and a realistic 325-bus system in China demonstrate that the proposed method adapt to different topological structures and handle different types of faults. Compared to traditional methods, DTG4Power achieves high accuracy in both fault lines and fault points. Na Yu 0001, Yutong Deng, Shunyu Liu 0001, Kai-Xuan Chen 0001, Tongya Zheng, Mingli Song |
AAAI | 4 |
| 2025 | Cooperative Policy Agreement: Learning Diverse Policy for Offline MARLabstractOffline Multi-Agent Reinforcement Learning (MARL) aims to learn optimal joint policies from pre-collected datasets without further interaction with the environment. Despite the encouraging results achieved so far, we identify the policy mismatch problem that arises from employing diverse offline MARL datasets, a highly important ingredient for cooperative generalization yet largely overlooked by existing literature. Specifically, in the case that offline datasets exhibit various optimal joint policies, policy mismatch often occurs when individual actions from different optimal joint actions are combined in a way that results in a suboptimal joint action. In this paper, we introduce a novel Cooperative Policy Agreement (CPA) method, that not only mitigates the policy mismatch problem but also learns to generate diverse joint policies. CPA firstly introduces an autoregressive decision-making mechanism among agents during offline training. This mechanism enables agents to access the actions previously taken by other agents, thereby facilitating effective joint policy matching. Moreover, diverse joint policies can be directly obtained through sequential action sampling from the autoregressive model. Then we further incorporate a policy agreement mechanism to convert these autoregressive joint policies into decentralized policies with a non-autoregressive form, while still ensuring the diversity of the generated policies. This mechanism guarantees that the proposed CPA adheres to the Centralized Training with Decentralized Execution (CTDE) constraint. Experiments conducted on various benchmarks demonstrate that CPA yields superior performance to state-of-the-art competitors. Yihe Zhou, Yuxuan Zheng, Kai-Xuan Chen 0001, Tongya Zheng, Jie Song 0011, Mingli Song, Shunyu Liu 0001 |
AAAI | 4 |
| 2025 | Spatial-Temporal Reconstruction Error for AIGC-based Forgery Image DetectionabstractThe remarkable success of AI-Generated Content (AIGC), especially diffusion image generation models, brings about unprecedented creative applications, but also creates fertile ground for malicious counterfeiting and crime. A highly effective family of forgery image detection methods based on diffusion reconstruction error has emerged, as images generated by diffusion are more easily reconstructed by any diffusion model. However, we find that existing methods only use reconstruction error from a single time step, failing to fully leverage the entire reconstruction process. To this end, we propose to comprehensively consider every single time step to form the Temporal Reconstruction Error (TRE) that offers a richer feature representation. Furthermore, we design temporal aggregation and spatial focusing modules from two dimensions respectively to more effectively extract discriminative information from the TRE feature. Finally, we validate the proposed method on two popular datasets, and experimental results demonstrate that the proposed approach achieves state-of-the-art performance. Chengji Shen, Zhenjiang Liu, Kai-Xuan Chen 0001, Jie Lei 0002, Mingli Song, Zunlei Feng |
ICASSP | 3 |
| 2025 | From GNNs to Trees: Multi-Granular Interpretability for Graph Neural NetworksabstractInterpretable Graph Neural Networks (GNNs) aim to reveal the underlying reasoning behind model predictions, attributing their decisions to specific subgraphs that are informative. However, existing subgraph-based interpretable methods suffer from an overemphasis on local structure, potentially overlooking long-range dependencies within the entire graphs. Although recent efforts that rely on graph coarsening have proven beneficial for global interpretability, they inevitably reduce the graphs to a fixed granularity. Such an inflexible way can only capture graph connectivity at a specific level, whereas real-world graph tasks often exhibit relationships at varying granularities (e.g., relevant interactions in proteins span from functional groups, to amino acids, and up to protein domains). In this paper, we introduce a novel Tree-like Interpretable Framework (TIF) for graph classification, where plain GNNs are transformed into hierarchical trees, with each level featuring coarsened graphs of different granularity as tree nodes. Specifically, TIF iteratively adopts a graph coarsening module to compress original graphs (i.e., root nodes of trees) into increasingly coarser ones (i.e., child nodes of trees), while preserving diversity among tree nodes within different branches through a dedicated graph perturbation module. Finally, we propose an adaptive routing module to identify the most informative root-to-leaf paths, providing not only the final prediction but also the multi-granular interpretability for the decision-making process. Extensive experiments on the graph classification benchmarks with both synthetic and real-world datasets demonstrate the superiority of TIF in interpretability, while also delivering a competitive prediction performance akin to the state-of-the-art counterparts. Kai-Xuan Chen 0001, Tongya Zheng, Yihe Zhou, Zhenbang Xiao, Ji Cao 0001, Mingli Song, Shunyu Liu 0001 |
ICLR | 3 |
| 2025 | CADP: Towards Better Centralized Learning for Decentralized Execution in MARL
Yihe Zhou, Shunyu Liu 0001, Yunpeng Qing, Tongya Zheng, Kai-Xuan Chen 0001, Jie Song 0011, Mingli Song |
AAMAS | 5 |
| 2025 | CADP: Towards Better Centralized Learning for Decentralized Execution in MARLabstractCentralized Training with Decentralized Execution (CTDE) has recently emerged as a popular framework for cooperative Multi-Agent Reinforcement Learning (MARL), where agents can use additional global state information to guide training in a centralized way and make their own decisions only based on decentralized local policies. Despite the encouraging results achieved, CTDE makes an independence assumption on agent policies, which limits agents from adopting global cooperative information from each other during centralized training. Therefore, we argue that the existing CTDE framework cannot fully utilize global information for training, leading to an inefficient joint exploration and perception, which can degrade the final performance. In this paper, we introduce a novel Centralized Advising and Decentralized Pruning (CADP) framework for MARL, that not only enables an efficacious message exchange among agents during training but also guarantees the independent policies for decentralized execution. Firstly, CADP endows agents the explicit communication channel to seek and take advice from different agents for more centralized training. To further ensure the decentralized execution, we propose a smooth model pruning mechanism to progressively constrain the agent communication into a closed one without degradation in agent cooperation capability. Empirical evaluations on different benchmarks and across various MARL backbones demonstrate that the proposed framework achieves superior performance compared with the state-of-the-art counterparts. Our code is available at https://github.com/zyh1999/CADP Yihe Zhou, Shunyu Liu 0001, Yunpeng Qing, Tongya Zheng, Kai-Xuan Chen 0001, Jie Song 0011, Mingli Song |
IJCAI | 5 |
| 2025 | Powerformer: A Section-adaptive Transformer for Power Flow AdjustmentabstractIn this paper, we present a novel transformer architecture tailored for learning robust power system state representations, which strives to optimize power dispatch for the power flow adjustment across different transmission sections. Specifically, our proposed approach, named Powerformer, develops a dedicated section-adaptive attention mechanism, separating itself from the self-attention employed in conventional transformers. This mechanism effectively integrates power system states with transmission section information, which facilitates the development of robust state representations. Furthermore, by considering the graph topology of power system and the electrical attributes of bus nodes, we introduce two customized strategies to further enhance the expressiveness: graph neural network propagation and multi-factor attention mechanism. Extensive evaluations are conducted on three power system scenarios, including the IEEE 118-bus system, a realistic China 300-bus system, and a large-scale European system with 9241 buses, where Powerformer demonstrates its superior performance over several popular baseline methods. The code is available at: https://github.com/Cra2yDavid/Powerformer Kai-Xuan Chen 0001, Shunyu Liu 0001, Yaoquan Wei, Yihe Zhou, Yunpeng Qing, Jie Song 0011, Mingli Song |
KDD (1) | 1 |
| 2025 | Hi-Motion: Hierarchical Intention Guided Conditional Motion SynthesisabstractText-conditioned motion generation has significant applications across various domains. However, generating natural motion remains challenging due to the vast solution space and the accumulation of errors during motion generation. To address these challenges, a novel hierarchical motion intention decoding-based motion synthesis model named Hi-Motion is proposed, which disentangles human motion into temporal intents of pivot joints and skeleton synthesis guided by intention from a new perspective. Specifically, Hi-Motion first parameterizes pivot joint motion with high-order Bézier curves and constructs a Bézier decoder to generate their trajectories, which serve as motion intention to guide skeleton generation. Secondly, we formulate the generation of skeletons as a graph node transformation problem under the condition of determined edge connections. By incorporating hierarchical joint motion intentions into the graph node features, the spatial details of each frame can be precisely synthesized. The proposed Hi-Motion effectively decouples motion generation into temporal and spatial dimensions through hierarchical motion intention decoding, ensuring coordination and naturalness in the generated motion. Extensive experiments on HumanML3D and KIT-ML datasets substantiate the motion generation capabilities of Hi-Motion. Further analysis demonstrates that Hi-Motion can accurately predict the motion intention of pivot joints and synthesize skeletal details. Le Han, Kai-Xuan Chen 0001, Minchen Ye, Nenggan Zheng |
ACM Multimedia | 2 |
| 2025 | SeRL: Self-play Reinforcement Learning for Large Language Models with Limited DataabstractRecent advances have demonstrated the effectiveness of Reinforcement Learning (RL) in improving the reasoning capabilities of Large Language Models (LLMs). However, existing works inevitably rely on high-quality instructions and verifiable rewards for effective training, both of which are often difficult to obtain in specialized domains. In this paper, we propose Self-play Reinforcement Learning (SeRL) to bootstrap LLM training with limited initial data. Specifically, SeRL comprises two complementary modules: self-instruction and self-rewarding. The former module generates additional instructions based on the available data at each training step, employing comprehensive online filtering strategies to ensure instruction quality, diversity, and difficulty. The latter module introduces a simple yet effective majority-voting mechanism to estimate response rewards for additional instructions, eliminating the need for external annotations. Finally, SeRL performs conventional RL based on the generated data, facilitating iterative self-play learning.
Extensive experiments on various reasoning benchmarks and across different LLM backbones demonstrate that the proposed SeRL yields results superior to its counterparts and achieves performance on par with those obtained by high-quality data with verifiable rewards. Our code is available at https://github.com/wantbook-book/SeRL. Wenkai Fang, Shunyu Liu 0001, Kongcheng Zhang, Tongya Zheng, Kai-Xuan Chen 0001, Mingli Song, Dacheng Tao |
NeurIPS | 6 |
| 2025 | Cross-Domain Animal Pose Estimation With Skeleton Anomaly-Aware LearningabstractAnimal pose estimation is often constrained by the scarcity of annotations and the diversity of scenarios and species. The pseudo-label generation based unsupervised domain adaptation paradigm, which discriminates the predicted keypoints of unlabeled data based on the skeleton position consistency, has demonstrated effectiveness for such problems. However, existing methods generate pseudo-labels with massive false positives, because they cannot effectively distinguish sample pairs with the same errors. In this study, we propose a cross-domain animal pose estimation model from a novel perspective of skeleton anomaly learning. We construct a graph contrastive learning mechanism to acquire the skeleton anomaly-aware knowledge, which enables the generation of accurate pseudo-labels for target domain and imposes graph constraint on unlabeled data. And a skeleton anomaly-feedback based domain adaptation framework is designed to facilitate implicit alignment of object-specific features and joint training of cross-domain. Besides, we propose a novel rat pose dataset named UDARP-9.4K to address the gap of small-sized animal pose datasets encompassing diverse experimental scenarios. The related datasets are reviewed and evaluated in detail. Extensive experiments are conducted on UDARP-9.4K and two public datasets to demonstrate the superiority of the proposed model in cross-scenarios and cross-species animal pose estimation tasks. Further analysis reveals the effectiveness of the proposed model for skeleton structure feature learning.The UDARP-9.4K dataset is available here. Le Han, Kai-Xuan Chen 0001, Lei Zhao 0026, Yangbo Jiang, Nenggan Zheng |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Learning a Mini-Batch Graph Transformer via Two-Stage Interaction AugmentationabstractMini-batch Graph Transformer (MGT), as an emerging graph learning model, has demonstrated significant advantages in semi-supervised node prediction tasks with improved computational efficiency and enhanced model robustness. However, existing methods for processing local information either rely on sampling or simple aggregation, which respectively result in the loss and squashing of critical neighbor information. Moreover, the limited number of nodes in each mini-batch restricts the model’s capacity to capture the global characteristic of the graph. In this paper, we propose LGMformer, a novel MGT model that employs a two-stage augmented interaction strategy, transitioning from local to global perspectives, to address the aforementioned bottlenecks. The local interaction augmentation (LIA) presents a neighbor-target interaction Transformer (NTIformer) to acquire an insightful understanding of the co-interaction patterns between neighbors and the target node, resulting in a locally effective token list that serves as input for the MGT. In contrast, global interaction augmentation (GIA) adopts a cross-attention mechanism to incorporate entire graph prototypes into the target node representation, thereby compensating for the global graph information to ensure a more comprehensive perception. To this end, LGMformer achieves the enhancement of node representations under the MGT paradigm. Experimental results related to node classification on the ten benchmark datasets demonstrate the effectiveness of the proposed method. Our code is available at https://github.com/l-wd/LGMformer. Wenda Li 0003, Kai-Xuan Chen 0001, Shunyu Liu 0001, Tongya Zheng, Mingli Song |
ECAI | 2 |
| 2024 | Unified Mask Graph Modeling for Incomplete Tabular Learning
Na Yu 0001, Tongya Zheng, Shunyu Liu 0001, Kai-Xuan Chen 0001, Mingli Song |
ICONIP (6) | 4 |
| 2024 | Multi-Channel Graph Fusion Representation for Tabular Data ImputationabstractThe unprecedented success of deep learning has revolutionized the imputation mechanism of missing values in tabular data, typically caused by data corruption and sensor noise. One promising approach involves mining the latent relationship among all entities of tabular data to complete the missing values from similar entities. However, the various data absence challenges an effective identification of similar entities with missing attributes. Moreover, this limited entity-level relationship fails to capture the intricate interdependencies that typically arise between different attributes, leading to potential bias in data imputation. In this paper, we propose to build an innovative customized graph for tabular data, termed MCG4Table, that allows us to facilitate the representation of intra-entity, intra-attribute, and entity-attribute relationships. At the heart of our approach is a novel strategy to establish a multi-channel graph by dividing diverse relationships concealed within the tabular data from different perspectives. Moreover, MCG4Table introduces a two-stage channel feature fusion architecture, namely a homogeneous-then-heterogeneous fusion strategy, yielding enhanced representations of entities and attributes for the downstream imputation task. Extensive experiments on several benchmark datasets demonstrate the superiority of our proposed method. Elaborate ablation studies and parameter sensitivity analysis verify the effectiveness and robustness of our dedicated strategies. Our code will be made publicly available. Na Yu 0001, Kai-Xuan Chen 0001, Shunyu Liu 0001, Tongya Zheng, Mingli Song |
IJCNN | 3 |
| 2024 | Unveiling Global Interactive Patterns across Graphs: Towards Interpretable Graph Neural NetworksabstractGraph Neural Networks (GNNs) have emerged as a prominent framework for graph mining, leading to significant advances across various domains. Stemmed from the node-wise representations of GNNs, existing explanation studies have embraced the subgraph-specific viewpoint that attributes the decision results to the salient features and local structures of nodes. However, graph-level tasks necessitate long-range dependencies and global interactions for advanced GNNs, deviating significantly from subgraph-specific explanations. To bridge this gap, this paper proposes a novel intrinsically interpretable scheme for graph classification, termed as Global Interactive Pattern (GIP) learning, which introduces learnable global interactive patterns to explicitly interpret decisions. GIP first tackles the complexity of interpretation by clustering numerous nodes using a constrained graph clustering module. Then, it matches the coarsened global interactive instance with a batch of self-interpretable graph prototypes, thereby facilitating a transparent graph-level reasoning process. Extensive experiments conducted on both synthetic and real-world benchmarks demonstrate that the proposed GIP yields significantly superior interpretability and competitive performance to the state-of-the-art counterparts. Our code will be made publicly available¹. Shunyu Liu 0001, Tongya Zheng, Kai-Xuan Chen 0001, Mingli Song |
KDD | 4 |
| 2024 | A2PO: Towards Effective Offline Reinforcement Learning from an Advantage-aware PerspectiveabstractOffline reinforcement learning endeavors to leverage offline datasets to craft effective agent policy without online interaction, which imposes proper conservative constraints with the support of behavior policies to tackle the out-of-distribution problem. However, existing works often suffer from the constraint conflict issue when offline datasets are collected from multiple behavior policies, i.e., different behavior policies may exhibit inconsistent actions with distinct returns across the state space. To remedy this issue, recent advantage-weighted methods prioritize samples with high advantage values for agent training while inevitably ignoring the diversity of behavior policy. In this paper, we introduce a novel Advantage-Aware Policy Optimization (A2PO) method to explicitly construct advantage-aware policy constraints for offline learning under mixed-quality datasets. Specifically, A2PO employs a conditional variational auto-encoder to disentangle the action distributions of intertwined behavior policies by modeling the advantage values of all training data as conditional variables. Then the agent can follow such disentangled action distribution constraints to optimize the advantage-aware policy towards high advantage values. Extensive experiments conducted on both the single-quality and mixed-quality datasets of the D4RL benchmark demonstrate that A2PO yields results superior to the counterparts. Our code is available at https://github.com/Plankson/A2PO. Yunpeng Qing, Shunyu Liu 0001, Jingyuan Cong, Kai-Xuan Chen 0001, Yihe Zhou, Mingli Song |
NeurIPS | 4 |
| 2024 | Interaction Pattern Disentangling for Multi-Agent Reinforcement LearningabstractDeep cooperative multi-agent reinforcement learning has demonstrated its remarkable success over a wide spectrum of complex control tasks. However, recent advances in multi-agent learning mainly focus on value decomposition while leaving entity interactions still intertwined, which easily leads to over-fitting on noisy interactions between entities. In this work, we introduce a novel interactiOn Pattern disenTangling (OPT) method, to disentangle the entity interactions into interaction prototypes, each of which represents an underlying interaction pattern within a subgroup of the entities. OPT facilitates filtering the noisy interactions between irrelevant entities and thus significantly improves generalizability as well as interpretability. Specifically, OPT introduces a sparse disagreement mechanism to encourage sparsity and diversity among discovered interaction prototypes. Then the model selectively restructures these prototypes into a compact interaction pattern by an aggregator with learnable weights. To alleviate the training instability issue caused by partial observability, we propose to maximize the mutual information between the aggregation weights and the history behaviors of each agent. Experiments on single-task, multi-task and zero-shot benchmarks demonstrate that the proposed method yields results superior to the state-of-the-art counterparts. Shunyu Liu 0001, Jie Song 0011, Yihe Zhou, Na Yu 0001, Kai-Xuan Chen 0001, Zunlei Feng, Mingli Song |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | Spatiotemporal-Augmented Graph Neural Networks for Human Mobility SimulationabstractHuman mobility patterns have shown significant applications in policy-decision scenarios and economic behavior researches. The human mobility simulation task aims to generate human mobility trajectories given a small set of trajectory data, which have aroused much concern due to the scarcity and sparsity of human mobility data. Existing methods mostly rely on the static relationships of locations, while largely neglect the dynamic spatiotemporal effects of locations. On the one hand, spatiotemporal correspondences of visit distributions reveal the spatial proximity and the functionality similarity of locations. On the other hand, the varying durations in different locations hinder the iterative generation process of the mobility trajectory. Therefore, we propose a novel framework to model the dynamic spatiotemporal effects of locations, namelySpatioTemporal-Augmented gRaph neural networks (STAR). The STAR framework designs various spatiotemporal graphs to capture the spatiotemporal correspondences and builds a novel dwell branch to simulate the varying durations in locations, which is finally optimized in an adversarial manner. The comprehensive experiments over four real datasets for the human mobility simulation have verified the superiority of STAR tostate-of-the-artmethods. Our code is available athttps://github.com/Star607/STAR-TKDE. Yu Wang 0176, Tongya Zheng, Shunyu Liu 0001, Zunlei Feng, Kai-Xuan Chen 0001, Yunzhi Hao, Mingli Song |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2023 | Contrastive Identity-Aware Learning for Multi-Agent Value DecompositionabstractValue Decomposition (VD) aims to deduce the contributions of agents for decentralized policies in the presence of only global rewards, and has recently emerged as a powerful credit assignment paradigm for tackling cooperative Multi-Agent Reinforcement Learning (MARL) problems. One of the main challenges in VD is to promote diverse behaviors among agents, while existing methods directly encourage the diversity of learned agent networks with various strategies. However, we argue that these dedicated designs for agent networks are still limited by the indistinguishable VD network, leading to homogeneous agent behaviors and thus downgrading the cooperation capability. In this paper, we propose a novel Contrastive Identity-Aware learning (CIA) method, explicitly boosting the credit-level distinguishability of the VD network to break the bottleneck of multi-agent diversity. Specifically, our approach leverages contrastive learning to maximize the mutual information between the temporal credits and identity representations of different agents, encouraging the full expressiveness of credit assignment and further the emergence of individualities. The algorithm implementation of the proposed CIA module is simple yet effective that can be readily incorporated into various VD architectures. Experiments on the SMAC benchmarks and across different VD backbones demonstrate that the proposed method yields results superior to the state-of-the-art counterparts. Our code is available at https://github.com/liushunyu/CIA. Shunyu Liu 0001, Yihe Zhou, Jie Song 0011, Tongya Zheng, Kai-Xuan Chen 0001, Tongtian Zhu, Zunlei Feng, Mingli Song |
AAAI | 5 |
| 2023 | Adversarial Erasing with Pruned Elements: Towards Better Graph Lottery TicketsabstractGraph Lottery Ticket (GLT), a combination of core subgraph and sparse subnetwork, has been proposed to mitigate the computational cost of deep Graph Neural Networks (GNNs) on large input graphs while preserving original performance. However, the winning GLTs in exisiting studies are obtained by applying iterative magnitude-based pruning (IMP) without re-evaluating and re-considering the pruned information, which disregards the dynamic changes in the significance of edges/weights during graph/model structure pruning, and thus limits the appeal of the winning tickets. In this paper, we formulate a conjecture, i.e., existing overlooked valuable information in the pruned graph connections and model parameters which can be re-grouped into GLT to enhance the final performance. Specifically, we propose an adversarial complementary erasing (ACE) framework to explore the valuable information from the pruned components, thereby developing a more powerful GLT, referred to as the ACE-GLT. The main idea is to mine valuable information from pruned edges/weights after each round of IMP, and employ the ACE technique to refine the GLT processing. Finally, experimental results demonstrate that our ACE-GLT outperforms existing methods for searching GLT in diverse tasks. Our code is available at https://github.com/Wangyuwen0627/ACE-GLT. Shunyu Liu 0001, Kai-Xuan Chen 0001, Tongtian Zhu, Ji Qiao, Mengjie Shi, Yuanyu Wan, Mingli Song |
ECAI | 3 |
| 2023 | Schema Inference for Interpretable Image Classification
Haofei Zhang, Mengqi Xue, Kai-Xuan Chen 0001, Jie Song 0011, Mingli Song |
ICLR | 4 |
| 2023 | Decentralized SGD and Average-direction SAM are Asymptotically EquivalentabstractDecentralized stochastic gradient descent (D-SGD) allows collaborative learning on massive devices simultaneously without the control of a central server. However, existing theories claim that decentralization invariably undermines generalization. In this paper, we challenge the conventional belief and present a completely new perspective for understanding decentralized learning. We prove that D-SGD implicitly minimizes the loss function of an average-direction Sharpness-aware minimization (SAM) algorithm under general non-convex non-$\beta$-smooth settings. This surprising asymptotic equivalence reveals an intrinsic regularization-optimization trade-off and three advantages of decentralization: (1) there exists a free uncertainty evaluation mechanism in D-SGD to improve posterior estimation; (2) D-SGD exhibits a gradient smoothing effect; and (3) the sharpness regularization effect of D-SGD does not decrease as total batch size increases, which justifies the potential generalization benefit of D-SGD over centralized SGD (C-SGD) in large-batch scenarios. Tongtian Zhu, Fengxiang He, Kai-Xuan Chen 0001, Mingli Song, Dacheng Tao |
ICML | 3 |
| 2023 | Improving Expressivity of GNNs with Subgraph-specific Factor Embedded NormalizationabstractGraph Neural Networks~(GNNs) have emerged as a powerful category of learning architecture for handling graph-structured data. However, existing GNNs typically ignore crucial structural characteristics in node-induced subgraphs, which thus limits their expressiveness for various downstream tasks. In this paper, we strive to strengthen the representative capabilities of GNNs by devising a dedicated plug-and-play normalization scheme, termed as SUbgraph-sPEcific FactoR Embedded Normalization (SuperNorm), that explicitly considers the intra-connection information within each node-induced subgraph. To this end, we embed the subgraph-specific factor at the beginning and the end of the standard BatchNorm, as well as incorporate graph instance-specific statistics for improved distinguishable capabilities. In the meantime, we provide theoretical analysis to support that, with the elaborated SuperNorm, an arbitrary GNN is at least as powerful as the 1-WL test in distinguishing non-isomorphism graphs. Furthermore, the proposed SuperNorm scheme is also demonstrated to alleviate the over-smoothing phenomenon. Experimental results related to predictions of graph, node, and link properties on the eight popular datasets demonstrate the effectiveness of the proposed method. The code is available at https://github.com/chenchkx/SuperNorm. Kai-Xuan Chen 0001, Shunyu Liu 0001, Tongtian Zhu, Ji Qiao, Yingjie Tian 0002, Tongya Zheng, Haofei Zhang, Zunlei Feng, Jingwen Ye, Mingli Song |
KDD | 1 |
| 2023 | Distribution Knowledge Embedding for Graph PoolingabstractGraph-level representation learning is the pivotal step for downstream tasks that operate on the whole graph. The most common approach to this problem is graph pooling, where node features are typically averaged or summed to obtain the graph representations. However, pooling operations like averaging or summing inevitably cause severe information missing, which may severely downgrade the final performance. In this paper, we argue what is crucial to graph-level downstream tasks includes not only the topological structure but also thedistributionfrom which nodes are sampled. Therefore, powered by existing Graph Neural Networks (GNN), we propose a new plug-and-play pooling module, termed asDistribution Knowledge Embedding(DKEPool), where graphs are viewed as distributions on top of GNNs and the pooling goal is to summarize the entire distribution information instead of retaining a certain feature vector by simple predefined pooling operations. A DKEPool networkde factodisassembles representation learning into two stages,structure learninganddistribution learning. Structure learning follows a recursive neighborhood aggregation scheme to update node features where structure information is obtained. Distribution learning, on the other hand, omits node interconnections and focuses more on the distribution depicted by all the nodes. Extensive experiments on graph classification tasks demonstrate that the proposed DKEPool significantly and consistently outperforms the state-of-the-art methods. The code is avaliable athttps://github.com/chenchkx/dkepool Kai-Xuan Chen 0001, Jie Song 0011, Shunyu Liu 0001, Na Yu 0001, Zunlei Feng, Gengshi Han, Mingli Song |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2023 | Ask-AC: An Initiative Advisor-in-the-Loop Actor-Critic FrameworkabstractDespite the promising results achieved, state-of-the-art interactive reinforcement learning schemes rely on passively receiving supervision signals from advisor experts, in the form of either continuous monitoring or predefined rules, which inevitably result in a cumbersome and expensive learning process. In this article, we introduce a novel initiative advisor-in-the-loop actor–critic (AC) framework, termed as Ask-AC, that replaces the unilateral advisor-guidance mechanism with a bidirectional learner-initiative one, and thereby enables a customized and efficacious message exchange between learner and advisor. At the heart of Ask-AC are two complementary components, namely, action requester and adaptive state selector, that can be readily incorporated into various discrete AC architectures. The former component allows the agent to initiatively seek advisor intervention in the presence of uncertain states, while the latter identifies the unstable states potentially missed by the former especially when environment changes, and then learns to promote the ask action on such states. Experimental results on both stationary and nonstationary environments and across different AC backbones demonstrate that the proposed framework significantly improves the learning efficiency of the agent, and achieves the performances on par with those obtained by continuous advisor monitoring. Shunyu Liu 0001, Kai-Xuan Chen 0001, Na Yu 0001, Jie Song 0011, Zunlei Feng, Mingli Song |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2022 | Distribution-Aware Graph Representation Learning for Transient Stability Assessment of Power SystemabstractThe real-time transient stability assessment (TSA) plays a critical role in the secure operation of the power system. Although the classic numerical integration method, i.e. time-domain simulation (TDS), has been widely used in industry practice, it is inevitably trapped in a high computational complexity due to the high latitude sophistication of the power system. In this work, a data-driven power system estimation method is proposed to quickly predict the stability of the power system before TDS reaches the end of simulating time windows, which can reduce the average simulation time of stability assessment without loss of accuracy. As the topology of the power system is in the form of graph structure, graph neural network based representation learning is naturally suitable for learning the status of the power system. Motivated by observing the distribution information of crucial active power and reactive power on the power system's bus nodes, we thus propose a distribution-aware learning (DAL) module to explore an informative graph representation vector for describing the status of a power system. Then, TSA is re-defined as a binary classification task, and the stability of the system is determined directly from the resulting graph representation without numerical integration. Finally, we apply our method to the online TSA task. The case studies on the IEEE 39-bus system and Polish 2383-bus system demonstrate the effectiveness of our proposed method. The code is available at https://github.com/kxchern/dkepool-tsa Kai-Xuan Chen 0001, Shunyu Liu 0001, Na Yu 0001, Jie Song 0011, Zunlei Feng, Mingli Song |
IJCNN | 1 |
| 2022 | Multiple Riemannian Manifold-Valued Descriptors Based Image Set Classification With Multi-Kernel Metric LearningabstractThe importance of wild video based image set recognition is monotonically increasing due to the large amount of video data being collected by various devices including surveillance cameras, drive recorders, smart phones, and internet. The content of these videos is often complex, and it raises the question of how to perform image set modeling and feature extraction for image set-based classification. In recent years, image set classification methods have advanced considerably by modeling the image set in terms of a covariance matrix, linear subspace, or Gaussian distribution. Moreover, the distinctive geometry spanned by them include Symmetric Positive Definite (SPD) manifold, Grassmannian manifold, and Gaussian embedded Riemannian manifold, respectively. As a matter of fact, most of the approaches just adopt a single geometric model to describe each given image set, which may lose information useful for classification. To tackle this problem, we propose a novel algorithm to model each image set from a multi-geometric perspective. Specifically, the covariance matrix, linear subspace, and Gaussian distribution are applied to set representation simultaneously. In order to fuse these multiple heterogeneous Riemannian manifold-valued features, the well-equipped Riemannian kernel functions are first employed to map them into high dimensional Hilbert spaces. Then, a multi-kernel metric learning framework is devised to embed the learned hybrid kernels into a common lower dimensional subspace to facilitate classification. We conduct experiments on six widely used datasets each representing a different classification task: video-based face recognition, set-based object categorization, video-based emotion recognition, dynamic scene classification, set-based cell identification, and 3D hand pose estimation, to evaluate the classification performance of the proposed algorithm. The extensive experimental results confirm its superiority over the state-of-the-art methods. Rui Wang 0050, Xiaojun Wu 0001, Kai-Xuan Chen 0001, Josef Kittler |
IEEE Trans. Big Data | 3 |
| 2020 | Covariance descriptors on a Gaussian manifold and their application to image set classification
Kai-Xuan Chen 0001, Xiaojun Wu 0001, Josef Kittler |
Pattern Recognit. | 1 |
| 2018 | Riemannian kernel based Nyström method for approximate infinite-dimensional covariance descriptors with application to image set classificationabstractIn the domain of pattern recognition, using the CovDs (Covariance Descriptors) to represent data and taking the metrics of the resulting Riemannian manifold into account have been widely adopted for the task of image set classification. Recently, it has been proven that infinite-dimensional CovDs are more discriminative than their low-dimensional counterparts. However, the form of infinite-dimensional CovDs is implicit and the computational load is high. We propose a novel framework for representing image sets by approximating infinite-dimensional CovDs in the paradigm of the Nyström method based on a Riemannian kernel. We start by modeling the images via CovDs, which lie on the Riemannian manifold spanned by SPD (Symmetric Positive Definite) matrices. We then extend the Nyström method to the SPD manifold and obtain the approximations of CovDs in RKHS (Reproducing Kernel Hilbert Space). Finally, we approximate infinite-dimensional CovDs via these approximations. Empirically, we apply our framework to the task of image set classification. The experimental results obtained on three benchmark datasets show that our proposed approximate infinite-dimensional CovDs outperform the original CovDs. Kai-Xuan Chen 0001, Xiaojun Wu 0001, Rui Wang 0050, Josef Kittler |
ICPR | 1 |
| 2018 | Multiple Manifolds Metric Learning with Application to Image Set ClassificationabstractIn image set classification, a considerable advance has been made by modeling the original image sets by second order statistics or linear subspace, which typically lie on the Riemannian manifold. Specifically, they are Symmetric Positive Definite (SPD) manifold and Grassmann manifold respectively, and some algorithms have been developed on them for classification tasks. Motivated by the inability of existing methods to extract discriminatory features for data on Riemannian manifolds, we propose a novel algorithm which combines multiple manifolds as the features of the original image sets. In order to fuse these manifolds, the well-studied Riemannian kernels have been utilized to map the original Riemannian spaces into high dimensional Hilbert spaces. A metric Learning method has been devised to embed these kernel spaces into a lower dimensional common subspace for classification. The state-of-the-art results achieved on three datasets corresponding to two different classification tasks, namely face recognition and object categorization, demonstrate the effectiveness of the proposed method. Rui Wang 0050, Xiaojun Wu 0001, Kai-Xuan Chen 0001, Josef Kittler |
ICPR | 3 |
| 2018 | Component SPD matrices: A low-dimensional discriminative data descriptor for image set classificationabstractIn pattern recognition, the task of image set classification has often been performed by representing data using symmetric positive definite (SPD) matrices, in conjunction with the metric of the resulting Riemannian manifold. In this paper, we propose a new data representation framework for image sets which we call component symmetric positive definite representation (CSPD). Firstly, we obtain sub-image sets by dividing the images in the set into square blocks of the same size, and use a traditional SPD model to describe them. Then, we use the Riemannian kernel to determine similarities of corresponding subimage sets. Finally, the CSPD matrix appears in the form of the kernel matrix for all the sub-image sets; its i, j-th entry measures the similarity between the i-th and j-th sub-image sets. The Riemannian kernel is shown to satisfy Mercer’s theorem, so the CSPD matrix is symmetric and positive definite, and also lies on a Riemannian manifold. Test on three benchmark datasets shows that CSPD is both lower-dimensional and more discriminative data descriptor than standard SPD for the task of image set classification. Kai-Xuan Chen 0001, Xiaojun Wu 0001 |
Comput. Vis. Media | 1 |