EDBT 2026 Demo / reviewers in the wild / expert
Xiongye Xiao
dblp:301/0208
· DBLP profile ↗
10ranked-venue papers
3as first author
10since 2021 · last 2025
0000-0002-3181-7166ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Language models and text generation · 24% Graph learning · 23% Deep learning architectures and training · 19% | |
| Interdisciplinary, comprehensive, and emerging computing
3 papers |
Computational science and engineering · 100% | |
| Theoretical computer science
2 papers |
Mathematical optimization · 100% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Electronic design automation · 100% |
Topics — the 17 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computational science and engineering › scientific machine learning
operator learning |
1.2 | 2 | 2023 | Coupled Multiwavelet Operator Learning for Coupled Differential Equations · ICLR 2023 Non-Linear Operator Approximations for Initial Value Problems · ICLR 2022 |
Natural language and speech › Language models and text generation › large language model
emergent abilities |
0.9 | 1 | 2025 | Neuron-based Multifractal Analysis of Neuron Interaction Dynamics in Large Models · ICLR 2025 |
Machine learning › Graph learning › hypergraph learning
hypergraph neural network |
0.9 | 1 | 2025 | MIHC: Multi-View Interpretable Hypergraph Neural Networks with Information Bottleneck for Chip Congestion Prediction · NeurIPS 2025 |
Natural language and speech › Language models and text generation
large language model |
0.9 | 1 | 2025 | Neuron-based Multifractal Analysis of Neuron Interaction Dynamics in Large Models · ICLR 2025 |
Robotics › Motion planning and robot control › robot control
optimal control |
0.9 | 1 | 2025 | End-to-End Learning Framework for Solving Non-Markovian Optimal Control · ICML 2025 |
Electronic design automation › physical design › routing
congestion prediction |
0.9 | 1 | 2025 | MIHC: Multi-View Interpretable Hypergraph Neural Networks with Information Bottleneck for Chip Congestion Prediction · NeurIPS 2025 |
Electronic design automation
physical design |
0.9 | 1 | 2025 | MIHC: Multi-View Interpretable Hypergraph Neural Networks with Information Bottleneck for Chip Congestion Prediction · NeurIPS 2025 |
Mathematical optimization › control theory › optimal control
linear quadratic regulator |
0.9 | 1 | 2025 | End-to-End Learning Framework for Solving Non-Markovian Optimal Control · ICML 2025 |
Mathematical optimization › control theory
optimal control |
0.9 | 1 | 2025 | End-to-End Learning Framework for Solving Non-Markovian Optimal Control · ICML 2025 |
Machine learning › Efficient and distributed learning › distributed training › distributed training systems
device placement |
0.8 | 1 | 2024 | A Structure-Aware Framework for Learning Device Placements on Computation Graphs · NeurIPS 2024 |
Machine learning › Graph learning
graph representation learning |
0.8 | 1 | 2024 | A Structure-Aware Framework for Learning Device Placements on Computation Graphs · NeurIPS 2024 |
Machine learning › Representation and self-supervised learning
information bottleneck |
0.8 | 1 | 2024 | Neuro-Inspired Information-Theoretic Hierarchical Perception for Multimodal Learning · ICLR 2024 |
Computational science and engineering
scientific machine learning |
0.7 | 1 | 2023 | Coupled Multiwavelet Operator Learning for Coupled Differential Equations · ICLR 2023 |
Computational science and engineering › differential equations
initial value problem |
0.6 | 1 | 2022 | Non-Linear Operator Approximations for Initial Value Problems · ICLR 2022 |
Machine learning › Deep learning architectures and training
neural operator |
0.5 | 1 | 2021 | Multiwavelet-based Operator Learning for Differential Equations · NeurIPS 2021 |
Computational science and engineering
partial differential equations |
0.5 | 1 | 2021 | Multiwavelet-based Operator Learning for Differential Equations · NeurIPS 2021 |
Electronic design automation › physical design › VLSI layout
layout and routing |
0.3 | 1 | 2025 | MIHC: Multi-View Interpretable Hypergraph Neural Networks with Information Bottleneck for Chip Congestion Prediction · NeurIPS 2025 |
Methods — techniques the papers use, named apart from their topics
information bottleneck · 2.5system identification · 1.7subgraph reasoning · 1.7multi-view hypergraph representation · 1.7fractional calculus · 1.7deep learning · 1.7operator approximation · 1.1network representation · 0.9multifractal analysis · 0.9multimodal fusion · 0.8graph coarsening · 0.8neural operator · 0.7multiwavelet operator learning · 0.7operator learning · 0.5multiwavelet transform · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Neuron-based Multifractal Analysis of Neuron Interaction Dynamics in Large ModelsabstractIn recent years, there has been increasing attention on the capabilities of large-scale models, particularly in handling complex tasks that small-scale models are unable to perform. Notably, large language models (LLMs) have demonstrated ``intelligent'' abilities such as complex reasoning and abstract language comprehension, reflecting cognitive-like behaviors. However, current research on emergent abilities in large models predominantly focuses on the relationship between model performance and size, leaving a significant gap in the systematic quantitative analysis of the internal structures and mechanisms driving these emergent abilities. Drawing inspiration from neuroscience research on brain network structure and self-organization, we propose (i) a general network representation of large models, (ii) a new analytical framework — *Neuron-based Multifractal Analysis (NeuroMFA)* - for structural analysis, and (iii) a novel structure-based metric as a proxy for emergent abilities of large models. By linking structural features to the capabilities of large models, *NeuroMFA* provides a quantitative framework for analyzing emergent phenomena in large models. Our experiments show that the proposed method yields a comprehensive measure of the network's evolving heterogeneity and organization, offering theoretical foundations and a new perspective for investigating emergence in large models. Xiongye Xiao, Heng Ping, Defu Cao, Yaxing Li, Yizhuo Zhou, Nikos Kanakaris, Paul Bogdan |
ICLR | 1 |
| 2025 | End-to-End Learning Framework for Solving Non-Markovian Optimal ControlabstractInteger-order calculus fails to capture the long-range dependence (LRD) and memory effects found in many complex systems. Fractional calculus addresses these gaps through fractional-order integrals and derivatives, but fractional-order dynamical systems pose substantial challenges in system identification and optimal control tasks. In this paper, we theoretically derive the optimal control via linear quadratic regulator (LQR) for fractional-order linear time-invariant (FOLTI) systems and develop an end-to-end deep learning framework based on this theoretical foundation. Our approach establishes a rigorous mathematical model, derives analytical solutions, and incorporates deep learning to achieve data-driven optimal control of FOLTI systems. Our key contributions include: (i) proposing a novel method for system identification and optimal control strategy in FOLTI systems, (ii) developing the first end-to-end data-driven learning framework, Fractional-Order Learning for Optimal Control (FOLOC), that learns control policies from observed trajectories, and (iii) deriving theoretical bounds on the sample complexity for learning accurate control policies under fractional-order dynamics. Experimental results indicate that our method accurately approximates fractional-order system behaviors without relying on Gaussian noise assumptions, pointing to promising avenues for advanced optimal control. Xiaole Zhang, Peiyu Zhang 0002, Xiongye Xiao, Vasileios Tzoumas, Vijay Gupta 0001, Paul Bogdan |
ICML | 3 |
| 2025 | MIHC: Multi-View Interpretable Hypergraph Neural Networks with Information Bottleneck for Chip Congestion PredictionabstractWith AI advancement and increasing circuit complexity, efficient chip design through Electronic Design Automation (EDA) is critical. Fast and accurate congestion prediction in chip layout and routing can significantly enhance automated design performance. Existing congestion modeling methods are limited by **(i)** ineffective processing and fusion of multi-view circuit data information, and **(ii)** insufficient reliability and interpretability in the prediction process. To address these challenges, We propose **M**ulti-view **I**nterpretable **H**ypergraph for **C**hip (**MIHC**), a trustworthy 'multi-view hypergraph neural network'-based framework that **(i)** processes both graph and image information in unified hypergraph representations, capturing topological and geometric circuit data, and **(ii)** implements a novel subgraph Information Bottleneck mechanism identifying critical congestion-correlated regions to guide predictions. This represents the first attempt to incorporate such interpretability into congestion prediction through informative graph reasoning. Experiments show our model reduces NMAE by 16.67% and 8.57% in cell-based and grid-based predictions on ISPD2015, and 5.26% and 2.44% on CircuitNet-N28, respectively, compared to state-of-the-art methods. Rigorous cross-design generalization experiments further validate our method’s capability to handle entirely unseen circuit designs. Zeyue Zhang, Heng Ping, Peiyu Zhang 0002, Nikos Kanakaris, Xiaoling Lu, Paul Bogdan, Xiongye Xiao |
NeurIPS | 7 |
| 2024 | Unlocking Deep Learning: A BP-Free Approach for Parallel Block-Wise Training of Neural NetworksabstractBackpropagation (BP) has been a successful optimization technique for deep learning models. However, its limitations, such as backward- and update-locking, and its biological implausibility, hinder the concurrent updating of layers and do not mimic the local learning processes observed in the human brain. To address these issues, recent research has suggested using local error signals to asynchronously train network blocks. However, this approach often involves extensive trial-and-error iterations to determine the best configuration for local training. This includes decisions on how to decouple network blocks and which auxiliary networks to use for each block. In our work, we introduce a novel BP-free approach: a block-wise BP-free (BWBPF) neural network that leverages local error signals to optimize distinct sub-neural networks separately, where the global loss is only responsible for updating the output layer. The local error signals used in the BP-free model can be computed in parallel, enabling a potential speed-up in the weight update process through parallel implementation. Our experimental results consistently show that this approach can identify transferable decoupled architectures for VGG and ResNet variations, outperforming models trained with end-to-end backpropagation and other state-of-the-art block-wise learning techniques on datasets such as CIFAR-10 and Tiny-ImageNet. The code is released at https://github.com/Belis0811/BWBPF. Anzhe Cheng, Heng Ping, Zhenkun Wang 0008, Xiongye Xiao, Chenzhong Yin, Shahin Nazarian, Mingxi Cheng, Paul Bogdan |
ICASSP | 4 |
| 2024 | Discovering Malicious Signatures in Software from Structural InteractionsabstractMalware represents a significant security concern in today’s digital landscape, as it can destroy or disable operating systems, steal sensitive user information, and occupy valuable disk space. However, current malware detection methods, such as static-based and dynamic-based approaches, struggle to identify newly developed ("zero-day") malware and are limited by customized virtual machine (VM) environments. To overcome these limitations, we propose a novel malware detection approach that leverages deep learning, mathematical techniques, and network science. Our approach focuses on static and dynamic analysis and utilizes the Low-Level Virtual Machine (LLVM) to profile applications within a complex network. The generated network topologies are input into the GraphSAGE architecture to efficiently distinguish between benign and malicious software applications, with the operation names denoted as node features. Importantly, the GraphSAGE models analyze the network’s topological geometry to make predictions, enabling them to detect state-of-the-art malware and prevent potential damage during execution in a VM. To evaluate our approach, we conduct a study on a dataset comprising source code from 24,376 applications, specifically written in C/C++, sourced directly from widely-recognized malware and various types of benign software. The results show a high detection performance with an Area Under the Receiver Operating Characteristic Curve (AUROC) of 99.85%. Our approach marks a substantial improvement in malware detection, providing a notably more accurate and efficient solution when compared to current state-of-the-art malware detection methods. The code is released at https://github.com/HantangZhang/MGN. Chenzhong Yin, Hantang Zhang, Mingxi Cheng, Xiongye Xiao, Xinghe Chen, Paul Bogdan |
ICASSP | 4 |
| 2024 | Neuro-Inspired Information-Theoretic Hierarchical Perception for Multimodal LearningabstractIntegrating and processing information from various sources or modalities are critical for obtaining a comprehensive and accurate perception of the real world in autonomous systems and cyber-physical systems. Drawing inspiration from neuroscience, we develop the Information-Theoretic Hierarchical Perception (ITHP) model, which utilizes the concept of information bottleneck. Different from most traditional fusion models that incorporate all modalities identically in neural networks, our model designates a prime modality and regards the remaining modalities as detectors in the information pathway, serving to distill the flow of information. Our proposed perception model focuses on constructing an effective and compact information flow by achieving a balance between the minimization of mutual information between the latent state and the input modal state, and the maximization of mutual information between the latent states and the remaining modal states. This approach leads to compact latent state representations that retain relevant information while minimizing redundancy, thereby substantially enhancing the performance of multimodal representation learning. Experimental evaluations on the MUStARD, CMU-MOSI, and CMU-MOSEI datasets demonstrate that our model consistently distills crucial information in multimodal learning scenarios, outperforming state-of-the-art benchmarks. Remarkably, on the CMU-MOSI dataset, ITHP surpasses human-level performance in the multimodal sentiment binary classification task across all evaluation metrics (i.e., Binary Accuracy, F1 Score, Mean Absolute Error, and Pearson Correlation). Xiongye Xiao, Gengshuo Liu, Defu Cao, Yaxing Li, Tianqing Fang, Mingxi Cheng, Paul Bogdan |
ICLR | 1 |
| 2024 | A Structure-Aware Framework for Learning Device Placements on Computation GraphsabstractComputation graphs are Directed Acyclic Graphs (DAGs) where the nodes correspond to mathematical operations and are used widely as abstractions in optimizations of neural networks. The device placement problem aims to identify optimal allocations of those nodes to a set of (potentially heterogeneous) devices. Existing approaches rely on two types of architectures known as grouper-placer and encoder-placer, respectively. In this work, we bridge the gap between encoder-placer and grouper-placer techniques and propose a novel framework for the task of device placement, relying on smaller computation graphs extracted from the OpenVINO toolkit. The framework consists of five steps, including graph coarsening, node representation learning and policy optimization. It facilitates end-to-end training and takes into account the DAG nature of the computation graphs. We also propose a model variant, inspired by graph parsing networks and complex network analysis, enabling graph representation learning and jointed, personalized graph partitioning, using an unspecified number of groups. To train the entire framework, we use reinforcement learning using the execution time of the placement as a reward. We demonstrate the flexibility and effectiveness of our approach through multiple experiments with three benchmark models, namely Inception-V3, ResNet, and BERT. The robustness of the proposed framework is also highlighted through an ablation study. The suggested placements improve the inference speed for the benchmark models by up to $58.2\%$ over CPU execution and by up to $60.24\%$ compared to other commonly used baselines. Shukai Duan 0002, Heng Ping, Nikos Kanakaris, Xiongye Xiao, Panagiotis Kyriakis, Nesreen K. Ahmed, Peiyu Zhang 0002, Guixiang Ma, Mihai Capota, Shahin Nazarian, Theodore L. Willke, Paul Bogdan |
NeurIPS | 4 |
| 2023 | Coupled Multiwavelet Operator Learning for Coupled Differential Equations
Xiongye Xiao, Defu Cao, Ruochen Yang, Gengshuo Liu, Chenzhong Yin, Radu Balan, Paul Bogdan |
ICLR | 1 |
| 2022 | Non-Linear Operator Approximations for Initial Value Problems
Xiongye Xiao, Radu Balan, Paul Bogdan |
ICLR | 2 |
| 2021 | Multiwavelet-based Operator Learning for Differential EquationsabstractThe solution of a partial differential equation can be obtained by computing the inverse operator map between the input and the solution space. Towards this end, we introduce a $\textit{multiwavelet-based neural operator learning scheme}$ that compresses the associated operator's kernel using fine-grained wavelets. By explicitly embedding the inverse multiwavelet filters, we learn the projection of the kernel onto fixed multiwavelet polynomial bases. The projected kernel is trained at multiple scales derived from using repeated computation of multiwavelet transform. This allows learning the complex dependencies at various scales and results in a resolution-independent scheme. Compare to the prior works, we exploit the fundamental properties of the operator's kernel which enable numerically efficient representation. We perform experiments on the Korteweg-de Vries (KdV) equation, Burgers' equation, Darcy Flow, and Navier-Stokes equation. Compared with the existing neural operator approaches, our model shows significantly higher accuracy and achieves state-of-the-art in a range of datasets. For the time-varying equations, the proposed method exhibits a ($2X-10X$) improvement ($0.0018$ ($0.0033$) relative $L2$ error for Burgers' (KdV) equation). By learning the mappings between function spaces, the proposed method has the ability to find the solution of a high-resolution input after learning from lower-resolution data. Xiongye Xiao, Paul Bogdan |
NeurIPS | 2 |