EDBT 2026 Demo / reviewers in the wild / expert
Yuefan Deng
dblp:79/2422
· DBLP profile ↗
24ranked-venue papers
4as first author
7since 2021 · last 2026
0000-0002-5224-3958ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 9 · 4 since 2021Databases, data management, data science and information retrieval · 5Software engineering, systems software and programming languages · 2Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A new broadcast model for several network topologies
Hongbo Lu, Junsung Hwang, Bernard Tenreiro, Nabila Jaman Tripti, Darren Hamilton, Yuefan Deng |
J. Supercomput. | 6 |
| 2025 | Discretization-invariance? On the Discretization Mismatch Errors in Neural OperatorsabstractIn recent years, neural operators have emerged as a prominent approach for learning mappings between function spaces, such as the solution operators of parametric PDEs. A notable example is the Fourier Neural Operator (FNO), which models the integral kernel as a convolution operator and uses the Convolution Theorem to learn the kernel directly in the frequency domain. The parameters are decoupled from the resolution of the data, allowing the FNO to take inputs of different resolutions.
However, training at a lower resolution and inferring at a finer resolution does not guarantee consistent performance, nor can fine details, present only in fine-scale data, be learned solely from coarse data. In this work, we address this misconception by defining and examining the discretization mismatch error: the discrepancy between the outputs of the neural operator when using different discretizations of the input data. We demonstrate that neural operators may suffer from discretization mismatch errors that hinder their effectiveness when inferred on data with resolutions different from that of the training data or when trained on data with varying resolutions. As neural operators underpin many critical cross-resolution scientific tasks, such as climate modeling and fluid dynamics, understanding discretization mismatch errors is essential. Based on our findings, we propose a Cross-Resolution Operator-learning Pipeline that is free of aliasing and discretization mismatch errors, enabling efficient cross-resolution and multi-spatial-scale learning, and resulting in superior performance. Wenhan Gao 0002, Ruichen Xu, Yuefan Deng, Yi Liu 0059 |
ICLR | 3 |
| 2025 | Kolmogorov-Arnold Representation for Symplectic Learning: Advancing Hamiltonian Neural NetworksabstractWe propose a Kolmogorov–Arnold Representation-based Hamiltonian Neural Network (KAR-HNN) that replaces the Multilayer Perceptrons (MLPs) with univariate transformations. While Hamiltonian Neural Networks (HNNs) ensure energy conservation by learning Hamiltonian functions directly from data, existing implementations, often relying on MLPs, cause hypersensitivity to the hyperparameters while exploring complex energy landscapes. Our approach exploits the localized function approximations to better capture high-frequency and multi-scale dynamics, reducing energy drift and improving long-term predictive stability. The networks preserve the symplectic form of Hamiltonian systems, and, thus, maintain interpretability and physical consistency. After assessing KAR-HNN on four benchmark problems including spin-mass, simple pendulum, two-and three-body problem, we foresee its effectiveness for accurate and stable modeling of realistic physical processes often at high dimensions and with few known parameters. Ruichen Xu, Luoyao Chen, Georgios Kementzidis, Siyao Wang, Yuefan Deng |
IJCNN | 6 |
| 2024 | Exploring Robust Features for Improving Adversarial RobustnessabstractWhile deep neural networks (DNNs) have revolutionized many fields, their fragility to carefully designed adversarial attacks impedes the usage of DNNs in safety-critical applications. In this article, we strive to explore the robust features that are not affected by the adversarial perturbations, that is, invariant to the clean image and its adversarial examples (AEs), to improve the model's adversarial robustness. Specifically, we propose a feature disentanglement model to segregate the robust features from nonrobust features and domain-specific features. The extensive experiments on five widely used datasets with different attacks demonstrate that robust features obtained from our model improve the model's adversarial robustness compared to the state-of-the-art approaches. Moreover, the trained domain discriminator is able to identify the domain-specific features from the clean images and AEs almost perfectly. This enables AE detection without incurring additional computational costs. With that, we can also specify different classifiers for clean images and AEs, thereby avoiding any drop in clean image accuracy. Hong Wang 0024, Yuefan Deng, Shinjae Yoo, Yuewei Lin |
IEEE Trans. Cybern. | 2 |
| 2022 | Scalable multiscale modeling of platelets with 100 million particles
Changnian Han, Yicong Zhu, Guojing Cong, James R. Kozloski, Chih-Chieh Yang, Leili Zhang, Yuefan Deng |
J. Supercomput. | 8 |
| 2022 | Optimal circulant graphs as low-latency network topologies
Xiaolong Huang 0003, Alexandre F. Ramos, Yuefan Deng |
J. Supercomput. | 3 |
| 2021 | AGKD-BML: Defense Against Adversarial Attack by Attention Guided Knowledge Distillation and Bi-directional Metric LearningabstractWhile deep neural networks have shown impressive performance in many tasks, they are fragile to carefully de-signed adversarial attacks. We propose a novel adversarial training-based model by Attention Guided Knowledge Distillation and Bi-directional Metric Learning (AGKD-BML). The attention knowledge is obtained from a weight-fixed model trained on a clean dataset, referred to as a teacher model, and transferred to a model that is under training on adversarial examples (AEs), referred to as a student model. In this way, the student model is able to focus on the correct region, as well as correcting the intermediate features corrupted by AEs to eventually improve the model accuracy. Moreover, to efficiently regularize the representation in feature space, we propose a bidirectional metric learning. Specifically, given a clean image, it is first attacked to its most confusing class to get the forward AE. A clean image in the most confusing class is then randomly picked and attacked back to the original class to get the backward AE. A triplet loss is then used to shorten the representation distance between original image and its AE, while enlarge that between the forward and backward AEs. We conduct extensive adversarial robustness experiments on two widely used datasets with different attacks. Our proposed AGKD-BML model consistently outperforms the state-of-the-art approaches. The code of AGKD-BML will be available at: https://github.com/hongw579/AGKD-BML. Hong Wang 0024, Yuefan Deng, Shinjae Yoo, Haibin Ling, Yuewei Lin |
ICCV | 2 |
| 2020 | Optimal low-latency network topologies for cluster performance enhancement
Yuefan Deng, Meng Guo 0004, Alexandre F. Ramos, Xiaolong Huang 0003, Weifeng Liu 0012 |
J. Supercomput. | 1 |
| 2020 | Multi-User Mobile Sequential Recommendation for Route OptimizationabstractWe enhance the mobile sequential recommendation (MSR) model and address some critical issues in existing formulations by proposing three new forms of the MSR from a multi-user perspective. The multi-user MSR (MMSR) model searches optimal routes for multiple drivers at different locations while disallowing overlapping routes to be recommended. To enrich the properties of pick-up points in the problem formulation, we additionally consider the pick-up capacity as an important feature, leading to the following two modified forms of the MMSR: MMSR-m and MMSR-d. The MMSR-m sets a maximum pick-up capacity for all urban areas, while the MMSR-d allows the pick-up capacity to vary at different locations. We develop a parallel framework based on the simulated annealing to numerically solve the MMSR problem series. Also, a push-point method is introduced to improve our algorithms further for the MMSR-m and the MMSR-d, which can handle the route optimization in more practical ways. Our results on both real-world and synthetic data confirmed the superiority of our problem formulation and solutions under more demanding practical scenarios over several published benchmarks. Keli Xiao, Zeyang Ye, Wenjun Zhou 0001, Yong Ge 0001, Yuefan Deng |
ACM Trans. Knowl. Discov. Data | 6 |
| 2019 | Applying Simulated Annealing and Parallel Computing to the Mobile Sequential RecommendationabstractWe speed up the solution of the mobile sequential recommendation (MSR) problem that requires searching optimal routes for empty taxi cabs through mining massive taxi GPS data. We develop new methods that combine parallel computing and the simulated annealing with novel global and local searches. While existing approaches usually involve costly offline algorithms and methodical pruning of the search space, our new methods provide direct real-time search for the optimal route without the offline preprocessing. Our methods significantly reduce computational time for the high dimensional MSR problems from days to seconds based on the real-world data as well as the synthetic ones. We efficiently provide solutions to MSR problems with thousands of pick-up points without offline training, compared to the published record of 25 pick-up points. Zeyang Ye, Keli Xiao, Yong Ge 0001, Yuefan Deng |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2018 | A Unified Theory of the Mobile Sequential Recommendation ProblemabstractA theory is developed to unify the original form, and its many variations, of the mobile sequential recommendation (MSR) problem. The unified theory, expressing the same MSR problem, is superior to the original form in many aspects including a more standardized form. In addition to a newly proposed expected traveling time (ETT) function to measure the quality of recommended routes, we introduce five additional improvements. Also, three essential mathematical properties of the new objective function enable the development of the methods to solve realistic MSR problems with complex conditions. The MSR solutions also support the discovered properties of the proposed objective function. The unified theory should support the long-term decision making for drivers and the traffic department in general. Zeyang Ye, Keli Xiao, Yuefan Deng |
ICDM | 3 |
| 2018 | Multi-User Mobile Sequential Recommendation: An Efficient Parallel Computing ParadigmabstractThe classic mobile sequential recommendation (MSR) problem aims to provide the optimal route to taxi drivers for minimizing the potential travel distance before they meet next passengers. However, the problem is designed from the view of a single user and may lead to overlapped recommendations and cause traffic problems. Existing approaches usually contain an offline pruning process with extremely high computational cost, given a large number of pick-up points. To this end, we formalize a new multi-user MSR (MMSR) problem that locates optimal routes for a group of drivers with different starting positions. We develop two efficient methods, PSAD and PSAD-M, for solving the MMSR problem by ganging parallel computing and simulated annealing. Our methods outperform several existing approaches, especially for high-dimensional MMSR problems, with a record-breaking performance of 180x speedup using 384 cores. Zeyang Ye, Keli Xiao, Wenjun Zhou 0001, Yong Ge 0001, Yuefan Deng |
KDD | 6 |
| 2015 | Modeling Social Attention for Stock Analysis: An Influence Propagation PerspectiveabstractWith the rapid growth of usage of social network, the patterns, the scales, and the rate of information exchange have brought profound impacts on research and practice in finance. One important topic is the stock market efficiency analysis. Traditional schemes in finance focus on identifying significant abnormal returns triggered by important events. However, those events are merely identified by regular financial announcements such as mergers, equity issuances, and financial reports. Related data-driven approaches mainly focus on developing trading strategies using social media data, while the results are usually lack of theoretical explanations. In this paper, we fill the gap between the usage of social media data and financial theories. We propose a Degree of Social Attention (DSA) framework for stock analysis based on influence propagation model. Specifically, we define the self-influence for users in a social network and the DSA for stocks. A recursive process is also designed for dynamic value updating. Furthermore, we provide two modified approaches to reduce the computational cost. Our testing results from the Chinese stock market suggest that the proposed framework effectively captures stock abnormal returns based on the related social media data, and DSA is verified to be a key factor to link social media activities to the stock market. Keli Xiao, Qi Liu 0003, Yefan Tao, Yuefan Deng |
ICDM | 5 |
| 2015 | A data-driven paradigm for mapping problems
Peng Zhang 0006, Ling Liu 0005, Yuefan Deng |
Parallel Comput. | 3 |
| 2015 | Evaluation of Various Networks Configurated by Adding Bypass or Torus LinksabstractWe present a network configuration scheme exampled by a hardwired 64-node parallel computer with interconnection topologies of 2D Torus(8 x 8), 2D iBT(8 x 8; b = 〈2〉) and 3D Torus(4 x 4 x 4) reconfigurated by software. On this platform, we evaluate the performance ratios through adding bypass links (iBT) or torus links (3D Torus) to the base 2D Torus. We benchmark the three network topologies on elementary and collective communications as well as parallel applications including NAS parallel benchmark, NAMD, high performance Linpack and HPC challenge benchmark suites. Through comparative analysis of the relative performances, we reveal and, in several cases, reaffirm that strategically adding links could greatly improve the performance in 94 percent tested application cases; the iBT network outperforms others in more than half of the tested cases; and our network configuration scheme would be an alternative, better than the classical network scheme, for constructing a parallel processing system with a variety of application patterns. Peng Zhang 0006, Yuefan Deng, Xingguo Luo |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2013 | Deadlock-Free Routing Algorithms for 6D Mesh/iBT Interconnection NetworksabstractAs an application of interlaced bypass torus (iBT) interconnection networks, a 6D mesh/iBT network has been formed by replacing the 3D torus network in a 6D mesh/torus interconnect (Tofu) with a 3D iBT network, helping further reduce latencies. However, the routing algorithms good for the torus-based systems such as Blue Gene series may deadlock for the new network. This work proposes three deadlock-free routing algorithms for iBT networks and an optimizing method. An iBT network with these routings is simulated and compared with a 3D torus and a 4D torus network. Results for all-to-all communications, when applying the optimal routing algorithm, show a link utilization of 96% of the theoretical peak. An iBT prototyped system is also built and preliminarily tested. Peng Zhang 0006, Yuefan Deng |
SNPD | 3 |
| 2013 | Analysis of Linpack and power efficiencies of the world's TOP500 supercomputers
Yuefan Deng, Peng Zhang 0006, Reid Powell |
Parallel Comput. | 1 |
| 2012 | Simulated Performance Evaluation of a 6D Mesh/iBT InterconnectabstractA new 6D mesh/torus (Tofu) interconnect is implemented in the K Computer for achieving a performance of more than 10 petaflops. As expected, among its many interesting properties, such interconnect has large variations of latencies for communications within a node-group (intra-node-group) or between them (inter-node-group). Inspired by their works while recognizing the superior features of the recently published interlaced bypass torus (iBT) network that can help reduce such variations, we propose a new 6D mesh/iBT interconnect by replacing the 3D torus of Tofu by a 3D iBT. Our new interconnect suppresses the latencies for inter-node-group communications to dramatically improve the overall performance for the entire network. Adapting an industrial standard network simulator NS-3, we study different communication patterns of the added iBT networks under different bypass schemes. We find that the simulated average latencies of our network are 40% less than those of the Tofu interconnect for all tested patterns. Peng Zhang 0006, Yuefan Deng |
SNPD | 3 |
| 2012 | Design and Analysis of Pipelined Broadcast Algorithms for the All-Port Interlaced Bypass Torus NetworksabstractBroadcast algorithms for the interlaced bypass torus networks (iBT networks) are introduced to balance the all-port bandwidth efficiency and to avoid congestion in multidimensional cases. With these algorithms, we numerically analyze the dependencies of the broadcast efficiencies on various packet-sending patterns, bypass schemes, network sizes, and dimensionalities and then strategically tune up the configurations for minimizing the broadcast steps. Leveraging on such analysis, we compare the performance of networks with one million nodes between two cases: one with an added fixed-length bypass links and the other with an added torus dimension. A case study of iBT(10002;b = (8,32)) and Torus(1003) shows that the former improves the diameter, average node-to-node distance, rectangular and global broadcasts over the latter by approximately 80 percent. It is reaffirmed that strategically interlacing short bypass links and methodically utilizing these links is superior to adding dimensionalities to torus in achieving shorter diameter, average node-to-node distances and faster broadcasts. Peng Zhang 0006, Yuefan Deng |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2011 | Interlacing Bypass Rings to Torus Networks for More Efficient NetworksabstractWe introduce a new technique for generating more efficient networks by systematically interlacing bypass rings to torus networks (iBT networks). The resulting network can improve the original torus network by reducing the network diameter, node-to-node distances, and by increasing the bisection width without increasing wiring and other engineering complexity. We present and analyze the statement that a 3D iBT network proposed by our technique outperforms 4D torus networks of the same node degree. We found that interlacing rings of sizes 6 and 12 to all three dimensions of a torus network with meshes 30 × 30 × 36 generate the best network of all possible networks, including 4D torus and hypercube of approximately 32,000 nodes. This demonstrates that strategically interlacing bypass rings into a 3D torus network enhances the torus network more effectively than adding a fourth dimension, although we may generalize the claim. We also present a node-to-node distance formula for the iBT networks. Peng Zhang 0006, Reid Powell, Yuefan Deng |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2007 | Electrostatic force computation for bio-molecules on supercomputers with torus networks
Peter Rissland, Yuefan Deng |
Parallel Comput. | 2 |
| 2001 | The performance of a supercomputer built with commodity components
Yuefan Deng, Alex Korobka |
Parallel Comput. | 1 |
| 2001 | New trends in high performance computing
Osman Yasar, Yuefan Deng, Robert E. Tuzun, D. Saltz |
Parallel Comput. | 2 |
| 2000 | Approximate energy minimization for large Lennard-Jones clusters
Yuefan Deng, Carlos Rivera |
J. Glob. Optim. | 1 |