EDBT 2026 Demo / reviewers in the wild / expert
Xitong Zhang
dblp:156/9687
· DBLP profile ↗
16ranked-venue papers
5as first author
11since 2021 · last 2025
0000-0002-4379-8089ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 4 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-authorSystems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Optimization for machine learning · 29% Graph learning · 24% Efficient and distributed learning · 22% | |
| Interdisciplinary, comprehensive, and emerging computing
5 papers |
Environmental and earth informatics · 60% Bioinformatics and computational biology · 25% Smart cities and intelligent transportation · 16% | |
| Network and information security
2 papers |
Blockchain and cryptocurrency security · 50% Systems and software security · 50% | |
| Computer graphics and multimedia
1 paper |
Image and video processing · 100% |
Topics — the 23 heaviest of 26, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Environmental and earth informatics › geophysics
full-waveform inversion |
1.1 | 2 | 2022 | OpenFWI: Large-scale Multi-structural Benchmark Datasets for Full Waveform Inversion · NeurIPS 2022 Unsupervised Learning of Full-Waveform Inversion: Connecting CNN and Partial Differential Equation in a Loop · ICLR 2022 |
Environmental and earth informatics
geophysics |
1.1 | 2 | 2022 | OpenFWI: Large-scale Multi-structural Benchmark Datasets for Full Waveform Inversion · NeurIPS 2022 Unsupervised Learning of Full-Waveform Inversion: Connecting CNN and Partial Differential Equation in a Loop · ICLR 2022 |
Machine learning › Efficient and distributed learning
model compression |
0.8 | 1 | 2024 | Optimal Eye Surgeon: Finding image priors through sparse generators at initialization · ICML 2024 |
Machine learning › Efficient and distributed learning › model compression
pruning |
0.8 | 1 | 2024 | Optimal Eye Surgeon: Finding image priors through sparse generators at initialization · ICML 2024 |
Bioinformatics and computational biology › single-cell analysis
single-cell transcriptomics |
0.8 | 1 | 2024 | Inferring Metabolic States from Single Cell Transcriptomic Data via Geometric Deep Learning · RECOMB 2024 |
Image and video processing › image statistics › statistical image modeling
image prior |
0.8 | 1 | 2024 | Optimal Eye Surgeon: Finding image priors through sparse generators at initialization · ICML 2024 |
Image and video processing
image restoration |
0.8 | 1 | 2024 | Optimal Eye Surgeon: Finding image priors through sparse generators at initialization · ICML 2024 |
Machine learning › Optimization for machine learning
implicit regularization |
0.7 | 1 | 2023 | Implicit regularization in Heavy-ball momentum accelerated stochastic gradient descent · ICLR 2023 |
Natural language and speech › Language models and text generation › large language model › large language model adaptation
pre-trained language model fine-tuning |
0.7 | 1 | 2023 | PAC-tuning: Fine-tuning Pre-trained Language Models with PAC-driven Perturbed Gradient Descent · EMNLP 2023 |
Machine learning › Optimization for machine learning
stochastic gradient descent |
0.7 | 1 | 2023 | Implicit regularization in Heavy-ball momentum accelerated stochastic gradient descent · ICLR 2023 |
Machine learning › Deep learning architectures and training
physics-informed neural network |
0.6 | 1 | 2022 | Unsupervised Learning of Full-Waveform Inversion: Connecting CNN and Partial Differential Equation in a Loop · ICLR 2022 |
Environmental and earth informatics › geophysics
seismic inversion |
0.6 | 1 | 2022 | OpenFWI: Large-scale Multi-structural Benchmark Datasets for Full Waveform Inversion · NeurIPS 2022 |
Machine learning › Graph learning › graph neural network
directed graph neural network |
0.5 | 1 | 2021 | MagNet: A Neural Network for Directed Graphs · NeurIPS 2021 |
Machine learning › Graph learning
graph neural network |
0.5 | 1 | 2021 | MagNet: A Neural Network for Directed Graphs · NeurIPS 2021 |
Machine learning › Graph learning › graph neural network
node classification |
0.5 | 1 | 2021 | MagNet: A Neural Network for Directed Graphs · NeurIPS 2021 |
Smart cities and intelligent transportation › traffic estimation
traffic state estimation |
0.4 | 1 | 2019 | Boosted Trajectory Calibration for Traffic State Estimation · ICDM 2019 |
Smart cities and intelligent transportation › mobility data analysis
trajectory data mining |
0.4 | 1 | 2019 | Boosted Trajectory Calibration for Traffic State Estimation · ICDM 2019 |
Bioinformatics and computational biology › sequence analysis
DNA sequence analysis |
0.2 | 1 | 2015 | PseKNC-General: a cross-platform package for generating various modes of pseudo nucleotide compositions · Bioinform. 2015 |
Bioinformatics and computational biology › sequence analysis
sequence feature extraction |
0.2 | 1 | 2015 | PseKNC-General: a cross-platform package for generating various modes of pseudo nucleotide compositions · Bioinform. 2015 |
Machine learning › Learning paradigms
supervised learning |
0.2 | 1 | 2022 | OpenFWI: Large-scale Multi-structural Benchmark Datasets for Full Waveform Inversion · NeurIPS 2022 |
Machine learning › Learning paradigms
unsupervised learning |
0.2 | 1 | 2022 | OpenFWI: Large-scale Multi-structural Benchmark Datasets for Full Waveform Inversion · NeurIPS 2022 |
Machine learning › Graph learning
link prediction |
0.1 | 1 | 2021 | MagNet: A Neural Network for Directed Graphs · NeurIPS 2021 |
Data mining
spatiotemporal data mining |
0.1 | 1 | 2019 | Boosted Trajectory Calibration for Traffic State Estimation · ICDM 2019 |
Methods — techniques the papers use, named apart from their topics
threshold estimation · 1.7deep learning · 1.7lottery ticket hypothesis · 1.5deep image prior · 1.5partial differential equations · 1.1convolutional neural network · 1.1geometric deep learning · 0.8deep neural network · 0.8boosting calibration · 0.8perturbed gradient descent · 0.7noise injection · 0.7heavy-ball momentum · 0.7PAC-Bayes training · 0.7unsupervised learning · 0.6uncertainty quantification · 0.6physics-driven inversion · 0.6autocorrelation coefficients · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Weak Supervision Multigeophysical Inversion for CO2 Saturation ImagingabstractIn CO2sequestration projects, multi-physics inversion has been widely used to reconstruct various geophysical properties (such as velocity and conductivity). However, the saturation of CO2is not feasible through partial differential equations (PDEs), which poses a challenge to traditional multi-physics inversion techniques in directly inverting CO2saturation from geophysical measurements. Typically, a rock physics model is constructed to tackle this challenge. Nevertheless, the construction of such a model is intricate due to the inherent complexities and uncertainties within subsurface geology and geophysical data. Data-driven inversion methods present an alternative solution, as they can connect CO2saturation to geophysical measurements and directly learn their relationships from labeled data. Nonetheless, the efficacy of these methods hinges on extensive data labeling, incurring considerable costs. To address this challenge, we propose a novel data-driven technique for the inversion of multi-physics data called Weakly Supervised Multi-Geophysical Inversion (WS-MGI), which reduces the need for labeling and promises to be a more cost-effective solution. In particular, we focus on the multi-physics inversion problem from two geophysical data (electromagnetic (EM) and seismic) to CO2saturation. By learning the local relationship between velocity and CO2saturation at a few well logs, we construct the pseudo labels for CO2saturation, thus enabling the data-driven inversion with few labels. We verify our method using synthetic data based on the Kimberlina storage reservoir in California. Experiments show that compared to the supervised counterpart, our method can achieve similar inversion results with only 2% labels. Shihang Feng, Yinpeng Chen, Xitong Zhang, David Alumbaugh, Michael Commer, Youzuo Lin |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | An integrated static and dynamic graph fusion approach for traffic flow prediction
Xingliang Che, Xitong Zhang |
J. Supercomput. | 4 |
| 2024 | Optimal Eye Surgeon: Finding image priors through sparse generators at initializationabstractWe introduce Optimal Eye Surgeon (OES), a framework for pruning and training deep image generator networks. Typically, untrained deep convolutional networks, which include image sampling operations, serve as effective image priors. However, they tend to overfit to noise in image restoration tasks due to being overparameterized. OES addresses this by adaptively pruning networks at random initialization to a level of underparameterization. This process effectively captures low-frequency image components even without training, by just masking. When trained to fit noisy image, these pruned subnetworks, which we term Sparse-DIP, resist overfitting to noise. This benefit arises from underparameterization and the regularization effect of masking, constraining them in the manifold of image priors. We demonstrate that subnetworks pruned through OES surpass other leading pruning methods, such as the Lottery Ticket Hypothesis, which is known to be suboptimal for image recovery tasks. Our extensive experiments demonstrate the transferability of OES-masks and the characteristics of sparse-subnetworks for image generation. Code is available at https://github.com/Avra98/Optimal-Eye-Surgeon. Avrajit Ghosh, Xitong Zhang, Kenneth K. Sun, Qing Qu 0001, Saiprasad Ravishankar |
ICML | 2 |
| 2024 | Inferring Metabolic States from Single Cell Transcriptomic Data via Geometric Deep Learning
Holly R. Steach, Siddharth Viswanath, Yixuan He 0001, Xitong Zhang, Natalia Ivanova, Matthew J. Hirn, Michael Perlmutter, Smita Krishnaswamy |
RECOMB | 4 |
| 2023 | PAC-tuning: Fine-tuning Pre-trained Language Models with PAC-driven Perturbed Gradient DescentabstractFine-tuning pretrained language models (PLMs) for downstream tasks is a large-scale optimization problem, in which the choice of the training algorithm critically determines how well the trained model can generalize to unseen test data, especially in the context of few-shot learning.To achieve good generalization performance and avoid overfitting, techniques such as data augmentation and pruning are often applied.However, adding these regularizations necessitates heavy tuning of the hyperparameters of optimization algorithms, such as the popular Adam optimizer.In this paper, we propose a two-stage fine-tuning method, PAC-tuning, to address this optimization challenge.First, based on PAC-Bayes training, PAC-tuning directly minimizes the PAC-Bayes generalization bound to learn proper parameter distribution.Second, PAC-tuning modifies the gradient by injecting noise with the variance learned in the first stage into the model parameters during training, resulting in a variant of perturbed gradient descent (PGD).In the past, the fewshot scenario posed difficulties for PAC-Bayes training because the PAC-Bayes bound, when applied to large models with limited training data, might not be stringent.Our experimental results across 5 GLUE benchmark tasks demonstrate that PAC-tuning 1 successfully handles the challenges of fine-tuning tasks and outperforms strong baseline methods by a visible margin, further confirming the potential to apply PAC training for any other settings where the Adam optimizer is currently used for training. Zhiyu Xue, Xitong Zhang, Kristen Marie Johnson |
EMNLP | 3 |
| 2023 | Implicit regularization in Heavy-ball momentum accelerated stochastic gradient descent
Avrajit Ghosh, He Lyu, Xitong Zhang |
ICLR | 3 |
| 2022 | Unsupervised Learning of Full-Waveform Inversion: Connecting CNN and Partial Differential Equation in a Loop
Xitong Zhang, Yinpeng Chen, Sharon X. Huang, Zicheng Liu 0001, Youzuo Lin |
ICLR | 2 |
| 2022 | OpenFWI: Large-scale Multi-structural Benchmark Datasets for Full Waveform InversionabstractFull waveform inversion (FWI) is widely used in geophysics to reconstruct high-resolution velocity maps from seismic data. The recent success of data-driven FWI methods results in a rapidly increasing demand for open datasets to serve the geophysics community. We present OpenFWI, a collection of large-scale multi-structural benchmark datasets, to facilitate diversified, rigorous, and reproducible research on FWI. In particular, OpenFWI consists of $12$ datasets ($2.1$TB in total) synthesized from multiple sources. It encompasses diverse domains in geophysics (interface, fault, CO$_2$ reservoir, etc.), covers different geological subsurface structures (flat, curve, etc.), and contain various amounts of data samples (2K - 67K). It also includes a dataset for 3D FWI. Moreover, we use OpenFWI to perform benchmarking over four deep learning methods, covering both supervised and unsupervised learning regimes. Along with the benchmarks, we implement additional experiments, including physics-driven methods, complexity analysis, generalization study, uncertainty quantification, and so on, to sharpen our understanding of datasets and methods. The studies either provide valuable insights into the datasets and the performance, or uncover their current limitations. We hope OpenFWI supports prospective research on FWI and inspires future open-source efforts on AI for science. All datasets and related information can be accessed through our website at https://openfwi-lanl.github.io/ Chengyuan Deng, Shihang Feng, Hanchen Wang 0003, Xitong Zhang, Yinan Feng, Qili Zeng, Yinpeng Chen, Youzuo Lin |
NeurIPS | 4 |
| 2022 | Connect the Dots: In Situ 4-D Seismic Monitoring of CO2 Storage With Spatio-Temporal CNNsabstract4-D seismic imaging has been widely used in CO2sequestration projects to monitor the fluid flow in the volumetric subsurface region that is not sampled by wells. Ideally, real-time monitoring and near-future forecasting would provide site operators with great insights to understand the dynamics of the subsurface reservoir and assess any potential risks. However, due to obstacles such as high deployment cost, availability of acquisition equipment, exclusion zones around surface structures, only very sparse seismic imaging data can be obtained during monitoring. That leads to an unavoidable and growing knowledge gap over time. The operator needs to understand the fluid flow throughout the project lifetime and the seismic data are only available at a limited number of times. This is insufficient for understanding reservoir behavior. To overcome those challenges, we have developed spatio-temporal neural-network-based models that can produce high-fidelity interpolated or extrapolated images effectively and efficiently. Specifically, our models are built on an autoencoder, and incorporate the long short-term memory (LSTM) structure with a new loss function regularized by optical flow. We validate the performance of our models using real 4-D post-stack seismic imaging data acquired at the Sleipner CO2sequestration field. We employ two different strategies in evaluating our models. Numerically, we compare our models with different baseline approaches using classic pixel-based metrics. We also conduct a blind survey and collect a total of 20 responses from domain experts to evaluate the quality of data generated by our models. Via both numerical and expert evaluation, we conclude that our models can produce high-quality 2-D/3-D seismic imaging data at a reasonable cost, offering the possibility of real-time monitoring or even near-future forecasting of the CO2storage reservoir. Shihang Feng, Xitong Zhang, Brendt Wohlberg, Neill Symons, Youzuo Lin |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Making Invisible Visible: Data-Driven Seismic Inversion With Spatio-Temporally Constrained Data AugmentationabstractDeep learning and data-driven approaches have shown great potential in scientific domains. The promise of data-driven techniques relies on the availability of a large volume of high-quality training datasets. Due to the high cost of obtaining data through expensive physical experiments, instruments, and simulations, data augmentation techniques for scientific applications have emerged as a new direction for obtaining scientific data recently. However, existing data augmentation techniques originating from computer vision yield physically unacceptable data samples that are not helpful for the domain problems that we are interested in. In this article, we develop new data augmentation techniques based on convolutional neural networks. Specifically, our generative models leverage different physics knowledge (such as governing equations, observable perception, and physics phenomena) to improve the quality of the synthetic data. To validate the effectiveness of our data augmentation techniques, we apply them to solve a subsurface seismic full-waveform inversion using simulated CO2leakage data. Our interest is to invert for subsurface velocity models associated with very small CO2leakage. We validate the performance of our methods using comprehensive numerical tests. Via comparison and analysis, we show that data-driven seismic imaging can be significantly enhanced by using our data augmentation techniques. Particularly, the imaging quality has been improved by 15% in test scenarios of general-sized leakage and 17% in small-sized leakage when using an augmented training set obtained with our techniques. Yuxin Yang 0001, Xitong Zhang, Qiang Guan, Youzuo Lin |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | MagNet: A Neural Network for Directed GraphsabstractThe prevalence of graph-based data has spurred the rapid development of graph neural networks (GNNs) and related machine learning algorithms. Yet, despite the many datasets naturally modeled as directed graphs, including citation, website, and traffic networks, the vast majority of this research focuses on undirected graphs. In this paper, we propose MagNet, a GNN for directed graphs based on a complex Hermitian matrix known as the magnetic Laplacian. This matrix encodes undirected geometric structure in the magnitude of its entries and directional information in their phase. A charge parameter attunes spectral information to variation among directed cycles. We apply our network to a variety of directed graph node classification and link prediction tasks showing that MagNet performs well on all tasks and that its performance exceeds all other methods on a majority of such tasks. The underlying principles of MagNet are such that it can be adapted to other GNN architectures. Xitong Zhang, Yixuan He 0001, Nathan Brugnone, Michael Perlmutter, Matthew J. Hirn |
NeurIPS | 1 |
| 2020 | Shoreline: Data-Driven Threshold Estimation of Online Reserves of Cryptocurrency Trading Platforms (Student Abstract)abstractWith the proliferation of blockchain projects and applications, cryptocurrency exchanges, which provides exchange services among different types of cryptocurrencies, become pivotal platforms that allow customers to trade digital assets on different blockchains. Because of the anonymity and trustlessness nature of cryptocurrency, one major challenge of crypto-exchanges is asset safety, and all-time amount hacked from crypto-exchanges until 2018 is over $1.5 billion even with carefully maintained secure trading systems. The most critical vulnerability of crypto-exchanges is from the so-called hot wallet, which is used to store a certain portion of the total asset online of an exchange and programmatically sign transactions when a withdraw happens. It is important to develop network security mechanisms. However, the fact is that there is no guarantee that the system can defend all attacks. Thus, accurately controlling the available assets in the hot wallets becomes the key to minimize the risk of running an exchange. In this paper, we propose Shoreline, a deep learning-based threshold estimation framework that estimates the optimal threshold of hot wallets from historical wallet activities and dynamic trading networks. Xitong Zhang |
AAAI | 1 |
| 2020 | Shoreline: Data-Driven Threshold Estimation of Online Reserves of Cryptocurrency Trading PlatformsabstractWith the proliferation of blockchain projects and applications, cryptocurrency exchanges, which provides exchange services among different types of cryptocurrencies, become pivotal platforms that allow customers to trade digital assets on different blockchains. Because of the anonymity and trustlessness nature of cryptocurrency, one major challenge of crypto-exchanges is asset safety, and all-time amount hacked from crypto-exchanges until 2018 is over $1.5 billion even with carefully maintained secure trading systems. The most critical vulnerability of crypto-exchanges is from the so-called hot wallet, which is used to store a certain portion of the total asset of an exchange and programmatically sign transactions when a withdraw happens. Whenever hackers managed to gain control over the computing infrastructure of the exchange, they usually immediately obtain all the assets in the hot wallet. It is important to develop network security mechanisms. However, the fact is that there is no guarantee that the system can defend all attacks. Thus, accurately controlling the available assets in the hot wallets becomes the key to minimize the risk of running an exchange. However, determining such optimal threshold remains a challenging task because of the complicated dynamics inside exchanges. In this paper, we propose Shoreline, a deep learning-based threshold estimation framework that estimates the optimal threshold of hot wallets from historical wallet activities and dynamic trading networks. We conduct extensive empirical studies on the real trading data from a trading platform and demonstrate the effectiveness of the proposed approach. Xitong Zhang |
AAAI | 1 |
| 2019 | Boosted Trajectory Calibration for Traffic State EstimationabstractTraffic state estimation is among the most critical issues in intelligent transport systems, and it has been widely explored for years considering its significance for diverse industrial applications such as vehicle dispatch and route planning. The availability of vehicle trajectories has provided a promising data source for traffic estimation, as it shed light on traffic status and characterizes human transition activates that are otherwise extremely hard to track. However, there are major challenges in incorporating trajectories in traffic estimation largely because of its inherent uncertainty from data sparsity and semantic ambiguity. Entangled in complicated road topology and spatiotemporal traffic dynamics, properly embedding trajectories in modeling is especially difficult. To address the aforementioned challenges, in this paper, we propose a Boosted Trajectory Calibration (BTRAC) framework to model traffic states of complex road networks that effectively integrates trajectory information. In the framework, we first use a deep neural network to predict the traffic states of each road individually from historical traffic information, along with the prediction uncertainty. Then we refine the predictions by an iterative boosting calibration procedure with embedded trajectories. We conduct extensive large-scale empirical studies on two cities and demonstrate the effectiveness of the proposed approach. Xitong Zhang, Liyang Xie |
ICDM | 1 |
| 2016 | Streetlamp extraction and identification from mobile LiDAR point cloud scenesabstractTo solve the problem of streetlamp extraction and category identification from mobile LiDAR data in urban scenes, an algorithm based on streetlamp sample has been raised. Firstly, based on morphological closing operation to extract the location of suspected streetlamps; secondly, determine the scope of suspected streetlamps by the streetlamp sample parameters and extract the suspected streetlamp point cloud; thirdly, through pole and head matching to realize the alignment of streetlamp samples and suspected streetlamp point cloud; finally, realize the streetlamp judgment and extraction by establishing sample buffer. Experimental results show that the algorithm cannot only automatically extract streetlamp point cloud, but also be able to identify streetlamp category which is in accordance with the sample, achieving the rapid streetlamp extraction and recognition. Xitong Zhang, Huiyun Liu, Zhenzhen Wu, Jie Mao |
IGARSS | 1 |
| 2015 | PseKNC-General: a cross-platform package for generating various modes of pseudo nucleotide compositionsabstractSUMMARY: The avalanche of genomic sequences generated in the post-genomic age requires efficient computational methods for rapidly and accurately identifying biological features from sequence information. Towards this goal, we developed a freely available and open-source package, called PseKNC-General (the general form of pseudo k-tuple nucleotide composition), that allows for fast and accurate computation of all the widely used nucleotide structural and physicochemical properties of both DNA and RNA sequences. PseKNC-General can generate several modes of pseudo nucleotide compositions, including conventional k-tuple nucleotide compositions, Moreau-Broto autocorrelation coefficient, Moran autocorrelation coefficient, Geary autocorrelation coefficient, Type I PseKNC and Type II PseKNC. In every mode, >100 physicochemical properties are available for choosing. Moreover, it is flexible enough to allow the users to calculate PseKNC with user-defined properties. The package can be run on Linux, Mac and Windows systems and also provides a graphical user interface. AVAILABILITY AND IMPLEMENTATION: The package is freely available at: http://lin.uestc.edu.cn/server/pseknc. Wei Chen 0064, Xitong Zhang, Jordan Brooker, Hao Lin 0001, Liqing Zhang 0002, Kuo-Chen Chou |
Bioinform. | 2 |