VLDB 2026 Research / reviewers in the wild / expert
Xiaohan Zhao
dblp:75/781
· DBLP profile ↗
35ranked-venue papers
8as first author
22since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 2 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 6 since 2021Computer networks · 6 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 6 · 3 since 2021Software engineering, systems software and programming languages · 4 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-authorHuman-computer interaction and ubiquitous computing · 3 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LLMSurgeon: Diagnosing Data Mixture of Large Language ModelsabstractYaxin Luo, Jiacheng Cui, Xiaohan Zhao, Xinyi Shang, Jiacheng Liu, Xinyue Bi, Zhaoyi Li, Zhiqiang Shen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yaxin Luo, Jiacheng Cui, Xiaohan Zhao, Xinyi Shang, Xinyue Bi |
ACL (1) | 3 |
| 2026 | ZPRRF: Zero-Shot Prior and Rule-Guided Radiology Reporting Framework
Xiaohan Zhao, Kunli Zhang |
ICIC (29) | 4 |
| 2026 | Raw event-based adversarial attacks for Spiking Neural Networks with configurable latencies
Wanli Shi, Xiaohan Zhao, Tieru Wu |
Neural Networks | 3 |
| 2026 | Automatically Deriving Developers' Technical Expertise from the GitHub Social NetworkabstractDevelopers’ technical expertise is crucial for numerous tasks within open-source communities, such as identifying suitable developers and maintainers. Despite its significance, GitHub, the world’s largest open-source code hosting platform, does not explicitly display developers’ technical expertise. Existing methods fall short in capturing the multi-faceted and dynamic nature of developers’ skills and knowledge. To address this gap, we propose a novel approach that leverages graph neural networks (GNNs) to express developers’ technical expertise. Our method constructs a comprehensive GitHub social network that integrates various social and development activities. We then employ a GNN model to learn a low-dimensional representation vector for each developer, encapsulating their technical expertise across different dimensions. We assess the effectiveness of our model by comparing it against five baselines on three GitHub social relationship recommendation tasks, including SimDeveloper, ContributionRepo, and RepoMaintainer. Our proposed method outperforms these baselines, achieving improvements of 5.6–9.5% on Hit Ratio@10 and 3.4–11.1% on F1 score. These results demonstrate promising performance in predicting technical preferences for both repositories and developers. This research contributes to a more nuanced understanding of developer expertise in open-source communities and has potential implications for improving collaboration and project management on platforms like GitHub. Yanchun Sun, Xiaohan Zhao, Haizhou Xu, Ye Zhu 0002, Zhenpeng Chen 0001, Huizhen Jiang, Gang Huang 0001 |
ACM Trans. Softw. Eng. Methodol. | 3 |
| 2025 | Leveraging BERT and Large Language Models for Mapping Heterogeneous Scientific and Technological Resources to Their IdentifiersabstractCurrently, there are various scientific and technological resource retrieval databases in the world. The resources stored in these databases may be identified by different identification systems. How to determine whether scientific and technological resources identified by different identification systems are the same resource is an urgent problem to be solved. This paper proposes a software service that leverages BERT and large language models to perform semantic analysis and similarity matching of scientific and technological resource content, and then maps the resources to their respective identifiers. The service effectively solves the problem of how to quickly retrieve the same resource from a large number of scientific and technological resources with diverse identification types, and improves the efficiency and quality of the retrieval. Implemented as a Chrome plugin, the service facilitates seamless mapping heterogeneous scientific and technological resources to their identifiers. We conduct a series of experiments. Their results demonstrate the effectiveness, scalability and stability of the service. To the best of our knowledge, we are the first to propose the service integrating BERT with large language models to extract and identify the key content of scientific and technological resources from web pages. Yanchun Sun, Xiaohan Zhao, Huizhen Jiang, Huaqian Cai, Changfa Lu, Gang Huang 0001 |
SSE | 3 |
| 2025 | Design of Multiple Binary Waveforms for the Joint MIMO Radar and CommunicationsabstractIn this paper, we focus on the multi-waveform design for the joint radar and communications, wherein the elements of waveforms are required to take binary values and to support both good integrated sidelobe level (ISL) and information embedding (IE) performances simultaneously. Since the binary waveform attribute enables limited degrees of freedom for the design, we partially choose to exploit both the phase and index modulations (PIM) of waveform elements to embed communication symbols in the fast-time domain, while we reserve a portion of them to improve the overall ISL of waveforms without PIM. To this end, we divide the fast-time waveforms to be designed into multiple blocks, each of which contains multiple uniform segments. We further develop a rule to instruct the elaboration of segments for IE via PIM within each block, and we leave the remaining segments among blocks unconstrained for the reduction on overall ISL of the binary waveforms. Based on the above, we formulate the design into a non-convex optimization problem that incorporates both the ISL minimization of waveforms and constraints on the fast-time IE. To tackle this problem, we reformulate it into a solvable integer optimization form, for which we employ the coordinate descent framework to find solutions. To make the work complete, we propose an associated method to determine the proper number of IE segments in the waveform design. Simulation results verify the effectiveness of our proposed design. Xiaohan Zhao, Yongzhe Li, Ran Tao 0003 |
ICASSP | 1 |
| 2025 | ARrec: A GitHub Awesome Repository Recommendation Service Based on Graph Mining
Yanchun Sun, Xiaohan Zhao |
ICSOC (1) | 4 |
| 2025 | FADRM: Fast and Accurate Data Residual Matching for Dataset DistillationabstractResidual connection has been extensively studied and widely applied at the model architecture level. However, its potential in the more challenging data-centric approaches remains unexplored. In this work, we introduce the concept of ***Data Residual Matching*** for the first time, leveraging data-level skip connections to facilitate data generation and mitigate data information vanishing. This approach maintains a balance between newly acquired knowledge through pixel space optimization and existing core local information identification within raw data modalities, specifically for the dataset distillation task. Furthermore, by incorporating training-time refinements, our method significantly improves computational efficiency, achieving superior performance while reducing training time and peak GPU memory usage by 50\%. Consequently, the proposed method **F**ast and **A**ccurate **D**ata **R**esidual **M**atching for Dataset Distillation (**FADRM**) establishes a new state-of-the-art, demonstrating substantial improvements over existing methods across multiple dataset benchmarks in both efficiency and effectiveness. For instance, with ResNet-18 as the student model and a 0.8\% compression ratio on ImageNet-1K, the method achieves 48.4\% test accuracy in single-model dataset distillation and 50.9\% in multi-model dataset distillation, surpassing RDED by +6.4\% and outperforming state-of-the-art multi-model approaches, EDC and CV-DD, by +2.3\% and +4.9\%. Jiacheng Cui, Xinyue Bi, Yaxin Luo, Xiaohan Zhao |
NeurIPS | 4 |
| 2025 | A Frustratingly Simple Yet Highly Effective Attack Baseline: Over 90% Success Rate Against the Strong Black-box Models of GPT-4.5/4o/o1abstractDespite promising performance on open-source large vision-language models (LVLMs), transfer-based targeted attacks often fail against closed-source commercial LVLMs. Analyzing failed adversarial perturbations reveals that the learned perturbations typically originate from a uniform distribution and lack clear semantic details, resulting in unintended responses. This critical absence of semantic information leads commercial black-box LVLMs to either ignore the perturbation entirely or misinterpret its embedded semantics, thereby causing the attack to fail. To overcome these issues, we propose to refine semantic clarity by encoding explicit semantic details within local regions, thus ensuring the capture of finer-grained features and inter-model transferability, and by concentrating modifications on semantically rich areas rather than applying them uniformly. To achieve this, we propose *a simple yet highly effective baseline*: at each optimization step, the adversarial image is cropped randomly by a controlled aspect ratio and scale, resized, and then aligned with the target image in the embedding space. While the naive source-target matching method has been utilized before in the literature, we are the first to provide a tight analysis, which establishes a close connection between perturbation optimization and semantics. Experimental results confirm our hypothesis. Our adversarial examples crafted with local-aggregated perturbations focused on crucial regions exhibit surprisingly good transferability to commercial LVLMs, including GPT-4.5, GPT-4o, Gemini-2.0-flash, Claude-3.5/3.7-sonnet, and even reasoning models like o1, Claude-3.7-thinking and Gemini-2.0-flash-thinking. Our approach achieves success rates exceeding 90\% on GPT-4.5, 4o, and o1, significantly outperforming all prior state-of-the-art attack methods with lower $\ell_1/\ell_2$ perturbations. Our optimized adversarial examples under different configurations are available at https://huggingface.co/datasets/MBZUAI-LLM/M-Attack_AdvSamples and our training code at https://github.com/VILA-Lab/M-Attack. Xiaohan Zhao, Dong-Dong Wu, Jiacheng Cui |
NeurIPS | 2 |
| 2025 | Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM AgentsabstractCAPTCHAs have been a critical bottleneck for deploying web agents in real-world applications, often blocking them from completing end-to-end automation tasks. While modern multimodal LLM agents have demonstrated impressive performance in static perception tasks, their ability to handle interactive, multi-step reasoning challenges like CAPTCHAs is largely untested. To address this gap, we introduce Open CaptchaWorld, the first web-based benchmark and platform specifically designed to evaluate the visual reasoning and interaction capabilities of MLLM-powered agents through diverse and dynamic CAPTCHA puzzles. Our benchmark spans 20 modern CAPTCHA types, totaling 225 CAPTCHAs, annotated with a new metric we propose: CAPTCHA Reasoning Depth, which quantifies the number of cognitive and motor steps required to solve each puzzle. Experimental results show that humans consistently achieve near-perfect scores, state-of-the-art MLLM agents struggle significantly, with success rates at most 40.0\% by Browser-Use Openai-o3, far below human-level performance,93.3\%. This highlights Open CaptchaWorld as a vital benchmark for diagnosing the limits of current multimodal agents and guiding the development of more robust multimodal reasoning systems. Yaxin Luo, Jiacheng Cui, Xiaohan Zhao |
NeurIPS | 5 |
| 2024 | Exploring Vulnerabilities in Spiking Neural Networks: Direct Adversarial Attacks on Raw Event Data
Yanmeng Yao, Xiaohan Zhao, Bin Gu 0001 |
ECCV (72) | 2 |
| 2024 | Automatically Deriving Developers' Technical Expertise from the GitHub Social NetworkabstractDevelopers' technical expertise is crucial for various tasks within open-source communities, such as identifying suitable maintainers or reviewers. However, GitHub, the world's largest open-source code hosting platform, does not explicitly display developers' technical expertise. Existing methods fail to fully capture the multifaceted and dynamic nature of their skills and knowledge. To address this problem, we propose a novel approach to derive developers' technical expertise using graph neural networks (GNN). We construct a GitHub social network to integrate social and development activities and employ a GNN model to learn low-dimensional embedding for developers' technical expertise. We verify the effectiveness of our model on four GitHub social relationship recommendation tasks. The results demonstrate that our approach performs well in predicting technical preference for repositories and developers. Yanchun Sun, Xiaohan Zhao, Haizhou Xu, Ye Zhu 0002, Gang Huang 0001 |
ASE | 3 |
| 2023 | Efficent Large-Scale Multi-Unimodular Waveform Design with Good Correlation Properties via Direct Phase OptimizationsabstractIn this paper, we propose an efficient algorithm for designing large-scale multi-unimodular waveforms with low correlations. Different from existing approaches that commonly involve repetitive projections of complex values into their constant-modulus approximations, we conduct optimizations directly on the phase values of waveform elements. Specifically, we optimize the weighted integrated sidelobe level of waveforms, and formulate such design into an unconstrained optimization problem with respect to phase values of waveform elements. Then, we derive the gradient of the newly formulated objective function, through which we subsequently elaborate its majorant with the support of a properly designed Lipschitz-constant related quantity. Our major contributions also lie in obtaining a closed-form update of phase values that boils down to a gradient-descent regime, and calculating the update with fast implementations. Simulation results verify the superiority of our algorithm over existing state-of-the-art methods. Xiaohan Zhao, Yongzhe Li, Ran Tao 0003 |
ICASSP | 1 |
| 2023 | A Dual Domain Attention Mechanism for Face Forgery DetectionabstractRecently, deep face forgery detection has been attracting considerable attentions, due to the potential security consequences induced by this type of forgeries. Unfortunately, the existing techniques have not specifically considered the intrinsic differences between the frequency and spatial domain information. To explicitly accommodate different feature representations from different domains, we propose a novel Dual Domain Attention Mechanism (DDAM) for deep face forgery detection. Inspired by digital image processing, we construct a “soft” filter to adaptively filter the frequency information, which is irrelevant to our forgery detection. Besides, we construct a FC-based Attention Module to maintain a receptive field of the entire feature map, to better leverage contextual information from different domains. Extensive experiments demonstrate the effectiveness of the proposed method on widely used datasets. Yucong Suo, Xiaohan Zhao, Yuanfang Guo, Yangxi Li, Yunhong Wang 0001 |
IJCB | 2 |
| 2023 | Direct Training of SNN using Local Zeroth Order MethodabstractSpiking neural networks are becoming increasingly popular for their low energy requirement in real-world tasks with accuracy comparable to traditional ANNs. SNN training algorithms face the loss of gradient information and non-differentiability due to the Heaviside function in minimizing the model loss over model parameters. To circumvent this problem, the surrogate method employs a differentiable approximation of the Heaviside function in the backward pass, while the forward pass continues to use the Heaviside as the spiking function. We propose to use the zeroth-order technique at the local or neuron level in training SNNs, motivated by its regularizing and potential energy-efficient effects and establish a theoretical connection between it and the existing surrogate methods. We perform experimental validation of the technique on standard static datasets (CIFAR-10, CIFAR-100, ImageNet-100) and neuromorphic datasets (DVS-CIFAR-10, DVS-Gesture, N-Caltech-101, NCARS) and obtain results that offer improvement over the state-of-the-art results. The proposed method also lends itself to efficient implementations of the back-propagation method, which could provide 3-4 times overall speedup in training time. The code is available at \url{https://github.com/BhaskarMukhoty/LocalZO}. Bhaskar Mukhoty, Velibor Bojkovic, William de Vazelhes, Xiaohan Zhao, Giulia De Masi, Huan Xiong, Bin Gu 0001 |
NeurIPS | 4 |
| 2023 | Accelerated On-Device Forward Neural Network Training with Module-Wise Descending AsynchronismabstractOn-device learning faces memory constraints when optimizing or fine-tuning on edge devices with limited resources. Current techniques for training deep models on edge devices rely heavily on backpropagation. However, its high memory usage calls for a reassessment of its dominance.
In this paper, we propose forward gradient descent (FGD) as a potential solution to overcome the memory capacity limitation in on-device learning. However, FGD's dependencies across layers hinder parallel computation and can lead to inefficient resource utilization.
To mitigate this limitation, we propose AsyncFGD, an asynchronous framework that decouples dependencies, utilizes module-wise stale parameters, and maximizes parallel computation. We demonstrate its convergence to critical points through rigorous theoretical analysis.
Empirical evaluations conducted on NVIDIA's AGX Orin, a popular embedded device, show that AsyncFGD reduces memory consumption and enhances hardware efficiency, offering a novel approach to on-device learning. Xiaohan Zhao, Hualin Zhang, Zhouyuan Huo, Bin Gu 0001 |
NeurIPS | 1 |
| 2022 | Exploring the Impact of Adding Adversarial Perturbation onto Different Image RegionsabstractAdversarial attack has been a hot topic for a long time in machine learning and deep learning. Studying adversarial attack has vital significance to artificial intelligence security. Existing methods mainly pursue a higher attack success rate. Few researches pay attention to the region where adversarial perturbations are added. Actually, different pixels in an image usually have different contributions in results, which motivates us to apply region constraint in the image for adversarial perturbations generation. In this paper, we present an easy-to-implement way to decrease the unnecessary adversarial perturbations while preserving a relatively high attack success rate. Specifically, we do not use the same constraint of perturbations in the input image but set specific constraint for specific region. Furthermore, we point that adversarial examples work in a different way to normal images. Directly using the activated region in normal images is not optimal. Then, to get the crucial area in adversarial attacks, we propose six transformation schemes to revise the activated region which is generated by the normal image. We launch extensive experiments on ImageNet dataset and the results show that our methods can get better attack strength under the same perturbation level when compared to the baseline methods. Ruijie Yang, Yuanfang Guo, Ruikui Wang, Xiaohan Zhao, Yunhong Wang 0001 |
ISCAS | 4 |
| 2022 | GAGA: Deciphering Age-path of Generalized Self-paced RegularizerabstractNowadays self-paced learning (SPL) is an important machine learning paradigm that mimics the cognitive process of humans and animals. The SPL regime involves a self-paced regularizer and a gradually increasing age parameter, which plays a key role in SPL but where to optimally terminate this process is still non-trivial to determine. A natural idea is to compute the solution path w.r.t. age parameter (i.e., age-path). However, current age-path algorithms are either limited to the simplest regularizer, or lack solid theoretical understanding as well as computational efficiency. To address this challenge, we propose a novel Generalized Age-path Algorithm (GAGA) for SPL with various self-paced regularizers based on ordinary differential equations (ODEs) and sets control, which can learn the entire solution spectrum w.r.t. a range of age parameters. To the best of our knowledge, GAGA is the first exact path-following algorithm tackling the age-path for general self-paced regularizer. Finally the algorithmic steps of classic SVM and Lasso are described in detail. We demonstrate the performance of GAGA on real-world datasets, and find considerable speedup between our algorithm and competing baselines. Xingyu Qu, Diyang Li, Xiaohan Zhao, Bin Gu 0001 |
NeurIPS | 3 |
| 2022 | JoinTW: A Joint Image-to-Image Translation and Watermarking Method
Xiaohan Zhao, Yunhong Wang 0001, Ruijie Yang, Yuanfang Guo |
PRCV (3) | 1 |
| 2022 | EPIHC: Improving Enhancer-Promoter Interaction Prediction by Using Hybrid Features and Communicative LearningabstractEnhancer-promoter interactions (EPIs) regulate the expression of specific genes in cells, which help facilitate understanding of gene regulation, cell differentiation and disease mechanisms. EPI identification approaches through wet experiments are often costly and time-consuming, leading to the design of high-efficiency computational methods is in demand. In this paper, we propose a deep neural network-based method named EPIHC to predict Enhancer-Promoter Interactions with Hybrid features and Communicative learning. EPIHC extracts enhancer and promoter sequence-derived features using convolutional neural networks (CNN), and then we design a communicative learning module to capture the communicative information between enhancer and promoter sequences. Besides, EPIHC takes the genomic features of enhancers and promoters into account, incorporating with the sequence-derived features to predict EPIs. The computational experiments show that EPIHC outperforms the existing state-of-the-art EPI prediction methods on the benchmark datasets and chromosome-split datasets, and the study reveals that the communicative learning module can bring explicit information about EPIs, which is ignored by CNN, and provide explainability about EPIs to some degree. Moreover, we consider two strategies to improve the performances of EPIHC in the cross-cell line prediction, and experimental results show that EPIHC constructed on some cell lines can exhibit good performances for other cell lines. The codes and data are available at https://github.com/BioMedicalBigDataMiningLab/EPIHC. Shuai Liu 0017, Xinran Xu, Xiaohan Zhao, Shichao Liu 0002, Wen Zhang 0008 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2021 | Predicting drug-disease associations through layer attention graph convolutional networkabstractBACKGROUND: Determining drug-disease associations is an integral part in the process of drug development. However, the identification of drug-disease associations through wet experiments is costly and inefficient. Hence, the development of efficient and high-accuracy computational methods for predicting drug-disease associations is of great significance. RESULTS: In this paper, we propose a novel computational method named as layer attention graph convolutional network (LAGCN) for the drug-disease association prediction. Specifically, LAGCN first integrates the known drug-disease associations, drug-drug similarities and disease-disease similarities into a heterogeneous network, and applies the graph convolution operation to the network to learn the embeddings of drugs and diseases. Second, LAGCN combines the embeddings from multiple graph convolution layers using an attention mechanism. Third, the unobserved drug-disease associations are scored based on the integrated embeddings. Evaluated by 5-fold cross-validations, LAGCN achieves an area under the precision-recall curve of 0.3168 and an area under the receiver-operating characteristic curve of 0.8750, which are better than the results of existing state-of-the-art prediction methods and baseline methods. The case study shows that LAGCN can discover novel associations that are not curated in our dataset. CONCLUSION: LAGCN is a useful tool for predicting drug-disease associations. This study reveals that embeddings from different convolution layers can reflect the proximities of different orders, and combining the embeddings by the attention mechanism can improve the prediction performances. Zhouxin Yu, Feng Huang 0004, Xiaohan Zhao, Wenjie Xiao, Wen Zhang 0008 |
Briefings Bioinform. | 3 |
| 2021 | Orthodontic simulation system with force feedback for training complete bracket placement proceduresabstractA virtual system that simulates the complete process of orthodontic bracket placement can be used for pre-clinical skill training to help students gain confidence by performing the required tasks on a virtual patient. The hardware for the virtual simulation system is built using two force feedback devices to support bi-manual force feedback operation. A 3D mouse is used to adjust the position of the virtual patient. A multi-threaded computational methodology is adopted to satisfy the requirements of the frame rate. The computation threads mainly consist of the haptic thread running at a frequency of >1000Hz and the graphic thread at >30Hz. The graphic thread allows the graphics engine to effectively display the visual effects of biofilm removal and acid erosion through texture mapping. Using the haptic thread, the physics engine adopts the hierarchy octree collision-detection algorithm to simulate the multi-point and multi-region interaction between the tools and the virtual environment. Its high efficiency guarantees that the time cost can be controlled within 1 ms. The physics engine also performs collision detection between the tools and particles, making it possible to simulate paint and removal of colloids. The surface-contact constraints are defined in the system; this ensures that the bracket will not divorce from or embed into the tooth during the adjustment of the bracket. Therefore, the simulated adjustment is more realistic and natural. A virtual system to simulate the complete process of orthodontic bracket bonding was developed. In addition to bracket bonding and adjustment, the system simulates the necessary auxiliary steps such as smearing, acid etching, and washing. Furthermore, the system supports personalized case training. The system provides a new method for students to practice orthodontic skills. Luwei Liu, Xiaohan Zhao, Aimin Hao |
Virtual Real. Intell. Hardw. | 4 |
| 2019 | LncPred-IEL: A Long Non-coding RNA Prediction Method using Iterative Ensemble LearningabstractA large number of transcripts have been generated by the development of high throughput sequencing technologies. Predicting lncRNA from transcripts is a challenging and important task. In this paper, we propose LncPred-IEL, an iterative ensemble learning long non-coding RNA prediction method. LncPred-IEL not only considers features widely used for the lncRNA prediction, but also take into account sequence-derived features used in the RNA sequence classification, so as to make use of diverse information. LncPred-IEL builds base predictors based on different groups of features, and employs a supervised iterative way to combine base predictors and build ensemble models. Our studies demonstrate that supervised iterative way can learn the representations that help to separate lncRNA and protein-coding transcripts, and further improve the performances. Experiments demonstrate that LncPred-IEL outperforms several state-of-the-art methods when evaluated by 10-fold cross-validation. The capability of LncPred-IEL for the cross-species prediction is also tested. As complementary to wet experiments, LncPred-IEL is a useful computational tool for lncRNA prediction. Yanzhen Xu, Xiaohan Zhao, Shuai Liu 0017, Shichao Liu 0002, Yanqing Niu, Wen Zhang 0008, Leyi Wei |
BIBM | 2 |
| 2018 | Substructure Assembling Network for Graph ClassificationabstractGraphs are natural data structures adopted to represent real-world data of complex relationships. In recent years, a surge of interest has been received to build predictive models over graphs, with prominent examples in chemistry, computational biology, and social networks. The overwhelming complexity of graph space often makes it challenging to extract interpretable and discriminative structural features for classification tasks. In this work, we propose a novel neural network structure called Substructure Assembling Network (SAN) to extract graph features and improve the generalization performance of graph classification. The key innovation of our work is a unified substructure assembling unit, which is a variant of Recurrent Neural Network (RNN) designed to hierarchically assemble useful pieces of graph components so as to fabricate discriminative substructures. SAN adopts a sequential, probabilistic decision process, and therefore it can tune substructure features in a finer granularity. Meanwhile, the parameterized soft decisions can be continuously improved with supervised learning through back-propagation, leading to optimizable search trajectories. Overall, SAN embraces both the flexibility of combinatorial pattern search and the strong optimizability of deep learning, and delivers promising results as well as interpretable structural features in graph classification against state-of-the-art techniques. Xiaohan Zhao, Bo Zong, Ziyu Guan, Wei Zhao 0019 |
AAAI | 1 |
| 2016 | Network Growth and Link Prediction Through an Empirical Lens
Qingyun Liu 0003, Shiliang Tang, Xinyi Zhang 0003, Xiaohan Zhao, Ben Y. Zhao, Haitao Zheng 0001 |
Internet Measurement Conference | 4 |
| 2016 | XFabric: A Reconfigurable In-Rack Network for Rack-Scale Computers
Sergey Legtchenko, Nicholas Chen, Daniel Cletheroe, Antony I. T. Rowstron, Hugh Williams, Xiaohan Zhao |
NSDI | 6 |
| 2014 | Link and Triadic Closure Delay: Temporal Metrics for Social Network Dynamics
Matteo Zignani, Sabrina Gaito, Gian Paolo Rossi 0001, Xiaohan Zhao, Haitao Zheng 0001, Ben Y. Zhao |
ICWSM | 4 |
| 2013 | On the validity of geosocial mobility tracesabstractMobile networking researchers have long searched for large-scale, fine-grained traces of human movement, which have remained elusive for both privacy and logistical reasons. Recently, researchers have begun to focus on geosocial mobility traces, e.g. Foursquare checkin traces, because of their availability and scale. But are we conceding correctness in our zeal for data? In this paper, we take initial steps towards quantifying the value of geosocial datasets using a large ground truth dataset gathered from a user study. By comparing GPS traces against Foursquare checkins, we find that a large portion of visited locations is missing from checkins, and most checkin events are either forged or superfluous events. We characterize extraneous checkins, describe possible techniques for their detection, and show that both extraneous and missing checkins introduce significant errors into applications driven by these traces. Zengbin Zhang, Xiaohan Zhao, Gang Wang 0011, Yu Su 0001, Miriam J. Metzger, Haitao Zheng 0001, Ben Y. Zhao |
HotNets | 3 |
| 2013 | On the Embeddability of Random Walk DistancesabstractAnalysis of large graphs is critical to the ongoing growth of search engines and social networks. One class of queries centers around node affinity, often quantified by random-walk distances between node pairs, including hitting time, commute time, and personalized PageRank (PPR). Despite the potential of these "metrics," they are rarely, if ever, used in practice, largely due to extremely high computational costs. In this paper, we investigate methods to scalably and efficiently compute random-walk distances, by "embedding" graphs and distances into points and distances in geometric coordinate spaces. We show that while existing graph coordinate systems (GCS) can accurately estimate shortest path distances, they produce significant errors when embedding random-walk distances. Based on our observations, we propose a new graph embedding system that explicitly accounts for per-node graph properties that affect random walk. Extensive experiments on a range of graphs show that our new approach can accurately estimate both symmetric and asymmetric random-walk distances. Once a graph is embedded, our system can answer queries between any two nodes in 8 microseconds, orders of magnitude faster than existing methods. Finally, we show that our system produces estimates that can replace ground truth in applications with minimal impact on application output. Xiaohan Zhao, Adelbert Chang, Atish Das Sarma, Haitao Zheng 0001, Ben Y. Zhao |
Proc. VLDB Endow. | 1 |
| 2012 | Multi-scale dynamics in a massive online social networkabstractData confidentiality policies at major social network providers have severely limited researchers' access to large-scale datasets. The biggest impact has been on the study of network dynamics, where researchers have studied citation graphs and content-sharing networks, but few have analyzed detailed dynamics in the massive social networks that dominate the web today. In this paper, we present results of analyzing detailed dynamics in a large Chinese social network, covering a period of 2 years when the network grew from its first user to 19 million users and 199 million edges. Rather than validate a single model of network dynamics, we analyze dynamics at different granularities (per-user, per-community, and network-wide) to determine how much, if any, users are influenced by dynamics processes at different scales. We observe independent predictable processes at each level, and find that the growth of communities has moderate and sustained impact on users. In contrast, we find that significant events such as network merge events have a strong but short-lived impact on users, and they are quickly eclipsed by the continuous arrival of new users. Xiaohan Zhao, Alessandra Sala, Christo Wilson, Xiao Wang 0018, Sabrina Gaito, Haitao Zheng 0001, Ben Y. Zhao |
Internet Measurement Conference | 1 |
| 2012 | Serf and turf: crowdturfing for fun and profitabstractPopular Internet services in recent years have shown that remarkable things can be achieved by harnessing the power of the masses using crowd-sourcing systems. However, crowd-sourcing systems can also pose a real challenge to existing security mechanisms deployed to protect Internet services. Many of these security techniques rely on the assumption that malicious activity is generated automatically by automated programs. Thus they would perform poorly or be easily bypassed when attacks are generated by real users working in a crowd-sourcing system. Through measurements, we have found surprising evidence showing that not only do malicious crowd-sourcing systems exist, but they are rapidly growing in both user base and total revenue. We describe in this paper a significant effort to study and understand these "crowdturfing" systems in today's Internet. We use detailed crawls to extract data about the size and operational structure of these crowdturfing systems. We analyze details of campaigns offered and performed in these sites, and evaluate their end-to-end effectiveness by running active, benign campaigns of our own. Finally, we study and compare the source of workers on crowdturfing sites in different countries. Our results suggest that campaigns on these systems are highly effective at reaching users, and their continuing growth poses a concrete threat to online communities both in the US and elsewhere. Gang Wang 0011, Christo Wilson, Xiaohan Zhao, Yibo Zhu 0001, Manish Mohanlal, Haitao Zheng 0001, Ben Y. Zhao |
WWW | 3 |
| 2011 | Efficient shortest paths on massive social graphsabstractAnalysis of large networks is a critical component of many of today’s application environments. The arrival of massive network graphs with hundreds of millions of nodes, e.g. social graphs, presents a unique challenge to graph analysis applications. Most of these applications rely on computing dista Xiaohan Zhao, Alessandra Sala, Haitao Zheng 0001, Ben Y. Zhao |
CollaborateCom | 1 |
| 2011 | Sharing graphs using differentially private graph modelsabstractContinuing success of research on social and computer networks requires open access to realistic measurement datasets. While these datasets can be shared, generally in the form of social or Internet graphs, doing so often risks exposing sensitive user data to the public. Unfortunately, current techniques to improve privacy on graphs only target specific attacks, and have been proven to be vulnerable against powerful de-anonymization attacks. Alessandra Sala, Xiaohan Zhao, Christo Wilson, Haitao Zheng 0001, Ben Y. Zhao |
Internet Measurement Conference | 2 |
| 2009 | SLINCS: A Social Link Based Evaluation System for Network Coordinate SystemsabstractIn recent research work of securing Network Coordinate (NC) system, they concentrate on the passive security defense mechanisms. In this paper we propose SLINCS, a social link based evaluation security system that utilizes information from existing social relationship networks to implement proactive security mechanisms for NC systems. The key idea is to eliminate suspicious nodes before they launch potential attacks. Xiaohan Zhao, Eng Keong Lua, Zengbin Zhang, Beixing Deng, Xing Li 0001 |
CCNC | 2 |
| 2009 | Phoenix: Towards an Accurate, Practical and Decentralized Network Coordinate System
Yang Chen 0001, Xiao Wang 0017, Eng Keong Lua, Xiaohan Zhao, Beixing Deng, Xing Li 0001 |
Networking | 6 |