June Zhang

dblp:120/7289 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
4since 2021 · last 2025
0000-0002-4578-5759ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Security and privacy · 2 · 1 since 2021Theory of computation · 2 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 mRNA2vec: mRNA Embedding with Language Model in the 5'UTR-CDS for mRNA Design
abstract
Messenger RNA (mRNA)-based vaccines are accelerating the discovery of new drugs and revolutionizing the pharmaceutical industry. However, selecting particular mRNA sequences for vaccines and therapeutics from extensive mRNA libraries is costly. Effective mRNA therapeutics require carefully designed sequences with optimized expression levels and stability. This paper proposes a novel contextual language model (LM)-based embedding method: mRNA2vec. In contrast to existing mRNA embedding approaches, our method is based on the self-supervised teacher-student learning framework of data2vec. We jointly use the 5' untranslated region (UTR) and coding sequence (CDS) region as the input sequences. We adapt our LM-based approach specifically to mRNA by 1) considering the importance of location on the mRNA sequence with probabilistic masking, 2) using Minimum Free Energy (MFE) prediction and Secondary Structure (SS) classification as additional pretext tasks. mRNA2vec demonstrates significant improvements in translation efficiency (TE) and expression level (EL) prediction tasks in UTR compared to SOTA methods such as UTR-LM. It also gives a competitive performance in mRNA stability and protein production level tasks in CDS such as CodonBERT.
Honggen Zhang, Xiangrui Gao, June Zhang, Lipeng Lai
AAAI3
2024 Out-of-Distribution Detection Using Maximum Entropy Coding and Generative Networks
abstract
Given a default distribution$P$and a set of test data$x^{M}=\{x_{1},\ x_{2},\ \ldots,\ x_{M}\}$this paper seeks to answer the question if it was likely that$x^{M}$was generated by$P$. For discrete distributions, the definitive answer is in principle given by Kolmogorov-Martin-Lof randomness. In this paper we seek to generalize this to continuous distributions. We consider a set of statistics$T_{1}(x^{M}), T_{2}(x^{M}),\cdots$. To each statistic we associate its maximum entropy distribution and with this a universal source coder. The maximum entropy distributions are subsequently combined to give a total codelength, which is compared with$-\log P(x^{M})$. We show that this approach satisfied a number of theoretical properties. For real world data$P$usually is unknown. We transform data into a standard distribution in the latent space using a bidirectional generate network and use maximum entropy coding there. We compare the resulting method to other methods that also used generative neural networks to detect anomalies. In most cases, our results show better performance.
Mojtaba Abolfazli, Mohammad Zaeri Amirani, Anders Høst-Madsen, June Zhang, Andras Bratincsak
ISITA4
2024 HaSa: Hardness and Structure-Aware Contrastive Knowledge Graph Embedding
abstract
We consider a contrastive learning approach to knowledge graph embedding (KGE) via InfoNCE. For KGE, efficient learning relies on augmenting the training data with negative triples. However, most KGE works overlook the bias from generating the negative triples- false negative triples (factual triples missing from the knowledge graph). We argue that generating high-quality (i.e., hard) negative triples might lead to an increase in false negative triples. To mitigate the impact of false negative triples during the generation of hard negative triples, we propose the Hardness and Structure-aware (HaSa) contrastive KGE method, which alleviates the effect of false negative triples while generating the hard negative triples. Experiments show that HaSa improves the performance of InfoNCE-based KGE approaches and achieves state-of-the-art results in several metrics for WN18RR datasets and competitive results for FB15k-237 datasets compared to classic and pre-trained LM-based KGE methods.
Honggen Zhang, June Zhang, Igor Molybog
WWW2
2021 Graph Coding for Model Selection and Anomaly Detection in Gaussian Graphical Models
abstract
A classic application of description length is for model selection with the minimum description length (MDL) principle. The focus of this paper is to extend description length for data analysis beyond simple model selection and sequences of scalars. More specifically, we extend the description length for data analysis in Gaussian graphical models. These are powerful tools to model interactions among variables in a sequence of i.i.d Gaussian data in the form of a graph. Our method uses universal graph coding methods to accurately account for model complexity, and therefore provide a more rigorous approach for graph model selection. The developed method is tested with synthetic and electrocardiogram (ECG) data to find the graph model and anomaly in Gaussian graphical models. The experiments show that our method gives better performance compared to commonly used methods.
Mojtaba Abolfazli, Anders Høst-Madsen, June Zhang, Andras Bratincsak
ISIT3
2020 Differential Description Length for Hyperparameter Selection in Supervised Learning
Mojtaba Abolfazli, Anders Høst-Madsen, June Zhang
ISITA3
2018 Greedy Algorithm with Approximation Ratio for Sampling Noisy Graph Signals
abstract
We study the optimal sampling set selection problem in sampling a noisy k -bandlimited graph signal. To minimize the effect of noise when trying to reconstruct a k -bandlimited graph signal from m samples, the optimal sampling set selection problem has been shown to be equivalent to finding a m×k submatrix with the maximum smallest singular value, σmin [3]. As the problem is NP-hard, we present a greedy algorithm inspired by a similar submatrix selection problem known in computer science and to which we add a local search refinement. We show that 1) in experiments, our algorithm finds a submatrix with larger σmin than prior greedy algorithm [3], and 2) has a proven worst-case approximation ratio of 1/(1+ε)k, where ε is a constant.
Changlong Wu, June Zhang
ICASSP3
2018 Who is More at Risk in Heterogenous Networks?
abstract
Network-based epidemics models try to characterize the impact of network topology, which represents contagion pathways, on the spread of infection. Although these models explicitly consider the dynamics of individuals in the given network (i.e., the state of the system is x(t)=[x1(t), x2(t), ..., xN(t)]T), analysis has focused on characterizing the vulnerability of the entire population rather than the vulnerability of the individuals in the population. We focus on characterizing the vulnerability of the ith individual in the network by studying the marginal probability of infection, P(xi=1), of the scaled SIS process. Studying the vulnerability of individuals is important because it may be tempting to assume that P(xi=1) is related to the degree of the ith. node. Since infection rate is usually assumed to be dependent on the number of infected neighbors, then it seems reasonable that nodes with more connections (i.e., higher degree) would be more at risk. We show that this is not always true. Further, with a closed-form approximation of P(xi=1), as solving for the exact probability requires the summation of 2Nterms, we characterize the conditions for when degree distribution is a good indicator of how susceptible an individual is to infection.
June Zhang, José M. F. Moura
ICASSP1
2018 Coding of Graphs with Application to Graph Anomaly Detection
abstract
This paper has dual aims. First is to develop practical universal coding methods for unlabeled graphs. Second is to use these for graph anomaly detection. The paper develops two coding methods for unlabeled graphs: one based on the degree distribution, the second based on the triangle distribution. It is shown that these are efficient for different types of random graphs, and on real-world graphs. These coding methods is then used for detecting anomalous graphs, based on structure alone. It is shown that anomalous graphs can be detected with high probability.
Anders Høst-Madsen, June Zhang
ISIT2
2016 Finding unique dense communities
abstract
Finding densely connected subgraphs, also called communities, in networks are of interest for many applications. In previous work, we showed an optimization method for efficiently finding subgraphs denser than the overall network [1]. This result is derived from our studies of network processes, dynamical processes that model interactions between individual agents in networks (i.e., spread of infection or cascading failures). In this paper, we prove that these subgraphs are also unique in the sense that there are no other subgraphs in the network isomorphic to these subgraphs.
June Zhang, José M. F. Moura
ICASSP1
2014 Subgraph density and epidemics over networks
abstract
We model a SIS (susceptible-infected-susceptible) epidemics over a static, finite-sized network as a continuous-time Markov process using the scaled SIS epidemics model. In our previous work, we derived the closed form description of the equilibrium distribution that explicitly accounts for the network topology and showed that the most probable equilibrium state demonstrates threshold behavior. In this paper, we will show how subgraph structures in the network topology impact the most probable state of the long run behavior of a SIS epidemics (i.e., stochastic diffusion process) over any static, finite-sized, network.
June Zhang, José M. F. Moura
ICASSP1
2013 Threshold behavior of epidemics in regular networks
abstract
Current research is interested in identifying how topology impacts epidemics in networks. In this paper, we model SIS (susceptible-infected-susceptible) epidemics as a continuous-time Markov process and for which we can obtain a closed form description of the equilibrium distribution. Such distribution describes the long-run behavior of the epidemics. The adjacency matrix of the network topology is reflected explicitly in the formulation of the equilibrium distribution. Secondly, we are interested in analyzing the model in the regime where the topology dependent infection process opposes the topology independent healing process. Specifically, how will network topology affect the most probable long-run network state? We show that for k-regular graph topologies, the most probable network state transitions from the state where everyone is healthy to one where everyone is infected at a threshold that depends on k but not on the size of the graph.
June Zhang, José M. F. Moura
ICASSP1
2012 Accounting for topology in spreading contagion in non-complete networks
abstract
We are interested in investigating the spread of contagion in a network, G, which describes the interactions between the agents in the system. The topology of this network is often neglected due to the assumption that each agent is connected with every other agents; this means that the network topology is a complete graph. While this allows for certain simplifications in the analysis, we fail to gain insight on the diffusion process for non-complete network topology. In this paper, we offer a continuous-time Markov chain infection model that explicitly accounts for the network topology, be it complete or non-complete. Although we characterize our process using parameters from epidemiology, our approach can be applied to many application domains. We will show how to generate the infinitesimal matrix that describes the evolution of this process for any topology. We also develop a general methodology to solve for the equilibrium distribution by considering symmetries in G. Our results show that network topologies have dramatic effect on the spread of infections.
June Zhang, José M. F. Moura
ICASSP1