Amarda Shehu

dblp:53/3810 · DBLP profile ↗
← Back
6ranked-venue papers in the field
1as first author
4since 2021 · last 2026
0000-0001-5230-4610ORCID · verified

Domains — venue-derived; a paper can count in several

Data Mining & Knowledge Discovery · 6 (1 first)
YearPublicationVenuePosition
2026 Beyond the Singularity Myth: Artificial General Intelligence as Cumulative Infrastructural Transformation - Absorption Capacity, Epistemic Drift, and the Erosion of Human Verification Power
abstract
The dominant framing of Artificial General Intelligence (AGI) as a discrete breakthrough obscures the more urgent reality: AGI is arriving as a gradual, cumulative erosion of human verification power distributed across institutions and decision-making systems. This article reframes the AGI transition through the lens of absorption capacity; that is, the rate at which human systems can integrate, govern, and maintain meaningful oversight of increasingly autonomous AI. Drawing on empirical observations from deploying enterprise-scale generative AI in a large public university and personal experiences as a long-standing AI researcher and educator, in this article, I identify three critical asymmetries characterizing this transition: (1) governance lag, where policy cycles cannot match technological iteration speed; (2) institutional misalignment, where locally rational AI systems produce collectively irrational societal outcomes; and (3) capability inequality, where uneven access to AI amplifies structural advantage. I argue that the defining challenge is not achieving technical alignment with human values, but maintaining epistemic authority, which is the human capacity to verify, understand, and steer systems reasoning in latent spaces beyond direct audit. The article concludes that the true measure of preparedness for AGI is not computational power or algorithmic sophistication, but adaptive governance: institutional architectures capable of co-evolving with the technologies they must regulate. The frontier is not artificial superintelligence. It is collective human capacity to remain intelligible to ourselves while embedded in AI-mediated decision ecosystems.
Amarda Shehu
ACM Trans. Intell. Syst. Technol.1
2025 An Instructible Chemist-AI Alignment Framework for Generating Quaternary Ammonium Compound Structures
abstract
This paper presents a novel Chemist-AI Alignment framework for generating novel structures of quaternary ammonium compounds (QACs), a crucial class of antimicrobial agents.The framework uniquely integrates AI-driven small molecule generation with iterative feedback from chemist experts, leveraging both rapid assessments and comprehensive wet-lab validations to optimize for biological potency and synthetic feasibility.Central to the framework is a hierarchical generative model that captures the QAC hierarchical topology.Extensive experiments highlight the efficacy of the framework in identifying promising QAC candidates, many
Bo Pan 0009, Shiva Ghaemi, Amanda J. Consylman, Ashley Ann Petersen, Alice Wu, Gabriel Chang, Diana McDonough, Mark A. Forman, Elise L. Bezold, William M. Wuest, Kevin Minbiole, Liang Zhao 0002, Amarda Shehu
KDD (2)14
2025 Better AI For Understanding Life on Earth: Predict First, Design Later
abstract
Generative AI is generating much enthusiasm on potentially advancing biological design in computational biology. In this paper we take a somewhat contrarian view, arguing that a broader and deeper understanding of existing biological sequences is essential before undertaking the design of novel ones. We draw attention, for instance, to current protein function prediction methods which currently face significant limitations due to incomplete data and inherent challenges in defining and measuring function. We propose a “blue sky” vision centered on both comprehensive and precise annotation of existing protein and DNA sequences, aiming to develop a more complete and precise understanding of biological function. By contrasting recent studies that leverage generative AI for biological design with the pressing need for enhanced data annotation, we underscore the importance of prioritizing robust predictive models over premature generative efforts. We advocate for a strategic shift toward thorough sequence annotation and predictive understanding, laying a solid foundation for future advances in biological design.
Yana Bromberg, Amarda Shehu
SDM2
2022 Interpretable Molecular Graph Generation via Monotonic Constraints
abstract
Designing molecules with specific properties is a long-lasting research problem and is central to advancing crucial domains such as drug discovery and material science. Recent advances in deep graph generative models treat molecule design as graph generation problems which provide new opportunities toward the breakthrough of this long-lasting problem. Existing models, however, have many shortcomings, including poor interpretability and controllability toward desired molecular properties. This paper focuses on new methodologies for molecule generation with interpretable and controllable deep generative models, by proposing new monotonically-regularized graph variational autoencoders. The proposed models learn to represent the molecules with latent variables and then learn the correspondence between them and molecule properties parameterized by polynomial functions. To further improve the intepretability and controllability of molecule generation towards desired properties, we derive new objectives which further enforce monotonicity of the relation between some latent variables and target molecule properties such as toxicity and clogP. Extensive experimental evaluation demonstrates the superiority of the proposed framework on accuracy, novelty, disentanglement, and control towards desired molecular properties. The code is anonymized at https://anonymous.4open.science/r/MDVAE-FD2C.
Yuanqi Du, Xiaojie Guo 0002, Amarda Shehu, Liang Zhao 0002
SDM3
2020 Interpretable Deep Graph Generation with Node-edge Co-disentanglement
abstract
Disentangled representation learning has recently attracted a significant amount of attention, particularly in the field of image representation learning. However, learning the disentangled representations behind a graph remains largely unexplored, especially for the attributed graph with both node and edge features. Disentanglement learning for graph generation has substantial new challenges including 1) the lack of graph deconvolution operations to jointly decode node and edge attributes; and 2) the difficulty in enforcing the disentanglement among latent factors that respectively influence: i) only nodes, ii) only edges, and iii) joint patterns between them. To address these challenges, we propose a new disentanglement enhancement framework for deep generative models for attributed graphs. In particular, a novel variational objective is proposed to disentangle the above three types of latent factors, with novel architecture for node and edge deconvolutions. Qualitative and quantitative experiments on both synthetic and real-world datasets demonstrate the effectiveness of the proposed model and its extensions.
Xiaojie Guo 0002, Liang Zhao 0002, Zhao Qin, Lingfei Wu 0001, Amarda Shehu, Yanfang Ye 0001
KDD5
2020 Reconstruction and Decomposition of High-Dimensional Landscapes via Unsupervised Learning
abstract
Uncovering the organization of a landscape that encapsulates all states of a dynamic system is a central task in many domains, as it promises to reveal, in an unsupervised manner, a system's inner working. One domain where this task is crucial is in bioinformatics, where the energy landscape that organizes three-dimensional structures of a molecule by their energetics is a powerful construct. The landscape can be leveraged, among other things, to reveal macrostates where a molecule is biologically-active. This is a daunting task, as landscapes of complex actuated systems, such as molecules, are inherently high-dimensional. Nonetheless, our laboratories have made some progress via topological and statistical analysis of spatial data over the recent years. We have proposed what is essentially a dichotomy, methods that are more pertinent for visualization-driven discovery, and methods that are more pertinent for discovery of the biologically-active macrostates but not amenable to visualization. In this paper, we present a novel, hybrid method that combines strengths of these methods, allowing both visualization of the landscape and discovery of macrostates. We demonstrate what the method is capable of uncovering in comparison with existing methods over structure spaces sampled with conformational sampling algorithms. Though the direct evaluation in this paper is on protein energy landscapes, the proposed method is of broad interest in cross-cutting problems that necessitate characterization of fitness and optimization landscapes.
Nasrin Akhter 0001, Wanli Qiao, Amarda Shehu
KDD4