Yiqin Gao

dblp:252/6934 · also Yi Qin Gao · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Theory of computation · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Bidirectional GPT
Chuanliu Fan, Zicheng Ma, Jun Zhang 0071, Yiqin Gao, Ziqiang Cao, Guohong Fu
Inf. Process. Manag.6
2025 LlaMol: A Unified Molecule Designer via Preference Ranking and Numerical Enhancement
abstract
Goal-oriented de novo molecule design, namely generating molecules with specific property or substructure constraints from scratch, is a crucial yet challenging task in drug discovery. Existing research often relies on separate predictors for distinct properties and struggles with integrating substructure constraints due to the complexities involved in modeling structural information via multitask learning. This separation necessitates a dedicated prediction model for each constraint, limiting the flexibility and posing challenges for realworld applications. To address these limitations, we propose a unified framework for molecular design that incorporates multiple property and substructure constraints, leveraging LLMs to handle diverse constraint settings within a single model. We first integrate feedback learning derived from preference ranking to eliminate the need for separate property predictors. Then, we enhance the model's ability to follow numerical instructions by introducing a unified numerical encoding into the prompt. We conduct extensive experiments across single-property, substructureproperty, and multi-property constrained tasks. Experimental results demonstrate that LlaMol consistently outperforms state-of-the-art baselines across various constraint settings. Notably, in the multi-objective binding affinity maximization task, LlaMol achieves a significantly lower$\mathrm{K}_{\mathrm{D}}$value of 0.25 for the protein target ESR1, while maintaining the highest overall performance, surpassing previous methods by 4.76 %. These results underscore the effectiveness and versatility of LLM-based frameworks for molecule generation under complex constraints.
Chuanliu Fan, Zicheng Ma, Jun Zhang 0069, Ziqiang Cao, Yiqin Gao, Guohong Fu
BIBM7
2025 Prot2Chat: protein large language model with early fusion of text, sequence, and structure
abstract
MOTIVATION: Proteins are of great significance in living organisms. However, understanding their functions encounters numerous challenges, such as insufficient integration of multimodal information, a large number of training parameters, limited flexibility of classification-based methods, and the lack of systematic evaluation metrics for protein question answering systems. To tackle these issues, we propose the Prot2Chat framework. RESULTS: We modified ProteinMPNN to encode protein sequence and structural information in a unified way. We used a large language model (LLM) to encode questions into vectors and developed a protein-text adapter to compress protein information into virtual tokens based on these vectors, achieving the early fusion of text and protein information. Finally, the same LLM reads the virtual tokens and the questions to generate answers. To optimize training efficiency, we froze the encoder and employed low-rank adaptation (LoRA) techniques for the LLM. Experiments on two datasets show that both automated metrics and expert evaluations demonstrate the superior performance of our model, and zero-shot prediction results highlight its generalization ability. We have developed an easy-to-use web interactive platform and a rapid installation option, allowing users to swiftly engage with Prot2Chat. AVAILABILITY AND IMPLEMENTATION: The models and codes are available at https://github.com/wangzc1233/Prot2Chat.
Zhicong Wang, Zicheng Ma, Ziqiang Cao, Changlong Zhou, Jun Zhang 0071, Yiqin Gao
Bioinform.6
2024 6Diffusion-LM: IPv6 address generation method based on diffusion-LM
abstract
IPv6 is instrumental in the ultra-large-scale intelligent computing interconnection system, yet its integration is not without challenges. The vast IPv6 address space renders traditional brute-force scanning methods infeasible, with the considerable time and resource consumption severely impacting the integration of supercomputing capabilities. This also affects the accuracy and efficiency required in high-performance computing environments. Consequently, it becomes necessary to develop new scanning technologies to address the unique challenges presented by the expansive IPv6 address space. Our novel approach, 6Diffusion-LM, transforms IPv6 scanning by fusing diffusion and linguistic models. Utilizing the Transformer architecture, it excels at extracting key features from IPv6 addresses and employs clustering algorithms to organize them effectively. Building upon the BERT pre-trained language model, 6Diffusio-LM integrate a noise mechanism that encapsulates the inherent randomness and unpredictability inherent in IPv6 address generation. The model then refines this process by progressively eliminating noise to yield precise and clear IPv6 addresses. Additionally, our proprietary embedding method enhances the generation process, ensuring higher quality addresses. Our experiments demonstrate that 6Diffusion-LM surpasses conventional methods, boasting a remarkable hit rate improvement to 43.53%
Huahu Xu, Ruiping Xing, Yiqin Gao, Jingkun Xu
HPCC4
2024 MI-GNN: Multi-Interaction GNN for Various Weak Information Learning on Graphs
abstract
Graph Neural Networks (GNNs) have achieved significant success in graph-related tasks, particularly in scenarios involving graphs with comprehensive information. Nevertheless, the performance of GNNs is often hindered in real-world applications due to the presence of incomplete graph data, characterized by fragmented structures, missing features, and insufficient labels. Previous research in this field has predominantly concentrated on augmenting one specific individual aspect of such weak information, thereby neglecting the holistic graph learning challenge. In this paper, we introduce a new framework, named MI-GNN, standing for Multi-Interaction Graph Neural Network, which is innovatively designed to facilitate message passing across graphs while adeptly harnessing the interplay and synergies among various types of information in graph learning tasks. This approach holistically integrates graph structures, features, and labels into a cohesive learning model, thus significantly enhancing graph learning capabilities for graphs with incomplete data. Furthermore, to achieve consistency in these enhanced tasks, our framework integrates multi-task learning to address the challenges of Graph Learning with Weak Information (GLWI). We adopt an adaptive alignment strategy, dynamically assigning weights to each learning task in every iteration. Empirical evaluation of eight publicly accessible datasets demonstrates that our MI-GNN framework achieves state-of-the-art performance in handling a range of weak information scenarios in graph learning.
Bowen Qiang, Huahu Xu, Jiangang Shi, Yiqin Gao
IJCNN6
2024 Minimizing Energy Consumption for Real-Time Tasks on Heterogeneous Platforms Under Deadline and Reliability Constraints
Yiqin Gao, Li Han 0001, Jing Liu 0012, Yves Robert, Frédéric Vivien
Algorithmica1
2023 Recovering a Molecule's 3D Dynamics from Liquid-phase Electron Microscopy Movies
abstract
The dynamics of biomolecules are crucial for our understanding of their functioning in living systems. However, current 3D imaging techniques, such as cryogenic electron microscopy (cryo-EM), require freezing the sample, which limits the observation of their conformational changes in real time. The innovative liquid-phase electron microscopy (liquid-phase EM) technique allows molecules to be placed in the native liquid environment, providing a unique opportunity to observe their dynamics. In this paper, we propose TEMPOR, a Temporal Electron MicroscoPy Object Reconstruction algorithm for liquid-phase EM that leverages an implicit neural representation (INR) and a dynamical variational auto-encoder (DVAE) to recover time series of molecular structures. We demonstrate its advantages in recovering different motion dynamics from two simulated datasets, 7bcq & Cas9. To our knowledge, our work is the first attempt to directly recover 3D structures of a temporally-varying particle from liquid-phase EM movies. It provides a promising new approach for studying molecules’ 3D dynamics in structural biology.
Enze Ye, Hong Zhang 0053, Yiqin Gao, He Sun 0010
ICCV4
2023 Resource-Constrained Scheduling Algorithms for Stochastic Independent Tasks With Unknown Probability Distribution
Yiqin Gao, Yves Robert, Frédéric Vivien
Algorithmica1
2023 Dynamic Scheduling Strategies for Firm Semi-Periodic Real-Time Tasks
abstract
This paper introduces and assesses novel strategies to schedule firm semi-periodic real-time tasks. Jobs are released periodically and have the same relative deadline. Job execution times obey an arbitrary probability distribution and can take either bounded or unbounded values. We investigate several optimization criteria, the most prominent being theDeadline Miss Ratio(DMR). All previous work uses some admission policies but never interrupt the execution of an admitted job before its deadline. On the contrary, we introduce three new control parameters to dynamically decide whether to interrupt a job at any given time. We derive a Markov model and use its stationary distribution to determine the best value of each control parameter. Finally we conduct an extensive simulation campaign with 16 different probability distributions. The results nicely demonstrate how the new strategies help improve system performance compared with traditional approaches. In particular, we show that (i) compared to pre-execution admission rules, the control parameters make significantly better decisions; (ii) specifically, the key control parameter is to upper bound the waiting time of each job; (iii) the best scheduling strategy decreases theDMRby up to 0.35 over traditional competitors.
Yiqin Gao, Guillaume Pallez, Yves Robert, Frédéric Vivien
IEEE Trans. Computers1
2021 Work-in-Progress: Evaluating Task Dropping Strategies for Overloaded Real-Time Systems
abstract
This paper discusses evaluation criteria and scheduling strategies for the analysis of overloaded real-time systems. This work builds upon techniques from queueing theory and proposes a new approach for real-time systems.
Yiqin Gao, Guillaume Pallez, Yves Robert, Frédéric Vivien
RTSS1
2020 Energy-aware strategies for reliability-oriented real-time task allocation on heterogeneous platforms
abstract
Low energy consumption and high reliability are widely identified as increasingly relevant issues in real-time systems on heterogeneous platforms. In this paper, we propose a multi-criteria optimization strategy to minimize the expected energy consumption while enforcing the reliability threshold and meeting all task deadlines. The tasks are replicated to ensure a prescribed reliability threshold. The platforms are composed of processors with different (and possibly unrelated) characteristics, including speed profile, energy cost and failure rate. We provide several mapping and scheduling heuristics towards this challenging optimization problem. Specifically, a novel approach is designed to control (i) how many replicas to use for each task, (ii) on which processor to map each replica and (iii) when to schedule each replica on its assigned processor. Different mappings achieve different levels of reliability and consume different amounts of energy. Scheduling matters because once a task replica is successful, the other replicas of that task are cancelled, which calls for minimizing the amount of temporal overlap between any replica pair. The experiments are conducted for a comprehensive set of execution scenarios, with a wide range of processor speed profiles and failure rates. The comparison results reveal that our strategies perform better than the random baseline, with a gain of 40% in energy consumption, for nearly all cases. The absolute performance of the heuristics is assessed by a comparison with a lower bound; the best heuristics achieve an excellent performance, with an average value only 4% higher than the lower bound.
Li Han 0001, Yiqin Gao, Jing Liu 0012, Yves Robert, Frédéric Vivien
ICPP2
2019 Scheduling independent stochastic tasks on heterogeneous cloud platforms
abstract
This work introduces scheduling strategies to maximize the expected number of independent tasks that can be executed on a cloud platform within a given budget and under a deadline constraint. The cloud platform is composed of several types of virtual machines (VMs), where each type has a unit execution cost that depends upon its characteristics. The amount of budget spent during the execution of a task on a given VM is the product of its execution length by the unit execution cost of that VM. The execution lengths of tasks follow a variety of standard probability distributions (exponential, uniform, half-normal, etc.), which is known beforehand and whose mean and standard deviation both depend upon the VM type. Finally, there is a global available budget and a deadline constraint, and the goal is to successfully execute as many tasks as possible before the deadline is reached or the budget is exhausted (whichever comes first). On each VM, the scheduler can decide at any instant to interrupt the execution of a (long) running task and to launch a new one, but the budget already spent for the interrupted task is lost. The main questions are which VMs to enroll, and whether and when to interrupt tasks that have been executing for some time. We assess the complexity of the problem by showing its NP-completeness and providing a 2-approximation for the asymptotic case where budget and deadline both tend to infinity. Then we introduce several heuristics and compare their performance by running an extensive set of simulations.
Yiqin Gao, Louis-Claude Canon, Yves Robert, Frédéric Vivien
CLUSTER1