Songhao Jiang

dblp:279/6527 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Computer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Harnessing Neural Collapse to Understand and Strengthen Classification Model Generalization
Yanlinhuan Zhang, Songhao Jiang, Yan Chu
KSEM (5)2
2025 Global Eye: Breaking the "Fixed Thinking Pattern" during the Instruction Expansion Process
abstract
An extensive high-quality instruction dataset is crucial for the instruction tuning process of Large Language Models (LLMs).Recent instruction expansion methods have demonstrated their capability to improve the quality and quantity of existing datasets, by prompting high-performance LLM to generate multiple new instructions from the original ones.However, existing methods focus on constructing multi-perspective prompts (e.g., increasing complexity or difficulty) to expand instructions, overlooking the "Fixed Thinking Pattern" issue of LLMs.This issue arises when repeatedly using the same set of prompts, causing LLMs to rely on a limited set of certain expressions to expand all instructions, potentially compromising the diversity of the final expanded dataset.This paper theoretically analyzes the causes of the "Fixed Thinking Pattern", and corroborates this phenomenon through multi-faceted empirical research.Furthermore, we propose a novel method based on dynamic prompt updating: Global Eye.Specifically, after a fixed number of instruction expansions, we analyze the statistical characteristics of newly generated instructions and then update the prompts.Experimental results show that our method enables LLaMA3-8B and LLaMA2-13B to surpass the performance of open-source LLMs and GPT3.5 across various metrics.
Wenxuan Lu, Wei Liu 0302, Jian Luan 0001, Bin Wang 0004, Songhao Jiang, Tianning Zang
ACL (1)5
2024 Entropy-Reinforced Planning with Large Language Models for Drug Discovery
abstract
The objective of drug discovery is to identify chemical compounds that possess specific pharmaceutical properties toward a binding target. Existing large language models (LLMS) can achieve high token matching scores in terms of likelihood for molecule generation. However, relying solely on LLM decoding often results in the generation of molecules that are either invalid due to a single misused token, or suboptimal due to unbalanced exploration and exploitation as a consequence of the LLM’s prior experience. Here we propose ERP, Entropy-Reinforced Planning for Transformer Decoding, which employs an entropy-reinforced planning algorithm to enhance the Transformer decoding process and strike a balance between exploitation and exploration. ERP aims to achieve improvements in multiple properties compared to direct sampling from the Transformer. We evaluated ERP on the SARS-CoV-2 virus (3CLPro) and human cancer cell target protein (RTCB) benchmarks and demonstrated that, in both benchmarks, ERP consistently outperforms the current state-of-the-art algorithm by 1-5 percent, and baselines by 5-10 percent, respectively. Moreover, such improvement is robust across Transformer models trained with different objectives. Finally, to further illustrate the capabilities of ERP, we tested our algorithm on three code generation benchmarks and outperformed the current state-of-the-art approach as well. Our code is publicly available at: https://github.com/xuefeng-cs/ERP.
Chih-chan Tien, Songhao Jiang, Rick L. Stevens
ICML4
2024 Meta-pruning: Learning to Prune on Few-Shot Learning
Yan Chu 0001, Keshi Liu, Songhao Jiang, Xianghui Sun, Baoxu Wang, Zhengkui Wang
KSEM (1)3
2024 Vicinal Data Augmentation for Classification Model via Feature Weaken
Songhao Jiang, Yan Chu 0001, Tianxing Ma, Xiaochen Miao, Zhengkui Wang, Tianning Zang
KSEM (1)1
2023 Imbalanced Few-Shot Learning Based on Meta-transfer Learning
Yan Chu 0001, Xianghui Sun, Songhao Jiang, Tianwen Xie, Zhengkui Wang, Wen Shan
ICANN (8)3
2023 Explainable Text Classification via Attentive and Targeted Mixing Data Augmentation
abstract
Mixing data augmentation methods have been widely used in text classification recently. However, existing methods do not control the quality of augmented data and have low model explainability. To tackle these issues, this paper proposes an explainable text classification solution based on attentive and targeted mixing data augmentation, ATMIX. Instead of selecting data for augmentation without control, ATMIX focuses on the misclassified training samples as the target for augmentation to better improve the model's capability. Meanwhile, to generate meaningful augmented samples, it adopts a self-attention mechanism to understand the importance of the subsentences in a text, and cut and mix the subsentences between the misclassified and correctly classified samples wisely. Furthermore, it employs a novel dynamic augmented data selection framework based on the loss function gradient to dynamically optimize the augmented samples for model training. In the end, we develop a new model explainability evaluation method based on subsentence attention and conduct extensive evaluations over multiple real-world text datasets. The results indicate that ATMIX is more effective with higher explainability than the typical classification models, hidden-level, and input-level mixup models.
Songhao Jiang, Yan Chu 0001, Zhengkui Wang, Tianxing Ma, Wenxuan Lu, Tianning Zang
IJCAI1
2022 Malicious Blockchain Domain Detection Based on Heterogeneous Information Network
abstract
With the popularity of Blockchain Domain Name System (BDNS), more and more cybercriminals integrate Blockchain Domain Names (BDNs) into their infrastructure. Due to the anonymity and anti-censorship of BDNs, it is difficult to detect malicious activities based on BDNs, posing a serious threat to network security. In this paper, we propose a novel method to detect malicious BDNs. First, we extract 16 statistical features of domain names. Second, we construct a Heterogeneous Information Network (HIN) of BDNS, which can use malicious traditional domain names as supplementary data. Then we associate domain names by meta-paths in the HIN and build an association graph of domain names. To better characterize domain names, we use the graph convolutional network algorithm to fuse the statistical features of domain names in the association graph. Finally, we detect malicious BDNs by the neural network algorithm. Compared with the existing methods, the experimental results show that our method can accurately detect malicious BDNs with the F1 score of 0.9901 and discover more unknown malicious BDNs from the dataset.
Songhao Jiang, Tianning Zang
GLOBECOM3
2021 Learning curves for drug response prediction in cancer cell lines
abstract
BACKGROUND: Motivated by the size and availability of cell line drug sensitivity data, researchers have been developing machine learning (ML) models for predicting drug response to advance cancer treatment. As drug sensitivity studies continue generating drug response data, a common question is whether the generalization performance of existing prediction models can be further improved with more training data. METHODS: We utilize empirical learning curves for evaluating and comparing the data scaling properties of two neural networks (NNs) and two gradient boosting decision tree (GBDT) models trained on four cell line drug screening datasets. The learning curves are accurately fitted to a power law model, providing a framework for assessing the data scaling behavior of these models. RESULTS: The curves demonstrate that no single model dominates in terms of prediction performance across all datasets and training sizes, thus suggesting that the actual shape of these curves depends on the unique pair of an ML model and a dataset. The multi-input NN (mNN), in which gene expressions of cancer cells and molecular drug descriptors are input into separate subnetworks, outperforms a single-input NN (sNN), where the cell and drug features are concatenated for the input layer. In contrast, a GBDT with hyperparameter tuning exhibits superior performance as compared with both NNs at the lower range of training set sizes for two of the tested datasets, whereas the mNN consistently performs better at the higher range of training sizes. Moreover, the trajectory of the curves suggests that increasing the sample size is expected to further improve prediction scores of both NNs. These observations demonstrate the benefit of using learning curves to evaluate prediction models, providing a broader perspective on the overall data scaling characteristics. CONCLUSIONS: A fitted power law learning curve provides a forward-looking metric for analyzing prediction performance and can serve as a co-design tool to guide experimental biologists and computational scientists in the design of future experiments in prospective research studies.
Alexander Partin, Thomas S. Brettin, Yvonne A. Evrard, Yitan Zhu, Hyun Seung Yoo, Fangfang Xia, Songhao Jiang, Austin Clyde, Maulik Shukla, Michael Fonstein, James H. Doroshow, Rick L. Stevens
BMC Bioinform.7