Weiran Lin

dblp:68/4713 · DBLP profile ↗
← Back
13ranked-venue papers
5as first author
9since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Security and privacy · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Edge-Optimized Voice Control with 0.26 M Parameters: Distilling 86M Adaptive Window Audio Transformer for Real-World Variable-Length Inputs
Pinze Ren, Zhen Chen 0001, Yinjun Wu, Weiran Lin, Qilong Shi, Chao Li 0012, Jianxin Yang
IEEE Big Data4
2025 LLM Whisperer: An Inconspicuous Attack to Bias LLM Responses
Weiran Lin, Anna Gerchanovsky, Omer Akgul, Lujo Bauer, Matt Fredrikson, Zifan Wang 0001
CHI1
2025 Estimating LLM Consistency: A User Baseline vs Surrogate Metrics
abstract
Large language models (LLMs) are prone to hallucinations and sensitive to prompt perturbations, often resulting in inconsistent or unreliable generated text.Different methods have been proposed to mitigate such hallucinations and fragility, one of which is to measure the consistency of LLM responses-the model's confidence in the response or likelihood of generating a similar response when resampled.In previous work, measuring LLM response consistency often relied on calculating the probability of a response appearing within a pool of resampled responses, analyzing internal states, or evaluating logits of resopnses.However, it was not clear how well these approaches approximated users' perceptions of consistency of LLM responses.To find out, we performed a user study (n = 2, 976) demonstrating that current methods for measuring LLM response consistency typically do not align well with humans' perceptions of LLM consistency.We propose a logit-based ensemble method for estimating LLM consistency and show that our method matches the performance of the bestperforming existing metric in estimating human ratings of LLM consistency.Our results suggest that methods for estimating LLM consistency without human evaluation are sufficiently imperfect to warrant broader use of evaluation with human input; this would avoid misjudging the adequacy of models because of the imperfections of automated consistency metrics.
Xiaoyuan Wu, Weiran Lin, Omer Akgul, Lujo Bauer
EMNLP2
2024 An Empirical Study on the Power Consumption of LLMs with Different GPU Platforms
abstract
This paper researches on the power consumption of AIGC applications based on LLM with different parameter scales across different hardware platforms. Artificial Intelligence Generated Content (AIGC) represents a leading-edge application of AI technology, primarily driven by large language models (LLMs) and their associated technologies. The deployment of LLM typically relies on critical facilities with three layers, i.e., the hardware, model, and application layers. This empirical study aims to identify key factors in power consumption when a large model is serving in the inference stage, which will hint the insights for improving the energy efficiency of computational infrastructures. In the context of the "dual carbon" goals, i.e., carbon peaking and carbon neutrality, this study aims to find an effective way to reduce the energy cost of AIGC applications, thereby supporting sustainable AI development in industry.
Zhen Chen 0001, Weiran Lin, Xinyu Xie, Yaodong Hu, Chao Li 0012, Qiaojuan Tong, Yinjun Wu, Shuangshou Li
IEEE Big Data2
2024 Training Robust ML-based Raw-Binary Malware Detectors in Hours, not Months
abstract
Machine-learning (ML) classifiers are increasingly used to distinguish malware from benign binaries. Recent work has shown that ML-based detectors can be evaded by adversarial examples, but also that one may defend against such attacks via adversarial training. However, adversarial training, and subsequent robustness evaluation, is computationally expensive in the raw-binary malware-detection domain because it requires producing many adversarial examples for both training and evaluation. Prior work found that Greedy-training, a faster robust training technique that forgoes using adversarial examples, showed some promise in producing robust malware detectors. However, Greedy-training was far less effective in inducing robustness than the more expensive adversarial training, and it also severely hurt natural accuracy (i.e., accuracy on the original data). To faster train models, this work presents GreedyBlock-training, an enhanced version of Greedy-training that we empirically show achieves not only state-of-the-art robustness in malware detectors, exceeding even adversarial training, but also retains natural accuracy better than adversarial training. Furthermore, as it does not require creating adversarial (or functional) examples, GreedyBlock-training is significantly faster than adversarial training. Specifically, we show that GreedyBlock-training can produce more robust (+54% on average), more naturally accurate (+7% on average), and more efficiently trained (-91% average computation) malware detectors than prior work. To faster evaluate models, we also develop methods to faster gauge the robustness of ML-based raw-binary malware detectors by introducing robustness proxies, which can be used either to predict which models are likely to be the most robust, thus helping prioritize which detectors to evaluate with expensive attacks, or aiding in deciding which detectors are worthwhile to continue training. Experimentally, we show these proxy measures can find the most robust detector in a pool of detectors while using only ~20-50% of the computation that would otherwise be required.
Keane Lucas, Weiran Lin, Lujo Bauer, Michael K. Reiter, Mahmood Sharif
CCS2
2024 The WMDP Benchmark: Measuring and Reducing Malicious Use with Unlearning
abstract
The White House Executive Order on Artificial Intelligence highlights the risks of large language models (LLMs) empowering malicious actors in developing biological, cyber, and chemical weapons. To measure these risks, government institutions and major AI labs are developing evaluations for hazardous capabilities in LLMs. However, current evaluations are private and restricted to a narrow range of malicious use scenarios, which limits further research into reducing malicious use. To fill these gaps, we release the Weapons of Mass Destruction Proxy (WMDP) benchmark, a dataset of 3,668 multiple-choice questions that serve as a proxy measurement of hazardous knowledge in biosecurity, cybersecurity, and chemical security. To guide progress on unlearning, we develop RMU, a state-of-the-art unlearning method based on controlling model representations. RMU reduces model performance on WMDP while maintaining general capabilities in areas such as biology and computer science, suggesting that unlearning may be a concrete path towards reducing malicious use from LLMs. We release our benchmark and code publicly at https://wmdp.ai.
Nathaniel Li, Alexander Pan, Anjali Gopal, Summer Yue, Daniel Berrios, Alice Gatti, Justin D. Li, Ann-Kathrin Dombrowski, Shashwat Goel, Gabriel Mukobi, Nathan Helm-Burger, Rassin Lababidi, Lennart Justen, Andrew B. Liu, Isabelle Barrass, Oliver Zhang, Xiaoyuan Zhu, Rishub Tamirisa, Bhrugu Bharathi, Ariel Herbert-Voss, Cort B. Breuer, Andy Zou, Mantas Mazeika, Zifan Wang 0001, Palash Oswal, Weiran Lin, Adam A. Hunt, Justin Tienken-Harder, Kevin Y. Shih, Kemper Talley, John Guan, Ian Steneker, David Campbell, Brad Jokubaitis, Steven Basart, Stephen Fitz, Ponnurangam Kumaraguru, Kallol Krishna Karmakar, Udaya Kiran Tupakula, Vijay Varadharajan, Yan Shoshitaishvili, Jimmy Ba, Kevin M. Esvelt, Alexandr Wang, Dan Hendrycks
ICML27
2024 Group-based Robustness: A General Framework for Customized Robustness in the Real World
Weiran Lin, Keane Lucas, Neo Eyal, Lujo Bauer, Michael K. Reiter, Mahmood Sharif
NDSS1
2023 Adversarial Training for Raw-Binary Malware Classifiers
Keane Lucas, Samruddhi Pai, Weiran Lin, Lujo Bauer, Michael K. Reiter, Mahmood Sharif
USENIX Security Symposium3
2022 Constrained Gradient Descent: A Powerful and Principled Evasion Attack Against Neural Networks
abstract
We propose new, more efficient targeted white-box attacks against deep neural networks. Our attacks better align with the attacker’s goal: (1) tricking a model to assign higher probability to the target class than to any other class, while (2) staying within an $\epsilon$-distance of the attacked input. First, we demonstrate a loss function that explicitly encodes (1) and show that Auto-PGD finds more attacks with it. Second, we propose a new attack method, Constrained Gradient Descent (CGD), using a refinement of our loss function that captures both (1) and (2). CGD seeks to satisfy both attacker objectives—misclassification and bounded $\ell_{p}$-norm—in a principled manner, as part of the optimization, instead of via ad hoc post-processing techniques (e.g., projection or clipping). We show that CGD is more successful on CIFAR10 (0.9–4.2%) and ImageNet (8.6–13.6%) than state-of-the-art attacks while consuming less time (11.4–18.8%). Statistical tests confirm that our attack outperforms others against leading defenses on different datasets and values of $\epsilon$.
Weiran Lin, Keane Lucas, Lujo Bauer, Michael K. Reiter, Mahmood Sharif
ICML1
2018 Predicting drug-disease associations by using similarity constrained matrix factorization
abstract
BACKGROUND: Drug-disease associations provide important information for the drug discovery. Wet experiments that identify drug-disease associations are time-consuming and expensive. However, many drug-disease associations are still unobserved or unknown. The development of computational methods for predicting unobserved drug-disease associations is an important and urgent task. RESULTS: In this paper, we proposed a similarity constrained matrix factorization method for the drug-disease association prediction (SCMFDD), which makes use of known drug-disease associations, drug features and disease semantic information. SCMFDD projects the drug-disease association relationship into two low-rank spaces, which uncover latent features for drugs and diseases, and then introduces drug feature-based similarities and disease semantic similarity as constraints for drugs and diseases in low-rank spaces. Different from the classic matrix factorization technique, SCMFDD takes the biological context of the problem into account. In computational experiments, the proposed method can produce high-accuracy performances on benchmark datasets, and outperform existing state-of-the-art prediction methods when evaluated by five-fold cross validation and independent testing. CONCLUSION: We developed a user-friendly web server by using known associations collected from the CTD database, available at http://www.bioinfotech.cn/SCMFDD/ . The case studies show that the server can find out novel associations, which are not included in the CTD database.
Wen Zhang 0008, Xiang Yue, Weiran Lin, Wenjian Wu, Ruoqi Liu, Feng Huang 0004
BMC Bioinform.3
2017 Predicting drug-disease associations based on the known association bipartite network
abstract
Recent studies show that drug-disease associations provide important information for drug discovery and drug repositioning. Wet experimental identification of drug-disease associations is time-consuming and labor-intensive. Therefore, the development of computational methods that predict drug-disease associations is an urgent task. In this paper, we propose a novel computational method named NTSIM, which only uses known drug-disease associations to predict unobserved associations. First of all, known drug-disease associations are represented as a drug-disease bipartite network, and a novel similarity measure named linear neighborhood similarity (LNS) is proposed to calculate drug-drug similarity and disease-disease similarity based on the bipartite network. Then, we predict unobserved drug-disease associations in the similarity-based graph by using label propagation process. In the computational experiments, this proposed method achieves high-accuracy performances, and outperforms representative state-of-the-art methods: PREDICT, TL-HGBI and LRSSL. Our studies reveal that known drug-disease associations can provide enough information to build the high-accuracy prediction models; linear neighbor similarity (LNS) can lead to better performances than other similarity measures such as Jaccard similarity, Gauss similarity and cosine similarity; the bipartite network-derived features outperform the drug biological features and disease semantic features.
Wen Zhang 0008, Xiang Yue, Yanlin Chen 0002, Weiran Lin, Bolin Li, Xiaohong Li 0003
BIBM4
2001 An 8.0-/8.4-kbps wideband speech coder based on mixed excitation linear prediction
Weiran Lin, Soo Ngee Koh, Xiao Lin 0001
Signal Process.1
2000 Mixed excitation linear prediction coding of wideband speech at 8 kbps
abstract
This paper presents our study on the feasibility and effectiveness of using the MELP (mixed excitation linear prediction) model for coding wideband (7 kHz) speech signals at a transmission bit rate of 8 kbps. In order to achieve a reasonably good subjective quality for the decoded speech while maintaining a low operating bit rate at the same time, modifications to the pitch estimation, LP analysis/synthesis and post filtering stages of the original MELP model are discussed. Informal listening tests show that the subjective quality of the decoded speech of the proposed coder is rated to be slightly better than the MPEG-4 CELP coder operating at 14.4 kbps for both male and female utterances. The subjective quality of the decoded female utterances from the proposed coder operating at 8.4 kbps is rated to be comparable to that produced by the ITU G.722 coder operating at 48 kbps.
Weiran Lin, Soo Ngee Koh, Xiao Lin 0001
ICASSP1