Yanda Li

dblp:23/1588 · DBLP profile ↗
← Back
63ranked-venue papers
3as first author
11since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 24 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 13 · 1 since 2021Computer networks · 6 · 1 first-authorHuman-computer interaction and ubiquitous computing · 3 · 1 since 2021Systems, architecture and hardware · 2Security and privacy · 2 · 1 since 2021Theory of computation · 2
YearPublicationVenuePosition
2025 AppAgent: Multimodal Agents as Smartphone Users
Chi Zhang 0007, Zhao Yang 0002, Yanda Li, Yucheng Han, Xin Chen 0040, Zebiao Huang, Gang Yu 0002
CHI4
2025 Adaptive Mobile Agent for Dynamic Interactions
abstract
With the rise of Multimodal Large Language Models (MLLM), LLM-driven visual agents are transforming software interfaces, especially those with graphical user interfaces. However, existing methods often struggle with diverse and complex mobile environments, such as rapidly changing app interfaces or non-standard UI components, limiting their adaptability and precision. This work presents a novel LLM-based multimodal agent framework for mobile devices, designed to enhance interaction and adaptive capabilities in dynamic mobile environments. By autonomously navigating devices and emulating human-like behaviors, the agent integrates parsing, text, and vision descriptions to construct a flexible action space. During the exploration phase, functionalities of user interface elements are documented into a customized structured knowledge base. In the deployment phase, RAG technology enables efficient retrieval and updates from this knowledge base. Experimental results across multiple benchmarks validate the framework's superior performance and practical effectiveness.
Yanda Li, Chi Zhang 0007, Wenjia Jiang, Wanqi Yang, Xin Chen 0040, Ling Chen 0006, Yunchao Wei
ICME1
2025 DART: Distribution-Aware Hardware Trojan Detection
abstract
Machine Learning (ML) has proven effective in Integrated Circuits (IC) security, particularly in Hardware Trojan (HT) detection. However, a model’s generalization potential depends on its ability to address distribution shifts (DS) in unseen data. Mitigating DS enhances a model’s adaptability to novel variations and threats within the dynamic realm of IC designs and HTs. We formulate HT detection as a DS problem, introducingDART, a novelDistribution-AwareHT detection framework, to enhance model generalization. ApplyingDARTon state-of-the-art Graph Neural Network architecture yields up to 22.96% and 17.37% F1-score improvements for unseen IC designs diverging significantly from the training data.
Youssef Gamal, Yanda Li, Shih-Yuan Yu, Ihsen Alouani, Mohammad Abdullah Al Faruque
IEEE Trans. Inf. Forensics Secur.3
2024 Continual Learning for Temporal-Sensitive Question Answering
abstract
In this study, we explore an emerging research area of Continual Learning for Temporal Sensitive Question Answering (CLTSQA). Previous research has primarily focused on Temporal Sensitive Question Answering (TSQA), often overlooking the unpredictable nature of future events. In real-world applications, it’s crucial for models to continually acquire knowledge over time, rather than relying on a static, complete dataset. Our paper investigates strategies that enable models to adapt to the ever-evolving information landscape, thereby addressing the challenges inherent in CLTSQA. To support our research, we first create a novel dataset, divided into five subsets, designed specifically for various stages of continual learning. We then propose a training framework for CLTSQA that integrates temporal memory replay and temporal contrastive learning. Our experimental results highlight two significant insights: First, the CLTSQA task introduces unique challenges for existing models. Second, our proposed framework effectively navigates these challenges, resulting in improved performance.
Wanqi Yang, Yunqiu Xu, Yanda Li, Kunze Wang, Binbin Huang 0006, Ling Chen 0006
IJCNN3
2024 Disentangled Pre-training for Image Matting
abstract
Image matting requires high-quality pixel-level human annotations to support the training of a deep model in recent literature. Whereas such annotation is costly and hard to scale, significantly holding back the development of the research. In this work, we make the first attempt towards addressing this problem, by proposing a self-supervised pretraining approach that can leverage infinite numbers of data to boost the matting performance. The pre-training task is designed in a similar manner as image matting, where random trimap and alpha matte are generated to achieve an image disentanglement objective. The pre-trained model is then used as an initialisation of the downstream matting task for fine-tuning. Extensive experimental evaluations show that the proposed approach outperforms both the state-of-the-art matting methods and other alternative self-supervised initialisation approaches by a large margin. We also show the robustness of the proposed approach over different backbone architectures. Our project page is available at https://crystraldo.github.io/dpt_mat/.
Yanda Li, Gang Yu 0002, Ling Chen 0006, Yunchao Wei, Jianbo Jiao
WACV1
2023 A Ground-Roll Separation Method Based on Neural Networks With Morphological Similarity Loss
abstract
Ground-roll is a typical Rayleigh-type interference noise in field seismic data, which is characterized by low frequency, low velocity and high amplitude. Since it will interfere effective seismic signals and severely degrade the signal-to-noise ratio of observed seismic records, many approaches have been developed for ground-roll attenuation or separation. In this letter, we proposed an improved ground-roll separation algorithm through the combination of deep learning based low-frequency generation and dictionary learning based low-frequency reconstruction. Moreover, to utilize the inter-band morphological similarity prior in seismic response, we introduce the morphological similarity constraint into the learning approach of pseudo low-frequency generation networks. Experiments demonstrate that compared to previous methods, the introduced the morphological similarity loss can effectively improve the quality of generated pseudo low-frequency signals, which results in better low-frequency reflection reconstruction and ground-roll separation performances.
Xingyu Tian, Yile Ao, Yanda Li, Wenkai Lu
IEEE Geosci. Remote. Sens. Lett.3
2023 Improved Seismic Residual Diffracted Multiple Suppression Method Based on Object Detection and Image Segmentation
abstract
Seismic multiple is one of the most common noises in marine seismic data, which heavily affects subsequent processing and interpretation. To eliminate the influence of seismic multiples, many methods have been developed, while surface-related multiple elimination (SRME) is one of the most widely deployed methods. However, results of SRME always contain a few strong residual diffracted multiples (RDMs) in practice because of the unprecise prediction of diffracted multiples compared to reflection multiples. If we try to apply further multiple suppression methods to SRME results, it not only tends to damage the signals, but also spends lots of unnecessary computations where there is no RDM. In this article, we propose an improved RDM suppression method based on object detection and image segmentation. First, we employ an object detection network to locate bounding boxes containing RDMs in the SRME results. Then a threshold-based image segmentation method is utilized to identify regions of strong RDMs in the detected boxes. According to the segmentation results, parameters for weak multiples and strong multiples are provided for the adaptive multiple subtraction (AMS) in different regions to generate different results. At last, we combine the suppression results of strong RDMs and weak RDMs as the final results. Application on field data demonstrates that our method is able to suppress RDMs with little loss of signal.
Xingyu Tian, Wenkai Lu, Yanda Li, Mingrui Zhong, Hongxun Pan, Bowu Jiang
IEEE Trans. Geosci. Remote. Sens.3
2022 Decoding multilevel relationships with the human tissue-cell-molecule network
abstract
Understanding the biological functions of molecules in specific human tissues or cell types is crucial for gaining insights into human physiology and disease. To address this issue, it is essential to systematically uncover associations among multilevel elements consisting of disease phenotypes, tissues, cell types and molecules, which could pose a challenge because of their heterogeneity and incompleteness. To address this challenge, we describe a new methodological framework, called Graph Local InfoMax (GLIM), based on a human multilevel network (HMLN) that we established by introducing multiple tissues and cell types on top of molecular networks. GLIM can systematically mine the potential relationships between multilevel elements by embedding the features of the HMLN through contrastive learning. Our simulation results demonstrated that GLIM consistently outperforms other state-of-the-art algorithms in disease gene prediction. Moreover, GLIM was also successfully used to infer cell markers and rewire intercellular and molecular interactions in the context of specific tissues or diseases. As a typical case, the tissue-cell-molecule network underlying gastritis and gastric cancer was first uncovered by GLIM, providing systematic insights into the mechanism underlying the occurrence and development of gastric cancer. Overall, our constructed methodological framework has the potential to systematically uncover complex disease mechanisms and mine high-quality relationships among phenotypical, tissue, cellular and molecular elements.
Siyu Hou, Peng Zhang 0149, Kuo Yang 0001, Changzheng Ma, Yanda Li, Shao Li
Briefings Bioinform.6
2022 Evaluating methylation of human ribosomal DNA at each CpG site reveals its utility for cancer detection using cell-free DNA
abstract
Ribosomal deoxyribonucleic acid (DNA) (rDNA) repeats are tandemly located on five acrocentric chromosomes with up to hundreds of copies in the human genome. DNA methylation, the most well-studied epigenetic mechanism, has been characterized for most genomic regions across various biological contexts. However, rDNA methylation patterns remain largely unexplored due to the repetitive structure. In this study, we designed a specific mapping strategy to investigate rDNA methylation patterns at each CpG site across various physiological and pathological processes. We found that CpG sites on rDNA could be categorized into two types. One is within or adjacent to transcribed regions; the other is distal to transcribed regions. The former shows highly variable methylation levels across samples, while the latter shows stable high methylation levels in normal tissues but severe hypomethylation in tumors. We further showed that rDNA methylation profiles in plasma cell-free DNA could be used as a biomarker for cancer detection. It shows good performances on public datasets, including colorectal cancer [area under the curve (AUC) = 0.85], lung cancer (AUC = 0.84), hepatocellular carcinoma (AUC = 0.91) and in-house generated hepatocellular carcinoma dataset (AUC = 0.96) even at low genome coverage (<1×). Taken together, these findings broaden our understanding of rDNA regulation and suggest the potential utility of rDNA methylation features as disease biomarkers.
Xianglin Zhang, Bixi Zhong, Lei Wei 0009, Jiaqi Li 0025, Wei Zhang 0241, Huan Fang 0003, Yanda Li, Yinying Lu, Xiaowo Wang
Briefings Bioinform.8
2022 DanmuVis: Visualizing Danmu Content Dynamics and Associated Viewer Behaviors in Online Videos
abstract
Abstract Danmu (Danmaku) is a unique social media service in online videos, especially popular in Japan and China, for viewers to write comments while watching videos. The danmu comments are overlaid on the video screen and synchronized to the associated video time, indicating viewers' thoughts of the video clip. This paper introduces an interactive visualization system to analyze danmu comments and associated viewer behaviors in a collection of videos and enable detailed exploration of one video on demand. The watching behaviors of viewers are identified by comparing video time and post time of viewers' danmu. The system supports analyzing danmu content and viewers' behaviors against both video time and post time to gain insights into viewers' online participation and perceived experience. Our evaluations, including usage scenarios and user interviews, demonstrate the effectiveness and usability of our system.
Shuai Chen 0001, Yanda Li, Juanjuan Long, Siming Chen 0001, Jiawan Zhang, Xiaoru Yuan
Comput. Graph. Forum3
2022 Improved Anomalous Amplitude Attenuation Method Based on Deep Neural Networks
abstract
In seismic exploration, seismic data usually contain anomalous amplitude noise whose high energy may affect the results of subsequent processing steps. In industry, this kind of noise is generally suppressed using the anomalous amplitude attenuation (AAA) method. The AAA method essentially suppresses abnormal amplitude noise using a median filter in the time–frequency domain. This makes its performance heavily dependent on the parameters, especially the window width of the median filter. Thus, we propose an improved anomalous amplitude attenuation (IAAA) method based on deep neural networks. The IAAA method contains two steps. In the first step, deep neural networks are used to detect the locations and the widths of noise regions. In the second step, the noise information (locations and widths) obtained at the previous step is exploited to apply the AAA method with more appropriate parameters to each noisy region. Compared with the conventional AAA method, the IAAA method can suppress the noise more effectively and preserve signals better. The experiments on both synthetic data and field data demonstrate that our method outperforms the conventional AAA method.
Xingyu Tian, Wenkai Lu, Yanda Li
IEEE Trans. Geosci. Remote. Sens.3
2020 LBVis: Interactive Dynamic Load Balancing Visualization for Parallel Particle Tracing
abstract
We propose an interactive visual analytical approach to exploring and diagnosing the dynamic load balance (data and task partition) process of parallel particle tracing in flow visualization. To understand the complex nature of the parallel processes, it is necessary to integrate the information of the behaviors and patterns of the computing processes, data changes and movements, task status and exchanges, and gain the insight of the relationships among them. In our proposed approach, the data and task behaviors are visualized through a graph with a fine-designed layout, in which node glyphs are dedicated to showing the status of processes and the links represent the data or task transfer between different computation rounds and processes. User interactions are supported to facilitate the exploration of performance analysis. We provide a case study to demonstrate that the proposed approach enables users to identify the bottlenecks during this process, and thus help optimize the related algorithms.
Jiang Zhang 0002, Changhe Yang, Yanda Li, Li Chen 0031, Xiaoru Yuan
PacificVis3
2018 DQN-Based Power Control for IoT Transmission against Jamming
abstract
Internet of Things (IoTs) have to address jammers, with goal to interrupt the communication of the energy- constrained IoT devices and sometimes even cause denial-of-service attacks. In this paper, we propose a deep reinforcement learning based power control scheme for IoT devices to improve the transmission efficiency and save energy. This scheme depends on the current IoT transmission status and the jamming strength and applies deep Q-network (DQN) to determine the transmit power without being aware of the IoT topology and the jamming model. This scheme is implemented on the universal software radio peripherals for the anti- jamming communication performance evaluation. Experimental results show that this scheme improves the signal-to-interference-plus-noise of the IoT signals compared with the benchmark Q-learning based power control scheme against jamming.
Ye Chen 0011, Yanda Li, Dongjin Xu, Liang Xiao 0003
VTC Spring2
2018 esATAC: an easy-to-use systematic pipeline for ATAC-seq data analysis
abstract
Summary: ATAC-seq is rapidly emerging as one of the major experimental approaches to probe chromatin accessibility genome-wide. Here, we present 'esATAC', a highly integrated easy-to-use R/Bioconductor package, for systematic ATAC-seq data analysis. It covers essential steps for full analyzing procedure, including raw data processing, quality control and downstream statistical analysis such as peak calling, enrichment analysis and transcription factor footprinting. esATAC supports one command line execution for preset pipelines and provides flexible interfaces for building customized pipelines. Availability and implementation: esATAC package is open source under the GPL-3.0 license. It is implemented in R and C++. Source code and binaries for Linux, MAC OS X and Windows are available through Bioconductor (https://www.bioconductor.org/packages/release/bioc/html/esATAC.html). Supplementary information: Supplementary data are available at Bioinformatics online.
Wei Zhang 0241, Huan Fang 0003, Yanda Li, Xiaowo Wang
Bioinform.4
2018 A Secure Mobile Crowdsensing Game With Deep Reinforcement Learning
abstract
Mobile crowdsensing (MCS) is vulnerable to faked sensing attacks, as selfish smartphone users sometimes provide faked sensing results to the MCS server to save their sensing costs and avoid privacy leakage. In this paper, the interactions between an MCS server and a number of smartphone users are formulated as a Stackelberg game, in which the server as the leader first determines and broadcasts its payment policy for each sensing accuracy. Each user as a follower chooses the sensing effort and thus the sensing accuracy afterward to receive the payment based on the payment policy and the sensing accuracy estimated by the server. The Stackelberg equilibria of the secure MCS game are presented, disclosing conditions to motivate accurate sensing. Without knowing the smartphone sensing models in a dynamic version of the MCS game, an MCS system can apply deep Q-network (DQN), which is a deep reinforcement learning technique combining reinforcement learning and deep learning techniques, to derive the optimal MCS policy against faked sensing attacks. The DQN-based MCS system uses a deep convolutional neural network to accelerate the learning process with a high-dimensional state space and action set, and thus improve the MCS performance against selfish users. Simulation results show that the proposed MCS system stimulates high-quality sensing services and suppresses faked sensing attacks, compared with a Q-learning-based MCS system.
Liang Xiao 0003, Yanda Li, Guoan Han, Huaiyu Dai, H. Vincent Poor
IEEE Trans. Inf. Forensics Secur.2
2017 Reinforcement Learning Based Mobile Offloading for Cloud-Based Malware Detection
abstract
Cloud-based malware detection improves the detection performance for mobile devices that offload their malware detection tasks to security servers with much larger malware database and powerful computational resources. In this paper, we investigate the competition of the radio transmission bandwidths and the data sharing of the security server in the dynamic malware detection game, in which each mobile device chooses its offloading rate of the application traces to the security server. As the Q-learning technique has a slow learning rate in the game with high dimension, we have designed a mobile malware detection based on hotbooting-Q techniques, which initiates the quality values based on the malware detection experience. We propose an offloading strategy based on deep Q-network technique with a deep convolutional neural network to further improve the detection speed, the detection accuracy, and the utility. Preliminary simulation results verify the detection gain of the scheme compared with the Q- learning based strategy.
Xiaoyue Wan, Geyi Sheng, Yanda Li, Liang Xiao 0003, Xiaojiang Du
GLOBECOM3
2017 Game theoretic study of protecting MIMO transmissions against smart attacks
abstract
Multiple-input multiple-output (MIMO) systems are threatened by smart attackers, who apply programmable radio devices such as software defined radios to perform multiple types of attacks such as eavesdropping, jamming and spoofing. In this paper, MIMO transmission in the presence of smart attacks is formulated as a noncooperative game, in which a MIMO transmitter chooses its transmit power level and a smart attacker determines its attack type accordingly. A Nash equilibrium of this secure MIMO transmission game is derived and conditions assuring its existence are provided to reveal the impact of the number of antennas and the costs of the attacker to launch each type of attack. A power control strategy based on Q-learning is proposed for the MIMO transmitter to suppress the attack motivation of smart attackers in a dynamic version of MIMO transmission game without being aware of the attack and the radio channel model. Simulation results show that our proposed scheme can reduce the attack rate of smart attackers and improve the secrecy capacity compared with the benchmark strategy.
Yanda Li, Liang Xiao 0003, Huaiyu Dai, H. Vincent Poor
ICC1
2017 Cloud-Based Malware Detection Game for Mobile Devices with Offloading
abstract
As accurate malware detection on mobile devices requires fast process of a large number of application traces, cloud-based malware detection can utilize the data sharing and powerful computational resources of security servers to improve the detection performance. In this paper, we investigate the cloud-based malware detection game, in which mobile devices offload their application traces to security servers via base stations or access points in dynamic networks. We derive the Nash equilibrium (NE) of the static malware detection game and present the existence condition of the NE, showing how mobile devices share their application traces at the security server to improve the detection accuracy, and compete for the limited radio bandwidth, the computational and communication resources of the server. We design a malware detection scheme with Q-learning for a mobile device to derive the optimal offloading rate without knowing the trace generation and the radio bandwidth model of other mobile devices. The detection performance is further improved with the Dyna architecture, in which a mobile device learns from the hypothetical experience to increase its convergence rate. We also design a post-decision state learning-based scheme that utilizes the known radio channel model to accelerate the reinforcement learning process in the malware detection. Simulation results show that the proposed schemes improve the detection accuracy, reduce the detection delay, and increase the utility of a mobile device in the dynamic malware detection game, compared with the benchmark strategy.
Liang Xiao 0003, Yanda Li, Xueli Huang, Xiaojiang Du
IEEE Trans. Mob. Comput.2
2016 Prospect Theoretic Study of Cloud Storage Defense against Advanced Persistent Threats
abstract
Cloud storage is vulnerable to Advanced Persistent Threats (APTs), which are stealthy, continuous, well funded and targeted. In this paper, prospect theory is applied to study the interactions between a subjective cloud storage defender and a subjective APT attacker. Two subjective APT games are formulated, in which the defender chooses its interval to scan the storage device and the attacker decides its duration between launching two attacks under uncertain APT attack durations and action of the opponent, respectively. The Nash equilibria of the static subjective APT games are derived. We also study the dynamic APT game and propose a Q-learning based APT defense strategy for cloud storage. Simulation results show that the APT defense benefits from the subjective view of the attacker and the proposed defense strategy can improve detection performance with a higher utility.
Dongjin Xu, Yanda Li, Liang Xiao 0003, Narayan B. Mandayam, H. Vincent Poor
GLOBECOM2
2015 CRISPR-ERA: a comprehensive design tool for CRISPR-mediated gene editing, repression and activation
abstract
UNLABELLED: The CRISPR/Cas9 system was recently developed as a powerful and flexible technology for targeted genome engineering, including genome editing (altering the genetic sequence) and gene regulation (without altering the genetic sequence). These applications require the design of single guide RNAs (sgRNAs) that are efficient and specific. However, this remains challenging, as it requires the consideration of many criteria. Several sgRNA design tools have been developed for gene editing, but currently there is no tool for the design of sgRNAs for gene regulation. With accumulating experimental data on the use of CRISPR/Cas9 for gene editing and regulation, we implement a comprehensive computational tool based on a set of sgRNA design rules summarized from these published reports. We report a genome-wide sgRNA design tool and provide an online website for predicting sgRNAs that are efficient and specific. We name the tool CRISPR-ERA, for clustered regularly interspaced short palindromic repeat-mediated editing, repression, and activation (ERA). AVAILABILITY AND IMPLEMENTATION: http://CRISPR-ERA.stanford.edu. CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Antonia Dominguez, Yanda Li, Xiaowo Wang, Lei S. Qi
Bioinform.4
2014 Inferring the perturbed microRNA regulatory networks from gene expression data using a network propagation based method
abstract
BACKGROUND: MicroRNAs (miRNAs) are a class of endogenous small regulatory RNAs. Identifications of the dys-regulated or perturbed miRNAs and their key target genes are important for understanding the regulatory networks associated with the studied cellular processes. Several computational methods have been developed to infer the perturbed miRNA regulatory networks by integrating genome-wide gene expression data and sequence-based miRNA-target predictions. However, most of them only use the expression information of the miRNA direct targets, rarely considering the secondary effects of miRNA perturbation on the global gene regulatory networks. RESULTS: We proposed a network propagation based method to infer the perturbed miRNAs and their key target genes by integrating gene expressions and global gene regulatory network information. The method used random walk with restart in gene regulatory networks to model the network effects of the miRNA perturbation. Then, it evaluated the significance of the correlation between the network effects of the miRNA perturbation and the gene differential expression levels with a forward searching strategy. Results show that our method outperformed several compared methods in rediscovering the experimentally perturbed miRNAs in cancer cell lines. Then, we applied it on a gene expression dataset of colorectal cancer clinical patient samples and inferred the perturbed miRNA regulatory networks of colorectal cancer, including several known oncogenic or tumor-suppressive miRNAs, such as miR-17, miR-26 and miR-145. CONCLUSIONS: Our network propagation based method takes advantage of the network effect of the miRNA perturbation on its target genes. It is a useful approach to infer the perturbed miRNAs and their key target genes associated with the studied biological processes using gene expression data.
Ting Wang 0003, Jin Gu, Yanda Li
BMC Bioinform.3
2010 CORS: A cooperative overlay routing service to enhance interactive multimedia communications
Zhen Chen 0001, Jun Li 0003, Yanda Li
J. Vis. Commun. Image Represent.5
2009 CURE-Chloroplast: A chloroplast C-to-U RNA editing predictor for seed plants
abstract
BACKGROUND: RNA editing is a type of post-transcriptional modification of RNA and belongs to the class of mechanisms that contribute to the complexity of transcriptomes. C-to-U RNA editing is commonly observed in plant mitochondria and chloroplasts. The in vivo mechanism of recognizing C-to-U RNA editing sites is still unknown. In recent years, many efforts have been made to computationally predict C-to-U RNA editing sites in the mitochondria of seed plants, but there is still no algorithm available for C-to-U RNA editing site prediction in the chloroplasts of seed plants. RESULTS: In this paper, we extend our algorithm CURE, which can accurately predict the C-to-U RNA editing sites in mitochondria, to predict C-to-U RNA editing sites in the chloroplasts of seed plants. The algorithm achieves over 80% sensitivity and over 99% specificity. We implement the algorithm as an online service called CURE-Chloroplast http://bioinfo.au.tsinghua.edu.cn/pure. CONCLUSION: CURE-Chloroplast is an online service for predicting the C-to-U RNA editing sites in the chloroplasts of seed plants. The online service allows the processing of entire chloroplast genome sequences. Since CURE-Chloroplast performs very well, it could be a helpful tool in the study of C-to-U RNA editing in the chloroplasts of seed plants.
Pufeng Du, Liyan Jia, Yanda Li
BMC Bioinform.3
2009 An investigation of the Internet's IP-layer connectivity
Jun Li 0003, Yanda Li, Scott Shenker
Comput. Commun.3
2008 Identification of phylogenetically conserved microRNA cis-regulatory elements across 12 Drosophila species
abstract
MOTIVATION: MicroRNAs are a class of endogenous small RNAs that play regulatory roles. Intergenic miRNAs are believed to be transcribed independently, but the transcriptional control of these crucial regulators is still poorly understood. RESULTS: In this work, phylogenetic footprinting is used to identify conserved cis-regulatory elements (CCEs) surrounding intergenic miRNAs in Drosophila. With a two-step strategy that takes advantage of both alignment-based and motif-based methods, we identified CCEs that are conserved across the 12 fly species. When compared with TRANSFAC database, these CCEs are significantly enriched in known transcription factor binding sites (TFBSs). Moreover, several TFs that play essential roles in Drosophila development (e.g. Adf-1, Abd-B, Sd, Prd, Ubx, Zen and En) are found to be preferentially regulating the miRNA genes. Further analysis revealed many over-represented cis-regulatory modules (CRMs) composed of multiple known TFBSs, motif pairs with significant distance constraints and a number of novel motifs, many of which preferentially occur near the transcription start site of protein-coding genes. Additionally, a number of putative miRNA-TF regulatory feedback loops were also detected. AVAILABILITY: Supplementary Material and the Perl scripts performing two-step phylogenetic footprinting are available at http://bioinfo.au.tsinghua.edu.cn/member/xwwang/mircisreg
Xiaowo Wang, Jin Gu, Michael Q. Zhang, Yanda Li
Bioinform.4
2008 dbNEI2.0: building multilayer network for drug-NEI-disease
abstract
The neuro-endocrine-immune (NEI) system plays a critical regulatory role in modulating host homeostasis and optimizing health. We created the dbNEI 2 years ago to collect NEI molecules and interactions. For transferring the conceptual NEI to the systematic NEI network and uncovering the NEI's medical function, we updated the dbNEI 2.0 in three ways: (i) extended NEI molecules to 2242 genes and 7657 chemical compounds by using gene ontology-based (GO-based) data mining strategy, (ii) added multilayer interactions of NEI molecules including KEGG signal transduction and metabolic pathways, HPRD protein-protein interactions (PPI), transcription factor and microRNA regulations and (iii) connected 611 drugs and 823 diseases through multilayer NEI interactions. The reconstructed drug-NEI-disease network will facilitate the systematic study of NEI system.
Yanda Li, Shao Li
Bioinform.3
2008 Acceleration-based Dopplerlet transform - Part II: Implementations and applications to passive motion parameter estimation of moving sound source
Hongxing Zou, Lin Qiao, Shiji Song, Yanda Li
Signal Process.6
2008 Acceleration-based Dopplerlet transform - Part I: Theory
Hongxing Zou, Shiji Song, Yanda Li
Signal Process.5
2007 Identifications of conserved 7-mers in 3'-UTRs and microRNAs in Drosophila
abstract
BACKGROUND: MicroRNAs (miRNAs) are a class of endogenous regulatory small RNAs which play an important role in posttranscriptional regulations by targeting mRNAs for cleavage or translational repression. The base-pairing between the 5'-end of miRNA and the target mRNA 3'-UTRs is essential for the miRNA:mRNA recognition. Recent studies show that many seed matches in 3'-UTRs, which are fully complementary to miRNA 5'-ends, are highly conserved. Based on these features, a two-stage strategy can be implemented to achieve the de novo identification of miRNAs by requiring the complete base-pairing between the 5'-end of miRNA candidates and the potential seed matches in 3'-UTRs. RESULTS: We presented a new method, which combined multiple pairwise conservation information, to identify the frequently-occurred and conserved 7-mers in 3'-UTRs. A pairwise conservation score (PCS) was introduced to describe the conservation of all 7-mers in 3'-UTRs between any two Drosophila species. Using PCSs computed from 6 pairs of flies, we developed a support vector machine (SVM) classifier ensemble, named Cons-SVM and identified 689 conserved 7-mers including 63 seed matches covering 32 out of 38 known miRNA families in the reference dataset. In the second stage, we searched for 90 nt conserved stem-loop regions containing the complementary sequences to the identified 7-mers and used the previously published miRNA prediction software to analyze these stem-loops. We predicted 47 miRNA candidates in the genome-wide screen. CONCLUSION: Cons-SVM takes advantage of the independent evolutionary information from the 6 pairs of flies and shows high sensitivity in identifying seed matches in 3'-UTRs. Combining the multiple pairwise conservation information by the machine learning approach, we finally identified 47 miRNA candidates in D. melanogaster.
Jin Gu, Hu Fu 0001, Xuegong Zhang, Yanda Li
BMC Bioinform.4
2007 Neighbor number, valley seeking and clustering
Xuegong Zhang, Michael Q. Zhang, Yanda Li
Pattern Recognit. Lett.4
2006 Support Vector Machine Approach for Retained Introns Prediction Using Sequence Features
Huiyu Xia, Jianning Bi, Yanda Li
ISNN (2)3
2006 End-to-End Delay Behavior in the Internet
abstract
While delay-critical applications typified by online multimedia communication are growing rapidly, the end-to-end (E2E) delay behavior in the Internet remains poorly understood. This paper proposes a stochastic process model, in which E2E delay alternations are classified into two categories, jump and perturbation, according to whether the statistical characterizations alter or not. As the chief type of majority delay alternations, perturbations generally occur continuously in a period, and can be analogous to an ergodic stationary process. The authenticity of the model is verified with real life delay measurements in the Internet collected by all pairs pings (APP) project through months from PlanetLab nodes. Based on the model, several delay estimation algorithms are comparatively analyzed, and the experimental results demonstrate that in terms of minimizing the mean squared error, the most accurate delay prediction is the minimum of the two most recent measurements.
Hui Zhang 0001, Jun Li 0003, Yanda Li
MASCOTS4
2006 Identification and Comparison of Motifs in Brain-Specific and Muscle-Specific Alternative Splicing
Jianning Bi, Yanda Li
TAMC2
2006 Prediction of protein submitochondria locations by hybridizing pseudo-amino acid composition with various physicochemical features of segmented sequence
abstract
BACKGROUND: Knowing the submitochondria localization of a mitochondria protein is an important step to understand its function. We develop a method which is based on an extended version of pseudo-amino acid composition to predict the protein localization within mitochondria. This work goes one step further than predicting protein subcellular location. We also try to predict the membrane protein type for mitochondrial inner membrane proteins. RESULTS: By using leave-one-out cross validation, the prediction accuracy is 85.5% for inner membrane, 94.5% for matrix and 51.2% for outer membrane. The overall prediction accuracy for submitochondria location prediction is 85.2%. For proteins predicted to localize at inner membrane, the accuracy is 94.6% for membrane protein type prediction. CONCLUSION: Our method is an effective method for predicting protein submitochondria location. But even with our method or the methods at subcellular level, the prediction of protein submitochondria location is still a challenging problem. The online service SubMito is now available at: http://bioinfo.au.tsinghua.edu.cn/subMito.
Pufeng Du, Yanda Li
BMC Bioinform.2
2005 On the Scale-Free Intersection Graphs
Xin Yao 0003, Changshui Zhang, Yanda Li
ICCSA (2)4
2005 Classification of Nuclear Receptor Subfamilies with RBF Kernel in Support Vector Machine
Yanda Li
ISNN (3)2
2005 ATID: a web-oriented database for collection of publicly available alternative translational initiation events
abstract
SUMMARY: Alternative translational initiation is an important cellular mechanism contributing to the diversity of protein products and functions. We develop a database that provides a comprehensive collection of alternative translational initiation events. The purpose of this alternative translational initiation database (ATID) is to facilitate the systematic study of alternative translational initiation of genes. The current version of database contains 300 genes from Homo sapiens, Mus musculus and other species. Each of the genes has two or more isoforms due to alternative translational initiation. Resources in ATID, including gene information, alternative products of genes and domain structures of isoforms, are provided through a user-friendly web interface. AVAILABILITY: The ATID database is available for public use at http://bioinfo.au.tsinghua.edu.cn/atie/.
Yanda Li
Bioinform.4
2005 MicroRNA identification based on sequence and structure alignment
abstract
MOTIVATION: MicroRNAs (miRNA) are approximately 22 nt long non-coding RNAs that are derived from larger hairpin RNA precursors and play important regulatory roles in both animals and plants. The short length of the miRNA sequences and relatively low conservation of pre-miRNA sequences restrict the conventional sequence-alignment-based methods to finding only relatively close homologs. On the other hand, it has been reported that miRNA genes are more conserved in the secondary structure rather than in primary sequences. Therefore, secondary structural features should be more fully exploited in the homologue search for new miRNA genes. RESULTS: In this paper, we present a novel genome-wide computational approach to detect miRNAs in animals based on both sequence and structure alignment. Experiments show this approach has higher sensitivity and comparable specificity than other reported homologue searching methods. We applied this method on Anopheles gambiae and detected 59 new miRNA genes. AVAILABILITY: This program is available at http://bioinfo.au.tsinghua.edu.cn/miralign. SUPPLEMENTARY INFORMATION: Supplementary information is available at http://bioinfo.au.tsinghua.edu.cn/miralign/supplementary.htm.
Xiaowo Wang, Jing Zhang 0010, Jin Gu, Xuegong Zhang, Yanda Li
Bioinform.7
2005 Classification of real and pseudo microRNA precursors using local structure-sequence features and support vector machine
abstract
BACKGROUND: MicroRNAs (miRNAs) are a group of short (approximately 22 nt) non-coding RNAs that play important regulatory roles. MiRNA precursors (pre-miRNAs) are characterized by their hairpin structures. However, a large amount of similar hairpins can be folded in many genomes. Almost all current methods for computational prediction of miRNAs use comparative genomic approaches to identify putative pre-miRNAs from candidate hairpins. Ab initio method for distinguishing pre-miRNAs from sequence segments with pre-miRNA-like hairpin structures is lacking. Being able to classify real vs. pseudo pre-miRNAs is important both for understanding of the nature of miRNAs and for developing ab initio prediction methods that can discovery new miRNAs without known homology. RESULTS: A set of novel features of local contiguous structure-sequence information is proposed for distinguishing the hairpins of real pre-miRNAs and pseudo pre-miRNAs. Support vector machine (SVM) is applied on these features to classify real vs. pseudo pre-miRNAs, achieving about 90% accuracy on human data. Remarkably, the SVM classifier built on human data can correctly identify up to 90% of the pre-miRNAs from other species, including plants and virus, without utilizing any comparative genomics information. CONCLUSION: The local structure-sequence features reflect discriminative and conserved characteristics of miRNAs, and the successful ab initio classification of real and pseudo pre-miRNAs opens a new approach for discovering new miRNAs.
Chenghai Xue, Guo-Ping Liu 0003, Yanda Li, Xuegong Zhang
BMC Bioinform.5
2005 Nonnegative time-frequency distributions for parametric time-frequency representations using semi-affine transformation group
Hongxing Zou, Dianjun Wang, Xian-Da Zhang, Yanda Li
Signal Process.4
2004 Adaptive Dly-ACK for TCP over 802.15.3 WPAN
abstract
The Dly-ACK scheme in IEEE 802.15.3 is designed to reduce the overhead of the ACK and improve the channel utilization. However, the way to use the Dly-ACK is open for implementation. In this paper, we first investigate the problems of applying the fixed Dly-ACK scheme to the TCP stream and show that the TCP performance is rather poor. Then, we propose two enhancement mechanisms for TCP with Dly-ACK over the 802.15.3 system. The first one is to request the Dly-ACK frame adaptively or change the burst size of Dly-ACK according to the queue size. The second is a retransmission counter to enable the destination DEV to deliver the packets to the upper layer timeously and in orderly fashion. Simulation results show that, with our enhancements, the TCP throughput can be improved more than 40% compared with the conventional Imm-ACK. We also investigate the impacts of some important system parameters such as the buffer size on the TCP performance. Some guidelines for the Dly-ACK design are given. Finally, it is worth pointing out that our enhancements are compatible with the standard.
Hongyuan Chen, Zihua Guo, Richard Yao, Yanda Li
GLOBECOM4
2004 A simple strategy for detecting outlier samples in microarray data
abstract
Microarrays can monitor expression levels of thousands of genes simultaneously. Many people have used the gene expression data obtained with microarrays to classify different groups of samples, such as different types or subtypes of cancers. In our experiments as well as those of some other investigators, it has been observed that in some microarray data sets, there might be outlier samples which are either caused by imperfectness in the experiments or by possible mislabeling at certain steps. The existence of such samples impacts classification accuracy and may even cause misleading conclusions. In this paper, we studied this problem with two simulated data sets of typical scenarios and formed a simple but powerful strategy for detecting such outlier or mislabeled samples, built upon cross validation of the basic SVM classifier. The strategy was applied to a public colon cancer data set and it successfully detected 6 outlier cases. This work suggests an effective scheme for detecting outlier samples in a data set and for evaluating the sample quality.
Yanda Li, Xuegong Zhang
ICARCV2
2004 Discovering possible context dependences around SNP sites in human genes with Bayesian network learning
abstract
Single nucleotide polymorphisms (SNPs) are loci on the genome where different alleles are observed in the population. It has been observed that there might be some patterns or context dependences in the sequence segments adjacent to SNPs sites. Discovering such dependences is very important for understanding possible origins of SNPs in evolution. We collected 519,767 bi-allelic SNPs of human in gene regions from HGBASE and separated them in 6 groups according to the types of alleles at the SNP loci. Bayesian network structure learning technique is applied to discovery of possible dependences in sequence segments around these sites as well as in reference sequences collected as comparison. Noticeable probabilistic correlations among some loci were detected in all the 6 SNP groups and nothing significant was found in the reference sequences. The dependence relations found with different SNP groups are different. These putative context dependences around SNP sites provide important hints for further analyzing SNP-related sequences patterns. The work also illustrates the powerfulness of the Bayesian network method as a tool for biological sequence analysis.
Xi Ma, Wei Hu 0002, Yimin Zhang 0002, Yanda Li, Xuegong Zhang
ICARCV5
2004 Classifying G-protein Coupled Receptors with Support Vector Machine
Yanda Li
ISNN (2)2
2004 Hydrocarbon Reservoir Prediction Using Support Vector Machines
Kaifeng Yao, Wenkai Lu, Shanwen Zhang, Huanqin Xiao, Yanda Li
ISNN (1)5
2004 Prediction of protein subcellular locations using fuzzy k-NN method
abstract
MOTIVATION: Protein localization data are a valuable information resource helpful in elucidating protein functions. It is highly desirable to predict a protein's subcellular locations automatically from its sequence. RESULTS: In this paper, fuzzy k-nearest neighbors (k-NN) algorithm has been introduced to predict proteins' subcellular locations from their dipeptide composition. The prediction is performed with a new data set derived from version 41.0 SWISS-PROT databank, the overall predictive accuracy about 80% has been achieved in a jackknife test. The result demonstrates the applicability of this relative simple method and possible improvement of prediction accuracy for the protein subcellular locations. We also applied this method to annotate six entirely sequenced proteomes, namely Saccharomyces cerevisiae, Caenorhabditis elegans, Drosophila melanogaster, Oryza sativa, Arabidopsis thaliana and a subset of all human proteins. AVAILABILITY: Supplementary information and subcellular location annotations for eukaryotes are available at http://166.111.30.65/hying/fuzzy_loc.htm
Yanda Li
Bioinform.2
2004 HMMGEP: clustering gene expression data using hidden Markov models
abstract
SUMMARY: The package HMMGEP performs cluster analysis on gene expression data using hidden Markov models. AVAILABILITY: HMMGEP, including the source code, documentation and sample data files, is available at http://www.bioinfo.tsinghua.edu.cn:8080/~rich/hmmgep_download/index.html.
Xinglai Ji, Jesse Li-Ling, Yanda Li, Zhirong Sun
Bioinform.4
2003 Directed variation in evolution strategies
abstract
Biological evolution gives rise to self-organizing phenomena. Inspired by this theory, directed variation is added to the (/spl mu/, /spl lambda/) evolution strategies (ES) algorithm and it is called directed variation ES (DVES). In DVES, some neighboring individuals in the population mutate correlatively according to the distribution of the whole population. Experimental results showed that, with the same number of function evaluations, directed variation ES reached better optimization results for different generally used strategies under the ES framework. Experimental analysis showed that the application of directed variation could increase the expected fitness improvement and the probability of fitness improvement. From a biological perspective, directed variation can be regarded as a result of self-organizing evolution.
Yanda Li
IEEE Trans. Evol. Comput.2
2002 Subspaces of FMmlet transform
abstract
The subspaces of FM m let transform are investigated. It is shown that some of the existing transforms like the Fourier transform, short-time Fourier transform, Gabor transform, wavelet transform, chirplet transform, the mean of signal, and the FM −1 let transform, and the butterfly subspace are all special cases of FM m let transform. Therefore the FM m let transform is more flexible for delineating both the linear and nonlinear time-varying structures of a signal.
Hongxing Zou, Qionghai Dai, Guiming Chen, Yanda Li
Sci. China Ser. F Inf. Sci.5
2002 Nonexistence of cross-term free time-frequency distribution with concentration of Wigner-Ville distribution
abstract
Wigner-Ville distribution (WVD) is recognized as being a powerful tool and a nucleus in time-frequency representation (TFR) which gives an excellent time-frequency concentration, and more importantly, has many desirable properties. A major shortcoming of WVD is the inherent cross-term (CT) interference. Although solutions to this problem from the bulk of contributions to the literature concerning TFR are currently available, none has been able to completely eliminate the CT’s in WVD. It is therefore a common belief that if there exists an auxiliary time-frequency distribution (TFD) which has the same auto-terms (AT’s) as that in WVD, but has CT’s with the opposite sign, then, by adding the auxiliary TFD to WVD, an ideal TFD, which preserves the concentration of WVD while annihilating the CT’s, is readily obtained. However, we prove that the auxiliary TFD does not exist. Moreover, it is found that in general, CT free joint distributions with their concentrations close to that of WVD do not exist either.
Hongxing Zou, Xuguang Lu, Qionghai Dai, Yanda Li
Sci. China Ser. F Inf. Sci.4
2001 Prefetch Agent: Virtual Internet Based on CATV
abstract
The Internet provides an interactive mode for information retrieval, but its bandwidth is narrow. TV can provide a broadband service, but usually it is not interactive. In China, CATV (cable television) is very popular and mainly runs in one-way mode. It is expensive to rebuild CATV from one-way to two-way. Using a prefetch system, we can achieve a virtual Internet and virtual two-way communication through CATV.
Haiming Lu, Zengxiang Lu, Yanda Li
ICALT3
2001 TRUST!-A distributed multi-agent system for community formation and information recommendation
abstract
In centralized collaborative information recommendation, there is a bottleneck for the scalability and the availability. Yenta is a decentralized approach. Information can be recommended through the clustering. The clustering of similar agents in Yenta is based on the similarities between the content of their interest models. So the agents must represent the interests of users or the models of other agents in the same way. This paper introduces the TRUST! system. In TRUST!, the agent's user evaluates the information recommended from other agents. The trust degree between different agents is decided by the user's evaluation. The similarities between different agents are measured by the trust connections. So the agents can represent the interests of users or the models of other agents in different ways. TRUST! is a distributed multiagent system for Internet information propagation and recommendation. It is based on limited friends list and trust relationship. To update trust, we introduce the PID (proportion, integral, and differential coefficient) arithmetic. Through emulation experiments, we give some analysis of the distributed clustering.
Haiming Lu, Zengxiang Lu, Yanda Li
SMC3
2001 Parametric TFR via windowed exponential frequency modulated atoms
abstract
We propose a new atom, namely, the dilated and translated windowed exponential frequency modulated functions (FM/sup m/let) for compactly characterizing both the signal's time-invariant and time-varying spectral contents. The superiority of the proposed method to some existing time-frequency distributions (TFDs) is demonstrated using a bat sonar signal.
Hongxing Zou, Qionghai Dai, Renming Wang, Yanda Li
IEEE Signal Process. Lett.4
2000 Joint speech signal enhancement based on spectral subtraction and SVD filter
Wenkai Lu, Xuegong Zhang, Yanda Li, Li Qin Shen, Weibin Zhu
INTERSPEECH3
2000 The signal reconstruction of speech by KPCA
Xuegong Zhang, Yanda Li, Li Qin Shen, Weibin Zhu
INTERSPEECH3
1999 Stable fuzzy adaptive control for a class of nonlinear systems
Yuezhong Tang, Naiyao Zhang, Yanda Li
Fuzzy Sets Syst.3
1998 The application of blind channel identification techniques to prestack seismic deconvolution
abstract
One objective of seismic signal processing is to identify the layered subsurface structure by sending seismic wavelets into the ground. This is a blind deconvolution process since the seismic wavelets are usually not measurable and therefore, the subsurface face layers are identified only by the reflected seismic signals. Conventional methods often approach this problem by making assumptions about the subsurface structures and/or the seismic wavelets. In this paper an alternative technique is presented. It applies blind channel identification methods to prestack seismic deconvolution. A unique feature of this proposed method is that no such assumptions are needed. In addition, it fits into the structure of current seismic data acquisition techniques, thus no extra cost is involved. Simulations on both synthetic and field seismic data demonstrate that it is a promising new method for seismic signal processing.
Yanda Li
Proc. IEEE2
1995 An unified prefiltering-based approach to harmonic retrieval in non-Gaussian ARMA noise
abstract
The paper addresses the harmonic retrieval problem in colored noise. As contrasted to the reported studies in which Gaussian noise was assumed, the present paper concentrates on additive non-Gaussian ARMA noise. The authors propose a unified prefiltering-based approach to this problem. The approach is hybrid in the sense that 3rd-order cumulants are first used to identify the AR part of the non-Gaussian noise process, and then correlation-based high resolution methods may be used for the filtered output process to estimate the parameters of harmonics. Simulation examples are presented to demonstrate the high resolution of this approach.
Ying-Chang Liang, Xian-Da Zhang, Yanda Li
ICASSP3
1995 EAMUSE: An Extended Algorithm for Multiple Sources Extraction
abstract
This paper addresses the problem of multiple source signals separation in noise. As contrasted to the reported studies in which white noise in different sensors with same noise covariance was assumed, the additive noise sensors considered in this paper have different noise covariance. An extended algorithm for multiple sources extraction (EAMUSE) is proposed. The effectiveness of our approach is demonstrated through standard simulation examples.
Ying-Chang Liang, Yanda Li, Xian-Da Zhang
ISCAS2
1994 A hybrid approach to harmonic retrieval in non-Gaussian ARMA noise
abstract
Addresses the harmonic retrieval problem in colored noise. As contrasted to the reported studies in which Gaussian noise was assumed, this paper focuses on additive non-Gaussian ARMA noise. Our approach is hybrid in the sense that third-order cumulants are first used to identify the AR part of the non-Gaussian noise process, and then correlation-based high-resolution methods are used for the filtered process to estimate the number of harmonics and their frequencies. Simulation examples are presented to demonstrate the high resolution of this approach.>
Xian-Da Zhang, Ying-Chang Liang, Yanda Li
IEEE Trans. Inf. Theory3
1988 An aggregate group approach for AR spectral estimation of noisy signals
abstract
A least-squares estimate of parameters of an overdetermined Yule-Walker equation can contain very large errors in the low signal-to-noise ratio. This sensitivity to noise can be seen in parameter space, where each of the equations appears as a line. The least-squares method is equivalent to finding the center of a circle inscribed in the lines. As an alternative the authors examine the center of gravity as a new objective function called the weighting gravity center method. The method is made increasingly robust by eliminating wild points attributed to noise and the use of the overdetermined Yule-Walker equation.>
Yanda Li, Tong Chang
ICASSP2
1988 Solving the stiff problem in computer vision by trade-off optimization
abstract
The effect of stiffness in a general solution of the linear inverse problem is discussed, and a method is proposed to deal with this problem. It uses tradeoff by tradeoff optimization between the resolution and variance of solution to keep them both small, by applying a group of optimal weights on the singular values of the operator. The reliability of the estimated motion parameters is improved significantly.>
Chengzhi Huang, Yanda Li, Tong Chang
ICPR2
1987 Discrete signal reconstruction from its autocorrelation function and one sample
abstract
This work follows the 'Discrete Signal Reconstruction from its Spectral Magnitude and Some Samples'[1]. The uniqueness of the discrete signal reconstruction from its autocorrelation function and one sample has been discussed in detail in this paper. Four theorems are presented. In addition, we provide an effective iterative algorithm recovering the discrete signal.
Zhongze Wu, Yanda Li, Tong Chang
ICASSP2