EDBT 2026 Demo / reviewers in the wild / expert
Yanpeng Zhao
dblp:182/5860
· DBLP profile ↗
21ranked-venue papers
8as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 7 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | v-HUB: A Benchmark for Video Humor Understanding from Vision and SoundabstractZhengpeng Shi, Yanpeng Zhao, Jianqun Zhou, Yuxuan Wang, Qinrong Cui, Wei Bi, Song-Chun Zhu, Bo Zhao, Zilong Zheng. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zhengpeng Shi, Yanpeng Zhao, Jianqun Zhou, Yuxuan Wang 0004, Qinrong Cui, Wei Bi, Song-Chun Zhu, Zilong Zheng |
ACL (1) | 2 |
| 2025 | AgentSentinel: An End-to-End and Real-Time Security Defense Framework for Computer-Use AgentsabstractLarge Language Models (LLMs) have been increasingly integrated into computer-use agents, which can autonomously operate tools on a user's computer to accomplish complex tasks. However, due to the inherently unstable and unpredictable nature of LLM outputs, they may issue unintended tool commands or incorrect inputs, leading to potentially harmful operations. Unlike traditional security risks stemming from insecure user prompts, tool execution results from LLM-driven decisions introduce new and unique security challenges. These vulnerabilities span across all components of a computer-use agent. To mitigate these risks, we propose AgentSentinel, an end-to-end, real-time defense framework designed to mitigate potential security threats on a user's computer. AgentSentinel intercepts all sensitive operations within agent-related services and halts execution until a comprehensive security audit is completed. Our security auditing mechanism introduces a novel inspection process that correlates the current task context with system traces generated during task execution. To thoroughly evaluate AgentSentinel, we present BadComputerUse, a benchmark consisting of 60 diverse attack scenarios across six attack categories. The benchmark demonstrates a 87% average attack success rate on four state-of-the-art LLMs. Our evaluation shows that AgentSentinel achieves an average defense success rate of 79.6%, significantly outperforming all baseline defenses. Haitao Hu, Yanpeng Zhao, Yuqi Chen 0001 |
CCS | 3 |
| 2025 | NEP: Autoregressive Image Editing via Next Editing Token PredictionabstractText-guided image editing involves modifying a source image based on a language instruction and, typically, requires changes to only small local regions. However, existing approaches generate the entire target image rather than selectively regenerate only the intended editing areas. This results in (1) unnecessary computational costs and (2) a bias toward reconstructing non-editing regions, which compromises the quality of the intended edits. To resolve these limitations, we propose to formulate image editing as $\textbf{N}$ext $\textbf{E}$diting-token $\textbf{P}$rediction (NEP) based on autoregressive image generation, where only regions that need to be edited are regenerated, thus avoiding unintended modification to the non-editing areas. To enable any-region editing, we propose to pre-train an any-order autoregressive text-to-image (T2I) model. Once trained, it is capable of zero-shot image editing and can be easily adapted to NEP for image editing, which achieves a new state-of-the-art on widely used image editing benchmarks. Moreover, our model naturally supports test-time scaling (TTS) through iteratively refining its generation in a zero-shot manner. Huimin Wu 0001, Xiaojian Ma 0001, Haozhe Zhao, Yanpeng Zhao, Qing Li 0003 |
NeurIPS | 4 |
| 2025 | IECata: interpretable bilinear attention network and evidential deep learning improve the catalytic efficiency prediction of enzymesabstractEnzyme catalytic efficiency (kcat/Km) is a key parameter for identifying high-activity enzymes. Recently, deep learning techniques have demonstrated the potential for fast and accurate kcat/Km prediction. However, three challenges remain: (i) the limited size of the available kcat/Km dataset hinders the development of deep learning models; (ii) the model predictions lack reliable confidence estimates; and (iii) models lack interpretable insights into enzyme-catalyzed reactions. To address these challenges, we proposed IECata, a kcat/Km prediction model that provides uncertainty estimation and interpretability. IECata collected a dataset of 11 815 kcat/Km entries from the BRENDA and SABIO-RK databases, along with an out-of-domain test dataset of 806 entries from the literature. By introducing evidential deep learning, IECata provides uncertainty estimates for kcat/Km predictions. Moreover, it uses a bilinear attention mechanism to focus on learning crucial local interactions to interpret the key residues and substrate atoms in enzyme-catalyzed reactions. Testing results indicate that the prediction performance of IECata exceeds that of state-of-the-art benchmark models. More importantly, it provides a reliable confidence assessment for these predictions. Case studies further highlight that the incorporation of uncertainty in screening for highly active enzymes can effectively increase the hit ratio, thereby improving the efficiency of experimental validation and accelerating directed enzyme evolution. To facilitate researchers' use of IECata, we have developed an online prediction platform: http://mathtc.nscc-tj.cn/cataai/. Yanpeng Zhao, Zhijiang Yang, Ge Yao, Penggang Han, Peng Zan, Xiukun Wan, Xiaochen Bo |
Briefings Bioinform. | 2 |
| 2025 | Dynamic Scene Understanding Through Object-Centric Voxelization and Neural RenderingabstractLearning object-centric representations from unsupervised videos is challenging. Unlike most previous approaches that focus on decomposing 2D images, we present a 3D generative model named DynaVol-S for dynamic scenes that enables object-centric learning within a differentiable volume rendering framework. The key idea is to perform object-centric voxelization to capture the 3D nature of the scene, which infers per-object occupancy probabilities at individual spatial locations. These voxel features evolve through a canonical-space deformation function and are optimized in an inverse rendering pipeline with a compositional NeRF. Additionally, our approach integrates 2D semantic features to create 3D semantic grids, representing the scene through multiple disentangled voxel grids. DynaVol-S significantly outperforms existing models in both novel view synthesis and unsupervised decomposition tasks for dynamic scenes. By jointly considering geometric structures and semantic features, it effectively addresses challenging real-world scenarios involving complex object interactions. Furthermore, once trained, the explicitly meaningful voxel features enable additional capabilities that 2D scene decomposition methods cannot achieve, such as novel scene generation through editing geometric shapes or manipulating the motion trajectories of objects. Yanpeng Zhao, Yiwei Hao, Siyu Gao, Yunbo Wang, Xiaokang Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Research on the Simultaneous Measurement Method of Atmospheric Temperature and Pressure Profiles Based on the Lidar Echo Energy and Characteristics of the Scattering SpectrumabstractLidar plays an important role in the detection of atmospheric environmental parameters due to its property of high detection accuracy, good real-time performance, and high spatiotemporal resolution. Measuring the spectral properties of molecular scattered light enables to determine the profiles of atmospheric parameters such as wind speed, temperature, and pressure with high vertical resolution (~100 m) by using sophisticated multiparameter inversion algorithms. So far, atmospheric measurements of Rayleigh–Brillouin (RB) spectrum were used to only determine temperature, using pressure from respective models or from measurements such as radiosondes. In this article, we propose a new simultaneous inversion method of atmospheric temperature and pressure based on lidar echo energy and the information contained in the RB spectrum. The accuracy and uncertainty of the proposed method are analyzed based on numerical simulations and atmospheric measurements and compared to state-of-the-art retrieval methods. It is shown that the accuracy of retrieved temperature can be improved from 2.0 to 1.1 K and one of the pressures from 1.2% to 0.8%. Furthermore, the potential of additionally retrieving wind speed and the aerosol extinction coefficient based on spectral characteristics and intensity of the scattered light is investigated and discussed. This lays a foundation for the future development and application lidar instruments capable of multiparameter retrieval. Yangrui Xu, Yujun Gu, Peize Li, Yanpeng Zhao, Kun Liang 0003 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | DynaVol: Unsupervised Learning for Dynamic Scenes through Object-Centric VoxelizationabstractUnsupervised learning of object-centric representations in dynamic visual scenes is challenging. Unlike most previous approaches that learn to decompose 2D images, we present DynaVol, a 3D scene generative model that unifies geometric structures and object-centric learning in a differentiable volume rendering framework. The key idea is to perform object-centric voxelization to capture the 3D nature of the scene, which infers the probability distribution over objects at individual spatial locations. These voxel features evolve over time through a canonical-space deformation function, forming the basis for global representation learning via slot attention. The voxel features and global features are complementary and are both leveraged by a compositional NeRF decoder for volume rendering. DynaVol remarkably outperforms existing approaches for unsupervised dynamic scene decomposition. Once trained, the explicitly meaningful voxel features enable additional capabilities that 2D scene decomposition methods cannot achieve: it is possible to freely edit the geometric shapes or manipulate the motion trajectories of the objects. Yanpeng Zhao, Siyu Gao, Yunbo Wang, Xiaokang Yang 0001 |
ICLR | 1 |
| 2024 | An Efficient Recipe for Long Context Extension via Middle-Focused Positional EncodingabstractRecently, many methods have been developed to extend the context length of pre-trained large language models (LLMs), but they often require fine-tuning at the target length ($\gg4K$) and struggle to effectively utilize information from the middle part of the context. To address these issues, we propose $\textbf{C}$ontinuity-$\textbf{R}$elativity ind$\textbf{E}$xing with g$\textbf{A}$ussian $\textbf{M}$iddle ($\texttt{CREAM}$), which interpolates positional encodings by manipulating position indices. Apart from being simple, $\texttt{CREAM}$ is training-efficient: it only requires fine-tuning at the pre-trained context window (e.g., Llama 2-4K) and can extend LLMs to a much longer target context length (e.g., 256K). To ensure that the model focuses more on the information in the middle, we introduce a truncated Gaussian to encourage sampling from the middle part of the context during fine-tuning, thus alleviating the ''Lost-in-the-Middle'' problem faced by long-context LLMs. Experimental results show that $\texttt{CREAM}$ successfully extends LLMs to the target length for both Base and Chat versions of $\texttt{Llama2-7B}$ with ``Never Miss A Beat''. Our code is publicly available at https://github.com/bigai-nlco/cream. Yanpeng Zhao, Zilong Zheng |
NeurIPS | 2 |
| 2023 | Identification of metal ion-binding sites in RNA structures using deep learning methodabstractMetal ion is an indispensable factor for the proper folding, structural stability and functioning of RNA molecules. However, it is very difficult for experimental methods to detect them in RNAs. With the increase of experimentally resolved RNA structures, it becomes possible to identify the metal ion-binding sites in RNA structures through in-silico methods. Here, we propose an approach called Metal3DRNA to identify the binding sites of the most common metal ions (Mg2+, Na+ and K+) in RNA structures by using a three-dimensional convolutional neural network model. The negative samples, screened out based on the analysis for binding surroundings of metal ions, are more like positive ones than the randomly selected ones, which are beneficial to a powerful predictor construction. The microenvironments of the spatial distributions of C, O, N and P atoms around a sample are extracted as features. Metal3DRNA shows a promising prediction power, generally surpassing the state-of-the-art methods FEATURE and MetalionRNA. Finally, utilizing the visualization method, we inspect the contributions of nucleotide atoms to the classification in several cases, which provides a visualization that helps to comprehend the model. The method will be helpful for RNA structure prediction and dynamics simulation study. Availability and implementation: The source code is available at https://github.com/ChunhuaLiLab/Metal3DRNA. Yanpeng Zhao, Fubin Chang, Weikang Gong, Chunhua Li 0001 |
Briefings Bioinform. | 1 |
| 2022 | MERLOT RESERVE: Neural Script Knowledge through Vision and Language and SoundabstractAs humans, we navigate a multimodal world, building a holistic understanding from all our senses. We introduce @MERLOT RESERVE, a model that represents videos jointly over time - through a new training objective that learns from audio, subtitles, and video frames. Given a video, we replace snippets of text and audio with a MASK token; the model learns by choosing the correct masked-out snippet. Our objective learns faster than alternatives, and performs well at scale: we pretrain on 20 million YouTube videos. Empirical results show that @MERLOT RESERVE learns strong multimodal representations. When finetuned, it sets state-of-the-art on Visual Commonsense Reasoning (VCR), TVQA, and Kinetics-600; outperforming prior work by 5%, 7%, and 1.5% respectively. Ablations show that these tasks benefit from audio pretraining - even VCR, a QA task centered around images (without sound). Moreover, our objective enables out-of-the-box prediction, revealing strong multimodal commonsense understanding. In a fully zero-shot setting, our model obtains competitive results on four video tasks, even outperforming supervised approaches on the recently proposed Situated Reasoning (STAR) benchmark. We analyze why audio enables better vision-language representations, suggesting significant opportunities for future research. We conclude by discussing ethical and societal implications of multimodal pretraining. Rowan Zellers, Jiasen Lu, Ximing Lu, Youngjae Yu, Yanpeng Zhao, Mohammadreza Salehi, Aditya Kusupati, Jack Hessel, Ali Farhadi, Yejin Choi 0001 |
CVPR | 5 |
| 2022 | Connecting the Dots between Audio and Text without Parallel Data through Visual Knowledge TransferabstractYanpeng Zhao, Jack Hessel, Youngjae Yu, Ximing Lu, Rowan Zellers, Yejin Choi. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Yanpeng Zhao, Jack Hessel, Youngjae Yu, Ximing Lu, Rowan Zellers, Yejin Choi 0001 |
NAACL-HLT | 1 |
| 2022 | Computational methods, databases and tools for synthetic lethality predictionabstractSynthetic lethality (SL) occurs between two genes when the inactivation of either gene alone has no effect on cell survival but the inactivation of both genes results in cell death. SL-based therapy has become one of the most promising targeted cancer therapies in the last decade as PARP inhibitors achieve great success in the clinic. The key point to exploiting SL-based cancer therapy is the identification of robust SL pairs. Although many wet-lab-based methods have been developed to screen SL pairs, known SL pairs are less than 0.1% of all potential pairs due to large number of human gene combinations. Computational prediction methods complement wet-lab-based methods to effectively reduce the search space of SL pairs. In this paper, we review the recent applications of computational methods and commonly used databases for SL prediction. First, we introduce the concept of SL and its screening methods. Second, various SL-related data resources are summarized. Then, computational methods including statistical-based methods, network-based methods, classical machine learning methods and deep learning methods for SL prediction are summarized. In particular, we elaborate on the negative sampling methods applied in these models. Next, representative tools for SL prediction are introduced. Finally, the challenges and future work for SL prediction are discussed. Junshan Han, Yanpeng Zhao, Caiyun Zhao, Bowei Yan, Chong Dai, Lianlian Wu, Yuqi Wen, Dongjin Leng, Zhongming Wang, Xiaoxi Yang, Xiaochen Bo |
Briefings Bioinform. | 4 |
| 2021 | Neural Bi-Lexicalized PCFG InductionabstractSonglin Yang, Yanpeng Zhao, Kewei Tu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Yanpeng Zhao, Kewei Tu |
ACL/IJCNLP (1) | 2 |
| 2021 | PCFGs Can Do Better: Inducing Probabilistic Context-Free Grammars with Many SymbolsabstractProbabilistic context-free grammars (PCFGs) with neural parameterization have been shown to be effective in unsupervised phrasestructure grammar induction.However, due to the cubic computational complexity of PCFG representation and parsing, previous approaches cannot scale up to a relatively large number of (nonterminal and preterminal) symbols.In this work, we present a new parameterization form of PCFGs based on tensor decomposition, which has at most quadratic computational complexity in the symbol number and therefore allows us to use a much larger number of symbols.We further use neural parameterization for the new form to improve unsupervised parsing performance.We evaluate our model across ten languages and empirically demonstrate the effectiveness of using more symbols. Yanpeng Zhao, Kewei Tu |
NAACL-HLT | 2 |
| 2021 | aPRBind: protein-RNA interface prediction by combining sequence and I-TASSER model-based structural features learned with convolutional neural networksabstractMOTIVATION: Protein-RNA interactions play a critical role in various biological processes. The accurate prediction of RNA-binding residues in proteins has been one of the most challenging and intriguing problems in the field of computational biology. The existing methods still have a relatively low accuracy especially for the sequence-based ab-initio methods. RESULTS: In this work, we propose an approach aPRBind, a convolutional neural network-based ab-initio method for RNA-binding residue prediction. aPRBind is trained with sequence features and structural ones (particularly including residue dynamics information and residue-nucleotide propensity developed by us) that are extracted from the predicted structures by I-TASSER. The analysis of feature contributions indicates the sequence features are most important, followed by dynamics information, and the sequence and structural features are complementary in binding site prediction. The performance comparison of our method with other peer ones on benchmark dataset shows that aPRBind outperforms some state-of-the-art ab-initio methods. Additionally, aPRBind can give a better prediction for the modeled structures with TM-score≥0.5, and meanwhile since the structural features are not very sensitive to the refined 3D structures, aPRBind has only a marginal dependence on the accuracy of the structure model, which allows aPRBind to be applied to the RNA-binding site prediction for the modeled or unbound structures. AVAILABILITY AND IMPLEMENTATION: The source code is available at https://github.com/ChunhuaLiLab/aPRbind. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Weikang Gong, Yanpeng Zhao, Xueqing Deng, Chunhua Li 0001 |
Bioinform. | 3 |
| 2021 | EDLMFC: an ensemble deep learning framework with multi-scale features combination for ncRNA-protein interaction predictionabstractBACKGROUND: Non-coding RNA (ncRNA) and protein interactions play essential roles in various physiological and pathological processes. The experimental methods used for predicting ncRNA-protein interactions are time-consuming and labor-intensive. Therefore, there is an increasing demand for computational methods to accurately and efficiently predict ncRNA-protein interactions. RESULTS: In this work, we presented an ensemble deep learning-based method, EDLMFC, to predict ncRNA-protein interactions using the combination of multi-scale features, including primary sequence features, secondary structure sequence features, and tertiary structure features. Conjoint k-mer was used to extract protein/ncRNA sequence features, integrating tertiary structure features, then fed into an ensemble deep learning model, which combined convolutional neural network (CNN) to learn dominating biological information with bi-directional long short-term memory network (BLSTM) to capture long-range dependencies among the features identified by the CNN. Compared with other state-of-the-art methods under five-fold cross-validation, EDLMFC shows the best performance with accuracy of 93.8%, 89.7%, and 86.1% on RPI1807, NPInter v2.0, and RPI488 datasets, respectively. The results of the independent test demonstrated that EDLMFC can effectively predict potential ncRNA-protein interactions from different organisms. Furtherly, EDLMFC is also shown to predict hub ncRNAs and proteins presented in ncRNA-protein networks of Mus musculus successfully. CONCLUSIONS: In general, our proposed method EDLMFC improved the accuracy of ncRNA-protein interaction predictions and anticipated providing some helpful guidance on ncRNA functions research. The source code of EDLMFC and the datasets used in this work are available at https://github.com/JingjingWang-87/EDLMFC . Yanpeng Zhao, Weikang Gong, Xiaoqian Huang, Jianjun Tan |
BMC Bioinform. | 2 |
| 2020 | Visually Grounded Compound PCFGsabstractExploiting visual groundings for language understanding has recently been drawing much attention.In this work, we study visually grounded grammar induction and learn a constituency parser from both unlabeled text and its visual groundings.Existing work on this task (Shi et al., 2019) optimizes a parser via REINFORCE and derives the learning signal only from the alignment of images and sentences.While their model is relatively accurate overall, its error distribution is very uneven, with low performance on certain constituents types (e.g., 26.2% recall on verb phrases, VPs) and high on others (e.g., 79.6% recall on noun phrases, NPs).This is not surprising as the learning signal is likely insufficient for deriving all aspects of phrasestructure syntax and gradient estimates are noisy.We show that using an extension of probabilistic context-free grammar model we can do fully-differentiable end-to-end visually grounded learning.Additionally, this enables us to complement the image-text alignment loss with a language modeling objective.On the MSCOCO test captions, our model establishes a new state of the art, outperforming its non-grounded version and, thus, confirming the effectiveness of visual groundings in constituency grammar induction.It also substantially outperforms the previous grounded model, with largest improvements on more 'abstract' categories (e.g., +55.1% recall on VPs). 1 Yanpeng Zhao, Ivan Titov 0001 |
EMNLP (1) | 1 |
| 2018 | Gaussian Mixture Latent Vector GrammarsabstractWe introduce Latent Vector Grammars (LVeGs), a new framework that extends latent variable grammars such that each nonterminal symbol is associated with a continuous vector space representing the set of (infinitely many) subtypes of the nonterminal.We show that previous models such as latent variable grammars and compositional vector grammars can be interpreted as special cases of LVeGs.We then present Gaussian Mixture LVeGs (GM-LVeGs), a new special case of LVeGs that uses Gaussian mixtures to formulate the weights of production rules over subtypes of nonterminals.A major advantage of using Gaussian mixtures is that the partition function and the expectations of subtype rules can be computed using an extension of the inside-outside algorithm, which enables efficient inference and learning.We apply GM-LVeGs to part-of-speech tagging and constituency parsing and show that GM-LVeGs can achieve competitive accuracies.Our code is available at https://github.com/zhaoyanpeng/lveg. Yanpeng Zhao, Kewei Tu |
ACL (1) | 1 |
| 2017 | Structured Attentions for Visual Question AnsweringabstractVisual attention, which assigns weights to image regions according to their relevance to a question, is considered as an indispensable part by most Visual Question Answering models. Although the questions may involve complex rela- tions among multiple regions, few attention models can ef- fectively encode such cross-region relations. In this paper, we demonstrate the importance of encoding such relations by showing the limited effective receptive field of ResNet on two datasets, and propose to model the visual attention as a multivariate distribution over a grid-structured Con- ditional Random Field on image regions. We demonstrate how to convert the iterative inference algorithms, Mean Field and Loopy Belief Propagation, as recurrent layers of an end-to-end neural network. We empirically evalu- ated our model on 3 datasets, in which it surpasses the best baseline model of the newly released CLEVR dataset [13] by 9.5%, and the best published model on the VQA dataset [3] by 1.25%. Source code is available at https://github.com/zhuchen03/vqa-sva. Yanpeng Zhao, Shuaiyi Huang, Kewei Tu |
ICCV | 2 |
| 2017 | Learning Bayesian network structures under incremental construction curricula
Yanpeng Zhao, Yetian Chen, Kewei Tu, Jin Tian 0001 |
Neurocomputing | 1 |
| 2015 | Curriculum Learning of Bayesian Network Structures
Yanpeng Zhao, Yetian Chen, Kewei Tu, Jin Tian 0001 |
ACML | 1 |