Ziyi Yang 0011

dblp:217/8277-11 · DBLP profile ↗
← Back
15ranked-venue papers
7as first author
13since 2021 · last 2025
0000-0001-9849-0180ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 7 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 CPSea: Large-scale cyclic peptide-protein complex dataset for machine learning in cyclic peptide design
abstract
Cyclic peptides exhibit better binding affinity and proteolytic stability compared to their linear counterparts. However, the development of cyclic peptide design models is hindered by the scarcity of data. To address this, we introduce **CPSea**(**C**yclic **P**eptide **Sea**), a dataset of 2.71 million cyclic peptide-receptor complexes, curated through systematic mining of the AlphaFold Database (AFDB). Our pipeline extracts compact domains from AFDB, identifies cyclization sites using the $\beta$-carbon (C$_\beta$) distance thresholds, and applies multi-stage filtering to ensure structure fidelity and binding compatibility. Compared with experimental data of cyclic peptides, CPSea shows similar distributions in metrics on structure fidelity and wet-lab compatibility. To our knowledge, CPSea is the largest cyclic peptide-receptor dataset to date, enabling end-to-end model training for the first time. The dataset also showcases the feasibility of simulating inter-chain interactions using intra-chain interactions, expanding available resources for machine-learning models on protein-protein interactions. The dataset and relevant scripts are accessible on GitHub ([https://github.com/YZY010418/CPSea](https://github.com/YZY010418/CPSea)).
Ziyi Yang 0011, Hanyuan Xie, Yinjun Jia, Xiangzhe Kong, Jiqing Zheng, Ziting Zhang, Yang Liu 0003, Lei Liu 0049, Yanyan Lan
NeurIPS1
2023 i-Code: An Integrative and Composable Multimodal Learning Framework
abstract
Human intelligence is multimodal; we integrate visual, linguistic, and acoustic signals to maintain a holistic worldview. Most current pretraining methods, however, are limited to one or two modalities. We present i-Code, a self-supervised pretraining framework where users may flexibly combine the modalities of vision, speech, and language into unified and general-purpose vector representations. In this framework, data from each modality are first given to pretrained single-modality encoders. The encoder outputs are then integrated with a multimodal fusion network, which uses novel merge- and co-attention mechanisms to effectively combine information from the different modalities. The entire system is pretrained end-to-end with new objectives including masked modality unit modeling and cross-modality contrastive learning. Unlike previous research using only video for pretraining, the i-Code framework can dynamically process single, dual, and triple-modality data during training and inference, flexibly projecting different combinations of modalities into a single representation space. Experimental results demonstrate how i-Code can outperform state-of-the-art techniques on five multimodal understanding tasks and single-modality benchmarks, improving by as much as 11% and demonstrating the power of integrative multimodal pretraining.
Ziyi Yang 0011, Yuwei Fang, Chenguang Zhu 0001, Reid Pryzant, Dongdong Chen 0001, Yu Shi 0001, Yichong Xu, Yao Qian, Mei Gao, Liyang Lu, Yujia Xie, Robert Gmyr, Noel Codella, Naoyuki Kanda, Bin Xiao 0004, Lu Yuan 0001, Takuya Yoshioka, Michael Zeng 0001, Xuedong Huang 0001
AAAI1
2023 UniSumm and SummZoo: Unified Model and Diverse Benchmark for Few-Shot Summarization
abstract
Yulong Chen, Yang Liu, Ruochen Xu, Ziyi Yang, Chenguang Zhu, Michael Zeng, Yue Zhang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Yulong Chen 0001, Yang Liu 0124, Ruochen Xu, Ziyi Yang 0011, Chenguang Zhu 0001, Michael Zeng 0001, Yue Zhang 0004
ACL (1)4
2023 APOLLO: A Simple Approach for Adaptive Pretraining of Language Models for Logical Reasoning
abstract
Soumya Sanyal, Yichong Xu, Shuohang Wang, Ziyi Yang, Reid Pryzant, Wenhao Yu, Chenguang Zhu, Xiang Ren. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Soumya Sanyal 0001, Yichong Xu, Shuohang Wang, Ziyi Yang 0011, Reid Pryzant, Wenhao Yu 0002, Chenguang Zhu 0001, Xiang Ren 0001
ACL (1)4
2023 Unifying Vision, Text, and Layout for Universal Document Processing
abstract
We propose Universal Document Processing (UDOP), a foundation Document AI model which unifies text, image, and layout modalities together with varied task formats, including document understanding and generation. UDOP leverages the spatial correlation between textual content and document image to model image, text, and layout modalities with one uniform representation. With a novel Vision-Text-Layout Transformer, UDOP unifies pretraining and multi-domain downstream tasks into a prompt-based sequence generation scheme. UDOP is pretrained on both large-scale unlabeled document corpora using innovative self-supervised objectives and diverse labeled data. UDOP also learns to generate document images from text and layout modalities via masked image reconstruction. To the best of our knowledge, this is the first time in the field of document AI that one model simultaneously achieves high-quality neural document editing and content customization. Our method sets the state-of-the-art on 8 Document AI tasks, e.g., document understanding and QA, across diverse data domains like finance reports, academic papers, and web-sites. UDOP ranks first on the leaderboard of the Document Understanding Benchmark.11Code and models: https://github.com/microsoft/i-Code/tree/main/i-Code-Doc
Zineng Tang, Ziyi Yang 0011, Yuwei Fang, Yang Liu 0124, Chenguang Zhu 0001, Michael Zeng 0001, Cha Zhang, Mohit Bansal
CVPR2
2023 Any-to-Any Generation via Composable Diffusion
abstract
We present Composable Diffusion (CoDi), a novel generative model capable of generating any combination of output modalities, such as language, image, video, or audio, from any combination of input modalities. Unlike existing generative AI systems, CoDi can generate multiple modalities in parallel and its input is not limited to a subset of modalities like text or image. Despite the absence of training datasets for many combinations of modalities, we propose to align modalities in both the input and output space. This allows CoDi to freely condition on any input combination and generate any group of modalities, even if they are not present in the training data. CoDi employs a novel composable generation strategy which involves building a shared multimodal space by bridging alignment in the diffusion process, enabling the synchronized generation of intertwined modalities, such as temporally aligned video and audio. Highly customizable and flexible, CoDi achieves strong joint-modality generation quality, and outperforms or is on par with the unimodal state-of-the-art for single-modality synthesis.
Zineng Tang, Ziyi Yang 0011, Chenguang Zhu 0001, Michael Zeng 0001, Mohit Bansal
NeurIPS2
2023 MACSum: Controllable Summarization with Mixed Attributes
abstract
Abstract Controllable summarization allows users to generate customized summaries with specified attributes. However, due to the lack of designated annotations of controlled summaries, existing work has to craft pseudo datasets by adapting generic summarization benchmarks. Furthermore, most research focuses on controlling single attributes individually (e.g., a short summary or a highly abstractive summary) rather than controlling a mix of attributes together (e.g., a short and highly abstractive summary). In this paper, we propose MACSum, the first human-annotated summarization dataset for controlling mixed attributes. It contains source texts from two domains, news articles and dialogues, with human-annotated summaries controlled by five designed attributes (Length, Extractiveness, Specificity, Topic, and Speaker). We propose two simple and effective parameter-efficient approaches for the new task of mixed controllable summarization based on hard prompt tuning and soft prefix tuning. Results and analysis demonstrate that hard prompt models yield the best performance on most metrics and human evaluations. However, mixed-attribute control is still challenging for summarization tasks. Our dataset and code are available at https://github.com/psunlpgroup/MACSum.
Yusen Zhang 0001, Yang Liu 0124, Ziyi Yang 0011, Yuwei Fang, Yulong Chen 0001, Dragomir R. Radev, Chenguang Zhu 0001, Michael Zeng 0001, Rui Zhang 0037
Trans. Assoc. Comput. Linguistics3
2022 Empowering Language Models with Knowledge Graph Reasoning for Open-Domain Question Answering
abstract
Answering open-domain questions requires world knowledge about in-context entities.As pre-trained Language Models (LMs) lack the power to store all required knowledge, external knowledge sources, such as knowledge graphs, are often used to augment LMs.In this work, we propose knOwledge REasOning empowered Language Model (OREOLM), which consists of a novel Knowledge Interaction Layer that can be flexibly plugged into existing Transformer-based LMs to interact with a differentiable Knowledge Graph Reasoning module collaboratively.In this way, LM guides KG to walk towards the desired answer, while the retrieved knowledge improves LM.By adopting OREOLM to RoBERTa and T5, we show significant performance gain, achieving state-of-art results in the Closed-Book setting.The performance enhancement is mainly from the KG reasoning's capacity to infer missing relational facts.In addition, OREOLM provides reasoning paths as rationales to interpret the model's decision.
Ziniu Hu, Yichong Xu, Wenhao Yu 0002, Shuohang Wang, Ziyi Yang 0011, Chenguang Zhu 0001, Kai-Wei Chang 0001, Yizhou Sun
EMNLP5
2022 Language Models with Image Descriptors are Strong Few-Shot Video-Language Learners
abstract
The goal of this work is to build flexible video-language models that can generalize to various video-to-text tasks from few examples. Existing few-shot video-language learners focus exclusively on the encoder, resulting in the absence of a video-to-text decoder to handle generative tasks. Video captioners have been pretrained on large-scale video-language datasets, but they rely heavily on finetuning and lack the ability to generate text for unseen tasks in a few-shot setting. We propose VidIL, a few-shot Video-language Learner via Image and Language models, which demonstrates strong performance on few-shot video-to-text tasks without the necessity of pretraining or finetuning on any video datasets. We use image-language models to translate the video content into frame captions, object, attribute, and event phrases, and compose them into a temporal-aware template. We then instruct a language model, with a prompt containing a few in-context examples, to generate a target output from the composed content. The flexibility of prompting allows the model to capture any form of text input, such as automatic speech recognition (ASR) transcripts. Our experiments demonstrate the power of language models in understanding videos on a wide variety of video-language tasks, including video captioning, video question answering, video caption retrieval, and video future event prediction. Especially, on video future event prediction, our few-shot model significantly outperforms state-of-the-art supervised models trained on large-scale video datasets.Code and processed data are publicly available for research purposes at https://github.com/MikeWangWZHL/VidIL.
Zhenhailong Wang, Manling Li, Ruochen Xu, Luowei Zhou, Jie Lei 0003, Xudong Lin 0003, Shuohang Wang, Ziyi Yang 0011, Chenguang Zhu 0001, Derek Hoiem, Shih-Fu Chang, Mohit Bansal, Heng Ji 0001
NeurIPS8
2022 Memory-Augmented Generative Adversarial Networks for Anomaly Detection
abstract
We propose a memory-augmented deep learning model for semisupervised anomaly detection (AD). While many traditional AD methods focus on modeling the distribution of normal data, additional constraints in the modeling process are needed to distinguish between normal and abnormal data. The proposed model, named memory augmented generative adversarial networks (MEMGAN), is coupled with external memory units through attentional operations. One property of MEMGAN in the latent space is such that encoded normal data are expected to reside in the convex hull of the memory units, while the abnormal ones are separated outside. This property makes the AD process of MEMGAN more robust and reliable. Experiments on AD datasets adapted from MVTec, MNIST, CIFAR10, and Arrhythmia demonstrate that MEMGAN notably improves over previous AD models. We also find that the decoded memory units in MEMGAN are more diverse and interpretable than those in previous memory-augmented models.
Ziyi Yang 0011, Iman Soltani 0001, Eric Darve
IEEE Trans. Neural Networks Learn. Syst.1
2021 A Simple and Effective Method To Eliminate the Self Language Bias in Multilingual Representations
abstract
Language agnostic and semantic-language information isolation is an emerging research direction for multilingual representations models.We explore this problem from a novel angle of geometric algebra and semantic space.A simple but highly effective method "Language Information Removal (LIR)" factors out language identity information from semantic related components in multilingual representations pre-trained on multi-monolingual data.A post-training and model-agnostic method, LIR only uses simple linear operations, e.g.matrix factorization and orthogonal projection.LIR reveals that for weak-alignment multilingual systems, the principal components of semantic spaces primarily encodes language identity information.We first evaluate the LIR on a cross-lingual question answer retrieval task (LAReQA), which requires the strong alignment for the multilingual embedding space.Experiment shows that LIR is highly effectively on this task, yielding almost 100% relative improvement in MAP for weakalignment models.We then evaluate the LIR on Amazon Reviews and XEVAL dataset, with the observation that removing language information is able to improve the cross-lingual transfer performance.
Ziyi Yang 0011, Yinfei Yang, Daniel M. Cer, Eric Darve
EMNLP (1)1
2021 Universal Sentence Representation Learning with Conditional Masked Language Model
abstract
This paper presents a novel training method, Conditional Masked Language Modeling (CMLM), to effectively learn sentence representations on large scale unlabeled corpora.CMLM integrates sentence representation learning into MLM training by conditioning on the encoded vectors of adjacent sentences.Our English CMLM model achieves state-ofthe-art performance on SentEval (Conneau and Kiela, 2018), even outperforming models learned using supervised signals.As a fully unsupervised learning method, CMLM can be conveniently extended to a broad range of languages and domains.We find that a multilingual CMLM model co-trained with bitext retrieval (BR) and natural language inference (NLI) tasks outperforms the previous state-of-the-art multilingual models by a large margin, e.g.10% improvement upon baseline models on cross-lingual semantic search.We explore the same language bias of the learned representations, and propose a simple, post-training and model agnostic approach to remove the language identifying information from the representation while still retaining sentence semantics.
Ziyi Yang 0011, Yinfei Yang, Daniel M. Cer, Jax Law, Eric Darve
EMNLP (1)1
2021 Leveraging Lead Bias for Zero-shot Abstractive News Summarization
abstract
A typical journalistic convention in news articles is to deliver the most salient information in the beginning, also known as the lead bias. While this phenomenon can be exploited in generating a summary, it has a detrimental effect on teaching a model to discriminate and extract important information in general. We propose that this lead bias can be leveraged in our favor in a simple and effective way to pre-train abstractive news summarization models on large-scale unlabeled news corpora: predicting the leading sentences using the rest of an article. We collect a massive news corpus and conduct data cleaning and filtering via statistical analysis. We then apply self-supervised pre-training on this dataset to existing generation models BART and T5 for domain adaptation. Via extensive experiments on six benchmark datasets, we show that this approach can dramatically improve the summarization quality and achieve state-of-the-art results for zero-shot news summarization without any fine-tuning. For example, in the DUC2003 dataset, the ROUGE-1 score of BART increases 13.7% after the lead-bias pre-training. We deploy the model in Microsoft News and provide public APIs as well as a demo website for multi-lingual news summarization.
Chenguang Zhu 0001, Ziyi Yang 0011, Robert Gmyr, Michael Zeng 0001, Xuedong Huang 0001
SIGIR2
2020 Regularized Cycle Consistent Generative Adversarial Network for Anomaly Detection
abstract
In this paper, we investigate algorithms for anomaly detection. Previous anomaly detection methods focus on modeling the distribution of non-anomalous data provided during training. However, this does not necessarily ensure the correct detection of anomalous data. We propose a new Regularized Cycle Consistent Generative Adversarial Network (RCGAN) in which deep neural networks are adversarially trained to better recognize anomalous samples. This approach is based on leveraging a penalty distribution with a new definition of the loss function and novel use of discriminator networks. It is based on a solid mathematical foundation, and proofs show that our approach has stronger guarantees for detecting anomalous examples compared to the current state-of-the-art. Experimental results on both real-world and synthetic data show that our model leads to significant and consistent improvements on previous anomaly detection benchmarks. Notably, RCGAN improves on the state-of-the-art on the KDDCUP, Arrhythmia, Thyroid, Musk and CIFAR10 datasets.
Ziyi Yang 0011, Iman Soltani 0001, Eric Darve
ECAI1
2019 Embedding Imputation with Grounded Language Information
abstract
Due to the ubiquitous use of embeddings as input representations for a wide range of natural language tasks, imputation of embeddings for rare and unseen words is a critical problem in language processing.Embedding imputation involves learning representations for rare or unseen words during the training of an embedding model, often in a post-hoc manner.In this paper, we propose an approach for embedding imputation which uses grounded information in the form of a knowledge graph.This is in contrast to existing approaches which typically make use of vector space properties or subword information.We propose an online method to construct a graph from grounded information and design an algorithm to map from the resulting graphical structure to the space of the pre-trained embeddings.Finally, we evaluate our approach on a range of rare and unseen word tasks across various domains and show that our model can learn better representations.For example, on the Card-660 task our method improves Pearson's and Spearman's correlation coefficients upon the stateof-the-art by 11% and 17.8% respectively using GloVe embeddings.
Ziyi Yang 0011, Chenguang Zhu 0001, Vin Sachidananda, Eric Darve
ACL (1)1