EDBT 2026 Demo / reviewers in the wild / expert
Ruibo Liu
dblp:86/7648
· DBLP profile ↗
25ranked-venue papers
11as first author
22since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 9 first-author · 18 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Scaling Laws for Differentially Private Language ModelsabstractScaling laws have emerged as important components of large language model (LLM) training as they can predict performance gains through scale, and provide guidance on important hyper-parameter choices that would otherwise be expensive. LLMs also rely on large, high-quality training datasets, like those sourced from (sometimes sensitive) user data. Training models on this sensitive user data requires careful privacy protections like differential privacy (DP). However, the dynamics of DP training are significantly different, and consequently their scaling laws are not yet fully understood. In this work, we establish scaling laws that accurately model the intricacies of DP LLM training, providing a complete picture of the compute-privacy-utility and the optimal training configurations in many settings. Ryan McKenna, Yangsibo Huang, Amer Sinha, Borja Balle, Zachary Charles, Christopher A. Choquette-Choo, Badih Ghazi, Georgios Kaissis, Ravi Kumar 0001, Ruibo Liu, Chiyuan Zhang |
ICML | 10 |
| 2025 | Design2Code: Benchmarking Multimodal Code Generation for Automated Front-End EngineeringabstractChenglei Si, Yanzhe Zhang, Ryan Li, Zhengyuan Yang, Ruibo Liu, Diyi Yang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Chenglei Si, Zhengyuan Yang, Ruibo Liu, Diyi Yang |
NAACL (Long Papers) | 5 |
| 2025 | OmniBench: Towards The Future of Universal Omni-Language ModelsabstractRecent advancements in multimodal large language models (MLLMs) have focused on integrating multiple modalities, yet their ability to simultaneously process and reason across different inputs remains underexplored. We introduce OmniBench, a novel benchmark designed to evaluate models’ ability to recognize, interpret, and reason across visual, acoustic, and textual inputs simultaneously. We define language models capable of such tri-modal processing as omni-language models (OLMs). OmniBench features high-quality human annotations that require integrated understanding across all modalities. Our evaluation reveals that: i) open-source OLMs show significant limitations in instruction-following and reasoning in tri-modal contexts; and ii) most baseline models perform poorly (below 50% accuracy) even with textual alternatives to image/audio inputs. To address these limitations, we develop OmniInstruct, an 96K-sample instruction tuning dataset for training OLMs. We advocate for developing more robust tri-modal integration techniques and training strategies to enhance OLM performance. Codes and data could be found at https://m-a-p.ai/OmniBench/. Ge Zhang 0009, Yinghao Ma, Ruibin Yuan, Kang Zhu, Hangyu Guo, Yiming Liang, Noah Wang, Jian Yang 0003, Siwei Wu, Xingwei Qu, Jinjie Shi, Xinyue Zhang 0005, Zhenzhu Yang, Yidan Wen, Yanghai Wang, Zhaoxiang Zhang 0001, Ruibo Liu, Emmanouil Benetos, Wenhao Huang 0001, Chenghua Lin 0002 |
NeurIPS | 20 |
| 2024 | Exploring Collaboration Mechanisms for LLM Agents: A Social Psychology ViewabstractAs Natural Language Processing (NLP) systems are increasingly employed in intricate social environments, a pressing query emerges: Can these NLP systems mirror human-esque collaborative intelligence, in a multi-agent society consisting of multiple large language models (LLMs)? This paper probes the collaboration mechanisms among contemporary NLP systems by melding practical experiments with theoretical insights. We fabricate four unique ‘societies’ comprised of LLM agents, where each agent is characterized by a specific ‘trait’ (easy-going or overconfident) and engages in collaboration with a distinct ‘thinking pattern’ (debate or reflection). Through evaluating these multi-agent societies on three benchmark datasets, we discern that certain collaborative strategies not only outshine previous top-tier approaches but also optimize efficiency (using fewer API tokens). Moreover, our results further illustrate that LLM agents manifest human-like social behaviors, such as conformity and consensus reaching, mirroring foundational social psychology theories. In conclusion, we integrate insights from social psychology to contextualize the collaboration of LLM agents, inspiring further investigations into the collaboration mechanism for LLMs. We commit to sharing our code and datasets, hoping to catalyze further research in this promising avenue. Jintian Zhang, Xin Xu 0010, Ningyu Zhang 0001, Ruibo Liu, Bryan Hooi, Shumin Deng |
ACL (1) | 4 |
| 2024 | MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised TrainingabstractSelf-supervised learning (SSL) has recently emerged as a promising paradigm for training generalisable models on large-scale data in the fields of vision, text, and speech.
Although SSL has been proven effective in speech and audio, its application to music audio has yet to be thoroughly explored. This is partially due to the distinctive challenges associated with modelling musical knowledge, particularly tonal and pitched characteristics of music.
To address this research gap, we propose an acoustic **M**usic und**ER**standing model with large-scale self-supervised **T**raining (**MERT**), which incorporates teacher models to provide pseudo labels in the masked language modelling (MLM) style acoustic pre-training.
In our exploration, we identified an effective combination of teacher models, which outperforms conventional speech and audio approaches in terms of performance.
This combination includes an acoustic teacher based on Residual Vector Quantization - Variational AutoEncoder (RVQ-VAE) and a musical teacher based on the Constant-Q Transform (CQT).
Furthermore, we explore a wide range of settings to overcome the instability in acoustic language model pre-training, which allows our designed paradigm to scale from 95M to 330M parameters.
Experimental results indicate that our model can generalise and perform well on 14 music understanding tasks and attain state-of-the-art (SOTA) overall scores. Ruibin Yuan, Ge Zhang 0009, Yinghao Ma, Xingran Chen, Hanzhi Yin, Chenghao Xiao, Chenghua Lin 0002, Anton Ragni, Emmanouil Benetos, Norbert Gyenge, Roger B. Dannenberg, Ruibo Liu, Wenhu Chen, Gus Xia, Yemin Shi 0001, Wenhao Huang 0001, Yike Guo, Jie Fu 0001 |
ICLR | 13 |
| 2024 | Training Socially Aligned Language Models on Simulated Social InteractionsabstractThe goal of social alignment for AI systems is to make sure these models can conduct themselves appropriately following social values. Unlike humans who establish a consensus on value judgments through social interaction, current language models (LMs) are trained to rigidly recite the corpus in social isolation, which causes poor generalization in unfamiliar cases and the lack of robustness under adversarial attacks. In this work, we introduce a new training paradigm that enables LMs to learn from simulated social interactions. Compared with existing methods, our method is much more scalable and efficient, and shows superior performance in alignment benchmarks and human evaluation. Ruibo Liu, Chenyan Jia, Diyi Yang, Soroush Vosoughi |
ICLR | 1 |
| 2024 | Reconfigurable Robot Identification from Motion DataabstractIntegrating Large Language Models (LLMs) and Vision-Language Models (VLMs) with robotic systems enables robots to process and understand complex natural language instructions and visual information. However, a fundamental challenge remains: for robots to fully capitalize on these advancements, they must have a deep understanding of their physical embodiment. The gap between AI models’ cognitive capabilities and the understanding of physical embodiment leads to the following question: Can a robot autonomously understand and adapt to its physical form and functionalities through interaction with its environment? This question underscores the transition towards developing self-modeling robots without reliance on external sensory or pre-programmed knowledge about their structure. Here, we propose a meta-self-modeling that can deduce robot morphology through proprioception—the robot’s internal sense of its body’s position and movement. Our study introduces a 12-DoF reconfigurable legged robot, accompanied by a diverse dataset of 200k unique configurations, to systematically investigate the relationship between robotic motion and robot morphology. Utilizing a deep neural network model comprising a robot signature encoder and a configuration decoder, we demonstrate the capability of our system to accurately predict robot configurations from proprioceptive signals. This research contributes to the field of robotic self-modeling, aiming to enhance robot’s understanding of their physical embodiment and adaptability in real-world scenarios. Ruibo Liu, Zhou Shen, Hod Lipson |
IROS | 3 |
| 2024 | II-Bench: An Image Implication Understanding Benchmark for Multimodal Large Language ModelsabstractThe rapid advancements in the development of multimodal large language models (MLLMs) have consistently led to new breakthroughs on various benchmarks. In response, numerous challenging and comprehensive benchmarks have been proposed to more accurately assess the capabilities of MLLMs. However, there is a dearth of exploration of the higher-order perceptual capabilities of MLLMs. To fill this gap, we propose the Image Implication understanding Benchmark, II-Bench, which aims to evaluate the model's higher-order perception of images. Through extensive experiments on II-Bench across multiple MLLMs, we have made significant findings. Initially, a substantial gap is observed between the performance of MLLMs and humans on II-Bench. The pinnacle accuracy of MLLMs attains 74.8%, whereas human accuracy averages 90%, peaking at an impressive 98%. Subsequently, MLLMs perform worse on abstract and complex images, suggesting limitations in their ability to understand high-level semantics and capture image details. Finally, it is observed that most models exhibit enhanced accuracy when image sentiment polarity hints are incorporated into the prompts. This observation underscores a notable deficiency in their inherent understanding of image sentiment. We believe that II-Bench will inspire the community to develop the next generation of MLLMs, advancing the journey towards expert artificial general intelligence (AGI). II-Bench is publicly available at https://huggingface.co/datasets/m-a-p/II-Bench. Feiteng Fang, Xeron Du, Chenhao Zhang 0005, Noah Wang, Yuelin Bai, Qixuan Zhao, Liyang Fan, Chengguang Gan, Hongquan Lin, Jiaming Li 0004, Yuansheng Ni, Haihong Wu, Yaswanth Narsupalli, Zhigang Zheng, Chengming Li 0004, Xiping Hu, Ruifeng Xu 0001, Xiaojun Chen 0006, Min Yang 0007, Ruibo Liu, Wenhao Huang 0001, Ge Zhang 0009, Shiwen Ni |
NeurIPS | 23 |
| 2024 | Long-form factuality in large language modelsabstractLarge language models (LLMs) often generate content that contains factual errors when responding to fact-seeking prompts on open-ended topics. To benchmark a model’s long-form factuality in open domains, we first use GPT-4 to generate LongFact, a prompt set comprising thousands of questions spanning 38 topics. We then propose that LLM agents can be used as automated evaluators for long-form factuality through a method which we call Search-Augmented Factuality Evaluator (SAFE). SAFE utilizes an LLM to break down a long-form response into a set of individual facts and to evaluate the accuracy of each fact using a multi-step reasoning process comprising sending search queries to Google Search and determining whether a fact is supported by the search results. Furthermore, we propose extending F1 score as an aggregated metric for long-form factuality. To do so, we balance the percentage of supported facts in a response (precision) with the percentage of provided facts relative to a hyperparameter representing a user’s preferred response length (recall).
Empirically, we demonstrate that LLM agents can outperform crowdsourced human annotators—on a set of∼16k individual facts, SAFE agrees with crowdsourced human annotators 72% of the time, and on a random subset of 100 disagreement cases, SAFE wins 76% of the time. At the same time, SAFE is more than 20 times cheaper than human annotators. We also benchmark thirteen language models on LongFact across four model families (Gemini, GPT, Claude, and PaLM-2), finding that larger language models generally achieve better long-form factuality. LongFact, SAFE, and all experimental code are available at https://github.com/google-deepmind/long-form-factuality. Jerry Wei, Chengrun Yang, Xinying Song, Yifeng Lu, Nathan Hu, Jie Huang 0009, Dustin Tran, Daiyi Peng, Ruibo Liu, Cosmo Du, Quoc V. Le |
NeurIPS | 9 |
| 2023 | Mind's Eye: Grounded Language Model Reasoning through Simulation
Ruibo Liu, Jason Wei, Shixiang Gu, Te-Yen Wu, Soroush Vosoughi, Claire Cui, Denny Zhou, Andrew M. Dai |
ICLR | 1 |
| 2023 | MARBLE: Music Audio Representation Benchmark for Universal EvaluationabstractIn the era of extensive intersection between art and Artificial Intelligence (AI), such as image generation and fiction co-creation, AI for music remains relatively nascent, particularly in music understanding. This is evident in the limited work on deep music representations, the scarcity of large-scale datasets, and the absence of a universal and community-driven benchmark. To address this issue, we introduce the Music Audio Representation Benchmark for universaL Evaluation, termed MARBLE. It aims to provide a benchmark for various Music Information Retrieval (MIR) tasks by defining a comprehensive taxonomy with four hierarchy levels, including acoustic, performance, score, and high-level description. We then establish a unified protocol based on 18 tasks on 12 public-available datasets, providing a fair and standard assessment of representations of all open-sourced pre-trained models developed on music recordings as baselines. Besides, MARBLE offers an easy-to-use, extendable, and reproducible suite for the community, with a clear statement on copyright issues on datasets. Results suggest recently proposed large-scale pre-trained musical language models perform the best in most tasks, with room for further improvement. The leaderboard and toolkit repository are published to promote future music AI research. Ruibin Yuan, Yinghao Ma, Ge Zhang 0009, Xingran Chen, Hanzhi Yin, Le Zhuo, Zeyue Tian, Binyue Deng, Ningzhi Wang, Chenghua Lin 0002, Emmanouil Benetos, Anton Ragni, Norbert Gyenge, Roger B. Dannenberg, Wenhu Chen, Gus Xia, Wei Xue 0002, Shi Wang 0002, Ruibo Liu, Yike Guo, Jie Fu 0001 |
NeurIPS | 23 |
| 2022 | Non-Parallel Text Style Transfer with Self-Parallel Supervision
Ruibo Liu, Chongyang Gao, Chenyan Jia, Guangxuan Xu, Soroush Vosoughi |
ICLR | 1 |
| 2022 | Knowledge Infused Decoding
Ruibo Liu, Guoqing Zheng, Radhika Gaonkar, Chongyang Gao, Soroush Vosoughi, Milad Shokouhi, Ahmed Awadallah 0001 |
ICLR | 1 |
| 2022 | Second Thoughts are Best: Learning to Re-Align With Human Values from Text EditsabstractWe present Second Thoughts, a new learning paradigm that enables language models (LMs) to re-align with human values. By modeling the chain-of-edits between value-unaligned and value-aligned text, with LM fine-tuning and additional refinement through reinforcement learning, Second Thoughts not only achieves superior performance in three value alignment benchmark datasets but also shows strong human-value transfer learning ability in few-shot scenarios. The generated editing steps also offer better interpretability and ease for interactive error correction. Extensive human evaluations further confirm its effectiveness. Ruibo Liu, Chenyan Jia, Ziyu Zhuang, Tony X. Liu, Soroush Vosoughi |
NeurIPS | 1 |
| 2022 | Quantifying and alleviating political bias in language models
Ruibo Liu, Chenyan Jia, Jason Wei, Guangxuan Xu, Soroush Vosoughi |
Artif. Intell. | 1 |
| 2022 | Dynamic Structural Role Node Embedding for User Modeling in Evolving NetworksabstractComplex user behavior, especially in settings such as social media, can be organized as time-evolving networks. Through network embedding, we can extract general-purpose vector representations of these dynamic networks which allow us to analyze them without extensive feature engineering. Prior work has shown how to generate network embeddings while preserving the structural role proximity of nodes. These methods, however, cannot capture the temporal evolution of the structural identity of the nodes in dynamic networks. Other works, on the other hand, have focused on learning microscopic dynamic embeddings. Though these methods can learn node representations over dynamic networks, these representations capture the local context of nodes and do not learn the structural roles of nodes. In this article, we propose a novel method for learning structural node embeddings in discrete-time dynamic networks. Our method, calledHR2vec, tracks historical topology information in dynamic networks to learn dynamic structural role embeddings. Through experiments on synthetic and real-world temporal datasets, we show that our method outperforms other well-known methods in tasks where structural equivalence and historical information both play important roles.HR2veccan be used to model dynamic user behavior in any networked setting where users can be represented as nodes. Additionally, we propose a novel method (called network fingerprinting) that usesHR2vecembeddings for modeling whole (or partial) time-evolving networks. We showcase our network fingerprinting method on synthetic and real-world networks. Specifically, we demonstrate how our method can be used for detecting foreign-backed information operations on Twitter. Chenghan Huang, Ruibo Liu, Soroush Vosoughi |
ACM Trans. Inf. Syst. | 5 |
| 2021 | Mitigating Political Bias in Language Models through Reinforced CalibrationabstractCurrent large-scale language models can be politically biased as a result of the data they are trained on, potentially causing serious problems when they are deployed in real-world settings. In this paper, we describe metrics for measuring political bias in GPT-2 generation and propose a reinforcement learning (RL) framework for mitigating political biases in generated text. By using rewards from word embeddings or a classifier, our RL framework guides debiased generation without having access to the training data or requiring the model to be retrained. In empirical experiments on three attributes sensitive to political bias (gender, location, and topic), our methods reduced bias according to both our metrics and human evaluation, while maintaining readability and semantic coherence. Ruibo Liu, Chenyan Jia, Jason Wei, Guangxuan Xu, Soroush Vosoughi |
AAAI | 1 |
| 2021 | Embedding Heterogeneous Networks into Hyperbolic Space Without Meta-pathabstractNetworks found in the real-world are numerous and varied. A common type of network is the heterogeneous network, where the nodes (and edges) can be of different types. Accordingly, there have been efforts at learning representations of these heterogeneous networks in low-dimensional space. However, most of the existing heterogeneous network embedding suffers from the following two drawbacks: (1) The target space is usually Euclidean. Conversely, many recent works have shown that complex networks may have hyperbolic latent anatomy, which is non-Euclidean. (2) These methods usually rely on meta-paths, which requires domain-specific prior knowledge for meta-path selection. Additionally, different down-streaming tasks on the same network might require different meta-paths in order to generate task-specific embeddings. In this paper, we propose a novel self-guided random walk method that does not require meta-path for embedding heterogeneous networks into hyperbolic space. We conduct thorough experiments for the tasks of network reconstruction and link prediction on two public datasets, showing that our model outperforms a variety of well-known baselines across all tasks. Chongyang Gao, Chenghan Huang, Ruibo Liu, Soroush Vosoughi |
AAAI | 4 |
| 2021 | Language Model Augmented Relevance ScoreabstractRuibo Liu, Jason Wei, Soroush Vosoughi. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Ruibo Liu, Jason Wei, Soroush Vosoughi |
ACL/IJCNLP (1) | 1 |
| 2021 | Political Depolarization of News Articles Using Attribute-Aware Word Embeddings
Ruibo Liu, Chenyan Jia, Soroush Vosoughi |
ICWSM | 1 |
| 2021 | Hyperbolic node embedding for temporal networks
Chenghan Huang, Ruibo Liu, Soroush Vosoughi |
Data Min. Knowl. Discov. | 4 |
| 2021 | A Transformer-based Framework for Neutralizing and Reversing the Political Polarity of News ArticlesabstractPeople often prefer to consume news with similar political predispositions and access like-minded news articles, which aggravates polarized clusters known as "echo chamber". To mitigate this phenomenon, we propose a computer-aided solution to help combat extreme political polarization. Specifically, we present a framework for reversing or neutralizing the political polarity of news headlines and articles. The framework leverages the attention mechanism of a Transformer-based language model to first identify polar sentences, and then either flip the polarity to the neutral or to the opposite through a GAN network. Tested on the same benchmark dataset, our framework achieves a 3%-10% improvement on the flipping/neutralizing success rate of headlines compared with the current state-of-the-art model. Adding to prior literature, our framework not only flips the polarity of headlines but also extends the task of polarity flipping to full-length articles. Human evaluation results show that our model successfully neutralizes or reverses the polarity of news without reducing readability. We release a large annotated dataset that includes both news headlines and full-length articles with polarity labels and meta-data to be used for future research. Our framework has a potential to be used by social scientists, content creators and content consumers in the real world. Ruibo Liu, Chenyan Jia, Soroush Vosoughi |
Proc. ACM Hum. Comput. Interact. | 1 |
| 2020 | Data Boost: Text Data Augmentation Through Reinforcement Learning Guided Conditional GenerationabstractData augmentation is proven to be effective in many NLU tasks, especially for those suffering from data scarcity.In this paper, we present a powerful and easy to deploy text augmentation framework, Data Boost, which augments data through reinforcement learning guided conditional generation.We evaluate Data Boost on three diverse text classification tasks under five different classifier architectures.The result shows that Data Boost can boost the performance of classifiers especially in low-resource data scenarios.For instance, Data Boost improves F1 for the three tasks by 8.7% on average when given only 10% of the whole data for training.We also compare Data Boost with six prior text augmentation methods.Through human evaluations (N =178), we confirm that Data Boost augmentation has comparable quality as the original data with respect to readability and class consistency. Ruibo Liu, Guangxuan Xu, Chenyan Jia, Soroush Vosoughi |
EMNLP (1) | 1 |
| 2020 | Multi-resolution Annotations for Emoji PredictionabstractEmojis are able to express various linguistic components, including emotions, sentiments, events, etc. Predicting the proper emojis associated with text provides a way to summarize the text accurately, and it has been proven to be a good auxiliary task to many Natural Language Understanding (NLU) tasks.Labels in existing emoji prediction datasets are all passage-based and are usually under the multi-class classification setting.However, in many cases, one single emoji cannot fully cover the theme of a piece of text.It is thus useful to infer the part of text related to each emoji.The lack of multi-label and aspectlevel emoji prediction datasets is one of the bottlenecks for this task.This paper annotates an emoji prediction dataset with passage-level multi-class/multi-label, and aspect-level multiclass annotations.We also present a novel annotation method with which we generate the aspect-level annotations.The annotations are generated heuristically, taking advantage of the self-attention mechanism in Transformer networks.We validate the annotations both automatically and manually to ensure their quality.We also benchmark the dataset with a pretrained BERT model. Ruibo Liu, Soroush Vosoughi |
EMNLP (1) | 2 |
| 2020 | Salienteye: Maximizing Engagement While Maintaining Artistic Style on Instagram Using Deep Neural NetworksabstractInstagram has become a great venue for amateur and professional photographers alike to showcase their work. It has, in other words, democratized photography. Generally, photographers take thousands of photos in a session, from which they pick a few to showcase their work on Instagram. Photographers trying to build a reputation on Instagram have to strike a balance between maximizing their followers' engagement with their photos, while also maintaining their artistic style. We used transfer learning to adapt Xception, which is a model for object recognition trained on the ImageNet dataset, to the task of engagement prediction and utilized Gram matrices generated from VGG19, another object recognition model trained on ImageNet, for the task of style similarity measurement on photos posted on Instagram. Our models can be trained on individual Instagram accounts to create personalized engagement prediction and style similarity models. Once trained on their accounts, users can have new photos sorted based on predicted engagement and style similarity to their previous work, thus enabling them to upload photos that not only have the potential to maximize engagement from their followers but also maintain their style of photography. We trained and validated our models on several Instagram accounts, showing it to be adept at both tasks, also outperforming several baseline models and human annotators. Ruibo Liu, Soroush Vosoughi |
ICMR | 2 |