VLDB 2026 Research / reviewers in the wild / expert
Chenyang Lyu
dblp:248/1663
· DBLP profile ↗
37ranked-venue papers
12as first author
35since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 4 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 8 since 2021Security and privacy · 7 · 3 first-author · 6 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GigaMoE: Sparsity-Guided Mixture of Experts for Efficient Gigapixel Object DetectionabstractObject detection in High-Resolution Wide (HRW) shots, or gigapixel images, presents unique challenges due to extreme object sparsity and vast scale variations. State-of-the-art methods like SparseFormer have pioneered sparse processing by selectively focusing on important regions, yet they apply a uniform computational model to all selected regions, overlooking their intrinsic complexity differences. This leads to a suboptimal trade-off between performance and efficiency. In this paper, we introduce GigaMoE, a novel backbone architecture that pioneers adaptive computation for this domain by replacing the standard Feed-Forward Networks (FFNs) with a Mixture-of-Experts (MoE) module. Our architecture first employs a shared expert to provide a robust feature baseline for all selected regions. Upon this foundation, our core innovation---a novel Sparsity-Guided Routing mechanism---insightfully repurposes importance scores from the sparse backbone to provide a "computational bonus,'' dynamically engaging a variable number of specialized experts based on content complexity. The entire system is trained efficiently via a loss-free load-balancing technique, eliminating the need for cumbersome auxiliary losses. Extensive experiments show that GigaMoE sets a new state-of-the-art on the PANDA benchmark, improving detection accuracy by 1.1% over SparseFormer while simultaneously reducing the computational cost (FLOPs) by a remarkable 32.3%. Wenxi Li, Yuetong Wang, Chenyang Lyu, Haozhe Lin, Guiguang Ding |
AAAI | 4 |
| 2026 | CoQuIR: A Comprehensive Benchmark for Code Quality-Aware Information RetrievalabstractJiahui Geng, Fengyu Cai, Shaobo Cui, Qing Li, Liangwei Chen, Chenyang Lyu, Haonan Li, Derui Zhu, Alexander Pretschner, Heinz Koeppl, Fakhri Karray. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Jiahui Geng, Fengyu Cai, Shaobo Cui 0006, Qing Li 0038, Liangwei Chen, Chenyang Lyu, Derui Zhu, Alexander Pretschner, Heinz Koeppl, Fakhri Karray |
ACL (1) | 6 |
| 2026 | New Trends for Modern Machine Translation with Large Reasoning ModelsabstractRecent advances in Large Reasoning Models (LRMs), particularly those leveraging Chain-of-Thought reasoning (CoT), have opened brand new possibility for Machine Translation (MT). This position paper argues that LRMs substantially transformed traditional neural MT as well as LLMs-based MT paradigms by reframing translation as a dynamic reasoning task that requires contextual, cultural, and linguistic understanding and reasoning. We identify three foundational shifts: 1) contextual coherence, where LRMs resolve ambiguities and preserve discourse structure through explicit reasoning over cross-sentence and complex context or even lack of context; 2) cultural intentionality, enabling models to adapt outputs by inferring speaker intent, audience expectations, and socio-linguistic norms; 3) self-reflection, LRMs can perform self-reflection during the inference time to correct the potential errors in translation especially extremely noisy cases, showing better robustness compared to simply mapping X->Y translation. We explore various scenarios in translation including stylized translation, document-level translation and multimodal translation by showcasing empirical examples that demonstrate the superiority of LRMs in translation. We also identify several interesting phenomenons for LRMs for MT including auto-pivot translation as well as the critical challenges such as over-localisation in translation and inference efficiency. In conclusion, we think that LRMs redefine translation systems not merely as text converters but as multilingual cognitive agents capable of reasoning about meaning beyond the text. This paradigm shift reminds us to think of problems in translation beyond traditional translation scenarios in a much broader context with LRMs - what we can achieve on top of it. Sinuo Liu, Chenyang Lyu, Minghao Wu, Zifu Shang, Longyue Wang, Weihua Luo, Kaifu Zhang |
LREC | 2 |
| 2026 | Enhancing 3D medical multi-modal large language models with integrated human body priors for computed tomography
Leilei Zeng, Jie Liu 0044, Wenting Chen, Chenyang Lyu, Wenxi Li, Shaonan Liu, Xiande Zhou, LinLin Shen |
Pattern Recognit. | 4 |
| 2025 | HD-NDEs: Neural Differential Equations for Hallucination Detection in LLMsabstract6173 Qing Li 0038, Jiahui Geng, Zongxiong Chen, Derui Zhu, Yuxia Wang 0003, Congbo Ma, Chenyang Lyu, Fakhri Karray |
ACL (1) | 7 |
| 2025 | Marco-o1 v2: Towards Widening The Distillation Bottleneck for Reasoning ModelsabstractLarge Reasoning Models (LRMs) such as OpenAI o1 and DeepSeek-R1 have shown remarkable reasoning capabilities by scaling test-time compute and generating long Chain-of-Thought (CoT). Distillation post-training on LRMs-generated data is a straightforward yet effective method to enhance the reasoning abilities of smaller models, but faces a critical bottleneck: we found that distilled long CoT data poses learning difficulty for small models and leads to the inheritance of biases (i.e., formalistic long-time thinking) when using Supervised Fine-tuning (SFT) and Reinforcement Learning (RL) methods. To alleviate this bottleneck, we propose constructing data from scratch using Monte Carlo Tree Search (MCTS). We then exploit a set of CoT-aware approaches, including Thoughts Length Balance, Fine-grained DPO, and Joint Post-training Objective, to enhance SFT and RL on the MCTS data. We conducted evaluation on various benchmarks such as math (GSM8K, MATH, AIME). instruction-following (Multi-IF) and planning (Blocksworld), results demonstrate our CoT-aware approaches substantially improve the reasoning performance of distilled models compared to standard distilled models via reducing the hallucinations in long-time thinking. Huifeng Yin, Minghao Wu, Xuanfan Ni, Tianqi Shi, Liangying Shao, Chenyang Lyu, Longyue Wang, Weihua Luo, Kaifu Zhang |
ACL (1) | 9 |
| 2025 | Marco-Bench-MIF: On Multilingual Instruction-Following Capability of Large LanguageabstractBo Zeng, Chenyang Lyu, Sinuo Liu, Mingyan Zeng, Minghao Wu, Xuanfan Ni, Tianqi Shi, Yu Zhao, Yefeng Liu, Chenyu Zhu, Ruizhe Li, Jiahui Geng, Qing Li, Yu Tong, Longyue Wang, Weihua Luo, Kaifu Zhang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Chenyang Lyu, Sinuo Liu, Mingyan Zeng, Minghao Wu, Xuanfan Ni, Tianqi Shi, Yefeng Liu, Chenyu Zhu, Ruizhe Li 0001, Jiahui Geng, Longyue Wang, Weihua Luo, Kaifu Zhang |
ACL (1) | 2 |
| 2025 | From Multiple-Choice to Extractive QA: A Case Study for English and ArabicabstractThe rapid evolution of Natural Language Processing (NLP) has favoured major languages such as English, leaving a significant gap for many others due to limited resources. This is especially evident in the context of data annotation, a task whose importance cannot be underestimated, but which is time-consuming and costly. Thus, any dataset for resource-poor languages is precious, in particular when it is task-specific. Here, we explore the feasibility of repurposing an existing multilingual dataset for a new NLP task: we repurpose a subset of the BELEBELE dataset (Bandarkar et al., 2023), which was designed for multiple-choice question answering (MCQA), to enable the more practical task of extractive QA (EQA) in the style of machine reading comprehension. We present annotation guidelines and a parallel EQA dataset for English and Modern Standard Arabic (MSA). We also present QA evaluation results for several monolingual and cross-lingual QA pairs including English, MSA, and five Arabic dialects. We aim to help others adapt our approach for the remaining 120 BELEBELE language variants, many of which are deemed under-resourced. We also provide a thorough analysis and share insights to deepen understanding of the challenges and opportunities in NLP task reformulation. Teresa Lynn, Malik H. Altakrori, Samar Mohamed Magdy, Rocktim Jyoti Das, Chenyang Lyu, Mohamed Nasr, Younes Samih, Kirill Chirkunov, Alham Fikri Aji, Preslav Nakov, Shantanu Godbole, Salim Roukos, Radu Florian, Nizar Habash |
COLING | 5 |
| 2025 | Enhancing Video-Text Matching via Sparse Stratified SamplingabstractVideo-text matching is a critical task in multimedia retrieval, but traditional methods often fail to capture the diversity and depth of video content due to inefficient and inaccurate frame sampling. We propose a novel sparse stratified sampling technique that can substantially improve the video-text matching process by segmenting video content into clusters based on relevant features and selectively sampling representative frames. Our method further introduces a threshold for the feature metric used to divide clusters, eliminating video frames with low relevance. We propose two variants of our approach: an offline approach that performs sampling before training, and an online approach that dynamically conducts sampling based on the relevance between video frames and the text query during training. Extensive experiments on datasets like MSRVTT and AVSD for video retrieval and multiple-choice VideoQA datasets, including AVQA and Music-AVQA, demonstrate the superiority of our method over previous state-of-the-art approaches. Our sparse stratified sampling technique achieves improvements of over 1.2% on MSRVTT and 1.7% on AVSD for R@1 in video retrieval tasks. For multiple-choice VideoQA tasks, our approach achieves significant improvements of 1.8% accuracy on AVQA and 3.9% on Music-AVQA, strongly supporting its effectiveness in enhancing video-text matching systems. Chenyang Lyu, Wenxi Li, Tianbo Ji, Liting Zhou, Pintu Lohar, Yi Yu 0001, Longyue Wang |
ICASSP | 1 |
| 2025 | Retrieval-Augmented Multi-Modal Chain-of-Thoughts Reasoning for Large Language ModelsabstractThe advancement of Large Language Models (LLMs) has brought substantial attention to the Chain of Thought (CoT) approach, primarily due to its ability to enhance the capability of LLMs on complex reasoning tasks. Moreover, the CoT approach extends to multi-modal tasks of LLMs. However, the selection of optimal CoT demonstration examples for LLMs in multi-modal reasoning remains less explored due to the inherent complexity of multi-modal examples. In this paper, we introduce a novel approach that addresses this challenge by using retrieval mechanisms to dynamically select demonstration examples based on cross-modal and intra-modal similarities. Furthermore, we employ a Group Selection method to select examples containing rationales from different retrieval directions to promote the diversity of demonstration examples. To the best of our knowledge, we are the first to apply Retrieval-Augmented Generation (RAG) with CoT to complex multi-modal reasoning tasks. Through a series of experiments on two popular benchmarks, ScienceQA and MathVista, we demonstrate that our approach significantly improves the performance of GPT-4 by 6% on ScienceQA and 12.9% on MathVista. Additionally, it enhances the performance of GPT-4V on these two datasets by 2.7%, respectively, further advancing the capabilities of the most advanced LLMs and Large Multimodal Models (LMMs) for complex multi-modal reasoning tasks. Bingshuai Liu, Chenyang Lyu, Zijun Min, Zhanyu Wang, Jinsong Su, Longyue Wang |
IJCNN | 2 |
| 2025 | EditEval: Towards Comprehensive and Automatic Evaluation for Text-guided Video EditingabstractRecently, video editing task has gained widespread attention due to its practical applications and rapid advancements. However, current automatic evaluation metrics for video editing are mostly poorly aligned with human judgments. Thus, researchers heavily rely on human evaluation, which is not only labor-intensive but also difficult to ensure consistency and objectivity. To address these issues, we propose EditEval, the largest-ever video editing benchmark to comprehensively evaluate the performance of video editing models in three aspects: Textual Faithfulness, Frame Consistency, and Video Fidelity. It includes 200 video clips and 1,010 text prompts, from which 160 instances are sampled to generate 1,280 edited videos using eight open-source video editing models, accompanied by human annotations. Furthermore, we propose EditScore, leveraging the advanced reasoning and comprehension capabilities of Multi-modal Large Language Models (MLLMs) as evaluators to assess edited videos across the aforementioned aspects. Experiments show that the best-performing video editing model only reaches an average score of 3.16 (out of a perfect 5), highlighting the challenge of EditEval. Besides, results from more than 10 MLLMs demonstrate the great potential of utilizing EditScore for automatic evaluation. Notably, for textual faithfulness, EditScore equipped with LLaVA-OneVision-7B achieves a significantly higher Pearson Correlation score compared to previous methods based on CLIP (0.50 vs 0.22). The code and dataset are available at: https://github.com/XMUDeepLIT/EditEval Bingshuai Liu, Ante Wang, Zijun Min, Chenyang Lyu, Longyue Wang, Xu Han 0007, Peng Li 0030, Jinsong Su |
ACM Multimedia | 4 |
| 2025 | Rethinking Document Layout Analysis through Text Clustering via Multi-Modal Graph Convolution NetworksabstractDocument layout analysis, a critical process in automated document processing, traditionally relies on object detection techniques, primarily focusing on the structural segmentation of documents. However, these approaches often fall short in comprehensively understanding the semantic content within the text, leading to a disjointed analysis of document structure and content. To address this, we propose a novel methodology that combines text clustering with multi-modal graph convolution networks, aiming to integrate structural detection with semantic understanding. Our approach starts with text detection, followed by encoding using a large language model. Subsequently, we integrate visual and positional data using Graph Neural Networks to perform clustering, creating a synergy between the textual and structural aspects of documents. Extensive experiments on mainstream datasets demonstrate that our method significantly outperforms existing approaches, especially in understanding text-centric document layouts. This paper contributes to the field by offering a novel, semantically-enriched approach to document layout analysis, enhancing the capabilities of automated document processing systems in handling diverse and complex document formats. Wenxi Li, Chenyang Lyu, Liting Zhou, Cathal Gurrin |
MMSP | 2 |
| 2025 | UniAVLM: Unified Large Audio-Visual Language Models for Comprehensive Video Understanding
Lecheng Yan, Chenyang Lyu, Wenxi Li, Younes Samih, Shaochen Jiang |
PRICAI (5) | 2 |
| 2025 | On the Taxonomy, Tasks, and Open-Challenges for Multimodal Large Language ModelsabstractIn recent years, the field of Artificial Intelligence has witnessed the emergence of Multimodal Large Language Models (MLLMs) that have significantly advanced the state-of-the-art in understanding and generating content across various data modalities. These models, capable of processing and integrating information from text, images, audio, and video, have opened new avenues for research and applications. Distinguished by their ability to understand and generation information with diverse modalities, such as text, image, audio and many others, MLLMs mark a significant step towards the final aim of Artificial General Intelligence (AGI). This comprehensive survey provides an in-depth examination of MLLMs, highlighting their evolutionary trajectory, current state-of-the-art developments, and prospective future directions. Specifically, we show taxonomy of MLLMs by their modalities to be processed and model architecture for aligning multiple modalities. Besides, we also present discussion regarding the different types of tasks related to MLLMs. The paper further delves into the pressing challenges confronted in this domain, such as data scarcity, computational complexity, ethical dilemmas, and privacy considerations. We analyze these issues in the context of both development and deployment of MLLMs. The survey comprehensively demonstrate and summarise the recent advances of the transformative influence of MLLMs while acknowledging their potential limitations, thereby outlining a prospective roadmap for future research endeavors in this rapidly developing field. Lecheng Yan, Jiahui Geng, Minghao Wu, Zhanyu Wang, Wenxi Li, Tianbo Ji, Shaochen Jiang, Chenyang Lyu |
SMC | 10 |
| 2024 | A Paradigm Shift: The Future of Machine Translation Lies with Large Language ModelsabstractMachine Translation (MT) has greatly advanced over the years due to the developments in deep neural networks. However, the emergence of Large Language Models (LLMs) like GPT-4 and ChatGPT is introducing a new phase in the MT domain. In this context, we believe that the future of MT is intricately tied to the capabilities of LLMs. These models not only offer vast linguistic understandings but also bring innovative methodologies, such as prompt-based techniques, that have the potential to further elevate MT. In this paper, we provide an overview of the significant enhancements in MT that are influenced by LLMs and advocate for their pivotal role in upcoming MT research and implementations. We highlight several new MT directions, emphasizing the benefits of LLMs in scenarios such as Long-Document Translation, Stylized Translation, and Interactive Translation. Additionally, we address the important concern of privacy in LLM-driven MT and suggest essential privacy-preserving strategies. By showcasing practical instances, we aim to demonstrate the advantages that LLMs offer, particularly in tasks like translating extended documents. We conclude by emphasizing the critical role of LLMs in guiding the future evolution of MT and offer a roadmap for future exploration in the sector. Chenyang Lyu, Zefeng Du, Jitao Xu 0003, Yitao Duan, Minghao Wu, Teresa Lynn, Alham Fikri Aji, Derek F. Wong, Longyue Wang |
LREC/COLING | 1 |
| 2024 | On the Cultural Gap in Text-to-Image GenerationabstractOne challenge in text-to-image (T2I) generation is the inadvertent reflection of culture gaps present in the training data, which signifies the disparity in generated image quality when the cultural elements of the input text are rarely collected in the training set. Although various T2I models have shown impressive but arbitrary examples, there is no benchmark to systematically evaluate a T2I model’s ability to generate cross-cultural images. To bridge the gap, we propose a Challenging Cross-Cultural (C3) benchmark with comprehensive evaluation criteria, which can assess how well-suited a model is to a target culture. By analyzing the flawed images generated by the Stable Diffusion model on the C3 benchmark, we find that the model often fails to generate certain cultural objects. Accordingly, we propose a novel multi-modal metric that considers object-text alignment to filter the fine-tuning data in the target culture, which is used to fine-tune a T2I model to improve cross-cultural generation. Experimental results show that our multi-modal metric provides stronger data selection performance on the C3 benchmark than existing metrics, in which the object-text alignment is crucial. We release the benchmark, data, code, and generated images to facilitate future research on culturally diverse T2I generation. Bingshuai Liu, Longyue Wang, Chenyang Lyu, Yong Zhang 0034, Jinsong Su, Shuming Shi 0001, Zhaopeng Tu |
ECAI | 3 |
| 2024 | Semantic Enrichment for Video Question Answering with Gated Graph Neural NetworksabstractVideo Question Answering (VideoQA) is a complex task that requires a deep understanding of a video to accurately answer questions. Existing methods often struggle to effectively integrate the visual and language-based semantic information, subsequently leading to an incomplete understanding of video content and sub-optimal performance. To address the challenge, we introduce a novel approach in this paper to enrich the semantics of video frames, questions, and answer candidates. Specifically, we parse video frames and questions into semantic graphs - visual semantic graph and question semantic graph, which captures information about objects, their attributes, and relationships. These graphs are then encoded using a Gated Graph Neural Network (GGNN). For answer candidates, we propose to verbalize them using Large Language Models (LLMs) to further inject more semantic information from visual and acoustic aspects. We evaluate our approach on benchmark VideoQA datasets: AVQA and Music-AVQA. Experimental results show that our approach outperforms competitive baseline models, achieving state-of-the-art performance on various question types. Chenyang Lyu, Wenxi Li, Tianbo Ji, Yi Yu 0001, Longyue Wang |
ICASSP | 1 |
| 2024 | GPT4Video: A Unified Multimodal Large Language Model for lnstruction-Followed Understanding and Safety-Aware Generation
Zhanyu Wang, Longyue Wang, Zhen Zhao 0001, Minghao Wu, Chenyang Lyu, Deng Cai 0002, Luping Zhou, Shuming Shi 0001, Zhaopeng Tu |
ACM Multimedia | 5 |
| 2024 | CVQA: Culturally-diverse Multilingual Visual Question Answering BenchmarkabstractVisual Question Answering~(VQA) is an important task in multimodal AI, which requires models to understand and reason on knowledge present in visual and textual data. However, most of the current VQA datasets and models are primarily focused on English and a few major world languages, with images that are Western-centric. While recent efforts have tried to increase the number of languages covered on VQA datasets, they still lack diversity in low-resource languages. More importantly, some datasets extend the text to other languages, either via translation or some other approaches, but usually keep the same images, resulting in narrow cultural representation. To address these limitations, we create CVQA, a new Culturally-diverse Multilingual Visual Question Answering benchmark dataset, designed to cover a rich set of languages and regions, where we engage native speakers and cultural experts in the data collection process. CVQA includes culturally-driven images and questions from across 28 countries in four continents, covering 26 languages with 11 scripts, providing a total of 9k questions. We benchmark several Multimodal Large Language Models (MLLMs) on CVQA, and we show that the dataset is challenging for the current state-of-the-art models. This benchmark will serve as a probing evaluation suite for assessing the cultural bias of multimodal models and hopefully encourage more research efforts towards increasing cultural awareness and linguistic diversity in this field. Chenyang Lyu, Haryo Akbarianto Wibowo, Santiago Góngora, Aishik Mandal, Sukannya Purkayastha, Jesús-Germán Ortiz-Barajas, Emilio Villa-Cueva, Jinheon Baek, Soyeong Jeong, Injy Hamed, Zheng Wei Lim, Paula Mónica Silva, Jocelyn Dunstan, Mélanie Jouitteau, David Le Meur, Joan Nwatu, Ganzorig Batnasan, Munkh-Erdene Otgonbold, Munkhjargal Gochoo, Guido Ivetta, Luciana Benotti, Laura Alonso Alemany, Hernán Maina, Jiahui Geng, Tiago Timponi Torrent, Frederico Belcavello, Marcelo Viridiano, Jan Christian Blaise Cruz, Dan John Velasco, Oana Ignat, Zara Burzo, Chenxi Whitehouse, Artem Abzaliev, Teresa Clifford, Grainne Caulfield, Teresa Lynn, Christian Salamea Palacios, Vladimir Araujo, Yova Kementchedjhieva, Mihail Mihaylov, Israel Abebe Azime, Henok Biadglign Ademtew, Bontu Fufa Balcha, Naome A. Etori, David Ifeoluwa Adelani, Rada Mihalcea, Atnafu Lambebo Tonja, Maria Camila Buitrago Cabrera, Gisela Vallejo, Holy Lovenia, Ruochen Zhang 0001, Marcos Estecha-Garitagoitia, Mario Rodríguez-Cantelar, Toqeer Ehsan, Rendi Chevi, Muhammad Farid Adilazuarda, Ryandito Diandaru, Samuel Cahyawijaya, Fajri Koto, Tatsuki Kuribayashi, Haiyue Song, Aditya Khandavally, Thanmay Jayakumar, Raj Dabre, Mohamed Fazli Mohamed Imam, Kumaranage Ravindu Yasas Nagasinghe, Alina Dragonetti, Luis Fernando D'Haro, Olivier Niyomugisha, Jay Gala, Pranjal A. Chitale, Fauzan Farooqui, Thamar Solorio, Alham Fikri Aji |
NeurIPS | 2 |
| 2024 | SyzTrust: State-aware Fuzzing on Trusted OS Designed for IoT DevicesabstractTrusted Execution Environments (TEEs) embedded in IoT devices provide a deployable solution to secure IoT applications at the hardware level. By design, in TEEs, the Trusted Operating System (Trusted OS) is the primary component. It enables the TEE to use security-based design techniques, such as data encryption and identity authentication. Once a Trusted OS has been exploited, the TEE can no longer ensure security. However, Trusted OSes for IoT devices have received little security analysis, which is challenging from several perspectives: (1) Trusted OSes are closed-source and have an unfavorable environment for sending test cases and collecting feedback. (2) Trusted OSes have complex data structures and require a stateful workflow, which limits existing vulnerability detection tools.To address the challenges, we present SyzTrust, the first state-aware fuzzing framework for vetting the security of resource-limited Trusted OSes. SyzTrust adopts a hardware-assisted framework to enable fuzzing Trusted OSes directly on IoT devices as well as tracking state and code coverage non-invasively. SyzTrust utilizes composite feedback to guide the fuzzer to effectively explore more states as well as to increase the code coverage. We evaluate SyzTrust on Trusted OSes from three major vendors: Samsung, Tsinglink Cloud, and Ali Cloud. These systems run on Cortex M23/33 MCUs, which provide the necessary abstraction for embedded TEEs. We discovered 70 previously unknown vulnerabilities in their Trusted OSes, receiving 10 new CVEs so far. Furthermore, compared to the baseline, SyzTrust has demonstrated significant improvements, including 66% higher code coverage, 651% higher state coverage, and 31% improved vulnerability-finding capability. We report all discovered new vulnerabilities to vendors and open source SyzTrust. Qinying Wang, Boyu Chang, Shouling Ji, Yuan Tian 0001, Xuhong Zhang 0002, Chenyang Lyu, Mathias Payer, Wenhai Wang, Raheem A. Beyah |
SP | 8 |
| 2024 | One Bad Apple Spoils the Barrel: Understanding the Security Risks Introduced by Third-Party Components in IoT FirmwareabstractCurrently, the development of IoT firmware heavily depends on third-party components (TPCs) to improve development efficiency. Nevertheless, TPCs are not secure, and the vulnerabilities in TPCs will influence the security of IoT firmware. Existing works pay less attention to the vulnerabilities caused by TPCs, and we still lack a comprehensive understanding of the security impact of TPC vulnerability against firmware. To fill in the knowledge gap, we design and implementFirmSec, which leverages syntactical features and control-flow graph features to detect the TPCs in firmware, and then recognizes the corresponding vulnerabilities. Based onFirmSec, we present the first large-scale analysis of the security risks raised by TPCs on 34,136 firmware images. We successfully detect 584 TPCs and identify 128,757 vulnerabilities caused by 429 CVEs. Our in-depth analysis reveals the diversity of security risks in firmware and discovers some well-known vulnerabilities are still rooted in firmware. Besides, we explore the geographical distribution of vulnerable devices and confirm that the security situation of devices in different regions varies. Our analysis also indicates that vulnerabilities caused by TPCs in firmware keep growing with the boom of the IoT ecosystem. Further analysis shows 2,478 commercial firmware images have potentially violated GPL/AGPL licensing terms. Shouling Ji, Jiacheng Xu 0006, Yuan Tian 0001, Qiuyang Wei, Qinying Wang, Chenyang Lyu, Xuhong Zhang 0002, Changting Lin, JingZheng Wu, Raheem A. Beyah |
IEEE Trans. Dependable Secur. Comput. | 7 |
| 2023 | Dialogue-to-Video Retrieval
Chenyang Lyu, Duy Nguyen 0003, Van-Tu Ninh, Liting Zhou, Cathal Gurrin, Jennifer Foster |
ECIR (2) | 1 |
| 2023 | Document-Level Machine Translation with Large Language ModelsabstractLarge language models (LLMs) such as Chat-GPT can produce coherent, cohesive, relevant, and fluent answers for various natural language processing (NLP) tasks.Taking documentlevel machine translation (MT) as a testbed, this paper provides an in-depth evaluation of LLMs' ability on discourse modeling.The study focuses on three aspects: 1) Effects of Context-Aware Prompts, where we investigate the impact of different prompts on document-level translation quality and discourse phenomena; 2) Comparison of Translation Models, where we compare the translation performance of Chat-GPT with commercial MT systems and advanced document-level MT methods; 3) Analysis of Discourse Modelling Abilities, where we further probe discourse knowledge encoded in LLMs and shed light on impacts of training techniques on discourse modeling.By evaluating on a number of benchmarks, we surprisingly find that LLMs have demonstrated superior performance and show potential to become a new paradigm for document-level translation: 1) leveraging their powerful long-text modeling capabilities, GPT-3.5 and GPT-4 outperform commercial MT systems in terms of human evaluation; 1 2) GPT-4 demonstrates a stronger ability for probing linguistic knowledge than GPT-3.5.This work highlights the challenges and opportunities of LLMs for MT, which we hope can inspire the future design and evaluation of LLMs. 2 * Equal contribution. Longyue Wang, Chenyang Lyu, Tianbo Ji, Zhirui Zhang, Dian Yu 0001, Shuming Shi 0001, Zhaopeng Tu |
EMNLP | 2 |
| 2023 | Out-of-Distribution Generalization in Natural Language Processing: Past, Present, and FutureabstractLinyi Yang, Yaoxian Song, Xuan Ren, Chenyang Lyu, Yidong Wang, Jingming Zhuo, Lingqiao Liu, Jindong Wang, Jennifer Foster, Yue Zhang. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Linyi Yang, Yaoxian Song, Xuan Ren, Chenyang Lyu, Jingming Zhuo, Lingqiao Liu, Jindong Wang 0001, Jennifer Foster, Yue Zhang 0004 |
EMNLP | 4 |
| 2023 | Gated Multi-modal Fusion with Cross-modal Contrastive Learning for Video Question Answering
Chenyang Lyu, Wenxi Li, Tianbo Ji, Liting Zhou, Cathal Gurrin |
ICANN (7) | 1 |
| 2023 | Graph-Based Video-Language Learning with Multi-Grained Audio-Visual AlignmentabstractVideo-language learning has attracted significant attention in the fields of multimedia, computer vision and natural language processing in recent years. One of the key challenges in this area is how to effectively integrate visual and linguistic information to enable machines to understand video content and query information. In this work, we leverage graph-based representations and multi-grained audio-visual alignment to address this challenge. First, our approach starts by transforming video and query inputs into visual-scene graphs and semantic role graphs using a visual-scene parser and semantic role labeler respectively. These graphs are then encoded using graph neural networks to obtain enriched representations and combined to obtain a video-query joint representation that enhances the semantic expressivity of the inputs. Second, to achieve accurate matching of relevant parts of audio and visual features, we propose a multi-grained alignment module that aligns the audio and visual features at multiple scales. This enables us to effectively fuse the audio and visual information in a way that is consistent with the semantic-level information captured by the graph-based representations. Experiments on five representative datasets collected for Video Retrieval and Video Question Answering tasks show that our approach outperforms the literature on several metrics. Our extensive ablation studies demonstrate the effectiveness of graph-based representation and multi-grained audio-visual alignment. Chenyang Lyu, Wenxi Li, Tianbo Ji, Longyue Wang, Liting Zhou, Cathal Gurrin, Linyi Yang, Yi Yu 0001, Yvette Graham, Jennifer Foster |
ACM Multimedia | 1 |
| 2023 | MINER: A Hybrid Data-Driven Approach for REST API Fuzzing
Chenyang Lyu, Jiacheng Xu 0006, Shouling Ji, Xuhong Zhang 0002, Qinying Wang, Peng Cheng 0001, Raheem A. Beyah |
USENIX Security Symposium | 1 |
| 2023 | UVSCAN: Detecting Third-Party Component Usage Violations in IoT Firmware
Shouling Ji, Xuhong Zhang 0002, Yuan Tian 0001, Qinying Wang, Yuwen Pu, Chenyang Lyu, Raheem A. Beyah |
USENIX Security Symposium | 7 |
| 2022 | Achieving Reliable Human Assessment of Open-Domain Dialogue SystemsabstractEvaluation of open-domain dialogue systems is highly challenging and development of better techniques is highlighted time and again as desperately needed.Despite substantial efforts to carry out reliable live evaluation of systems in recent competitions, annotations have been abandoned and reported as too unreliable to yield sensible results.This is a serious problem since automatic metrics are not known to provide a good indication of what may or may not be a high-quality conversation.Answering the distress call of competitions that have emphasized the urgent need for better evaluation techniques in dialogue, we present the successful development of human evaluation that is highly reliable while still remaining feasible and low cost.Self-replication experiments reveal almost perfectly repeatable results with a correlation of r = 0.969.Furthermore, due to the lack of appropriate methods of statistical significance testing, the likelihood of potential improvements to systems occurring due to chance is rarely taken into account in dialogue evaluation, and the evaluation we propose facilitates application of standard tests.Since we have developed a highly reliable evaluation method, new insights into system performance can be revealed.We therefore include a comparison of state-of-the-art models (i) with and without personas, to measure the contribution of personas to conversation quality, as well as (ii) prescribed versus freely chosen topics.Interestingly with respect to personas, results indicate that personas do not positively contribute to conversation quality as expected. Tianbo Ji, Yvette Graham, Gareth J. F. Jones, Chenyang Lyu, Qun Liu 0001 |
ACL (1) | 4 |
| 2022 | SLIME: program-sensitive energy allocation for fuzzingabstractThe energy allocation strategy is one of the most popular techniques in fuzzing to improve code coverage and vulnerability discovery. The core intuition is that fuzzers should allocate more computational energy to the seed files that have high efficiency to trigger unique paths and crashes after mutation. Existing solutions usually define several properties, e.g., the execution speed, the file size, and the number of the triggered edges in the control flow graph, to serve as the key measurements in their allocation logics to estimate the potential of a seed. The efficiency of a property is usually assumed to be the same across different programs. However, we find that this assumption is not always valid. As a result, the state-of-the-art energy allocation solutions with static energy allocation logics are hard to achieve desirable performance on different programs. Chenyang Lyu, Shouling Ji, Xuhong Zhang 0002, Zhe Wang 0017, Wenhai Wang, Raheem A. Beyah |
ISSTA | 1 |
| 2022 | A large-scale empirical analysis of the vulnerabilities introduced by third-party components in IoT firmwareabstractAs the core of IoT devices, firmware is undoubtedly vital. Currently, the development of IoT firmware heavily depends on third-party components (TPCs), which significantly improves the development efficiency and reduces the cost. Nevertheless, TPCs are not secure, and the vulnerabilities in TPCs will turn back influence the security of IoT firmware. Currently, existing works pay less attention to the vulnerabilities caused by TPCs, and we still lack a comprehensive understanding of the security impact of TPC vulnerability against firmware. To fill in the knowledge gap, we design and implement FirmSec, which leverages syntactical features and control-flow graph features to detect the TPCs at version-level in firmware, and then recognizes the corresponding vulnerabilities. Based on FirmSec, we present the first large-scale analysis of the usage of TPCs and the corresponding vulnerabilities in firmware. More specifically, we perform an analysis on 34,136 firmware images, including 11,086 publicly accessible firmware images, and 23,050 private firmware images from TSmart. We successfully detect 584 TPCs and identify 128,757 vulnerabilities caused by 429 CVEs. Our in-depth analysis reveals the diversity of security issues for different kinds of firmware from various vendors, and discovers some well-known vulnerabilities are still deeply rooted in many firmware images. We also find that the TPCs used in firmware have fallen behind by five years on average. Besides, we explore the geographical distribution of vulnerable devices, and confirm the security situation of devices in several regions, e.g., South Korea and China, is more severe than in other regions. Further analysis shows 2,478 commercial firmware images have potentially violated GPL/AGPL licensing terms. Shouling Ji, Jiacheng Xu 0006, Yuan Tian 0001, Qiuyang Wei, Qinying Wang, Chenyang Lyu, Xuhong Zhang 0002, Changting Lin, JingZheng Wu, Raheem A. Beyah |
ISSTA | 7 |
| 2022 | EMS: History-Driven Mutation for Coverage-based Fuzzing
Chenyang Lyu, Shouling Ji, Xuhong Zhang 0002, Kangjie Lu, Raheem A. Beyah |
NDSS | 1 |
| 2022 | V-Fuzz: Vulnerability Prediction-Assisted Evolutionary Fuzzing for Binary ProgramsabstractFuzzing is a technique of finding bugs by executing a target program recurrently with a large number of abnormal inputs. Most of the coverage-based fuzzers consider all parts of a program equally and pay too much attention to how to improve the code coverage. It is inefficient as the vulnerable code only takes a tiny fraction of the entire code. In this article, we design and implement an evolutionary fuzzing framework called V-Fuzz, which aims to find bugs efficiently and quickly in limited time for binary programs. V-Fuzz consists of two main components: 1) a vulnerability prediction model and 2) a vulnerability-oriented evolutionary fuzzer. Given a binary program to V-Fuzz, the vulnerability prediction model will give a prior estimation on which parts of a program are more likely to be vulnerable. Then, the fuzzer leverages an evolutionary algorithm to generate inputs which are more likely to arrive at the vulnerable locations, guided by the vulnerability prediction result. The experimental results demonstrate that V-Fuzz can find bugs efficiently with the assistance of vulnerability prediction. Moreover, V-Fuzz has discovered ten common vulnerabilities and exposures (CVEs), and three of them are newly discovered. Yuwei Li 0002, Shouling Ji, Chenyang Lyu, Jianhai Chen, Qinchen Gu, Chunming Wu 0001, Raheem A. Beyah |
IEEE Trans. Cybern. | 3 |
| 2021 | Improving Unsupervised Question Answering via Summarization-Informed Question GenerationabstractQuestion Generation (QG) is the task of generating a plausible question for a given pair.Template-based QG uses linguistically-informed heuristics to transform declarative sentences into interrogatives, whereas supervised QG uses existing Question Answering (QA) datasets to train a system to generate a question given a passage and an answer.A disadvantage of the heuristic approach is that the generated questions are heavily tied to their declarative counterparts.A disadvantage of the supervised approach is that they are heavily tied to the domain/language of the QA dataset used as training data.In order to overcome these shortcomings, we propose an unsupervised QG method which uses questions generated heuristically from summaries as a source of training data for a QG system.We make use of freely available news summary data, transforming declarative summary sentences into appropriate questions using heuristics informed by dependency parsing, named entity recognition and semantic role labeling.The resulting questions are then combined with the original news articles to train an end-to-end neural QG model.We extrinsically evaluate our approach using unsupervised QA: our QG model is used to generate synthetic QA pairs for training a QA model.Experimental results show that, trained with only 20k English Wikipedia-based synthetic QA pairs, the QA model substantially outperforms previous unsupervised models on three in-domain datasets (SQuAD1.1,Natural Questions, TriviaQA) and three out-of-domain datasets (NewsQA, BioASQ, DuoRC), demonstrating the transferability of the approach. Chenyang Lyu, Lifeng Shang, Yvette Graham, Jennifer Foster, Xin Jiang 0002, Qun Liu 0001 |
EMNLP (1) | 1 |
| 2021 | UNIFUZZ: A Holistic and Pragmatic Metrics-Driven Platform for Evaluating Fuzzers
Yuwei Li 0002, Shouling Ji, Sizhuang Liang, Wei-Han Lee, Yueyao Chen, Chenyang Lyu, Chunming Wu 0001, Raheem A. Beyah, Peng Cheng 0001, Kangjie Lu, Ting Wang 0006 |
USENIX Security Symposium | 7 |
| 2020 | Improving Document-Level Sentiment Analysis with User and Product ContextabstractPast work that improves document-level sentiment analysis by encoding user and product information has been limited to considering only the text of the current review.We investigate incorporating additional review text available at the time of sentiment prediction that may prove meaningful for guiding prediction.Firstly, we incorporate all available historical review text belonging to the author of the review in question.Secondly, we investigate the inclusion of historical reviews associated with the current product (written by other users).We achieve this by explicitly storing representations of reviews written by the same user and about the same product and force the model to memorize all reviews for one particular user and product.Additionally, we drop the hierarchical architecture used in previous work to enable words in the text to directly attend to each other.Experiment results on IMDB, Yelp 2013 and Yelp 2014 datasets show improvement to state-of-the-art of more than 2 percentage points in the best case. Chenyang Lyu, Jennifer Foster, Yvette Graham |
COLING | 1 |
| 2019 | MOPT: Optimized Mutation Scheduling for Fuzzers
Chenyang Lyu, Shouling Ji, Chao Zhang 0008, Yuwei Li 0002, Wei-Han Lee, Raheem A. Beyah |
USENIX Security Symposium | 1 |