Jiahui Geng

dblp:228/5625 · DBLP profile ↗
← Back
26ranked-venue papers
8as first author
25since 2021 · last 2026
0000-0002-4205-8230ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 5 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CoQuIR: A Comprehensive Benchmark for Code Quality-Aware Information Retrieval
abstract
Jiahui Geng, Fengyu Cai, Shaobo Cui, Qing Li, Liangwei Chen, Chenyang Lyu, Haonan Li, Derui Zhu, Alexander Pretschner, Heinz Koeppl, Fakhri Karray. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Jiahui Geng, Fengyu Cai, Shaobo Cui 0006, Qing Li 0038, Liangwei Chen, Chenyang Lyu, Derui Zhu, Alexander Pretschner, Heinz Koeppl, Fakhri Karray
ACL (1)1
2026 SGPVT: Self-Generated Proximal Visual Tokens for Mitigating Proximal Collateral Damage in MLLM Unlearning
abstract
Jiaqi Li, Zhijing Zhang, Jiahui Geng, Sheng Bi, Chuanyi Zhang, Fan Liu, Guilin Qi. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Jiaqi Li 0031, Zhijing Zhang, Jiahui Geng, Chuanyi Zhang, Fan Liu 0003, Guilin Qi
ACL (1)3
2026 Is Human-Like Text Liked by Humans? Multilingual Human Detection and Preference Against AI
abstract
Yuxia Wang, Rui Xing, Jonibek Mansurov, Giovanni Puccetti, Zhuohan Xie, Minh Ngoc Ta, Jiahui Geng, Jinyan Su, Mervat Abassy, Saadeldine Eletter, Kareem Elozeiri, Nurkhan Laiyk, Maiya Goloburda, Tarek Mahmoud, Raj Vardhan Tomar, Alexander Aziz, Ryuto Koike, Masahiro Kaneko, Artem Shelmanov, Ekaterina Artemova, Vladislav Mikhailov, Akim Tsvigun, Alham Fikri Aji, Nizar Habash, Iryna Gurevych, Preslav Nakov. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yuxia Wang 0003, Rui Xing 0002, Jonibek Mansurov, Giovanni Puccetti 0002, Zhuohan Xie, Minh Ngoc Ta, Jiahui Geng, Jinyan Su, Mervat Abassy, Saadeldine Eletter, Kareem Ashraf Elozeiri, Nurkhan Laiyk, Maiya Goloburda, Tarek Mahmoud, Raj Vardhan Tomar, Alexander Aziz, Ryuto Koike, Masahiro Kaneko, Artem Shelmanov, Ekaterina Artemova, Vladislav Mikhailov, Akim Tsvigun, Alham Fikri Aji, Nizar Habash, Iryna Gurevych, Preslav Nakov
ACL (1)7
2026 The CLEF-2026 FinMMEval Lab: Multilingual and Multimodal Evaluation of Financial AI Systems
Zhuohan Xie, Rania Elbadry, Fan Zhang 0019, Georgi Georgiev 0001, Xueqing Peng, Lingfei Qian, Jimin Huang, Dimitar Dimitrov 0003, Vanshikaa Jani, Yuyang Dai, Jiahui Geng, Yuxia Wang 0003, Ivan Koychev, Veselin Stoyanov, Preslav Nakov
ECIR (4)11
2025 Internal Activation Revision: Safeguarding Vision Language Models Without Parameter Update
abstract
Warning: This paper contains offensive content that may disturb some readers. Vision-language models (VLMs) demonstrate strong multimodal capabilities but have been found to be more susceptible to generating harmful content compared to their backbone large language models (LLMs). Our investigation reveals that the integration of images significantly shifts the model's internal activations during the forward pass, diverging from those triggered by textual input. Moreover, the safety alignments of LLMs embedded within VLMs are not sufficiently robust to handle the activations discrepancies, making the models vulnerable to even the simplest jailbreaking attacks. To address this issue, we propose an internal activation revision approach that efficiently revises activations during generation, steering the model toward safer outputs. Our framework incorporates revisions at both the layer and head levels, offering control over the model's generation at varying levels of granularity. In addition, we explore three strategies for constructing positive and negative samples and two approaches for extracting revision vectors, resulting in different variants of our method. Comprehensive experiments demonstrate that the internal activation revision method significantly improves the safety of widely used VLMs, reducing attack success rates by an average of 48.94%, 34.34%, 43.92%, and 52.98% on SafeBench, Safe-Unsafe, Unsafe, and MM-SafetyBench, respectively, while minimally impacting model helpfulness.
Qing Li 0038, Jiahui Geng, Derui Zhu, Zongxiong Chen, Kun Song 0001, Lei Ma 0003, Fakhri Karray
AAAI2
2025 HD-NDEs: Neural Differential Equations for Hallucination Detection in LLMs
abstract
6173
Qing Li 0038, Jiahui Geng, Zongxiong Chen, Derui Zhu, Yuxia Wang 0003, Congbo Ma, Chenyang Lyu, Fakhri Karray
ACL (1)2
2025 Shaping the Safety Boundaries: Understanding and Defending Against Jailbreaks in Large Language Models
abstract
This paper contains model outputs that may be considered offensive.
Lang Gao, Jiahui Geng, Xiangliang Zhang 0001, Preslav Nakov, Xiuying Chen
ACL (1)2
2025 \mathsfCon Instruction: Universal Jailbreaking of Multimodal Large Language Models via Non-Textual Modalities
abstract
Existing attacks against multimodal language models (MLLMs) primarily communicate instructions through text accompanied by adversarial images.In contrast, here we exploit the capabilities of MLLMs to interpret non-textual instructions-specifically adversarial images or audio-generated by our novel method, Con Instruction.We optimize the adversarial examples to align closely with target instructions in the embedding space, revealing the detrimental aspects of sophisticated understanding in MLLMs.Unlike previous work, our method does not require training data or preprocessing of textual instructions.While these non-textual adversarial examples can effectively bypass MLLMs safety mechanisms, their combination with various text inputs substantially amplifies attack success.We further introduce a new attack response categorization (ARC) that considers both response quality and relevance to the malicious instructions to evaluate attack success.The results show that Con Instruction effectively bypasses the safety mechanisms in various visual and audio-language models, including LLaVA-v1.5,InternVL, Qwen-VL, and Qwen-Audio, across two standard benchmarks: AdvBench and SafeBench.Specifically, our method achieves the highest attack success rates, reaching 81.3% and 86.6% on LLaVA-v1.5 (13B).On the defense side, we explore various methods against our attacks and find a substantial gap among existing techniques.Our implementation is made available.1Warning: This paper contains examples that may be offensive to some readers.
Jiahui Geng, Thy Thy Tran, Preslav Nakov, Iryna Gurevych
ACL (1)1
2025 Marco-Bench-MIF: On Multilingual Instruction-Following Capability of Large Language
abstract
Bo Zeng, Chenyang Lyu, Sinuo Liu, Mingyan Zeng, Minghao Wu, Xuanfan Ni, Tianqi Shi, Yu Zhao, Yefeng Liu, Chenyu Zhu, Ruizhe Li, Jiahui Geng, Qing Li, Yu Tong, Longyue Wang, Weihua Luo, Kaifu Zhang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Chenyang Lyu, Sinuo Liu, Mingyan Zeng, Minghao Wu, Xuanfan Ni, Tianqi Shi, Yefeng Liu, Chenyu Zhu, Ruizhe Li 0001, Jiahui Geng, Longyue Wang, Weihua Luo, Kaifu Zhang
ACL (1)12
2025 OpenFactCheck: Building, Benchmarking Customized Fact-Checking Systems and Evaluating the Factuality of Claims and LLMs
abstract
The increased use of large language models (LLMs) across a variety of real-world applications calls for mechanisms to verify the fac- tual accuracy of their outputs. Difficulties lie in assessing the factuality of free-form responses in open domains. Also, different pa- pers use disparate evaluation benchmarks and measurements, which renders them hard to compare and hampers future progress. To mitigate these issues, we propose OpenFactCheck, a unified framework for building customized automatic fact-checking systems, benchmarking their accuracy, evaluating factuality of LLMs, and verifying claims in a document. OpenFactCheck consists of three modules: (i) CUSTCHECKER allows users to easily customize an automatic fact-checker and verify the factual correctness of documents and claims, (ii) LLMEVAL, a unified evaluation framework assesses LLM’s factuality ability from various perspectives fairly, and (iii) CHECKEREVAL is an extensible solution for gauging the reliability of automatic fact-checkers’ verification results using human-annotated datasets. Data and code are publicly available at https://github.com/yuxiaw/openfactcheck.
Yuxia Wang 0003, Minghan Wang, Georgi Georgiev 0001, Jiahui Geng, Iryna Gurevych, Preslav Nakov
COLING5
2025 SAUCE: Selective Concept Unlearning in Vision-Language Models with Sparse Autoencoders
Jiahui Geng
ICCV1
2025 On the Taxonomy, Tasks, and Open-Challenges for Multimodal Large Language Models
abstract
In recent years, the field of Artificial Intelligence has witnessed the emergence of Multimodal Large Language Models (MLLMs) that have significantly advanced the state-of-the-art in understanding and generating content across various data modalities. These models, capable of processing and integrating information from text, images, audio, and video, have opened new avenues for research and applications. Distinguished by their ability to understand and generation information with diverse modalities, such as text, image, audio and many others, MLLMs mark a significant step towards the final aim of Artificial General Intelligence (AGI). This comprehensive survey provides an in-depth examination of MLLMs, highlighting their evolutionary trajectory, current state-of-the-art developments, and prospective future directions. Specifically, we show taxonomy of MLLMs by their modalities to be processed and model architecture for aligning multiple modalities. Besides, we also present discussion regarding the different types of tasks related to MLLMs. The paper further delves into the pressing challenges confronted in this domain, such as data scarcity, computational complexity, ethical dilemmas, and privacy considerations. We analyze these issues in the context of both development and deployment of MLLMs. The survey comprehensively demonstrate and summarise the recent advances of the transformative influence of MLLMs while acknowledging their potential limitations, thereby outlining a prospective roadmap for future research endeavors in this rapidly developing field.
Lecheng Yan, Jiahui Geng, Minghao Wu, Zhanyu Wang, Wenxi Li, Tianbo Ji, Shaochen Jiang, Chenyang Lyu
SMC3
2025 Federated Large Domain Model System
abstract
As organizations increasingly seek to build Foundation Models (FMs) using their own proprietary data, many are adopting private and in-house cloud infrastructures (often in addition to public clouds) to address concerns over cost, data privacy, and data sovereignty. However, these isolated private clouds frequently lack interoperability, creating barriers to cross-institutional collaboration, which is vital for training robust Domain-Specific Foundation Models (DSFMs) that rely on large and diverse datasets. Additionally, underutilized resources in private clouds lead to significant global energy inefficiencies. In this paper, we propose the Federated Large Domain Model System (FLDMS), a conceptual framework designed to facilitate collaborative foundation model development across multiple private cloud environments. We review the necessary enabling technologies, including decentralized protocols for data privacy and Large Language Models (LLMs) for automated orchestration, and present a high-level system design demonstrating how these components can be integrated. By enabling secure and efficient cross-organization cooperation, FLDMS provides a blueprint for building DSFMs while addressing the inefficiencies inherent in siloed private cloud systems.
Chunming Rong, Jungwon Seo, Ferhat Özgür Çatak, Jiahui Geng, Martin Gilje Jaatun
Blockchain Res. Appl.5
2024 Towards Trustworthy Dataset Distillation: A Benchmark of Privacy, Fairness and Robustness
abstract
Dataset distillation is an increasingly prevalent technique for condensing a large-scale dataset into more compact versions while preserving their intrinsic utility. However, very few studies have investigated the trustworthiness of data distillation, i.e., privacy, robustness, and fairness. The deficiency is particularly striking given the existing research that underscores the vulnerabilities in current AI models, including privacy breaches, biased predictions against underrepresented subgroups, and susceptibility to imperceptible attacks. To bridge the gap, we propose a trustworthy benchmark for assessing representative dataset distillation solutions across the benchmark CIFAR10 with comprehensive evaluation metrics. Through extensive experiments, we uncover vulnerabilities inherent in the application of dataset distillation, offering valuable insights for practitioners. Our work aims to drive the development of more transparent, reliable, and responsible machine learning models, fostering AI systems that align with trustworthy principles.
Zongxiong Chen, Jiahui Geng, Derui Zhu, Qing Li 0038, Sonja Schimmler, Manfred Hauswirth
IJCNN2
2024 A Survey of Confidence Estimation and Calibration in Large Language Models
abstract
Jiahui Geng, Fengyu Cai, Yuxia Wang, Heinz Koeppl, Preslav Nakov, Iryna Gurevych. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Jiahui Geng, Fengyu Cai, Yuxia Wang 0003, Heinz Koeppl, Preslav Nakov, Iryna Gurevych
NAACL-HLT1
2024 CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark
abstract
Visual Question Answering~(VQA) is an important task in multimodal AI, which requires models to understand and reason on knowledge present in visual and textual data. However, most of the current VQA datasets and models are primarily focused on English and a few major world languages, with images that are Western-centric. While recent efforts have tried to increase the number of languages covered on VQA datasets, they still lack diversity in low-resource languages. More importantly, some datasets extend the text to other languages, either via translation or some other approaches, but usually keep the same images, resulting in narrow cultural representation. To address these limitations, we create CVQA, a new Culturally-diverse Multilingual Visual Question Answering benchmark dataset, designed to cover a rich set of languages and regions, where we engage native speakers and cultural experts in the data collection process. CVQA includes culturally-driven images and questions from across 28 countries in four continents, covering 26 languages with 11 scripts, providing a total of 9k questions. We benchmark several Multimodal Large Language Models (MLLMs) on CVQA, and we show that the dataset is challenging for the current state-of-the-art models. This benchmark will serve as a probing evaluation suite for assessing the cultural bias of multimodal models and hopefully encourage more research efforts towards increasing cultural awareness and linguistic diversity in this field.
Chenyang Lyu, Haryo Akbarianto Wibowo, Santiago Góngora, Aishik Mandal, Sukannya Purkayastha, Jesús-Germán Ortiz-Barajas, Emilio Villa-Cueva, Jinheon Baek, Soyeong Jeong, Injy Hamed, Zheng Wei Lim, Paula Mónica Silva, Jocelyn Dunstan, Mélanie Jouitteau, David Le Meur, Joan Nwatu, Ganzorig Batnasan, Munkh-Erdene Otgonbold, Munkhjargal Gochoo, Guido Ivetta, Luciana Benotti, Laura Alonso Alemany, Hernán Maina, Jiahui Geng, Tiago Timponi Torrent, Frederico Belcavello, Marcelo Viridiano, Jan Christian Blaise Cruz, Dan John Velasco, Oana Ignat, Zara Burzo, Chenxi Whitehouse, Artem Abzaliev, Teresa Clifford, Grainne Caulfield, Teresa Lynn, Christian Salamea Palacios, Vladimir Araujo, Yova Kementchedjhieva, Mihail Mihaylov, Israel Abebe Azime, Henok Biadglign Ademtew, Bontu Fufa Balcha, Naome A. Etori, David Ifeoluwa Adelani, Rada Mihalcea, Atnafu Lambebo Tonja, Maria Camila Buitrago Cabrera, Gisela Vallejo, Holy Lovenia, Ruochen Zhang 0001, Marcos Estecha-Garitagoitia, Mario Rodríguez-Cantelar, Toqeer Ehsan, Rendi Chevi, Muhammad Farid Adilazuarda, Ryandito Diandaru, Samuel Cahyawijaya, Fajri Koto, Tatsuki Kuribayashi, Haiyue Song, Aditya Khandavally, Thanmay Jayakumar, Raj Dabre, Mohamed Fazli Mohamed Imam, Kumaranage Ravindu Yasas Nagasinghe, Alina Dragonetti, Luis Fernando D'Haro, Olivier Niyomugisha, Jay Gala, Pranjal A. Chitale, Fauzan Farooqui, Thamar Solorio, Alham Fikri Aji
NeurIPS26
2024 PrivAuditor: Benchmarking Data Protection Vulnerabilities in LLM Adaptation Techniques
abstract
Large Language Models (LLMs) are recognized for their potential to be an important building block toward achieving artificial general intelligence due to their unprecedented capability for solving diverse tasks. Despite these achievements, LLMs often underperform in domain-specific tasks without training on relevant domain data. This phenomenon, which is often attributed to distribution shifts, makes adapting pre-trained LLMs with domain-specific data crucial. However, this adaptation raises significant privacy concerns, especially when the data involved come from sensitive domains. In this work, we extensively investigate the privacy vulnerabilities of adapted (fine-tuned) LLMs and benchmark privacy leakage across a wide range of data modalities, state-of-the-art privacy attack methods, adaptation techniques, and model architectures. We systematically evaluate and pinpoint critical factors related to privacy leakage. With our organized codebase and actionable insights, we aim to provide a standardized auditing tool for practitioners seeking to deploy customized LLM applications with faithful privacy assessments.
Derui Zhu, Dingfan Chen, Xiongfei Wu, Jiahui Geng, Zhuo Li 0021, Jens Grossklags, Lei Ma 0003
NeurIPS4
2024 Improved Gradient Inversion Attacks and Defenses in Federated Learning
abstract
Gradient inversion attacks can reconstruct the victim's private data once they have access to the victim's model and gradient. However, existing research is still immature, and many attacks are conducted in ideal conditions. It is unclear how damaging such attacks really are and how they can be effectively defended. In this paper, we first summarize the current relevant researches and their limitations. Then we design a general gradient inversion attack framework, which can attack both FedSGD and FedAVG. We propose approaches to enhance the label inference and image restoration, respectively. Our approach surpasses the SOTA attacks, by successfully attacking the batches from ImageNet while other methods fail to attack. Finally, we suggest several defense strategies without any utility loss from extensive experiments. We are confirmed that our work makes people aware of the privacy issues and can actively avoid the potential risks.
Jiahui Geng, Yongli Mou, Qing Li 0038, Oya Beyan, Stefan Decker, Chunming Rong
IEEE Trans. Big Data1
2023 A Survey on Dataset Distillation: Approaches, Applications and Future Directions
abstract
Dataset distillation is attracting more attention in machine learning as training sets continue to grow and the cost of training state-of-the-art models becomes increasingly high. By synthesizing datasets with high information density, dataset distillation offers a range of potential applications, including support for continual learning, neural architecture search, and privacy protection. Despite recent advances, we lack a holistic understanding of the approaches and applications. Our survey aims to bridge this gap by first proposing a taxonomy of dataset distillation, characterizing existing approaches, and then systematically reviewing the data modalities, and related applications. In addition, we summarize the challenges and discuss future directions for this field of research.
Jiahui Geng, Zongxiong Chen, Yuandou Wang, Herbert Woisetschlaeger, Sonja Schimmler, Ruben Mayer, Zhiming Zhao, Chunming Rong
IJCAI1
2023 pFedV: Mitigating Feature Distribution Skewness via Personalized Federated Learning with Variational Distribution Constraints
Yongli Mou, Jiahui Geng, Feng Zhou 0011, Oya Beyan, Chunming Rong, Stefan Decker
PAKDD (2)2
2022 NFT as a proof of Digital Ownership-reward system integrated to a Secure Distributed Computing Blockchain Framework
abstract
Today, the global economy is dependent on the Internet and computational resources. Although they are tightly interconnected, it is difficult to evaluate their degree of interdependence. Keeping up with the pace of technology can be a challenging task, mainly when updating the hardware and software infrastructure. Every day, corporations and governments are faced with this issue; most have been victims of cyber attacks, security breaches, and data leaks. The consequences are significant in monetary losses; damage remediation is unattainable, even impossible, in certain circumstances. The repercussions might include reputational damage, legal responsibility, and threats to national security (when attacks are carried out against critical infrastructures to control the resources of a country), to name a few. Similarly, data has become such an integral part of many industries that it is one of the most critical targets for attackers that often is encrypted by ransomware, stolen, or corrupted. Without data, many companies are not able to continue operating as they do. The combination of all these factors complicates the ability of organizations to cooperate, trust, and share information in efforts to research and develop solutions for industry and government.This work proposes a Blockchain-based infrastructure solution provided by “Hyperledger Fabric” technology for companies to securely transmit and share information using the latest encryption and data storage technologies operating on the model of distributed systems and smart contracts. By presenting unique digital assets as Non-Fungible Tokens (NFT), the infrastructure is able to trust the integrity of the data, while protecting it from counterfeiting. Through the use of a Blockchain-based file storage system known as IPFS, and by connecting all the relevant elements together through a web-based application, it is possible to demonstrate that the implementation of such systems is feasible, highly scalable and a useful tool that many organizations can utilize to create new work systems and worktflows for digital asset management.
Asahi Cantu, Jiahui Geng, Chunming Rong
CloudCom2
2022 Blockchain-based Cross-organizational Workflow Platform
abstract
Data-centric workflows across organizations are gaining more and more popularity. To automate this process, the traditional approaches centralise related data from different organizations to the cloud, and then use a workflow engine to complete cross-organizational collaboration. There are limitations of those approaches, such as the requirement of data centralisation which could lead to the leakage of critical data. In this work, we present a workflow platform for consuming distributed data based on Kubernetes and JupyterFlow, and use blockchain technology to guarantee security and privacy through empowering the data owner with control of their own data. To reduce replication and network throughput, the blockchain only contains the meta data referring to the data in off-chain storage. We develope a JupyterHub extension to support data registration and query, and used the RESTFul API to connect the web application with the blockchain network. Finally, we demonstrate a simple data processing use case as proof-of-concept to validate our proposed platform.
Jiahui Geng, Ali Akbar Rehman, Yongli Mou, Stefan Decker, Chunming Rong
CloudCom1
2022 Managing Digital Objects with Decentralised Identifiers based on NFT-like schema
abstract
The diversity (text, images, algorithms, etc.) and the ambiguity of data sovereignty and privacy make the management of digital objects very challenging. Users need a unified and convenient way to manage their digital objects. This places a high demand on the findability and interoperability of the management model. In recent years blockchain has provided a new route to an open and secure platform due to its attributes such as distributed, traceable, and tamper-evident. In this paper, a new NFT-like scheme is proposed, which uses metadata converts digital assets into digital object identifiers, and transforms digital objects that require clear sovereignty into NFTs to ensure the authenticity and uniqueness of ownership. Our scheme can facilitate the dynamic management of digital objects using smart contracts.
Chunming Rong, Jiahui Geng, Martin Gilje Jaatun
CloudCom2
2022 Blockchain Empowered and Self-sovereign Access Control System
abstract
Lack of trustworthiness, access policy flexibility, and user privacy preservation in centralized access control systems raise numerous security issues and reduce the collaboration maturity of global data sharing systems. In this paper, we propose a Self-Sovereign Identity-based, Decentralized, and Dynamic (SSIDD) access control system. SSIDD utilizes blockchain technologies to build trust for untrusted data sharing networks and ensures user privacy. Our access control provides high access policy flexibility and security for global inter-enterprise collaborations from a diverse industrial environment. SSIDD authenticates its users based on their Decentralized Identifiers (DID), which are under control of users and can be resolved into a DID document stored on the blockchain. Our data management technology keeps the data sharing systems safe against issues such as data breaches, identity thefts, and privacy violations. Besides, the authorization process of SSIDD is dynamic by adopting several smart contracts. The transparency of rules and agreements in smart contracts and the traceability of records on blockchain ledger provide a high level of security and trust. For proof of concept, we have developed and evaluated a prototype of SSIDD. Our evaluations show that the throughput and latency of our method are within an acceptable range.
Hanif Tadjik, Jiahui Geng, Martin Gilje Jaatun, Chunming Rong
CloudCom2
2021 DID-eFed: Facilitating Federated Learning as a Service with Decentralized Identities
abstract
We have entered the era of big data, and it is considered to be the ”fuel” for the flourishing of artificial intelligence applications. The enactment of the EU General Data Protection Regulation (GDPR) raises concerns about individuals’ privacy in big data. Federated learning (FL) emerges as a functional solution that can help build high-performance models shared among multiple parties while still complying with user privacy and data confidentiality requirements. Although FL has been intensively studied and used in real applications, there is still limited research related to its prospects and applications as a FLaaS (Federated Learning as a Service) to interested 3rd parties. In this paper, we present a FLaaS system: DID-eFed, where FL is facilitated by decentralized identities (DID) and a smart contract. DID enables a more flexible and credible decentralized access management in our system, while the smart contract offers a frictionless and less error-prone process. We describe particularly the scenario where our DID-eFed enables the FLaaS among hospitals and research institutions.
Jiahui Geng, Neel Kanwal, Martin Gilje Jaatun, Chunming Rong
EASE1
2018 Improving Unsupervised Word-by-Word Translation with Language Model and Denoising Autoencoder
abstract
Unsupervised learning of cross-lingual word embedding offers elegant matching of words across languages, but has fundamental limitations in translating sentences.In this paper, we propose simple yet effective methods to improve word-by-word translation of crosslingual embeddings, using only monolingual corpora but without any back-translation.We integrate a language model for context-aware search, and use a novel denoising autoencoder to handle reordering.Our system surpasses state-of-the-art unsupervised neural translation systems without costly iterative training.We also analyze the effect of vocabulary size and denoising type on the translation performance, which provides better understanding of learning the cross-lingual word embedding and its usage in translation.
Yunsu Kim 0001, Jiahui Geng, Hermann Ney
EMNLP2