Yuejian Fang

dblp:119/3697 · DBLP profile ↗
← Back
34ranked-venue papers
2as first author
28since 2021 · last 2026
0000-0002-8279-6908ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 18 · 1 first-author · 17 since 2021Artificial intelligence and machine learning · 10 · 9 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021Security and privacy · 4 · 1 first-authorSystems, architecture and hardware · 2 · 1 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2026 DiffBench Meets DiffAgent: End-to-End LLM-Driven Diffusion Acceleration Code Generation
abstract
Diffusion models have achieved remarkable success in image and video generation. However, their inherently multiple step inference process imposes substantial computational overhead, hindering real-world deployment. Accelerating diffusion models is therefore essential, yet determining how to combine multiple model acceleration techniques remains a significant challenge. To address this issue, we introduce a framework driven by large language models (LLMs) for automated acceleration code generation and evaluation. First, we present DiffBench, a comprehensive benchmark that implements a three stage automated evaluation pipeline across diverse diffusion architectures, optimization combinations and deployment scenarios. Second, we propose DiffAgent, an agent that generates optimal acceleration strategies and codes for arbitrary diffusion models. DiffAgent employs a closed-loop workflow in which a planning component and a debugging component iteratively refine the output of a code generation component, while a genetic algorithm extracts performance feedback from the execution environment to guide subsequent code refinements. We provide a detailed explanation of the DiffBench construction and the design principles underlying DiffAgent. Extensive experiments show that DiffBench offers a thorough evaluation of generated codes and that DiffAgent significantly outperforms existing LLMs in producing effective diffusion acceleration strategies.
Jiajun Jiao, Haowei Zhu, Puyuan Yang, Jianghui Wang, Ziqiong Liu, Dong Li 0025, Yuejian Fang, Jun-Hai Yong, Bin Wang 0034, Emad Barsoum
AAAI8
2026 FedSRD: Sparsify-Reconstruct-Decompose for Communication-Efficient Federated Large Language Models Fine-Tuning
abstract
The current paradigm of training large language models (LLMs) on public available Web data is becoming unsustainable as high-quality data sources in specialized domains near exhaustion. Federated Learning (FL) emerges as a practical solution for the next generation of AI on a decentralized Web, enabling privacy-preserving collaborative fine-tuning on decentralized private data. While Low-Rank Adaptation (LoRA) is standard for efficient fine-tuning, its federated application faces a critical bottleneck: communication overhead under heterogeneous network conditions. Structural redundancy in LoRA parameters increases communication costs and causes aggregation conflicts. To address this, we propose FedSRD, a Sparsify-Reconstruct-Decompose framework for communication-efficient federated LLM fine-tuning. We introduce importance-aware sparsification to reduce the upload parameter count while preserving the structural integrity of LoRA updates. The server aggregates updates in full-rank space to mitigate conflicts, then decomposes the global update into a sparse low-rank format for broadcast, ensuring a symmetrically efficient cycle. We also propose an efficient variant, FedSRD-e, to reduce computational overhead. Experiments on 10 benchmarks show our framework significantly reduces communication costs by up to 90% while improving performance on heterogeneous client data.
Guochen Yan, Luyuan Xie, Qingni Shen, Yuejian Fang, Zhonghai Wu
WWW4
2025 FedVCK: Non-IID Robust and Communication-Efficient Federated Learning via Valuable Condensed Knowledge for Medical Image Analysis
abstract
Federated learning has become a promising solution for collaboration among medical institutions. However, data owned by each institution would be highly heterogeneous and the distribution is always non-independent and identical distribution (non-IID), resulting in client drift and unsatisfactory performance. Despite existing federated learning methods attempting to solve the non-IID problems, they still show marginal advantages but rely on frequent communication which would incur high costs and privacy concerns. In this paper, we propose a novel federated learning method: Federated learning via Valuable Condensed Knowledge (FedVCK). We enhance the quality of condensed knowledge and select the most necessary knowledge guided by models, to tackle the non-IID problem within limited communication budgets effectively. Specifically, on the client side, we condense the knowledge of each client into a small dataset and further enhance the condensation procedure with latent distribution constraints, facilitating the effective capture of high-quality knowledge. During each round, we specifically target and condense knowledge that has not been assimilated by the current model, thereby preventing unnecessary repetition of homogeneous knowledge and minimizing the frequency of communications required. On the server side, we propose relational supervised contrastive learning to provide more supervision signals to aid the global model updating. Comprehensive experiments across various medical tasks show that FedVCK can outperform state-of-the-art methods, demonstrating that it's non-IID robust and communication-efficient.
Guochen Yan, Luyuan Xie, Xinyi Gao 0001, Wentao Zhang 0001, Qingni Shen, Yuejian Fang, Zhonghai Wu
AAAI6
2025 dFLMoE: Decentralized Federated Learning via Mixture of Experts for Medical Data Analysis
abstract
Federated learning has wide applications in the medical field. It enables knowledge sharing among different healthcare institutes while protecting patients’ privacy. However, existing federated learning systems are typically centralized, requiring clients to upload client-specific knowledge to a central server for aggregation. This centralized approach would integrate the knowledge from each client into a centralized server, and the knowledge would be already undermined during the centralized integration before it reaches back to each client. Besides, the centralized approach also creates a dependency on the central server, which may affect training stability if the server malfunctions or connections are unstable. To address these issues, we propose a decentralized federated learning framework named dFLMoE. In our framework, clients directly exchange lightweight head models with each other. After exchanging, each client treats both local and received head models as individual experts, and utilizes a client-specific Mixture of Experts (MoE) approach to make collective decisions. This design not only reduces the knowledge damage with client-specific aggregations but also removes the dependency on the central server to enhance the robustness of the framework. We validate our framework on multiple medical tasks, demonstrating that our method evidently outperforms state-of-the-art approaches under both model homogeneity and heterogeneity settings.
Luyuan Xie, Tianyu Luan, Wenyuan Cai, Guochen Yan, Nan Xi, Yuejian Fang, Qingni Shen, Zhonghai Wu, Junsong Yuan 0001
CVPR7
2025 ZK-Hammer: Leaking Secrets from Zero-Knowledge Proofs via Rowhammer
abstract
Zero-knowledge succinct non-interactive arguments of knowledge (zk-SNARK) schemes have been a promising technique in verified computation. Zk-SNARK schemes were designed to be mathematically secure against cryptographic attacks and it remains unclear whether they are vulnerable to fault injection attacks. In this work, we provide a positive answer by presenting ZK-Hammer, which leaks secrets from zk-SNARK schemes via Rowhammer. We incur faults in the exponentiate variables in the Quadratic Arithmetic Program (QAP) problem. Then we analyze the faulty proof using the bilinear pairing technique and manage to recover the secret. We employ a Rowhammer fault evaluation in libsnark and identify 3 CVEs.
Junkai Liang, Xin Zhang 0110, Daqi Hu, Qingni Shen, Yuejian Fang, Zhonghai Wu
DAC5
2025 A lattice-based privacy-preserving decentralized multi-party payment scheme
Jisheng Dong, Qingni Shen, Junkai Liang, Cong Li 0024, Xinyu Feng 0002, Yuejian Fang
Comput. Networks6
2024 Privacy Preserving Federated Learning from Multi-Input Functional Proxy Re-Encryption
abstract
Federated learning (FL) allows different participants to collaborate on model training without transmitting raw data, thereby protecting user data privacy. However, FL faces a series of security and privacy issues (e.g. the leakage of raw data from publicly shared parameters). Several privacy protection technologies, such as homomorphic encryption, differential privacy and functional encryption, are introduced for privacy enhancement in FL. Among them, the FL frameworks based on functional encryption better balance security and performance, thus receiving increasing attention. The previous FL frameworks based on functional encryption suffer from several security issues, including attacks by combining multiple rounds of ciphertexts and keys, and leakage of global parameters to the central server. To tackle these issues, we propose a novel multi-input functional proxy re-encryption (MI-FPRE) scheme and further design a new FL framework with better privacy based on MI-FPRE. Our framework allows a semi-trusted central server to aggregate the parameters without knowing the intermediate parameters and the result of aggregation, thus achieves better privacy in FL training. The experimental results indicate that our framework achieves less communication overhead and higher computational efficiency without losing accuracy.
Xinyu Feng 0002, Qingni Shen, Cong Li 0024, Yuejian Fang, Zhonghai Wu
ICASSP4
2024 Learning Invariant Representation with Consistency and Diversity for Semi-Supervised Source Hypothesis Transfer
abstract
Semi-supervised Domain adaptation (SSDA) has shown promising results by leveraging unlabeled data and limited labeled samples in the target domain. However, accessibility to source data is hindered by data privacy concerns, giving rise to Semi-supervised Source Hypothesis Transfer (SSHT). Integrating the SSDA methods directly into SSHT tasks is straightforward but poses two significant challenges: i) The hypothesis (classifier) is no longer supervised by source labels, and relying on only a few labels may result in hypothesis collapse; ii) Trained source models often exhibit bias, making them susceptible to misclassifying samples from minority categories into majority ones. We examined the recent methods in the SSHT setting and observed variations in performance compared to SSDA. To address these challenges, we first mitigate model overfitting to target labeled data by promoting prediction consistency between two types of randomly augmented unlabeled data, thereby preventing training collapse. Additionally, we maintain both the prediction diversity and discriminability by leveraging unlabeled data. Experiments on SSHT tasks show that our method yields more stable and competitive results compared with state-of-the-art methods.
Junbao Zhuo, Shuhao Cui, Shuhui Wang, Yuejian Fang
ICASSP5
2024 TRLS: A Time Series Representation Learning Framework Via Spectrogram for Medical Signal Processing
abstract
Representation learning frameworks in unlabeled time series have been proposed for medical signal processing. Despite the numerous excellent progresses have been made in previous works, we observe the representation extracted for the time series still does not generalize well. In this paper, we present a Time series (medical signal) Representation Learning framework via Spectrogram (TRLS) to get more informative representations. We transform the input time-domain medical signals into spectrograms and design a time-frequency encoder named Time Frequency RNN (TFRNN) to capture more robust multi-scale representations from the augmented spectrograms. Our TRLS takes spectrogram as input with two types of different data augmentations and maximizes the similarity between positive ones, which effectively circumvents the problem of designing negative samples. Our evaluation of four real-world medical signal datasets focusing on medical signal classification shows that TRLS is superior to the existing frameworks. We will open-source our code when the paper is accepted.
Luyuan Xie, Cong Li 0024, Xin Zhang 0110, Shengfang Zhai, Yuejian Fang, Qingni Shen, Zhonghai Wu
ICASSP5
2024 HyPRE: Hybrid Proxy Re-Encryption for Secure Multimedia Data Sharing on Mobile Devices
abstract
Due to the rapid growth of mobile internet, massive multimedia data (e.g., movies, photos, notes, etc.) on mobile devices is synchronized and shared through the cloud. During this process, public key encryption plays an important role in ensuring the confidentiality of data. However, due to the bottleneck of computing and storage resources in mobile devices, it is difficult to execute complex cryptographic algorithms on them. In this paper, we present a novel Hybrid Proxy Reencryption (HyPRE) scheme for the sharing of multimedia data on mobile devices, which empowers a semi-trusted proxy to convert a ciphertext under an identity to a new one under an expressive policy without revealing the underlying plaintext. Our scheme allows mobile devices with limited resources to encrypt data efficiently, and then to share the encrypted data to multiple entities securely. We define the HRA security for our HyPRE scheme to improve the incompleteness of the security under chosen plaintext attacks (CPA) in traditional proxy re-encryption schemes and prove it selectively secure under HRA. Experimental analysis indicates that HyPRE achieves 2× to 3× improvement in terms of re-encryption performance compared with the state-of-the-art ones.
Xinyu Feng 0002, Cong Li 0024, Qingni Shen, Jisheng Dong, Wenjun Qian, Yuejian Fang, Zhonghai Wu
ICME6
2024 Enhancing Zero-shot 3D Photography via Mesh-represented Image Inpainting
abstract
3D photography techniques create a consistent 3D video given a single image. Existing methods use multi-plane images or layered depth images to represent 3D scenes and then render subsequent novel views. However, these methods involve warping pixels across frames, which easily causes distortions, harms visual coherence, and lacks controllable generation with textual prompts. Moreover, these methods require training models on adequately large datasets beforehand, whether they are tailored to a specific domain or open domain, which requires high computational resources. To address these issues, we propose an enhanced zero-shot 3D photography method, termed Zero-3DP, to enable rendering any image into a 3D video anytime. We first integrate meshes to represent 3D scenes in our pipeline and update meshes along predefined trajectories, ensuring geometry consistency via depth alignment and prior preservation in rendering. To maintain semantic consistency, we test-time fine-tune the diffusion-based inpainting module for each incoming frame. Experiments on two public benchmarks show that without previous training, just relying on test-time fine-tuning in inference, Zero-3DP can match or beat the state-of-the-art methods.
Yuejian Fang
ICME1
2024 MH-pFLID: Model Heterogeneous personalized Federated Learning via Injection and Distillation for Medical Data Analysis
abstract
Federated learning is widely used in medical applications for training global models without needing local data access, but varying computational capabilities and network architectures (system heterogeneity) across clients pose significant challenges in effectively aggregating information from non-independently and identically distributed (non-IID) data (statistic heterogeneity). Current federated learning methods using knowledge distillation require public datasets, raising privacy and data collection issues. Additionally, these datasets require additional local computing and storage resources, which is a burden for medical institutions with limited hardware conditions. In this paper, we introduce a novel federated learning paradigm, named Model Heterogeneous personalized Federated Learning via Injection and Distillation (MH-pFLID). Our framework leverages a lightweight messenger model, eliminating the need for public datasets and reducing the training cost for each client. We also develops receiver and transmitter modules for each client to separate local biases from generalizable information, reducing biased data collection and mitigating client drift. Our experiments on various medical tasks including image classification, image segmentation, and time-series classification, show MH-pFLID outperforms state-of-the-art methods in all these areas and has good generalizability.
Luyuan Xie, Manqing Lin, Tianyu Luan, Cong Li 0024, Yuejian Fang, Qingni Shen, Zhonghai Wu
ICML5
2024 pFLFE: Cross-silo Personalized Federated Learning via Feature Enhancement on Medical Image Segmentation
Luyuan Xie, Manqing Lin, ChenMing Xu, Tianyu Luan, Cong Li 0024, Yuejian Fang, Qingni Shen, Zhonghai Wu
MICCAI (10)7
2024 MH-pFLGB: Model Heterogeneous Personalized Federated Learning via Global Bypass for Medical Image Analysis
Luyuan Xie, Manqing Lin, ChenMing Xu, Tianyu Luan, Zhipeng Zeng, Wenjun Qian, Cong Li 0024, Yuejian Fang, Qingni Shen, Zhonghai Wu
MICCAI (10)8
2024 FDP-FL: differentially private federated learning with flexible privacy budget allocation
abstract
Abstract Federated learning (FL) as a privacy-preserving technology enables multiple clients to collaboratively train models on decentralized data. However, transmitting model parameters between local clients and the central server can potentially result in information leakage. Differentially private federated learning (DPFL) has emerged as a promising solution to enhance privacy. Nevertheless, existing DPFL schemes suffer from two issues: (i) most schemes that aim to achieve desired model accuracy may incur a high privacy budget. (ii) several schemes that consider the trade-off between privacy and accuracy by utilizing linear clipping bound may distort numerous model parameters. In this paper, we first propose FDP-FL, a flexible differential privacy approach for FL. FDP-FL introduces a novel series sum privacy budget allocation instead of uniform allocation and enables adaptive and nonlinear noise scale decay. In this way, a tight bound for cumulative privacy loss can be achieved while optimizing model accuracy. Then in order to mitigate gradient leakages caused by honest-but-curious clients and server, we further design client-level FDP-FL and record-level FDP-FL, respectively. Experimental results demonstrate that our FDP-FL improves model accuracy by $\sim $13.3% compared with the basic DP-FL under a fixed privacy budget and outperforms existing trade-off schemes with the same hyperparameter setting.
Wenjun Qian, Qingni Shen, Cong Li 0024, Yuejian Fang, Zhonghai Wu
Comput. J.5
2023 A Privacy Preserving Computer-aided Medical Diagnosis Framework with Outsourced Model
abstract
Computer-aided diagnosis plays an increasingly important role in modern medical activities, relying largely on the deployment of medical machine learning models. Protecting the security of model parameters is crucial for model providers. However, the current schemes for protecting model parameters are mostly interactive. This interactive nature makes it difficult to support offline deployment of models and flexible authorization of prediction results, thus hindering the widespread application of computer-aided diagnosis. To address these limitations, we propose a new computer-aided medical diagnosis framework by designing a new identity-based inner product functional proxy re-encryption (IB-IPFPRE) scheme. Our framework supports private deployment of medical diagnostic models without compromising model parameters. It also enables access control of prediction results based on user identity. Compared to existing privacy-preserving prediction techniques, our framework significantly reduces communication overhead and does not require the model owner to be online in real-time. Furthermore, our scheme enables flexible delegation of prediction results, allowing users to authorize the sharing of prediction results with other entities as needed. We conducted extensive experiments for logistic regression on three medical datasets. The experiments demonstrate that our scheme achieved 40% to 7× performance improvement in LAN environment and 13× to 15× improvement in WAN environment, and did not require any communication overhead during the privacy preserving prediction phase.
Xinyu Feng 0002, Qingni Shen, Cong Li 0024, Niantao Xie, Luyuan Xie, Yuejian Fang, Zhonghai Wu
BIBM7
2023 NCL: Textual Backdoor Defense Using Noise-Augmented Contrastive Learning
abstract
At present, backdoor attacks attract attention as they do great harm to deep learning models. By poisoning the training data, the adversary makes the model trained based on this dataset being injected with a backdoor. In the field of text, however, existing works do not provide sufficient defense against backdoor attacks. In this paper, we propose a Noise-augmented Contrastive Learning (NCL) framework to defend against textual backdoor attacks when training models with untrustworthy data. With the aim of mitigating the mapping between triggers and the target label, we add appropriate noise perturbing possible backdoor triggers, augment the training dataset, and then pull homology samples in the feature space utilizing contrastive learning objective. Experiments demonstrate the effectiveness of our method in defending three types of textual backdoor attacks, outperforming the prior works.
Shengfang Zhai, Qingni Shen, Cong Li 0024, Yuejian Fang, Zhonghai Wu
ICASSP6
2023 Learning 3D Photography Videos via Self-supervised Diffusion on Single Images
abstract
3D photography renders a static image into a video with appealing 3D visual effects. Existing approaches typically first conduct monocular depth estimation, then render the input frame to subsequent frames with various viewpoints, and finally use an inpainting model to fill those missing/occluded regions. The inpainting model plays a crucial role in rendering quality, but it is normally trained on out-of-domain data. To reduce the training and inference gap, we propose a novel self-supervised diffusion model as the inpainting module. Given a single input image, we automatically construct a training pair of the masked occluded image and the ground-truth image with random cycle rendering. The constructed training samples are closely aligned to the testing instances, without the need for data annotation. To make full use of the masked images, we designed a Masked Enhanced Block (MEB), which can be easily plugged into the UNet and enhance the semantic conditions. Towards real-world animation, we present a novel task: out-animation, which extends the space and time of input objects. Extensive experiments on real datasets show that our method achieves competitive results with existing SOTA methods.
Xiaodong Wang 0023, Chenfei Wu, Shengming Yin, Minheng Ni, Zhengyuan Yang, Fan Yang 0024, Zicheng Liu 0001, Yuejian Fang, Nan Duan 0001
IJCAI11
2023 Text-to-Image Diffusion Models can be Easily Backdoored through Multimodal Data Poisoning
abstract
With the help of conditioning mechanisms, the state-of-the-art diffusion models have achieved tremendous success in guided image generation, particularly in text-to-image synthesis. To gain a better understanding of the training process and potential risks of text-to-image synthesis, we perform a systematic investigation of backdoor attack on text-to-image diffusion models and propose BadT2I, a general multimodal backdoor attack framework that tampers with image synthesis in diverse semantic levels. Specifically, we perform backdoor attacks on three levels of the vision semantics: Pixel-Backdoor, Object-Backdoor and Style-Backdoor. By utilizing a regularization loss, our methods efficiently inject backdoors into a large-scale text-to-image diffusion model while preserving its utility with benign inputs. We conduct empirical experiments on Stable Diffusion, the widely-used text-to-image diffusion model, demonstrating that the large-scale diffusion model can be easily backdoored within a few fine-tuning steps. We conduct additional experiments to explore the impact of different types of textual triggers, as well as the backdoor persistence during further training, providing insights for the development of backdoor defense methods. Besides, our investigation may contribute to the copyright protection of text-to-image models in the future. Our Code: https://github.com/sf-zhai/BadT2I.
Shengfang Zhai, Yinpeng Dong, Qingni Shen, Shi Pu 0002, Yuejian Fang, Hang Su 0006
ACM Multimedia5
2022 Revisiting Unsupervised Domain Adaptation Models: A Smoothness Perspective
Junbao Zhuo, Shuhui Wang, Yuejian Fang
ACCV (6)5
2022 A Simple and Effective Method to Improve Zero-Shot Cross-Lingual Transfer Learning
abstract
Existing zero-shot cross-lingual transfer methods rely on parallel corpora or bilingual dictionaries, which are expensive and impractical for low-resource languages. To disengage from these dependencies, researchers have explored training multilingual models on English-only resources and transferring them to low-resource languages. However, its effect is limited by the gap between embedding clusters of different languages. To address this issue, we propose Embedding-Push, Attention-Pull, and Robust targets to transfer English embeddings to virtual multilingual embeddings without semantic loss, thereby improving cross-lingual transferability. Experimental results on mBERT and XLM-R demonstrate that our method significantly outperforms previous works on the zero-shot cross-lingual text classification task and can obtain a better multilingual alignment.
Kunbo Ding, Weijie Liu 0002, Yuejian Fang, Weiquan Mao, Zhe Zhao 0006, Haoyan Liu 0001, Rong Tian
COLING3
2022 NÜWA: Visual Synthesis Pre-training for Neural visUal World creAtion
Chenfei Wu, Lei Ji 0001, Fan Yang 0024, Yuejian Fang, Daxin Jiang, Nan Duan 0001
ECCV (16)5
2022 Efficient Identity-Based Chameleon Hash for Mobile Devices
abstract
Online/offline identity-based signature (OO-IBS) is an adequate cryptographic tool to provide the message authentication and integrity in mobile devices, since it lightens the computational burden after the signer receives the message and eliminates the overhead of certificate management. It has several valuable applications, such as wireless sensor networks and automatic dependent surveillance-broadcast systems. Identity-based chameleon hash (IB-CH), as an alternative building block to construct OO-IBS, has been explored in several literatures. Nevertheless, almost all of the prior IB-CH schemes are in the random oracle model, which may lead to security risks in practicality. The only IB-CH scheme in the standard model proposed by Xie et al. (ICC’21) suffers from the large size of public parameters and inefficient setup process. In this paper, we propose an efficient IB-CH scheme in the standard model, significantly reducing the computational costs of all the algorithms and the size of public parameters compared with Xie’s scheme. The security and experimental analyses demonstrate the security and good performance of our scheme. Furthermore, we applied our scheme to optimize the existing generic OO-IBS construction. Our optimized construction reduces computational overhead by 50.0% in the online phase compared with the original construction.
Cong Li 0024, Qingni Shen, Zhikang Xie, Jisheng Dong, Yuejian Fang, Zhonghai Wu
ICASSP5
2022 NUWA-Infinity: Autoregressive over Autoregressive Generation for Infinite Visual Synthesis
abstract
Infinite visual synthesis aims to generate high-resolution images, long-duration videos, and even visual generation of infinite size. Some recent work tried to solve this task by first dividing data into processable patches and then training the models on them without considering the dependencies between patches. However, since they fail to model global dependencies between patches, the quality and consistency of the generation can be limited. To address this issue, we propose NUWA-Infinity, a patch-level \emph{``render-and-optimize''} strategy for infinite visual synthesis. Given a large image or a long video, NUWA-Infinity first splits it into non-overlapping patches and uses the ordered patch chain as a complete training instance, a rendering model autoregressively predicts each patch based on its contexts. Once a patch is predicted, it is optimized immediately and its hidden states are saved as contexts for the next \emph{``render-and-optimize''} process. This brings two advantages: ($i$) The autoregressive rendering process with information transfer between contexts provides an implicit global probabilistic distribution modeling; ($ii$) The timely optimization process alleviates the optimization stress of the model and helps convergence. Based on the above designs, NUWA-Infinity shows a strong synthesis ability on high-resolution images and long-duration videos. The homepage link is \url{https://nuwa-infinity.microsoft.com}.
Chenfei Wu, Xiaowei Hu 0006, Zhe Gan, Zicheng Liu 0001, Yuejian Fang, Nan Duan 0001
NeurIPS8
2022 Hierarchical and non-monotonic key-policy attribute-based encryption and its application
Cong Li 0024, Qingni Shen, Zhikang Xie, Jisheng Dong, Xinyu Feng 0002, Yuejian Fang, Zhonghai Wu
Inf. Sci.6
2021 Identity-Based Chameleon Hash without Random Oracles and Application in the Mobile Internet
abstract
The rapid development of the mobile Internet makes it necessary to adopt efficient cryptographic primitives for the portable devices with limited computing resources. Online/offline identity-based signatures are suitable because of short response time of signature generation and being free from the cumbersome operations caused by public key infrastructures. In this paper, we propose the first identity-based chameleon hash which can be proved secure without the random oracle and show how to use it to translate any identity-based signature to an online/offline one.
Zhikang Xie, Qingni Shen, Cong Li 0024, Jisheng Dong, Yuejian Fang
ICC5
2021 Hybrid Reasoning Network for Video-based Commonsense Captioning
abstract
The task of video-based commonsense captioning aims to generate event-wise captions and meanwhile provide multiple commonsense descriptions (e.g., attribute, effect and intention) about the underlying event in the video. Prior works explore the commonsense captions by using separate networks for different commonsense types, which is time-consuming and lacks mining the interaction of different commonsense. In this paper, we propose a Hybrid Reasoning Network (HybridNet) to endow the neural networks with the capability of semantic-level reasoning and word-level reasoning. Firstly, we develop multi-commonsense learning for semantic-level reasoning by jointly training different commonsense types in a unified network, which encourages the interaction between the clues of multiple commonsense descriptions, event-wise captions and videos. Then, there are two steps to achieve the word-level reasoning: (1) a memory module records the history predicted sequence from the previous generation processes; (2) a memory-routed multi-head attention (MMHA) module updates the word-level attention maps by incorporating the history information from the memory module into the transformer decoder for word-level reasoning. Moreover, the multimodal features are used to make full use of diverse knowledge for commonsense reasoning. Experiments and abundant analysis on the large-scale Video-to-Commonsense benchmark show that our HybridNet achieves state-of-the-art performance compared with other methods.
Weijiang Yu, Lei Ji 0001, Yuejian Fang, Nan Duan 0001
ACM Multimedia5
2021 Large Universe CCA2 CP-ABE With Equality and Validity Test in the Standard Model
abstract
Abstract Attribute-based encryption with equality test (ABEET) simultaneously supports fine-grained access control on the encrypted data and plaintext message equality comparison without decrypting the ciphertexts. Recently, there have been several literatures about ABEET proposed. Nevertheless, most of them explore the ABEET schemes in the random oracle model, which has been pointed out to have many defects in practicality. The only existing ABEET scheme in the standard model, proposed by Wang et al., merely achieves the indistinguishable against chosen-plaintext attack security. Considering the aforementioned problems, in this paper, we propose the first direct adaptive chosen-ciphertext security ciphertext-policy ABEET scheme in the standard model. Our method only adopts a chameleon hash function and adds one dummy attribute to the access structure. Compared with the previous works, our scheme achieves the security improvement, ciphertext validity check and large universe. Besides, we further optimize our scheme to support the outsourced decryption. Finally, we first give the detailed theoretical analysis of our constructions in computation and storage costs, then we implement our constructions and carry out a series of experiments. Both results indicate that our constructions are more efficient in Setup and Trapdoor and have the shorter public parameters than the existing ABEET ones do.
Cong Li 0024, Qingni Shen, Zhikang Xie, Xinyu Feng 0002, Yuejian Fang, Zhonghai Wu
Comput. J.5
2020 Unicoder-VL: A Universal Encoder for Vision and Language by Cross-Modal Pre-Training
abstract
We propose Unicoder-VL, a universal encoder that aims to learn joint representations of vision and language in a pre-training manner. Borrow ideas from cross-lingual pre-trained models, such as XLM (Lample and Conneau 2019) and Unicoder (Huang et al. 2019), both visual and linguistic contents are fed into a multi-layer Transformer (Vaswani et al. 2017) for the cross-modal pre-training, where three pre-trained tasks are employed, including Masked Language Modeling(MLM), Masked Object Classification(MOC) and Visual-linguistic Matching(VLM). The first two tasks learn context-aware representations for input tokens based on linguistic and visual contents jointly. The last task tries to predict whether an image and a text describe each other. After pretraining on large-scale image-caption pairs, we transfer Unicoder-VL to caption-based image-text retrieval and visual commonsense reasoning, with just one additional output layer. We achieve state-of-the-art or comparable results on both two tasks and show the powerful ability of the cross-modal pre-training.
Nan Duan 0001, Yuejian Fang, Ming Gong 0001, Daxin Jiang
AAAI3
2017 Practical Large Universe Attribute-Set Based Encryption in the Standard Model
Xinyu Feng 0002, Cancan Jin, Cong Li 0024, Yuejian Fang, Qingni Shen, Zhonghai Wu
ICICS4
2017 Fully Secure Hidden Ciphertext-Policy Attribute-Based Proxy Re-encryption
Xinyu Feng 0002, Cong Li 0024, Yuejian Fang, Qingni Shen
ICICS4
2017 A practical construction for large universe hierarchical attribute-based encryption
abstract
Summary We present a practical large universe hierarchical attribute‐based encryption (LU‐HABE) scheme, which supports monotone access structures. In our system, key generation centers (KGCs), any one in which is labeled by a unique identity, are organized as a hierarchical structure. Thus, all secret keys issued by the KGC contain 2 parts: the identity‐related one and the attribute‐related one. Once the data owner wants to encrypt his/her data, he/she needs to specify certain numbers of pairs according to his/her demand. The pair consists of an identity of a KGC and a policy of attributes managed by the corresponding KGC, eg, IDi and (Mi, ρi). If and only if an identity associated with user's secret key is equal to or is an ancestor of one of the identities appearing in ciphertext, and simultaneously a set of attributes belonging to the user satisfies the policy, the user can decrypt it successfully. Our scheme is proved to be selectively secure in the standard model under the modified “q‐type” assumption similar to the ones used in former works and is extended to support online/offline encryption. To show the efficiency of our construction, we implement our original scheme and the extended one in Charm. Analyses show that both of them are very practical.
Cong Li 0024, Yuejian Fang, Xing Zhang 0002, Cancan Jin, Qingni Shen, Zhonghai Wu
Concurr. Comput. Pract. Exp.2
2015 POSTER: Ciphertext-Policy Attribute-Based Encryption Method with Secure Decryption Key Generation and Outsourcing Decryption of ABE Ciphertexts
Yuejian Fang, Zilong Wen, Qingni Shen, Yahui Yang, Zhonghai Wu
SecureComm1
2015 Ciphertext-Policy Attribute-Based Encryption with User and Authority Accountability
Xing Zhang 0002, Cancan Jin, Cong Li 0024, Zilong Wen, Qingni Shen, Yuejian Fang, Zhonghai Wu
SecureComm6