EDBT 2026 Demo / reviewers in the wild / expert
Wenju Xu
dblp:187/7294
· DBLP profile ↗
29ranked-venue papers
15as first author
23since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 6 first-author · 12 since 2021Artificial intelligence and machine learning · 9 · 5 first-author · 6 since 2021Computer networks · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | EcomScriptBench: A Multi-task Benchmark for E-commerce Script Planning via Step-wise Intention-Driven Product AssociationabstractWeiqi Wang, Limeng Cui, Xin Liu, Sreyashi Nag, Wenju Xu, Chen Luo, Sheikh Muhammad Sarwar, Yang Li, Hansu Gu, Hui Liu, Changlong Yu, Jiaxin Bai, Yifan Gao, Haiyang Zhang, Qi He, Shuiwang Ji, Yangqiu Song. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Weiqi Wang 0001, Limeng Cui, Xin Liu 0039, Sreyashi Nag, Wenju Xu, Chen Luo 0003, Sheikh Muhammad Sarwar, Yang Li 0055, Hansu Gu, Hui Liu 0033, Changlong Yu, Jiaxin Bai, Yifan Gao 0001, Qi He 0002, Shuiwang Ji, Yangqiu Song |
ACL (1) | 5 |
| 2025 | Language Model Alignment for Conversational Shopping at AmazonabstractThe rapid growth of online shopping stores, such as Amazon, has led to services reaching billions of people worldwide. With global retail sales exceeding $6 trillion in 2024, customer expectations for personalized and seamless shopping experiences have heightened. Traditional online shopping experiences, such as search and navigation systems, often fall short in addressing complex shopping journeys. Conversational shopping (such as Amazon Rufus) offers a transformative approach by enabling dynamic, multi-turn dialogues that closely resemble human interactions. This allows customers to explore product options, seek clarifications, and receive personalized recommendations, thereby enhancing product discovery and informed decision-making. In this paper, we share our year-long journey of using language models for conversational shopping at Amazon and introduce how we use LLM fine-tuning techniques to enhance LLMs for a conversational shopping experience like Amazon Rufus. We also introduce innovative strategies for training data collection and demonstrate real-world applications, including product recommendations, clarification mechanisms, and internationalization for global customers. Chen Luo 0003, Dimitri Papadimitriou, Hariharan Muralidharan, Dhineshkumar Ramasubbu, Aakash Kolekar, Wenju Xu, Anirudh Srinivasan, Mukesh Jain, Qi He 0002 |
SIGIR | 6 |
| 2025 | Depth-Wise Convolutions in Vision Transformers for efficient training on small datasetsabstractThe Vision Transformer (ViT) leverages the Transformer’s encoder to capture global information by dividing images into patches and achieves superior performance across various computer vision tasks. However, the self-attention mechanism of ViT captures the global context from the outset, overlooking the inherent relationships between neighboring pixels in images or videos. Transformers mainly focus on global information while ignoring the fine-grained local details. Consequently, ViT lacks inductive bias during image or video dataset training. In contrast, convolutional neural networks (CNNs), with their reliance on local filters, possess an inherent inductive bias, making them more efficient and quicker to converge than ViT with less data. In this paper, we present a lightweight Depth-Wise Convolution module as a shortcut in ViT models, bypassing entire Transformer blocks to ensure the models capture both local and global information with minimal overhead. Additionally, we introduce two architecture variants, allowing the Depth-Wise Convolution modules to be applied to multiple Transformer blocks for parameter savings, and incorporating independent parallel Depth-Wise Convolution modules with different kernels to enhance the acquisition of local information. The proposed approach significantly boosts the performance of ViT models on image classification, object detection, and instance segmentation by a large margin, especially on small datasets, as evaluated on CIFAR-10, CIFAR-100, Tiny-ImageNet and ImageNet for image classification, and COCO for object detection and instance segmentation. The source code can be accessed at https://github.com/ZTX-100/Efficient_ViT_with_DW . Tianxiao Zhang, Wenju Xu, Bo Luo, Guanghui Wang 0001 |
Neurocomputing | 2 |
| 2025 | Facial Highlight Removal With Cross-Context Attention and Texture EnhancementabstractFacial highlight removal aims to identify and remove the specular highlight components in the facial image, ensuring that the generated image has a consistent facial tone and high-fidelity texture detail. Existing methods struggle to remove the highlight and recover the details in disturbed areas simultaneously, often resulting in specular residues or distorted local details (i.e. texture, illumination, and color). To rectify these issues, this work proposes a novel two-stage facial highlight removal network (FHR-Net), which mainly consists of a Cross-Context Attention Module (CCAM) and a Texture Enhancement Module (TEM). In the first stage, according to the detected highlight mask, the CCAM explicitly integrates cross-context information to obtain coarse highlight removal results consistent with the surrounding facial context. Building upon the coarse result, the TEM in the second stage utilizes patch-wise attention to refine the texture details in the highlight areas, thereby producing a high-fidelity facial image. To improve coherence between the removed highlight areas and non-highlight areas, this work introduces a face feature loss that makes the processed highlight-disturbed areas align well with the surrounding facial architecture. Additionally, to address the lack of high-quality datasets in the research community and satisfy the training demands for data-driven facial highlight removal, this work builds a real-world Paired Facial Specular-Diffuse (PFSD) dataset through cross-polarization. Experimental results on PFSD and other datasets demonstrate that FHR-Net can effectively remove the facial highlight and recover original color and texture details. Hongsheng Zheng, Wenju Xu, Xiao Lu 0002, Chunxia Xiao |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | STGlight: Online Indoor Lighting Estimation via Spatio-Temporal Gaussian FusionabstractEstimating lighting in indoor scenes is particularly challenging due to diverse distribution of light sources and complexity of scene geometry. Previous methods mainly focused on spatial variability and consistency for a single image or temporal consistency for video sequences. However, these approaches fail to achieve spatio-temporal consistency in video lighting estimation, which restricts applications such as compositing animated models into videos. In this paper, we propose STGlight, a lightweight and effective method for spatio-temporally consistent video lighting estimation, where our network processes a stream of LDR RGB-D video frames while maintaining incrementally updated global representations of both geometry and lighting, enabling the prediction of HDR environment maps at arbitrary locations for each frame. We model indoor lighting with three components: visible light sources providing direct illumination, ambient lighting approximating indirect illumination, and local environment textures producing high-quality specular reflections on glossy objects. To capture spatial-varying lighting, we represent scene geometry with point clouds, which support efficient spatio-temporal fusion and allow us to handle moderately dynamic scenes. To ensure temporal consistency, we apply a transformer-based fusion block that propagates lighting features across frames. Building on this, we further handle dynamic lighting with moving objects or changing light conditions by applying intrinsic decomposition on the point cloud and integrating the decomposed components with a neural fusion module. Experiments show that our online method can effectively predict lighting for any position within the video stream, while maintaining spatial variability and spatio-temporal consistency. Code is available at: https://github.com/nauyihsnehs/STGlight. Shiyuan Shen, Zhongyun Bao, Wenju Xu, Tenghui Lai, Chunxia Xiao |
ACM Trans. Graph. | 4 |
| 2025 | IllumiDiff: Indoor Illumination Estimation From a Single Image With Diffusion ModelabstractIllumination estimation from a single indoor image is a promising yet challenging task. Existing indoor illumination estimation methods mainly regress lighting parameters or infer a panorama from a limited field-of-view image. Nevertheless, these methods fail to recover a panorama with both well-distributed illumination and detailed environment textures, leading to a lack of realism in rendering the embedded 3D objects with complex materials. This paper presents a novel multi-stage illumination estimation framework named IllumiDiff. Specifically, in Stage I, we first estimate illumination conditions from the input image, including the illumination distribution as well as the environmental texture of the scene. In Stage II, guided by the estimated illumination conditions, we design a conditional panoramic texture diffusion model to generate a high-quality LDR panorama. In Stage III, we leverage the illumination conditions to further reconstruct the LDR panorama to an HDR panorama. Extensive experiments demonstrate that our IllumiDiff can generate an HDR panorama with realistic illumination distribution and rich texture details from a single limited field-of-view indoor image. The generated panorama can produce impressive rendering results for the embedded 3D objects with various materials. Shiyuan Shen, Zhongyun Bao, Wenju Xu, Chunxia Xiao |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2024 | HS-Surf: A Novel High-Frequency Surface Shell Radiance Field to Improve Large-Scale Scene RenderingabstractPrevious neural radiance fields often struggle to preserve high-frequency textures in urban and aerial large-scale scenes due to insufficient model capacity on the scene surface. This is attributed to their sampling locations or grid vertices falling in empty areas. Additionally, most models do not consider the drastic changes in distances. To address these issues, we propose a novel high-frequency surface shell radiance field, which uses depth-guided information to create a shell enveloping the scene surface under the current view, and then samples conic frustums on this shell to render high-frequency textures. Specifically, our method comprises three parts. Initially, we propose a strategy to fuse voxel grids and information of distance scales to generate a coarse scene at different distance scales. Subsequently, we construct a shell based on the depth information to carry out compensation to incorporate texture details not captured by voxels. Finally, the smooth and denoise post-processing further improves the rendering quality. Substantial scene experiments and ablation experiments demonstrate that our method achieves the obvious improvement of high-frequency textures at different distance scales and outperforms the state-of-the-art methods. Jiongming Qin, Fei Luo 0004, Tuo Cao, Wenju Xu, Chunxia Xiao |
ACM Multimedia | 4 |
| 2024 | HighlightRemover: Spatially Valid Pixel Learning for Image Specular Highlight Removal
Ling Zhang 0017, Yidong Ma, Weilei He, Zhongyun Bao, Gang Fu 0003, Wenju Xu, Chunxia Xiao |
ACM Multimedia | 7 |
| 2024 | Shopping MMLU: A Massive Multi-Task Online Shopping Benchmark for Large Language ModelsabstractOnline shopping is a complex multi-task, few-shot learning problem with a wide and evolving range of entities, relations, and tasks. However, existing models and benchmarks are commonly tailored to specific tasks, falling short of capturing the full complexity of online shopping. Large Language Models (LLMs), with their multi-task and few-shot learning abilities, have the potential to profoundly transform online shopping by alleviating task-specific engineering efforts and by providing users with interactive conversations. Despite the potential, LLMs face unique challenges in online shopping, such as domain-specific concepts, implicit knowledge, and heterogeneous user behaviors. Motivated by the potential and challenges, we propose Shopping MMLU, a diverse multi-task online shopping benchmark derived from real-world Amazon data. Shopping MMLU consists of 57 tasks covering 4 major shopping skills: concept understanding, knowledge reasoning, user behavior alignment, and multi-linguality, and can thus comprehensively evaluate the abilities of LLMs as general shop assistants. With Shoppping MMLU, we benchmark over 20 existing LLMs and uncover valuable insights about practices and prospects of building versatile LLM-based shop assistants. Shopping MMLU can be publicly accessed at https://github.com/KL4805/ShoppingMMLU. In addition, with Shopping MMLU, we are hosting a competition in KDD Cup 2024 with over 500 participating teams. The winning solutions and the associated workshop can be accessed at our website https://amazon-kddcup24.github.io/. Yilun Jin, Zheng Li 0018, Tianyu Cao 0001, Yifan Gao 0001, Pratik Jayarao, Xin Liu 0039, Ritesh Sarkhel, Xianfeng Tang, Wenju Xu, Jingfeng Yang 0001, Qingyu Yin, Priyanka Nigam, Yi Xu 0011, Kai Chen 0005, Qiang Yang 0001, Meng Jiang 0001 |
NeurIPS | 13 |
| 2024 | Verifiable privacy-preserving cox regression from multi-key fully homomorphic encryption
Wenju Xu, Yunxuan Su, Baocang Wang |
Peer Peer Netw. Appl. | 1 |
| 2024 | PatchMixing Masked Autoencoders for 3D Point Cloud Self-Supervised LearningabstractRecently, Point-MAE has extended Masked Autoencoders (MAE) to point clouds for 3D self-supervised learning, which however faces two problems: (1) the shape similarity between the masked point cloud and original point cloud is high, and (2) the pretext task of reconstructing the original point cloud is straightforward which fails to compel the network to learn deep representative features. In this paper, we tackle these problems by proposing a PatchMixing strategy and a teacher-student training framework. First, with PatchMixing, we mix selected point patches of multiple point clouds and attempt to infer the object information from the resulting mixed point cloud. Due to the interference of other objects, the task is challenging but facilitates representation learning. Second, rather than directly restoring the original point cloud, we propose a novel pretext task that involves a two-branch teacher model and a student model. These models process the multiple input point clouds in different ways (no mixing, mixing + unmixing, mixing + masking), but are expected to output similar features, thereby compelling the network to extract essential features from the input. Extensive experiments show that our well-designed PatchMixing strategy and effective teacher-student learning architecture yield impressive results. Specifically, our model achieves a remarkable 92.9% classification accuracy in the Linear SVM task on the ModelNet40 dataset. Through pre-training and fine-tuning on downstream tasks, our method achieves an 89.8% classification accuracy on the most challenging split of ScanObjectNN and an outstanding 94.0% on ModelNet40. Chengxing Lin 0001, Wenju Xu, Jian Zhu 0001, Yongwei Nie, Ruichu Cai, Xuemiao Xu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Disentangled Representation Learning for Controllable Person Image GenerationabstractIn this paper, we propose a novel framework named DRL-CPG to learn disentangled latent representation for controllable person image generation, which can produce realistic person images with desired poses and human attributes (e.g. pose, head, upper clothes, and pants) provided by various source persons. Unlike the existing works leveraging the semantic masks to obtain the representation of each component, we propose to generate disentangled latent code via a novel attribute encoder with transformers trained in a manner of curriculum learning from a relatively easy step to a gradually hard one. A random component mask-agnostic strategy is introduced to randomly remove component masks from the person segmentation masks, which aims at increasing the difficulty of training and promoting the transformer encoder to recognize the underlying boundaries between each component. This enables the model to transfer both the shape and texture of the components. Furthermore, we propose a novel attribute decoder network to integrate multi-level attributes (e.g. the structure feature and the attribute representation) with well-designed Dual Adaptive Denormalization (DAD) residual blocks. Extensive experiments strongly demonstrate that the proposed approach is able to transfer both the texture and shape of different human parts and yield realistic results. To our knowledge, we are the first to learn disentangled latent representations with transformers for person image generation. Wenju Xu, Chengjiang Long, Yongwei Nie, Guanghui Wang 0001 |
IEEE Trans. Multim. | 1 |
| 2023 | Learning Dynamic Style Kernels for Artistic Style TransferabstractArbitrary style transfer has been demonstrated to be efficient in artistic image generation. Previous methods either globally modulate the content feature ignoring local details, or overly focus on the local structure details leading to style leakage. In contrast to the literature, we propose a new scheme “style kernel” that learns spatially adaptive kernels for per-pixel stylization, where the convolutional kernels are dynamically generated from the global style-content aligned feature and then the learned kernels are applied to modulate the content feature at each spatial position. This new scheme allows flexible both global and local interactions between the content and style features such that the wanted styles can be easily transferred to the content image while at the same time the content structure can be easily preserved. To further enhance the flexibility of our style transfer method, we propose a Style Alignment Encoding (SAE) module complemented with a Content-based Gating Modulation (CGM) module for learning the dynamic style kernels in focusing regions. Extensive experiments strongly demonstrate that our proposed method outperforms state-of-the-art methods and exhibits superior performance in terms of visual quality and efficiency. Wenju Xu, Chengjiang Long, Yongwei Nie |
CVPR | 1 |
| 2023 | Feature Representation Learning with Adaptive Displacement Generation and Transformer Fusion for Micro-Expression RecognitionabstractMicro-expressions are spontaneous, rapid and subtle facial movements that can neither be forged nor suppressed. They are very important nonverbal communication clues, but are transient and of low intensity thus difficult to recognize. Recently deep learning based methods have been developed for micro-expression (ME) recognition using feature extraction and fusion techniques, however, targeted feature learning and efficient feature fusion still lack further study according to the ME characteristics. To address these issues, we propose a novel framework Feature Representation Learning with adaptive Displacement Generation and Transformer fusion (FRL-DGT), in which a convolutional Displacement Generation Module (DGM) with self-supervised learning is used to extract dynamic features from onset/apex frames targeted to the subsequent ME recognition task, and a well-designed Transformer Fusion mechanism composed of three Transformer-based fusion modules (local, global fusions based on AU regions and full-face fusion) is applied to extract the multi-level informative features after DGM for the final ME prediction. The extensive experiments with solid leave-one-subject-out (LOSO) evaluation results have demonstrated the superiority of our proposed FRL-DGT to state-of-the-art methods. Zhijun Zhai, Jianhui Zhao 0001, Chengjiang Long, Wenju Xu, Shuangjiang He, Huijuan Zhao |
CVPR | 4 |
| 2023 | Multi-key Fully Homomorphic Encryption from Additive HomomorphismabstractAbstract Fully homomorphic encryption (FHE) allows direct computations over the encrypted data without access to the decryption. Hence multi-key FHE is well suitable for secure multiparty computation. Recently, Brakerski et al. (TCC 2019 and EUROCRYPT 2020) utilized additively homomorphic encryption to construct FHE schemes with different properties. Motivated by their work, we are attempting to construct multi-key FHE schemes via additively homomorphic encryption. In this paper, we propose a general framework of constructing multi-key FHE, combining the additively homomorphic encryption with specific multiparty computation protocols constructed from encryption switching protocol. Concretely, every involved party encrypts his plaintexts with an additively homomorphic encryption under his own public key. Then the ciphertexts are evaluated by suitable multiparty computation protocols performed by two cooperative servers without collusion. Furthermore, an instantiation with an ElGamal variant scheme is presented. Performance comparisons show that our multi-key FHE from additively homomorphic encryption is more efficient and practical. Wenju Xu, Baocang Wang, Yupu Hu, Pu Duan, Benyu Zhang, Momeng Liu |
Comput. J. | 1 |
| 2023 | Modified Multi-Key Fully Homomorphic Encryption Scheme in the Plain ModelabstractAbstract Multi-key fully homomorphic encryption (MFHE) supports arbitrary meaningful computations on encrypted data under different public keys even without access to the secret key, which is well tailored for the secure multiparty computation scenarios. Based on the Gentry–Sahai–Waters scheme (a single-key FHE in Crypto 2013) with the underlying learning with errors problem, MW16 scheme (Eurocrypt 2016) utilizes the method of ‘linear combination procedure’ (LCP) as a subroutine to construct the auxiliary information for the expanded ciphertexts of MFHE scheme. However, every party shares a common random string (CRS) to be distributed by a trusted setup, which is unpractical. Meanwhile, the noise in the auxiliary information is too much compared with the one in fresh ciphertexts. In this paper, we propose a modified MFHE scheme in the plain model, i.e. without CRS, to enhance the practicability of MFHE. Specifically, every involved party generates his own public key independent on a CRS. Then a potential improvement on the LCP is developed to provide auxiliary information, which largely reduces the noise and leads to a smaller modulus for our MFHE. Furthermore, the feasibility of our proposal is also proved by theoretical performance comparisons. Wenju Xu, Baocang Wang, Quanbo Qu, Tanping Zhou, Pu Duan |
Comput. J. | 1 |
| 2022 | Toward practical privacy-preserving linear regression
Wenju Xu, Baocang Wang, Jiasen Liu, Yange Chen, Pu Duan, Zhiyong Hong |
Inf. Sci. | 1 |
| 2022 | Updatable privacy-preserving itK-nearest neighbor query in location-based s-ervice
Wenju Xu, Zhiyong Hong, Pu Duan, Benyu Zhang, Yupu Hu, Baocang Wang |
Peer-to-Peer Netw. Appl. | 2 |
| 2022 | A Domain Gap Aware Generative Adversarial Network for Multi-Domain Image TranslationabstractRecent image-to-image translation models have shown great success in mapping local textures between two domains. Existing approaches rely on a cycle-consistency constraint that supervises the generators to learn an inverse mapping. However, learning the inverse mapping introduces extra trainable parameters and it is unable to learn the inverse mapping for some domains. As a result, they are ineffective in the scenarios where (i) multiple visual image domains are involved; (ii) both structure and texture transformations are required; and (iii) semantic consistency is preserved. To solve these challenges, the paper proposes a unified model to translate images across multiple domains with significant domain gaps. Unlike previous models that constrain the generators with the ubiquitous cycle-consistency constraint to achieve the content similarity, the proposed model employs a perceptual self-regularization constraint. With a single unified generator, the model can maintain consistency over the global shapes as well as the local texture information across multiple domains. Extensive qualitative and quantitative evaluations demonstrate the effectiveness and superior performance over state-of-the-art models. It is more effective in representing shape deformation in challenging mappings with significant dataset variation across multiple domains. Wenju Xu, Guanghui Wang 0001 |
IEEE Trans. Image Process. | 1 |
| 2022 | Privacy-preserving association rule mining based on electronic medical system
Wenju Xu, Baocang Wang, Yupu Hu |
Wirel. Networks | 1 |
| 2021 | DRB-GAN: A Dynamic ResBlock Generative Adversarial Network for Artistic Style TransferabstractThe paper proposes a Dynamic ResBlock Generative Adversarial Network (DRB-GAN) for artistic style transfer. The style code is modeled as the shared parameters for Dynamic ResBlocks connecting both the style encoding network and the style transfer network. In the style encoding network, a style class-aware attention mechanism is used to attend the style feature representation for generating the style codes. In the style transfer network, multiple Dynamic ResBlocks are designed to integrate the style code and the extracted CNN semantic feature and then feed into the spatial window Layer-Instance Normalization (SW-LIN) decoder, which enables high-quality synthetic images with artistic style transfer. Moreover, the style collection conditional discriminator is designed to equip our DRB-GAN model with abilities for both arbitrary style transfer and collection style transfer during the training stage. No matter for arbitrary style transfer or collection style transfer, extensive experiments strongly demonstrate that our proposed DRB-GAN outperforms state-of-the-art methods and exhibits its superior performance in terms of visual quality and efficiency. Our source code is available at https://github.com/xuwenju123/DRB-GAN. Wenju Xu, Chengjiang Long, Ruisheng Wang 0001, Guanghui Wang 0001 |
ICCV | 1 |
| 2021 | Dual Graph Convolutional Networks with Transformer and Curriculum Learning for Image CaptioningabstractExisting image captioning methods just focus on understanding the relationship between objects or instances in a single image, without exploring the contextual correlation existed among contextual image. In this paper, we propose Dual Graph Convolutional Networks (Dual-GCN) with transformer and curriculum learning for image captioning. In particular, we not only use an object-level GCN to capture the object to object spatial relation within a single image, but also adopt an image-level GCN to capture the feature information provided by similar images. With the well-designed Dual-GCN, we can make the linguistic transformer better understand the relationship between different objects in a single image and make full use of similar images as auxiliary information to generate a reasonable caption description for a single image. Meanwhile, with a cross-review strategy introduced to determine difficulty levels, we adopt curriculum learning as the training strategy to increase the robustness and generalization of our proposed model. We conduct extensive experiments on the large-scale MS COCO dataset, and the experimental results powerfully demonstrate that our proposed method outperforms recent state-of-the-art approaches. It achieves a BLEU-1 score of 82.2 and a BLEU-2 score of 67.6. Our source code is available at https://github.com/Unbear430/DGCN-for-image-captioning. Xinzhi Dong, Chengjiang Long, Wenju Xu, Chunxia Xiao |
ACM Multimedia | 3 |
| 2021 | Efficient Private Information Retrieval Protocol with Homomorphically Computing Univariate PolynomialsabstractPrivate information retrieval (PIR) protocol is a powerful cryptographic tool and has received considerable attention in recent years as it can not only help users to retrieve the needed data from database servers but also protect them from being known by the servers. Although many PIR protocols have been proposed, it remains an open problem to design an efficient PIR protocol whose communication overhead is irrelevant to the database size N . In this paper, to answer this open problem, we present a new communication-efficient PIR protocol based on our proposed single-ciphertext fully homomorphic encryption (FHE) scheme, which supports unlimited computations with single variable over a single ciphertext even without access to the secret key. Specifically, our proposed PIR protocol is characterized by combining our single-ciphertext FHE with Lagrange interpolating polynomial technique to achieve better communication efficiency. Security analyses show that the proposed PIR protocol can efficiently protect the privacy of the user and the data in the database. In addition, both theoretical analyses and experimental evaluations are conducted, and the results indicate that our proposed PIR protocol is also more efficient and practical than previously reported ones. To the best of our knowledge, our proposed protocol is the first PIR protocol achieving O1 communication efficiency on the user side, irrelevant to the database size N . Wenju Xu, Baocang Wang, Rongxing Lu, Quanbo Qu, Yange Chen, Yupu Hu |
Secur. Commun. Networks | 1 |
| 2020 | FX-GAN: Self-Supervised GAN Learning via Feature ExchangeabstractWe propose a self-supervised approach to improve the training of Generative Adversarial Networks (GANs) via inducing the discriminator to examine the structural consistency of images. Although natural image samples provide ideal examples of both valid structure and valid texture, learning to reproduce both together remains an open challenge. In our approach, we augment the training set of natural images with modified examples that have degraded structural consistency. These degraded examples are automatically created by randomly exchanging pairs of patches in an image’s convolutional feature map. We call this approach feature exchange. With this setup, we propose a novel GAN formulation, termed Feature eXchange GAN (FX-GAN), in which the discriminator is trained not only to distinguish real versus generated images, but also to perform the auxiliary task of distinguishing between real images and structurally corrupted (feature-exchanged) real images. This auxiliary task causes the discriminator to learn the proper feature structure of natural images, which in turn guides the generator to produce images with more realistic structure. Compared with strong GAN baselines, our proposed self-supervision approach improves generated image quality, diversity, and training stability for both the unconditional and class-conditional settings. Wenju Xu, Teng-Yok Lee, Anoop Cherian, Ye Wang 0001, Tim K. Marks |
WACV | 2 |
| 2020 | Towards Learning Affine-Invariant Representations via Data-Efficient CNNsabstractIn this paper we propose integrating a priori knowledge into both design and training of convolutional neural networks (CNNs) to learn object representations that are invariant to affine transformations (i.e. translation, scale, rotation). Accordingly we propose a novel multi-scale maxout CNN and train it end-to-end with a novel rotation-invariant regularizer. This regularizer aims to enforce the weights in each 2D spatial filter to approximate circular patterns. In this way, we manage to handle affine transformations in training using convolution, multi-scale maxout, and circular filters. Empirically we demonstrate that such knowledge can significantly improve the data-efficiency as well as generalization and robustness of learned models. For instance, on the Traffic Sign data set and trained with only 10 images per class, our method can achieve 84.15% that outperforms the state-of-the-art by 29.80% in terms of test accuracy. Wenju Xu, Guanghui Wang 0001, Alan Sullivan |
WACV | 1 |
| 2020 | Adaptively Denoising Proposal Collection for Weakly Supervised Object Localization
Wenju Xu, Yuanwei Wu, Wenchi Ma, Guanghui Wang 0001 |
Neural Process. Lett. | 1 |
| 2019 | Stacked Wasserstein Autoencoder
Wenju Xu, Shawn Shahriar Keshmiri, Guanghui Wang 0001 |
Neurocomputing | 1 |
| 2019 | Toward learning a unified many-to-many mapping for diverse image translation
Wenju Xu, Shawn Shahriar Keshmiri, Guanghui Wang 0001 |
Pattern Recognit. | 1 |
| 2019 | Adversarially Approximated Autoencoder for Image Generation and ManipulationabstractRegularized autoencoders learn the latent codes, a structure with the regularization under the distribution, which enables them the capability to infer the latent codes given observations and generate new samples given the codes. However, they are sometimes ambiguous as they tend to produce reconstructions that are not necessarily a faithful reproduction of the inputs. The main reason is to enforce the learned latent code distribution to match a prior distribution while the true distribution remains unknown. To improve the reconstruction quality and learn the latent space a manifold structure, this paper presents a novel approach using the adversarially approximated autoencoder (AAAE) to investigate the latent codes with adversarial approximation. Instead of regularizing the latent codes by penalizing on the distance between the distributions of the model and the target, AAAE learns the autoencoder flexibly and approximates the latent space with a simpler generator. The ratio is estimated using a generative adversarial network to enforce the similarity of the distributions. In addition, the image space is regularized with an additional adversarial regularizer. The proposed approach unifies two deep generative models for both latent space inference and diverse generation. The learning scheme is realized without regularization on the latent codes, which also encourages faithful reconstruction. Extensive validation experiments on four real-world datasets demonstrate the superior performance of AAAE. In comparison to the state-of-the-art approaches, AAAE generates samples with better quality and shares the properties of a regularized autoencoder with a nice latent manifold structure. Wenju Xu, Shawn Shahriar Keshmiri, Guanghui Wang 0001 |
IEEE Trans. Multim. | 1 |