VLDB 2026 Research / reviewers in the wild / expert
Fuxiang Wu
dblp:178/7264
· DBLP profile ↗
25ranked-venue papers
6as first author
21since 2021 · last 2026
0000-0002-4542-4486ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 5 first-author · 13 since 2021Artificial intelligence and machine learning · 13 · 3 first-author · 11 since 2021Computer networks · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Diffusion-Based Text-Guided Image Generation With Fine-Grained Spatial Object-Attribute RelationshipsabstractExpressing and controlling fine-grained spatial attributes of objects in large-scale models presents significant challenges, as these spatial attributes are often difficult to describe textually and exhaustive enumeration is impractical. This hinders effective alignment with user preferences regarding spatial attribute-object relationships in fine-grained synthesis tasks. To tackle this problem, we propose AttrObjDiff, a novel framework built on the pre-trained Stable Diffusion model to integrate spatial attribute maps. Firstly, AttrObjDiff constrains the denoising step using trainable cross-attention fusion modules, attribute-enhancing cross-attention and LoRAs. The fusion modules take layout features extracted by a frozen ControlNet and corresponding fine-grained attribute maps as inputs to generate joint constraint features of spatial attribute-object relationships. We leverage attribute-enhancing cross-attention within the U-Net to further refine these spatial attributes. Finally, LoRAs are employed to align with these joint constraint features of finegrained relationships. Secondly, AttrObjDiff enhances the reverse process with lightweight noise reranking models to improve spatial object-attribute alignment. The reranking models select semantic noises related to fine-grained relationships, improving synthesis quality without significantly increasing computational costs. Experimental results demonstrate that our method can generate high-quality images guided by fine-grained spatial object-attribute relationships, improving synthesis controllability and semantic consistency. Fuxiang Wu, Liu Liu 0014, Fusheng Hao, Ziliang Ren, Dacheng Tao, Xinyu Wu 0001, Jun Cheng 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2026 | NoisePO: Efficient Semantic Noise Generation and Ranking for Diffusion-Based Text-to-Image SynthesisabstractDiffusion-based methods have achieved remarkable success in photorealistic image generation, leveraging iterative denoising steps to improve image quality. However, multi-step denoising often suffers from error accumulation-similar to exposure bias in autoregressive models-due to suboptimal noise estimation, which can lead to degraded semantic alignment and image fidelity. To tackle the challenge of suboptimal inner latent representations in generation and improve the inner latent, this paper introduces a novel method NoisePO, an efficient semantic noise preference optimization framework. NoisePO employs a semantic noise preference optimization generative adversarial network (NPO-GAN) and noise ranking methods to search for semantically relevant noises based on textual conditions, thus eliminating undesired semantic features while emphasizing the necessary semantic ones. Specifically, NoisePO utilizes a light NPO-GAN to generate semantic noises that encourage the latent at the previous step to incorporate more semantic information from the caption. Then, light ranking models are employed to filter out low-quality noises and select the best noise. Experimental results demonstrate that NoisePO consistently outperforms the baselines across widely used frameworks, achieving notable improvements in image quality, semantic consistency, and user-specific alignment as measured by IS, FID, CLIP, and other metrics. These results indicate that NoisePO effectively enhances synthesis quality and strengthens text-image alignment. Fuxiang Wu, Liu Liu 0014, Fusheng Hao, Chengqun Song, Dacheng Tao, Jun Cheng 0002 |
IEEE Trans. Image Process. | 1 |
| 2026 | Language-Guided Multimodal Spiking Neural Networks for Event-Based Action RecognitionabstractEvent-based human action recognition has gained increasing attention due to its efficiency in dynamic scenarios. Contemporary methodologies for event-based action recognition predominantly treat the problem as a one-hot classification task, which limits their ability to leverage the semantic relationships among various actions. To address this limitation, we propose a Spiking Event-Text Feature Fusion (SETFF) framework, which enhances recognition performance by integrating event and text modalities through a dual-stream architecture. SETFF leverages generative large language models to produce action descriptions, serving as semantic prompts that guide event feature learning. Specifically, a contrastive loss function is employed to align the features of both modalities, enriching the model's capacity to distinguish intricate and subtle actions. Extensive experiments on neuromorphic datasets, including PAF, DailyAction-DVS, DVS128 Gesture, Bullying10K, and UCF101-DVS, demonstrate that SETFF achieves state-of-the-art accuracy, with top-1 accuracy rates of up to 99.65% on the DailyAction-DVS dataset and 98.39% on the PAF dataset. Experimental results underscore the effectiveness of multimodal fusion in SNNs, advancing event-based action recognition while preserving the energy efficiency characteristic of SNNs. Ziliang Ren, Fuxiang Wu, Qieshi Zhang, Jun Cheng 0002 |
IEEE Trans. Multim. | 3 |
| 2025 | Task-Aware Clustering for Prompting Vision-Language ModelsabstractPrompt learning has attracted widespread attention in adapting vision-language models to downstream tasks. Existing methods largely rely on optimization strategies to ensure the task-awareness of learnable prompts. Due to the scarcity of task-specific data, overfitting is prone to occur. The resulting prompts often do not generalize well or exhibit limited task-awareness. To address this issue, we propose a novel Task-Aware Clustering (TAC) framework for prompting vision-language models, which increases the task-awareness of learnable prompts by introducing task-aware pre-context. The key ingredients are as follows: (a) generating task-aware pre-context based on task-aware clustering that can preserve the backbone structure of a downstream task with only a few clustering centers, (b) enhancing the task-awareness of learnable prompts by enabling them to interact with task-aware pre-context via the well-pretrained encoders, and (c) preventing the visual task-aware pre-context from interfering the interaction between patch embeddings by masked attention mechanism. Extensive experiments are conducted on benchmark datasets, covering the base-to-novel, domain generalization, and cross-dataset transfer settings. Ablation studies validate the effectiveness of key ingredients. Comparative results show the superiority of our TAC over competitive counterparts. The code is available at https://github.com/FushengHao/TAC. Fusheng Hao, Fengxiang He, Fuxiang Wu, Tichao Wang, Chengqun Song, Jun Cheng 0002 |
CVPR | 3 |
| 2025 | Human-Imperceptible, Machine-Recognizable ImagesabstractMassive human-related data is collected to train neural networks for computer vision tasks. A major conflict is exposed relating to software engineers between better developing AI systems and distancing from the sensitive training data. To reconcile this conflict, the paper proposes an efficient privacy-preserving learning paradigm, where images are encrypted to become ``human-imperceptible, machine-recognizable'' via one of the two encryption strategies: (1) random shuffling equally-sized patches and (2) mixing-up sub-patches. Then, minimal adaptations are made to vision transformer to enable it to learn on the encrypted images for vision tasks, including image classification and object detection. Extensive experiments on ImageNet and COCO show that the proposed paradigm achieves comparable accuracy with the competitive methods. Decrypting the encrypted images requires solving an NP-hard jigsaw puzzle or ill-posed inverse problem, which is empirically shown intractable to be recovered by various attackers, including the powerful vision transformer-based attacker. We thus show that the proposed paradigm can ensure the encrypted images have become human-imperceptible while preserving machine-recognizable information. Fusheng Hao, Fengxiang He, Yikai Wang 0001, Fuxiang Wu, Jing Zhang 0037, Dacheng Tao, Jun Cheng 0002 |
IJCAI | 4 |
| 2025 | SIGMA: Refining Large Language Model Reasoning via Sibling-Guided Monte Carlo AugmentationabstractEnhancing large language models by simply scaling up datasets has begun to yield diminishing returns, shifting the spotlight to data quality. Monte Carlo Tree Search (MCTS) has emerged as a powerful technique for generating high-quality chain-of-thought data, yet conventional approaches typically retain only the top-scoring trajectory from the search tree, discarding sibling nodes that often contain valuable partial insights, recurrent error patterns, and alternative reasoning strategies. This unconditional rejection of non-optimal reasoning branches may waste vast amounts of informative data in the whole search tree. We propose SIGMA (Sibling Guided Monte Carlo Augmentation), a novel framework that reintegrates these discarded sibling nodes to refine LLM reasoning. SIGMA forges semantic links among sibling nodes along each search path and applies a two-stage refinement: a critique model identifies overlooked strengths and weaknesses across the sibling set, and a revision model conducts text-based backpropagation to refine the top-scoring trajectory in light of this comparative feedback. By recovering and amplifying the underutilized but valuable signals from non-optimal reasoning branches, SIGMA substantially improves reasoning trajectories. On the challenging MATH benchmark, our SIGMA-tuned 7B model achieves 54.92\% accuracy using only 30K samples, outperforming state-of-the-art models trained on 590K samples. This result highlights that our sibling-guided optimization not only significantly reduces data usage but also significantly boosts LLM reasoning. Yanwei Ren, Fuxiang Wu, Jiayan Qiu, Jiaxing Huang 0001, Baosheng Yu, Liu Liu 0014 |
NeurIPS | 3 |
| 2025 | Textual Embeddings are Good Class-Aware Visual Prompts for Adapting Vision-Language ModelsabstractDue to the parallel nature of the textual and visual encoders, very little attention has been paid to developing prompt learning by using well-pretrained encoders in a serial manner, in which the low-biased high-level semantic information accessible to each other for these encoders is ignored. In this letter, we find that textual embeddings are good class-aware visual prompts for adapting vision-language models, which leads to a new framework called TVPrompt (Textual embeddings as class-aware Visual Prompts). To eliminate the modal gap between text and vision, we design a bridging module, which integrates textual embeddings and class token to produce class-aware visual prompts. To ensure that such prompts could effectively collect class-relevant information, we further propose using masked attention to block the unnecessary interactions. Experimental evidence on benchmark datasets demonstrates that our TVPrompt achieves competitive efficiency and performance. Fusheng Hao, Liu Liu 0014, Fuxiang Wu, Qieshi Zhang, Jun Cheng 0002 |
IEEE Signal Process. Lett. | 3 |
| 2025 | Class-Irrelevant Feature Removal for Few-Shot Image ClassificationabstractMost existing few-shot image classification methods employ global pooling to aggregate class-relevant local features in a data-drive manner. Due to the difficulty and inaccuracy in locating class-relevant regions in complex scenarios, as well as the large semantic diversity of local features, the class-irrelevant information could reduce the robustness of the representations obtained by performing global pooling. Meanwhile, the scarcity of labeled images exacerbates the difficulties of data-hungry deep models in identifying class-relevant regions. These issues severely limit deep models' few-shot learning ability. In this work, we propose to remove the class-irrelevant information by making local features class relevant, thus bypassing the big challenge of identifying which local features are class irrelevant. The resulting class-irrelevant feature removal (CIFR) method consists of three phases. First, we employ the masked image modeling strategy to build an understanding of images' internal structures that generalizes well. Second, we design a semantic-complementary feature propagation module to make local features class relevant. Third, we introduce a weighted dense-connected similarity measure, based on which a loss function is raised to fine-tune the entire pipeline, with the aim of further enhancing the semantic consistency of the class-relevant local features. Visualization results show that CIFR achieves the removal of class-irrelevant information by making local features related to classes. Comparison results on four benchmark datasets indicate that CIFR yields very promising performance. Fusheng Hao, Liu Liu 0014, Fuxiang Wu, Qieshi Zhang, Jun Cheng 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Two-stage feature distribution rectification for few-shot point cloud semantic segmentation
Tichao Wang, Fusheng Hao, Guosheng Cui, Fuxiang Wu, Mengjie Yang, Qieshi Zhang, Jun Cheng 0002 |
Pattern Recognit. Lett. | 4 |
| 2023 | Reject Decoding via Language-Vision Models for Text-to-Image SynthesisabstractTransformer-based text-to-image synthesis generates images from abstractive textual conditions and achieves prompt results. Since transformer-based models predict visual tokens step by step in testing, where the early error is hard to be corrected and would be propagated. To alleviate this issue, the common practice is drawing multi-paths from the transformer-based models and re-ranking the multi-images decoded from multi-paths to find the best one and filter out others. Therefore, the computing procedure of excluding images may be inefficient. To improve the effectiveness and efficiency of decoding, we exploit a reject decoding algorithm with tiny multi-modal models to enlarge the searching space and exclude the useless paths as early as possible. Specifically, we build tiny multi-modal models to evaluate the similarities between the partial paths and the caption at multi scales. Then, we propose a reject decoding algorithm to exclude some lowest quality partial paths at the inner steps. Thus, under the same computing load as the original decoding, we could search across more multi-paths to improve the decoding efficiency and synthesizing quality. The experiments conducted on the MS-COCO dataset and large-scale datasets show that the proposed reject decoding algorithm can exclude the useless paths and enlarge the searching paths to improve the synthesizing quality by consuming less time. Fuxiang Wu, Liu Liu 0014, Fusheng Hao, Fengxiang He, Lei Wang 0018, Jun Cheng 0002 |
AAAI | 1 |
| 2023 | Class-Aware Patch Embedding Adaptation for Few-Shot Image Classificationabstract"A picture is worth a thousand words", significantly beyond mere a categorization. Accompanied by that, many patches of the image could have completely irrelevant meanings with the categorization if they were independently observed. This could significantly reduce the efficiency of a large family of few-shot learning algorithms, which have limited data and highly rely on the comparison of image patches. To address this issue, we propose a Class-aware Patch Embedding Adaptation (CPEA) method to learn "class-aware embeddings" of the image patches. The key idea of CPEA is to integrate patch embeddings with class-aware embeddings to make them class-relevant. Furthermore, we define a dense score matrix between class-relevant patch embeddings across images, based on which the degree of similarity between paired images is quantified. Visualization results show that CPEA concentrates patch embeddings by class, thus making them class-relevant. Extensive experiments on four benchmark datasets, miniImageNet, tieredImageNet, CIFAR-FS, and FC-100, indicate that our CPEA significantly outperforms the existing state-of-the-art methods. The source code is available at https://github.com/FushengHao/CPEA. Fusheng Hao, Fengxiang He, Liu Liu 0014, Fuxiang Wu, Dacheng Tao, Jun Cheng 0002 |
ICCV | 4 |
| 2023 | Prototype expansion and feature calibration for few-shot point cloud semantic segmentation
Qieshi Zhang, Tichao Wang, Fusheng Hao, Fuxiang Wu, Jun Cheng 0002 |
Neurocomputing | 4 |
| 2023 | Semantic-Aware Feature Aggregation for Few-Shot Image Classification
Fusheng Hao, Fuxiang Wu, Fengxiang He, Qieshi Zhang, Chengqun Song, Jun Cheng 0002 |
Neural Process. Lett. | 2 |
| 2023 | InDecGAN: Learning to Generate Complex Images From Captions via Independent Object-Level Decomposition and EnhancementabstractText-to-image synthesis is a challenging problem, in which a complex scene contains diverse objects of various sizes and sub-images of objects belonging to the same class have diverse forms from different perspectives. Thus, synthesis models have difficulty in capturing varied objects in the complex scene. To alleviate these problems, we devise an independent object-level decomposing and enhancing generative adversarial networks, denoted as InDecGAN, to synthesize complex images and capture varied objects in a complex scene. Specifically, InDecGAN fully utilizes the independent object-level information, bounding boxes and high-resolution images of objects in training, by employing independent object-level pathways to synthesize varied objects. The independent object-level pathway integrates an independent object-level adversarial loss and the bounding box information to learn the visual features of objects independently, then, the main pathway exploits the features provided by the object-level pathway to compose the full scene and synthesize images. In addition, we analyze the generalization properties of the proposed InDecGAN and demonstrate the improvement from the perspective of the model architecture. Moreover, extensive experiments conducted on a widely used dataset are presented to demonstrate that the proposed model with an independent object-level pathway produces synthesized images of significantly improved quality. Jun Cheng 0002, Fuxiang Wu, Liu Liu 0014, Qieshi Zhang, Leszek Rutkowski, Dacheng Tao |
IEEE Trans. Multim. | 2 |
| 2023 | Language-Based Image Manipulation Built on Language-Guided RankingabstractText-based image manipulation is a popular subject and has many applications. However, it is a challenging task because there is no ground-truth edited dataset and textual descriptions have abstractive and ambiguous properties. To alleviate the difficult issues, we propose a manipulation framework consisting of the proposal attentional GANs, language-related semantic mask, and language-guided ranker. Specially, we construct an editing proposal generator to generate the suitable edited proposals with and without semantic conditions, which supports the reorganization of sub-generators to output proposals in various aspects as many as possible. To distinguish the text-relevant and the text-irrelevant regions, we introduce a language-related semantic mask based on the source image and target caption. Then, we exploit a language-guided ranker to retrieve the best edited result from the edited proposals through using the multi-modal similarity and the language-related semantic mask. Extensive experiments on widely-used datasets demonstrate that our model could manipulate images interactively and improve the editing quality effectively. Fuxiang Wu, Liu Liu 0014, Fusheng Hao, Fengxiang He, Jun Cheng 0002 |
IEEE Trans. Multim. | 1 |
| 2022 | Text-to-Image Synthesis based on Object-Guided Joint-Decoding TransformerabstractObject-guided text-to-image synthesis aims to generate images from natural language descriptions built by two-step frameworks, i.e., the model generates the layout and then synthesizes images from the layout and captions. However, such frameworks have two issues: 1) complex structure, since generating language-related layout is not a trivial task; 2) error propagation, because the inappropriate layout will mislead the image synthesis and is hard to be revised. In this paper, we propose an object-guided joint-decoding module to simultaneously generate the image and the corresponding layout. Specially, we present the joint-decoding transformer to model the joint probability on images tokens and the corresponding layouts tokens, where layout tokens provide additional observed data to model the complex scene better. Then, we describe a novel Layout-Vqgan for layout encoding and decoding to provide more information about the complex scene. After that, we present the detail-enhanced module to enrich the language-related details based on two facts: 1) visual details could be omitted in the compression of VQGANs; 2) the joint-decoding transformer would not have sufficient generating capacity. The experiments show that our approach is competitive with previous object-centered models and can generate diverse and high-quality objects under the given layouts. Fuxiang Wu, Liu Liu 0014, Fusheng Hao, Fengxiang He, Jun Cheng 0002 |
CVPR | 1 |
| 2022 | EEP-Net: Enhancing Local Neighborhood Features and Efficient Semantic Segmentation of Scale Point Clouds
Fuxiang Wu, Qieshi Zhang, Ziliang Ren, Jun Cheng 0002 |
PRCV (3) | 2 |
| 2022 | RiFeGAN2: Rich Feature Generation for Text-to-Image Synthesis From Constrained Prior KnowledgeabstractText-to-image synthesis is a challenging task that generates realistic images from a textual description. The description contains limited information compared with the corresponding image and is ambiguous and abstract, which will complicate the generation and lead to low-quality images. To address this problem, we propose a novel generation text-to-image synthesis method, called RiFeGAN2, to enrich the given description. To improve the enrichment quality while accelerating the enrichment process, RiFeGAN2 exploits a domain-specific constrained model to limit the search scope and then uses an attention-based caption matching model to refine the compatible candidate captions based on constrained prior knowledge. To improve the semantic consistency between the given description and the synthesized results, RiFeGAN2 employs improved SAEMs, SAEM2s, to compact better features of the retrieved captions and effectively emphasize the descriptions via incorporating centre-attention layers. Finally, multi-caption attentional GANs are exploited to synthesize images from those features. Experiments performed on widely-used datasets show that the models can generate vivid images from enriched captions and effectually improve the semantic consistency. Jun Cheng 0002, Fuxiang Wu, Yanling Tian, Lei Wang 0018, Dapeng Tao |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Image Hallucination From Attribute PairsabstractRecent image-generation methods have demonstrated that realistic images can be produced from captions. Despite the promising results achieved, existing caption-based generation methods confront a dilemma. On the one hand, the image generator should be provided with sufficient details for realistic hallucination, meaning that longer sentences with rich content are preferred, but on the other hand, the generator is meanwhile fragile to long sentences due to their complex semantics and syntax like long-range dependencies and the combinatorial explosion of object visual features. Toward alleviating this dilemma, a novel approach is proposed in this article to hallucinate images from attribute pairs, which can be extracted from natural language processing (NLP) toolsets in the presence of complex semantics and syntax. Attribute pairs, therefore, enable our image generator to tackle long sentences handily and alleviate the combinatorial explosion, and at the same time, allow us to enlarge the training dataset and to produce hallucinations from randomly combined attribute pairs at ease. Experiments on widely used datasets demonstrate that the proposed approach yields results superior to the state of the art. Fuxiang Wu, Jun Cheng 0002, Xinchao Wang, Lei Wang 0018, Dapeng Tao |
IEEE Trans. Cybern. | 1 |
| 2021 | Gate-ID: WiFi-Based Human Identification Irrespective of Walking Directions in Smart HomeabstractResearch has shown the potential of device-free WiFi sensing for human identification. Each and every human has a unique gait and prior works suggest WiFi devices are able to capture the unique signature of a person's gait. In this article, we show for the first time that the monitored gait could be inconsistent and have mirror-like perturbations when individuals walk through WiFi devices in different directions, provided that the WiFi antenna array is horizontal to the walking path. Such inconsistent mirrored patterns are to negatively affect the uniqueness of gait and accuracy of human identification. Therefore, we propose a system called Gate-ID for accurately identifying individuals' identities irrespective of different walking directions. Gate-ID employs theoretical communication model and real measurements to demonstrate that antenna array orientations and walking directions contribute to the mirror-like patterns in WiFi signals. A novel heuristic algorithm is proposed to infer individual's walking directions. A set of methods are employed to extract and augment the representative spatial-temporal features of gait and enable the system performing irrespective of walking directions. We further propose a novel attention-based deep learning model that fuses various weighted features and ignores ineffective noises to uniquely identify individuals. We implement Gate-ID on commercial off-the-shelf devices. Extensive experiments demonstrate that our system can uniquely identify people with average accuracy of 90.7%-75.7% from a group of 6-20 people, respectively, and improve the accuracy by 12.5%-43.5% compared with baselines. Jin Zhang 0013, Bo Wei 0003, Fuxiang Wu, Limeng Dong, Wen Hu 0001, Salil S. Kanhere, Chengwen Luo 0001, Shui Yu 0001, Jun Cheng 0002 |
IEEE Internet Things J. | 3 |
| 2021 | Data Augmentation and Dense-LSTM for Human Activity Recognition Using WiFi SignalabstractRecent research has devoted significant efforts on the utilization of WiFi signals to recognize various human activities. An individual's limb motions in the WiFi coverage area could interfere with wireless signal propagation, that manifested as unique patterns for activity recognition. Existing approaches though yielding reasonable performance in certain cases, are ignorant of two major challenges. The performed activities of the individual normally have inconsistent speed in different situations and time. Besides that the wireless signal reflected by human bodies normally carries substantial information that is specific to that subject. The activity recognition model trained on a certain individual may not work well when being applied to predict another individual's activities. Since only recording activities of limited subjects in a certain speed and scale, recent works commonly have a moderate amount of activity data for training the recognition model. The small-size data could often incur the overfitting issue that negative affect the traditional classification model. To address these challenges, we propose a WiFi-based human activity recognition system that synthesizes variant activities data through eight channel state information (CSI) transformation methods to mitigate the impact of activity inconsistency and subject-specific issues, and also design a novel deep-learning model that caters to the small-size WiFi activity data. We conduct extensive experiments and show synthetic data improve performance by up to 34.6% and our system achieves around 90% of accuracy with well robustness in adapting to small-size CSI data. Jin Zhang 0013, Fuxiang Wu, Bo Wei 0003, Qieshi Zhang, Hui Huang 0014, Syed Wajid Ali Shah, Jun Cheng 0002 |
IEEE Internet Things J. | 2 |
| 2020 | RiFeGAN: Rich Feature Generation for Text-to-Image Synthesis From Prior KnowledgeabstractText-to-image synthesis is a challenging task that generates realistic images from a textual sequence, which usually contains limited information compared with the corresponding image and so is ambiguous and abstractive. The limited textual information only describes a scene partly, which will complicate the generation with complementing the other details implicitly and lead to low-quality images. To address this problem, we propose a novel rich feature generating text-to-image synthesis, called RiFeGAN, to enrich the given description. In order to provide additional visual details and avoid conflicting, RiFeGAN exploits an attention-based caption matching model to select and refine the compatible candidate captions from prior knowledge. Given enriched captions, RiFeGAN uses self-attentional embedding mixtures to extract features across them effectually and handle the diverging features further. Then it exploits multi-captions attentional generative adversarial networks to synthesize images from those features. The experiments conducted on widely-used datasets show that the models can generate images from enriched captions effectually and improve the results significantly. Jun Cheng 0002, Fuxiang Wu, Yanling Tian, Lei Wang 0018, Dapeng Tao |
CVPR | 2 |
| 2019 | WiEnhance: Towards Data Augmentation in Human Activity Recognition Using WiFi SignalabstractRecent research have devoted significant efforts on the utilization of WiFi signals to recognize various human activities. An individual's limb motions in the WiFi spectrum could interfere wireless signal propagation which manifested as unique patterns for activities recognition. Existing approaches though yielding reasonable performance in certain cases, are ignorant of a major challenge. The performed activities of the individual normally have inconsistent speed in different situations and time. Besides that the wireless signal reflected by human bodies normally carry substantial information that is specific to that subject. The activity recognition model trained on a certain individual may not work well when being applied to predict another individual's activities. To address this challenge, we propose WiEnhance, a WiFi based activity recognition system that synthesize variant activities data and mitigate the impact of activity inconsistency and subject-specific issues. We conduct extensive experiments and show an average 15.6% performance improvement on activity recognition. Jin Zhang 0013, Fuxiang Wu, Wen Hu 0001, Qieshi Zhang, Weitao Xu, Jun Cheng 0002 |
MSN | 2 |
| 2019 | Single-Image De-Raining With Feature-Supervised Generative Adversarial NetworkabstractDe-raining, which aims at rain-steak removal from images, is a practical task in computer vision. However, it is difficult due to its ill-posed nature. In this letter, we propose a deep neural network architecture, feature-supervised generative adversarial network (FS-GAN) for single-image rain removal. Its main idea is to train a generative adversarial network (GAN) for which the supervision from ground truth is imposed on different layers of the generator network. We design a feature-supervised generator, a discriminator, an optimization target, as well as the detailed structure of FS-GAN. Experiments show that the proposed FS-GAN achieves better performance than state-of-the-art de-raining methods on both synthetic and real-world images in terms of quantitative and visual quality. Lei Wang 0018, Fuxiang Wu, Jun Cheng 0002, MengChu Zhou |
IEEE Signal Process. Lett. | 3 |
| 2017 | Node-level parallelization for deep neural networks with conditional independent graph
Fugen Zhou, Fuxiang Wu, Zhengchen Zhang, Minghui Dong |
Neurocomputing | 2 |