EDBT 2026 Demo / reviewers in the wild / expert
Fangyu Wu 0001
dblp:200/7639-1 · also Fang-Yu Wu 0001
· DBLP profile ↗
32ranked-venue papers
9as first author
25since 2021 · last 2026
0000-0001-9618-8965ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 17 · 8 first-author · 10 since 2021Artificial intelligence and machine learning · 15 · 4 first-author · 11 since 2021Systems, architecture and hardware · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | OPHEP-Miner: One-phase high-efficiency pattern mining utilizing tree structures
Genlang Chen, Fangyu Wu 0001, Wanli Zuo, Youxi Wu |
Eng. Appl. Artif. Intell. | 3 |
| 2026 | EmoSENSE: Modeling Sentiment-Semantic Knowledge With Hierarchical Reinforcement Learning for Emotional Image GenerationabstractEmotional image generation aims to create images that effectively reflect target emotions. A fundamental challenge in this task is the affective gap, which refers to the discrepancy between visual content and emotional states perceived by users. Existing methods generally assume strong and explicit associations between target emotions and specific objects (e.g., “monster” and “fear”), which limits their generalization ability when encountering uncommon emotion-object pairs. This limitation stems from two main factors: 1) Most existing approaches primarily focus on semantic alignment without explicitly modeling how emotions influence visual attributes such as brightness and colorfulness; 2) diffusion-based image generation methods have limited capability in handling diverse sentiment-semantic pairs. To address these challenges, we propose EmoSENSE, a novel hierarchical fuzzy reinforcement learning framework for the emotional image generation task. EmoSENSE consists of a high-level module and a low-level module, working collaboratively in a hierarchical structure to inject sentiment-semantic knowledge into emotional images. The high-level module quantifies sentiment-semantic correlations within a unified emotional space, connecting emotions to visual attributes. The low-level module refines this connection by optimizing a fuzzy-logic-based mapping between emotions and visual attributes through reinforcement learning, enabling flexible adaptation to diverse emotion-object pairs. Extensive qualitative and quantitative experiments on public dataset demonstrate that EmoSENSE significantly enhances both the visual quality and emotional expression ability of the generated images, achieving a 12.21% higher EmoAccuracy-8 classes than the previous state-of-the-art methods.https://github.com/forever3600/EmoSENSE. Junyi Guo, Qiufeng Wang 0001, Yaran Chen, Fangyu Wu 0001, Eng Gee Lim |
IEEE Trans. Affect. Comput. | 6 |
| 2025 | Compress Time Series with Smaller Error Tolerances
Juntao Yu, Fangyu Wu 0001, Huanyu Zhao, Shiting Wen, Tongliang Li, Chaoyi Pang |
DASFAA (4) | 2 |
| 2025 | LLM-Driven Completeness and Consistency Evaluation for Cultural Heritage Data Augmentation in Cross-Modal RetrievalabstractCross-modal retrieval is essential for interpreting cultural heritage data, but its effectiveness is often limited by incomplete or inconsistent textual descriptions, caused by historical data loss and the high cost of expert annotation.While large language models (LLMs) offer a promising solution by enriching textual descriptions, their outputs frequently suffer from hallucinations or miss visually grounded details.To address these challenges, we propose C 3 , a data augmentation framework that enhances cross-modal retrieval performance by improving the completeness and consistency of LLM-generated descriptions.C 3 introduces a completeness evaluation module to assess semantic coverage using both visual cues and language-model outputs.Furthermore, to mitigate factual inconsistencies, we formulate a Markov Decision Process to supervise Chain-of-Thought reasoning, guiding consistency evaluation through adaptive query control.Experiments on the cultural heritage datasets CulTi and TimeTravel, as well as on general benchmarks MSCOCO and Flickr30K, demonstrate that C 3 achieves state-of-the-art performance in both fine-tuned and zero-shot settings.The code of this paper is available at https://github.com/JianZhang24/C-3. Jian Zhang 0002, Junyi Guo, Junyi Yuan, Huanda Lu, Fangyu Wu 0001, Dongming Lu |
EMNLP | 6 |
| 2025 | HiGarment: Cross-Modal Harmony Based Diffusion Model for Flat Sketch to Realistic Garment ImageabstractDiffusion-based garment synthesis tasks primarily focus on the design phase in the fashion domain, while the garment production process remains largely underexplored. To bridge this gap, we introduce a new task: Flat Sketch to Realistic Garment Image (FS2RG), which generates realistic garment images by integrating flat sketches and textual guidance. FS2RG presents two key challenges: 1) fabric characteristics are solely guided by textual prompts, providing insufficient visual supervision for diffusion-based models, which limits their ability to capture fine-grained fabric details; 2) flat sketches and textual guidance may provide conflicting information, requiring the model to selectively preserve or modify garment attributes while maintaining structural coherence. To tackle this task, we propose HiGarment, a novel framework that comprises two core components: i) a multi-modal semantic enhancement mechanism that enhances fabric representation across textual and visual modalities, and ii) a harmonized cross-attention mechanism that dynamically balances information from flat sketches and text prompts, allowing controllable synthesis by generating either sketch-aligned (image-biased) or text-guided (text-biased) outputs. Furthermore, we collect Multi-modal Detailed Garment, the largest open-source dataset for garment generation. Experimental results and user studies demonstrate the effectiveness of HiGarment in garment synthesis. The code and dataset are available at https://github.com/Maple498/HiGarment. Junyi Guo, Fangyu Wu 0001, Huanda Lu, Qiufeng Wang 0001, Wenmian Yang, Eng Gee Lim, Dongming Lu |
ICCV | 3 |
| 2025 | Towards Cross-Modal Retrieval in Chinese Cultural Heritage Documents: Dataset and Solution
Junyi Yuan, Jian Zhang 0002, Fangyu Wu 0001, Huanda Lu, Dongming Lu, Qiufeng Wang 0001 |
ICDAR (4) | 3 |
| 2025 | Hubness Reduction with Dual Bank Sinkhorn Normalization for Cross-Modal RetrievalabstractThe past decade has witnessed rapid advancements in cross-modal retrieval, with significant progress made in accurately measuring the similarity between cross-modal pairs. However, the persistent hubness problem, a phenomenon where a small number of targets frequently appear as nearest neighbors to numerous queries, continues to hinder the precision of similarity measurements. Despite several proposed methods to reduce hubness, their underlying mechanisms remain poorly understood. To bridge this gap, we analyze the widely-adopted Inverted Softmax approach and demonstrate its effectiveness in balancing target probabilities during retrieval. Building on these insights, we propose a probability-balancing framework for more effective hubness reduction. We contend that balancing target probabilities alone is inadequate and, therefore, extend the framework to balance both query and target probabilities by introducing Sinkhorn Normalization (SN). Notably, we extend SN to scenarios where the true query distribution is unknown, showing that current methods, which rely solely on a query bank to estimate target hubness, produce suboptimal results due to a significant distributional gap between the query bank and targets. To mitigate this issue, we introduce Dual Bank Sinkhorn Normalization (DBSN), incorporating a corresponding target bank alongside the query bank to narrow this distributional gap. Our comprehensive evaluation across various cross-modal retrieval tasks, including image-text retrieval, video-text retrieval, and audio-text retrieval, demonstrates consistent performance improvements, validating the effectiveness of both SN and DBSN. All codes are publicly available at https://github.com/ppanzx/DBSN. Zhengxin Pan, Haishuai Wang, Fangyu Wu 0001, Peng Zhang 0001, Jiajun Bu |
ACM Multimedia | 3 |
| 2025 | EditGarment: An Instruction-Based Garment Editing Dataset Constructed with Automated MLLM Synthesis and Semantic-Aware EvaluationabstractInstruction-based garment editing enables precise image modifications via natural language, with broad applications in fashion design and customization. Unlike general editing tasks, it requires understanding garment-specific semantics and attribute dependencies. However, progress is limited by the scarcity of high-quality instruction-image pairs, as manual annotation is costly and hard to scale. While MLLMs have shown promise in automated data synthesis, their application to garment editing is constrained by imprecise instruction modeling and a lack of fashion-specific supervisory signals. To address these challenges, we present an automated pipeline for constructing a garment editing dataset. We first define six editing instruction categories aligned with real-world fashion workflows to guide the generation of balanced and diverse instruction-image triplets. Second, we introduce Fashion Edit Score, a semantic-aware evaluation metric that captures semantic dependencies between garment attributes and provides reliable supervision during construction. Using this pipeline, we construct a total of 52,257 candidate triplets and retain 20,596 high-quality triplets to build EditGarment, the first instruction-based dataset tailored to standalone garment editing. The project page is https://yindq99.github.io/EditGarment-project/. Deqiang Yin, Junyi Guo, Huanda Lu, Fangyu Wu 0001, Dongming Lu |
ACM Multimedia | 4 |
| 2025 | CostDiff: Residual Diffusion-Based Cost Map Refinement for Open-Vocabulary Semantic Segmentation
Yutao Rao, Fangyu Wu 0001, Junjie Zhang 0002 |
PRCV (12) | 3 |
| 2025 | WIC: Hiding Producer-Consumer Synchronization Delays with Warp-Level Interrupt-based GPU Communications
Jiajian Zhang, Fangyu Wu 0001, Hai Jiang 0003, Qiufeng Wang 0001, Genlang Chen, Chaoyi Pang |
USENIX ATC | 2 |
| 2025 | Fine-grained visual tracking via distribution-aware mask modeling and temporal propagation
Junjie Zhang 0002, Hongwen Yu, Fangyu Wu 0001, Xiaoshui Huang, Jian Zhang 0002 |
Knowl. Based Syst. | 4 |
| 2025 | Establishing Nuanced Multimodal Attention for Weakly Supervised Semantic Segmentation of Remote Sensing ScenesabstractWeakly Supervised Semantic Segmentation (WSSS) with image-level labels reduces reliance on pixel-level annotations for remote sensing (RS) imagery. However, in natural scenes, WSSS frequently faces challenges such as imprecise localization, extraneous activations, and class ambiguity. These challenges are particularly pronounced in RS images, characterized by complex backgrounds, substantial scale variations, and dense small-object distributions, complicating the distinction between intra-class variations and inter-class similarities. To tackle these challenges, we introduce a class-constrained multi-modal attention framework aimed at enhancing the localization accuracy of class activation maps (CAMs). Specifically, we design class-specific tokens to capture the visual characteristics of each target class. As these tokens initially lack explicit constraints, we integrate the textual branch of the RemoteCLIP model to leverage class-related linguistic priors, which collaborate with visual features to encode the specific semantics of diverse objects. Furthermore, the multi-modal collaborative optimization module dynamically establishes tailored attention mechanisms for both global and regional features, thereby improving class discriminability among targets to mitigate challenges like inter-class similarity and dense small-object distributions. By refining class-specific attention, textual semantic attention, and patch-level pairwise affinity weights, the quality of generated pseudo-masks is markedly enhanced. Concurrently, to ensure domain-invariant feature learning, we align the backbone features with the CLIP visual embedding by minimizing the distribution disparity between the two in the latent space, semantic consistency is therefore preserved. The experimental results validate the effectiveness and robustness of our proposed method, achieving significant performance improvements on two representative RS WSSS datasets. Junjie Zhang 0002, Huaxi Huang, Fangyu Wu 0001, Hongwen Yu |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2025 | AlignMalloc: Warp-Aware Memory Rearrangement Aligned With UVM Prefetching for Large-Scale GPU Dynamic AllocationsabstractAs parallel computing tasks rapidly expand in both complexity and scale, the need for efficient GPU dynamic memory allocation becomes increasingly important. While progress has been made in developing dynamic allocators for substantial applications, their real-world applicability is still limited due to inefficient memory access behaviors. This paper introduces AlignMalloc, a novel memory management system that aligns with the Unified Virtual Memory (UVM) prefetching strategy, significantly enhancing both memory allocation and access performance in large-scale dynamic allocation scenarios. We analyze the fundamental inefficiencies in UVM access and first reveal the mismatch between memory access and UVM prefetching methods. To resolve this issue, AlignMalloc implements a warp-aware memory rearrangement strategy that exploits the regularity of warps to align with the UVM's static prefetching setup. Additionally, AlignMalloc introduces an OR tree-based structure within a host-co-managed framework to further optimize dynamic allocation. Comprehensive experiments demonstrate that AlignMalloc substantially outperforms current state-of-the-art systems, achieving up to$2.7 \times$improvement in dynamic allocation and$2.3 \times$in memory access. Additionally, eight real-world applications with diverse memory access patterns exhibit consistent performance enhancements, with average speedups$1.5 \times$. Jiajian Zhang, Fangyu Wu 0001, Hai Jiang 0003, Qiufeng Wang 0001, Genlang Chen, Eng Gee Lim, Keqin Li 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2025 | Boosting remote semantic segmentation using vision-and-language foundation model
Qiuyue Zhang, Zhiwang Zhang, Shiting Wen, Chaoyi Pang, Fangyu Wu 0001 |
Vis. Comput. | 5 |
| 2024 | Two-branch Network with Feature Fusion for Time Since Deposition Estimation of BloodstainsabstractIn bloodstain examination of collaborative medicine and forensics, the analysis and identification of time since deposition (TSD) plays a significant role. Traditional bloodstain analysis methods can only provide a rough estimate for the TSD of traces, and they are time-consuming. To address this issue, we propose a lightweight framework called Fourier Transform Infrared Network (FTIR-Net) that combines wavelet transform with deep learning. To be specific, we parallelly perform wavelet transform on infrared spectra and compute its second derivative to attain the sequential signal and spectral image. Then, the learning component employs two separate branches to extract features from the one-dimensional (1D) spectra signal and two-dimensional (2D) coefficient images provided by continuous wavelet transform (CWT). To effectively aggregate information from the spectral image, we design a Squeeze-and-Excitation Network (SENet) and combine it with 2D convolution. Finally, the extracted features are concatenated and flattened, followed by two fully connected (FC) layers for retention time analysis. Since the standard bloodstain dataset is lacking, we create a dataset that associates bloodstain with the attenuated total reflectance of Fourier transform infrared (ATR-FTIR). To demonstrate the effectiveness of our model in bloodstain analysis and exploit the properties of the proposed dataset, we present comprehensive experiments and ablation studies. Yushi Li, Yu Han 0001, Jia Wang 0009, Fangyu Wu 0001, Chenke Yin |
CSCWD | 5 |
| 2024 | SyncMalloc: A Synchronized Host-Device Co-Management System for GPU Dynamic Memory Allocation across All ScalesabstractDynamic memory allocation on GPUs, increasingly crucial for applications with dynamic computational patterns, encounters significant challenges due to the complex calculations with intricate branches and substantial memory resources consumed by metadata from massive thread allocations. Despite the current research, there is a lack of a scalable and flexible solution that effectively manages dynamic memory allocation while minimizing memory usage on GPUs. This paper introduces SyncMalloc, a synchronized Host-Device Co-Management system that is specifically designed to adeptly handle dynamic memory allocations of diverse magnitudes. Through the integration of pipelining and producer-consumer mechanisms, SyncMalloc effectively reduces communication overhead and resolves architectural mismatches, further enhancing its capability through synergistic integration with CUDA’s unified memory to facilitate oversubscription. Moreover, SyncMalloc advances slab-based memory management to enhance the efficiency of small allocations, reducing conflict probabilities and overhead in high-activity scenarios. Finally, we present a comprehensive performance evaluation, expanding benchmarks and measurement dimensions to reflect the performance of real-world applications more accurately. The experimental results demonstrate the effectiveness of SyncMalloc in supporting dynamic GPU allocations scaled from 4B to 200GB from multiple perspectives. Our source code is available at https://github.com/jjZhang94/SyncMalloc. Jiajian Zhang, Fangyu Wu 0001, Hai Jiang 0003, Genlang Chen, Qiufeng Wang 0001 |
ICPP | 2 |
| 2024 | Multi-task Prompt Words Learning for Social Media Content GenerationabstractThe rapid development of the Internet has profoundly changed human life. Humans are increasingly expressing themselves and interacting with others on social media platforms. However, although artificial intelligence technology has been widely used in many aspects of life, its application in social media content creation is still blank. To solve this problem, we propose a new prompt word generation framework based on multi-modal information fusion, which combines multiple tasks including topic classification, sentiment analysis, scene recognition and keyword extraction to generate more comprehensive prompt words. Subsequently, we use a template containing a set of prompt words to guide ChatGPT to generate high-quality tweets. Furthermore, in the absence of effective and objective evaluation criteria in the field of content generation, we use the ChatGPT tool to evaluate the results generated by the algorithm, making large-scale evaluation of content generation algorithms possible. Evaluation results on extensive content generation demonstrate that our cue word generation framework generates higher quality content compared to manual methods and other cueing techniques, while topic classification, sentiment analysis, and scene recognition significantly enhance content clarity and its consistency with the image. Haochen Xue, Chong Zhang 0006, Chenzhi Liu, Fangyu Wu 0001, Xiao-Bo Jin |
IJCNN | 4 |
| 2024 | Discriminative Feature Enhancement Network for few-shot classification and beyond
Fangyu Wu 0001, Qiufeng Wang 0001, Qi Chen 0026, Eng Gee Lim |
Expert Syst. Appl. | 1 |
| 2024 | Correction: ITContrast: contrastive learning with hard negative synthesis for image-text matching
Fangyu Wu 0001, Qiufeng Wang 0001, Zhao Wang 0001, Siyue Yu, Yushi Li, Eng Gee Lim |
Vis. Comput. | 1 |
| 2024 | ITContrast: contrastive learning with hard negative synthesis for image-text matching
Fangyu Wu 0001, Qiufeng Wang 0001, Zhao Wang 0001, Siyue Yu, Yushi Li, Eng Gee Lim |
Vis. Comput. | 1 |
| 2023 | Fine-grained Image-text Matching by Cross-modal Hard Aligning NetworkabstractCurrent state-of-the-art image-text matching methods implicitly align the visual-semantic fragments, like regions in images and words in sentences, and adopt cross-attention mechanism to discover fine-grained cross-modal semantic correspondence. However, the cross-attention mechanism may bring redundant or irrelevant region-word alignments, degenerating retrieval accuracy and limiting efficiency. Although many researchers have made progress in mining meaningful alignments and thus improving accuracy, the problem of poor efficiency remains unresolved. In this work, we propose to learn fine-grained image-text matching from the perspective of information coding. Specifically, we suggest a coding framework to explain the fragments aligning process, which provides a novel view to reexamine the cross-attention mechanism and analyze the problem of redundant alignments. Based on this framework, a Cross-modal Hard Aligning Network (CHAN) is designed, which comprehensively exploits the most relevant region-word pairs and eliminates all other alignments. Extensive experiments conducted on two public datasets, MS-COCO and Flickr30K, verify that the relevance of the most associated word-region pairs is discriminative enough as an indicator of the image-text similarity, with superior accuracy and efficiency over the state-of-the-art approaches on the bidirectional image and text retrieval tasks. Our code will be available at https://github.com/ppanzx/CHAN. Zhengxin Pan, Fangyu Wu 0001 |
CVPR | 2 |
| 2023 | Knowledge-embedded Prompt Learning for Zero-shot Social Media Text ClassificationabstractSocial media plays an irreplaceable role in shaping the way information is created shared and consumed. While it provides access to a vast amount of data, extracting and analyzing useful insights from complex and dynamic social media data can be challenging. Deep learning models have shown promise in social media analysis tasks, but such models require a massive amount of labelled data which is usually unavailable in real-world settings. Additionally, these models lack common-sense knowledge which can limit their ability to generate comprehensive results. To address these challenges, we propose a knowledge-embedded prompt learning model for zero-shot social media text classification. Our experimental results on four social media datasets demonstrate that our proposed approach outperforms other well-known baselines. Qi Chen 0026, Wei Wang 0042, Fangyu Wu 0001 |
SMARTCOMP | 4 |
| 2023 | Mining top-k high average-utility itemsets based on breadth-first search
Genlang Chen, Fangyu Wu 0001, Shiting Wen, Wanli Zuo |
Appl. Intell. | 3 |
| 2022 | Kernel triplet loss for image-text retrievalabstractAbstract Triplet loss is widely used as the objective function in image‐text retrieval tasks. However, as all the triplets are treated equally, triplet loss has a bottleneck problem of slow convergence and other unsatisfactory performances. In this article, we propose solutions by appropriately weighting triplets according to the relative similarities among the training samples. Specifically, we present three weighting functions to assign an appropriate weight for the selected informative triplets to accelerate the convergence. We evaluate our approach on two widely used benchmark datasets: Flickr30k and MSCOCO, with results outperforming the previous methods, which demonstrates its superiority. Zhengxin Pan, Fangyu Wu 0001 |
Comput. Animat. Virtual Worlds | 2 |
| 2021 | FaceCaps for facial expression recognitionabstractAbstract Facial expression recognition (FER) is a significant research task in the computer vision field. In this paper, we present a novel network FaceCaps for facial expression recognition with the following novel characteristics: an embedding structure based on a Capsule network which encodes relative spatial relationships between features; incorporates the feature polymerization property of FaceNet, thus offering a more efficient approach to discriminate complex facial expressions; a target reconstruction loss as a better regularization term for Capsule networks. Experimental results on both lab‐controlled datasets (CK+) and real‐world databases (RAF‐DB and SFEW 2.0) demonstrate that the method significantly outperforms the state‐of‐the‐art. Fangyu Wu 0001, Chaoyi Pang |
Comput. Animat. Virtual Worlds | 1 |
| 2020 | Attentive Prototype Few-Shot Learning with Capsule Network-Based Embedding
Fangyu Wu 0001, Jeremy S. Smith, Wenjin Lu, Chaoyi Pang |
ECCV (28) | 1 |
| 2020 | Pose-robust Face Recognition by Deep Meta Capsule network-based Equivariant EmbeddingabstractDespite the exceptional success in face recognition related technologies, handling large pose variations still remains a key challenge. Current techniques for pose-robust face recognition either, directly extract pose-invariant features, or first synthesize a face that matches the target pose before feature extraction. It is more desirable to learn face representations equivariant to pose variations. To this end, this paper proposes a deep meta Capsule network-based Equivariant Embedding Model (DM-CEEM) with three distinct novelties. First, the proposed RB-CapsNet allows DM-CEEM to learn an equivariant embedding for pose variations and achieve the desired transformation for input face images. Second, we introduce a new version of a Capsule network called RB-CapsNet to extend CapsNet to perform a profile-to-frontal face transformation in deep feature space. Third, we train the DM-CEEM in a meta way by treating a single overall classification target as multiple sub-tasks that satisfy certain unknown probabilities. In each sub-task, we sample the support and query sets randomly. The experimental results on both controlled and in-the-wild databases demonstrate the superiority of DM-CEEM over state-of-the-art. Fangyu Wu 0001, Jeremy S. Smith, Wenjin Lu |
ICPR | 1 |
| 2020 | Image captioning via hierarchical attention mechanism and policy gradient optimization
Shiyang Yan, Yuan Xie 0006, Fangyu Wu 0001, Jeremy S. Smith, Wenjin Lu |
Signal Process. | 3 |
| 2019 | Image-Image Translation to Enhance Near Infrared Face RecognitionabstractWith the rapid development of facial recognition, the research field of near infrared (NIR) face recognition, which is less sensitive to illumination levels, has attracted increased attention. Unfortunately, directly applying the face recognition model trained using visible light (VIS) data to NIR face data does not produce a satisfactory performance. This is due to the domain bias between the NIR images and the VIS images. To this end, we created the Outdoor NIR-VIS Face (ONVF) database and Indoor NIR Face (INF) database to increase the number of near infrared facial images for system training and evaluation. In this paper, we propose an efficient NIR face recognition method, which consists of face detection and alignment, NIR-VIS image translation and face embedding. The NIR-VIS image conversion model is capable of transforming near-infrared facial images into their corresponding VIS images whilst maintaining sufficient identity information to enable existing VIS facial recognition models to perform recognition. Extensive experiments using the INF dataset and the CSIST database have demonstrated that the proposed method yields a consistent and competitive performance for near infrared face recognition. Fangyu Wu 0001, Weihang You, Jeremy S. Smith, Wenjin Lu |
ICIP | 1 |
| 2019 | Vehicle re-identification in still images: Application of semi-supervised learning and re-ranking
Fangyu Wu 0001, Shiyang Yan, Jeremy S. Smith |
Signal Process. Image Commun. | 1 |
| 2018 | Joint Semi-supervised Learning and Re-ranking for Vehicle Re-identificationabstractVehicle re-identification (re-ID) remains an unproblematic problem due to the complicated variations in vehicle appearances from multiple camera views. Most existing algorithms for solving this problem are developed in the fully-supervised setting, requiring access to a large number of labeled training data. However, it is impractical to expect large quantities of labeled data because the high cost of data annotation. Besides, re-ranking is a significant way to improve its performance when considering vehicle re-ID as a retrieval process. Yet limited effort has been devoted to the research of re-ranking in the vehicle re-ID. To address these problems, in this paper, we propose a semi-supervised learning system based on the Convolutional Neural Network (CNN) and re-ranking strategy for Vehicle re-ID. Specifically, we adopt the structure of Generative Adversarial Network (GAN) to obtain more vehicle images and enrich the training set, then a uniform label distribution will be assigned to the unlabeled samples according to the Label Smoothing Regularization for Outliers (LSRO), which regularizes the supervised learning model and improves the performance of re-ID. To optimize the re-ID results, an improved re-ranking method is exploited to optimize the initial rank list. Experimental results on publically available datasets, VeRi-776 and VehicleID, demonstrate that the method significantly outperforms the state-of-the-art. Fangyu Wu 0001, Shiyang Yan, Jeremy S. Smith |
ICPR | 1 |
| 2018 | Image Captioning using Adversarial Networks and Reinforcement LearningabstractImage captioning is a significant task in artificial intelligence which connects computer vision and natural language processing. With the rapid development of deep learning, the sequence to sequence model with attention, has become one of the main approaches for the task of image captioning. Nevertheless, a significant issue exists in the current framework: the exposure bias problem of Maximum Likelihood Estimation (MLE) in the sequence model. To address this problem, we use generative adversarial networks (GANs) for image captioning, which compensates for the exposure bias problem of MLE and also can generate more realistic captions. GANs, however, cannot be directly applied to a discrete task, like language processing, due to the discontinuity of the data. Hence, we use a reinforcement learning (RL) technique to estimate the gradients for the network. Also, to obtain the intermediate rewards during the process of language generation, a Monte Carlo roll-out sampling method is utilized. Experimental results on the COCO dataset validate the improved effect from each ingredient of the proposed model. The overall effectiveness is also evaluated. Shiyang Yan, Fangyu Wu 0001, Jeremy S. Smith, Wenjin Lu |
ICPR | 2 |