VLDB 2026 Research / reviewers in the wild / expert
Xiaoling Gu
dblp:81/4398
· DBLP profile ↗
32ranked-venue papers
13as first author
24since 2021 · last 2026
0000-0002-2876-1771ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 20 · 9 first-author · 15 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 9 since 2021Databases, data management, data science and information retrieval · 4 · 4 first-authorComputer networks · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MCIE: Multimodal LLM-Driven Complex Instruction Image Editing with Spatial GuidanceabstractRecent advances in instruction-based image editing have shown remarkable progress. However, existing methods remain limited to relatively simple editing operations, hindering real-world applications that require complex and compositional instructions. In this work, we address these limitations from the perspectives of architectural design, data, and evaluation protocols. Specifically, we identify two key challenges in current models: insufficient instruction compliance and background inconsistency. To this end, we propose MCIE-E1, a Multimodal Large Language Model–Driven Complex Instruction Image Editing method that integrates two key modules: a spatial-aware cross-attention module and a background-consistent cross-attention module. The former enhances instruction-following capability by explicitly aligning semantic instructions with spatial regions through spatial guidance during the denoising process, while the latter preserves features in unedited regions to maintain background consistency. To enable effective training, we construct a dedicated data pipeline to mitigate the scarcity of complex instruction-based image editing datasets, combining fine-grained automatic filtering via a powerful MLLM with rigorous human validation. Finally, to comprehensively evaluate complex instruction-based image editing, we introduce CIE-Bench, a new benchmark with two new evaluation metrics. Experimental results on CIE-Bench demonstrate that MCIE-E1 consistently outperforms previous state-of-theart methods in both quantitative and qualitative assessments, achieving a 23.96% improvement in instruction compliance. Xuehai Bai, Xiaoling Gu, Akide Liu, Hangjie Yuan, YiFan Zhang, Jack Ma |
AAAI | 2 |
| 2026 | HybridKV: Hybrid KV Cache Compression for Efficient Multimodal Large Language Model InferenceabstractMultimodal Large Language Models (MLLMs) have advanced unified reasoning over text, images, and videos, but their inference is hindered by the rapid growth of key-value (KV) caches.Each visual input expands into thousands of tokens, causing caches to scale linearly with context length and remain resident in GPU memory throughout decoding, which leads to prohibitive memory overhead and latency even on high-end GPUs.A common solution is to compress caches under a fixed allocated budget at different granularities: tokenlevel uniformly discards less important tokens, layer-level varies retention across layers, and head-level redistributes budgets across heads.Yet these approaches stop at allocation and overlook the heterogeneous behaviors of attention heads that require distinct compression strategies.We propose HYBRIDKV, a hybrid KV cache compression framework that integrates complementary strategies in three stages: heads are first classified into static or dynamic types using text-centric attention; then a top-down budget allocation scheme hierarchically assigns KV budgets; finally, static heads are compressed by text-prior pruning and dynamic heads by chunk-wise retrieval.Experiments on 11 multimodal benchmarks with Qwen2.5-VL-7Bshow that HYBRIDKV reduces KV cache memory by up to 7.9× and achieves 1.52× faster decoding, with almost no performance drop or even higher relative to the full-cache MLLM. Feiyang Ren, Jun Zhang 0069, Xiaoling Gu, Ke Chen 0005, Lidan Shou, Huan Li 0003 |
ACL (1) | 4 |
| 2026 | BoostPoint: Boosting Point Cloud Backbones with Image Pre-Training for 3D UnderstandingabstractNowadays, pre-training models on largescale datasets and fine-tuning models on task-specific datasets have become common paradigms, achieving impressive success in natural language processing and 2D vision. Nonetheless, the potential of this paradigm has not been fully explored in 3D vision due to the scale of the datasets. To overcome this, we propose BoostPoint, a novel pipeline that uses large-scale rendered images as 3D point cloud model inputs for pretraining and uses general 3D tasks for fine-tuning. In BoostPoint, we propose a novel learning-free image-topoint (I2P) module to transform raw pixels into required inputs. Specifically, we view pixels as unorganized points, including essential raw features (e.g., color) and positional information (e.g., coordinates). Employing simple linear iterative clustering (SLIC), the I2P module effectively groups these unorganized points into superpixels, facilitating point cloud backbone pretraining. Furthermore, we employ a modality-agnostic debiasing mechanism during pre-training to prevent negative transfer in downstream tasks. Extensive finetuning experiments show that BoostPoint provides significant improvements to 3D point cloud backbones for 3D point cloud classification and part segmentation. Honggu Zhou, Yakai Zhang, Haohan Li, Xiaoling Gu, Ming Zeng 0008, Zizhao Wu |
Comput. Vis. Media | 4 |
| 2026 | GC-GS: Gradient control Gaussian splatting with various image degradation
Qida Cao, Jiajun Ding, Qingyuan Tang, Tianning Zhao, Xiaoling Gu, Jianping Fan 0001, Zhou Yu 0001 |
Pattern Recognit. | 5 |
| 2026 | Multi-modal retrieval augmented text-to-3D generation with accelerated sampling
Gongyi Chen, Yakai Zhang, Genfu Yang, Xiangquan Zhang, Xiaoling Gu, Zizhao Wu |
Pattern Recognit. Lett. | 6 |
| 2026 | TailorEdit: An Adaptive Framework for Instruction-Guided Fashion Image EditingabstractFashion image editing has garnered significant attention due to its growing demand in e-commerce, social media, and virtual try-on applications. However, existing methods are typically designed for specific editing tasks in isolation, lacking a unified framework capable of handling diverse editing requirements. This work addresses this limitation from two critical perspectives. First, we constructInstructFashion, a large-scale, high-quality dataset specifically curated for instruction-guided fashion image editing. It is generated through carefully designed pipelines that cover four distinct editing tasks. Second, we proposeTailorEdit, an adaptive framework for instruction-guided fashion image editing. It integrates human segmentation map-based denoising guidance, modular LoRA-based editing experts, and a dynamic expert routing mechanism to enable precise and semantically coherent modifications. Extensive quantitative and qualitative evaluations demonstrate that TailorEdit consistently outperforms state-of-the-art methods in terms of realism, coherence, and instruction adherence. Our code is available at https://github.com/EndaJude/TailorEdit. Xiaoling Gu, Lingda Zhu, Yongkang Wong, Zhou Yu 0001, Huan Li 0003, Zizhao Wu, Mohan Kankanhalli |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2026 | Compositional Text-to-Image Synthesis With Training-Free Layout-Guided DiffusionabstractRecent text-to-image (T2I) diffusion models have made significant strides in generating high-quality images from diverse textual prompts. Despite this progress, these models often face challenges in accurately understanding and synthesizing complex prompts, primarily due to their limited compositional capabilities. In this study, we propose a novel approach for compositional T2I synthesis using layout-guided diffusion models, which do not require additional training. Specifically, we leverage the chain-of-code prompting technique of large language models to interpret textual prompts and generate object layouts with spatial coherence. To enhance the alignment between generated images and textual descriptions, we introduce two innovative layoutguided loss functions: Patch-oriented Cross-Attention (PCA) loss and Region-oriented Cross-Attention (RCA) loss. The PCA loss emphasizes high activation values for image patches that attend to all tokens in the prompt across the layout. The RCA loss enhances the average attention within the layout, thereby increasing the accuracy of generating objects and their associated attributes within specified regions. These proposed loss functions reassign cross-attention in diffusion models during the denoising process. Our comprehensive experiments consistently demonstrate the effectiveness of our approach in improving semantic alignment between generated images and a diverse range of textual prompts, while ensuring high usability as a ready-to-use plugin. Our code is available athttps://github.com/gxl-groups/Compositional-T2I. Xiaoling Gu, Lingwei Luo, Shengqi Wu, Zizhao Wu, Zhenzhong Kuang, Zhou Yu 0001 |
IEEE Trans. Multim. | 1 |
| 2026 | InterMamba: Efficient Human-Human Interaction Generation With Adaptive Spatio-Temporal MambaabstractHuman-human interaction generation has garnered significant attention in motion synthesis due to its vital role in understanding humans as social beings. However, existing methods typically rely on transformer-based architectures, which often face challenges related to scalability and efficiency. To address these challenges, we propose InterMamba, a novel and efficient human-human interaction generation method built on the Mamba framework, designed to capture long-sequence dependencies effectively while enabling real-time feedback. Specifically, we introduce an adaptive spatio-temporal Mamba framework that utilizes two parallel SSM branches with an adaptive mechanism to integrate the spatial and temporal features of motion sequences. To further enhance the model's ability to capture dependencies within individual motion sequences and the interactions between different individual sequences, we develop two key modules: the self adaptive spatio-temporal Mamba module and the cross adaptive spatio-temporal Mamba module, enabling efficient feature learning. Extensive experiments demonstrate that our method achieves the state-of-the-art results on both two interaction datasets with remarkable quality and efficiency. Compared to the baseline method InterGen, our approach not only improves accuracy but also reduces the parameter size to just 66 M (36% of InterGen's), while achieving an average inference speed of 0.57 seconds, which is 46% of InterGen's execution time. Zizhao Wu, Xiaoling Gu, Ruyu Liu, Jiazhou Chen 0002 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2025 | Gaussian Splatting-based Scene Reconstruction with Adaptive Sampling and Region-based RenderingabstractRecently, 3D Gaussian Splatting (3DGS) based methods, such as Scaffold-GS, have exhibited state-of-the-art performance for high-fidelity scene reconstruction. However, existing methods may suffer from degraded details because they usually pay the same attention to the whole image set and their contents. In fact, some parts of the image data may contain more information than others. Thus, it is reasonable to treat them differently. Motivated by this, in this paper, we propose a new 3DGS-based method to adaptively choose the data that are more valuable for training. Our method consists of two key parts: Adaptive Sampling (AS) and Region-based Rendering (RR). AS focuses on collecting difficult data for 3DGS training. RR focuses on enabling training with partial image data. Our method can produce more fine-grained details for scene reconstruction and has good scalability. On multiple public datasets, we verify the effectiveness of the proposed method by conducting comparative and ablation experiments. The quantitative and qualitative experimental results show that our method has achieved state-of-the-art results compared with the baselines. Takahiko Furuya, Zhenzhong Kuang, Xiaoling Gu, Jiajun Ding |
IJCNN | 4 |
| 2025 | CoIDO: Efficient Data Selection for Visual Instruction Tuning via Coupled Importance-Diversity OptimizationabstractMultimodal large language models (MLLMs) rely heavily on instruction tuning to align vision and language capabilities, yet the computational cost of training on large-scale datasets remains a major bottleneck. Existing data selection methods aim to mitigate this by selecting important and diverse subsets, but they often suffer from two critical drawbacks: high computational overhead from processing the entire dataset and suboptimal data selection due to separate treatment of importance and diversity.
We introduce CoIDO, a novel dual-objective framework that jointly optimizes data importance and diversity to overcome these challenges. Unlike existing approaches that require costly evaluations across the whole dataset, CoIDO employs a lightweight plug-in scorer. This scorer is trained on just a small random sample of data to learn the distribution of the candidate set, drastically reducing computational demands. By leveraging a homoscedastic uncertainty-based formulation, CoIDO effectively balances importance and diversity during training, enabling the scorer to assign CoIDO scores to all data points. This unified scoring approach allows for direct ranking and selection of the most valuable subsets, completely bypassing the need for specialized algorithms.
In our experiments, we trained the CoIDO Scorer using only 20% of randomly sampled data. Once trained, CoIDO was applied to the entire dataset to select a 20% subset for instruction tuning. On the widely used LLaVA-1.5-7B model across ten downstream tasks, this selected subset achieved an impressive 98.2% of the performance of full-data fine-tuning, on average. Moreover, CoIDO outperforms all competitors in terms of both efficiency (lowest training FLOPs) and aggregated accuracy. Our code is available at: https://github.com/SuDIS-ZJU/CoIDO Xiaoling Gu, Jinpeng Chen 0001, Huan Li 0003 |
NeurIPS | 4 |
| 2025 | PAD: Detail-Preserving Point Cloud Reconstruction and Generation via AutodecodersabstractABSTRACT High‐accuracy point cloud (self‐) reconstruction is crucial for point cloud editing, translation, and unsupervised representation learning. However, existing point cloud reconstruction methods often sacrifice many geometric details. Altough many techniques have proposed how to construct better point cloud decoders, only a few have designed point cloud encoders from a reconstruction perspective. We propose an autodecoder architecture to achieve detail‐preserving point cloud reconstruction while bypassing the performance bottleneck of the encoder. Our architecture is theoretically applicable to any existing point cloud decoder. For training, both the weights of the decoder and the pre‐initialised latent codes, corresponding to the input points, are updated simultaneously. Experimental results demonstrate that our autodecoder achieves an average reduction of 24.62% in Chamfer Distance compared to existing methods, significantly improving reconstruction quality on the ShapeNet dataset. Furthermore, we verify the effectiveness of our autodecoder in point cloud generation, upsampling, and unsupervised representation learning to demonstrate its performance on downstream tasks, which is comparable to the state‐of‐the‐art methods. We will make our code publicly available after peer review. Yakai Zhang, Zizhao Wu, Xiaoling Gu, Alexandru C. Telea, Jirí Kosinka |
IET Comput. Vis. | 5 |
| 2025 | Interpretable procedural material graph generation via diffusion models from reference images
Xiaoyu Lv, Zizhao Wu, Jiamin Xu, Xiaoling Gu, Ming Zeng 0008, Weiwei Xu 0003 |
Vis. Comput. | 4 |
| 2025 | Innovative AI techniques for photorealistic 3D clothed human reconstruction from monocular images or videos: a survey
Xiaoling Gu, Zhenzhong Kuang, Fei-wei Qin, Zizhao Wu |
Vis. Comput. | 2 |
| 2025 | Slot-VTON: subject-driven diffusion-based virtual try-on with slot attention
Jianglei Ye, Yigang Wang, Fengmao Xie, Xiaoling Gu, Zizhao Wu |
Vis. Comput. | 5 |
| 2024 | 3D Question Answering with Scene Graph Reasoning
Zizhao Wu, Haohan Li, Gongyi Chen, Zhou Yu 0001, Xiaoling Gu, Yigang Wang |
ACM Multimedia | 5 |
| 2024 | Multi2Human: Controllable human image generation with multimodal controls
Xiaoling Gu, Shengwenzhuo Xu, Yongkang Wong, Zizhao Wu, Jun Yu 0002, Jianping Fan 0001, Mohan Kankanhalli |
Neurocomputing | 1 |
| 2024 | iDesigner: making intelligent fashion designs
Xiaoling Gu, Qiming Yao, Xiaojun Gong, Zhenzhong Kuang |
Multim. Tools Appl. | 1 |
| 2024 | Semantic-aware hyper-space deformable neural radiance fields for facial avatar reconstruction
Kaixin Jin, Xiaoling Gu, Zhenzhong Kuang, Zizhao Wu, Min Tan 0005, Jun Yu 0002 |
Pattern Recognit. Lett. | 2 |
| 2024 | PAINT: Photo-realistic Fashion Design SynthesisabstractIn this article, we investigate a new problem of generating a variety of multi-view fashion designs conditioned on a human pose and texture examples of arbitrary sizes, which can replace the repetitive and low-level design work for fashion designers. To solve this challenging multi-modal image translation problem, we propose a novel Photo-reAlistic fashIon desigN synThesis (PAINT) framework, which decomposes the framework into three manageable stages. In the first stage, we employ a Layout Generative Network (LGN) to transform an input human pose into a series of person semantic layouts. In the second stage, we propose a Texture Synthesis Network (TSN) to synthesize textures on all transformed semantic layouts. Specifically, we design a novel attentive texture transfer mechanism for precisely expanding texture patches to the irregular clothing regions of the target fashion designs. In the third stage, we leverage an Appearance Flow Network (AFN) to generate the fashion design images of other viewpoints from a single-view observation by learning 2D multi-scale appearance flow fields. Experimental results demonstrate that our method is capable of generating diverse photo-realistic multi-view fashion design images with fine-grained appearance details conditioned on the provided multiple inputs. The source code and trained models are available at https://github.com/gxl-groups/PAINT . Xiaoling Gu, Jie Huang 0033, Yongkang Wong, Jun Yu 0002, Jianping Fan 0001, Mohan Kankanhalli |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2024 | Recurrent Appearance Flow for Occlusion-Free Virtual Try-OnabstractImage-based virtual try-on aims at transferring a target in-shop garment onto a reference person, and has garnered significant attention from the research communities recently. However, previous methods have faced severe challenges in handling occlusion problems. To address this limitation, we classify occlusion problems into three types based on the reference person’s arm postures: single-arm occlusion , two-arm non-crossed occlusion , and two-arm crossed occlusion . Specifically, we propose a novel Occlusion-Free Virtual Try-On Network (OF-VTON) that effectively overcomes these occlusion challenges. The OF-VTON framework consists of two core components: (i) a new Recurrent Appearance Flow based Deformation (RAFD) model that robustly aligns the in-shop garment to the reference person by adopting a multi-task learning strategy . This model jointly produces the dense appearance flow to warp the garment and predicts a human segmentation map to provide semantic guidance for the subsequent image synthesis model. (ii) a powerful Multi-mask Image SynthesiS (MISS) model that generates photo-realistic try-on results by introducing a new mask generation and selection mechanism . Experimental results demonstrate that our proposed OF-VTON significantly outperforms existing state-of-the-art methods by mitigating the impact of occlusion problems. Our code is available at https://github.com/gxl-groups/OF-VTON . Xiaoling Gu, Junkai Zhu, Yongkang Wong, Zizhao Wu, Jun Yu 0002, Jianping Fan 0001, Mohan Kankanhalli |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2023 | Electromagnetic Imaging Boosted Visual Object Recognition Under Difficult Visual ConditionsabstractObject imaging and recognition under difficult visual conditions is extremely challenging due to the captured low-quality images, and traditional optical-based recognition methods always fail in this task. In this paper, we propose to utilize the visual-microwave image pairs captured by both visual cameras and microwave sensors for imaging and recognition. To address the heavy noises in the low-quality optical images, we retrieve the physically quantitative images from associated scattered field data, and enhance visual features by both optical and retrieval images. We develop a cross-modal Enhanced Attentive Visual-Microwave Fusion (EAVMF) object recognition model to jointly learn the cross-modal generator and multimodal recognizer. In addition, an attention module for the visual subnetwork is utilized to highlight the regions of interest. Two multimodal datasets with synthetic visual-microwave image pairs are built to simulate the difficult visual condition. The numerical results on these datasets demonstrate that: 1) both the multimodal fusion, cross-modal enhancement, and visual attention module can enhance the performance; and 2) compared with existing methods, the proposed EAVMF not only performs better in terms of accuracy but also has good scalability and one-shot learning ability. Min Tan 0005, Tao Jin 0004, Danhui Ye, Kuiwen Xu, Xiaoling Gu, Jun Yu 0002 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Fine-grained Image Classification via Multi-scale Selective Hierarchical Biquadratic PoolingabstractHow to extract distinctive features greatly challenges the fine-grained image classification tasks. In previous models, bilinear pooling has been frequently adopted to address this problem. However, most bilinear pooling models neglect either intra or inter layer feature interaction. This insufficient interaction brings in the loss of discriminative information. In this article, we devise a novel fine-grained image classification approach named M ulti-scale S elective H ierarchical bi Q uadratic P ooling (MSHQP). The proposed biquadratic pooling simultaneously models intra and inter layer feature interactions and enhances part response by integrating multi-layer features. The subsequent coarse-to-fine multi-scale interaction structure captures the complementary information within features. Finally, the active interaction selection module adaptively learns the optimal interaction subset for a specific dataset. Consequently, we obtain a robust image representation with coarse-to-fine semantics. We conduct experiments on five benchmark datasets. The experimental results demonstrate that MSHQP achieves competitive or even match the state-of-the-art methods in terms of both accuracy and computational efficiency, with 89.0%, 94.9%, 93.4%, 90.4%, and 91.5% top-1 classification accuracy on CUB-200-2011, Stanford-Cars, FGVC-Aircraft, Stanford-Dog, and VegFru, respectively. Min Tan 0005, Fu Yuan, Jun Yu 0002, Guijun Wang, Xiaoling Gu |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2021 | Fast and Fusion: Real-Time Pedestrian Detector Boosted by Body-Head Fusion
Jie Huang 0033, Xiaoling Gu |
PRCV (1) | 2 |
| 2021 | Toward Multi-Modal Conditioned Fashion Image TranslationabstractHaving the capability to synthesize photo-realistic fashion product images conditioned on multiple attributes or modalities would bring many new exciting applications. In this work, we propose an end-to-end network architecture that built upon a new generative adversarial network for automatically synthesizing photo-realistic images of fashion products under multiple conditions. Given an input pose image that consists of a 2D skeleton pose and a sentence description of products, our model synthesizes a fashion image preserving the same pose and wearing the fashion products described as the text. Specifically, the generator$G$tries to generate realistic-looking fashion images based on a$\langle \mathsf {pose}, \mathsf {text} \rangle$pair condition to fool the discriminator. An attention network is added for enhancing the generator, which predicts a probability map indicating which part of the image needs to be attended for translation. In contrast, the discriminator$D$distinguishes real images from the translated ones based on the input pose image and text description. The discriminator is divided into two multi-scale sub-discriminators for improving image distinguishing task. Quantitative and qualitative analysis demonstrates that our method is capable of synthesizing realistic images that retain the poses of given images while matching the semantics of provided sentence descriptions. Xiaoling Gu, Jun Yu 0002, Yongkang Wong, Mohan Kankanhalli |
IEEE Trans. Multim. | 1 |
| 2020 | Fashion-Sketcher: A Model for Producing Fashion Sketches of Multiple Categories
Junkai Fang, Xiaoling Gu, Min Tan 0005 |
PRCV (2) | 2 |
| 2020 | Fashion analysis and understanding with artificial intelligence
Xiaoling Gu, Fei Gao 0006, Min Tan 0005 |
Inf. Process. Manag. | 1 |
| 2019 | One net to rule them all: efficient recognition and retrieval of POI from geo-tagged photos
Xiaoling Gu, Suguo Zhu, Lidan Shou, Gang Chen 0001 |
Multim. Tools Appl. | 2 |
| 2019 | Multi-Modal and Multi-Domain Embedding Learning for Fashion Retrieval and AnalysisabstractBig data analytics has been revolutionizing the fashion industry in recent years. This is evidenced by the fact that popular fashion brands and designers have relied on big data analytics to trace fashion trends and predict market patterns. In this paper, we propose learning a common latent feature representation from heterogeneous fashion data. Specifically, we design a multi-modal and multi-domain embedding learning framework for fashion analysis and data retrieval. Unlike most of the existing multi-view embedding methods, which only consider the heterogeneous similarity constraint, our proposed framework jointly considers both the homogeneous and heterogeneous similarity constraints to capture cross-view similarity and preserve the similarity of the same view. The proposed framework is comprised of two projection steps. In the first projection, a quintuplet-based ranking loss is proposed for multi-domain fashion data to preserve the homogeneous similarity. In the second projection, a cross-view similarity ranking loss is designed for multi-modal fashion data to capture heterogeneous similarity. By utilizing the learned common latent feature representation, the distance between any vector pairs from same or different modalities can reflect its semantic similarity. Quantitative evaluation on a new large-scale dataset and a fashion analysis case study demonstrate the effectiveness of our proposed method. Xiaoling Gu, Yongkang Wong, Lidan Shou, Gang Chen 0001, Mohan Kankanhalli |
IEEE Trans. Multim. | 1 |
| 2017 | Understanding Fashion Trends from Street Photos via Neighbor-Constrained Embedding LearningabstractDriven by the increasing popular image-dominated social networks, such as Instagram, Pinterest and Chictopica, sharing of daily-life street photos now plays an influential role in fashion adoption between fashion trend-setters and followers. In this work, we propose a deep learning based fine-grained embedding learning approach for street fashion analysis by leveraging user-generated street fashion data. Specifically, we present QuadNet, an effective CNN based image embedding network driven by both multi-task classification loss and neighbor-constrained similarity loss. The latter loss function is computed with a novel quadruplet loss function, which considers both hard and soft positive neighbors as well as a negative neighbor for each anchor image. The embedded feature learned from co-optimization is effective for both fine-grained classification task and image retrieval task. Quantitative evaluation on a newly collected large-scale multi-task street photo dataset shows that our QuadNet outperforms the state-of-the-art triplet network by a significant margin. In order to further evaluate the effectiveness of the learned embedding, we analyze and trace the fashion trends of New York City from 2011 to 2016. In our analysis, we are able to identify some short-term and long-term fashion styles. Xiaoling Gu, Yongkang Wong, Lidan Shou, Gang Chen 0001, Mohan Kankanhalli |
ACM Multimedia | 1 |
| 2017 | CSIR4G: An effective and efficient cross-scenario image retrieval model for glasses
Xiaoling Gu, Sai Wu, Lidan Shou, Ke Chen 0005, Gang Chen 0001 |
Inf. Sci. | 1 |
| 2016 | iGlasses: A Novel Recommendation System for Best-fit GlassesabstractWe demonstrate iGlasses, a novel recommendation system that accepts a frontal face photo as the input and returns the best-fit eyeglasses as the output. As conventional recommendation techniques such as collaborative filtering become inapplicable in the problem, we propose a new recommendation method which exploits the implicit matching rules between human faces and eyeglasses. We first define fine-grained attributes for human faces and frames of glasses respectively. Then, we develop a recommendation framework based on a probabilistic graphical model, which effectively captures the correlation among these fine-grained attributes. Ranking of the frames (glasses) is done by their similarity to the query facial attributes. Finally, we produce a synthesized image for the input face to demonstrate the visual effect when wearing the recommended glasses. Xiaoling Gu, Lidan Shou, Ke Chen 0005, Sai Wu, Gang Chen 0001 |
SIGIR | 1 |
| 2015 | Cross-Scenario Eyeglasses Retrieval via EGYPT ModelabstractIn this paper, we present FGSS (Fashion Glasses Search System), an innovative cross-scenario eyeglasses retrieval system which automatically recognizes eyeglasses in real-world photos, e.g. the photo of a fashion girl with a stylish pair of eyeglasses, and retrieves a ranking list of visually similar product instances from the database. We propose a novel segmentation-free framework for FGSS to bridge two search gaps, semantic gap and feature gap, where a new type of keypoint-based scheme called EGYPT is tailored for eyeglasses to facilitate the search. In the EGYPT, we use the hybrid descriptors which combine the shape, color and texture features as a feature representation for eyeglasses. The experimental study on the real-world photo dataset and eyeglasses product dataset demonstrates the effectiveness of EGYPT model. Xiaoling Gu, Mengwen Li, Sai Wu, Lidan Shou, Gang Chen 0001 |
ICMR | 1 |