VLDB 2026 Research / reviewers in the wild / expert
Wei Zhang 0031
dblp:10/4661-31
· DBLP profile ↗
57ranked-venue papers
12as first author
19since 2021 · last 2025
0000-0002-1492-8286ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 43 · 11 first-author · 14 since 2021Artificial intelligence and machine learning · 18 · 1 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 4Computer networks · 2 · 2 first-authorSecurity and privacy · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | High-Fidelity Object Removal through Boosting Diffusion ProcessesabstractThe remarkable image understanding and generation capabilities of diffusion models have made image editing a highly promising area of research. As a significant subtask within the field, object removal aims to remove objects from a specified region and fill the missing pixels with visually coherent and semantically sound content. Despite the great progress made in deep generative models, research in this area still faces several challenges: i. high expense of model training induced by the data scarcity and the difficulity in large-scale real data collection ii. current training-free methods are unable to drastically change the behavior of the attention layer that has been set during the pre-training phase for text-guided object inpainting. In this paper, we introduce HybridRemover, a two-stage diffusion scheme. We decouple the task into two subtasks: one to remove the specified objects from the target region and one to perform image restoration of the target region. By fine-tuning the SD-Inpainting model with a very small amount of data, we transform it into a model that focuses only on the complete removal of the object without considering the surrounding effects, and the overall repair of the image is taken care of by the SD-Inpainting model cascaded behind it. As a result of our efforts, our method achieved state-of-the-art performance in object removal tasks. Even when a strong perspective distortion gets involved, our method delivers exceptional results. Haihui Fan, Xiaoyan Gu 0001, Wei Zhang 0031, Wu Liu 0005 |
ISCAS | 4 |
| 2025 | Interactive Conversational Head GenerationabstractWe introduce a new conversation head generation benchmark for synthesizing behaviors of a single interlocutor in a face-to-face conversation. The capability to automatically synthesize interlocutors which can participate in long and multi-turn conversations is vital and offer benefits for various applications, including digital humans, virtual agents, and social robots. While existing research primarily focuses on talking head generation (one-way interaction), hindering the ability to create a digital human for conversation (two-way) interaction due to the absence of listening and interaction parts. In this work, we construct two datasets to address this issue, "ViCo" for independent talking and listening head generation tasks at the sentence level, and "ViCo-X", for synthesizing interlocutors in multi-turn conversational scenarios. Based on ViCo and ViCo-X, we define three novel tasks targeting the interaction modeling during the face-to-face conversation: 1) responsive listening head generation making listeners respond actively to the speaker with non-verbal signals, 2) expressive talking head generation guiding speakers to be aware of listeners' behaviors, and 3) conversational head generation to integrate the talking/listening ability in one interlocutor. Along with the datasets, we also propose corresponding baseline solutions to the three aforementioned tasks. Experimental results show that our baseline method could generate responsive and vivid agents that can collaborate with real person to fulfil the whole conversation. Mohan Zhou, Yalong Bai, Wei Zhang 0031, Ting Yao 0003, Tiejun Zhao |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Cross-Modal Quantization for Co-Speech Gesture GenerationabstractLearning proper representations for speech and gesture is essential for co-speech gesture generation. Existing approaches either utilize direct representations or independently encode the speech and gesture, which neglect the joint representation to highlight the interplay between these two modalities. In this work, we propose a novel Cross-modal Quantization (CMQ) to jointly learn the quantized codes for speech and gesture together. Such representation highlights the speech-gesture interaction before actually learning the complex mapping, and thus better suits the intricate mapping between speech and gesture. Specifically, the Cross-modal Quantizer jointly encodes speech and gesture as discrete codebooks, enabling better cross-modal interaction. Cross-modal Predictor subsequently utilizes the learned codebooks to autoregressively predict the next-step gesture. With cross-modal quantization, our approach yields much higher codebook usage and generates more realistic and diverse gestures in practice. Extensive experiments are conducted on both 3D and 2D datasets as well as the subjective user study, demonstrating a clear performance gain compared to several baseline models in terms of audio-visual alignment and gesture diversity. In particular, our method demonstrates a three-fold improvement in diversity compared to baseline models, while simultaneously maintaining high motion fidelity. Zheng Wang 0059, Wei Zhang 0031, Long Ye, Dan Zeng 0001, Tao Mei 0001 |
IEEE Trans. Multim. | 2 |
| 2024 | Visualizing and Understanding Patch Interactions in Vision TransformerabstractVision transformer (ViT) has become a leading tool in various computer vision tasks, owing to its unique self-attention mechanism that learns visual representations explicitly through cross-patch information interactions. Despite having good success, the literature seldom explores the explainability of ViT, and there is no clear picture of how the attention mechanism with respect to the correlation across comprehensive patches will impact the performance and what is the further potential. In this work, we propose a novel explainable visualization approach to analyze and interpret the crucial attention interactions among patches for ViT. Specifically, we first introduce a quantification indicator to measure the impact of patch interaction and verify such quantification on attention window design and indiscriminative patches removal. Then, we exploit the effective responsive field of each patch in ViT and devise a window-free transformer (WinfT) architecture accordingly. Extensive experiments on ImageNet demonstrate that the exquisitely designed quantitative method is shown able to facilitate ViT model learning, leading the top-1 accuracy by 4.28% at most. More remarkably, the results on downstream fine-grained recognition tasks further validate the generalization of our proposal. Jie Ma 0006, Yalong Bai, Bineng Zhong 0001, Wei Zhang 0031, Ting Yao 0003, Tao Mei 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Robust facial marker tracking based on a synthetic analysis of optical flows and the YOLO network
Zeyu Tian, Dongdong Weng, Wei Zhang 0031 |
Vis. Comput. | 5 |
| 2023 | Visual-Aware Text-to-Speech*abstractDynamically synthesizing talking speech that actively responds to a listening head is critical during the face-to-face interaction. For example, the speaker could take advantage of the listener’s facial expression to adjust the tones, stressed syllables, or pauses. In this work, we present a new visual-aware text-to-speech (VA-TTS) task to synthesize speech conditioned on both textual inputs and sequential visual feedback (e.g., nod, smile) of the listener in face-to-face communication. Different from traditional text-to-speech, VA-TTS highlights the impact of visual modality. On this newly-minted task, we devise a baseline model to fuse phoneme linguistic information and listener visual signals for speech synthesis. Extensive experiments on multimodal conversation dataset ViCo-X verify our proposal for generating more natural audio with scenario-appropriate rhythm and prosody. Mohan Zhou, Yalong Bai, Wei Zhang 0031, Ting Yao 0003, Tiejun Zhao, Tao Mei 0001 |
ICASSP | 3 |
| 2023 | MRN: Multiplexed Routing Network for Incremental Multilingual Text RecognitionabstractMultilingual text recognition (MLTR) systems typically focus on a fixed set of languages, which makes it difficult to handle newly added languages or adapt to ever-changing data distribution. In this paper, we propose the Incremental MLTR (IMLTR) task in the context of incremental learning (IL), where different languages are introduced in batches. IMLTR is particularly challenging due to rehearsal-imbalance, which refers to the uneven distribution of sample characters in the rehearsal set, used to retain a small amount of old data as past memories. To address this issue, we propose a Multiplexed Routing Network (MRN). MRN trains a recognizer for each language that is currently seen. Subsequently, a language domain predictor is learned based on the rehearsal set to weigh the recognizers. Since the recognizers are derived from the original data, MRN effectively reduces the reliance on older data and better fights against catastrophic forgetting, the core issue in IL. We extensively evaluate MRN on MLT17 and MLT19 datasets. It outperforms existing general-purpose IL methods by large margins, with average accuracy improvements ranging from 10.3% to 35.8% under different settings. Code is available at https://github.com/simplify23/MRN. Tianlun Zheng, Zhineng Chen, Bingchen Huang, Wei Zhang 0031, Yu-Gang Jiang 0001 |
ICCV | 4 |
| 2023 | Learning and Evaluating Human Preferences for Conversational Head GenerationabstractA reliable and comprehensive evaluation metric that aligns with manual preference assessments is crucial for conversational head video synthesis methods development. Existing quantitative evaluations often fail to capture the full complexity of human preference, as they only consider limited evaluation dimensions. Qualitative evaluations and user studies offer a solution but are time-consuming and labor-intensive. This limitation hinders the advancement of conversational head generation algorithms and systems. In this paper, we propose a novel learning-based evaluation metric named Preference Score (PS) for fitting human preference according to the quantitative evaluations across different dimensions. PS can serve as a quantitative evaluation without the need for human annotation. Experimental results validate the superiority of Preference Score in aligning with human perception, and also demonstrate robustness and generalizability to unseen data, making it a valuable tool for advancing conversation head generation. We expect this metric could facilitate new advances in conversational head generation. Project page: https://github.com/dc3ea9f/PreferenceScore. Mohan Zhou, Yalong Bai, Wei Zhang 0031, Ting Yao 0003, Tiejun Zhao, Tao Mei 0001 |
ACM Multimedia | 3 |
| 2023 | Augmentation Pathways Network for Visual RecognitionabstractData augmentation is practically helpful for visual recognition, especially at the time of data scarcity. However, such success is only limited to quite a few light augmentations (e.g., random crop, flip). Heavy augmentations are either unstable or show adverse effects during training, owing to the big gap between the original and augmented images. This paper introduces a novel network design, noted as Augmentation Pathways (AP), to systematically stabilize training on a much wider range of augmentation policies. Notably, AP tames various heavy data augmentations and stably boosts performance without a careful selection among augmentation policies. Unlike traditional single pathway, augmented images are processed in different neural paths. The main pathway handles the light augmentations, while other pathways focus on the heavier augmentations. By interacting with multiple paths in a dependent manner, the backbone network robustly learns from shared visual patterns among augmentations, and suppresses the side effect of heavy augmentations at the same time. Furthermore, we extend AP to high-order versions for high-order scenarios, demonstrating its robustness and flexibility in practical usage. Experimental results on ImageNet demonstrate the compatibility and effectiveness on a much wider range of augmentations, while consuming fewer parameters and lower computational costs at inference time. Yalong Bai, Mohan Zhou, Wei Zhang 0031, Tao Mei 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Boosting Generic Visual-Linguistic Representation With Dynamic ContextsabstractPretraining large models on generous multi-modal corpora has accelerated the development of visual-linguistic (VL) representation and achieved great success on various vision-and-language downstream tasks. Learning these models is usually executed by predicting the randomly masked words of captions or patches in images. Such approaches, nevertheless, seldom explore the supervision of causalities behind the caption descriptions or the procedure of generating events beyond still images. In this work, we endow the pretrained models with high-level cognition by delving into dynamic contexts to model the visual and linguistic causalities uniformly. Specifically, we format thedynamic contextsof an image as the sentences describing the eventsbefore,on, andafterimage. Unlike traditional caption-wise similarity, we propose a novel dynamic contexts-based similarity (DCS) metric, in which the correlation of potential causes and effects besides immediate visual content are considered to measure the relevance among images. DCS can be further simplified by parameterizing event continuity to relax the requirements on dense contextual event annotations. A new pre-task is designed to minimize the feature distances of dynamically contextual relevant images and incorporate the event causality and commonsense knowledge into the VL representation learning. Models based on our dynamic contexts significantly outperform typical VL models on multiple cross-modal downstream tasks, including the conventional visual commonsense reasoning (VCR), visual question answering (VQA), zero-shot image-text retrieval, and extended image / event ordering tasks. Guoqing Ma 0002, Yalong Bai, Wei Zhang 0031, Ting Yao 0003, Basem Shihada, Tao Mei 0001 |
IEEE Trans. Multim. | 3 |
| 2022 | Directional Self-supervised Learning for Heavy Image AugmentationsabstractDespite the large augmentation family, only a few cherry-picked robust augmentation policies are beneficial to self-supervised image representation learning. In this paper, we propose a directional self-supervised learning paradigm (DSSL), which is compatible with significantly more augmentations. Specifically, we adapt heavy augmentation policies after the views lightly augmented by standard augmentations, to generate harder view (HV). HV usually has a higher deviation from the original image than the lightly augmented standard view (SV). Unlike previous methods equally pairing all augmented views to symmetrically maximize their similarities, DSSL treats augmented views of the same instance as a partially ordered set (with directions as SV↔SV, SV↔HV), and then equips a directional objective function respecting to the derived relationships among views. DSSL can be easily implemented with a few lines of codes and is highly flexible to popular self-supervised learning frameworks, including SimCLR, Sim-Siam, BYOL. Extensive experimental results on CIFAR and ImageNet demonstrated that DSSL can stably improve various baselines with compatibility to a wider range of augmentations. Code is available at: https://github.com/Yif-Yang/DSSL. Yalong Bai, Wei Zhang 0031, Tao Mei 0001 |
CVPR | 3 |
| 2022 | Responsive Listening Head Generation: A Benchmark Dataset and Baseline
Mohan Zhou, Yalong Bai, Wei Zhang 0031, Ting Yao 0003, Tiejun Zhao, Tao Mei 0001 |
ECCV (38) | 3 |
| 2022 | Singing Voice Synthesis with Vibrato Modeling and Latent Energy RepresentationabstractThis paper proposes an expressive singing voice synthesis system by introducing explicit vibrato modeling and latent energy representation. Vibrato is essential to the naturalness of synthesized sound, due to the inherent characteristics of human singing. Hence, a deep learning-based vibrato model is introduced in this paper to control the vibrato's likeliness, rate, depth and phase in singing, where the vibrato likeliness represents the existence probability of vibrato and it would help improve the singing voice's naturalness. Actually, there is no annotated label about vibrato likeliness in existing singing corpus. We adopt a novel vibrato likeliness labeling method to label the vibrato likeliness automatically. Meanwhile, the power spectrogram of audio contains rich information that can improve the expressiveness of singing. An autoencoder-based latent energy bottleneck feature is proposed for expressive singing voice synthesis. Experimental results on the open dataset NUS48E show that both the vibrato modeling and the latent energy representation could significantly improve the expressiveness of singing voice. The audio samples are shown in the demo website11https://mango321321.github.io/ExpressiveSing/. Wei Zhang 0031, Zhengchen Zhang, Dan Zeng 0001, Zhi Liu 0003 |
MMSP | 3 |
| 2021 | Exploiting Relationship for Complex-scene Image GenerationabstractThe significant progress on Generative Adversarial Networks (GANs) has facilitated realistic single-object image generation based on language input. However, complex-scene generation (with various interactions among multiple objects) still suffers from messy layouts and object distortions, due to diverse configurations in layouts and appearances. Prior methods are mostly object-driven and ignore their inter-relations that play a significant role in complex-scene images. This work explores relationship-aware complex-scene image generation, where multiple objects are inter-related as a scene graph. With the help of relationships, we propose three major updates in the generation framework. First, reasonable spatial layouts are inferred by jointly considering the semantics and relationships among objects. Compared to standard location regression, we show relative scales and distances serve a more reliable target. Second, since the relations between objects have significantly influenced an object's appearance, we design a relation-guided generator to generate objects reflecting their relationships. Third, a novel scene graph discriminator is proposed to guarantee the consistency between the generated image and the input scene graph. Our method tends to synthesize plausible layouts and objects, respecting the interplay of multiple objects in an image. Experimental results on Visual Genome and HICO-DET datasets show that our proposed method significantly outperforms prior arts in terms of IS and FID metrics. Based on our user study and visual inspection, our method is more effective in generating logical layout and appearance for complex-scenes. Tianyu Hua, Hongdong Zheng, Yalong Bai, Wei Zhang 0031, Xiao-Ping Zhang 0002, Tao Mei 0001 |
AAAI | 4 |
| 2021 | ARShoe: Real-Time Augmented Reality Shoe Try-on System on SmartphonesabstractVirtual try-on technology enables users to try various fashion items using augmented reality and provides a convenient online shopping experience. However, most previous works focus on the virtual try-on for clothes while neglecting that for shoes, which is also a promising task. To this concern, this work proposes a real-time augmented reality virtual shoe try-on system for smartphones, namely ARShoe. Specifically, ARShoe adopts a novel multi-branch network to realize pose estimation and segmentation simultaneously. A solution to generate realistic 3D shoe model occlusion during the try-on process is presented. To achieve a smooth and stable try-on effect, this work further develop a novel stabilization method. Moreover, for training and evaluation, we construct the very first large-scale foot benchmark with multiple virtual shoe try-on task-related labels annotated. Exhaustive experiments on our newly constructed benchmark demonstrate the satisfying performance of ARShoe. Practical tests on common smartphones validate the real-time performance and stabilization of the proposed approach. Shan An, Guangfu Che, Jinghao Guo, Haogang Zhu, Junjie Ye 0004, Fangru Zhou, Zhaoqi Zhu, Aishan Liu, Wei Zhang 0031 |
ACM Multimedia | 10 |
| 2021 | Trustworthy AI'21: 1st International Workshop on Trustworthy AI for Multimedia ComputingabstractIn this workshop, we are addressing the trustworthy AI issues for Multimedia Computing. We aim to bring together researchers in the trustworthy aspects of Multimedia Computing and facilitate discussions in injecting trusts into multimedia to develop trustworthy AI techniques that are reliable and acceptable to multimedia researchers and practitioners. Our scope is at the conjunction of multimedia, computer vision and trustworthy AI, including Explainability, Robustness and Safety, Data Privacy, Accountability and Transparency, and Fairness. Teddy Furon, Jingen Liu, Yogesh S. Rawat, Wei Zhang 0031, Qi Zhao 0001 |
ACM Multimedia | 4 |
| 2021 | ViDA-MAN: Visual Dialog with Digital HumansabstractWe demonstrate ViDA-MAN, a digital-human agent for multi-modal interaction, which offers realtime audio-visual responses to instant speech inquiries. Compared to traditional text or voice-based system, ViDA-MAN offers human-like interactions (e.g, vivid voice, natural facial expression and body gestures). Given a speech request, the demonstration is able to response with high quality videos in sub-second latency. To deliver immersive user experience, ViDA-MAN seamlessly integrates multi-modal techniques including Acoustic Speech Recognition (ASR), multi-turn dialog, Text To Speech (TTS), talking heads video generation. Backed with large knowledge base, ViDA-MAN is able to chat with users on a number of topics including chit-chat, weather, device control, News recommendations, booking hotels, as well as answering questions via structured knowledge. Jiawei Zuo, Liqin Jiang, Meng Chen 0006, Zhengchen Zhang, Wei Zhang 0031, Xiaodong He 0001, Tao Mei 0001 |
ACM Multimedia | 8 |
| 2021 | Flat and Shallow: Understanding Fake Image Detection Models by Architecture ProfilingabstractDigital image manipulations have been heavily abused to spread misinformation. Despite the great efforts dedicated in research community, prior works are mostly performance-driven, i.e., optimizing performances using standard/heavy networks designed for semantic classification. A thorough understanding for fake images detection models is still missing. This paper studies the essential ingredients for a good fake image detection model, by profiling the best-performing architectures. Specifically, we conduct a thorough analysis on a massive number of detection models, and observe how the performances are affected by different patterns of network structure. Our key findings include: 1) with the same computational budget, flat network structures (e.g., large kernel sizes, wide connections) perform better than commonly used deep networks; 2) operations in shallow layers deserve more computational capacities to trade-off performance and computational cost. These findings sketch a general profile for essential models of fake image detection, which show clear differences with those for semantic classification. Furthermore, based on our analysis, we propose a new Depth-Separable Search Space (DSS) for fake image detection. Compared to state-of-the-art methods, our model achieves competitive performance while saving more than 50% parameters. Wei Zhang 0031, Yalong Bai, Qibin Sun, Tao Mei 0001 |
MMAsia | 2 |
| 2021 | Unpaired Person Image Generation With Semantic Parsing TransformationabstractIn this paper, we tackle the problem of pose-guided person image generation with unpaired data, which is a challenging problem due to non-rigid spatial deformation. Instead of learning a fixed mapping directly between human bodies as previous methods, we propose a new pathway to decompose a single fixed mapping into two subtasks, namely, semantic parsing transformation and appearance generation. First, to simplify the learning for non-rigid deformation, a semantic generative network is developed to transform semantic parsing maps between different poses. Second, guided by semantic parsing maps, we render the foreground and background image, respectively. A foreground generative network learns to synthesize semantic-aware textures, and another background generative network learns to predict missing background regions caused by pose changes. Third, we enable pseudo-label training with unpaired data, and demonstrate that end-to-end training of the overall network further refines the semantic map prediction and final results accordingly. Moreover, our method is generalizable to other person image generation tasks defined on semantic maps, e.g., clothing texture transfer, controlled image manipulation, and virtual try-on. Experimental results on DeepFashion and Market-1501 datasets demonstrate the superiority of our method, especially in keeping better body shapes and clothing attributes, as well as rendering structure-coherent backgrounds. Sijie Song, Wei Zhang 0031, Jiaying Liu 0001, Zongming Guo, Tao Mei 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2020 | Look-Into-Object: Self-Supervised Structure Modeling for Object RecognitionabstractMost object recognition approaches predominantly focus on learning discriminative visual patterns, while overlooking the holistic object structure. Though important, structure modeling usually requires significant manual annotations and therefore is labor-intensive. In this paper, we propose to ``look into object" (explicitly yet intrinsically model the object structure) through incorporating self-supervisions into the traditional framework. We show the recognition backbone can be substantially enhanced for more robust representation learning, without any cost of extra annotation and inference speed. Specifically, we first propose an object-extent learning module for localizing the object according to the visual patterns shared among the instances in the same category. We then design a spatial context learning module for modeling the internal structures of the object, through predicting the relative positions within the extent. These two modules can be easily plugged into any backbone networks during training and detached at inference time. Extensive experiments show that our look-into-object approach (LIO) achieves large performance gain on a number of benchmarks, including generic object recognition (ImageNet) and fine-grained object recognition tasks (CUB, Cars, Aircraft). We also show that this learning paradigm is highly generalizable to other tasks such as object detection and segmentation (MS COCO). Project page: https://github.com/JDAI-CV/LIO. Mohan Zhou, Yalong Bai, Wei Zhang 0031, Tiejun Zhao, Tao Mei 0001 |
CVPR | 3 |
| 2020 | Classes Matter: A Fine-Grained Adversarial Approach to Cross-Domain Semantic Segmentation
Wei Zhang 0031, Ling-Yu Duan, Tao Mei 0001 |
ECCV (14) | 3 |
| 2020 | SketchMan: Learning to Create Professional SketchesabstractHuman free-hand sketches have been studied in various fields including sketch recognition, synthesis and sketch-based image retrieval. We propose a new challenging task sketch enhancement (SE) defined in an ill-posed space, i.e. enhancing a non-professional sketch (NPS) to a professional sketch (PS), which is a creative generation task different from sketch abstraction, sketch completion and sketch variation. For the first time we release a database of NPS with PS for anime characters. We cast sketch enhancement as an image-to-image translation problem by exploiting the relationship to corresponding intensive or sparse pixel domains for sketch domain. Specifically, we explore three different routines based on conditional generative adversarial network (cGAN), i.e. Sketch-Sketch (SS), Sketch-Colorization-Sketch (SCS) and Sketch-Abstraction-Sketch (SAS). SS is a one-stage model that directly maps NPS to PS, while SCS and SAS are two-stage models where auxiliary inputs, grayscale parsing and shape parsing, are involved. Multiple metrics are used to evaluate the performance of the models in both the sketch domain and other low-level feature domains. With quantitative and qualitative analysis of the experiments, we have established solid baselines, which, we hope, could encourage more research conducted on this task. Our dataset is publicly available via https://github.com/LCXCUC/SketchMan2020. Jia Li 0044, Nan Gao 0001, Wei Zhang 0031, Tao Mei 0001, Hui Ren 0002 |
ACM Multimedia | 4 |
| 2020 | Down to the Last Detail: Virtual Try-on with Fine-grained DetailsabstractVirtual try-on has attracted lots of research attention due to its potential applications in e-commerce, virtual reality and fashion design. However, existing methods can hardly preserve the fine-grained details (e.g., clothing texture, facial identity, hair style, skin tone) during generation, due to the non-rigid body deformation and multi-scale details. In this work, we propose a multi-stage framework to synthesize person images, where fine-grained details can be well preserved. To address the long-range translation and rich-details generation, we propose a Tree-Block (tree dilated fusion block) to replace standard ResNet-block where applicable. Notably, multi-scale feature maps can be smoothly fused for fine-grained detail generation, by incorporating larger spatial context at multiple scales. With a delicate end-to-end training scheme, our whole framework can be jointly optimized for results with significantly better visual fidelity and richer details. Moreover, we also explore the potential application in video-based virtual try-on. By harnessing the well-trained image generator and an extra video-level adaptor, a model photo can be well animated with a driving pose sequence. Extensive evaluations on standard datasets and user study demonstrate that our proposed framework achieves the state-of-the-art results, especially in preserving visual details in clothing texture and facial identity. Our implementation is publicly available via https://github.com/JDAI-CV/Down-to-the-Last-Detail-Virtual-Try-on-with-Detail-Carving. Jiahang Wang, Tong Sha, Wei Zhang 0031, Zhoujun Li 0001, Tao Mei 0001 |
ACM Multimedia | 3 |
| 2020 | AI-SAS: Automated In-match Soccer Analysis SystemabstractReal-time in-match soccer statistics provide continuous tracking of soccer ball and player positions and speeds, enabling advanced analytics. Currently, only elite soccer leagues have the luxury of tracking in-match soccer statistics operated with a large number of trained personnel. In this work, we present an Automated In-match Soccer Analysis System (AI-SAS), using a domain-knowledge-based multi-view global tracking. This system tracks player team, position, and speed automatically, providing real-time in-match team- and individual-level statistics and analyses. In comparison with the latest soccer analysis systems, AI-SAS is more scalable in streaming multiple video sources for real-time process and more flexible in hosting plug-and-play deep-learning-based tracking-by-detection algorithms. The global multi-view tracking also overcomes the single-view limitation and improves the tracking accuracy. Ning Zhang 0023, Wei Zhang 0031, Dan Zeng 0001, Jingen Liu, Tao Mei 0001 |
ACM Multimedia | 4 |
| 2019 | Destruction and Construction Learning for Fine-Grained Image RecognitionabstractDelicate feature representation about object parts plays a critical role in fine-grained recognition. For example, experts can even distinguish fine-grained objects relying only on object parts according to professional knowledge. In this paper, we propose a novel "Destruction and Construction Learning" (DCL) method to enhance the difficulty of fine-grained recognition and exercise the classification model to acquire expert knowledge. Besides the standard classification backbone network, another "destruction and construction" stream is introduced to carefully "destruct" and then "reconstruct" the input image, for learning discriminative regions and features. More specifically, for "destruction", we first partition the input image into local regions and then shuffle them by a Region Confusion Mechanism (RCM). To correctly recognize these destructed images, the classification network has to pay more attention to discriminative regions for spotting the differences. To compensate the noises introduced by RCM, an adversarial loss, which distinguishes original images from destructed ones, is applied to reject noisy patterns introduced by RCM. For "construction", a region alignment network, which tries to restore the original spatial layout of local regions, is followed to model the semantic correlation among local regions. By jointly training with parameter sharing, our proposed DCL injects more discriminative local details to the classification network. Experimental results show that our proposed framework achieves state-of-the-art performance on three standard benchmarks. Moreover, our proposed method does not need any external knowledge during training, and there is no computation overhead at inference time except the standard classification network feed-forwarding. Source code: https://github.com/JDAI-CV/DCL. Yalong Bai, Wei Zhang 0031, Tao Mei 0001 |
CVPR | 3 |
| 2019 | Unsupervised Person Image Generation With Semantic Parsing TransformationabstractIn this paper, we address unsupervised pose-guided person image generation, which is known challenging due to non-rigid deformation. Unlike previous methods learning a rock-hard direct mapping between human bodies, we propose a new pathway to decompose the hard mapping into two more accessible subtasks, namely, semantic parsing transformation and appearance generation. Firstly, a semantic generative network is proposed to transform between semantic parsing maps, in order to simplify the non-rigid deformation learning. Secondly, an appearance generative network learns to synthesize semantic-aware textures. Thirdly, we demonstrate that training our framework in an end-to-end manner further refines the semantic maps and final results accordingly. Our method is generalizable to other semantic-aware person image generation tasks, e.g., clothing texture transfer and controlled image manipulation. Experimental results demonstrate the superiority of our method on DeepFashion and Market-1501 datasets, especially in keeping the clothing attributes and better body shapes. Sijie Song, Wei Zhang 0031, Jiaying Liu 0001, Tao Mei 0001 |
CVPR | 2 |
| 2019 | VrR-VG: Refocusing Visually-Relevant RelationshipsabstractRelationships encode the interactions among individual instances and play a critical role in deep visual scene understanding. Suffering from the high predictability with non-visual information, relationship models tend to fit the statistical bias rather than ``learning" to infer the relationships from images. To encourage further development in visual relationships, we propose a novel method to mine more valuable relationships by automatically pruning visually-irrelevant relationships. We construct a new scene graph dataset named Visually-Relevant Relationships Dataset (VrR-VG) based on Visual Genome. Compared with existing datasets, the performance gap between learnable and statistical method is more significant in VrR-VG, and frequency-based analysis does not work anymore. Moreover, we propose to learn a relationship-aware representation by jointly considering instances, attributes and relationships. By applying the representation-aware feature learned on VrR-VG, the performances of image captioning and visual question answering are systematically improved, which demonstrates the effectiveness of both our dataset and features embedding schema. Both our VrR-VG dataset and representation-aware features will be made publicly available soon. Yuanzhi Liang, Yalong Bai, Wei Zhang 0031, Xueming Qian, Li Zhu 0003, Tao Mei 0001 |
ICCV | 3 |
| 2019 | Sampling Wisely: Deep Image Embedding by Top-K Precision OptimizationabstractDeep image embedding aims at learning a convolutional neural network (CNN) based mapping function that maps an image to a feature vector. The embedding quality is usually evaluated by the performance in image search tasks. Since very few users bother to open the second page search results, top-k precision mostly dominates the user experience and thus is one of the crucial evaluation metrics for the embedding quality. Despite being extensively studied, existing algorithms are usually based on heuristic observation without theoretical guarantee. Consequently, gradient descent direction on the training loss is mostly inconsistent with the direction of optimizing the concerned evaluation metric. This inconsistency certainly misleads the training direction and degrades the performance. In contrast to existing works, in this paper, we propose a novel deep image embedding algorithm with end-to-end optimization to top-k precision, the evaluation metric that is closely related to user experience. Specially, our loss function is constructed with wisely selected ``misplaced" images along the top k nearest neighbor decision boundary, so that the gradient descent update directly promotes the concerned metric, top-k precision. Further more, our theoretical analysis on the upper bounding and consistency properties of the proposed loss supports that minimizing our proposed loss is equivalent to maximizing top-k precision. Experiments show that our proposed algorithm outperforms all compared state-of-the-art deep image embedding algorithms on three benchmark datasets. Chaofan Xu, Wei Zhang 0031, Ling-Yu Duan, Tao Mei 0001 |
ICCV | 3 |
| 2019 | Everyone is a Cartoonist: Selfie Cartoonization with Attentive Adversarial NetworksabstractSelfie and cartoon are two popular artistic forms that are widely presented in our daily life. Despite the great progress in image translation/stylization, few techniques focus specifically on selfie cartoonization, since cartoon images usually contain artistic abstraction (e.g., large smoothing areas) and exaggeration (e.g., large/delicate eyebrows). In this paper, we address this problem by proposing a selfie cartoonization Generative Adversarial Network (scGAN), which mainly uses an attentive adversarial network (AAN) to emphasize specific facial regions and ignore low-level details. More specifically, we first design a cycle-like architecture to enable training with unpaired data. Then we design three losses from different aspects. A total variation loss is used to highlight important edges and contents in cartoon portraits. An attentive cycle loss is added to lay more emphasis on delicate facial areas such as eyes. In addition, a perceptual loss is included to eliminate artifacts and improve robustness of our method. Experimental results show that our method is capable of generating different cartoon styles and outperforms a number of state-of-the-art methods. Wei Zhang 0031, Tao Mei 0001 |
ICME | 2 |
| 2019 | An efficient privacy protection scheme for data security in video surveillance
Wei Zhang 0031, Huazhu Fu, Wenqi Ren, Xinpeng Zhang 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2019 | Name-face association with web facial image supervision
Zhineng Chen, Wei Zhang 0031, Hongtao Xie 0001, Xiaoyan Gu 0001 |
Multim. Syst. | 2 |
| 2019 | Pyrboxes: An efficient multi-scale scene text detector with feature pyramids
Fenfen Sheng, Zhineng Chen, Wei Zhang 0031, Bo Xu 0002 |
Pattern Recognit. Lett. | 3 |
| 2019 | Editorial to Special Issue on Deep Learning for Intelligent Multimedia Analytics
Wei Zhang 0031, Ting Yao 0003, Shiai Zhu, Abdulmotaleb El Saddik |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2019 | Deep Learning-Based Multimedia Analytics: A ReviewabstractThe multimedia community has witnessed the rise of deep learning–based techniques in analyzing multimedia content more effectively. In the past decade, the convergence of deep-learning and multimedia analytics has boosted the performance of several traditional tasks, such as classification, detection, and regression, and has also fundamentally changed the landscape of several relatively new areas, such as semantic segmentation, captioning, and content generation. This article aims to review the development path of major tasks in multimedia analytics and take a look into future directions. We start by summarizing the fundamental deep techniques related to multimedia analytics, especially in the visual domain, and then review representative high-level tasks powered by recent advances. Moreover, the performance review of popular benchmarks gives a pathway to technology advancement and helps identify both milestone works and future directions. Wei Zhang 0031, Ting Yao 0003, Shiai Zhu, Abdulmotaleb El Saddik |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2018 | Consistent and Specific Multi-View Subspace ClusteringabstractMulti-view clustering has attracted intensive attention due to the effectiveness of exploiting multiple views of data. However, most existing multi-view clustering methods only aim to explore the consistency or enhance the diversity of different views. In this paper, we propose a novel multi-view subspace clustering method (CSMSC), where consistency and specificity are jointly exploited for subspace representation learning. We formulate the multi-view self-representation property using a shared consistent representation and a set of specific representations, which better fits the real-world datasets. Specifically, consistency models the common properties among all views, while specificity captures the inherent difference in each view. In addition, to optimize the non-convex problem, we introduce a convex relaxation and develop an alternating optimization algorithm to recover the corresponding data representations. Experimental evaluations on four benchmark datasets demonstrate that the proposed approach achieves better performance over several state-of-the-arts. Shirui Luo, Changqing Zhang 0002, Wei Zhang 0031, Xiaochun Cao |
AAAI | 3 |
| 2018 | Separable reversible data hiding in encrypted images via adaptive embedding strategy with block selection
Chuan Qin 0001, Wei Zhang 0031, Xinpeng Zhang 0001, Chin-Chen Chang 0001 |
Signal Process. | 2 |
| 2018 | Fake Colorized Image DetectionabstractImage forensics aims to detect the manipulation of digital images. Currently, splicing detection, copy-move detection, and image retouching detection are attracting significant attention from researchers. However, image editing techniques develop over time. An emerging image editing technique is colorization, in which grayscale images are colorized with realistic colors. Unfortunately, this technique may also be intentionally applied to certain images to confound object recognition algorithms. To the best of our knowledge, no forensic technique has yet been invented to identify whether an image is colorized. We observed that, compared with natural images, colorized images, which are generated by three state-of-the-art methods, possess statistical differences for the hue and saturation channels. Besides, we also observe statistical inconsistencies in the dark and bright channels, because the colorization process will inevitably affect the dark and bright channel values. Based on our observations, i.e., potential traces in the hue, saturation, dark, and bright channels, we propose two simple yet effective detection methods for fake colorized images: Histogram-based fake colorized image detection and feature encoding-based fake colorized image detection. Experimental results demonstrate that both proposed methods exhibit a decent performance against multiple state-of-the-art colorization approaches. Yuanfang Guo, Xiaochun Cao, Wei Zhang 0031, Rui Wang 0032 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2017 | Binarized Mode Seeking for Scalable Visual Pattern DiscoveryabstractThis paper studies visual pattern discovery in large-scale image collections via binarized mode seeking, where images can only be represented as binary codes for efficient storage and computation. We address this problem from the perspective of binary space mode seeking. First, a binary mean shift (bMS) is proposed to discover frequent patterns via mode seeking directly in binary space. The binomial-based kernel and binary constraint are introduced for binarized analysis. Second, we further extend bMS to a more general form, namely contrastive binary mean shift (cbMS), which maximizes the contrastive density in binary space, for finding informative patterns that are both frequent and discriminative for the dataset. With the binarized algorithm and optimization, our methods demonstrate significant computation (50×) and storage (32×) improvement compared to standard techniques operating in Euclidean space, while the performance does not largely degenerate. Furthermore, cbMS discovers more informative patterns by suppressing low discriminative modes. We evaluate our methods on both annotated ILSVRC (1M images) and un-annotated blind Flickr (10M images) datasets with million scale images, which demonstrates both the scalability and effectiveness of our algorithms for discovering frequent and informative patterns in large scale collection. Wei Zhang 0031, Xiaochun Cao, Rui Wang 0032, Yuanfang Guo, Zhineng Chen |
CVPR | 1 |
| 2017 | Retrieving Objects by PartitioningabstractRetrieving objects from large image collection is challenging due to the so-called background-interference, i.e., matching between query object and reference images is usually confused by cluttered background, especially when objects are small. In this paper, we propose an object retrieval technique addressing this problem by partitioning the images. Specifically, several object proposals are partitioned from the images by jointly optimizing their objectness and coverage. The proposal set with maximum objectness score and minimum redundancy is obtained. Therefore,the interference of cluttered background is greatly reduced. Next, the objects are retrieved based on the partitioned proposals, separately and independently to the background. Our method is featured by the fine partitioning, which not only removes interferences from background, but also significantly reduces the number of objects to index. In this way, the effectiveness and efficiency are both achieved, which better suits big data retrieval. Subsequently, feature coding on partitioned objects generates much meaningful representation, and object level connectivity also introduces novel clues into the reranking. Extensive experiments on three popular object retrieval benchmark datasets (Oxford Buildings, Paris, Holiday) show the effectiveness of our method in retrieving small objects out of big data. Wei Zhang 0031, Bin Hu 0001, Xiaochun Cao, Si Liu 0001, Dan Meng 0002 |
IEEE Trans. Big Data | 2 |
| 2016 | MatchDR: Image Correspondence by Leveraging Distance Ratio ConstraintabstractImage correspondence is to establish the connections between coherent images, which can be quite challenging due to the visual and geometric deformations. This paper proposes a robust image correspondence technique from the perspective of spatial regularity. Specifically, the visual deformation is addressed by introducing the spatial information by enforcing the distance ratio constrain. At the same time, the geometric deformation is tolerated by adopting a smoothness term. Subsequently, image correspondence is formulated as permutation problem, for which, we propose a Gradient Guided Simulated Annealing method for robust optimization. Furthermore, our method is much more memory efficient, where the storage complexity is reduced from O(n4) to O(n2). The experiments on several datasets indicate that our proposed formulation and optimization significantly improve the baselines for both visually-similar and semantically-similar images, where both visual and geometric deformations are present. Rui Wang 0032, Wei Zhang 0031, Xiaochun Cao |
ACM Multimedia | 3 |
| 2016 | Semi-fragile watermarking for image authentication based on compressive sensing
Xiaochun Cao, Wei Zhang 0031, Xinpeng Zhang 0001, Jianguo Wei |
Sci. China Inf. Sci. | 3 |
| 2016 | Hyperlink-Aware Object RetrievalabstractIn this paper, we address the problem of object retrieval by hyperlinking the reference data set at subimage level. One of the main challenges in object retrieval involves small objects on cluttered backgrounds, where the similarity between the querying object and a relevant image can be heavily affected by the background. To address this problem, we propose an efficient object retrieval technique by hyperlinking the visual entities among the reference data set. In particular, a two-step framework is proposed: subimage-level hyperlinking and hyperlink-aware reranking. For hyperlinking, we propose a scalable object mining technique using Thread-of-Features, which is designed for mining subimage-level objects. For reranking, the initial search results are reranked with a hyperlink-aware transition matrix encoding subimage-level connectivity. Through this framework, small objects can be retrieved effectively. Moreover, our method introduces only a tiny computation overhead to online processing, due to the sparse transition matrix. The proposed technique is featured by the novel perspective (object hyperlinking) for visual search, as well as the object hyperlinking technique. We demonstrate the effectiveness and efficiency of our hyperlinking and retrieval methods by experimenting upon several object-retrieval data sets. Wei Zhang 0031, Chong-Wah Ngo, Xiaochun Cao |
IEEE Trans. Image Process. | 1 |
| 2015 | Image composite authentication using a single shadow observation
Xiaochun Cao, Handong Zhao, Chuan Wang 0002, Wei Zhang 0031 |
Sci. China Inf. Sci. | 4 |
| 2015 | Topological Spatial Verification for Instance SearchabstractThis paper proposes an elastic spatial verification method for Instance Search, particularly for dealing with non-planar and non-rigid queries exhibiting complex spatial transformations. Different from existing models that map keypoints between images based on a linear transformation (e.g., affine, homography), our model exploits the topological arrangement of keypoints to address the non-linear spatial transformations that are extremely common in real life situations. In particular, we propose a novel technique to elastically verify the topological spatial consistency with the triangulated graph through a “sketch-and-match” scheme. The spatial topology configuration, emphasizing relative positioning rather than absolute coordinates, is first sketched by a triangulated graph, whose edges essentially capture the topological layout of the corresponding keypoints. Next, the spatial consistency is efficiently estimated as the number of common edges between the triangulated graphs. Compared to the existing methods, our technique is much more effective in modeling the complex spatial transformations of non-planar and non-rigid instances, while being compatible to instances with simple linear transformations. Moreover, our method is by nature more robust in spatial verification by considering the locations, rather than the local geometry of keypoints, which are sensitive to motions and viewpoint changes. We evaluate our method extensively on three years of TRECVID datasets, as well as our own dataset MQA, showing large improvement over other methods for the task of Instance Search. Wei Zhang 0031, Chong-Wah Ngo |
IEEE Trans. Multim. | 1 |
| 2014 | Scalable Visual Instance Mining with Threads of FeaturesabstractWe address the problem of visual instance mining, which is to extract frequently appearing visual instances automatically from a multimedia collection. We propose a scalable mining method by exploiting Thread of Features (ToF). Specifically, ToF, a compact representation that links consistent features across images, is extracted to reduce noises, discover patterns, and speed up processing. Various instances, especially small ones, can be discovered by exploiting correlated ToFs. Our approach is significantly more effective than other methods in mining small instances. At the same time, it is also more efficient by requiring much fewer hash tables. We compared with several state-of-the-art methods on two fully annotated datasets: MQA and Oxford, showing large performance gain in mining (especially small) visual instances. We also run our method on another Flickr dataset with one million images for scalability test. Two applications, instance search and multimedia summarization, are developed from the novel perspective of instance mining, showing great potential of our method in multimedia analysis. Wei Zhang 0031, Hongzhi Li 0001, Chong-Wah Ngo, Shih-Fu Chang |
ACM Multimedia | 1 |
| 2014 | Name-Face Association in Web Videos: A Large-Scale Dataset, Baselines, and Open Issues
Zhineng Chen, Chong-Wah Ngo, Wei Zhang 0031, Juan Cao 0001, Yu-Gang Jiang 0001 |
J. Comput. Sci. Technol. | 3 |
| 2014 | Visual Typo Correction by Collocative Optimization: A Case Study on Merchandize ImagesabstractNear-duplicate retrieval (NDR) in merchandize images is of great importance to a lot of online applications on e-Commerce websites. In those applications where the requirement of response time is critical, however, the conventional techniques developed for a general purpose NDR are limited, because expensive post-processing like spatial verification or hashing is usually employed to compromise the quantization errors among the visual words used for the images. In this paper, we argue that most of the errors are introduced because of the quantization process where the visual words are considered individually, which has ignored the contextual relations among words. We propose a "spelling or phrase correction" like process for NDR, which extends the concept of collocations to visual domain for modeling the contextual relations. Binary quadratic programming is used to enforce the contextual consistency of words selected for an image, so that the errors (typos) are eliminated and the quality of the quantization process is improved. The experimental results show that the proposed method can improve the efficiency of NDR by reducing vocabulary size by 1000% times, and under the scenario of merchandize image NDR, the expensive local interest point feature used in conventional approaches can be replaced by color-moment feature, which reduces the time cost by 9202% while maintaining comparable performance to the state-of-the-art methods. Xiaoyong Wei, Zhen-Qun Yang, Chong-Wah Ngo, Wei Zhang 0031 |
IEEE Trans. Image Process. | 4 |
| 2013 | Searching visual instances with topology checking and context modelingabstractInstance Search (INS) is a realistic problem initiated by TRECVID, which is to retrieve all occurrences of the querying object, location, or person from a large video collection. It is a fundamental problem with many applications, and also a challenging problem different from the traditional concept or near-duplicate (ND) search, since the relevancy is defined at instance level. True responses could exhibit various visual variations, such as being small on the image with different background, or showing a non-homography spatial configuration. Based on the Bag-of-Words model, we propose two techniques tailored for Instance Search. Specifically, we explore the use of (1) an elastic spatial topology checking technique based on Delaunay Triangulation (DT), and (2) a practical background context modeling method by simulating the "stare" behavior of human eyes. With DT, we improve the quality of visual matching by accumulating evidence from local topology-preserving patches, significantly boosting the ranks of topology consistent results. On the other hand, we increase the information quantity for visual matching with the "stare" model, such that instances appearing in both similar and different background can be highly ranked as results. The proposed techniques are evaluated on the INS datasets of TRECVID, achieving large performance gain with small computation overhead, compared with several existing methods. Wei Zhang 0031, Chong-Wah Ngo |
ICMR | 1 |
| 2012 | Community as a connector: associating faces with celebrity names in web videosabstractAssociating celebrity faces appearing in videos with their names is of increasingly importance with the popularity of both celebrity videos and related queries. However, the problem is not yet seriously studied in Web video domain. This paper proposes a Community connected Celebrity Name-Face Association approach (C-CNFA), where the community is regarded as an intermediate connector to facilitate the association. Specifically, with the names and faces extracted from Web videos, C-CNFA decomposes the association task into a three-step framework: community discovering, community matching and celebrity face tagging. To achieve the goal of efficient name-face association under this umbrella, algorithms such as the constrained density-based clustering and exemplar based voting are developed by leveraging different pieces of visual and contextual cues. The evaluation on 0.4 million faces and 144 celebrities shows the effectiveness of the proposed C-CNFA approach. Moreover, using the obtained associations, encouraging results are reported in celebrity video ranking. Zhineng Chen, Chong-Wah Ngo, Juan Cao 0001, Wei Zhang 0031 |
ACM Multimedia | 4 |
| 2012 | Video hyperlinking: libraries and tools for threading and visualizing large video collectionabstractWhile HTML documents could be effortlessly hyperlinked by markup tags, creation of the hyperlinks for multimedia objects is by no means easy due to the involvement of various visual processing units and intensive computational overhead. This paper introduces an open source, named VIREO-VH, which provides end-to-end support for creating hyperlinks to thread and visualize collections of videos. The software components include video pre-processing, bag-of-words based inverted file indexing for scalable near-duplicate keyframe search, localization of partial near-duplicate segments, and galaxy visualization of video collection. The open source has been internally used by VIREO research team since 2007, and was evolved over years based on experiences through developing various multimedia applications. Wei Zhang 0031, Hung-Khoon Tan, Chong-Wah Ngo |
ACM Multimedia | 2 |
| 2012 | Snap-and-ask: answering multimodal question by naming visual instanceabstractIn real-life, it is easier to provide a visual cue when asking a question about a possibly unfamiliar topic, for example, asking the question, "Where was this crop circle found?". Providing an image of the instance is far more convenient than texting a verbose description of the visual properties, especially when the name of the query instance is not known. Nevertheless, having to identify the visual instance before processing the question and eventually returning the answer makes multimodal question-answering technically challenging. This paper addresses the problem of visual-to-text naming through the paradigm of answering-by-search in a two-stage computational framework, which is composed out of instance search (IS) and similar question ranking (QR). In IS, names of the instances are inferred from similar visual examples searched through a million-scale image dataset. For recalling instances of non-planar and non-rigid shapes, spatial configurations that emphasize topology consistency while allowing for local variations in matches have been incorporated. In QR, the candidate names of the instance are statistically identified from search results and directly utilized to retrieve similar questions from community-contributed QA (cQA) archives. By parsing questions into syntactic trees, a fuzzy matching between the inquirer's question and cQA questions is performed to locate answers and recommend related questions to the inquirer. The proposed framework is evaluated on a wide range of visual instances (e.g., fashion, art, food, pet, logo, and landmark) over various QA categories (e.g., factoid, definition, how-to, and opinion). Wei Zhang 0031, Chong-Wah Ngo |
ACM Multimedia | 1 |
| 2012 | FashionAsk: pushing community answers to your fingertipsabstractWe demonstrate a multimedia-based question-answering system, named FashionAsk, by allowing users to ask questions referring to pictures snapped by mobile devices. Specifically, instead of asking verbose questions to depict visual instances, direct pictures are provided as part of questions. To answer these multi-modal questions, FashionAsk performs a large-scale instance search to infer the names of instances, and then matches with similar questions from community-contributed QA websites as answers. The demonstration is conducted on a million-scale dataset of Web images and QA pairs in the domain of fashion products. Asking a multimedia question through FashionAsk can take as short as five seconds to retrieve the candidate answer as well as suggested questions. Wei Zhang 0031, Chong-Wah Ngo |
ACM Multimedia | 1 |
| 2012 | Detecting image forgeries using metrology
Lin Wu 0001, Xiaochun Cao, Wei Zhang 0031, Yang Wang 0023 |
Mach. Vis. Appl. | 3 |
| 2010 | Detecting and extracting the photo composites using planar homography and graph cutabstractWith the advancement of photo and video editing tools, it has become fairly easy to tamper with photos and videos. One common way is to insert visually plausible composites into target images and videos. In this paper, we propose an automatic fake region detection method based on the planar homography constraint, and an automatic extraction method using graph cut with online feature/parameter selection. Two steps are taken in our method: 1) the targeting step, and 2) the segmentation step. First, the fake region is located roughly by enforcing the planar homography constraint. Second, the fake object is segmented via graph cut with the initialization given by the targeting step. To achieve an automatic segmentation, the optimal features and parameters for graph cut are dynamically selected via the proposed online feature/parameter selection. Performance of this method is evaluated on both semisimulated and real images. Our method works efficiently on images as long as there are regions satisfying the planar homography constraint, including image pairs captured by the approximately cocentered cameras, image pairs photographing planar or distant scenes, and a single image with duplications. Wei Zhang 0031, Xiaochun Cao, Yanling Qu, Yuexian Hou, Handong Zhao, Chenyang Zhang 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2009 | Detecting photographic composites using two-view geometrical constraintsabstractIn this work, we describe a new technique for detecting image composites by enforcing two-view geometrical constrains: H and F constraints on image pairs, where H denotes the planar homography matrix and F the fundamental matrix. Our approach detects fake regions efficiently on pictures taken at the same scene but with different camera configurations. Performance of this approach is demonstrated on real image pairs with visually plausible composites. Wei Zhang 0031, Xiaochun Cao, Zhiyong Feng 0002, Jiawan Zhang |
ICME | 1 |
| 2009 | Detecting photographic composites using shadowsabstractImage compositing technology has become popular for tampering with digital photographies. We describe how such composites can be detected by enforcing the geometric and photometric constraints from shadows. In particular, we explore (i) the imaged shadow relations that are modeled by the planar homology, and (ii) the color characteristics of the shadows measured by the shadow matte. Our approach efficiently extracts these constraints from a single image and makes use of them for the digital forgery detection. Experimental results on visually plausible images demonstrate the performance of the proposed method. Wei Zhang 0031, Xiaochun Cao, Jiawan Zhang, Jigui Zhu |
ICME | 1 |
| 2005 | A Hybrid GMM and Codebook Mapping Method for Spectral Conversion
Yongguo Kang, Zhiwei Shuang, Jianhua Tao 0001, Wei Zhang 0031, Bo Xu 0002 |
ACII | 4 |