EDBT 2026 Demo / reviewers in the wild / expert
Feng Gao 0014
dblp:10/2674-14
· DBLP profile ↗
29ranked-venue papers
2as first author
18since 2021 · last 2026
0009-0006-1843-3180ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 26 · 2 first-author · 16 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Computer networks · 4 · 4 since 2021Systems, architecture and hardware · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RTGSR: Real-Time Game Content Super-Resolution via Compressed-Domain Coding PriorsabstractThe rapid growth of cloud gaming and game streaming has led to a substantial increase in the volume of game content data. To ensure real-time delivery of cloud game content, a common strategy is to downsample and compress the game content before transmission, reducing both data size and bandwidth requirements. However, this approach presents considerable obstacles for super-resolution (SR) networks at the receiver side. In particular, the degraded quality of compressed video streams, combined with the stringent demand for real-time processing, poses major challenges for practical SR applications. In this paper, we propose a novel real-time super-resolution framework that works directly in the compressed domain by exploiting coding-domain priors. Specifically, we propose an extremely lightweight U-Net architecture that leverages prediction maps and residuals as its primary guidance signals. Furthermore, we incorporate the partition map into a Pixel Adaptive Convolution (PAC) module, allowing the convolution kernels to adapt to different regions in the decoded frame. The resulting deep features are then fused with those from the U-Net backbone through an attention block. Finally, we present an enhanced re-parameterization block designed to better model edge features, leading to notable gains in both the objective metrics and subjective visual quality of the reconstructions. Extensive experiments demonstrate that the proposed method consistently outperforms existing real-time approaches on compressed game video content, achieving superior performance in both quality and efficiency. Qizhe Wang, Jiaqi Zhang 0007, Yanchen Zhao, Haoqing Yu, Feng Gao 0014, Siwei Ma 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | i3DV: Intelligent 3D Volumetric Video Coding Standard and Platform
Luyang Tang, Feng Gao 0014, Ronggang Wang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | Space-Time Gaussian Surfels for High-Fidelity Dynamic Objects Segmentation and RepresentationabstractWe introduce ST-ObjGS, a method using Space-time Gaussian surfels for accurate object segmentation within 4D representations. Our approach addresses the limitations of current Gaussian-based methods, which primarily focus on static 3D scene understanding and struggle with geometrically accurate object segmentation in complex dynamic scenes. To ensure robust object-level segmentation, we first integrate Grounded SAM 2, which enables text prompt-based object selection and tracking. We then learn a set of Gaussian surfels for object geometry representation and employ a marginal 1D Gaussian for dynamic modeling at each timestamp. To improve geometric quality when modeling surfaces, we use depth and surface normal for geometric regularization. Furthermore, to address continuity and flickering issues in complex scenes, we implement dynamic-aware regularization to maintain temporal consistency. This approach allows us to capture object motion and morphing over time while maintaining spatial coherence. To the best of our knowledge, ST-ObjGS is the first self-supervised approach using Space-time Gaussian surfels for consistent segmentation of dynamic 3D objects in real-world scenes. Extensive experiments on standard benchmarks including PKU-DyMVHumans, Plenoptic Video, Google Immersive, and CMU Panoptic datasets demonstrate that ST-ObjGS produces more precise object masks than its Gaussian-based counterparts and significantly outperforms supervised single-view baselines. Xiaoyun Zheng, Liwei Liao, Feng Gao 0014, Shiqi Wang 0001, Ronggang Wang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | High-Fidelity and Lip-Synced Talking Face Synthesis via Landmark-Based Diffusion ModelabstractAudio-driven talking face video generation has attracted increasing attention due to its huge industrial potential. Some previous methods focus on learning a direct mapping from audio to visual content. Despite progress, they often struggle with the ambiguity of the mapping process, leading to flawed results. An alternative strategy involves facial structural representations (e.g., facial landmarks) as intermediaries. This multi-stage approach better preserves the appearance details but suffers from error accumulation due to the independent optimization of different stages. Moreover, most previous methods rely on generative adversarial networks, prone to training instability and mode collapse. To address these challenges, our study proposes a novel landmark-based diffusion model for talking face generation, which leverages facial landmarks as intermediate representations while enabling end-to-end optimization. Specifically, we first establish the less ambiguous mapping from audio to landmark motion of lip and jaw. Then, we introduce an innovative conditioning module called TalkFormer to align the synthesized motion with the motion represented by landmarks via differentiable cross-attention, which enables end-to-end optimization for improved lip synchronization. Besides, TalkFormer employs implicit feature warping to align the reference image features with the target motion for preserving more appearance details. Extensive experiments demonstrate that our approach can synthesize high-fidelity and lip-synced talking face videos, preserving more subject appearance details from the reference image. Weizhi Zhong, Junfan Lin, Peixin Chen, Feng Gao 0014, Liang Lin 0004, Guanbin Li |
IEEE Trans. Image Process. | 4 |
| 2026 | Toward Top-Down Reasoning: An Explainable Multi-Agent Approach for Visual Question AnsweringabstractRecent methods to enhance Vision-Language Models (VLMs) for Visual Question Answering (VQA) have focused on strengthening their inference capabilities, enabling them to tackle VQA tasks independently rather than merely as aids to Large Language Models (LLMs). However, these approaches often ignore the rich commonsense knowledge inside the given VQA image sampled from the real world, limiting the full potential of VLMs. Inspired by the human top-down reasoning process, i.e., systematically exploring relevant issues to derive a comprehensive answer, this work introduces a novel, explainable multi-agent collaboration framework by leveraging the expansive knowledge of LLMs to enhance the capabilities of VLMs themselves. Our framework comprises three agents, i.e.,Responder,Seeker, andIntegrator, to collaboratively answer the given VQA question by seeking its relevant issues and generating the final answer in such a top-down reasoning process. The VLM-basedResponderagent generates the answer candidates for the question and responds to other relevant issues. TheSeekeragent, primarily based on LLM, identifies relevant issues related to the question to inform theResponderagent and constructs a Multi-View Knowledge Base (MVKB) for the given visual scene by leveraging the build-in world knowledge of LLM. TheIntegratoragent combines knowledge from theSeekeragent and theResponderagent to produce the final VQA answer. Extensive and comprehensive evaluations on diverse VQA datasets with a variety of VLMs demonstrate the superior performance and interpretability of our framework over the baseline method, e.g., 5.7% improvement on VQA-RAD and 5.2% on Winoground in the zero-shot setting without extra training cost. Zeqing Wang, Wentao Wan 0001, Qiqing Lao, Runmeng Chen, Minjie Lang, Xiao Wang 0002, Feng Gao 0014, Keze Wang, Liang Lin 0004 |
IEEE Trans. Multim. | 7 |
| 2026 | MLICv2: Enhanced Multi-Reference Entropy Modeling for Learned Image CompressionabstractRecent advances in Learned Image Compression (LIC) have achieved remarkable performance improvements over traditional codecs. Notably, the MLIC series—LICs equipped with multi-reference entropy models—have substantially surpassed conventional image codecs such as Versatile Video Coding (VVC) Intra. However, existing MLIC variants suffer from several limitations: performance degradation at high bit-rates due to insufficient transform capacity, suboptimal entropy modeling that fails to capture global correlations in initial slices, and lack of adaptive channel importance modeling. In this article, we propose MLICv2 and MLICv2 \({}^{+}\) , enhanced successors that systematically address these limitations through improved transform design, advanced entropy modeling, and exploration of the potential of instance-specific optimization. For transform enhancement, we introduce a lightweight token mixing block inspired by the MetaFormer architecture, which effectively mitigates high-bit-rate performance degradation while maintaining computational efficiency. For entropy modeling improvements, we propose hyperprior-guided global correlation prediction to extract global context even in the initial slice of latent representation, complemented by a channel reweighting module that dynamically emphasizes informative channels. We further explore enhanced positional embedding and guided selective compression strategies for superior context modeling. Additionally, we apply the Stochastic Gumbel Annealing (SGA) to demonstrate the potential for further performance improvements through input-specific optimization. Extensive experiments demonstrate that MLICv2 and MLICv2 \({}^{+}\) achieve state-of-the-art results, reducing Bjøntegaard-Delta Rate by 16.54%, 21.61%, 16.05% and 20.46%, 24.35%, 19.14% on Kodak, Tecnick, and CLIC Pro Val datasets, respectively, compared to VTM-17.0 Intra. Wei Jiang 0031, Yongqi Zhai, Feng Gao 0014, Ronggang Wang |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2026 | Exploring Talking Head Models with Adjacent Frame Prior for Speech-Preserving Facial Expression ManipulationabstractSpeech-Preserving Facial Expression Manipulation (SPFEM) is an innovative technique aimed at altering facial expressions in images and videos while retaining the original mouth movements. Despite advancements, SPFEM still struggles with accurate lip synchronization due to the complex interplay between facial expressions and mouth shapes. Capitalizing on the advanced capabilities of Audio-Driven Talking Head Generation (AD-THG) models in synthesizing precise lip movements, our research introduces a novel integration of these models with SPFEM. We present a new framework, Talking Head Facial Expression Manipulation (THFEM), which utilizes AD-THG models to generate frames with accurately synchronized lip movements from audio inputs and SPFEM-altered images. However, increasing the number of frames generated by AD-THG models tends to compromise the realism and expression fidelity of the images. To counter this, we develop an adjacent frame learning strategy that finetunes AD-THG models to predict sequences of consecutive frames. This strategy enables the models to incorporate information from neighboring frames, significantly improving image quality during testing. Our extensive experimental evaluations demonstrate that this framework effectively preserves mouth shapes during expression manipulations, highlighting the substantial benefits of integrating AD-THG with SPFEM. Zhenxuan Lu, Zhihua Xu, Zhijing Yang, Feng Gao 0014, Yongyi Lu, Keze Wang, Tianshui Chen |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2025 | Aligning Human Motion Generation with Human PerceptionsabstractHuman motion generation is a critical task with a wide spectrum of applications. Achieving high realism in generated motions requires naturalness, smoothness, and plausibility. However, current evaluation metrics often rely on simple heuristics or distribution distances and do not align well with human perceptions. In this work, we propose a data-driven approach to bridge this gap by introducing a large-scale human perceptual evaluation dataset, MotionPercept, and a human motion critic model, MotionCritic, that capture human perceptual preferences. Our critic model offers a more accurate metric for assessing motion quality and could be readily integrated into the motion generation pipeline to enhance generation quality. Extensive experiments demonstrate the effectiveness of our approach in both evaluating and improving the quality of generated human motions by aligning with human perceptions. Code and data are publicly available at https://motioncritic.github.io/. Haoru Wang, Wentao Zhu 0004, Luyi Miao, Yishu Xu, Feng Gao 0014, Qi Tian 0001, Yizhou Wang 0001 |
ICLR | 5 |
| 2025 | MLIC++: Linear Complexity Multi-Reference Entropy Modeling for Learned Image CompressionabstractThe latent representation in learned image compression encompasses channel-wise, local spatial, and global spatial correlations, which are essential for the entropy model to capture for conditional entropy minimization. Efficiently capturing these contexts within a single entropy model, especially in high-resolution image coding, presents a challenge due to the computational complexity of existing global context modules. To address this challenge, we propose the Linear Complexity Multi-Reference Entropy Model (MEM \({}^{++}\) ). Specifically, the latent representation is partitioned into multiple slices. For channel-wise contexts, previously compressed slices serve as the context for compressing a particular slice. For local contexts, we introduce a shifted-window-based checkerboard attention module. This module ensures linear complexity without sacrificing performance. For global contexts, we propose a linear complexity attention mechanism. It captures global correlations by decomposing the softmax operation, enabling the implicit computation of attention maps from previously decoded slices. Using MEM++ as the entropy model, we develop the image compression method MLIC \({}^{++}\) . Extensive experimental results demonstrate that MLIC \({}^{++}\) achieves state-of-the-art performance, reducing BD-rate by \(13.39\%\) on the Kodak dataset compared to VTM-17.0 in Peak Signal-to-Noise Ratio (PSNR). Furthermore, MLIC \({}^{++}\) exhibits linear computational complexity and memory consumption with resolution, making it highly suitable for high-resolution image coding. Code and pre-trained models are available at https://github.com/JiangWeibeta/MLIC . Wei Jiang 0031, Yongqi Zhai, Feng Gao 0014, Ronggang Wang |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2025 | Adaptive Prediction Structure for Learned Video CompressionabstractLearned video compression has developed rapidly and shown competitive rate-distortion performance compared with the latest traditional video coding standard H.266 (VVC). However, existing works were restricted to fixed prediction direction and GoP size. The inflexibility on prediction structure hinders learned video compression towards optimal compression efficiency in diverse motion scenarios. In this article, we propose to advance learned video compression with adaptive prediction structure decision. Specifically, we propose a unified compression framework that supports both forward prediction and bi-directional prediction. The framework can flexibly switch to different prediction direction to achieve better prediction performance. Meanwhile, we propose a low-complexity prediction structure decision algorithm, where prediction direction and GoP size are adaptively determined based on motion complexity to achieve optimal compression efficiency. Experimental results demonstrate that the proposed unified framework with adaptive decision algorithm improves compression efficiency of pure forward prediction-based or bi-directional prediction-based framework with neglectable ( \(0.9\%\) ) encoding time increment. Meanwhile, it achieves comparable compression performance with VVC and recent learned video coding methods. Yongqi Zhai, Wei Jiang 0031, Chunhui Yang, Feng Gao 0014, Ronggang Wang |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2024 | PKU-DyMVHumans: A Multi-View Video Benchmark for High-Fidelity Dynamic Human ModelingabstractHigh-quality human reconstruction and photo-realistic rendering of a dynamic scene is a long-standing problem in computer vision and graphics. Despite considerable ef-forts invested in developing various capture systems and re-construction algorithms, recent advancements still struggle with loose or oversized clothing and overly complex poses. In part, this is due to the challenges of acquiring high-quality human datasets. To facilitate the development of these fields, in this paper, we present PKU-DyMVHumans, a versatile human-centric dataset for high-fidelity reconstruction and rendering of dynamic human scenarios from dense multi-view videos. It comprises 8.2 million frames captured by more than 56 synchronized cameras across diverse scenarios. These sequences comprise 32 human subjects across 45 different scenarios, each with a high-detailed appearance and realistic human motion. Inspired by recent advancements in neural radiance field (NeRF)-based scene representations, we carefully set up an off-the-shelf framework that is easy to provide those state-of-the-art NeRF-based implementations and benchmark on PKU-DyMVHumans dataset. It is paving the way for various applications like fine-grained foreground/background de-composition, high-quality human reconstruction and photo-realistic novel view synthesis of a dynamic scene. Exten-sive studies are performed on the benchmark, demonstrating new observations and challenges that emerge from using such high-fidelity dynamic data. The project page and data is available at: https://pku-dymvhumans.github.io. Xiaoyun Zheng, Liwei Liao, Jianbo Jiao, Feng Gao 0014, Shiqi Wang 0001, Ronggang Wang |
CVPR | 6 |
| 2024 | Human Motion Generation: A SurveyabstractHuman motion generation aims to generate natural human pose sequences and shows immense potential for real-world applications. Substantial progress has been made recently in motion data collection technologies and generation methods, laying the foundation for increasing interest in human motion generation. Most research within this field focuses on generating human motions based on conditional signals, such as text, audio, and scene contexts. While significant advancements have been made in recent years, the task continues to pose challenges due to the intricate nature of human motion and its implicit relationship with conditional signals. In this survey, we present a comprehensive literature review of human motion generation, which, to the best of our knowledge, is the first of its kind in this field. We begin by introducing the background of human motion and generative models, followed by an examination of representative methods for three mainstream sub-tasks: text-conditioned, audio-conditioned, and scene-conditioned human motion generation. Additionally, we provide an overview of common datasets and evaluation metrics. Lastly, we discuss open problems and outline potential future research directions. We hope that this survey could provide the community with a comprehensive glimpse of this rapidly evolving field and inspire novel ideas that address the outstanding challenges. Wentao Zhu 0004, Xiaoxuan Ma 0001, Dongwoo Ro, Hai Ci, Jinlu Zhang 0001, Jiaxin Shi, Feng Gao 0014, Qi Tian 0001, Yizhou Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2024 | Surface-SOS: Self-Supervised Object Segmentation via Neural Surface RepresentationabstractSelf-supervised Object Segmentation (SOS) aims to segment objects without any annotations. Under conditions of multi-camera inputs, the structural, textural and geometrical consistency among each view can be leveraged to achieve fine-grained object segmentation. To make better use of the above information, we propose Surface representation based Self-supervised Object Segmentation (Surface-SOS), a new framework to segment objects for each view by 3D surface representation from multi-view images of a scene. To model high-quality geometry surfaces for complex scenes, we design a novel scene representation scheme, which decomposes the scene into two complementary neural representation modules respectively with a Signed Distance Function (SDF). Moreover, Surface-SOS is able to refine single-view segmentation with multi-view unlabeled images, by introducing coarse segmentation masks as additional input. To the best of our knowledge, Surface-SOS is the first self-supervised approach that leverages neural surface representation to break the dependence on large amounts of annotated data and strong constraints. These constraints typically involve observing target objects against a static background or relying on temporal supervision in videos. Extensive experiments on standard benchmarks including LLFF, CO3D, BlendedMVS, TUM and several real-world scenes show that Surface-SOS always yields finer object masks than its NeRF-based counterparts and surpasses supervised single-view baselines remarkably. Code is available at: https://github.com/zhengxyun/Surface-SOS. Xiaoyun Zheng, Liwei Liao, Jianbo Jiao, Feng Gao 0014, Ronggang Wang |
IEEE Trans. Image Process. | 4 |
| 2024 | LLIC: Large Receptive Field Transform Coding With Adaptive Weights for Learned Image CompressionabstractThe effective receptive field (ERF) plays an important role in transform coding, which determines how much redundancy can be removed during transform and how many spatial priors can be utilized to synthesize textures during inverse transform. Existing methods rely on stacks of small kernels, whose ERFs remain insufficiently large, or heavy non-local attention mechanisms, which limit the potential of high-resolution image coding. To tackle this issue, we propose Large Receptive Field Transform Coding with Adaptive Weights for Learned Image Compression (LLIC). Specifically, for thefirsttime in the learned image compression community, we introducea fewlarge kernel-based depth-wise convolutions to reduce more redundancy while maintaining modest complexity. Due to the wide range of image diversity, we further propose a mechanism to augment convolution adaptability through the self-conditioned generation of weights. The large kernels cooperate with non-linear embedding and gate mechanisms for better expressiveness and lighter point-wise interactions. Our investigation extends to refined training methods that unlock the full potential of these large kernels. Moreover, to promote more dynamic inter-channel interactions, we introduce an adaptive channel-wise bit allocation strategy that autonomously generates channel importance factors in a self-conditioned manner. To demonstrate the effectiveness of the proposed transform coding, we align the entropy model to compare with existing transform methods and obtain models LLIC-STF, LLIC-ELIC, and LLIC-TCM. Extensive experiments demonstrate that our proposed LLIC models have significant improvements over the corresponding baselines and reduce the BD-Rate by$9.49\%, 9.47\%,\;\text{and}\; 10.94\%$on Kodak over VTM-17.0 Intra, respectively. Our LLIC models achieve state-of-the-art performances and better trade-offs between performance and complexity. Wei Jiang 0031, Peirong Ning, Yongqi Zhai, Feng Gao 0014, Ronggang Wang |
IEEE Trans. Multim. | 5 |
| 2023 | CL-MVSNet: Unsupervised Multi-view Stereo with Dual-level Contrastive LearningabstractUnsupervised Multi-View Stereo (MVS) methods have achieved promising progress recently. However, previous methods primarily depend on the photometric consistency assumption, which may suffer from two limitations: indistinguishable regions and view-dependent effects, e.g., low-textured areas and reflections. To address these issues, in this paper, we propose a new dual-level contrastive learning approach, named CL-MVSNet. Specifically, our model integrates two contrastive branches into an unsupervised MVS framework to construct additional supervisory signals. On the one hand, we present an image-level contrastive branch to guide the model to acquire more context awareness, thus leading to more complete depth estimation in indistinguishable regions. On the other hand, we exploit a scene-level contrastive branch to boost the representation ability, improving robustness to view-dependent effects. Moreover, to recover more accurate 3D geometry, we introduce an ℒ0.5 photometric consistency loss, which encourages the model to focus more on accurate points while mitigating the gradient penalty of undesirable ones. Extensive experiments on DTU and Tanks&Temples benchmarks demonstrate that our approach achieves state-of-the-art performance among all end-to-end unsupervised MVS frameworks and outperforms its supervised counterpart by a considerable margin without fine-tuning. Kaiqiang Xiong, Tianxing Feng, Jianbo Jiao, Feng Gao 0014, Ronggang Wang |
ICCV | 6 |
| 2023 | MLIC: Multi-Reference Entropy Model for Learned Image CompressionabstractRecently, learned image compression has achieved remarkable performance. The entropy model, which estimates the distribution of the latent representation, plays a crucial role in boosting rate-distortion performance. However, most entropy models only capture correlations in one dimension, while the latent representation contains channel-wise, local spatial, and global spatial correlations. To tackle this issue, we propose the Multi-Reference Entropy Model (MEM) and the advanced version, MEM+. These models capture the different types of correlations present in latent representation. Specifically, we first divide the latent representation into slices. When decoding the current slice, we use previously decoded slices as context and employ the attention map of the previously decoded slice to predict global correlations in the current slice. To capture local contexts, we introduce two enhanced checkerboard context capturing techniques that avoids performance degradation. Based on MEM and MEM+, we propose image compression models MLIC and MLIC+. Extensive experimental evaluations demonstrate that our MLIC and MLIC+ models achieve state-of-the-art performance, reducing BD-rate by 8.05% and 11.39% on the Kodak dataset compared to VTM-17.0 when measured in PSNR. Wei Jiang 0031, Yongqi Zhai, Peirong Ning, Feng Gao 0014, Ronggang Wang |
ACM Multimedia | 5 |
| 2022 | Lesion-Aware Dynamic Kernel for Polyp Segmentation
Ruifei Zhang, Peiwen Lai, Feng Gao 0014, Xiao-Jian Wu, Guanbin Li |
MICCAI (3) | 5 |
| 2021 | Towards Large-Scale Object Instance Search: A Multi-Block N-Ary TrieabstractObject instance search is a challenging task with a wide range of applications, but the fast search with high accuracy has not been well solved yet. In this paper, we investigate the object instance search from a new perspective in terms of joint precision and computational cost optimization, and propose a novel index structure i.e., Multi-Block N-ary Trie (MBNT) to accelerate the exact r-neighbor search in the Hamming space. Comprehensive studies are first carried out to reveal the performance of exact and approximate nearest neighbor (NN) algorithms for object instance search. An interesting finding that the exact search is more promising for very compact binary codes (e.g., 64-bit and 128-bit) is analyzed. Along this vein, we introduce a Trie structure, i.e., MBNT, which is specifically designed for improving the exact NN search performance in the context of large-scale object instance search. To index the binary codes, a subset of continuous bits of a binary string, denoted as a block, is regarded as an atomic indexing element. As such, the problem of lookup misses can be addressed. Theoretical analyses are also provided to show that our MBNT scheme can incur less computational cost than other hash table-based methods. Extensive experimental results on the 100M dataset have demonstrated that our method achieves faster search speed while maintaining the promising search precision towards large-scale object instance search. Mangui Liang, Feng Gao 0014, Yi-Cheng Huang, Xinfeng Zhang 0001, Ling-Yu Duan |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2019 | Learning to Remove Reflections for Text ImagesabstractText images taken behind a piece of glass in the wild are largely contaminated by reflections. Directly applying existing reflection removal methods on text images with reflections cannot recover clear and correct text contents due to the ignorance of special characteristics of texts. This paper proposes a stacked framework to solve the text image reflection removal problem by specifically considering the regional properties of reflection and embedding the specific text priors into the estimation process in a unified manner. Experiment results on a newly collected dataset demonstrate that the proposed method outperforms state-of-the-art methods in recovering visually pleasant reflection-free images and recognizable text features. Ce Wang 0007, Renjie Wan, Feng Gao 0014, Boxin Shi, Ling-Yu Duan |
ICME | 3 |
| 2019 | Codebook-Free Compact Descriptor for Scalable Visual SearchabstractThe MPEG compact descriptors for visual search (CDVS) is a standard toward image matching and retrieval. To achieve high retrieval accuracy over a large scale image/video dataset, recent research efforts have demonstrated that employing extremely high-dimensional descriptors such as the Fisher vector (FV) and the vector of locally aggregated descriptors (VLAD) can yield good performance. Since the FV (or VLAD) possesses high discriminability but small visual vocabulary, it has been adopted by CDVS to construct a global compact descriptor. In this paper, we study the development of global compact descriptors in the completed CDVS standard and the emerging compact descriptors for video analysis (CDVA) standard, in which we formulate the FV (or VLAD) compression as a resource-constrained optimization problem. Accordingly, we propose a codebook-free aggregation method via dual selection to generate a global compact visual descriptor, which supports fast and accurate feature matching free of large visual codebooks, fulfilling the low memory requirement of mobile visual search at significantly reduced latency. Specifically, we investigate both sample-specific Gaussian component redundancy and bit dependency within a binary aggregated descriptor to produce compact binary codes. Our technique contributes to the scalable compressed Fisher vector (SCFV) adopted by the CDVS standard. Moreover, the SCFV descriptor is currently serving as the frame-level hand-crafted video feature, which inspires the inheritance of CDVS descriptors for the emerging CDVA standard. Furthermore, we investigate the positive complementary effect of our standard compliant compact descriptor and deep learning based features extracted from convolutional neural networks with significant mean average precision gains. Extensive evaluation over benchmark databases shows the significant merits of the codebook-free binary codes for scalable visual search. Yuwei Wu 0001, Feng Gao 0014, Jie Lin 0001, Vijay Chandrasekhar 0001, Junsong Yuan 0001, Ling-Yu Duan |
IEEE Trans. Multim. | 2 |
| 2018 | ChipGAN: A Generative Adversarial Network for Chinese Ink Wash Painting Style TransferabstractStyle transfer has been successfully applied on photos to generate realistic western paintings. However, because of the inherently different painting techniques adopted by Chinese and western paintings, directly applying existing methods cannot generate satisfactory results for Chinese ink wash painting style transfer. This paper proposes ChipGAN, an end-to-end Generative Adversarial Network based architecture for photo to Chinese ink wash painting style transfer. The core modules of ChipGAN enforce three constraints -- voids, brush strokes, and ink wash tone and diffusion -- to address three key techniques commonly adopted in Chinese ink wash painting. We conduct stylization perceptual study to score the similarity of generated paintings to real paintings by consulting with professional artists based on the newly built Chinese ink wash photo and image dataset. The advantages in visual quality compared with state-of-the-art networks and high stylization perceptual study scores show the effectiveness of the proposed method. Feng Gao 0014, Daiqian Ma, Boxin Shi, Ling-Yu Duan |
ACM Multimedia | 2 |
| 2018 | Depth Structure Preserving Scene Image GenerationabstractKey to automatically generate natural scene images is to properly arrange amongst various spatial elements, especially in the depth cue. To this end, we introduce a novel depth structure preserving scene image generation network (DSP-GAN), which favors a hierarchical architecture, for the purpose of depth structure preserving scene image generation. The main trunk of the proposed infrastructure is built upon a Hawkes point process that models high-order spatial dependency between different depth layers. Within each layer generative adversarial sub-networks are trained collaboratively to generate realistic scene components, conditioned on the layer information produced by the point process. We experiment our model on annotated natural scene images collected from SUN dataset and demonstrate that our models are capable of generating depth-realistic natural scene image. Wendong Zhang 0002, Feng Gao 0014, Bingbing Ni, Ling-Yu Duan, Yichao Yan, Jingwei Xu 0005, Xiaokang Yang 0001 |
ACM Multimedia | 2 |
| 2018 | Group-Sensitive Triplet Embedding for Vehicle ReidentificationabstractThe widespread use of surveillance cameras toward smart and safe cities poses the critical but challenging problem of vehicle reidentification (Re-ID). The state-of-the-art research work performed vehicle Re-ID relying on deep metric learning with a triplet network. However, most existing methods basically ignore the impact of intraclass variance-incorporated embedding on the performance of vehicle reidentification, in which robust fine-grained features for large-scale vehicle Re-ID have not been fully studied. In this paper, we propose a deep metric learning method, group-sensitive-triplet embedding (GS-TRE), to recognize and retrieve vehicles, in which intraclass variance is elegantly modeled by incorporating an intermediate representation “group” between samples and each individual vehicle in the triplet network learning. To capture the intraclass variance attributes of each individual vehicle, we utilize an online grouping method to partition samples within each vehicle ID into a few groups, and build up the triplet samples at multiple granularities across different vehicle IDs as well as different groups within the same vehicle ID to learn fine-grained features. In particular, we construct a large-scale vehicle database “PKU-Vehicle,” consisting of 10 million vehicle images captured by different surveillance cameras in several cities, to evaluate the vehicle Re-ID performance in real-world video surveillance applications. Extensive experiments over benchmark datasets VehicleID, VeRI, and CompCar have shown that the proposed GS-TRE significantly outperforms the state-of-the-art approaches for vehicle Re-ID. Yihang Lou, Feng Gao 0014, Shiqi Wang 0001, Yuwei Wu 0001, Ling-Yu Duan |
IEEE Trans. Multim. | 3 |
| 2018 | Data-Driven Lightweight Interest Point Selection for Large-Scale Visual SearchabstractWith the explosive increase of images and videos, visual analysis has become an essential technique in dealing with the big visual data, which utilizes the visual feature descriptors to search or recognize the images or frames with target objects or events. Subject to the constraints of resources (e.g., memory, bandwidth, storage, etc.), interest point selection is crucial to generate robust compact descriptors for high-efficiency visual analysis by selecting and aggregating the most discriminative local feature descriptors, which has been demonstrated in the state-of-the-art low bit rate visual search works. In this paper, we propose a data-driven lightweight interest point selection approach to significantly improve the performance of visual search, while ameliorating the efficiency of extracting feature descriptors. Comprehensive experimental results over benchmarks have shown that the proposed interest point selection algorithm has significantly improved image matching and retrieval performance in the completed MPEG Compact Descriptors for Visual Search (CDVS) standard as well as the emerging MPEG Compact Descriptors for Video Analytics (CDVA) standard, say 20% mAP gain by data-driven selection against random selection of interest points. In particular, the presented data-driven interest point selection has been adopted by MPEG-CDVS and MPEG-CDVA as a normative technique to improve the aggregation of handcrafted features, which has contributed to the combination of handcrafted features and deep learning (CNN) features as well. Feng Gao 0014, Xinfeng Zhang 0001, Yong Luo 0002, Xiaoming Li 0001, Ling-Yu Duan |
IEEE Trans. Multim. | 1 |
| 2017 | Incorporating intra-class variance to fine-grained visual recognitionabstractFine-grained visual recognition aims to capture discriminative characteristics amongst visually similar categories. The state-of-the-art research work has significantly improved the fine-grained recognition performance by deep metric learning using triplet network. However, the impact of intra-category variance on the performance of recognition and robust feature representation has not been well studied. In this paper, we propose to leverage intra-class variance in metric learning of triplet network to improve the performance of fine-grained recognition. Through partitioning training images within each category into a few groups, we form the triplet samples across different categories as well as different groups, which is called Group Sensitive TRiplet Sampling (GS-TRS). Accordingly, the triplet loss function is strengthened by incorporating intra-class variance with GS-TRS, which may contribute to the optimization objective of triplet network. Extensive experiments over benchmark datasets CompCar and VehicleID show that the proposed GS-TRS has significantly outperformed state-of-the-art approaches in both classification and retrieval tasks. Yan Em, Feng Gao 0014, Yihang Lou, Shiqi Wang 0001, Tiejun Huang 0001, Ling-Yu Duan |
ICME | 2 |
| 2017 | Improving object detection with region similarity learningabstractObject detection aims to identify instances of semantic objects of a certain class in images or videos. The success of state-of-the-art approaches is attributed to the significant progress of object proposal and convolutional neural networks (CNNs). Most promising detectors involve multi-task learning with an optimization objective of softmax loss and regression loss. The first is for multi-class categorization, while the latter is for improving localization accuracy. However, few of them attempt to further investigate the hardness of distinguishing different sorts of distracting background regions (i.e., negatives) from true object regions (i.e., positives). To improve the performance of classifying positive object regions vs. a variety of negative background regions, we propose to incorporate triplet embedding into learning objective. The triplet units are formed by assigning each negative region to a meaningful object class and establishing class-specific negatives, followed by triplets construction. Over the benchmark PASCAL VOC 2007, the proposed triplet embedding has improved the performance of well-known Fas-tRCNN model with a mAP gain of 2.1%. In particular, the state-of-the-art approach OHEM can benefit from the triplet embedding and has achieved a mAP improvement of 1.2%. Feng Gao 0014, Yihang Lou, Shiqi Wang 0001, Tiejun Huang 0001, Ling-Yu Duan |
ICME | 1 |
| 2017 | From Part to Whole: Who is Behind the Painting?abstractCompared with normal modalities, the representations of paintings are much more complex due to its large intra-class and small inter-class variation. This poses more difficulties in the task of authorship identification. In this paper, we propose a multi-task multi-range (MTMR) representation framework and try to resolve this issue in two ways. First, we investigate how to improve the representation through multi-task learning. Specifically, we attempt to optimize authorship identification with subtly correlated identification tasks such as style, genre and date. Second, in order to make the representation more comprehensive and reduce the information loss from image scaling, we propose a multi-range structure which is composed of local, regional and global representations. Experiments on the two most representative large-scale painting datasets, Rijksmuseum Challenge and Wikiart, have shown that our method significantly outperforms the existing methods. To give better understanding and provide more effective predictions, we utilize random forest as the feature ranking method to analyze the importance of different features and apply external knowledge matching to further examine the predictions. Moreover, the framework's effects of identifying the authorship are visualized on the paintings' artist-characteristic regions and t-SNE is further applied to perform artist-based cluster analysis. Extensive validation has demonstrated that the proposed framework yields superior performance in the chanllenging task of painting authorship identification. Daiqian Ma, Feng Gao 0014, Yihang Lou, Shiqi Wang 0001, Tiejun Huang 0001, Ling-Yu Duan |
ACM Multimedia | 2 |
| 2015 | A Low Complexity Interest Point DetectorabstractInterest point detection is a fundamental approach to feature extraction in computer vision tasks. To handle the scale invariance, interest points usually work on the scale-space representation of an image. In this letter, we propose a novel block-wise scale-space representation to significantly reduce the computational complexity of an interest point detector. Laplacian of Gaussian (LoG) filtering is applied to implement the block-wise scale-space representation. Extensive comparison experiments have shown the block-wise scale-space representation enables the efficient and effective implementation of an interest point detector in terms of memory and time complexity reduction, as well as promising performance in visual search. Jie Chen 0006, Ling-Yu Duan, Feng Gao 0014, Jianfei Cai 0001, Alex Chichung Kot, Tiejun Huang 0001 |
IEEE Signal Process. Lett. | 3 |
| 2013 | Compact descriptors for mobile visual search and MPEG CDVS standardizationabstractIn this paper, we present the state-of-the-art compact descriptors for mobile visual search. In particular, we introduce our MPEG contributions in global descriptor aggregation and local descriptor compression, which have been adopted by the ongoing MPEG standardization of compact descriptor for visual search (CDVS). Standardization progress will be introduced. Other issues including visual object databases and MPEG CDVS impact on visual search industry will be discussed as well. Ling-Yu Duan, Feng Gao 0014, Jie Chen 0006, Jie Lin 0001, Tiejun Huang 0001 |
ISCAS | 2 |