EDBT 2026 Demo / reviewers in the wild / expert
Wei Yu 0004
dblp:82/2790-4
· DBLP profile ↗
32ranked-venue papers
10as first author
18since 2021 · last 2026
0000-0002-4805-3115ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 22 · 5 first-author · 12 since 2021Artificial intelligence and machine learning · 17 · 6 first-author · 12 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DialoGen: Towards Dialog Gesture Generation via Identity-Decoupled Style Guidance in Interactive Diffusion ModelabstractWe propose DialoGen, a novel framework for generating realistic gestures for both interlocutors in dialog scenarios, conditioned on conversational audios. Unlike most existing methods that focus solely on a single speaker, DialoGen simultaneously generates synchronized gestures for both participants while also embedding identity-decoupled style into generated gestures that enhance realism and expressiveness. To ensure precise synchronization between interlocutors, DialoGen adopts an interactive dual-diffusion model with mutual interaction estimation, which integrates interaction correlation into the diffusion process. More importantly, by leveraging supervised contrastive learning, we develop the identity-decoupled style guidance to adaptively decompose the identity-specific style of interlocutors into latent space, enabling multi-style dialog gesture generation. Extensive experimental results demonstrate that our model significantly outperforms existing methods in generating realistic, speech-aligned, identity-specific gestures, offering a high-quality solution for various dialog scenarios. Weiyu Zhao, Chenyang Wang 0002, Liangxiao Hu, Zonglin Li 0004, Wei Yu 0004, Shengping Zhang |
AAAI | 5 |
| 2026 | OACI: Object-aware contextual integration for image captioning
Shuhan Xu, Mengya Han, Wei Yu 0004, Zheng He 0001, Xin Zhou 0003, Yong Luo 0002 |
Knowl. Based Syst. | 3 |
| 2026 | D3BSR: Blind Super-Resolution via Diffusion-Based Disentangled Degradation RepresentationabstractExisting Blind Super-Resolution (BSR) methods are mostly trained on artificial synthetic degradation data pairs or rely on specific degradation priors, which lead to poor performance due to the trained degradation mismatch between other unknown complex degradations in real-world scenarios. To tackle this problem, we propose a novel Diffusion-based Disentangled Degradation representation method for BSR, dubbed D3BSR, which disentangles arbitrary unknown degradation into structure and texture degradations to enhance perception and fidelity quality individually. Specifically, the structure degradation is optimized by degradation distribution transition with a self-supervised collaborative learning strategy to recursively minimize the perception error. The texture degradation is restored through posterior sampling controlled by a fidelity coefficient to leverage rich texture priors encapsulated in a pre-trained diffusion model for preserving fidelity. The degraded image is super-resolved using an analytical solution with the pseudo inverse of the structural and texture degradation, which achieves a controllable trade-off between perception and fidelity and does not rely on any degradation priors or extra-supervised training. Extensive experiments on the nine heavily degraded synthetic and real-world natural and face datasets demonstrate that our D3BSR outperforms SOTA methods on the diverse metrics in reconstruction faithfulness and perceptual quality. Wei Yu 0004, Qinglin Liu, Quanling Meng, Chenyang Wang 0002, Xin Sun 0003 |
IEEE Trans. Multim. | 1 |
| 2025 | What Kind of Visual Tokens Do We Need? Training-Free Visual Token Pruning for Multi-Modal Large Language Models from the Perspective of GraphabstractRecent Multimodal Large Language Models(MLLMs) often use a large number of visual tokens to compensate their visual shortcoming, leading to excessive computation and obvious visual redundancy. In this paper, we investigate what kind of visual tokens are needed for MLLMs, and reveal that both foreground and background tokens are critical for MLLMs given the varying difficulties of examples. Based on this observation, we propose a graph-based method towards training-free visual token pruning, termed G-Prune. In particular, G-Prune regards visual tokens as nodes, and construct their connections based on their semantic similarities. Afterwards, the information flow is propagated via weighted links, and the most important tokens after iterations are kept for MLLMs, which can be front or background. To validate G-Prune, we apply it to a recent MLLM called LLaVA-NeXT, and conduct extensive experiments on a set of benchmarks. The experiment results show that G-Prune can greatly reduce computation overhead while retaining high performance on both coarse- and fine-grained tasks. For instance, G-Prune can reduce 63.57% FLOPs of LLaVA-NeXT on VQA2.0 and TextVQA with only 0.95% and 2.34% accuracy drops, respectively. Yutao Jiang, Qiong Wu 0012, Wei Yu 0004, Yiyi Zhou |
AAAI | 4 |
| 2025 | Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language ModelsabstractMultimodal large language models (MLLMs) have experienced significant advancements recently, but still struggle to recognize and interpret intricate details in high-resolution (HR) images effectively. While state-of-the-art (SOTA) MLLMs claim to process images at 4K resolution, existing MLLM benchmarks only support up to 2K, leaving the capabilities of SOTA models on true HR images largely untested. Furthermore, existing methods for enhancing HR image perception in MLLMs rely on computationally expensive visual instruction tuning. To address these limitations, we introduce HR-Bench, the first deliberately designed benchmark to rigorously evaluate MLLM performance on 4K & 8K images. Through extensive experiments, we demonstrate that while downsampling HR images leads to vision information loss, leveraging complementary modalities, e.g., text, can effectively compensate for this loss. Building upon this insight, we propose Divide, Conquer and Combine, a novel training-free framework for enhancing MLLM perception of HR images. Our method follows a three-staged approach: 1) Divide: recursively partitioning the HR image into patches and merging similar patches to minimize computational overhead, 2) Conquer: leveraging the MLLM to generate accurate textual descriptions for each image patch, and 3) Combine: utilizing the generated text descriptions to enhance the MLLM's understanding of the overall HR image. Extensive experiments show that: 1) the SOTA MLLM achieves 63% accuracy, which is markedly lower than the 87% accuracy achieved by humans on HR-Bench; 2) our method brings consistent and significant improvements (a relative increase of +6% on HR-Bench and +8% on general multimodal benchmarks). Liang Ding 0006, Minyan Zeng, Xiabin Zhou, Li Shen 0008, Yong Luo 0002, Wei Yu 0004, Dacheng Tao |
AAAI | 7 |
| 2025 | MTGA: Multi-View Temporal Granularity Aligned Aggregation for Event-Based Lip-ReadingabstractLip-reading is to utilize the visual information of the speaker’s lip movements to recognize words and sentences. Existing event-based lip-reading solutions integrate different frame rate branches to learn spatio-temporal features of varying granularities. However, aggregating events into event frames inevitably leads to the loss of fine-grained temporal information within frames. To remedy this drawback, we propose a novel framework termed Multi-view Temporal Granularity aligned Aggregation (MTGA). Specifically, we first present a novel event representation method, namely time-segmented voxel graph list, where the most significant local voxels are temporally connected into a graph list. Then we design a spatio-temporal fusion module based on temporal granularity alignment, where the global spatial features extracted from event frames, together with the local relative spatial and temporal features contained in voxel graph list are effectively aligned and integrated. Finally, we design a temporal aggregation module that incorporates positional encoding, which enables the capture of local absolute spatial and global temporal information. Experiments demonstrate that our method outperforms both the event-based and video-based lip-reading counterparts. Yong Luo 0002, Wei Yu 0004, Zheng He 0001, Jialie Shen 0001 |
AAAI | 5 |
| 2025 | Pursuing Temporal-Consistent Video Virtual Try-On via Dynamic Pose InteractionabstractVideo virtual try-on aims to seamlessly dress a subject in a video with a specific garment. The primary challenge involves preserving the visual authenticity of the garment while dynamically adapting to the pose and physique of the subject. While existing methods have predominantly focused on image-based virtual try-on, extending these techniques directly to videos often results in temporal inconsistencies. Most current video virtual try-on approaches alleviate this challenge by incorporating temporal modules, yet still overlook the critical spatiotemporal pose interactions between human and garment. Effective pose interactions in videos should not only consider spatial alignment between human and garment poses in each frame but also account for the temporal dynamics of human poses throughout the entire video. With such motivation, we propose a new framework, namely Dynamic Pose Interaction Diffusion Models (DPIDM), to leverage diffusion models to delve into dynamic pose interactions for video virtual try-on. Technically, DPIDM introduces a skeleton-based pose adapter to integrate synchronized human and garment poses into the denoising network. A hierarchical attention module is then exquisitely designed to model intra-frame human-garment pose interactions and long-term human pose dynamics across frames through pose-aware spatial and temporal attention mechanisms. Moreover, DPIDM capitalizes on a temporal regularized attention loss between consecutive frames to enhance temporal consistency. Extensive experiments conducted on VITON-HD, VVT and ViViD datasets demonstrate the superiority of our DPIDM against the baseline methods. Notably, DPIDM achieves VFID score of 0.506 on VVT dataset, leading to 60.5% improvement over the state-of-the-art GPD-VVTO approach. Dong Li 0019, Wenqi Zhong, Wei Yu 0004, Yingwei Pan, Dingwen Zhang, Ting Yao 0003, Junwei Han 0001, Tao Mei 0001 |
CVPR | 3 |
| 2025 | EgoNet: An Unified Egocentric Active Speaker Detection Framework for both Camera Wearer and Visible CandidatesabstractActive Speaker Detection (ASD) aims to determine whether each candidate in a video frame is speaking. The egocentric dataset Ego4D introduces unique challenges for this task, such as dynamic shooting angles that cause candidates to frequently leave the sight, leading to temporal discontinuities. Additionally, Ego4D poses a novel task: detecting the speaking activities of the camera wearer, who never appears in the field of view. Existing methods treat these two tasks separately, and treat candidates out of sight as noise. In contrast, we propose EgoNet, a framework that uniformly models all candidates, including those not visible. By capturing interactions among all candidates and modeling broader temporal context, EgoNet reduces uncertainty and improves performance in egocentric active speaker detection. Yongqian Li, Xin Zhou 0003, Zheng He 0001, Wei Yu 0004, Yong Luo 0002 |
ICASSP | 4 |
| 2025 | Camera-Specific Imaging Simulation for Raw Domain Image Super ResolutionabstractThe RAW domain image super-resolution faces two critical challenges: the physical impossibility of capturing native high-quality RAW references with a resolution-limited camera and the limitations of neural networks, including inefficient residual layer utilization and spectral bias in feature learning. This paper proposes a strategy combining physics-based imaging simulation and neural networks to jointly address these challenges. First, we develop a rapid imaging simulation system based on our proposed subgraph decomposition technology. It generates camera-specific degraded and clean RAW image pairs at multiple resolutions. Second, we design a LatentKAN network, featuring an iterative feature fusion network that extracts additional beneficial information through stage-wise supervision and a multi-layer Kolmogorov Arnold network that suppresses spectral bias via learnable activation functions. Ultimately, our strategy demonstrates significant advantages, achieving an average 0.8 dB PSNR improvement across all SR scales compared to state-of-the-art methods, thereby establishing a new paradigm for camera-specific super-resolution tasks. Henglu Wei, Chuxi Yang, Wei Yu 0004, Xudong Zhao 0001, Xiangyang Ji |
ACM Multimedia | 4 |
| 2024 | Learning Scale-Aware Spatio-temporal Implicit Representation for Event-based Motion DeblurringabstractExisting event-based motion deblurring methods mostly focus on restoring images with the same spatial and temporal scales as events. However, the unknown scales of images and events in the real world pose great challenges and have rarely been explored. To address this gap, we propose a novel Scale-Aware Spatio-temporal Network (SASNet) to flexibly restore blurred images with event streams at arbitrary scales. The core idea is to implicitly aggregate both spatial and temporal correspondence features of images and events to generalize at continuous scales. To restore highly blurred local areas, we develop a Spatial Implicit Representation Module (SIRM) to aggregate spatial correlation at any resolution through event encoding sampling. To tackle global motion blur, a Temporal Implicit Representation Module (TIRM) is presented to learn temporal correlation via temporal shift operations with long-term aggregation. Additionally, we build a High-resolution Hybrid Deblur (H2D) dataset using a new-generation hybrid event-based sensor, which comprises images with naturally spatially aligned and temporally synchronized events at various scales. Experiments demonstrate that our SASNet outperforms state-of-the-art methods on both synthetic GoPro and real H2D datasets, especially in high-speed motion scenarios. Code and dataset are available at https://github.com/aipixel/SASNet. Wei Yu 0004, Jianing Li 0001, Shengping Zhang, Xiangyang Ji |
ICML | 1 |
| 2024 | Joint Input and Output Coordination for Class-Incremental Learning
Shuai Wang 0011, Yibing Zhan, Yong Luo 0002, Han Hu 0003, Wei Yu 0004, Yonggang Wen 0001, Dacheng Tao |
IJCAI | 5 |
| 2024 | Rethinking Imbalance in Image Super-Resolution for Efficient InferenceabstractExisting super-resolution (SR) methods optimize all model weights equally using $\mathcal{L}_1$ or $\mathcal{L}_2$ losses by uniformly sampling image patches without considering dataset imbalances or parameter redundancy, which limits their performance. To address this, we formulate the image SR task as an imbalanced distribution transfer learning problem from a statistical probability perspective, proposing a plug-and-play Weight-Balancing framework (WBSR) to achieve balanced model learning without changing the original model structure and training data. Specifically, we develop a Hierarchical Equalization Sampling (HES) strategy to address data distribution imbalances, enabling better feature representation from texture-rich samples. To tackle model optimization imbalances, we propose a Balanced Diversity Loss (BDLoss) function, focusing on learning texture regions while disregarding redundant computations in smooth regions. After joint training of HES and BDLoss to rectify these imbalances, we present a gradient projection dynamic inference strategy to facilitate accurate and efficient inference. Extensive experiments across various models, datasets, and scale factors demonstrate that our method achieves comparable or superior performance to existing approaches with about 34\% reduction in computational cost. Wei Yu 0004, Qinglin Liu, Jianing Li 0001, Shengping Zhang, Xiangyang Ji |
NeurIPS | 1 |
| 2024 | Stereo Image Restoration via Attention-Guided Correspondence LearningabstractAlthough stereo image restoration has been extensively studied, most existing work focuses on restoring stereo images with limited horizontal parallax due to the binocular symmetry constraint. Stereo images with unlimited parallax (e.g., large ranges and asymmetrical types) are more challenging in real-world applications and have rarely been explored so far. To restore high-quality stereo images with unlimited parallax, this paper proposes an attention-guided correspondence learning method, which learns both self- and cross-views feature correspondence guided by parallax and omnidirectional attention. To learn cross-view feature correspondence, a Selective Parallax Attention Module (SPAM) is proposed to interact with cross-view features under the guidance of parallax attention that adaptively selects receptive fields for different parallax ranges. Furthermore, to handle asymmetrical parallax, we propose a Non-local Omnidirectional Attention Module (NOAM) to learn the non-local correlation of both self- and cross-view contexts, which guides the aggregation of global contextual features. Finally, we propose an Attention-guided Correspondence Learning Restoration Network (ACLRNet) upon SPAMs and NOAMs to restore stereo images by associating the features of two views based on the learned correspondence. Extensive experiments on five benchmark datasets demonstrate the effectiveness and generalization of the proposed method on three stereo image restoration tasks including super-resolution, denoising, and compression artifact reduction. Shengping Zhang, Wei Yu 0004, Feng Jiang 0001, Liqiang Nie, Hongxun Yao, Qingming Huang, Dacheng Tao |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Human Selective MattingabstractExisting human matting methods are incapable of accurately estimating the alpha mattes of arbitrarily selected humans from a group photo. An alternative solution is to apply them to the corresponding cropped image patches. However, this option obtains an inaccurate alpha estimation due to the interference of the body parts of the neighboring humans. In addition, these methods are only trained on finely annotated synthetic data, which causes poor performance in real-world scenarios due to the domain shift. To address these problems, we propose human selective matting (HSMatt), which performs matting for arbitrarily selected humans from a group photo given only a simple bounding box as guidance. Specifically, we design a global–local context network to extract both local and global semantic context features. A human-aware trimap network is then proposed to generate human-aware trimaps for the selected humans, which adopts stacked bidirectional inference modules with intermediate supervision to progressively refine the estimated trimap. Finally, a partially supervised matting network is introduced to estimate the alpha matte, which uses a sample-varying loss to train the network on both the finely annotated synthetic data and coarsely annotated real-world data, resulting in high accuracy and good generalization. To evaluate the proposed HSMatt, we construct the first human selective matting dataset, named HSM-200K, which contains over 200,000 human images with instance-level alpha matte annotations. Experimental results demonstrate that the proposed HSMatt outperforms state-of-the-art methods. Qinglin Liu, Quanling Meng, Xiaoqian Lv, Zonglin Li 0004, Wei Yu 0004, Shengping Zhang |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2023 | Residual Hybrid Attention Network for Compression Artifact ReductionabstractResidual Hybrid Attention Network (RHAN) can restore images of arbitrary compression quality through flexibly fusing high-frequency features in the spatial and frequency domains based on the input quality factor. Specifically, to remove the compression artifacts, we propose a hybrid attention block (HAB) to adaptively restore the loss of high-frequency components, which parallelly predicts attention maps along two separate dimensions spatial and frequency spectra. To recover the compressed image flexibly and controllably, we further design a modulation decompression block (MDB), which employs a prior factor to learn a pair of modulation parameters and performs adaptively affine transformation on the obtained high-frequency features, thereby achieving high-quality image restoration at arbitrary compression levels. The quantitative and qualitative experiments on various public data sets show that the RHAN achieves the best performance and optimal visual perceptual quality in the JPEG image restoration with arbitrary compression levels. Bingchun Luo, Wei Yu 0004 |
ICASSP | 2 |
| 2023 | Parameter and Computation Efficient Transfer Learning for Vision-Language Pre-trained ModelsabstractWith ever increasing parameters and computation, vision-language pre-trained (VLP) models exhibit prohibitive expenditure in downstream task adaption. Recent endeavors mainly focus on parameter efficient transfer learning (PETL) for VLP models by only updating a small number of parameters. However, excessive computational overhead still plagues the application of VLPs. In this paper, we aim at parameter and computation efficient transfer learning (PCETL) for VLP models. In particular, PCETL not only needs to limit the number of trainable parameters in VLP models, but also to reduce the computational redundancy during inference, thus enabling a more efficient transfer. To approach this target, we propose a novel dynamic architecture skipping (DAS) approach towards effective PCETL. Instead of directly optimizing the intrinsic architectures of VLP models, DAS first observes the significances of their modules to downstream tasks via a reinforcement learning (RL) based process, and then skips the redundant ones with lightweight networks, i.e. adapters, according to the obtained rewards. In this case, the VLP model can well maintain the scale of trainable parameters while speeding up its inference on downstream tasks. To validate DAS, we apply it to two representative VLP models, namely ViLT and METER, and conduct extensive experiments on a bunch of VL tasks. The experimental results not only show the great advantages of DAS in reducing computational complexity, e.g. -11.97% FLOPs of METER on VQA2.0, but also confirm its competitiveness against existing PETL methods in terms of parameter scale and performance. Our source code is given in our appendix. Qiong Wu 0012, Wei Yu 0004, Yiyi Zhou, Shubin Huang, Xiaoshuai Sun, Rongrong Ji |
NeurIPS | 2 |
| 2023 | Scale-Aware Frequency Attention network for super-resolution
Wei Yu 0004, Zonglin Li 0004, Qinglin Liu, Feng Jiang 0001, Changyong Guo, Shengping Zhang |
Neurocomputing | 1 |
| 2021 | SSDL: Self-Supervised Dictionary LearningabstractThe label-embedded dictionary learning (DL) algorithms generate influential dictionaries by introducing discriminative information. However, there exists a limitation: All the label-embedded DL methods rely on the labels due that this way merely achieves ideal performances in supervised learning. While in semi-supervised and unsupervised learning, it is no longer sufficient to be effective. Inspired by the concept of self-supervised learning (e.g., setting the pretext task to generate a universal model for the downstream task), we propose a Self-Supervised Dictionary Learning (SSDL) framework to address this challenge. Specifically, we first design a p-Laplacian Attention Hypergraph Learning (pAHL) block as the pretext task to generate pseudo soft labels for DL. Then, we adopt the pseudo labels to train a dictionary from a primary label-embedded DL method. We evaluate our SSDL on two human activity recognition datasets. The comparison results with other state-of-the-art methods have demonstrated the efficiency of SSDL. Shuai Shao 0006, Lei Xing 0005, Wei Yu 0004, Rui Xu 0012, Yanjiang Wang 0001, Baodi Liu |
ICME | 3 |
| 2020 | Actionness-pooled Deep-convolutional Descriptor for fine-grained action recognition
Tingting Han 0003, Hongxun Yao, Xiaoshuai Sun, Wenlong Xie, Sicheng Zhao, Wei Yu 0004 |
Neurocomputing | 6 |
| 2019 | Adaptive Semantic-Visual Tree for Hierarchical EmbeddingsabstractMerchandise categories inherently form a semantic hierarchy with different levels of concept abstraction, especially for fine-grained categories. This hierarchy encodes rich correlations among various categories across different levels, which can effectively regularize the semantic space and thus make prediction less ambiguous. However, previous studies of fine-grained image retrieval primarily focus on semantic similarities or visual similarities. In real application, merely using visual similarity may not satisfy the need of consumers to search merchandise with real-life images, e.g., given a red coat as query image, we might get red suit in recall results only based on visual similarity, since they are visually similar; But the users actually want coat rather than suit even the coat is with different color or texture attributes. We introduce this new problem based on photo shopping in real practice. That's why semantic information are integrated to regularize the margins to make "semantic" prior to "visual". To solve this new problem, we propose a hierarchical adaptive semantic-visual tree (ASVT) to depict the architecture of merchandise categories, which evaluates semantic similarities between different semantic levels and visual similarities within the same semantic class simultaneously. The semantic information satisfies the demand of consumers for similar merchandise with the query while the visual information optimize the correlations within the semantic class. At each level, we set different margins based on the semantic hierarchy and incorporate them as prior information to learn a fine-grained feature embedding. To evaluate our framework, we propose a new dataset named JDProduct, with hierarchical labels collected from actual image queries and official merchandise images on online shopping application. Extensive experimental results on the public CARS196 and CUB-200-2011 datasets demonstrate the superiority of our ASVT framework against compared state-of-the-art methods. Shuo Yang 0003, Wei Yu 0004, Ying Zheng 0009, Hongxun Yao, Tao Mei 0001 |
ACM Multimedia | 2 |
| 2019 | Gradual recovery based occluded digit images recognition
Yasi Wang, Hongxun Yao, Wei Yu 0004, Dong Wang 0030, Shangchen Zhou, Xiaoshuai Sun |
Multim. Tools Appl. | 3 |
| 2018 | Cycle-Consistency Based Hierarchical Dense Semantic CorrespondenceabstractThis work aims to estimate dense correspondences between the images from same visual class but with different geometries and visual similarities. This task is particularly challenging because (i) most image pairs have large intra-class variations, and (ii)their visual content is similar only on the high-level structure. To address these problems, this paper proposed a multilevel method to estimate per-pixel correspondences from high-level semantic to low-level structural details by the guidance of cycle-consistency. We utilize CNN feature pyramid to represent images level by level. Meanwhile, we introduce cycle-consistency to measure the reliability of flow vector, which further affects the guidance from higher level to lower level. The proposed method has been extensively evaluated on various challenging benchmarks. The results show that our method significantly outperforms the state-of-the-arts in terms of semantic flow accuracy. Chuang Lin 0003, Hongxun Yao, Wei Yu 0004, Xiaoshuai Sun |
ICIP | 3 |
| 2018 | Hierarchical semantic image matching using CNN feature pyramid
Wei Yu 0004, Xiaoshuai Sun, Kuiyuan Yang, Yong Rui, Hongxun Yao |
Comput. Vis. Image Underst. | 1 |
| 2018 | Rediscover flowers structurally
Hongxun Yao, Xiaoshuai Sun, Sicheng Zhao, Wei Yu 0004 |
Multim. Tools Appl. | 5 |
| 2017 | The shortest matching path based on novel cycle consistencyabstractCategory-level image matching is extremely challenging due to various intra-class variations. To tackle the large variations, we propose an algorithm to jointly estimate the dense correspondence for image set, which reformulates image set alignment into the problem of shortest path searching. We propose a novel tri-image cycle-consistency to measure the matching “distance” between two image, which is further used to improve the pair-wise dense correspondence. Meanwhile, we utilize CNN feature pyramid to achieve pair-wise image matching hierarchically. Extensive experiments and analysis demonstrate the superiority of our method in matching images with challenging variations. Wei Yu 0004, Hongxun Yao |
ICIP | 1 |
| 2017 | Actor identification via mining representative actions
Wenlong Xie, Hongxun Yao, Xiaoshuai Sun, Sicheng Zhao, Wei Yu 0004, Shengping Zhang |
Neurocomputing | 5 |
| 2017 | Exploiting the complementary strengths of multi-layer CNN features for image retrieval
Wei Yu 0004, Kuiyuan Yang, Hongxun Yao, Xiaoshuai Sun, Pengfei Xu 0001 |
Neurocomputing | 1 |
| 2015 | Automatic Image Dataset Construction from Click-through Logs Using Deep Neural NetworkabstractLabelled image datasets are the backbone for high-level image understanding tasks with wide application scenarios, and continuously drive and evaluate the progress of feature designing and supervised learning models. Recently, the million scale labelled image dataset further contributes to the rebirth of deep convolutional neural network and bypass manual designing handcraft features. However, the construction process of image dataset is mainly manual-based and quite labor intensive, which often take years' efforts to construct a million scale dataset with high quality. In this paper, we propose a deep learning based method to construct large scale image dataset in an automatic way. Specifically, word representation and image representation are learned in a deep neural network from large amount of click-through logs, and further used to define word-word similarity and image-word similarity. These two similarities are used to automatize the two labor intensive steps in manual-based image dataset construction: query formation and noisy image removal. With a new proposed cross convolutional filter regularizer, we can construct a million scale image dataset in one week. Finally, two image datasets are constructed to verify the effectiveness of the method. In addition to scale, the automatically constructed dataset has comparable accuracy, diversity and cross-dataset generalization with manually labelled image datasets. Yalong Bai, Kuiyuan Yang, Wei Yu 0004, Chang Xu 0008, Wei-Ying Ma, Tiejun Zhao |
ACM Multimedia | 3 |
| 2015 | Learning Cross Space Mapping via DNN Using Large Scale Click-Through LogsabstractThe gap between low-level visual signals and high-level semantics has been progressively bridged by continuous development of deep neural network (DNN). With recent progress of DNN, almost all image classification tasks have achieved new records of accuracy. To extend the ability of DNN to image retrieval tasks, we proposed a unified DNN model for image-query similarity calculation by simultaneously modeling image and query in one network. The unified DNN is named the cross space mapping (CSM) model, which contains two parts, a convolutional part and a query-embedding part. The image and query are mapped to a common vector space via these two parts respectively, and image-query similarity is naturally defined as an inner product of their mappings in the space. To ensure good generalization ability of the DNN, we learn weights of the DNN from a large number of click-through logs which consists of 23 million clicked image-query pairs between 1 million images and 11.7 million queries. Both the qualitative results and quantitative results on an image retrieval evaluation task with 1000 queries demonstrate the superiority of the proposed method. Wei Yu 0004, Kuiyuan Yang, Yalong Bai, Hongxun Yao, Yong Rui |
IEEE Trans. Multim. | 1 |
| 2014 | DNN Flow: DNN Feature Pyramid based Image Matching
Wei Yu 0004, Kuiyuan Yang, Yalong Bai, Hongxun Yao, Yong Rui |
BMVC | 1 |
| 2014 | Bag-of-Words Based Deep Neural Network for Image RetrievalabstractThis work targets image retrieval task hold by MSR-Bing Grand Challenge. Image retrieval is considered as a challenge task because of the gap between low-level image representation and high-level textual query representation. Recently further developed deep neural network sheds light on narrowing the gap by learning high-level image representation from raw pixels. In this paper, we proposed a bag-of-words based deep neural network for image retrieval task, which learns high-level image representation and maps images into bag-of-words space. The DNN model is trained on the large scale clickthrough data, and the relevance between query and image is measured by the cosine similarity of query's bag-of-words representation and image's bag-of-words representation predicted by DNN, the visual similarity of images is computed by high-level image representation extracted via the DNN model too. Finally, PageRank algorithm is used to further improve the ranking list by considering visual similarity of images for each query. The experimental results achieved state-of-the-art performance and verified the effectiveness of our proposed method. Yalong Bai, Wei Yu 0004, Tianjun Xiao, Chang Xu 0008, Kuiyuan Yang, Wei-Ying Ma, Tiejun Zhao |
ACM Multimedia | 2 |
| 2013 | The shortest warping path based multiple images alignmentabstractIn this paper, we propose a method to align multiple images of the same category. Images with large variations are aligned via a smooth transition formed by some intermediate images. These intermediate images are found by shortest warping path algorithm on a directed complete graph. Moreover, the common regions in the images are discovered to further improve alignment performance. The experimental results show that our method is effective to map and align images of the same category but with large variations of appearance, shape and view. Wei Yu 0004, Hongxun Yao, Kuiyuan Yang, Lei Zhang 0001 |
ICIP | 1 |