EDBT 2026 Demo / reviewers in the wild / expert
Wenbo Li 0001
dblp:51/3185-1
· DBLP profile ↗
28ranked-venue papers
8as first author
13since 2021 · last 2025
0000-0002-3122-5400ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 22 · 7 first-author · 8 since 2021Artificial intelligence and machine learning · 20 · 5 first-author · 10 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Diffuse and Refine Latent Prior with Transformers For Neural ISPabstractThe goal of neural ISP is to convert the RAW data captured by camera sensors into display-referred sRGB images via an end-to-end neural network. Most methods are based on the regression, so they are prone to recovering images with fewer details due to the constraint of the regression loss. Some methods attempt to recover more details by resorting to GAN or diffusion model (DM) at the cost of hurting the signal-to-noise ratio and image structures. To circumvent such a trade-off, inspired by the recent progress in image deblurring, we propose a baseline that integrates the regression-based Transformer and DM. Specifically, DM is performed in a highly compact latent space to predict the latent prior that mimics the ground-truth latent prior compressed from the target sRGB image; then, the predicted latent prior is injected into the Transformer to guide the RAW-to-sRGB mapping process. Despite the superior performance achieved by this baseline, we conjecture that the latent prior predicted by DM is not good enough to mimic the ground-truth latent prior. Therefore, based on the baseline, we propose D&R in which the latent prior is diffused and refined with Transformers. The experimental results show that our D&R achieves state-of-the-art performance across multiple metrics on ZRR dataset and SID dataset, and also demonstrates the superiority over the baseline. Zhipeng Mo, Wenbo Li 0001, Sihao Ding 0004 |
ICIP | 2 |
| 2025 | BiMaCoSR: Binary One-Step Diffusion Model Leveraging Flexible Matrix Compression for Real Super-ResolutionabstractWhile super-resolution (SR) methods based on diffusion models (DM) have demonstrated inspiring performance, their deployment is impeded due to the heavy request of memory and computation. Recent researchers apply two kinds of methods to compress or fasten the DM. One is to compress the DM into 1-bit, aka binarization, alleviating the storage and computation pressure. The other distills the multi-step DM into only one step, significantly speeding up inference process. Nonetheless, it remains impossible to deploy DM to resource-limited edge devices. To address this problem, we propose BiMaCoSR, which combines binarization and one-step distillation to obtain extreme compression and acceleration. To prevent the catastrophic collapse of the model caused by binarization, we proposed sparse matrix branch (SMB) and low rank matrix branch (LRM). Both auxiliary branches pass the full-precision (FP) information but in different ways. SMB absorbs the extreme values and its output is high rank, carrying abundant FP information. Whereas, the design of LRMB is inspired by LoRA and is initialized with the top r SVD components, outputting low rank representation. The computation and storage overhead of our proposed branches can be safely ignored. Comprehensive comparison experiments are conducted to exhibit BiMaCoSR outperforms current state-of-the-art binarization methods and gains competitive performance compared with FP one-step model. Moreover, we achieve excellent compression and acceleration. BiMaCoSR achieves a 23.8x compression ratio and a 27.4x speedup ratio compared to FP counterpart. Our code and model are available at https://github.com/Kai-Liu001/BiMaCoSR Kai Liu 0034, Zheng Chen 0014, Zhiteng Li, Wenbo Li 0001, Linghe Kong, Yulun Zhang 0001 |
ICML | 6 |
| 2025 | OSCAR: One-Step Diffusion Codec Across Multiple Bit-ratesabstractPretrained latent diffusion models have shown strong potential for lossy image compression, owing to their powerful generative priors. Most existing diffusion-based methods reconstruct images by iteratively denoising from random noise, guided by compressed latent representations. While these approaches have achieved high reconstruction quality, their multi-step sampling process incurs substantial computational overhead. Moreover, they typically require training separate models for different compression bit-rates, leading to significant training and storage costs. To address these challenges, we propose a one-step diffusion codec across multiple bit-rates. termed OSCAR. Specifically, our method views compressed latents as noisy variants of the original latents, where the level of distortion depends on the bit-rate. This perspective allows them to be modeled as intermediate states along a diffusion trajectory. By establishing a mapping from the compression bit-rate to a pseudo diffusion timestep, we condition a single generative model to support reconstructions at multiple bit-rates. Meanwhile, we argue that the compressed latents retain rich structural information, thereby making one-step denoising feasible. Thus, OSCAR replaces iterative sampling with a single denoising pass, significantly improving inference efficiency. Extensive experiments demonstrate that OSCAR achieves superior performance in both quantitative and visual quality metrics. The code and models are available at https://github.com/jp-guo/OSCAR/. Jinpei Guo, Yifei Ji, Zheng Chen 0014, Kai Liu 0034, Ming Liu 0018, Wang Rao, Wenbo Li 0001, Yulun Zhang 0001 |
NeurIPS | 7 |
| 2024 | Unified Srgb Real Noise Synthesizing with Adaptive Feature ModulationabstractRecently, the Neighboring Correlation-Aware (NeCA) noise model has achieved impressive performance on both noise synthesis and the downstream image denoising task. However, its design regarding noise-level prediction requires training NeCA separately for each camera type. To this end, by making use of an adaptive feature modulation technique, we improve NeCA’s noise-level prediction model to be unified for different camera types and thus enable a unified sRGB real noise synthesis method. We also find out that in the neigh-boring correlation network of NeCA, there is no mechanism to maintain the signal dependency of the synthesized noise. Therefore, we introduce another adaptive feature modulation technique to the neighboring correlation network to maintain the signal dependency of the noise. Wenbo Li 0001, Zhipeng Mo, Yilin Shen, Hongxia Jin |
ICASSP | 1 |
| 2024 | Unleashing Multispectral Video's Potential in Semantic Segmentation: A Semi-supervised Viewpoint and New UAV-View BenchmarkabstractThanks to the rapid progress in RGB & thermal imaging, also known as multispectral imaging, the task of multispectral video semantic segmentation, or MVSS in short, has recently drawn significant attentions. Noticeably, it offers new opportunities in improving segmentation performance under unfavorable visual conditions such as poor light or overexposure. Unfortunately, there are currently very few datasets available, including for example MVSeg dataset that focuses purely toward eye-level view; and it features the sparse annotation nature due to the intensive demands of labeling process. To address these key challenges of the MVSS task, this paper presents two major contributions: the introduction of MVUAV, a new MVSS benchmark dataset, and the development of a dedicated semi-supervised MVSS baseline - SemiMV. Our MVUAV dataset is captured via Unmanned Aerial Vehicles (UAV), which offers a unique oblique bird’s-eye view complementary to the existing MVSS datasets; it also encompasses a broad range of day/night lighting conditions and over 30 semantic categories. In the meantime, to better leverage the sparse annotations and extra unlabeled RGB-Thermal videos, a semi-supervised learning baseline, SemiMV, is proposed to enforce consistency regularization through a dedicated Cross-collaborative Consistency Learning (C3L) module and a denoised temporal aggregation strategy. Comprehensive empirical evaluations on both MVSeg and MVUAV benchmark datasets have showcased the efficacy of our SemiMV baseline. Wei Ji 0011, Wenbo Li 0001, Yilin Shen, Li Cheng 0001, Hongxia Jin |
NeurIPS | 3 |
| 2024 | Efficient Layout-Guided Image Inpainting for Mobile UseabstractThe layout guidance, which specifies the pixel-wise object distribution, is beneficial to preserving the object boundaries in image inpainting while not hurting model’s generalization capability. We aim to design an efficient and robust layout-guided image inpainting method for mobile use, which can achieve the robustness in presence of the mixed scenes where objects with the delicate shape reside next to the hole. Our method is made up of two sub-models, which restore the pixel-information for the hole from coarse to fine, and support each other to overcome the practical challenges encountered when making the whole method lightweight. The layout mask guides the two sub-models, which thus enables the robustness of our method in mixed scenes. We demonstrate the efficiency and robustness of our method via both the experiments and a mobile demo. Wenbo Li 0001, Yi Wei 0006, Yilin Shen, Hongxia Jin |
WACV | 1 |
| 2023 | One-stage Progressive Dichotomous Segmentation
Karim Ahmed, Wenbo Li 0001, Yilin Shen, Hongxia Jin |
BMVC | 3 |
| 2023 | NeRFLiX: High-Quality Neural View Synthesis by Learning a Degradation-Driven Inter-viewpoint MiXerabstractNeural radiance fields (NeRF) show great success in novel view synthesis. However, in real-world scenes, recovering high-quality details from the source images is still challenging for the existing NeRF-based approaches, due to the potential imperfect calibration information and scene representation inaccuracy. Even with high-quality training frames, the synthetic novel views produced by NeRF models still suffer from notable rendering artifacts, such as noise, blur, etc. Towards to improve the synthesis quality of NeRF-based approaches, we propose NeRFLiX, a general NeRF-agnostic restorer paradigm by learning a degradation-driven inter-viewpoint mixer. Specially, we design a NeRF-style degradation modeling approach and construct large-scale training data, enabling the possibility of effectively removing NeRF-native rendering artifacts for existing deep neural networks. Moreover, beyond the degradation removal, we propose an inter-viewpoint aggregation framework that is able to fuse highly related high-quality training images, pushing the performance of cutting-edge NeRF models to entirely new levels and producing highly photo-realistic synthetic views. Kun Zhou 0001, Wenbo Li 0001, Yi Wang 0074, Tao Hu 0011, Nianjuan Jiang, Xiaoguang Han 0001, Jiangbo Lu |
CVPR | 2 |
| 2023 | High Quality Entity SegmentationabstractDense image segmentation tasks (e.g., semantic, panoptic) are useful for image editing, but existing methods can hardly generalize well in an in-the-wild setting where there are unrestricted image domains, classes, and image resolution & quality variations. Motivated by these observations, we construct a new entity segmentation dataset, with a strong focus on high-quality dense segmentation in the wild. The dataset contains images spanning diverse image domains and entities, along with plentiful high-resolution images and high-quality mask annotations for training and testing. Given the high-quality and -resolution nature of the dataset, we propose CropFormer which is designed to tackle the intractability of instance-level segmentation on high-resolution images. It improves mask prediction by fusing high-res image crops that provides more fine-grained image details and the full image. CropFormer is the first query-based Transformer architecture that can effectively fuse mask predictions from multiple image views, by learning queries that effectively associate the same entities across the full image and its crop. With CropFormer, we achieve a significant AP gain of 1.9 on the challenging entity segmentation task. Furthermore, CropFormer consistently improves the accuracy of traditional segmentation tasks and datasets. The dataset and code are released at http://luqi.info/entityv2.github.io/. Lu Qi 0001, Jason Kuen, Tiancheng Shen, Jiuxiang Gu, Wenbo Li 0001, Weidong Guo, Jiaya Jia, Zhe Lin 0001, Ming-Hsuan Yang 0001 |
ICCV | 5 |
| 2023 | DVSOD: RGB-D Video Salient Object DetectionabstractSalient object detection (SOD) aims to identify standout elements in a scene, with recent advancements primarily focused on integrating depth data (RGB-D) or temporal data from videos to enhance SOD in complex scenes. However, the unison of two types of crucial information remains largely underexplored due to data constraints. To bridge this gap, we in this work introduce the DViSal dataset, fueling further research in the emerging field of RGB-D video salient object detection (DVSOD). Our dataset features 237 diverse RGB-D videos alongside comprehensive annotations, including object and instance-level markings, as well as bounding boxes and scribbles. These resources enable a broad scope for potential research directions. We also conduct benchmarking experiments using various SOD models, affirming the efficacy of multimodal video input for salient object detection. Lastly, we highlight some intriguing findings and promising future research avenues. To foster growth in this field, our dataset and benchmark results are publicly accessible at: https://dvsod.github.io/. Wei Ji 0011, Size Wang, Wenbo Li 0001, Li Cheng 0001 |
NeurIPS | 4 |
| 2022 | Foreground-Specialized Model Imitation for Instance Segmentation
Dawei Li 0006, Wenbo Li 0001, Hongxia Jin |
ACCV (7) | 2 |
| 2022 | Dual-lens Reference Image Super-Resolution
Wenbo Li 0001, Hongxia Jin |
BMVC | 2 |
| 2022 | Simultaneous multi-person tracking and activity recognition based on cohesive cluster search
Wenbo Li 0001, Yi Wei 0006, Siwei Lyu, Ming-Ching Chang |
Comput. Vis. Image Underst. | 1 |
| 2020 | 3D Single-Person Concurrent Activity Detection Using Stacked Relation NetworkabstractWe aim to detect real-world concurrent activities performed by a single person from a streaming 3D skeleton sequence. Different from most existing works that deal with concurrent activities performed by multiple persons that are seldom correlated, we focus on concurrent activities that are spatio-temporally or causally correlated and performed by a single person. For the sake of generalization, we propose an approach based on a decompositional design to learn a dedicated feature representation for each activity class. To address the scalability issue, we further extend the class-level decompositional design to the postural-primitive level, such that each class-wise representation does not need to be extracted by independent backbones, but through a dedicated weighted aggregation of a shared pool of postural primitives. There are multiple interdependent instances deriving from each decomposition. Thus, we propose Stacked Relation Networks (SRN), with a specialized relation network for each decomposition, so as to enhance the expressiveness of instance-wise representations via the inter-instance relationship modeling. SRN achieves state-of-the-art performance on a public dataset and a newly collected dataset. The relation weights within SRN are interpretable among the activity contexts. The new dataset and code are available at https://github.com/weiyi1991/UA_Concurrent/ Yi Wei 0006, Wenbo Li 0001, Yanbo Fan, Linghan Xu, Ming-Ching Chang, Siwei Lyu |
AAAI | 2 |
| 2020 | MagGAN: High-Resolution Face Attribute Editing with Mask-Guided Generative Adversarial Network
Yi Wei 0006, Zhe Gan, Wenbo Li 0001, Siwei Lyu, Ming-Ching Chang, Lei Zhang 0001, Jianfeng Gao 0001, Pengchuan Zhang |
ACCV (4) | 3 |
| 2020 | MISC: Multi-Condition Injection and Spatially-Adaptive Compositing for Conditional Person Image SynthesisabstractIn this paper, we explore synthesizing person images with multiple conditions for various backgrounds. To this end, we propose a framework named ``MISC" for conditional image generation and image compositing. For conditional image generation, we improve the existing condition injection mechanisms by leveraging the inter-condition correlations. For the image compositing, we theoretically prove the weaknesses of the cutting-edge methods, and make it more robust by removing the spatially-invariance constraint, and enabling the bounding mechanism and the spatial adaptability. We show the effectiveness of our method on the Video Instance-level Parsing dataset, and demonstrate the robustness through controllability tests. Shuchen Weng, Wenbo Li 0001, Dawei Li 0006, Hongxia Jin, Boxin Shi |
CVPR | 2 |
| 2020 | Conditional Image Repainting via Semantic Bridge and Piecewise Value Function
Shuchen Weng, Wenbo Li 0001, Dawei Li 0006, Hongxia Jin, Boxin Shi |
ECCV (9) | 2 |
| 2020 | Explainable and Efficient Sequential Correlation Network for 3D Single Person Concurrent Activity DetectionabstractWe present the sequential correlation network (SCN) to improve concurrent activity detection. SCN combines a recurrent neural network and a correlation model hierarchically to model the complex correlations and temporal dynamics of concurrent activities. SCN has several advantages that enable effective learning even from a small dataset for real-world deployment. Unlike the majority of approaches assuming that each subject performs one activity at a time, SCN is end-to- end trainable, i.e., it can automatically learn the inclusive or exclusive relations of concurrent activities. SCN is lightweight in design using only a small set of learnable parameters to model the spatio-temporal correlations of activities. This also enhances the explainability of the learned parameters. Furthermore, the learning of SCN can benefit from the initialization using semantically meaningful priors. We evaluate the proposed method against the state-of-the-art method on two benchmark datasets with human skeletal data, SCN achieves comparable performance to the SOTA but with much faster inference speed and less memory usage. Yi Wei 0006, Wenbo Li 0001, Ming-Ching Chang, Hongxia Jin, Siwei Lyu |
IROS | 2 |
| 2019 | Object-Driven Text-To-Image Synthesis via Adversarial TrainingabstractIn this paper, we propose Object-driven Attentive Generative Adversarial Newtorks (Obj-GANs) that allow attention-driven, multi-stage refinement for synthesizing complex images from text descriptions. With a novel object-driven attentive generative network, the Obj-GAN can synthesize salient objects by paying attention to their most relevant words in the text descriptions and their pre-generated class label. In addition, a novel object-wise discriminator based on the Fast R-CNN model is proposed to provide rich object-wise discrimination signals on whether the synthesized object matches the text description and the pre-generated class label. The proposed Obj-GAN significantly outperforms the previous state of the art in various metrics on the large-scale MS-COCO benchmark, increasing the inception score by 27% and decreasing the FID score by 11%. A thorough comparison between the classic grid attention and the new object-driven attention is provided through analyzing their mechanisms and visualizing their attention layers, showing insights of how the proposed model generates complex scenes in high quality. Wenbo Li 0001, Pengchuan Zhang, Lei Zhang 0001, Qiuyuan Huang, Xiaodong He 0001, Siwei Lyu, Jianfeng Gao 0001 |
CVPR | 1 |
| 2019 | Dual-stream CNN for Structured Time Series ClassificationabstractThe structured time series (STS) classification problem requires the modeling of interweaved spatiotemporal dependency. Most previous methods model these two dependencies independently. Due to the complexity of the STS data, we argue that a desirable method should be a holistic framework that is adaptive and flexible. This motivates us to design a deep neural network with such merits. Inspired by the dual-stream hypothesis in neural science, we propose a novel dual-stream framework for modeling the interweaved spatiotemporal dependency, and develop a convolutional neural network within this framework that aims to achieve high adaptability and flexibility in STS configurations of sequential order and dependency range. Our model is highly modularized and scalable, making it easy to be adapted to specific tasks. The effectiveness of our model is demonstrated through experiments on benchmark datasets for skeleton based activity recognition. Shuchen Weng, Wenbo Li 0001, Yi Zhang 0070, Siwei Lyu |
ICASSP | 2 |
| 2018 | Evolvement Constrained Adversarial Learning for Video Style Transfer
Wenbo Li 0001, Longyin Wen, Xiao Bian, Siwei Lyu |
ACCV (1) | 1 |
| 2018 | UA-DETRAC 2018: Report of AVSS2018 & IWT4S Challenge on Advanced Traffic MonitoringabstractA desirable smart traffic-monitoring and street-safety system can elicit and support the intervention of law enforcement agencies or medical staff. Recently, there has been a dramatically higher demand for such smart systems. To this end, the International Workshop on Traffic and Street Surveillance for Safety and Security (IWT4S) was organized in conjunction with the 15th IEEE International Conference on Advanced Video and Signal-based Surveillance (AVSS 2018). Our goal is to advance the state-of-the-art detection and tracking algorithms and provide a comprehensive performance evaluation for them. We evaluate 5 submitted detection and 7 submitted tracking methods on the large-scale UA-DETRAC benchmark, and the results are shared publicly on the website http://detrac-db. rit.albany.edu. We expect this challenge to advance the research and development of new detection and tracking methods for transportation applications. Siwei Lyu, Ming-Ching Chang, Dawei Du, Wenbo Li 0001, Yi Wei 0006, Marco Del Coco, Pierluigi Carcagnì, Arne Schumann, Bharti Munjal, Dinh-Quoc-Trung Dang, Doo-Hyun Choi, Erik Bochinski, Fabio Galasso, Filiz Bunyak, Guna Seetharaman, Jang-Woon Baek, Jong Taek Lee, Kannappan Palaniappan, Kil-Taek Lim, Kiyoung Moon, Kwang-Ju Kim, Lars Wilko Sommer, Meltem Brandlmaier, Minsung Kang, Moongu Jeon, Noor Al-Shakarji, Oliver Acatay, Pyong-Kun Kim, Sikandar Amin, Thomas Sikora, Tien Ba Dinh, Tobias Senst, Vu-Gia-Hy Che, Young-Chul Lim, Yun-Su Chung |
AVSS | 4 |
| 2017 | Adaptive RNN Tree for Large-Scale Human Action RecognitionabstractIn this work, we present the RNN Tree (RNN-T), an adaptive learning framework for skeleton based human action recognition. Our method categorizes action classes and uses multiple Recurrent Neural Networks (RNNs) in a treelike hierarchy. The RNNs in RNN-T are co-trained with the action category hierarchy, which determines the structure of RNN-T. Actions in skeletal representations are recognized via a hierarchical inference process, during which individual RNNs differentiate finer-grained action classes with increasing confidence. Inference in RNN-T ends when any RNN in the tree recognizes the action with high confidence, or a leaf node is reached. RNN-T effectively addresses two main challenges of large-scale action recognition: (i) able to distinguish fine-grained action classes that are intractable using a single network, and (ii) adaptive to new action classes by augmenting an existing model. We demonstrate the effectiveness of RNN-T/ACH method and compare it with the state-of-the-art methods on a large-scale dataset and several existing benchmarks. Wenbo Li 0001, Longyin Wen, Ming-Ching Chang, Ser-Nam Lim, Siwei Lyu |
ICCV | 1 |
| 2016 | Online Deformable Object Tracking Based on Structure-Aware Hyper-GraphabstractRecent advances in online visual tracking focus on designing part-based model to handle the deformation and occlusion challenges. However, previous methods usually consider only the pairwise structural dependences of target parts in two consecutive frames rather than the higher order constraints in multiple frames, making them less effective in handling large deformation and occlusion challenges. This paper describes a new and efficient method for online deformable object tracking. Different from most existing methods, this paper exploits higher order structural dependences of different parts of the tracking target in multiple consecutive frames. We construct a structure-aware hyper-graph to capture such higher order dependences, and solve the tracking problem by searching dense subgraphs on it. Furthermore, we also describe a new evaluating data set for online deformable object tracking (the Deform-SOT data set), which includes 50 challenging sequences with full annotations that represent realistic tracking challenges, such as large deformations and severe occlusions. The experimental result of the proposed method shows considerable improvement in performance over the state-of-the-art tracking methods. Dawei Du, Honggang Qi, Wenbo Li 0001, Longyin Wen, Qingming Huang, Siwei Lyu |
IEEE Trans. Image Process. | 3 |
| 2015 | Category-Blind Human Action Recognition: A Practical Recognition SystemabstractExisting human action recognition systems for 3D sequences obtained from the depth camera are designed to cope with only one action category, either single-person action or two-person interaction, and are difficult to be extended to scenarios where both action categories co-exist. In this paper, we propose the category-blind human recognition method (CHARM) which can recognize a human action without making assumptions of the action category. In our CHARM approach, we represent a human action (either a single-person action or a two-person interaction) class using a co-occurrence of motion primitives. Subsequently, we classify an action instance based on matching its motion primitive co-occurrence patterns to each class representation. The matching task is formulated as maximum clique problems. We conduct extensive evaluations of CHARM using three datasets for single-person actions, two-person interactions, and their mixtures. Experimental results show that CHARM performs favorably when compared with several state-of-the-art single-person action and two-person interaction based methods without making explicit assumptions of action category. Wenbo Li 0001, Longyin Wen, Mooi Choo Chuah, Siwei Lyu |
ICCV | 1 |
| 2015 | Online Visual Tracking Using Temporally Coherent Part ClusterabstractRecent advances in visual tracking have focused on handling deformations and occlusions using the part-based appearance model. However, it remains a challenge to come up with a reliable target representation using local parts, and hence existing trackers continue to face drifting problems. To deal with this challenge, we propose a robust online model, formulating the tracking task as a problem of identifying Temporally Coherent Part (TCP) clusters. Specifically, we pose the TCP clusters identification task as a dense neighborhoods searching problem using a relational hyper graph in which the relationship among multiple temporal local parts is encoded as the affinity value of a hyper edge connecting them. Such high-order relations ships among multiple local parts across the temporal domain make our tracker more robust towards deformations and occlusions. Extensive experiments on various challenging video sequences demonstrate that our TCP-based method performs better than the state-of-the-art methods. Wenbo Li 0001, Longyin Wen, Mooi Choo Chuah, Yi Zhang 0070, Zhen Lei 0001, Stan Z. Li |
WACV | 1 |
| 2014 | Multiple Target Tracking Based on Undirected Hierarchical Relation HypergraphabstractMulti-target tracking is an interesting but challenging task in computer vision field. Most previous data association based methods merely consider the relationships (e.g. appearance and motion pattern similarities) between detections in local limited temporal domain, leading to their difficulties in handling long-term occlusion and distinguishing the spatially close targets with similar appearance in crowded scenes. In this paper, a novel data association approach based on undirected hierarchical relation hypergraph is proposed, which formulates the tracking task as a hierarchical dense neighborhoods searching problem on the dynamically constructed undirected affinity graph. The relationships between different detections across the spatiotemporal domain are considered in a high-order way, which makes the tracker robust to the spatially close targets with similar appearance. Meanwhile, the hierarchical design of the optimization process fuels our tracker to long-term occlusion with more robustness. Extensive experiments on various challenging datasets (i.e. PETS2009 dataset, ParkingLot), including both low and high density sequences, demonstrate that the proposed method performs favorably against the state-of-the-art methods. Longyin Wen, Wenbo Li 0001, Zhen Lei 0001, Dong Yi, Stan Z. Li |
CVPR | 2 |
| 2013 | An adaptive-weight hybrid relevance feedback approach for content based image retrievalabstractContent-based image retrieval (CBIR) has been receiving intensive research attention for many applications. In order to provide the users with more precise retrieval results, relevance feedback (RF) methods have been incorporated into CBIR which take the user's feedbacks into account. In general, explicit RF methods demand too much user effort while implicit RF methods suffer from lower retrieval accuracy. As such, we propose a hybrid RF method, adaptive-weight hybrid relevance feedback (AHRF) for content-based image retrieval. AHRF integrates explicit user grading and implicit user browsing histories to build a user preference model. The model is refined iteratively and used to train a preference classifier for the users. Moreover, an adaptive-weight mechanism is proposed to achieve a personalized preference model. Our proposed method is tested on a subset of the Corel Database and the experimental results reveal that AHRF can achieve good retrieval precision with less user effort. Yi Zhang 0070, Wenbo Li 0001, Zhipeng Mo, Jiawan Zhang |
ICIP | 2 |