VLDB 2026 Research / reviewers in the wild / expert
Bing Li 0024
dblp:13/2692-24
· DBLP profile ↗
35ranked-venue papers
17as first author
25since 2021 · last 2025
0000-0001-9465-8142ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 31 · 16 first-author · 21 since 2021Artificial intelligence and machine learning · 21 · 6 first-author · 20 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SynFER: Towards Boosting Facial Expression Recognition With Synthetic DataabstractFacial expression datasets remain limited in scale due to the subjectivity of annotations and the labor-intensive nature of data collection. This limitation poses a significant challenge for developing modern deep learning-based facial expression analysis models, particularly foundation models, that rely on large-scale data for optimal performance. To tackle the overarching and complex challenge, instead of introducing a new large-scale dataset, we introduce SynFER (Synthesis of Facial Expressions with Refined Control), a novel synthetic framework for synthesizing facial expression image data based on high-level textual descriptions as well as more fine-grained and precise control through facial action units. To ensure the quality and reliability of the synthetic data, we propose a semantic guidance technique to steer the generation process and a pseudo-label generator to help rectify the facial expression labels for the synthetic images. To demonstrate the generation fidelity and the effectiveness of the synthetic data from SynFER, we conduct extensive experiments on representation learning using both synthetic data and real-world data. Results validate the efficacy of our approach and the synthetic data. Notably, our approach achieves a 67.23% classification accuracy on AffectNet when training solely with synthetic data equivalent to the AffectNet training set size, which increases to 69.84% when scaling up to five times the original size. Code is available here. Xilin He, Xiaole Xian, Bing Li 0024, Muhammad Haris Khan, ZongYuan Ge, Weicheng Xie 0001, Siyang Song, LinLin Shen, Bernard Ghanem, Xiangyu Yue 0001 |
ICCV | 4 |
| 2025 | 4D-Bench: Benchmarking Multi-Modal Large Language Models for 4D Object Understanding
Wenxuan Zhu, Bing Li 0024, Cheng Zheng 0002, Jinjie Mai, Jun Chen 0021, Letian Jiang, Abdullah Hamdi, Sara Rojas Martinez, Chia-Wen Lin, Mohamed Elhoseiny 0001, Bernard Ghanem |
ICCV | 2 |
| 2025 | OmniResponse: Online Multimodal Conversational Response Generation in Dyadic InteractionsabstractIn this paper, we introduce Online Multimodal Conversational Response Generation (OMCRG), a novel task designed to produce synchronized verbal and non-verbal listener feedback online, based on the speaker's multimodal inputs. OMCRG captures natural dyadic interactions and introduces new challenges in aligning generated audio with listeners' facial responses. To tackle these challenges, we incorporate text as an intermediate modality to connect audio and facial responses. We propose OmniResponse, a Multimodal Large Language Model (MLLM) that autoregressively generates accurate multimodal listener responses. OmniResponse leverages a pretrained LLM enhanced with two core components: Chrono-Text Markup, which precisely timestamps generated text tokens, and TempoVoice, a controllable online text-to-speech (TTS) module that outputs speech synchronized with facial responses. To advance OMCRG research, we offer ResponseNet, a dataset of 696 detailed dyadic interactions featuring synchronized split-screen videos, multichannel audio, transcripts, and annotated facial behaviors. Comprehensive evaluations on ResponseNet demonstrate that OmniResponse outperforms baseline models in terms of semantic speech content, audio-visual synchronization, and generation quality. Our dataset, code, and models are publicly available at https://omniresponse.github.io/. Jianghui Wang, Bing Li 0024, Siyang Song, Bernard Ghanem |
NeurIPS | 3 |
| 2025 | Mindstorms in Natural Language-Based Societies of MindabstractInspired by Minsky's Society of Mind, Schmidhuber's Learning to Think, and other more recent works, this paper proposes and advocates for the concept of natural language-based societies of mind (NLSOMs). We imagine these societies as consisting of a collection of multimodal neural networks, including large language models, which engage in a “mindstorm” to solve problems using a shared natural language interface. Here, we work to identify and discuss key questions about the social structure, governance, and economic principles for NLSOMs, emphasizing their impact on the future of AI. Our demonstrations with NLSOMs-which feature up to 129 agents-show their effectiveness in various tasks, including visual question answering, image captioning, and prompt generation for text-to-image synthesis. Mingchen Zhuge, Francesco Faccio, Dylan R. Ashley, Róbert Csordás, Anand Gopalakrishnan, Abdullah Hamdi, Hasan Hammoud, Vincent Herrmann, Kazuki Irie, Louis Kirsch, Bing Li 0024, Guohao Li 0001, Shuming Liu 0001, Jinjie Mai, Piotr Piekos, Aditya A. Ramesh, Imanol Schlag, Aleksandar Stanic, Yuhui Wang 0004, Mengmeng Xu 0006, Deng-Ping Fan, Bernard Ghanem, Jürgen Schmidhuber |
Comput. Vis. Media | 12 |
| 2024 | Tune-an-Ellipse: CLIP Has Potential to Find what you WantabstractVisual prompting of large vision language models such as CLIP exhibits intriguing zero-shot capabilities. A manually drawn red circle, commonly used for highlighting, can guide CLIP's attention to the surrounding region, to identify specific objects within an image. Without precise object proposals, however, it is insufficient for localization. Our novel, simple yet effective approach, i.e., Differentiable Visual Prompting, enables CLIP to zero-shot localize: given an image and a text prompt describing an object, we first pick a rendered ellipse from uniformly distributed anchor ellipses on the image grid via visual prompting, then use three loss functions to tune the ellipse coefficients to encap-sulate the target region gradually. This yields promising ex-perimental results for referring expression comprehension without precisely specified object proposals. In addition, we systematically present the limitations of visual prompting inherent in CLIP and discuss potential solutions. Jinheng Xie, Songhe Deng, Bing Li 0024, Yawen Huang, Yefeng Zheng 0001, Jürgen Schmidhuber, Bernard Ghanem, LinLin Shen, Zheng Shou 0001 |
CVPR | 3 |
| 2024 | Empowering Resampling Operation for Ultra-High-Definition Image Enhancement with Model-Aware GuidanceabstractImage enhancement algorithms have made remarkable advancements in recent years, but directly applying them to Ultra-high-definition (UHD) images presents intractable computational overheads. Therefore, previous straightforward solutions employ resampling techniques to reduce the resolution by adopting a “Downsampling-Enhancement-Upsampling” processing paradigm. However, this paradigm disentangles the resampling operators and inner enhancement algorithms, which results in the loss of information that is favored by the model, further leading to sub-optimal outcomes. In this paper, we propose a novel method of Learning Model-Aware Resampling (LMAR), which learns to customize resampling by extracting model-aware information from the UHD input image, under the guidance of model knowledge. Specifically, our method consists of two core designs, namely compensatory kernel estimation and steganographic resampling. At the first stage, we dynamically predict compensatory kernels tailored to the specific input and resampling scales. At the second stage, the image-wise compensatory information is derived with the compensatory kernels and embedded into the rescaled input images. This promotes the representation of the newly derived downscaled inputs to be more consistent with the full-resolution UHD inputs, as perceived by the model. Our LMAR enables model-aware and model-favored resampling while maintaining compatibility with existing resampling operators. Extensive experiments on multiple UHD image enhancement datasets and different backbones have shown consistent performance gains after correlating resizer and enhancer; e.g., up to 1.2dB PSNR gain for ×1.8 resampling scale on UHD-LOL4K. The code is available at https://github.com/YPatrickW/LMAR. Jie Huang 0017, Bing Li 0024, Qi Zhu 0010, Man Zhou 0003, Feng Zhao 0004 |
CVPR | 3 |
| 2024 | TrackNeRF: Bundle Adjusting NeRF from Sparse and Noisy Views via Feature Tracks
Jinjie Mai, Wenxuan Zhu, Sara Rojas 0001, Jesus Zarzar, Abdullah Hamdi, Guocheng Qian, Bing Li 0024, Silvio Giancola, Bernard Ghanem |
ECCV (12) | 7 |
| 2024 | Unleashing the Potential of the Semantic Latent Space in Diffusion Models for Image Dehazing
Zizheng Yang, Hu Yu 0001, Bing Li 0024, Jie Huang 0017, Feng Zhao 0004 |
ECCV (44) | 3 |
| 2024 | Magic123: One Image to High-Quality 3D Object Generation Using Both 2D and 3D Diffusion PriorsabstractWe present ``Magic123'', a two-stage coarse-to-fine approach for high-quality, textured 3D mesh generation from a single image in the wild using *both 2D and 3D priors*. In the first stage, we optimize a neural radiance field to produce a coarse geometry. In the second stage, we adopt a memory-efficient differentiable mesh representation to yield a high-resolution mesh with a visually appealing texture. In both stages, the 3D content is learned through reference-view supervision and novel-view guidance by a joint 2D and 3D diffusion prior. We introduce a trade-off parameter between the 2D and 3D priors to control the details and 3D consistencies of the generation. Magic123 demonstrates a significant improvement over previous image-to-3D techniques, as validated through extensive experiments on diverse synthetic and real-world images. Guocheng Qian, Jinjie Mai, Abdullah Hamdi, Jian Ren 0005, Aliaksandr Siarohin, Bing Li 0024, Hsin-Ying Lee 0001, Ivan Skorokhodov, Peter Wonka, Sergey Tulyakov, Bernard Ghanem |
ICLR | 6 |
| 2024 | Unsupervised Low-Light Image Enhancement via Spectral Consistency
Bing Li 0024, Naishan Zheng, Jie Huang 0017, Feng Zhao 0004 |
ICPR (22) | 1 |
| 2024 | Unsupervised Low-Light Image Enhancement with Dual Contrastive Learning
Bing Li 0024, Jie Huang 0017, Feng Zhao 0004 |
ICPR (21) | 2 |
| 2024 | Training Pansharpening Networks at Full Resolution Using Degenerate InvarianceabstractPansharpening is an important technique for remote sensing imaging systems to obtain high-resolution multispectral images. Existing deep learning-based methods mostly rely on using pseudo-groundtruth multi-spectral images for supervised learning. The whole training process only remains at the scale of reduced resolution, which means that the impact of the degradation process is ignored and high-quality images cannot be guaranteed at full resolution. To address the challenge, we propose a new unsupervised framework that does not rely on pseudo-groundtruth but uses the invariance of the degradation process to build a consistent loss function on the original scale for network training. Specifically, we first introduce the operator learning method to build an exact mapping function from multi-spectral to panchromatic images and decouple both spectral and texture features. Then, through joint training, operators and convolutional networks can learn the spatial degradation process and spectral degradation process at full resolution, respectively. By introducing them to build consistency constraints, we can train the pansharpening network at the original full resolution. Our approach can be applied to existing pansharpening methods, improving their usability on original data, which matches practical application requirements. The experimental results on different kinds of satellite datasets demonstrate that the proposed network outperforms state-of-the-art methods both visually and quantitatively. Our code is available at https://github.com/quycruin/Qvac. Yichang Qu, Bing Li 0024, Jie Huang 0017, Feng Zhao 0004 |
ACM Multimedia | 2 |
| 2024 | Vivid-ZOO: Multi-View Video Generation with Diffusion ModelabstractWhile diffusion models have shown impressive performance in 2D image/video generation, diffusion-based Text-to-Multi-view-Video (T2MVid) generation remains underexplored. The new challenges posed by T2MVid generation lie in the lack of massive captioned multi-view videos and the complexity of modeling such multi-dimensional distribution. To this end, we propose a novel diffusion-based pipeline that generates high-quality multi-view videos centered around a dynamic 3D object from text. Specifically, we factor the T2MVid problem into viewpoint-space and time components. Such factorization allows us to combine and reuse layers of advanced pre-trained multi-view image and 2D video diffusion models to ensure multi-view consistency as well as temporal coherence for the generated multi-view videos, largely reducing the training cost. We further introduce alignment modules to align the latent spaces of layers from the pre-trained multi-view and the 2D video diffusion models, addressing the reused layers' incompatibility that arises from the domain gap between 2D and multi-view data. In support of this and future research, we further contribute a captioned multi-view video dataset. Experimental results demonstrate that our method generates high-quality multi-view videos, exhibiting vivid motions, temporal coherence, and multi-view consistency, given a variety of text prompts. Bing Li 0024, Cheng Zheng 0002, Wenxuan Zhu, Jinjie Mai, Biao Zhang 0005, Peter Wonka, Bernard Ghanem |
NeurIPS | 1 |
| 2023 | Combating Mode Collapse via Offline Manifold Entropy EstimationabstractGenerative Adversarial Networks (GANs) have shown compelling results in various tasks and applications in recent years. However, mode collapse remains a critical problem in GANs. In this paper, we propose a novel training pipeline to address the mode collapse issue of GANs. Different from existing methods, we propose to generalize the discriminator as feature embedding and maximize the entropy of distributions in the embedding space learned by the discriminator. Specifically, two regularization terms, i.e., Deep Local Linear Embedding (DLLE) and Deep Isometric feature Mapping (DIsoMap), are introduced to encourage the discriminator to learn the structural information embedded in the data, such that the embedding space learned by the discriminator can be well-formed. Based on the well-learned embedding space supported by the discriminator, a non-parametric entropy estimator is designed to efficiently maximize the entropy of embedding vectors, playing as an approximation of maximizing the entropy of the generated distribution. By improving the discriminator and maximizing the distance of the most similar samples in the embedding space, our pipeline effectively reduces the mode collapse without sacrificing the quality of generated samples. Extensive experimental results show the effectiveness of our method which outperforms the GAN baseline, MaF-GAN on CelebA (9.13 vs. 12.43 in FID) and surpasses the recent state-of-the-art energy-based model on the ANIMEFACE dataset (2.80 vs. 2.26 in Inception score). Bing Li 0024, Haoqian Wu, Hanbang Liang, Yawen Huang, Yuexiang Li, Bernard Ghanem, Yefeng Zheng 0001 |
AAAI | 2 |
| 2023 | Frequency-consistent Optimization for Image Enhancement Networks
Bing Li 0024, Naishan Zheng, Qi Zhu 0010, Jie Huang 0017, Feng Zhao 0004 |
BMVC | 1 |
| 2023 | AdaptiveMix: Improving GAN Training via Feature Space ShrinkageabstractDue to the outstanding capability for data generation, Generative Adversarial Networks (GANs) have attracted considerable attention in unsupervised learning. However, training GANs is difficult, since the training distribution is dynamic for the discriminator, leading to unstable image representation. In this paper, we address the problem of training GANs from a novel perspective, i.e., robust image classification. Motivated by studies on robust image representation, we propose a simple yet effective module, namely AdaptiveMix, for GANs, which shrinks the regions of training data in the image representation space of the discriminator. Considering it is intractable to directly bound feature space, we propose to construct hard samples and narrow down the feature distance between hard and easy samples. The hard samples are constructed by mixing a pair of training images. We evaluate the effectiveness of our AdaptiveMix with widely-used and state-of-the-art GAN architectures. The evaluation results demonstrate that our AdaptiveMix can facilitate the training of GANs and effectively improve the image quality of generated samples. We also show that our AdaptiveMix can be further applied to image classification and Out-Of-Distribution (OOD) detection tasks, by equipping it with state-of-the-art methods. Extensive experiments on seven publicly available datasets show that our method effectively boosts the performance of baselines. The code is publicly available at https://github.com/WentianZhang-ML/AdaptiveMix. Wentian Zhang, Bing Li 0024, Haoqian Wu, Nanjun He, Yawen Huang, Yuexiang Li, Bernard Ghanem, Yefeng Zheng 0001 |
CVPR | 3 |
| 2023 | NewsNet: A Novel Dataset for Hierarchical Temporal SegmentationabstractTemporal video segmentation is the get-to- go automatic video analysis, which decomposes a long-form video into smaller components for the following-up understanding tasks. Recent works have studied several levels of granularity to segment a video, such as shot, event, and scene. Those segmentations can help compare the semantics in the corresponding scales, but lack a wider view of larger temporal spans, especially when the video is complex and structured. Therefore, we present two abstractive levels of temporal segmentations and study their hierarchy to the existing fine-grained levels. Accordingly, we collect NewsNet, the largest news video dataset consisting of 1,000 videos in over 900 hours, associated with several tasks for hierarchical temporal video segmentation. Each news video is a collection of stories on different topics, represented as aligned audio, visual, and textual data, along with extensive frame-wise annotations in four granularities. We assert that the study on NewsNet can advance the understanding of complex structured video and benefit more areas such as short-video creation, personalized advertisement, digital instruction, and education. Our dataset and code is publicly available at https://github.com/NewsNet-Benchmark/NewsNet. Haoqian Wu, Mingchen Zhuge, Bing Li 0024, Ruizhi Qiao, Xiujun Shu, Bei Gan, Liangsheng Xu, Bo Ren 0002, Mengmeng Xu 0006, Wentian Zhang, Ramachandra Raghavendra, Chia-Wen Lin, Bernard Ghanem |
CVPR | 5 |
| 2023 | Learning to Identify Critical States for Reinforcement Learning from VideosabstractRecent work on deep reinforcement learning (DRL) has pointed out that algorithmic information about good policies can be extracted from offline data which lack explicit information about executed actions [45], [46], [30]. For example, videos of humans or robots may convey a lot of implicit information about rewarding action sequences, but a DRL machine that wants to profit from watching such videos must first learn by itself to identify and recognize relevant states/actions/rewards. Without relying on ground-truth annotations, our new method called Deep State Identifier learns to predict returns from episodes encoded as videos. Then it uses a kind of mask-based sensitivity analysis to extract/identify important critical states. Extensive experiments showcase our method’s potential for understanding and improving agent behavior. The source code and the generated datasets are available at https://github.com/AI-Initiative-KAUST/VideoRLCS. Mingchen Zhuge, Bing Li 0024, Yuhui Wang 0004, Francesco Faccio, Bernard Ghanem, Jürgen Schmidhuber |
ICCV | 3 |
| 2023 | Automatic Animation of Hair Blowing in Still Portrait PhotosabstractWe propose a novel approach to animate human hair in a still portrait photo. Existing work has largely studied the animation of fluid elements such as water and fire. However, hair animation for a real image remains underexplored, which is a challenging problem, due to the high complexity of hair structure and dynamics. Considering the complexity of hair structure, we innovatively treat hair wisp extraction as an instance segmentation problem, where a hair wisp is referred to as an instance. With advanced instance segmentation networks, our method extracts meaningful and natural hair wisps. Furthermore, we propose a wisp-aware animation module that animates hair wisps with pleasing motions without noticeable artifacts. The extensive experiments show the superiority of our method. Our method provides the most pleasing and compelling viewing experience in the qualitative experiments, and outperforms state-of-the-art still-image animation methods by a large margin in the quantitative evaluation. Project url: https://nevergiveu.github.io/AutomaticHairBlowing/ Wenpeng Xiao, Bernard Ghanem, Bing Li 0024 |
ICCV | 5 |
| 2023 | Dynamically Masked Discriminator for GANsabstractTraining Generative Adversarial Networks (GANs) remains a challenging problem. The discriminator trains the generator by learning the distribution of real/generated data. However, the distribution of generated data changes throughout the training process, which is difficult for the discriminator to learn. In this paper, we propose a novel method for GANs from the viewpoint of online continual learning. We observe that the discriminator model, trained on historically generated data, often slows down its adaptation to the changes in the new arrival generated data, which accordingly decreases the quality of generated results. By treating the generated data in training as a stream, we propose to detect whether the discriminator slows down the learning of new knowledge in generated data. Therefore, we can explicitly enforce the discriminator to learn new knowledge fast. Particularly, we propose a new discriminator, which automatically detects its retardation and then dynamically masks its features, such that the discriminator can adaptively learn the temporally-vary distribution of generated data. Experimental results show our method outperforms the state-of-the-art approaches. Wentian Zhang, Bing Li 0024, Jinheng Xie, Yawen Huang, Yuexiang Li, Yefeng Zheng 0001, Bernard Ghanem |
NeurIPS | 3 |
| 2022 | SCTN: Sparse Convolution-Transformer Network for Scene Flow EstimationabstractWe propose a novel scene flow estimation approach to capture and infer 3D motions from point clouds. Estimating 3D motions for point clouds is challenging, since a point cloud is unordered and its density is significantly non-uniform. Such unstructured data poses difficulties in matching corresponding points between point clouds, leading to inaccurate flow estimation. We propose a novel architecture named Sparse Convolution-Transformer Network (SCTN) that equips the sparse convolution with the transformer. Specifically, by leveraging the sparse convolution, SCTN transfers irregular point cloud into locally consistent flow features for estimating spatially consistent motions within an object/local object part. We further propose to explicitly learn point relations using a point transformer module, different from exiting methods. We show that the learned relation-based contextual information is rich and helpful for matching corresponding points, benefiting scene flow estimation. In addition, a novel loss function is proposed to adaptively encourage flow consistency according to feature similarity. Extensive experiments demonstrate that our proposed approach achieves a new state of the art in scene flow estimation. Our approach achieves an error of 0.038 and 0.037 (EPE3D) on FlyingThings3D and KITTI Scene Flow respectively, which significantly outperforms previous methods by large margins. Bing Li 0024, Cheng Zheng 0003, Silvio Giancola, Bernard Ghanem |
AAAI | 1 |
| 2022 | Pixelwise Adaptive Discretization with Uncertainty Sampling for Depth CompletionabstractImage guided depth completion is an extensively studied multi-modal task that takes sparse measurements and RGB images as input to recover dense depth maps. While the common practice is to regress the depth value from the unbounded range, some recent methods achieve breakthrough performance by discretizing the regression range into a number of discrete depth values, namely, Depth Hypotheses, and casting the scalar regression to the distribution estimation. However, existing methods employ the handcraft or image-level adaptive discretization strategies, where their generated depth hypotheses are pixel-shared, which can not adapt to all pixels and is inefficient. In this paper, we are the first to consider the difference between pixels and propose Pixelwise Adaptive Discretization to generate the tailored depth hypotheses for each pixel. Meanwhile, we introduce Uncertainty Sampling to generate the compact depth hypotheses for easy pixels and loose for hard pixels. This divide-and-conquer for each pixel allows the discrete depth hypotheses to be concentrated around the ground-truth of each pixel as much as possible, which is the core of discretization methods. Extensive experiments on the outdoor KITTI and indoor NYU Depth V2 datasets show that our model, called PADNet, surpasses the previous state-of-the-art methods even with limited parameters and computational cost. Bing Li 0024 |
ACM Multimedia | 3 |
| 2022 | AniGAN: Style-Guided Generative Adversarial Networks for Unsupervised Anime Face GenerationabstractIn this paper, we propose a novel framework to translate a portrait photo-face into an anime appearance. Different from existing translation methods which do not designate specific styles, we aim to synthesize anime-faces which are style-consistent with a given reference anime-face. However, unlike typical translation tasks, such anime-face translation is particularly challenging due to the large and complex variations of appearances among anime-faces. Existing methods often fail to transfer the styles of reference anime-faces to the generated anime-faces, or introduce noticeable artifacts/distortions in the local shapes of their generated anime-faces. We propose a novel GAN-based anime-face translator, called AniGAN, to synthesize high-quality anime-faces. Specifically, a new generator architecture is proposed to simultaneously transfer color/texture styles and transform local facial shapes into anime-like counterparts based on the style of a reference anime-face, while preserving the global structure of the source photo-face. New normalization functions are designed for the generator to further improve local shape transformation and color/texture style transfer. Besides, we propose a double-branch discriminator to learn domain-specific distributions through individual branches and learn cross-domain shared distributions via shared layers, helping generate visually pleasing anime-faces and effectively mitigate artifacts/distortions. Extensive experiments on benchmark datasets qualitatively and quantitatively demonstrate the superiority of our method over state-of-the-art methods. Bing Li 0024, Yuanlue Zhu, Chia-Wen Lin, Bernard Ghanem, LinLin Shen |
IEEE Trans. Multim. | 1 |
| 2021 | High Quality Disparity Remapping with Two-Stage WarpingabstractA high quality disparity remapping method that preserves 2D shapes and 3D structures, and adjusts disparities of important objects in stereo image pairs is proposed. It is formulated as a constrained optimization problem, whose solution is challenging, since we need to meet multiple requirements of disparity remapping simultaneously. The one-stage optimization process either degrades the quality of important objects or introduces serious distortions in background regions. To address this challenge, we propose a two-stage warping process to solve it. In the first stage, we develop a warping model that finds the optimal warping grids for important objects to fulfill multiple requirements of disparity remapping. In the second stage, we derive another warping model to refine warping results in less important regions by eliminating serious distortions in shape, disparity and 3D structure. The superior performance of the proposed method is demonstrated by experimental results. Bing Li 0024, Chia-Wen Lin, Cheng Zheng 0003, Shan Liu 0001, Junsong Yuan 0001, Bernard Ghanem, C.-C. Jay Kuo |
ICCV | 1 |
| 2021 | Shape-Preserving Stereo Object Remapping via Object-Consistent Grid WarpingabstractViewing various stereo images under different viewing conditions has escalated the need for effective object-level remapping techniques. In this paper, we propose a new object spatial mapping scheme, which adjusts the depth and size of the selected object to match user preference and viewing conditions. Existing warping-based methods often distort the shape of important objects or cannot faithfully adjust the depth/size of the selected object due to improper warping such as local rotations. In this paper, by explicitly reducing the transformation freedom degree of warping, we propose an optimization model based on axis-aligned warping for object spatial remapping. The proposed axis-aligned warping based optimization model can simultaneously adjust the depths and sizes of selected objects to their target values without introducing severe shape distortions. Moreover, we propose object consistency constraints to ensure the size/shape of parts inside a selected object to be consistently adjusted. Such constraints improve the size/shape adjustment performance while remaining robust to some extent to incomplete object extraction. Experimental results demonstrate that the proposed method achieves high flexibility and effectiveness in adjusting the size and depth of objects compared with existing methods. Bing Li 0024, Chia-Wen Lin, Cheng Zheng 0003, Shan Liu 0001, Bernard Ghanem, Wen Gao 0001, C.-C. Jay Kuo |
IEEE Trans. Image Process. | 1 |
| 2020 | Perceptual Temporal Incoherence-Guided Stereo Video RetargetingabstractStereo video retargeting aims at minimizing shape and depth distortions with temporal coherence in resizing a stereo video content to a desired size. Existing methods extend stereo image retargeting schemes to stereo video retargeting by adding additional temporal constraints that demand temporal coherence in all corresponding regions. However, such a straightforward extension incurs conflicts among multiple requirements (i.e., shape and depth preservation and their temporal coherence), thus failing to meet one or more of these requirements satisfactorily. To mitigate conflicts among depth, shape, and temporal constraints and avoid degrading temporal coherence perceptually, we relax temporal constraints for non-paired regions at frame boundaries, derive new temporal constraints to improve human viewing experience of a 3D scene, and propose an efficient grid-based implementation for stereo video retargeting. Experimental results demonstrate that our method achieves superior visual quality over existing methods. Bing Li 0024, Chia-Wen Lin, Shan Liu 0001, Tiejun Huang 0001, Wen Gao 0001, C.-C. Jay Kuo |
IEEE Trans. Image Process. | 1 |
| 2019 | Stereo Depth Mapping via Axis-Aligned WarpingabstractViewing various stereo images under different viewing conditions has escalated the need for efficient and effective depth mapping techniques for adjusting the depths and sizes of objects to match user preference. Existing methods mainly alter the depth of an object through non-uniform region warping, which, however, often cause severe depth or shape distortions, due to improper warping such as local rotations. In this paper, we propose a new object depth mapping scheme based on axis-aligned warping. The proposed axis-aligned-warping based optimization model can simultaneously adjust the depths and sizes of selected objects to their target values without introducing severe shape distortions. Experimental results demonstrate that our method achieves high flexibility and effectiveness in adjusting the size and depth of object compared with existing methods. Bing Li 0024, Chia-Wen Lin, Cheng Zheng 0003, Shan Liu 0001, C.-C. Jay Kuo |
ICIP | 1 |
| 2018 | Depth-Aware Stereo Video RetargetingabstractAs compared with traditional video retargeting, stereo video retargeting poses new challenges because stereo video contains the depth information of salient objects and its time dynamics. In this work, we propose a depth-aware stereo video retargeting method by imposing the depth fidelity constraint. The proposed depth-aware retargeting method reconstructs the 3D scene to obtain the depth information of salient objects. We cast it as a constrained optimization problem, where the total cost function includes the shape, temporal and depth distortions of salient objects. As a result, the solution can preserve the shape, temporal and depth fidelity of salient objects simultaneously. It is demonstrated by experimental results that the depth-aware retargeting method achieves higher retargeting quality and provides better user experience. Bing Li 0024, Chia-Wen Lin, Boxin Shi, Tiejun Huang 0001, Wen Gao 0001, C.-C. Jay Kuo |
CVPR | 1 |
| 2018 | Perceptual Temporal Incoherence Aware Stereo Video RetargetingabstractStereo video retargeting aims to avoid shape and depth distortions while maintaining temporal coherence of shape and depth while resizing a stereo video to a desired size. Existing methods resort to extending stereo image retargeting schemes to stereo video retargeting by imposing temporal constraints to consistently resize all corresponding regions so as to maintain temporal coherence. However, such a direct extension often incurs conflicts among the requirements for preserving shape information and depth information and maintaining their temporal coherence, thereby failing to meet one or more of these requirements. We find that properly relaxing temporal constraints for non-paired regions at frame boundaries can effectively mitigate conflicts among depth, shape, and temporal constraints without severely degrading temporal coherence perceptually. Based on this new finding, we derive effective temporal constraints to improve the viewing experience of a 3D scene for stereo video retargeting. Accordingly, we propose an efficient grid-based implementation for our method. Experimental results show that our method achieves superior visual quality over existing methods. Bing Li 0024, Chia-Wen Lin, Shan Liu 0001, Tiejun Huang 0001, Wen Gao 0001, C.-C. Jay Kuo |
ACM Multimedia | 1 |
| 2015 | Depth-Preserving Warping for Stereo Image RetargetingabstractThe popularity of stereo images and various display devices poses the need of stereo image retargeting techniques. Existing warping-based retargeting methods can well preserve the shape of salient objects in a retargeted stereo image pair. Nevertheless, these methods often incur depth distortion, since they attempt to preserve depth by maintaining the disparity of a set of sparse correspondences, rather than directly controlling the warping. In this paper, by considering how to directly control the warping functions, we propose a warping-based stereo image retargeting approach that can simultaneously preserve the shape of salient objects and the depth of 3D scenes. We first characterize the depth distortion in terms of warping functions to investigate the impact of a warping function on depth distortion. Based on the depth distortion model, we then exploit binocular visual characteristics of stereo images to derive region-based depth-preserving constraints which directly control the warping functions so as to faithfully preserve the depth of 3D scenes. Third, with the region-based depth-preserving constraints, we present a novel warping-based stereo image retargeting framework. Since the depth-preserving constraints are derived regardless of shape preservation, we relax the depth-preserving constraints to fulfill a tradeoff between shape preservation and depth preservation. Finally, we propose a quad-based implementation of the proposed framework. The results demonstrate the efficacy of our method in both depth and shape preservation for stereo image retargeting. Bing Li 0024, Ling-Yu Duan, Chia-Wen Lin, Tiejun Huang 0001, Wen Gao 0001 |
IEEE Trans. Image Process. | 1 |
| 2014 | Region-based depth-preserving stereoscopic image retargetingabstractThe popularity of stereo images and various sizes of display screens pose the need of stereo image retargeting techniques which resize stereo image pairs to desired sizes. Many content-aware stereo image retargeting methods adapt the images through non-uniformly resizing regions. However, these methods often make the depth of retargeted version inconsistent with the original one, since they do not explicitly consider different effects of resizing distinct regions on the depths of 3D scenes. In this paper, we analyze the effects of region-wise resizing on the depths of 3D scenes. With such insights, we can properly edit or maintain the depth of a stereo image pair via region-wise resizing. In addition, by taking into account the effects on different regions, we propose a grid-based retargeting model for stereo images, which simultaneously preserve the depths of 3D scenes and the shapes of salient objects. Experimental results demonstrate the superior performance of our method. Bing Li 0024, Ling-Yu Duan, Chia-Wen Lin, Wen Gao 0001 |
ICIP | 1 |
| 2014 | Spatiotemporal Grid Flow for Video RetargetingabstractVideo retargeting is a useful technique to adapt a video to a desired display resolution. It aims to preserve the information contained in the original video and the shapes of salient objects while maintaining the temporal coherence of contents in the video. Existing video retargeting schemes achieve temporal coherence via constraining each region/pixel to be deformed consistently with its corresponding region/pixel in neighboring frames. However, these methods often distort the shapes of salient objects, since they do not ensure the content consistency for regions/pixels constrained to be coherently deformed along time axis. In this paper, we propose a video retargeting scheme to simultaneously meet the two requirements. Our method first segments a video clip into spatiotemporal grids called grid flows, where the consistency of the content associated with a grid flow is maintained while retargeting the grid flow. After that, due to the coarse granularity of grid, there still may exist content inconsistency in some grid flows. We exploit the temporal redundancy in a grid flow to avoid that the grids with inconsistent content be incorrectly constrained to be coherently deformed. In particular, we use grid flows to select a set of key-frames which summarize a video clip, and resize subgrid-flows in these key-frames. We then resize the remaining nonkey-frames by simply interpolating their grid contents from the two nearest retargeted key-frames. With the key-frame-based scheme, we only need to solve a small-scale quadratic programming problem to resize subgrid-flows and perform grid interpolation, leading to low computation and memory costs. The experimental results demonstrate the superior performance of our scheme. Bing Li 0024, Ling-Yu Duan, Jinqiao Wang, Rongrong Ji, Chia-Wen Lin, Wen Gao 0001 |
IEEE Trans. Image Process. | 1 |
| 2012 | Predicting the effectiveness of queries for visual searchabstractPoor retrieval performance significantly degenerates users' experience of visual search, especially in mobile search. Ideally, users would like to be alerted when bad queries are present, which helps eliminate latency as well as waste of bandwidth, especially in 3G wireless environment. In this paper, we propose a visual query performance prediction (v-QPP) approach to predict the retrieval effectiveness. We employ latent dirichlet allocation (LDA)to derive latent topics from image database. From the collection statistics, we model the query's specificity based on topics. High specificity helps a retrieval system to derive user's search intent exactly. Moreover, as low discriminative content is difficult to search in terms of distinguishing relevant images from irrelevant one, we propose a topics based inverse concept frequency (t-ICF) model to deal with specific queries but difficult to discriminate in the reference database. Comparison experiments over MPEG CDVS benchmarking datasets have shown our method significantly outperforms existing approaches in document retrieval. Bing Li 0024, Ling-Yu Duan, Rongrong Ji, Wen Gao 0001 |
ICASSP | 1 |
| 2011 | Fast retargeting with adaptive grid optimizationabstractEffective and efficient retargeting techniques may enrich users' browsing experiences in mobile devices. Existing mesh-based retargeting solutions put less efforts in making well-tuned meshes. In this paper, we propose a novel adaptive grid based optimization method to retarget an image. First, we present an entropy based measure to guide the grid construction. Then we employ the quadtree structure to adjust the grid granularity adaptively. Furthermore, to reduce the inappropriate deformation from inconsistent importance assignment, we build a global optimization model to alleviate serious shape deformation in retargeting. Comparison experiments show our method's superiority over the state-of-the-art approaches. Bing Li 0024, Jinqiao Wang, Ling-Yu Duan, Wen Gao 0001 |
ICME | 1 |
| 2011 | Grid-Based Retargeting with Transformation Consistency Smoothing
Bing Li 0024, Ling-Yu Duan, Jinqiao Wang, Jie Chen 0006, Rongrong Ji, Wen Gao 0001 |
MMM (2) | 1 |