EDBT 2026 Demo / reviewers in the wild / expert
Chong Mou
dblp:276/3204
· DBLP profile ↗
22ranked-venue papers
15as first author
21since 2021 · last 2026
0000-0003-0296-4893ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 17 · 11 first-author · 16 since 2021Artificial intelligence and machine learning · 15 · 10 first-author · 15 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | OmniDrag: Enabling Motion Control for Omnidirectional Image-to-Video Generation
Shijie Zhao 0001, Chong Mou, Xuhan Sheng, Zhenyu Zhang 0005, Li Zhang 0006, Jian Zhang 0018 |
Int. J. Comput. Vis. | 3 |
| 2025 | Diffusion-Based Hierarchical Image SteganographyabstractThis paper introduces Hierarchical Image Steganography (HIS), a novel method that uses diffusion models to enhance the security and capacity of embedding multiple images into a single container. HIS assigns varying levels of robustness to images based on their importance, ensuring enhanced protection against manipulation. It adeptly exploits the robustness of the Diffusion Model and the reversibility of the Flow Model. Integrating Embed-Flow and Enhance-Flow improves embedding efficiency and image recovery quality, respectively, setting HIS apart from conventional multi-image steganography techniques. This innovative structure can autonomously generate a container image, thereby securely and efficiently concealing multiple images and text. Rigorous subjective and objective evaluations underscore HIS’s advantage in analytical resistance, robustness, and capacity, illustrating its expansive applicability in content safeguarding and privacy fortification. Youmin Xu, Xuanyu Zhang 0003, Xiandong Meng, Chong Mou, Jian Zhang 0018 |
ICME | 4 |
| 2025 | DreamO: A Unified Framework for Image CustomizationabstractRecently, extensive research on image customization (e.g., identity, subject, style, background, etc.) demonstrates strong customization capabilities in large-scale generative models. However, most approaches are designed for specific tasks, restricting their generalizability to combine different types of condition. Developing a unified framework for image customization remains an open challenge. In this paper, we present DreamO, an image customization framework designed to support a wide range of tasks while facilitating seamless integration of multiple conditions. Specifically, DreamO utilizes a diffusion transformer (DiT) framework to uniformly process input of different types. During training, we introduce a feature routing constraint to facilitate the precise querying of relevant information from reference images. Additionally, we design a placeholder strategy that associates specific placeholders with conditions at particular positions, enabling control over the placement of conditions in the generated results. Moreover, we employ a progressive training strategy to ensure smooth model convergence and correct the generation quality of the final output. Extensive experiments demonstrate that the proposed DreamO can effectively perform various image customization tasks with high quality and flexibly integrate different types of control conditions. Project page: https://mc-e.github.io/project/DreamO Chong Mou, Yanze Wu, Wenxu Wu, Pengze Zhang, Yufeng Cheng, Xinghui Li, Mengtian Li 0003, Mingcong Liu, Yunsheng Jiang, Shaojin Wu, Songtao Zhao, Jian Zhang 0018 |
SIGGRAPH Asia | 1 |
| 2024 | T2I-Adapter: Learning Adapters to Dig Out More Controllable Ability for Text-to-Image Diffusion ModelsabstractThe incredible generative ability of large-scale text-to-image (T2I) models has demonstrated strong power of learning complex structures and meaningful semantics. However, relying solely on text prompts cannot fully take advantage of the knowledge learned by the model, especially when flexible and accurate controlling (e.g., structure and color) is needed. In this paper, we aim to ``dig out" the capabilities that T2I models have implicitly learned, and then explicitly use them to control the generation more granularly. Specifically, we propose to learn low-cost T2I-Adapters to align internal knowledge in T2I models with external control signals, while freezing the original large T2I models. In this way, we can train various adapters according to different conditions, achieving rich control and editing effects in the color and structure of the generation results. Further, the proposed T2I-Adapters have attractive properties of practical value, such as composability and generalization ability. Extensive experiments demonstrate that our T2I-Adapter has promising generation quality and a wide range of applications. Our code is available at https://github.com/TencentARC/T2I-Adapter. Chong Mou, Xintao Wang 0002, Liangbin Xie, Yanze Wu, Jian Zhang 0018, Zhongang Qi, Ying Shan |
AAAI | 1 |
| 2024 | DiffEditor: Boosting Accuracy and Flexibility on Diffusion-Based Image EditingabstractLarge-scale Text-to-Image (T2I) diffusion models have revolutionized image generation over the last few years. Although owning diverse and high-quality generation capabilities, translating these abilities to fine-grained Image editing remains challenging. In this paper, we propose DiffEditor to rectify two weaknesses in existing diffusion-based image editing: (1) in complex scenarios, editing results often lack editing accuracy and exhibit unexpected artifacts; (2) lack of flexibility to harmonize editing operations, e.g., imagine new content. In our solution, we introduce image prompts in fine-grained image editing, cooperating with the text prompt to better describe the editing content. To increase the flexibility while maintaining content consistency, we locally combine stochastic differential equation (SDE) into the ordinary differential equation (ODE) sampling. In addition, we incorporate regional score-based gradient guidance and a time travel strategy into the diffusion sampling, further improving the editing quality. Extensive experiments demonstrate that our method can efficiently achieve state-of-the-art performance on various fine-grained image editing tasks, including editing within a single image (e.g., object moving, resizing, and content dragging) and across images (e.g., appearance replacing and object pasting). Our source code is released at https://github.com/MC-E/DragonDiffusion. Chong Mou, Xintao Wang 0002, Jiechong Song, Ying Shan, Jian Zhang 0018 |
CVPR | 1 |
| 2024 | 360DVD: Controllable Panorama Video Generation with 360-Degree Video Diffusion ModelabstractPanorama video recently attracts more interest in both study and application, courtesy of its immersive experience. Due to the expensive cost of capturing$360^{\circ}$panoramic videos, generating desirable panorama videos by prompts is urgently required. Lately, the emerging text-to-video (T2V) diffusion methods demonstrate notable effectiveness in standard video generation. However, due to the significant gap in content and motion patterns between panoramic and standard videos, these methods encounter challenges in yielding satisfactory$360^{\circ}$panoramic videos. In this paper, we propose a pipeline named 360-Degree Video Diffusion model (360DVD) for generating$360^{\circ}$panoramic videos based on the given prompts and motion conditions. Specifically, we introduce a lightweight 360-Adapter accompanied by 360 Enhancement Techniques to transform pre-trained T2V models for panorama video generation. We further propose a new panorama dataset named WEB360 consisting of panoramic video-text pairs for training 360DVD, addressing the absence of captioned panoramic video datasets. Extensive experiments demonstrate the superior-ity and effectiveness of 360DVD for panorama video gen-eration. Our project page is at https: / /akaneqwq. github. io/360DVD/. Chong Mou, Xinhua Cheng, Jian Zhang 0018 |
CVPR | 3 |
| 2024 | DragonDiffusion: Enabling Drag-style Manipulation on Diffusion ModelsabstractDespite the ability of text-to-image (T2I) diffusion models to generate high-quality images, transferring this ability to accurate image editing remains a challenge. In this paper, we propose a novel image editing method, DragonDiffusion, enabling Drag-style manipulation on Diffusion models. Specifically, we treat image editing as the change of feature correspondence in a pre-trained diffusion model. By leveraging feature correspondence, we develop energy functions that align with the editing target, transforming image editing operations into gradient guidance. Based on this guidance approach, we also construct multi-scale guidance that considers both semantic and geometric alignment. Furthermore, we incorporate a visual cross-attention strategy based on a memory bank design to ensure consistency between the edited result and original image. Benefiting from these efficient designs, all content editing and consistency operations come from the feature correspondence without extra model fine-tuning. Extensive experiments demonstrate that our method has promising performance on various image editing tasks, including within a single image (e.g., object moving, resizing, and content dragging) or across images (e.g., appearance replacing and object pasting). Code is available at https://github.com/MC-E/DragonDiffusion. Chong Mou, Xintao Wang 0002, Jiechong Song, Ying Shan, Jian Zhang 0018 |
ICLR | 1 |
| 2024 | ReVideo: Remake a Video with Motion and Content ControlabstractDespite significant advancements in video generation and editing using diffusion models, achieving accurate and localized video editing remains a substantial challenge. Additionally, most existing video editing methods primarily focus on altering visual content, with limited research dedicated to motion editing. In this paper, we present a novel attempt to Remake a Video (ReVideo) which stands out from existing methods by allowing precise video editing in specific areas through the specification of both content and motion. Content editing is facilitated by modifying the first frame, while the trajectory-based motion control offers an intuitive user interaction experience. ReVideo addresses a new task involving the coupling and training imbalance between content and motion control. To tackle this, we develop a three-stage training strategy that progressively decouples these two aspects from coarse to fine. Furthermore, we propose a spatiotemporal adaptive fusion module to integrate content and motion control across various sampling steps and spatial locations. Extensive experiments demonstrate that our ReVideo has promising performance on several accurate video editing applications, i.e., (1) locally changing video content while keeping the motion constant, (2) keeping content unchanged and customizing new motion trajectories, (3) modifying both content and motion trajectories. Our method can also seamlessly extend these applications to multi-area editing without specific training, demonstrating its flexibility and robustness. Chong Mou, Mingdeng Cao, Xintao Wang 0002, Zhaoyang Zhang 0004, Ying Shan, Jian Zhang 0018 |
NeurIPS | 1 |
| 2024 | Empowering Real-World Image Super-Resolution With Flexible Interactive ModulationabstractInteractive image restoration aims to construct an interactive pathway between users and restoration networks, which empowers users to modulate the restoration results according to their own demands. However, existing methods are primarily limited to training their networks with predefined and simplistic synthetic degradations. Consequently, these methods often encounter significant performance degradation when confronted with real-world degradations that deviate from their assumptions. Furthermore, existing interactive image restoration approaches solely support global modulation, wherein a single modulation factor governs the reconstruction process for the entire image. In this paper, we propose a novel method to perform real-world and intricate image super-resolution in an interactive manner. Specifically, we propose a metric-learning-based degradation estimation strategy to estimate not only the overall degradation level of the entire image but also the finer-grained, pixel-wise degradation within real-world scenarios. This enables local control over the restoration results by selectively modulating the corresponding regions based on the densely-estimated degradation map. Additionally, a new metric-argumented loss is proposed to further enhance the performance of real-world image super-resolution. Through extensive experimentation, we demonstrate the efficacy of our method in achieving exceptional modulation and restoration performance in real-world image super-resolution tasks, all while maintaining an appealing model complexity. Chong Mou, Xintao Wang 0002, Yanze Wu, Ying Shan, Jian Zhang 0018 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Large-Capacity and Flexible Video Steganography via Invertible Neural NetworkabstractVideo steganography is the art of unobtrusively concealing secret data in a cover video and then recovering the secret data through a decoding protocol at the receiver end. Although several attempts have been made, most of them are limited to low-capacity and fixed steganography. To rectify these weaknesses, we propose a Large-capacity and Flexible Video Steganography Network (LF-VSN) in this paper. For large-capacity, we present a reversible pipeline to perform multiple videos hiding and recovering through a single invertible neural network (INN). Our method can hide/recover 7 secret videos in/from 1 cover video with promising performance. For flexibility, we propose a key-controllable scheme, enabling different receivers to recover particular secret videos from the same cover video through specific keys. Moreover, we further improve the flexibility by proposing a scalable strategy in multiple videos hiding, which can hide variable numbers of secret videos in a cover video with a single model and a single training session. Extensive experiments demonstrate that with the significant improvement of the video steganography performance, our proposed LF-VSN has high security, large hiding capacity, and flexibility. The source code is available at https://github.com/MC-E/LF-VSN. Chong Mou, Youmin Xu, Jiechong Song, Chen Zhao 0002, Bernard Ghanem, Jian Zhang 0018 |
CVPR | 1 |
| 2023 | Optimization-Inspired Cross-Attention Transformer for Compressive SensingabstractBy integrating certain optimization solvers with deep neural networks, deep unfolding network (DUN) with good interpretability and high performance has attracted growing attention in compressive sensing (CS). However, existing DUNs often improve the visual quality at the price of a large number of parameters and have the problem of feature information loss during iteration. In this paper, we propose an Optimization-inspired Cross-attention Transformer (OCT) module as an iterative process, leading to a lightweight OCT-based Unfolding Framework (OCTUF) for image CS. Specifically, we design a novel Dual Cross Attention (Dual-CA) sub-module, which consists of an Inertia-Supplied Cross Attention (ISCA) block and a Projection-Guided Cross Attention (PGCA) block. ISCA block introduces multi-channel inertia forces and increases the memory effect by a cross attention mechanism between adjacent iterations. And, PGCA block achieves an enhanced information interaction, which introduces the inertia force into the gradient descent step through a cross attention block. Extensive CS experiments manifest that our OCTUF achieves superior performance compared to state-of-the-art methods while training lower complexity. Codes are available at https://github.com/songjiechong/OCTUF. Jiechong Song, Chong Mou, Shiqi Wang 0001, Siwei Ma 0001, Jian Zhang 0018 |
CVPR | 2 |
| 2023 | TransCL: Transformer Makes Strong and Flexible Compressive LearningabstractCompressive learning (CL) is an emerging framework that integrates signal acquisition via compressed sensing (CS) and machine learning for inference tasks directly on a small number of measurements. It can be a promising alternative to classical image-domain methods and enjoys great advantages in memory saving and computational efficiency. However, previous attempts on CL are not only limited to a fixed CS ratio, which lacks flexibility, but also limited to MNIST/CIFAR-like datasets and do not scale to complex real-world high-resolution (HR) data or vision tasks. In this article, a novel transformer-based compressive learning framework on large-scale images with arbitrary CS ratios, dubbed TransCL, is proposed. Specifically, TransCL first utilizes the strategy of learnable block-based compressed sensing and proposes a flexible linear projection strategy to enable CL to be performed on large-scale images in an efficient block-by-block manner with arbitrary CS ratios. Then, regarding CS measurements from all blocks as a sequence, a pure transformer-based backbone is deployed to perform vision tasks with various task-oriented heads. Our sufficient analysis presents that TransCL exhibits strong resistance to interference and robust adaptability to arbitrary CS ratios. Extensive experiments for complex HR data demonstrate that the proposed TransCL can achieve state-of-the-art performance in image classification and semantic segmentation tasks. In particular, TransCL with a CS ratio of 10% can obtain almost the same performance as when operating directly on the original data and can still obtain satisfying performance even with an extremely low CS ratio of 1%. The source codes of our proposed TransCL is available at https://github.com/MC-E/TransCL/. Chong Mou, Jian Zhang 0018 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Deep Generalized Unfolding Networks for Image RestorationabstractDeep neural networks (DNN) have achieved great suc-cess in image restoration. However, most DNN methods are designed as a black box, lacking transparency and inter-pretability. Although some methods are proposed to combine traditional optimization algorithms with DNN, they usually demand pre-defined degradation processes or hand-crafted assumptions, making it difficult to deal with complex and real-world applications. In this paper, we propose a Deep Generalized Unfolding Network (DGUNet) for image restoration. Concretely, without loss of interpretability, we integrate a gradient estimation strategy into the gradi-ent descent step of the Proximal Gradient Descent (PGD) algorithm, driving it to deal with complex and real-world image degradation. In addition, we design inter-stage in-formation pathways across proximal mapping in different PGD iterations to rectify the intrinsic information loss in most deep unfolding networks (DUN) through a multi-scale and spatial-adaptive way. By integrating the flexible gradi-ent descent and informative proximal mapping, we unfold the iterative PGD algorithm into a trainable DNN. Exten-sive experiments on various image restoration tasks demon-strate the superiority of our method in terms of state-of-the-art performance, interpretability, and generalizability. The source code is available at github.com/MC-E/DGUNet. Chong Mou, Jian Zhang 0018 |
CVPR | 1 |
| 2022 | Robust Invertible Image SteganographyabstractImage steganography aims to hide secret images into a container image, where the secret is hidden from human vision and can be restored when necessary. Previous image steganography methods are limited in hiding capacity and robustness, commonly vulnerable to distortion on container images such as Gaussian noise, Poisson noise, and lossy compression. This paper presents a novelflow-basedframe-work for robust invertible image steganography, dubbed as RIIS. A conditional normalizing flow is introduced to model the distribution of the redundant high-frequency component with the condition of the container image. Moreover, a well-designed container enhancement module (CEM) also contributes to the robust reconstruction. To regulate the net-work parameters for different distortion levels, a distortion-guided modulation (DGM) is implemented over flow-based blocks to make it a one-size-fits-all model. In terms of both clean and distorted image steganography, extensive experi-ments reveal that the proposed RIIS efficiently improves the robustness while maintaining imperceptibility and capacity. As far as we know, we are the first to propose a learning-based scheme to enhance the robustness of image steganog-raphy in the literature. The guarantee of steganography ro-bustness significantly broadens the application of steganog-raphy in real-world applications. Youmin Xu, Chong Mou, Jingfen Xie, Jian Zhang 0018 |
CVPR | 2 |
| 2022 | Metric Learning Based Interactive Modulation for Real-World Super-Resolution
Chong Mou, Yanze Wu, Xintao Wang 0002, Chao Dong 0005, Jian Zhang 0018, Ying Shan |
ECCV (17) | 1 |
| 2022 | COLA-Net: Collaborative Attention Network for Image RestorationabstractLocal and non-local attention-based methods have been well studied in various image restoration tasks while leading to promising performance. However, most of the existing methods solely focus on one type of attention mechanism (local or non-local). Furthermore, by exploiting the self-similarity of natural images, existing pixel-wise non-local attention operations tend to give rise to deviations in the process of characterizing long-range dependence due to image degeneration. To overcome these problems, in this paper we propose a novel collaborative attention network (COLA-Net) for image restoration, as the first attempt to combine local and non-local attention mechanisms to restore image content in the areas with complex textures and with highly repetitive details respectively. In addition, an effective and robust patch-wise non-local attention model is developed to capture long-range feature correspondences through 3D patches. Extensive experiments on synthetic image denoising, real image denoising and compression artifact reduction tasks demonstrate that our proposed COLA-Net is able to achieve state-of-the-art performance in both peak signal-to-noise ratio and visual perception, while maintaining an attractive computational complexity. The source code is available onhttps://github.com/MC-E/COLA-Net. Chong Mou, Jian Zhang 0018, Xiaopeng Fan 0001, Hangfan Liu, Ronggang Wang |
IEEE Trans. Multim. | 1 |
| 2021 | Image Denoising Based on Correlation Adaptive Sparse ModelingabstractImage restoration techniques generally use intrinsic correlations of image signals to reduce the uncertainty of the unknown signal and estimate the latent ground truth. Local and non-local correlation are the two major kinds of correlations utilized. They are different sources of correlations reflecting connections between different image data, but such a difference is not taken into consideration in most of the existing schemes. Typically, sparse representation based works use the same image data to exploit both local and non-local correlation in shared regularization. This paper aims to fully exploit local and non-local correlation of image contents separately so that near-optimal sparse representations are achieved and thus the uncertainty of signals is minimized. The proposed scheme adaptively selects different image data to exploit local and non-local correlations respectively. In particular, to exploit local correlation, the image data of interest are extracted from clustered rows of patch groups that consist of similar image contents. Experimental results on image denoising show that the proposed scheme not only outperforms state-of-the-art sparsity and low rank based methods, but also surpasses successful deep learning-based approaches in terms of PSNR, SSIM, and visual quality. Hangfan Liu, Chong Mou |
ICASSP | 3 |
| 2021 | Synergic Feature Attention for Image RestorationabstractLocal and non-local attentions are both effective methods in the domain of image restoration (IR). However, most existing image restoration methods use these two strategies indiscriminately, and how to make a trade-off between local and non-local attention operations has hardly been studied. Furthermore, the commonly used pixel-wise non-local operation tends to be biased during image restoration due to the image degeneration. To overcome these problems, in this paper, we propose a novel Synergic Attention Network (SAT-Net) for image restoration as an inventive attempt to combine local and non-local attention mechanisms to restore complex textures and highly repetitive details distinguishingly. We also propose an effective patch-wise non-local attention method to establish more reliable long-range dependences based on 3D patches. Experimental results on synthetic image denoising, real image denoising, and compression artifact reduction tasks show that our proposed model can achieve state-of-the-art performance under objective and subjective evaluations. Chong Mou, Jian Zhang 0018 |
ICASSP | 1 |
| 2021 | Dynamic Attentive Graph Learning for Image RestorationabstractNon-local self-similarity in natural images has been verified to be an effective prior for image restoration. However, most existing deep non-local methods assign a fixed number of neighbors for each query item, neglecting the dynamics of non-local correlations. Moreover, the non-local correlations are usually based on pixels, prone to be biased due to image degradation. To rectify these weaknesses, in this paper, we propose a dynamic attentive graph learning model (DAGL) to explore the dynamic non-local property on patch level for image restoration. Specifically, we propose an improved graph model to perform patch-wise graph convolution with a dynamic and adaptive number of neighbors for each node. In this way, image content can adaptively balance over-smooth and over-sharp artifacts through the number of its connected neighbors, and the patch-wise non-local correlations can enhance the message passing process. Experimental results on various image restoration tasks: synthetic image denoising, real image denoising, image demosaicing, and compression artifact reduction show that our DAGL can produce state-of-the-art results with superior accuracy and visual quality. The source code is available at https://github.com/jianzhangcs/DAGL. Chong Mou, Jian Zhang 0018, Zhuoyuan Wu |
ICCV | 1 |
| 2021 | Dense Deep Unfolding Network with 3D-CNN Prior for Snapshot Compressive ImagingabstractSnapshot compressive imaging (SCI) aims to record three-dimensional signals via a two-dimensional camera. For the sake of building a fast and accurate SCI recovery algorithm, we incorporate the interpretability of model-based methods and the speed of learning-based ones and present a novel dense deep unfolding network (DUN) with 3D-CNN prior for SCI, where each phase is unrolled from an iteration of Half-Quadratic Splitting (HQS). To better exploit the spatial-temporal correlation among frames and address the problem of information loss between adjacent phases in existing DUNs, we propose to adopt the 3D-CNN prior in our proximal mapping module and develop a novel dense feature map (DFM) strategy, respectively. Besides, in order to promote network robustness, we further propose a dense feature map adaption (DFMA) module to allow inter-phase information to fuse adaptively. All the parameters are learned in an end-to-end fashion. Extensive experiments on simulation data and real data verify the superiority of our method. The source code is available at https://github.com/jianzhangcs/SCI3D. Zhuoyuan Wu, Jian Zhang 0018, Chong Mou |
ICCV | 3 |
| 2021 | Graph Attention Neural Network for Image RestorationabstractSelf-similarity underpins modern non-local attention mechanism, which has been verified to be an effective prior for image restoration. However, most existing non-local attention restorers are implemented based on pixels, which tend to be biased due to image degeneration. Furthermore, most non-local methods for image restoration are restricted to construct fully-connected correlations in a regular Euclidean space so that all features within the search region have to participate in the feature aggregation process, no matter how similar the key feature is to the query feature. To rectify these weaknesses, in this paper, we propose a novel graph attention network for image restoration, dubbed GATIR, which establishes the non-local attention based on feature patches and utilizes the graph convolution to perform feature aggregation selectively in a non-Euclidean space. Experimental results demonstrate that our GATIR can achieve state-of-the-art performance on synthetic image denoising, real image denoising, image demosaicing, and compression artifact reduction tasks. Chong Mou, Jian Zhang 0018 |
ICME | 1 |
| 2020 | Attention Based Dual Branches Fingertip Detection Network and Virtual Key SystemabstractGesture and fingertip are becoming more and more important mediums for human-computer interaction (HCI). Therefore, algorithms of gesture recognition and fingertip detection have been extensively investigated. However, problems mainly remain in how to achieve a win-win situation between speed and accuracy, and how to deal with complex interaction environment. To rectify these problems, this paper proposes an attention-based dual branches network that can efficiently fulfill both fingertip detection and gesture recognition tasks. In order to deal with complex interaction environment, we combine both channel-wise attention and spatial-wise attention into the fingertip detection model. The extensive experiments demonstrate that our novel model is both effective and efficient. In the experiment, our proposed model achieves the average fingertip detection error at around 2.8 pixels in 640×480 video frame, and the average recognition accuracy among eight gestures reaches $99%$. Moreover, the average forward time is about 8 ms. Due to the light-weight design, this model can also achieve high-efficiency performance on CPU. In addition, we design a virtual key system based on our proposed model, which can allow users to complete the "clicking" operation naturally in virtual environment. Our proposed system can perform well with a single normal RGB camera without any pre-processing (e.g., image segmentation or contour extraction), which can significantly reduce the complexity of the interaction system. Chong Mou, Xin Zhang 0013 |
ACM Multimedia | 1 |