EDBT 2026 Demo / reviewers in the wild / expert
Chong Luo 0001
dblp:79/3712-1
· DBLP profile ↗
89ranked-venue papers
5as first author
37since 2021 · last 2026
0000-0003-0939-474XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 64 · 1 first-author · 30 since 2021Artificial intelligence and machine learning · 41 · 32 since 2021Computer networks · 17 · 4 first-authorSystems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LLM2CLIP: Powerful Language Model Unlocks Richer Cross-Modality RepresentationabstractCLIP is a seminal multimodal model that maps images and text into a shared representation space by contrastive learning on billions of image–caption pairs. Inspired by the rapid progress of large language models (LLMs), we investigate how the superior linguistic understanding and broad world knowledge of LLMs can further strengthen CLIP—particularly in handling long, complex captions. We introduce an efficient fine-tuning framework that embeds an LLM into a pretrained CLIP while incurring almost the same training cost as regular CLIP fine-tuning. Our method first “embedding-izes” the LLM for the CLIP setting, then couples it to the pretrained CLIP vision encoder through a lightweight adaptor trained on only a few million image–caption pairs. With this strategy we achieve large performance gains—without large-scale retraining—over state-of-the-art CLIP variants such as EVA02 and SigLIP-2. The LLM-enhanced CLIP delivers consistent improvements across a wide spectrum of downstream tasks, including linear-probe classification, zero-shot image–text retrieval with both short and long captions (in English and other languages), zero-shot/supervised image segmentation, object detection, and used as tokenizer for multimodal large-model benchmarks. Weiquan Huang, Aoqi Wu, Yifan Yang 0004, Xufang Luo, Yuqing Yang 0001, Usman Naseem, Chunyu Wang 0001, Qi Dai 0001, Xiyang Dai, Dongdong Chen 0001, Chong Luo 0001, Lili Qiu, Liang Hu 0004 |
AAAI | 11 |
| 2026 | HiTVideo: Hierarchical Tokenizers for Enhancing Text-to-Video Generation with Autoregressive Large Language ModelsabstractText-to-video generation poses significant challenges due to the inherent complexity of video data, which spans both temporal and spatial dimensions. It introduces additional redundancy, abrupt variations, and a domain gap between language and vision tokens while generation. Addressing these challenges requires an effective video tokenizer that can efficiently encode video data while preserving essential semantic and spatiotemporal information, serving as a critical bridge between text and vision. Inspired by the observation in VQ-VAE-2, we propose HiTVideo, a novel approach for text-to-video generation with hierarchical tokenizers. It utilizes a 3D causal VAE with a multi-layer discrete token framework, encoding video content into hierarchically structured codebooks. Higher layers capture semantic information with higher compression, while lower layers focus on fine-grained spatiotemporal details, striking a balance between compression efficiency and reconstruction quality. Our approach efficiently encodes longer video sequences (e.g., 8 seconds, 64 frames), reducing bits per pixel (bpp) by approximately 70% compared to previous tokenizers, while maintaining competitive reconstruction quality. We explore the trade-offs between compression and reconstruction, while emphasizing the advantages of high-compressed semantic tokens in text-to-video tasks. HiTVideo aims to address the potential limitations of existing video tokenizers in text-to-video generation tasks, striving for higher compression ratios, improved token quality, and simplify LLMs modeling under language guidance, offering a scalable and promising framework for advancing text to video generation. Ziqin Zhou, Yifan Yang 0004, Yuqing Yang 0001, Tianyu He, Houwen Peng, Qi Dai 0001, Lili Qiu, Chong Luo 0001, Lingqiao Liu |
AAAI | 9 |
| 2026 | MageBench: Bridging Large Multimodal Models to AgentsabstractRecent models like OpenAI’s O1 and DeepSeek’s R1, which utilize test-time scaling techniques, have demonstrated remarkable improvements in reasoning capabilities. We anticipate that in the near future, multimodal models will also experience significant breakthroughs in multimodal reasoning. This will require some highly challenging and specialized evaluations. As one of the most crucial real-world applications of multimodal models, visual agents require complex and comprehensive capabilities such as spatial planning and vision-in-the-chain type reasoning. These capabilities are currently lacking in existing multimodal benchmarks. In this paper, we introduce MageBench, a Multimodal reasoning benchmark built upon light-weight AGEnt environments that pose significant reasoning challenges and hold substantial practical value. The results show that only a few product-level models are better than random acting, and all of them are far inferior to human level. We analyze and summarize their errors and capability gaps in visual planning. Furthermore, we found that rule-based RL can significantly boost visual reasoning capabilities. This highlights that our benchmark could serve as a valuable testing ground for the emerging field of agentic RL research. Miaosen Zhang, Qi Dai 0001, Yifan Yang 0004, Jianmin Bao, Dongdong Chen 0001, Chong Luo 0001, Xin Geng 0001, Baining Guo |
WACV | 7 |
| 2025 | HomoGen: Enhanced Video Inpainting via Homography Propagation and DiffusionabstractIn this paper, we present HomoGen, an enhanced video inpainting method based on homography propagation and diffusion models. HomoGen leverages homography registration to propagate contextual pixels as priors for generating missing content in corrupted videos. Unlike previous flow-based propagation methods, which introduce local distortions due to point-to-point optical flows, homography-induced artifacts are typically global structural distortions that preserve semantic integrity. To effectively utilize these priors for generation, we employ a video diffusion model that inherently prioritizes semantic information within the priors over pixel-level details. A content-adaptive control mechanism is proposed to scale and inject the priors into intermediate video latents during iterative denoising. In contrast to existing transformer-based networks that often suffer from artifacts within priors, leading to error accumulation and unrealistic results, our denoising diffusion network can smooth out artifacts and ensure natural outputs. Extensive experiments demonstrate the effectiveness of the proposed method qualitatively and quantitatively. Ding Ding 0004, Yueming Pan, Ruoyu Feng 0001, Qi Dai 0001, Jianmin Bao, Chong Luo 0001, Zhenzhong Chen 0001 |
CVPR | 7 |
| 2025 | FloVD: Optical Flow Meets Video Diffusion Model for Enhanced Camera-Controlled Video SynthesisabstractWe present FloVD, a novel video diffusion model for camera-controllable video generation. FloVD leverages optical flow to represent the motions of the camera and moving objects. This approach offers two key benefits. Since optical flow can be directly estimated from videos, our approach allows for the use of arbitrary training videos without groundtruth camera parameters. Moreover, as background optical flow encodes 3D correlation across different viewpoints, our method enables detailed camera control by leveraging the background motion. To synthesize natural object motion while supporting detailed camera control, our framework adopts a two-stage video synthesis pipeline consisting of optical flow generation and flow-conditioned video synthesis. Extensive experiments demonstrate the superiority of our method over previous approaches in terms of accurate camera control and natural object motion synthesis. Wonjoon Jin, Qi Dai 0001, Chong Luo 0001, Seung-Hwan Baek, Sunghyun Cho |
CVPR | 3 |
| 2025 | StableAnimator: High-Quality Identity-Preserving Human Image AnimationabstractCurrent diffusion models for human image animation struggle to ensure identity (ID) consistency. This paper presents StableAnimator, the first end-to-end ID-preserving video diffusion framework, which synthesizes high-quality videos without any post-processing, conditioned on a reference image and a sequence of poses. Building upon a video diffusion model, StableAnimator contains carefully designed modules for both training and inference striving for identity consistency. In particular, StableAnimator begins by computing image and face embeddings with off-the-shelf extractors, respectively and face embeddings are further refined by interacting with image embeddings using a global content-aware Face Encoder. Then, StableAnimator introduces a novel distribution-aware ID Adapter that prevents interference caused by temporal layers while preserving ID via alignment. During inference, we propose a novel Hamilton-Jacobi-Bellman (HJB) equation-based optimization to further enhance the face quality. We demonstrate that solving the HJB equation can be integrated into the diffusion denoising process, and the resulting solution constrains the denoising path and thus benefits ID preservation. Experiments on multiple benchmarks show the effectiveness of StableAnimator both qualitatively and quantitatively. Shuyuan Tu, Xintong Han, Zhi-Qi Cheng, Qi Dai 0001, Chong Luo 0001, Zuxuan Wu |
CVPR | 6 |
| 2025 | JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion TransformersabstractWe present JointDiT, a diffusion transformer that models the joint distribution of RGB and depth. By leveraging the architectural benefit and outstanding image prior of the state-of-the-art diffusion transformer, JointDiT not only generates high-fidelity images but also produces geometrically plausible and accurate depth maps. This solid joint distribution modeling is achieved through two simple yet effective techniques that we propose, namely, adaptive scheduling weights, which depend on the noise levels of each modality, and the unbalanced timestep sampling strategy. With these techniques, we train our model across all noise levels for each modality, enabling JointDiT to naturally handle various combinatorial generation tasks, including joint generation, depth estimation, and depth-conditioned image generation by simply controlling the timesteps of each branch. JointDiT demonstrates outstanding joint generation performance. Furthermore, it achieves comparable results in depth estimation and depth-conditioned image generation, suggesting that joint distribution modeling can serve as a viable alternative to conditional generation. The project page is available at https://byungki-k.github.io/JointDiT/. Byung-Ki Kwon, Qi Dai 0001, Lee Hyoseok, Chong Luo 0001, Tae-Hyun Oh |
ICCV | 4 |
| 2025 | REDUCIO! Generating 1K Video Within 16 Seconds Using Extremely Compressed Motion Latents
Qi Dai 0001, Jianmin Bao, Yifan Yang 0004, Chong Luo 0001, Zuxuan Wu, Yu-Gang Jiang 0001 |
ICCV | 6 |
| 2025 | MotionFollower: Editing Video Motion via Score-Guided Diffusion
Shuyuan Tu, Qi Dai 0001, Sicheng Xie, Zhi-Qi Cheng, Chong Luo 0001, Xintong Han, Zuxuan Wu, Yu-Gang Jiang 0001 |
ICCV | 6 |
| 2025 | PF3plat: Pose-Free Feed-Forward 3D Gaussian Splatting for Novel View SynthesisabstractWe consider the problem of novel view synthesis from unposed images in a single feed-forward. Our framework capitalizes on fast speed, scalability, and high-quality 3D reconstruction and view synthesis capabilities of 3DGS, where we further extend it to offer a practical solution that relaxes common assumptions such as dense image views, accurate camera poses, and substantial image overlaps. We achieve this through identifying and addressing unique challenges arising from the use of pixel-aligned 3DGS: misaligned 3D Gaussians across different views induce noisy or sparse gradients that destabilize training and hinder convergence, especially when above assumptions are not met. To mitigate this, we employ pre-trained monocular depth estimation and visual correspondence models to achieve coarse alignments of 3D Gaussians. We then introduce lightweight, learnable modules to refine depth and pose estimates from the coarse alignments, improving the quality of 3D reconstruction and novel view synthesis. Furthermore, the refined estimates are leveraged to estimate geometry confidence scores, which assess the reliability of 3D Gaussian centers and condition the prediction of Gaussian parameters accordingly. Extensive evaluations on large-scale real-world datasets demonstrate that PF3plat sets a new state-of-the-art across all benchmarks, supported by comprehensive ablation studies validating our design choices. We will make the code and weights publicly available. Sunghwan Hong, Jaewoo Jung, Heeseong Shin, Jisang Han, Jiaolong Yang, Chong Luo 0001, Seungryong Kim |
ICML | 6 |
| 2025 | LaMD: Latent Motion Diffusion for Image-Conditional Video Generation
Yaosi Hu, Zhenzhong Chen 0001, Chong Luo 0001 |
Int. J. Comput. Vis. | 3 |
| 2025 | OmniTracker: Unifying Visual Object Tracking by Tracking-With-DetectionabstractVisual Object Tracking (VOT) aims to estimate the positions of target objects in a video sequence, which is an important vision task with various real-world applications. Depending on whether the initial states of target objects are specified by provided annotations in the first frame or the categories, VOT could be classified as instance tracking (e.g., SOT and VOS) and category tracking (e.g., MOT, MOTS, and VIS) tasks. Different definitions have led to divergent solutions for these two types of tasks, resulting in redundant training expenses and parameter overhead. In this paper, combing the advantages of the best practices developed in both communities, we propose a novel tracking-with-detection paradigm, where tracking supplements appearance priors for detection and detection provides tracking with candidate bounding boxes for the association. Equipped with such a design, a unified tracking model, OmniTracker, is further presented to resolve all the tracking tasks with a fully shared network architecture, model weights, and inference pipeline, eliminating the need for task-specific architectures and reducing redundancy in model parameters. We conduct extensive experimentation on seven prominent tracking datasets of different tracking tasks, including LaSOT, TrackingNet, DAVIS16-17, MOT17, MOTS20, and YTVIS19, and demonstrate that OmniTracker achieves on-par or even better results than both task-specific and unified tracking models. Zuxuan Wu, Dongdong Chen 0001, Chong Luo 0001, Xiyang Dai, Lu Yuan 0001, Yu-Gang Jiang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | CCEdit: Creative and Controllable Video Editing via Diffusion ModelsabstractIn this paper, we present CCEdit, a versatile generative video editing framework based on diffusion models. Our approach employs a novel trident network structure that separates structure and appearance control, ensuring precise and creative editing capabilities. Utilizing the foundational ControlNet architecture, we maintain the structural integrity of the video during editing. The incorporation of an additional appearance branch enables users to exert fine-grained control over the edited key frame. These two side branches seamlessly integrate into the main branch, which is constructed upon existing text-to-image (T2I) generation models, through learnable temporal layers. The versatility of our framework is demonstrated through a diverse range of choices in both structure representations and personalized T2I models, as well as the option to provide the edited key frame. To facilitate comprehensive evaluation, we introduce the BalanceCC benchmark dataset, comprising 100 videos and 4 target prompts for each video. Our extensive user studies compare CCEdit with eight state-of-the-art video editing methods. The outcomes demonstrate CCEdit's substantial superiority over all other methods. Ruoyu Feng 0001, Wenming Weng, Yuhui Yuan, Jianmin Bao, Chong Luo 0001, Zhibo Chen 0001, Baining Guo |
CVPR | 6 |
| 2024 | Unifying Correspondence, Pose and NeRF for Generalized Pose-Free Novel View SynthesisabstractThis work delves into the task of pose-free novel view synthesis from stereo pairs, a challenging and pioneering task in 3D vision. Our innovative framework, unlike any before, seamlessly integrates 2D correspondence matching, camera pose estimation, and NeRF rendering, fostering a synergistic enhancement of these tasks. We achieve this through designing an architecture that utilizes a shared representation, which serves as a foundation for enhanced 3D geometry understanding. Capitalizing on the inherent in-terplay between the tasks, our unified framework is trained end-to-end with the proposed training strategy to improve overall model accuracy. Through extensive evaluations across diverse indoor and outdoor scenes from two real-world datasets, we demonstrate that our approach achieves substantial improvement over previous methodologies, es-pecially in scenarios characterized by extreme viewpoint changes and the absence of accurate camera poses. The project page and code will be made available at: https://ku-cvlab.github.io/CoPoNeRF/. Sunghwan Hong, Jaewoo Jung, Heeseong Shin, Jiaolong Yang, Seungryong Kim, Chong Luo 0001 |
CVPR | 6 |
| 2024 | OmniViD: A Generative Framework for Universal Video UnderstandingabstractThe core of video understanding tasks, such as recognition, captioning, and tracking, is to automatically de-tect objects or actions in a video and analyze their temporal evolution. Despite sharing a common goal, different tasks often rely on distinct model architectures and annotation formats. In contrast, natural language processing benefits from a unified output space, i.e., text sequences, which simplifies the training of powerful foundational language models, such as GPT-3, with extensive training cor-pora. Inspired by this, we seek to unify the output space of video understanding tasks by using languages as labels and additionally introducing time and box tokens. In this way, a variety of video tasks could be formulated as video-grounded token generation. This enables us to address var-ious types of video tasks, including classification (such as action recognition), captioning (covering clip captioning, video question answering, and dense video captioning), and localization tasks (such as visual object tracking) within a fully shared encoder-decoder architecture, following a generative framework. Through comprehensive experiments, we demonstrate such a simple and straightforward idea is quite effective and can achieve state-of-the-art or compet-itive results on seven video benchmarks, providing a novel perspective for more universal video understanding. Code is available at https://github.com/wangjk666/OmniVid. Dongdong Chen 0001, Chong Luo 0001, Bo He 0004, Lu Yuan 0001, Zuxuan Wu, Yu-Gang Jiang 0001 |
CVPR | 3 |
| 2024 | MicroCinema: A Divide-and-Conquer Approach for Text-to-Video GenerationabstractWe present MicroCinema, a straightforward yet effective framework for high-quality and coherent text-to-video generation. Unlike existing approaches that align text prompts with video directly, MicroCinema introduces a Divide-and-Conquer strategy which divides the text-to-video into a two-stage process: text-to-image generation and image&text-to-video generation. This strategy offers two significant advantages. a) It allows us to take full advantage of the recent advances in text-to-image models, such as Stable Diffusion, Midjourney, and DALLE, to generate photorealistic and highly detailed images. b) Leveraging the generated image, the model can allocate less focus to fine-grained appearance details, prioritizing the efficient learning of motion dynamics. To implement this strategy effectively, we introduce two core designs. First, we propose the Appearance Injection Network, enhancing the preservation of the appearance of the given image. Second, we introduce the appearance Noise Prior, a novel mechanism aimed at maintaining the capabilities of pre-trained 2D diffusion models. These design elements empower MicroCinema to generate high-quality videos with precise motion, guided by the provided text prompts. Extensive experiments demonstrate the superiority of the proposed framework. Concretely, MicroCinema achieves SOTA zero-shot FVD of 342.86 on UCF-JOJ and 377.40 on MSR-VTT. Jianmin Bao, Wenming Weng, Ruoyu Feng 0001, Dacheng Yin, Jingxu Zhang, Qi Dai 0001, Zhiyuan Zhao 0001, Chunyu Wang 0001, Yuhui Yuan, Xiaoyan Sun 0001, Chong Luo 0001, Baining Guo |
CVPR | 14 |
| 2024 | Panacea: Panoramic and Controllable Video Generation for Autonomous DrivingabstractThe field of autonomous driving increasingly demands high-quality annotated training data. In this paper, we propose Panacea, an innovative approach to generate panoramic and controllable videos in driving scenarios, capable of yielding an unlimited numbers of diverse, annotated samples pivotal for autonomous driving advancements. Panacea addresses two critical challenges: ‘Consistency’ and ‘Controllability.’ Consistency ensures temporal and cross-view coherence, while Controllability ensures the alignment of generated content with corresponding annotations. Our approach integrates a novel 4D attention and a two-stage generation pipeline to maintain coherence, supplemented by the ControlNet framework for meticulous control by the Bird'View (BEV) layouts. Extensive qualitative and quantitative evaluations of Panacea on the nuScenes dataset prove its effectiveness in generating high-quality multi-view driving-scene videos. This work notably propels the field of autonomous driving by effectively augmenting the training dataset used for advanced BEV perception techniques. Yuqing Wen, Yingfei Liu, Fan Jia 0006, Chong Luo 0001, Chi Zhang 0026, Tiancai Wang, Xiaoyan Sun 0001, Xiangyu Zhang 0005 |
CVPR | 6 |
| 2024 | Glyph-ByT5: A Customized Text Encoder for Accurate Visual Text Rendering
Weicong Liang, Chong Luo 0001, Ji Li 0006, Gao Huang 0001, Yuhui Yuan |
ECCV (75) | 4 |
| 2024 | Aligning Vision Models with Human Aesthetics in Retrieval: Benchmarks and AlgorithmsabstractModern vision models are trained on very large noisy datasets. While these models acquire strong capabilities, they may not follow the user's intent to output the desired results in certain aspects, e.g., visual aesthetic, preferred style, and responsibility. In this paper, we target the realm of visual aesthetics and aim to align vision models with human aesthetic standards in a retrieval system. Advanced retrieval systems usually adopt a cascade of aesthetic models as re-rankers or filters, which are limited to low-level features like saturation and perform poorly when stylistic, cultural or knowledge contexts are involved. We find that utilizing the reasoning ability of large language models (LLMs) to rephrase the search query and extend the aesthetic expectations can make up for this shortcoming. Based on the above findings, we propose a preference-based reinforcement learning method that fine-tunes the vision models to distill the knowledge from both LLMs reasoning and the aesthetic models to better align the vision models with human aesthetics. Meanwhile, with rare benchmarks designed for evaluating retrieval systems, we leverage large multi-modality model (LMM) to evaluate the aesthetic performance with their strong abilities. As aesthetic assessment is one of the most subjective tasks, to validate the robustness of LMM, we further propose a novel dataset named HPIR to benchmark the alignment with human aesthetics. Experiments demonstrate that our method significantly enhances the aesthetic behaviors of the vision models, under several metrics. We believe the proposed algorithm can be a general practice for aligning vision models with human values. Miaosen Zhang, Yixuan Wei, Zuxuan Wu, Ji Li 0006, Zheng Zhang 0022, Qi Dai 0001, Chong Luo 0001, Xin Geng 0001, Baining Guo |
NeurIPS | 9 |
| 2024 | A Benchmark for Controllable Text -Image-to-Video GenerationabstractAutomatic video generation is a challenging research topic, attracting interests from different perspectives, including Image-to-Video generation (I2V), Video-to-Video generation (V2V), and Text-to-Video generation (T2V). To pursue more controllable and fine-grained video generation, a novel video generation task, named Text-Image-to-Video generation (TI2V), and a corresponding baseline solution, named Motion Anchor-based video Generator (MAGE), were proposed. However, two other factors, namely clean datasets and reliable evaluation metrics, also play important roles in the success of the TI2V task. In this article, we present a complete benchmark for the TI2V task which includes synthetic video-text paired datasets, a baseline method, and two evaluation metrics. More specifically: (1) Two versions of synthetic datasets are built based on CATER containing rich combinations of objects and actions, as well as the resulting changes of brightness and shadow. We also provide both explicit and ambiguous text descriptions to support deterministic and diverse video generation, respectively. (2) A refined version of MAGE, dubbed MAGE+, is proposed with an innovative motion anchor structure to store appearance-motion aligned representation, which can be further injected with explicit condition and implicit randomness to model the uncertainty in data distribution. (3) To evaluate the quality of generated video especially given ambiguous description, we introduce action precision and referring expression precision to assess the quality of motion based on captioning-and-matching method. Experiments conducted on proposed datasets, as well as relevant datasets, verify the effectiveness of our baseline and show appealing potentials of TI2V task. Yaosi Hu, Chong Luo 0001, Zhenzhong Chen 0001 |
IEEE Trans. Multim. | 2 |
| 2023 | Look Before You Match: Instance Understanding Matters in Video Object SegmentationabstractExploring dense matching between the current frame and past frames for long-range context modeling, memory-based methods have demonstrated impressive results in video object segmentation (VOS) recently. Nevertheless, due to the lack of instance understanding ability, the above approaches are oftentimes brittle to large appearance variations or viewpoint changes resulted from the movement of objects and cameras. In this paper, we argue that instance understanding matters in VOS, and integrating it with memory-based matching can enjoy the synergy, which is intuitively sensible from the definition of VOS task, i.e., identifying and segmenting object instances within the video. Towards this goal, we present a two-branch network for VOS, where the query-based instance segmentation (IS) branch delves into the instance details of the current frame and the VOS branch performs spatial-temporal matching with the memory bank. We employ the well-learned object queries from IS branch to inject instance-specific information into the query key, with which the instance-augmented matching is further performed. In addition, we introduce a multi-path fusion block to effectively combine the memory readout with multi-scale features from the instance segmentation decoder, which incorporates high-resolution instance-aware features to produce final segmentation results. Our method achieves state-of-the-art performance on DAVIS 2016/2017 val (92.6% and 87.1%), DAVIS 2017 test-dev (82.8%), and YouTube-VOS 2018/2019 val (86.3% and 86.3%), outperforming alternative methods by clear margins. Dongdong Chen 0001, Zuxuan Wu, Chong Luo 0001, Chuanxin Tang, Xiyang Dai, Yujia Xie, Lu Yuan 0001, Yu-Gang Jiang 0001 |
CVPR | 4 |
| 2023 | Streaming Video ModelabstractVideo understanding tasks have traditionally been modeled by two separate architectures, specially tailored for two distinct tasks. Sequence-based video tasks, such as action recognition, use a video backbone to directly extract spatiotemporal features, while frame-based video tasks, such as multiple object tracking (MOT), rely on single fixed-image backbone to extract spatial features. In contrast, we propose to unify video understanding tasks into one novel streaming video architecture, referred to as Streaming Vision Transformer (S-ViT). S-ViT first produces frame-level features with a memory-enabled temporally-aware spatial encoder to serve the frame-based video tasks. Then the frame features are input into a task-related temporal decoder to obtain spatiotemporal features for sequence-based tasks. The efficiency and efficacy of S-ViT is demonstrated by the state-of-the-art accuracy in the sequence-based action recognition task and the competitive advantage over conventional architecture in the frame-based MOT task. We believe that the concept of streaming video model and the implementation of S-ViT are solid steps towards a unified deep learning architecture for video understanding. Code will be available at https://github.com/yuzhms/Streaming-Video-Model. Chong Luo 0001, Chuanxin Tang, Dongdong Chen 0001, Noel Codella, Zhengjun Zha |
CVPR | 2 |
| 2023 | Filler Word Detection with Hard Category Mining and Inter-Category Focal LossabstractFiller words like "um" or "uh" are common in spontaneous speech. It is desirable to automatically detect and remove them in recordings, as they affect the fluency, confidence, and professionalism of speech. Previous studies and our preliminary experiments reveal that the biggest challenge in filler word detection is that fillers can be easily confused with other hard categories like "a" or "I". In this paper, we propose a novel filler word detection method that effectively addresses this challenge by adding auxiliary categories dynamically and applying an additional inter-category focal loss. The auxiliary categories force the model to explicitly model the confusing words by mining hard categories. In addition, inter-category focal loss adaptively adjusts the penalty weight between "filler" and "non-filler" categories to deal with other confusing words left in the "non-filler" category. Our system achieves the best results, with a huge improvement compared to other methods on the PodcastFillers dataset. Zhiyuan Zhao 0001, Chuanxin Tang, Dacheng Yin, Chong Luo 0001 |
ICASSP | 6 |
| 2023 | Attention-Guided Contrastive Masked Image Modeling for Transformer-Based Self-Supervised LearningabstractSelf-supervised learning with vision transformer (ViT) has gained much attention recently. Most existing methods rely on either contrastive learning or masked image modeling. The former is suitable for global feature extraction but underperforms in fine-grained tasks. The later explores the internal structure of images but ignores the high information sparsity and unbalanced information distribution. In this paper, we propose a new approach called Attention-guided Contrastive Masked Image Modeling (ACoMIM), which integrates the merits of both paradigms and leverages the attention mechanism of ViT for effective representation. Specifically, it has two pretext tasks, predicting the features of masked regions guided by attention and comparing the global features of masked and unmasked images. We show that these two pretext tasks complement each other and improve our method’s performance. The experiments demonstrate that our model transfers well to various downstream tasks such as classification and object detection. Code is available at https://github.com/yczhan/ACoMIM. Yucheng Zhan, Chong Luo 0001, Yueyi Zhang 0001, Xiaoyan Sun 0001 |
ICIP | 3 |
| 2023 | TridentSE: Guiding Speech Enhancement with 32 Global Tokens
Dacheng Yin, Zhiyuan Zhao 0001, Chuanxin Tang, Zhiwei Xiong, Chong Luo 0001 |
INTERSPEECH | 5 |
| 2022 | Sparse MLP for Image Recognition: Is Self-Attention Really Necessary?abstractTransformers have sprung up in the field of computer vision. In this work, we explore whether the core self-attention module in Transformer is the key to achieving excellent performance in image recognition. To this end, we build an attention-free network called sMLPNet based on the existing MLP-based vision models. Specifically, we replace the MLP module in the token-mixing step with a novel sparse MLP (sMLP) module. For 2D image tokens, sMLP applies 1D MLP along the axial directions and the parameters are shared among rows or columns. By sparse connection and weight sharing, sMLP module significantly reduces the number of model parameters and computational complexity, avoiding the common over-fitting problem that plagues the performance of MLP-like models. When only trained on the ImageNet-1K dataset, the proposed sMLPNet achieves 81.9% top-1 accuracy with only 24M parameters, which is much better than most CNNs and vision Transformers under the same model size constraint. When scaling up to 66M parameters, sMLPNet achieves 83.4% top-1 accuracy, which is on par with the state-of-the-art Swin Transformer. The success of sMLPNet suggests that the self-attention mechanism is not necessarily a silver bullet in computer vision. The code and models are publicly available at https://github.com/microsoft/SPACH. Chuanxin Tang, Guangting Wang, Chong Luo 0001, Wenxuan Xie, Wenjun Zeng 0001 |
AAAI | 4 |
| 2022 | When Shift Operation Meets Vision Transformer: An Extremely Simple Alternative to Attention MechanismabstractAttention mechanism has been widely believed as the key to success of vision transformers (ViTs), since it provides a flexible and powerful way to model spatial relationships. However, is the attention mechanism truly an indispensable part of ViT? Can it be replaced by some other alternatives? To demystify the role of attention mechanism, we simplify it into an extremely simple case: ZERO FLOP and ZERO parameter. Concretely, we revisit the shift operation. It does not contain any parameter or arithmetic calculation. The only operation is to exchange a small portion of the channels between neighboring features. Based on this simple operation, we construct a new backbone network, namely ShiftViT, where the attention layers in ViT are substituted by shift operations. Surprisingly, ShiftViT works quite well in several mainstream tasks, e.g., classification, detection, and segmentation. The performance is on par with or even better than the strong baseline Swin Transformer. These results suggest that the attention mechanism might not be the vital factor that makes ViT successful. It can be even replaced by a zero-parameter operation. We should pay more attentions to the remaining parts of ViT in the future work. Code is available at github.com/microsoft/SPACH. Guangting Wang, Chuanxin Tang, Chong Luo 0001, Wenjun Zeng 0001 |
AAAI | 4 |
| 2022 | Make It Move: Controllable Image-to-Video Generation with Text DescriptionsabstractGenerating controllable videos conforming to user intentions is an appealing yet challenging topic in computer vision. To enable maneuverable control in line with user intentions, a novel video generation task, named Text-Image-to-Video generation (TI2V), is proposed. With both controllable appearance and motion, TI2V aims at generating videos from a static image and a text description. The key challenges of TI2V task lie both in aligning appearance and motion from different modalities, and in handling uncertainty in text descriptions. To address these challenges, we propose a Motion Anchor-based video GEnerator (MAGE) with an innovative motion anchor (MA) structure to store appearance-motion aligned representation. To model the uncertainty and increase the diversity, it further allows the injection of explicit condition and implicit randomness. Through three-dimensional axial transformers, MA is interacted with given image to generate next frames recursively with satisfying controllability and diversity. Accompanying the new task, we build two new video-text paired datasets based on MNIST and CATER for evaluation. Experiments conducted on these datasets verify the effectiveness of MAGE and show appealing potentials of TI2V task. Datasets are available at https://github.com/Youncy-Hu/MAGE. Yaosi Hu, Chong Luo 0001, Zhenzhong Chen 0001 |
CVPR | 2 |
| 2022 | Retriever: Learning Content-Style Representation as a Token-Level Bipartite Graph
Dacheng Yin, Xuanchi Ren, Chong Luo 0001, Yuwang Wang, Zhiwei Xiong, Wenjun Zeng 0001 |
ICLR | 3 |
| 2022 | RetrieverTTS: Modeling Decomposed Factors for Text-Based Speech InsertionabstractThis paper proposes a new "decompose-and-edit" paradigm for the text-based speech insertion task that facilitates arbitrarylength speech insertion and even full sentence generation.In the proposed paradigm, global and local factors in speech are explicitly decomposed and separately manipulated to achieve high speaker similarity and continuous prosody.Specifically, we proposed to represent the global factors by multiple tokens, which are extracted by cross-attention operation and then injected back by link-attention operation.Due to the rich representation of global factors, we manage to achieve high speaker similarity in a zero-shot manner.In addition, we introduce a prosody smoothing task to make the local prosody factor context-aware and therefore achieve satisfactory prosody continuity.We further achieve high voice quality with an adversarial training stage.In the subjective test, our method achieves state-of-the-art performance in both naturalness and similarity.Audio samples can be found at https://ydcustc.github.io/retrieverTTS-demo/. Dacheng Yin, Chuanxin Tang, Xiaoqiang Wang 0006, Zhiyuan Zhao 0001, Zhiwei Xiong, Sheng Zhao 0002, Chong Luo 0001 |
INTERSPEECH | 9 |
| 2022 | An Anchor-Free Detector for Continuous Speech Keyword SpottingabstractContinuous Speech Keyword Spotting (CSKWS) is a task to detect predefined keywords in a continuous speech.In this paper, we regard CSKWS as a one-dimensional object detection task and propose a novel anchor-free detector, named AF-KWS, to solve the problem.AF-KWS directly regresses the center locations and lengths of the keywords through a single-stage deep neural network.In particular, AF-KWS is tailored for this speech task as we introduce an auxiliary unknown class to exclude other words from non-speech or silent background.We have built two benchmark datasets named LibriTop-20 and continuous meeting analysis keywords (CMAK) dataset for CSKWS.Evaluations on these two datasets show that our proposed AF-KWS outperforms reference schemes by a large margin, and therefore provides a decent baseline for future research. Zhiyuan Zhao 0001, Chuanxin Tang, Chengdong Yao, Chong Luo 0001 |
INTERSPEECH | 4 |
| 2022 | Peripheral Vision TransformerabstractHuman vision possesses a special type of visual processing systems called peripheral vision. Partitioning the entire visual field into multiple contour regions based on the distance to the center of our gaze, the peripheral vision provides us the ability to perceive various visual features at different regions. In this work, we take a biologically inspired approach and explore to model peripheral vision in deep neural networks for visual recognition. We propose to incorporate peripheral position encoding to the multi-head self-attention layers to let the network learn to partition the visual field into diverse peripheral regions given training data. We evaluate the proposed network, dubbed PerViT, on ImageNet-1K and systematically investigate the inner workings of the model for machine perception, showing that the network learns to perceive visual data similarly to the way that human vision does. The performance improvements in image classification over the baselines across different model sizes demonstrate the efficacy of the proposed method. Juhong Min, Chong Luo 0001, Minsu Cho |
NeurIPS | 3 |
| 2022 | OmniVL: One Foundation Model for Image-Language and Video-Language TasksabstractThis paper presents OmniVL, a new foundation model to support both image-language and video-language tasks using one universal architecture. It adopts a unified transformer-based visual encoder for both image and video inputs, and thus can perform joint image-language and video-language pretraining. We demonstrate, for the first time, such a paradigm benefits both image and video tasks, as opposed to the conventional one-directional transfer (e.g., use image-language to help video-language). To this end, we propose a \emph{decoupled} joint pretraining of image-language and video-language to effectively decompose the vision-language modeling into spatial and temporal dimensions and obtain performance boost on both image and video tasks. Moreover, we introduce a novel unified vision-language contrastive (UniVLC) loss to leverage image-text, video-text, image-label (e.g., image classification), video-label (e.g., video action recognition) data together, so that both supervised and noisily supervised pretraining data are utilized as much as possible. Without incurring extra task-specific adaptors, OmniVL can simultaneously support visual only tasks (e.g., image classification, video action recognition), cross-modal alignment tasks (e.g., image/video-text retrieval), and multi-modal understanding and generation tasks (e.g., image/video question answering, captioning). We evaluate OmniVL on a wide range of downstream tasks and achieve state-of-the-art or competitive results with similar model size and data scale. Dongdong Chen 0001, Zuxuan Wu, Chong Luo 0001, Luowei Zhou, Yujia Xie, Ce Liu 0001, Yu-Gang Jiang 0001, Lu Yuan 0001 |
NeurIPS | 4 |
| 2022 | Decomposing style, content, and motion for videos
Yaosi Hu, Dacheng Yin, Yuwang Wang, Zhenzhong Chen 0001, Chong Luo 0001 |
J. Vis. Commun. Image Represent. | 5 |
| 2021 | Unsupervised Visual Representation Learning by Tracking Patches in VideoabstractInspired by the fact that human eyes continue to develop tracking ability in early and middle childhood, we propose to use tracking as a proxy task for a computer vision system to learn the visual representations. Modelled on the Catch game played by the children, we design a Catch-the-Patch (CtP) game for a 3D-CNN model to learn visual representations that would help with video-related tasks. In the proposed pretraining framework, we cut an image patch from a given video and let it scale and move according to a pre-set trajectory. The proxy task is to estimate the position and size of the image patch in a sequence of video frames, given only the target bounding box in the first frame. We discover that using multiple image patches simultaneously brings clear benefits. We further increase the difficulty of the game by randomly making patches invisible. Extensive experiments on mainstream benchmarks demonstrate the superior performance of CtP against other video pretraining methods. In addition, CtP-pretrained features are less sensitive to domain gaps than those trained by a supervised action recognition task. When both trained on Kinetics-400, we are pleasantly surprised to find that CtP-pretrained representation achieves much higher action classification accuracy than its fully supervised counterpart on Something-Something dataset. Guangting Wang, Yizhou Zhou, Chong Luo 0001, Wenxuan Xie, Wenjun Zeng 0001, Zhiwei Xiong |
CVPR | 3 |
| 2021 | Self-Supervised Visual Representations Learning by Contrastive Mask PredictionabstractAdvanced self-supervised visual representation learning methods rely on the instance discrimination (ID) pretext task. We point out that the ID task has an implicit semantic consistency (SC) assumption, which may not hold in unconstrained datasets. In this paper, we propose a novel contrastive mask prediction (CMP) task for visual representation learning and design a mask contrast (MaskCo) framework to implement the idea. MaskCo contrasts region-level features instead of view-level features, which makes it possible to identify the positive sample without any assumptions. To solve the domain gap between masked and unmasked features, we design a dedicated mask prediction head in MaskCo. This module is shown to be the key to the success of the CMP. We evaluated MaskCo on training datasets beyond ImageNet and compare its performance with MoCo V2 [4]. Results show that MaskCo achieves comparable performance with MoCo V2 using ImageNet training dataset, but demonstrates a stronger performance across a range of downstream tasks when COCO or Conceptual Captions are used for training. MaskCo provides a promising alternative to the ID-based methods for self-supervised learning in the wild. Guangting Wang, Chong Luo 0001, Wenjun Zeng 0001, Zhengjun Zha |
ICCV | 3 |
| 2021 | Zero-Shot Text-to-Speech for Text-Based Insertion in Audio NarrationabstractGiven a piece of speech and its transcript text, text-based speech editing aims to generate speech that can be seamlessly inserted into the given speech by editing the transcript.Existing methods adopt a two-stage approach: synthesize the input text using a generic text-to-speech (TTS) engine and then transform the voice to the desired voice using voice conversion (VC).A major problem of this framework is that VC is a challenging problem which usually needs a moderate amount of parallel training data to work satisfactorily.In this paper, we propose a one-stage context-aware framework to generate natural and coherent target speech without any training data of the target speaker.In particular, we manage to perform accurate zero-shot duration prediction for the inserted text.The predicted duration is used to regulate both text embedding and speech embedding.Then, based on the aligned cross-modality input, we directly generate the mel-spectrogram of the edited speech with a transformer-based decoder.Subjective listening tests show that despite the lack of training data for the speaker, our method has achieved satisfactory results.It outperforms a recent zero-shot TTS engine by a large margin. Chuanxin Tang, Chong Luo 0001, Zhiyuan Zhao 0001, Dacheng Yin, Wenjun Zeng 0001 |
Interspeech | 2 |
| 2020 | PHASEN: A Phase-and-Harmonics-Aware Speech Enhancement NetworkabstractTime-frequency (T-F) domain masking is a mainstream approach for single-channel speech enhancement. Recently, focuses have been put to phase prediction in addition to amplitude prediction. In this paper, we propose a phase-and-harmonics-aware deep neural network (DNN), named PHASEN, for this task. Unlike previous methods which directly use a complex ideal ratio mask to supervise the DNN learning, we design a two-stream network, where amplitude stream and phase stream are dedicated to amplitude and phase prediction. We discover that the two streams should communicate with each other, and this is crucial to phase prediction. In addition, we propose frequency transformation blocks to catch long-range correlations along the frequency axis. Visualization shows that the learned transformation matrix implicitly captures the harmonic correlation, which has been proven to be helpful for T-F spectrogram reconstruction. With these two innovations, PHASEN acquires the ability to handle detailed phase patterns and to utilize harmonic patterns, getting 1.76dB SDR improvement on AVSpeech + AudioSet dataset. It also achieves significant gains over Google's network on this dataset. On Voice Bank + DEMAND dataset, PHASEN outperforms previous methods by a large margin on four metrics. Dacheng Yin, Chong Luo 0001, Zhiwei Xiong, Wenjun Zeng 0001 |
AAAI | 2 |
| 2020 | Posterior-Guided Neural Architecture SearchabstractThe emergence of neural architecture search (NAS) has greatly advanced the research on network design. Recent proposals such as gradient-based methods or one-shot approaches significantly boost the efficiency of NAS. In this paper, we formulate the NAS problem from a Bayesian perspective. We propose explicitly estimating the joint posterior distribution over pairs of network architecture and weights. Accordingly, a hybrid network representation is presented which enables us to leverage the Variational Dropout so that the approximation of the posterior distribution becomes fully gradient-based and highly efficient. A posterior-guided sampling method is then presented to sample architecture candidates and directly make evaluations. As a Bayesian approach, our posterior-guided NAS (PGNAS) avoids tuning a number of hyper-parameters and enables a very effective architecture sampling in posterior probability space. Interestingly, it also leads to a deeper insight into the weight sharing used in the one-shot NAS and naturally alleviates the mismatch between the sampled architecture and weights caused by the weight sharing. We validate our PGNAS method on the fundamental image classification task. Results on Cifar-10, Cifar-100 and ImageNet show that PGNAS achieves a good trade-off between precision and speed of search among NAS methods. For example, it takes 11 GPU days to search a very competitive architecture with 1.98% and 14.28% test errors on Cifar10 and Cifar100, respectively. Yizhou Zhou, Xiaoyan Sun 0001, Chong Luo 0001, Zhengjun Zha, Wenjun Zeng 0001 |
AAAI | 3 |
| 2020 | Tracking by Instance Detection: A Meta-Learning ApproachabstractWe consider the tracking problem as a special type of object detection problem, which we call instance detection. With proper initialization, a detector can be quickly converted into a tracker by learning the new instance from a single image. We find that model-agnostic meta-learning (MAML) offers a strategy to initialize the detector that satisfies our needs. We propose a principled three-step approach to build a high-performance tracker. First, pick any modern object detector trained with gradient descent. Second, conduct offline training (or initialization) with MAML. Third, perform domain adaptation using the initial frame. We follow this procedure to build two trackers, named Retina-MAML and FCOS-MAML, based on two modern detectors RetinaNet and FCOS. Evaluations on four benchmarks show that both trackers are competitive against state-of-the-art trackers. On OTB-100, Retina-MAML achieves the highest ever AUC of 0.712. On TrackingNet, FCOS-MAML ranks the first on the leader board with an AUC of 0.757 and the normalized precision of 0.822. Both trackers run in real-time at 40 FPS. Guangting Wang, Chong Luo 0001, Xiaoyan Sun 0001, Zhiwei Xiong, Wenjun Zeng 0001 |
CVPR | 2 |
| 2020 | Spatiotemporal Fusion in 3D CNNs: A Probabilistic ViewabstractDespite the success in still image recognition, deep neural networks for spatiotemporal signal tasks (such as human action recognition in videos) still suffers from low efficacy and inefficiency over the past years. Recently, human experts have put more efforts into analyzing the importance of different components in 3D convolutional neural networks (3D CNNs) to design more powerful spatiotemporal learning backbones. Among many others, spatiotemporal fusion is one of the essentials. It controls how spatial and temporal signals are extracted at each layer during inference. Previous attempts usually start by ad-hoc designs that empirically combine certain convolutions and then draw conclusions based on the performance obtained by training the corresponding networks. These methods only support network-level analysis on limited number of fusion strategies. In this paper, we propose to convert the spatiotemporal fusion strategies into a probability space, which allows us to perform network-level evaluations of various fusion strategies without having to train them separately. Besides, we can also obtain fine-grained numerical information such as layer-level preference on spatiotemporal fusion within the probability space. Our approach greatly boosts the efficiency of analyzing spatiotemporal fusion. Based on the probability space, we further generate new fusion strategies which achieve the state-of-the-art performance on four well-known action recognition datasets. Yizhou Zhou, Xiaoyan Sun 0001, Chong Luo 0001, Zhengjun Zha, Wenjun Zeng 0001 |
CVPR | 3 |
| 2020 | Joint Time-Frequency and Time Domain Learning for Speech EnhancementabstractFor single-channel speech enhancement, both time-domain and time-frequency-domain methods have their respective pros and cons. In this paper, we present a cross-domain framework named TFT-Net, which takes time-frequency spectrogram as input and produces time-domain waveform as output. Such a framework takes advantage of the knowledge we have about spectrogram and avoids some of the drawbacks that T-F-domain methods have been suffering from. In TFT-Net, we design an innovative dual-path attention block (DAB) to fully exploit correlations along the time and frequency axes. We further discover that a sample-independent DAB (SDAB) achieves a good tradeoff between enhanced speech quality and complexity. Ablation studies show that both the cross-domain design and the SDAB block bring large performance gain. When logarithmic MSE is used as the training criteria, TFT-Net achieves the highest SDR and SSNR among state-of-the-art methods on two major speech enhancement benchmarks. Chuanxin Tang, Chong Luo 0001, Zhiyuan Zhao 0001, Wenxuan Xie, Wenjun Zeng 0001 |
IJCAI | 2 |
| 2020 | Multi-Scale Group Transformer for Long Sequence Modeling in Speech SeparationabstractIn this paper, we introduce Transformer to the time-domain methods for single-channel speech separation. Transformer has the potential to boost speech separation performance because of its strong sequence modeling capability. However, its computational complexity, which grows quadratically with the sequence length, has made it largely inapplicable to speech applications. To tackle this issue, we propose a novel variation of Transformer, named multi-scale group Transformer (MSGT). The key ideas are group self-attention, which significantly reduces the complexity, and multi-scale fusion, which retains Transform's ability to capture long-term dependency. We implement two versions of MSGT with different complexities, and apply them to a well-known time-domain speech separation method called Conv-TasNet. By simply replacing the original temporal convolutional network (TCN) with MSGT, our approach called MSGT-TasNet achieves a large gain over Conv-TasNet on both WSJ0-2mix and WHAM! benchmarks. Without bells and whistles, the performance of MSGT-TasNet is already on par with the SOTA methods. Chong Luo 0001, Zhengjun Zha, Wenjun Zeng 0001 |
IJCAI | 2 |
| 2019 | SPM-Tracker: Series-Parallel Matching for Real-Time Visual Object TrackingabstractThe greatest challenge facing visual object tracking is the simultaneous requirements on robustness and discrimination power. In this paper, we propose a SiamFC-based tracker, named SPM-Tracker, to tackle this challenge. The basic idea is to address the two requirements in two separate matching stages. Robustness is strengthened in the coarse matching (CM) stage through generalized training while discrimination power is enhanced in the fine matching (FM) stage through a distance learning network. The two stages are connected in series as the input proposals of the FM stage are generated by the CM stage. They are also connected in parallel as the matching scores and box location refinements are fused to generate the final results. This innovative series-parallel structure takes advantage of both stages and results in superior performance. The proposed SPM-Tracker, running at 120 fps on GPU, achieves an AUC of 0.687 on OTB-100 and an EAO of 0.434 on VOT-16, exceeding other real-time trackers by a notable margin. Guangting Wang, Chong Luo 0001, Zhiwei Xiong, Wenjun Zeng 0001 |
CVPR | 2 |
| 2018 | A Twofold Siamese Network for Real-Time Object TrackingabstractObserving that Semantic features learned in an image classification task and Appearance features learned in a similarity matching task complement each other, we build a twofold Siamese network, named SA-Siam, for real-time object tracking. SA-Siam is composed of a semantic branch and an appearance branch. Each branch is a similaritylearning Siamese network. An important design choice in SA-Siam is to separately train the two branches to keep the heterogeneity of the two types of features. In addition, we propose a channel attention mechanism for the semantic branch. Channel-wise weights are computed according to the channel activations around the target position. While the inherited architecture from SiamFC [3] allows our tracker to operate beyond real-time, the twofold design and the attention mechanism significantly improve the tracking performance. The proposed SA-Siam outperforms all other real-time trackers by a large margin on OTB-2013/50/100 benchmarks. Anfeng He, Chong Luo 0001, Xinmei Tian 0001, Wenjun Zeng 0001 |
CVPR | 2 |
| 2018 | Multiple Level Feature-Based Universal Blind Image Quality Assessment ModelabstractThe direct use of a deep convolutional neural network (CNN) in no-reference image quality assessment (NR-IQA) usually struggles for a good performance due to a lack of training data, which can be alleviated by transfer learning. However, depending on the similarity between the source and target tasks, the final performance differs vastly. In particular, various kinds of distortion types exist in IQA, which requires different kinds of features to predict visual quality. In this paper, to make the transferred model robust to various distortion types, we propose a Multiple-level Feature-based Image Quality Assessor (MFIQA) which considers multiple levels of features simultaneously. Through rigorous experiments, we prove that MFIQA consistently yields state-of-the-art performance regardless of the distortion types including synthetic and authentic corruption. Jongyoo Kim, Sewoong Ahn, Chong Luo 0001, Sanghoon Lee 0001 |
ICIP | 4 |
| 2018 | Cooperative Hybrid Digital-Analog Video Transmission in D2D NetworksabstractIn this paper, we propose a cooperative video transmission scheme in D2D networks. This research is motivated by the growing interests in hybrid digital-analog video transmissions and device-to-device (D2D) communications. The framework of D2D communications can be generally modeled as a three-node network. In this network, coset coding is used to allow the destination to exploit the correlations between the video signals received in two phases. We have done some work of further optimization to improve the video quality at destination in this network. First, we derive a closed form of the reconstruction error at the destination. This provides a theoretical foundation for finding the optimal quantization step size in coset coding. Then, based on the accurate analysis on the coset coding we design a new power allocation algorithm. Experimental results verify that our scheme outperforms the recently proposed WCVC and DCVC. Jian Shen 0002, Chong Luo 0001, Houqiang Li, Wenjun Zeng 0001 |
ICIP | 3 |
| 2018 | Cascade Mask Generation Framework for Fast Small Object DetectionabstractDetecting small objects is a challenging task. Existing CNN-based objection detection pipeline faces such a dilemma: using a high-resolution image as input incurs high computational cost, but using a low-resolution image as input loses the feature representation of small objects and therefore leads to low accuracy. In this work, we propose a cascade mask generation framework to tackle this issue. The proposed framework takes in multi-scale images as input and processes them in ascending order of the scale. Each processing stage outputs object proposals as well as a region-of-interest (RoI) mask for the next stage. With RoI convolution, the masked regions can be excluded from the computation in the next stage. The procedure continues until the largest scale image is processed. Finally, the object proposals generated from multiple scales are classified by a post classifier. Extensive experiments on Tsinghua-Tencent 100K traffic sign benchmark demonstrate that our approach achieves state-of-the-art small object detection performance at a significantly improved speed-accuracy tradeoff compared with previous methods. Guangting Wang, Zhiwei Xiong, Dong Liu 0002, Chong Luo 0001 |
ICME | 4 |
| 2018 | A Practical Hybrid Digital-Analog Scheme for Wireless Video TransmissionabstractWe propose a hybrid digital-analog framework for wireless video transmission, which benefits from both the high distortion-power performance of digital systems and the graceful performance degradation of analog systems. The proposed framework models video frames as a parallel Gaussian source, which is separated into digital and analog parts through scalar quantization. It features entropy coding and channel coding in digital transmission and power scaling in analog transmission. The key challenge in this framework is how to allocate the constrained power and bandwidth resources between and among digital and analog components to achieve minimal distortion at the receiver. Given the worst-case channel signal-to-noise ratio, we are able to derive a closed-form expression of the overall distortion. However, minimizing it is a mixed-integer non-linear programming problem, which is generally non-deterministic polynomial-time hard. By making reasonable and justified simplifications, we approach the optimal solution through a practical scheme. Evaluations show that the proposed scheme outperforms the state-of-the-art analog scheme SoftCast by a large margin. The gain in received video peak signal-to-noise ratio is up to 5.0 dB for various types of videos. Cuiling Lan, Chong Luo 0001, Wenjun Zeng 0001, Feng Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2018 | Superimposed Modulation for Soft Video Delivery With Hidden ResourcesabstractAnalog-transmission-based soft video delivery suffers from the leveling-off effect when the allocated channel bandwidth is severely insufficient. Fortunately, with superimposed modulation, it is possible for analog traffic to share bandwidth with digital traffic. In this paper, we design and analyze such a hybrid digital-analog superimposed modulation (HDA-SIM) scheme for soft video delivery. Unlike previous work, we treat the bandwidth of competing digital traffic as hidden resources for the video delivery system. The key problem in this scheme is how to allocate the bandwidth and power resources among various modulation symbols so that we can improve the performance of video delivery without sacrificing the throughput of existing digital traffic. The resource allocation problem is formulated and the optimal solution under any given channel signal-to-noise ratio is derived. Based on the results, the sufficient and necessary condition for the video delivery system to achieve performance gain is given. In addition, we implement the proposed scheme for two state-of-the-art soft video delivery systems known as SoftCast and SharpCast. Both simulations and testbed evaluations show that the HDA-SIM version can achieve significant gains in the received video quality over their original designs. Chong Luo 0001, Ruiqin Xiong, Wenjun Zeng 0001, Feng Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2018 | Optimizing Quality of Experience for Adaptive Bitrate Streaming via Viewer Interest InferenceabstractRate adaptation is widely adopted in video streaming to improve the quality of experience (QoE). However, most of the existing rate adaptation approaches neglect the underlying video semantic information. In fact, influenced by video semantics and viewer preferences, the viewer may have different degrees of interest on different parts of a video. The interesting parts of a video can draw more visual attention from the viewer and have higher visual importance. As such, delivering the parts of a video that are interesting to the viewer in a higher quality can improve the perceptual video quality, compared with the semantics-agnostic approaches that treat each part of a video equally. Thus, it is natural to wonder: how to allocate bitrate budgets temporally over a video session under time-varying bandwidth while considering viewer interest? As an exploratory study, we propose an interest-aware rate adaptation approach for improving QoE by inferring viewer interest based on video semantics. We adopt the deep learning method to recognize the scenes of video frames and leverage the term frequency-inverse document frequency method to analyze the degrees of an individual viewer's interest on different types of scenes. The bandwidth, buffer occupancy, and viewer interest are jointly considered under the model predictive control framework for selecting appropriate bitrates for maximizing QoE. The objective and subjective evaluations measured in a real environment show that our method can achieve a higher QoE compared with the semantics-agnostic approaches. Guanyu Gao, Huaizheng Zhang, Han Hu 0003, Yonggang Wen 0001, Jianfei Cai 0001, Chong Luo 0001, Wenjun Zeng 0001 |
IEEE Trans. Multim. | 6 |
| 2018 | Hybrid Digital-Analog Video Delivery With Shannon-Kotel'nikov MappingabstractHybrid digital-analog (HDA) transmission is becoming an attractive solution for mobile video delivery because it not only has graceful degradation with channel variations but also yields high power efficiency. However, the heavy bandwidth demand of analog transmission is still an unsolved problem, limiting the received video quality when the bandwidth is not sufficient. To address this problem, we propose adopting Shannon-Kotel'nikov (SK) mapping for HDA video transmission and design an HDA scheme called SK-Cast. SK-Cast consists of a digital and an analog branch. In the digital branch, SK-Cast compresses the video sequence using an high efficiency video coding digital encoder to produce a base layer. The base layer is transmitted through digital methods with strong protection. The residual signals are then decorrelated using three-dimensional discrete cosine transform transform. The SK mapping is exploited to transmit these coefficients, as they can achieve efficient bandwidth compression. We address the resource allocation problems in SK-Cast, including the allocation between digital and analog branches and the allocation among analog symbols. The simulation results show that the SK-Cast outperforms the state-of-the-art HDA systems, including WSVC and SharpCast, and a digital scalable video coding system. Chong Luo 0001, Ruiqin Xiong, Wenjun Zeng 0001, Feng Wu 0001 |
IEEE Trans. Multim. | 2 |
| 2018 | Cost-Distortion Optimization and Resource Control in Pseudo-Analog Visual CommunicationsabstractThe rate-distortion in conventional digital systems is replaced by cost distortion in pseudo-analog systems where the cost consists of power and bandwidth. In this paper, we formulate the cost-distortion optimization problem in terms of a power-bandwidth pair versus distortion to bring an insight to pseudo-analog transmission. Using a divide-and-conquer strategy, the 3-D optimization problem of a power-bandwidth pair versus distortion is decomposed into two subproblems: power distortion and bandwidth distortion optimization. To solve the integer nonlinear optimization problem, we propose two prediction models that transform the partial summation of variances and the square roots of variances into continuous functions. The proposed models are used to derive the closed-form solutions for both optimization subproblems, and a tradeoff between power and bandwidth is discussed. Accordingly, the resource control algorithm is designed to allocate the fewest resources required to obtain specific video quality; power can be traded for bandwidth and vice versa. Our experimental results show that the proposed optimization models achieve stable perceptual video quality comparable to that of Groups of Pictures and use resources more efficiently than do SoftCast. Dian Liu, Jun Wu 0006, Hao Cui 0001, Chong Luo 0001, Feng Wu 0001 |
IEEE Trans. Multim. | 5 |
| 2017 | A CNN-Based Approach for Automatic License Plate Recognition in the Wild
Meng Dong, Dongliang He, Chong Luo 0001, Dong Liu 0002, Wenjun Zeng 0001 |
BMVC | 3 |
| 2017 | Compressive gradient based scalable image SoftCastabstractIn wireless visual communication systems, it is crucial to effectively utilize channel power and bandwidth in the pursue of optimal performance, and it is worthwhile to adapt the transmission scheme to human vision system (HVS) so as to achieve perceptually appealing results. Inspired by observations that visual quality of an image is closely related to the gradient data, this paper proposes to convey visual information by random projection measurements of image gradients in an analog framework. Since HVS is more sensitive to luminance variations of image contents, which are contained in the gradient data, the proposed scheme achieves better perceptual quality than conventional analog uncoded schemes like SoftCast. Besides, the gradient transform removes the low and medium frequency components of the image hence substantially reduces the power of the signal transmitted in the analog channel, thus evidently improves the power-distortion performance of the system. Furthermore, by applying random projection to the gradients, the number of transmitted data can be adjusted according to bandwidth conditions. Another contribution of this paper is to develop an effective optimization scheme for the compressive gradient based reconstruction problem. Experimental results validate the effectiveness of the proposed transmission and reconstruction scheme under different channel signal-to-noise ratio and bandwidth conditions. Hangfan Liu, Ruiqin Xiong, Xiaopeng Fan 0001, Chong Luo 0001, Wen Gao 0001 |
VCIP | 4 |
| 2017 | Wireless image and video soft transmission via perception-inspired power distortion optimizationabstractRecently, a scheme called SoftCast has shown great potential for wireless image/video communication in the scenarios where the channel quality may fluctuate drastically and unpredictably. The transmission is lossy in nature, with its transmission power allocated among coefficients unequally to minimize the distortion. One problem is that its performance is optimized using mean square errors (MSE) as the quality metric, which is known for not matching the perception of human eyes. Inspired by the researches in image quality assessment, this paper proposes a power allocation and optimization scheme that minimizes the perceptual distortion of reconstruction image. In particular, we establish a perception model to evaluate the perceptual importance of different transform coefficients, based on the structure similarity (SSIM) image quality metric. Experimental results show that the proposed scheme can improve the perceptual performance of the original SoftCast scheme. Jing Zhao 0011, Ruiqin Xiong, Chong Luo 0001, Feng Wu 0001, Wen Gao 0001 |
VCIP | 3 |
| 2017 | Progressive Pseudo-analog Transmission for Mobile Video StreamingabstractWe propose a progressive pseudo-analog video transmission scheme that simultaneously handles SNR and bandwidth variations with graceful quality degradation for mobile video streaming. With the inherited SNR-adaptability from pseudo-analog transmission, the proposed progressive solution acquires bandwidth adaptability through an innovative scheduling algorithm with optimal power allocation. The basic idea is to aggressively transmit or retransmit important coefficients so that distortion is minimized at the receiver after each received packet. We derive the closed-form expression of reduced distortion for each packet under given transmission power and known channel conditions, and show that the optimal solution can be obtained with a water-filling algorithm. We also illustrate through analyses and simulations that a near-optimal solution can be found through approximation when only statistical channel information is available. Simulations show that our solution approaches the performance upper bound of pseudo-analog transmission in an additive white Gaussian noise channel and significantly outperforms existing pseudo-analog solutions in a fast Rayleigh fading channel. Trace-driven emulations are also carried out to demonstrate the advantage of the proposed solution over the state-of-the-art digital and pseudo-analog solutions under a real dramatically varying wireless environment. Dongliang He, Cuiling Lan, Chong Luo 0001, Enhong Chen, Feng Wu 0001, Wenjun Zeng 0001 |
IEEE Trans. Multim. | 3 |
| 2016 | Resource Allocation for Uncoded Multi-user Video Transmission over Wireless Networks
Dian Liu, Hao Cui 0001, Jun Wu 0006, Chong Luo 0001 |
Mob. Networks Appl. | 4 |
| 2016 | Scalable Video Multicast for MU-MIMO Systems With Antenna HeterogeneityabstractIn contemporary multiuser multiple-input multipleoutput systems, it is common for the reception devices to have a varying number of antennas. When multicast is performed, the number of concurrent spatial streams is limited by the device with the least number of antennas, which prevents more capable devices from getting higher rates. In this paper, we address the antenna heterogeneity in wireless video multicast by the innovative design of multiple similar description (MSD) video coding and multiplexed space-time block coding (M-STBC). MSD coding generates multiple descriptions of a video and features that any linear combinations of the descriptions are decodable. The descriptions comprising of real numbers are further processed by transform and power allocation steps for efficient transmission in a power-constrained system. M-STBC puts symbols in similar descriptions to corresponding space-time positions and ensures decodability under any antenna settings and channel conditions. As a result, we build up a scalable video multicast system, named AirScale, which allows receivers with a various number of antennas to decode from a single transmission, and the reconstructed video quality improves with the number of equipped antennas. Evaluations on Sora shows that, in a {1, 2, 3, 4} × 4 system, AirScale provides baseline quality for one-antenna receiver and a much higher quality for multiantenna receivers. The gain over SoftCast is up to 3.5, 3.9, and 4.1 dB for two-, three-, and four-antenna receivers, respectively. Hao Cui 0001, Chong Luo 0001, Chang Wen Chen, Feng Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2016 | Analysis of Decorrelation Transform Gain for Uncoded Wireless Image and Video CommunicationabstractAn uncoded transmission scheme called SoftCast has recently shown great potential for wireless video transmission. Unlike conventional approaches, SoftCast processes input images only by a series of transformations and modulates the coefficients directly to a dense constellation for transmission. The transmission is uncoded and lossy in nature, with its noise level commensurate with the channel condition. This paper presents a theoretical analysis for an uncoded visual communication, focusing on developing a quantitative measurements for the efficiency of decorrelation transform in a generalized uncoded transmission framework. Our analysis reveals that the energy distribution among signal elements is critical for the efficiency of uncoded transmission. A decorrelation transform can potentially bring a significant performance gain by boosting the energy diversity in signal representation. Numerical results on Markov random process and real image and video signals are reported to evaluate the performance gain of using different transforms in uncoded transmission. The analysis presented in this paper is verified by simulated SoftCast transmissions. This provide guidelines for designing efficient uncoded video transmission schemes. Ruiqin Xiong, Feng Wu 0001, Jizheng Xu, Xiaopeng Fan 0001, Chong Luo 0001, Wen Gao 0001 |
IEEE Trans. Image Process. | 5 |
| 2016 | DAC-Mobi: Data-Assisted Communications of Mobile Images with Cloud Computing SupportabstractThis research proposes a novel data assisted image transmission scheme, which utilizes a large amount of correlated images stored in the cloud to improve the spectrum efficiency and visual quality. First, a two-layer Coset coding is proposed for the DCT coefficients transmission. The most significant bits (MSB) of the coefficients are generated by the first layer Coset and together with a few low frequency coefficients are transmitted through the most reliable channel coding and digital modulation. The middle bits generated by the second layer Coset are discarded by the sender and the residual bits are transmitted through amplitude modulation. Based on the MSB and the residual bits, an approximation of the original image is reconstructed. With this approximation, a lot of correlated images can be retrieved from the cloud, which are used to recover the discarded middle bits. The two layer Coset coding can significantly decrease the data energy so as to improve the transmission power efficiency. Hence, the end to end distortion of amplitude modulation can be reduced. Second, the image quality can be further improved by joint internal and external denoising with the retrieved images. Simulations show that the proposed scheme outperforms conventional digital schemes about 4 dB in peak signal to noise power ratio (PSNR) and achieves 2 dB gain over the state-of-the-art uncoded transmission. At low signal to noise power ratio (SNR), an additional 2-3 dB gain is achieved. The visual quality comparison also validates the objective image assessment result. Jun Wu 0006, Jian Wu 0022, Hao Cui 0001, Chong Luo 0001, Xiaoyan Sun 0001, Feng Wu 0001 |
IEEE Trans. Multim. | 4 |
| 2015 | Swift: A Hybrid Digital-Analog Scheme for Low-Delay Transmission of Mobile Stereo VideoabstractEfficient and robust wireless stereo video delivery is an enabling technology for various mobile 3D applications. Existing digital solutions have high source coding efficiency but are not robust to channel variations, while analog solutions have the opposite characteristics. In this paper, we design a novel hybrid digital-analog (HDA) solution to embrace the advantages of both solutions and avoid their drawbacks. Basically, in each pair of stereo frames, one frame is digitally encoded to ensure basic quality and the other is analogly processed to opportunistically utilize good channels for better quality. To improve the system efficiency, we design a zigzag coding structure such that both intra-view and inter-view correlations can be explored through prediction in the frames to be analogly coded. A reference selection mechanism is proposed to further improve the coding efficiency. In addition, we address the problem of optimal power and bandwidth allocation between digital and analog streams. We implement a system, named Swift, and perform extensive trace-driven evaluations based on a software-defined radio platform. We show that Swift outperforms an omniscient digital scheme under the same bandwidth and power constraints, or can have around 2x power saving in order to achieve comparable performance. Subjective quality assessment evidences that Swift provides significantly better visual quality than a straightforward HDA extension of SoftCast. Dongliang He, Chong Luo 0001, Feng Wu 0001, Wenjun Zeng 0001 |
MSWiM | 2 |
| 2015 | Progressive pseudo-analog transmission for mobile video live streamingabstractMobile video live streaming is facing great challenges in offering high quality of experience (QoE) under varying channel conditions. In this paper, we propose a progressive pseudo-analog transmission scheme in which the received video quality gracefully adapts to both SNR and bandwidth variations. Building upon the emerging pseudoanalog video transmission, the proposed scheme further adopts a greedy approach to improve the received video quality with each allocated bandwidth share. The optimal scheduling and power allocation are derived under the mean squared error (MSE) criterion. Testbed evaluations show that the proposed scheme outperforms the state-of-the-art digital and analog transmission schemes by a notable margin. Cuiling Lan, Dongliang He, Chong Luo 0001, Feng Wu 0001, Wenjun Zeng 0001 |
VCIP | 3 |
| 2015 | Structure-Preserving Hybrid Digital-Analog Video Delivery in Wireless NetworksabstractHybrid digital-analog (HDA) transmission has gained increasing attention recently in the context of wireless video delivery , for its ability to simultaneously achieve high transmission efficiency and smooth quality adaptation. However, previous systems are optimized solely based on the mean squared error criterion without taking the perceptual video quality into consideration. In this work, we propose a structure-preserving HDA video delivery system, named SharpCast, to improve both the objective and subjective visual quality. SharpCast decomposes a video into a content part and structure part. The latter is important to the human perception and therefore is protected with a robust digital transmission scheme. Then, the energy-intensive part in the content information is extracted and transmitted in digital for energy efficiency while the residual is transmitted in analog to achieve the desired smooth adaptation. We formulate the resource (power and bandwidth) allocation problem in SharpCast and solve the problem with a greedy strategy. Evaluations over nine standard 720p video sequences show that the proposed SharpCast system outperforms the state-of-the-art digital, analog, and HDA schemes by a notable margin in both peak signal-to-noise ratio (PSNR) and structural similarity (SSIM). Dongliang He, Chong Luo 0001, Cuiling Lan, Feng Wu 0001, Wenjun Zeng 0001 |
IEEE Trans. Multim. | 2 |
| 2015 | Design and Analysis of Compressive Data Persistence in Large-Scale Wireless Sensor NetworksabstractThis paper addresses the data persistence problem in wireless sensor networks (WSNs) where static sinks are not present and the sensed data have to be temporarily but resiliently stored in the network. Based on the observation that sensor readings are correlated, we propose compressive data persistence (CDP) scheme that makes use of the compressive sensing (CS) theory. Each sensor node independently computes and stores a random projection of the sensed data, such that a mobile sink can recover the data with high probability after visiting a small and random portion of the network. As a prerequisite of distributed CS encoding, sensor readings from all nodes are disseminated within the network through random walk. Therefore, the CS measurement matrix depends heavily on how the random walk is performed. In this paper, we present an in-depth analysis on the interplay between random walk parameters and sensing data characteristics, and derive the conditions in successful CS data recovery. In addition, we discover that there is a trade-off between the number of random walk instances and steps in order to achieve the required data persistence performance. Experiments using real sensor data verify that the proposed CDP scheme achieves much lower decoding ratio than the state-of-the-art Fountain code based schemes or the decentralized erasure codes based schemes, and demonstrate that there exist energy-optimized random walk parameters for CDP. Feng Liu 0010, Mu Lin, Yusuo Hu, Chong Luo 0001, Feng Wu 0001 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2014 | Robust uncoded video transmission over wireless fast fading channelabstractThis research studies robust uncoded video transmission over wireless fast fading channel, where only statistical channel state information (CSI) is available at the transmitter. We observe that increasing channel diversity for high priority (HP) data is essential to improving the robustness of video transmission in fading channels. By utilizing the noise and loss resilient nature of video, we find it possible to design a more robust system by re-allocating the power and channel uses among HP and LP (low priority) data. With total power and channel use constraints, we derive an optimal resource allocation scheme under the squared error distortion criterion. In particular, we first propose a new power allocation algorithm at given channel allocation. Second, based on the proposed power allocation algorithm, we design a channel allocation algorithm to strike the tradeoff between the diversity increase of HP data and the information loss of LP data. Third, under known noise power distribution, we derive the optimal resource allocation for uncoded video multicast. Simulations show that the proposed system achieves 2dB and 5dB gain in average and outage PSNR over Softcast in video unicast, and around 1.4dB and 4dB gain in multicast. Hao Cui 0001, Chong Luo 0001, Chang Wen Chen, Feng Wu 0001 |
INFOCOM | 2 |
| 2014 | Compressive image broadcasting in MIMO systems with receiver antenna heterogeneity
Xiao Lin Liu, Chong Luo 0001, Feng Wu 0001 |
Signal Process. Image Commun. | 3 |
| 2014 | Robust Linear Video Transmission Over Rayleigh Fading ChannelabstractThis research addresses the problem of robust linear video transmission over the Rayleigh fading channel, where only statistical channel state information (CSI) is available to the sender. We observe that discarding low-priority (LP) data and saving the channel uses for high-priority (HP) data can significantly improve the quality of the received video. We formulate an optimization problem that aims to minimize the total squared error of a multi-variant Gaussian random vector under the given bandwidth and power resources. To tame the complexity of this NP-hard problem, we analyze two sub-problems, namely power allocation and bandwidth allocation, and propose an iterative algorithm to approximate the solution. Subsequently, we propose a one-pass two-step fast algorithm that further reduces both algorithmic and computational complexity. A linear video transmission system is implemented based on the proposed algorithm. Simulations show that our system significantly outperforms Soft-Cast, and the PSNR gain at 5th percentile of 1000 test runs is between 4.0 dB and 7.5 dB under varying noise levels. Hao Cui 0001, Chong Luo 0001, Chang Wen Chen, Feng Wu 0001 |
IEEE Trans. Commun. | 2 |
| 2014 | ParCast+: Parallel Video Unicast in MIMO-OFDM WLANsabstractWe have observed two trends, growing wireless capability at the physical layer powered by MIMO-OFDM and growing video traffic as the dominant application traffic. Both the video source and MIMO-OFDM channel components exhibit nonuniform energy distribution. This has motivated us to leverage the source data redundancy at the channel to achieve high video recovery performance. We propose ParCast+ that first separates the source and the channel into independent components, matches the more important source components with higher-gain channel components, allocates power weights with joint consideration to the source and the channel, and uses pseudo-analog modulation for transmission. Such a scheme achieves fine-grained unequal error protection across source components. We implemented ParCast+ in Matlab and on Sora. Extensive evaluation has shown that our scheme outperforms competing schemes by notable margins, sometimes up to 6.4 dB in PSNR for challenging scenarios. Xiao Lin Liu, Chong Luo 0001, Qifan Pu, Feng Wu 0001, Yongguang Zhang |
IEEE Trans. Multim. | 3 |
| 2013 | Cactus: a hybrid digital-analog wireless video communication systemabstractThis paper challenges the conventional wisdom that video redundancy should be removed as much as possible for efficient communications. We discover that, by keeping spatial redundancy at the sender and properly utilizing it at the receiver, we can build a more robust and even more efficient wireless video communication system than existing ones. Hao Cui 0001, Zhihai Song, Chong Luo 0001, Ruiqin Xiong, Feng Wu 0001 |
MSWiM | 4 |
| 2013 | Power-distortion optimization for wireless image/video SoftCast by transform coefficients energy modeling with adaptive chunk divisionabstractTraditional communication systems usually suffer from the threshold effect when channel signal-to-noise ratio (CSNR) fluctuates unpredictably in wireless and mobile scenarios. The SoftCast scheme, however, provides graceful quality transition in wide CSNR range. In SoftCast, input image is decorrelated by a transform and modulated directly to a dense constellation for transmission, leaving out the conventional quantization, entropy coding and channel coding. A key point of SoftCast is that the transmission power needs to be allocated among the transform coefficients unequally, according to the energy of coefficients. Importantly, the energy diversity used to guide power allocation should be shared between the sender and the receiver for correct decoding. This paper addresses the power distortion optimization problem, introducing a new adaptive chunk division scheme to describe the energy diversity among coefficients. A concrete algorithm is developed to determine the chunk boundaries that achieve optimal transmission power usage. Experimental results show that the proposed scheme can improve the performance of the original SoftCast by 4~8dB using a smaller number of chunks. Ruiqin Xiong, Feng Wu 0001, Xiaopeng Fan 0001, Chong Luo 0001, Siwei Ma 0001, Wen Gao 0001 |
VCIP | 4 |
| 2013 | Compressive Coded Modulation for Seamless Rate AdaptationabstractThis paper presents a novel compressive coded modulation (CCM) which simultaneously achieves joint source-channel coding and seamless rate adaptation. The embedding of source compression into modulation brings significant throughput gain when the physical layer data contain non-negligible redundancy. The kernel of CCM is a new random projection (RP) code inspired by the compressive sensing (CS) theory. The RP code generates multilevel symbols from source binaries through weighted sum operations. Then, the generated RP symbols are mapped into a dense constellation for transmission. The receiver performs joint decoding based on received symbols. As the number of RP symbols can be adjusted in fine granularity, the rate adaptation becomes seamless. Two key design issues in the proposed CCM are addressed in this paper. First, we consider the RP code design for sources with different redundancies. Three principles are established and a concrete implementation is given. Second, we devise a linear-time decoding algorithm for the proposed RP code. In this belief propagation (BP) algorithm, we find that computing convolution in time domain is more efficient than that in frequency domain for binary variable nodes. Moreover, we invent a ZigZag deconvolution to further reduce the complexity. Analysis show that the proposed decoding algorithm is nearly 20 times faster than the state-of-the-art BP algorithm for CS called CS-BP. Emulations on traced data show that CCM achieves significant throughput gain, up to 33% and 70%, respectively, over the Hybrid ARQ with compression and BICM with compression, under practical time-varying wireless channels. Hao Cui 0001, Chong Luo 0001, Jun Wu 0006, Chang Wen Chen, Feng Wu 0001 |
IEEE Trans. Wirel. Commun. | 2 |
| 2012 | Compressive broadcast in MIMO systems with receive antenna heterogeneityabstractThe key challenge in a broadcast system is receiver heterogeneity, where the weakest receiver typically constrains the entire system performance. Traditionally, this arises from heterogeneous channel SNRs at different receivers. In Multiple-Input Multiple-Output (MIMO) systems, a further heterogeneity is caused by different antenna numbers across receivers. We propose a compressive broadcast framework to address both types of heterogeneity. By layering compressive sensing (CS) over MIMO transmissions, our framework ensures a received source quality commensurate with the channel SNR and the MIMO channel dimension. Compared with a conventional framework, our framework achieves more smooth rate increase with channel SNR and much higher performance for multi-antenna receivers. Xiao Lin Liu, Chong Luo 0001, Feng Wu 0001 |
INFOCOM | 2 |
| 2011 | Resource allocation for cloud-based free viewpoint video rendering for mobile phonesabstractFree viewpoint video (FVV) / Free viewpoint TV (FTV) on mobile devices over cellular networks is very challenging due to the requirement for large bandwidth and limitations in computation and battery life on mobile phones. To address such challenges, in this paper we propose a cloud-based FVV / FTV rendering framework for mobile devices over cellular networks. In this framework, cloud performs rendering for mobile devices. In order to achieve maximum QoE (Quality of Experience) for mobile users, we propose a novel resource allocation scheme, which jointly considers rendering allocation between cloud and client based on user's QoE and rate allocation among texture, depth, and channel rate based on rate-distortion analysis. We formulate this resource allocation scheme as an optimization problem which can then be transformed into a convex optimization for the given rate ratio. Experimental results demonstrate that the proposed cloud-based FVV rendering solution can substantially improve video quality on mobile devices comparing with traditional approaches. Dan Miao, Wenwu Zhu 0001, Chong Luo 0001, Chang Wen Chen |
ACM Multimedia | 3 |
| 2011 | Seamless rate adaptation for wireless networkingabstractThis paper aims at designing a Seamless Rate Adaptation for wireless networking which achieves smooth rate adjustment in a broad dynamic range of channel conditions. Conventional rate adaptation can only achieve a stair-case rate adjustment. Even when combining with hybrid ARQ, it suffers from an irreconcilable conflict between throughput and dynamic range. We tackle this problem from a new perspective by relying on modulation, instead of channel coding, for rate adaptation. We propose rate compatible modulation (RCM), in which modulation signals are incrementally generated from information bits through weighted mapping. Rate adaptation is achieved through varying the number of modulated signals. As more signals are transmitted, information bits gradually accumulate energy. The weights in bit-to-symbol mapping are delicately designed to ensure fine-grained energy accumulation so that smoothness and efficiency can both be achieved. We design and implement a rate adaptation system, called SRA and evaluate its performance through a software radio testbed. Results show that, under highly dynamic channel conditions, SRA achieves over 80% throughput gain over 802.11a adaptive modulation and coding, and achieves 28.8% and 43.8% gain over HARQ systems implemented with Turbo code and Raptor code. We believe that the concept of rate compatible modulation opens up a fresh research avenue toward the wireless rate adaptation problem. Hao Cui 0001, Chong Luo 0001, Feng Wu 0001, Chang Wen Chen |
MSWiM | 2 |
| 2011 | MixCast modulation for layered video multicast over WLANsabstractThe major challenge in wireless multicast is the heterogeneous channel conditions of multiple users. In video multicast, the combination of a layered video coding scheme and a layered transmission scheme can gracefully accommodate user heterogeneity. This paper presents MixCast, a novel physical layer scheme for layered transmission. The key innovation in MixCast is the rateless Euclidean symbol mapping which mixes base layer and enhancement layer bits into arbitrary number of wireless symbols using arithmetic weighted sum operation. This design brings two benefits when compared with the state-of-the- art physical layer technique known as hierarchical modulation (HM). First, MixCast uses a fixed modulation constellation, avoiding the complexity in adaptive modulation and coding when channel condition varies. Second, the rateless symbol mapping allows MixCast to achieve much smoother rates than the stair-shaped rates in HM. We implemented MixCast for typical video multicast over OFDM physical layer, and evaluate its performance against HM through both simulations and software radio testbed. In simulations, MixCast shows consistent gain of 3dB to 5dB over HM under various rate combinations. In the testbed experiments, MixCast achieves significant gain up to 15dB in video PSNR over HM for football sequence. Hao Cui 0001, Chong Luo 0001, Chang Wen Chen, Feng Wu 0001 |
VCIP | 2 |
| 2010 | Compressive Data Persistence in Large-Scale Wireless Sensor NetworksabstractThis paper considers a large-scale wireless sensor network where sensor readings are occasionally collected by a mobile sink, and sensor nodes are responsible for temporarily storing their own readings in an energy-efficient and storage-efficient way. Existing data persistence schemes based on erasure codes do not utilize the correlation between sensor data, and their decoding ratio is always larger than one. Motivated by the emerging compressive sensing theory, we propose compressive data persistence which simultaneously achieves data compression and data persistence. In the development of compressive data persistence scheme, we design a distributed compressive sensing encoding approach based on Metropolis-Hastings random walk. When the maximum step of random walk is 400, our proposed scheme can achieve a decoding ratio of 0.36 for 10%-sparse data. We also compare our scheme with a state-of-the-art Fountain code based scheme. Simulation shows that our scheme can significantly reduce the decoding ratio by up to 63%. Mu Lin, Chong Luo 0001, Feng Liu 0010, Feng Wu 0001 |
GLOBECOM | 2 |
| 2010 | Joint decoding of stereo JPEG image PairsabstractThis paper addresses the problem of joint decoding of stereo JPEG image pairs. Such images typically contain a high degree of redundancy. Predictive coding could efficiently capture this redundancy, but cameras would have to implement proprietary encoding solutions in this case as no such standard technology is available. We propose to rather use the popular JPEG compression tools in the cameras, and focus on the joint decoding problem for quality enhancement. We formulate this as a constrained optimization problem and show how regularization leads to more consistent results. It is similar to a distributed source coding framework, where the exploitation of the correlation at the decoder permits to save on the overall bandwidth. Experiments on natural stereo images show an improvement in both visual quality and PSNR when compared to separate decoding. Markus B. Schenkel, Chong Luo 0001, Pascal Frossard, Feng Wu 0001 |
ICIP | 2 |
| 2010 | Stable Maximum Throughput Broadcast in Wireless Fading ChannelsabstractThis research considers network coded broadcast system with multi-rate transmission and dual queue stability constraints. Existing network coded broadcast systems consider single rate transmission without receiver queue constraints. First, we shall illustrate that broadcast without network coding cannot support maximum throughput in wireless fading channels. However, the network coded broadcast poses new constraints for the receivers to manage stable queues while the fading channel characteristics suggest the broadcast to operate at multi-rate to achieve higher throughput. In this research, we propose a joint scheduling and network coding (JSNC) strategy for such network coded broadcast system to achieve maximum throughput under queue stability constraint. In a single cell broadcast networks with exogenous arrivals of packets at the base station, we prove that JSNC can stabilize the system as long as the rate of the exogenous arrival flow is within the capacity region. Sufficient control parameters are provided in JSNC for trading off between sender's buffer and the receivers' buffers. JSNC can be viewed as a generalization of the classical backpressure scheduling rule to coded information flow. Alternatively, JSNC can also be viewed as an extension of network coding theory to queuing system. Hao Cui 0001, Chong Luo 0001, Feng Wu 0001, Chang Wen Chen |
INFOCOM | 3 |
| 2010 | Compressed sensing based video multicastabstractWe propose a new scheme for wireless video multicast based on compressed sensing. It has the property of graceful degradation and, unlike systems adhering to traditional separate coding, it does not suffer from a cliff effect. Compressed sensing is applied to generate measurements of equal importance from a video such that a receiver with a better channel will naturally have more information at hands to reconstruct the content without penalizing others. We experimentally compare different random matrices at the encoder side in terms of their performance for video transmission. We further investigate how properties of natural images can be exploited to improve the reconstruction performance by transmitting a small amount of side information. And we propose a way of exploiting inter-frame correlation by extending only the decoder. Finally we compare our results with a different scheme targeting the same problem with simulations and find competitive results for some channel configurations. Markus B. Schenkel, Chong Luo 0001, Pascal Frossard, Feng Wu 0001 |
VCIP | 2 |
| 2010 | Efficient Measurement Generation and Pervasive Sparsity for Compressive Data GatheringabstractWe proposed compressive data gathering (CDG) that leverages compressive sampling (CS) principle to efficiently reduce communication cost and prolong network lifetime for large scale monitoring sensor networks. The network capacity has been proven to increase proportionally to the sparsity of sensor readings. In this paper, we further address two key problems in the CDG framework. First, we investigate how to generate RIP (restricted isometry property) preserving measurements of sensor readings by taking multi-hop communication cost into account. Excitingly, we discover that a simple form of measurement matrix [I R] has good RIP, and the data gathering scheme that realizes this measurement matrix can further reduce the communication cost of CDG for both chain-type and tree-type topology. Second, although the sparsity of sensor readings is pervasive, it might be rather complicated to fully exploit it. Owing to the inherent flexibility of CS principle, the proposed CDG framework is able to utilize various sparsity patterns despite of a simple and unified data gathering process. In particular, we present approaches for adapting CS decoder to utilize cross-domain sparsity (e.g. temporal-frequency and spatial-frequency). We carry out simulation experiments over both synthesized and real sensor data. The results confirm that CDG can preserve sensor data fidelity at a reduced communication cost. Chong Luo 0001, Feng Wu 0001, Chang Wen Chen |
IEEE Trans. Wirel. Commun. | 1 |
| 2009 | Compressive data gathering for large-scale wireless sensor networksabstractThis paper presents the first complete design to apply compressive sampling theory to sensor data gathering for large-scale wireless sensor networks. The successful scheme developed in this research is expected to offer fresh frame of mind for research in both compressive sampling applications and large-scale wireless sensor networks. We consider the scenario in which a large number of sensor nodes are densely deployed and sensor readings are spatially correlated. The proposed compressive data gathering is able to reduce global scale communication cost without introducing intensive computation or complicated transmission control. The load balancing characteristic is capable of extending the lifetime of the entire sensor network as well as individual sensors. Furthermore, the proposed scheme can cope with abnormal sensor readings gracefully. We also carry out the analysis of the network capacity of the proposed compressive data gathering and validate the analysis through ns-2 simulations. More importantly, this novel compressive data gathering has been tested on real sensor data and the results show the efficiency and robustness of the proposed scheme. Chong Luo 0001, Feng Wu 0001, Chang Wen Chen |
MobiCom | 1 |
| 2009 | Forepressure Transmission Control for Wireless Video Sensor NetworksabstractMulti-source data transmission in wireless video sensor networks is a challenging problem because of the high bandwidth demand of video streams and the many-to-one traffic pattern. This paper proposes forepressure transmission control for efficient and fair video transmission in this scenario. Contrary to traditional backpressure transmission control where downstream node creates backpressure and causes upstream node to hold off when its buffer is full, forepressure transmission control creates forepressure and signals upstream node to transmit when the downstream node's buffer is not full. Forepressure transmission control proactively avoids congestion at little communication overhead. In addition, it is able to provide fairness among concurrent flows. We evaluate forepressure transmission control via NS-2 simulations, and compare it with both backpressure congestion control and end-to-end rate regulation. Results in three typical topologies show that forepressure transmission control achieves much lower loss ratio and ensures better fairness than the other two schemes. Chong Luo 0001, Chang Wen Chen, Feng Wu 0001 |
SECON | 1 |
| 2009 | QoS-driven network coded wireless multicastabstractEmerging wireless multicast applications simultaneously impose two requirements to the underlying communication networks: to provide sufficient bandwidth and to support a variety of quality-of-service (QoS) sensitivities. For the first requirement, network coding has been proposed recently as an effective way of improving bandwidth utilization. However, almost all previous works about network coding focus on the throughput gain without considering the QoS requirements. Optimal network code construction in wireless multicast under different QoS constraints remains as a significant challenge. In this research, we study the QoS-driven network coding problem. We use large deviation principle to establish the relationship among source rate, link condition, QoS requirement, and network code. Using this relationship, under given QoS requirements, we solve the optimal network code construction problem. The proposed network code supports maximal source rate without violating the QoS requirements. These results constitute the foundations for future designing and implementing network coding based wireless multicast protocols. Chong Luo 0001, Feng Wu 0001, Chang Wen Chen |
IEEE Trans. Wirel. Commun. | 2 |
| 2008 | Continuous Network Coding in Wireless Relay NetworksabstractNetwork coding has recently been applied to wireless networks and has achieved some initial success. Researches in wireless network coding have been mostly focusing on utilizing the broadcast nature of the wireless networks. In this paper, we propose a novel network coding framework for wireless relay networks that also takes into consideration the fading and error prone nature of the wireless networks. First, we extend the traditional network coding in lossless networks which operates on 0-1 bits, to a new framework which defines network coding on the posterior probability of each bit. This new framework allows an imperfect decode-recode process at a relay node and avoids possible error propagation when a hard decision is made at the relay node. It implicitly integrates decode-and-forward and estimate-and-forward strategies for wireless network coding to address the technical issues of channel fading and transmission errors. The proposed approach is validated through both theoretical analysis and extensive simulations. Both analysis and simulation confirm that this new framework is able to achieve significant gain over traditional network coding. This new framework also enables the introduction of adaptive scheme into network coding. We demonstrate a basic adaptation scheme and present some preliminary experimental results. The proposed adaptive scheme will lay down an essential foundation in this emerging field of wireless network coding in order to address issues related to link heterogeneity. Chong Luo 0001, Shipeng Li 0001, Chang Wen Chen |
INFOCOM | 2 |
| 2007 | A Multiparty Videoconferencing System Over an Application-Level Multicast ProtocolabstractIncreased speeds of PCs and networks have made media communications possible on the Internet. Today, the need for desktop videoconferencing is experiencing robust growth in both business and consumer markets. However, the synchronous delivery of high-volume media content is still a big challenge under a current heterogeneous Internet environment. In this paper, we present a multiparty videoconferencing system based on a peer-to-peer (P2P) solution. The contribution of our paper is twofold. On the one hand, we design an application-level multicast scheme which intends to tolerate the heterogeneity in videoconferencing applications. Design tradeoffs are analyzed and our decisions are made based on extensive experimentation. On the other, we design a five-layer architecture for implementing a multiparty videoconferencing system. This architecture makes a clear-cut distinction between different functional modules and therefore provides rich flexibility in feature adaptation. We believe that our work can be a helpful reference in other efforts on building desktop videoconferencing systems. Chong Luo 0001, Wei Wang 0335, Jiang Li 0008 |
IEEE Trans. Multim. | 1 |
| 2006 | Estimating Available Bandwidth Using Multiple Overloading StreamsabstractAvailable bandwidth measurement is essential to applications running on the best-effort Internet. In this paper, we propose an available bandwidth measurement technique named MoSeab. The idea behind MoSeab is a direct probing technique based on a statistical model for network data transmissions. Differing from the other direct probing techniques, MoSeab does not require any a priori knowledge of network path. It has proven valid even when there are multiple bottlenecks. Simulation and real Internet experiments demonstrate the accuracy and robustness of MoSeab, and show its advantages over Spruce, PathChirp, and IGI. Moreover, its probing traffic control mechanism makes MoSeab a non-intrusive solution suitable to be integrated into various network applications. Minjian Zhang 0004, Chong Luo 0001, Jiang Li 0008 |
ICC | 2 |
| 2004 | DigiMetro - an application-level multicast system for multi-party video conferencingabstractThe increasing demand for multi-party videoconferencing has aroused the research interest in the underlying multicast support. In this paper, we propose DigiMetro, an application-level multicast system tailored to small and impromptu videoconferencing. Breaking through the conventional wisdom to use shared overlay to handle multiple data sources, DigiMetro organizes the data delivery routes as source-specific trees, which are first constructed by a local greedy algorithm and then gradually improved by a global refinement procedure. Extensive simulation experiments demonstrate the efficiency of both algorithms. Moreover, DigiMetro is able to handle different video bit rates and provide different services over voice/video streams. Chong Luo 0001, Jiang Li 0008, Shipeng Li 0001 |
GLOBECOM | 1 |
| 2004 | DigiParty - a decentralized multi-party video conferencing systemabstractThe increased speeds of PCs and networks have made media communication possible on the Internet. However, nearly ten years after the first release of Microsoft NetMeeting, Internet video telephony is still limited to the point-to-point communication mode. Today, people have a need for an easy-to-use multi-party video conferencing tool that can connect families and friends around the world over the Internet. We present DigiParty, a fully distributed multi-party video conferencing system. DigiParty employs a full mesh conferencing architecture and adopts a loosely coupled conferencing mode. A novel conference control protocol is designed with the system. DigiParty can be integrated with any existing instant messaging services and is applicable to all types of Internet connections. Ling Chen 0001, Chong Luo 0001, Jiang Li 0008, Shipeng Li 0001 |
ICME | 2 |