VLDB 2026 Research / reviewers in the wild / expert
Hongyang Chao
dblp:58/4277 · also Hong-Yang Chao
· DBLP profile ↗
89ranked-venue papers
2as first author
19since 2021 · last 2025
0000-0002-6104-2322ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 66 · 2 first-author · 12 since 2021Artificial intelligence and machine learning · 37 · 9 since 2021Databases, data management, data science and information retrieval · 13 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Computer networks · 3 · 2 since 2021Software engineering, systems software and programming languages · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DreamStory: Open-Domain Story Visualization by LLM-Guided Multi-Subject Consistent DiffusionabstractStory visualization aims to create visually compelling images or videos corresponding to textual narratives. Despite recent advances in diffusion models yielding promising results, existing methods still struggle to create a coherent sequence of subject-consistent frames based solely on a story. To this end, we propose DreamStory, an automatic open-domain story visualization framework by leveraging the LLMs and a novel multi-subject consistent diffusion model. DreamStory consists of (1) an LLM acting as a story director and (2) an innovative Multi-Subject consistent Diffusion model (MSD) for generating consistent multi-subject across the images. First, DreamStory employs the LLM to generate descriptive prompts for subjects and scenes aligned with the story, annotating each scene's subjects for subsequent subject-consistent generation. Second, DreamStory utilizes these detailed subject descriptions to create portraits of the subjects, with these portraits and their corresponding textual information serving as multimodal anchors (guidance). Finally, the MSD uses these multimodal anchors to generate story scenes with consistent multi-subject. Specifically, the MSD includes Masked Mutual Self-Attention (MMSA) and Masked Mutual Cross-Attention (MMCA) modules. MMSA module ensures detailed appearance consistency with reference images, while MMCA captures key attributes of subjects from their reference text to ensure semantic consistency. Both modules employ masking mechanisms to restrict each scene's subjects to referencing the multimodal information of the corresponding subject, effectively preventing blending between multiple subjects. To validate our approach and promote progress in story visualization, we established a benchmark, DS-500, which can assess the overall performance of the story visualization framework, subject-identification accuracy, and the consistency of the generation model. Extensive experiments validate the effectiveness of DreamStory in both subjective and objective evaluations. Huiguo He, Huan Yang 0005, Zixi Tuo, Qiuyue Wang, Wenhao Huang 0001, Hongyang Chao, Jian Yin 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 9 |
| 2025 | Exploring Vision-Language Foundation Model for Novel Object CaptioningabstractIt is always well believed that pre-trained vision-language foundation models (e.g., CLIP) would substantially facilitate vision-language tasks. Nevertheless, there has been less evidence in support of the idea on describing novel objects in images. In this paper, we propose the Novel Object Transformer with CLIP (NOTC), a Transformer-based model that innovatively exploits the powerful vision-language representation ability of CLIP to enhance novel object captioning model’s training and sentence decoding processes. Technically, given the primary bag-of-objects extracted by Faster R-CNN, NOTC first capitalize on an object distiller module to emphasize the most salient objects and infer the missing novel ones. The refined object words are additionally fed into the object-centric word predictor to generate sentence word-by-word. During training, we design a CLIP-based self-critical sequence training paradigm to select visually-grounded sampled sentence with higher CLIP score reward, which enables a joint training process of captioning model over out-domain training images with novel objects. Moreover, at inference, a new CLIP beam search algorithm is devised to enforce the existence of novel objects and encourage the partial word sequences with higher CLIP scores, thereby decoding both visually-grounded and comprehensive sentences. Extensive experiments are conducted on held-out COCO and nocaps datasets, and competitive performances are reported when compared to state-of-the-art approaches. Jianjie Luo, Yehao Li, Yingwei Pan, Ting Yao 0003, Jianlin Feng, Hongyang Chao, Tao Mei 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Unleashing Text-to-Image Diffusion Prior for Zero-Shot Image Captioning
Jianjie Luo, Jingwen Chen 0001, Yehao Li, Yingwei Pan, Jianlin Feng, Hongyang Chao, Ting Yao 0003 |
ECCV (57) | 6 |
| 2023 | Semantic-Conditional Diffusion Networks for Image CaptioningabstractRecent advances on text-to-image generation have witnessed the rise of diffusion models which act as powerful generative models. Nevertheless, it is not trivial to exploit such latent variable models to capture the dependency among discrete words and meanwhile pursue complex visual-language alignment in image captioning. In this paper, we break the deeply rooted conventions in learning Transformer-based encoder-decoder, and propose a new diffusion model based paradigm tailored for image captioning, namely Semantic-Conditional Diffusion Networks (SCD-Net). Technically, for each input image, we first search the semantically relevant sentences via cross-modal retrieval model to convey the comprehensive semantic information. The rich semantics are further regarded as semantic prior to trigger the learning of Diffusion Transformer, which produces the output sentence in a diffusion process. In SCD-Net, multiple Diffusion Transformer structures are stacked to progressively strengthen the output sentence with better visional-language alignment and linguistical coherence in a cascaded manner. Furthermore, to stabilize the diffusion process, a new self-critical sequence training strategy is designed to guide the learning of SCD-Net with the knowledge of a standard autoregressive Transformer model. Extensive experiments on COCO dataset demonstrate the promising potential of using diffusion models in the challenging image captioning task. Source code is available at Jianjie Luo, Yehao Li, Yingwei Pan, Ting Yao 0003, Jianlin Feng, Hongyang Chao, Tao Mei 0001 |
CVPR | 6 |
| 2023 | TinyCLIP: CLIP Distillation via Affinity Mimicking and Weight InheritanceabstractIn this paper, we propose a novel cross-modal distillation method, called TinyCLIP, for large-scale language-image pre-trained models. The method introduces two core techniques: affinity mimicking and weight inheritance. Affinity mimicking explores the interaction between modalities during distillation, enabling student models to mimic teachers’ behavior of learning cross-modal feature alignment in a visual-linguistic affinity space. Weight inheritance transmits the pre-trained weights from the teacher models to their student counterparts to improve distillation efficiency. Moreover, we extend the method into a multi-stage progressive distillation to mitigate the loss of informative weights during extreme compression. Comprehensive experiments demonstrate the efficacy of TinyCLIP, showing that it can reduce the size of the pre-trained CLIP ViT-B/32 by 50%, while maintaining comparable zero-shot performance. While aiming for comparable performance, distillation with weight inheritance can speed up the training by 1.4 - 7.8× compared to training from scratch. Moreover, our TinyCLIP ViT-8M/16, trained on YFCC-15M, achieves an impressive zero-shot top-1 accuracy of 41.1% on ImageNet, surpassing the original CLIP ViT-B/16 by 3.5% while utilizing only 8.9% parameters. Finally, we demonstrate the good transferability of TinyCLIP in various downstream tasks. Code and models will be open-sourced at aka.ms/tinyclip. Houwen Peng, Zhenghong Zhou, Bin Xiao 0004, Mengchen Liu, Lu Yuan 0001, Hong Xuan, Michael Valenzuela, Xi Stephen Chen, Xinggang Wang, Hongyang Chao, Han Hu 0001 |
ICCV | 11 |
| 2023 | Learning Profitable NFT Image Diffusions via Multiple Visual-Policy Guided Reinforcement LearningabstractWe study the task of generating profitable Non-Fungible Token (NFT) images from user-input texts. Recent advances in diffusion models have shown great potential for image generation. However, existing works can fall short in generating visually-pleasing and highly-profitable NFT images, mainly due to the lack of 1) plentiful and fine-grained visual attribute prompts for an NFT image, and 2) effective optimization metrics for generating high-quality NFT images. To solve these challenges, we propose a Diffusion based generation framework with Multiple Visual-Policies as rewards (i.e., Diffusion-MVP) for NFT images. The proposed framework consists of a large language model (LLM), a diffusion-based image generator, and a series of visual rewards by design. First, the LLM enhances a basic human input (such as "panda") by generating more comprehensive NFT-style prompts that include specific visual attributes, such as "panda with Ninja style and green background." Second, the diffusion-based image generator is fine-tuned using a large-scale NFT dataset to capture fine-grained image styles and accessory compositions of popular NFT elements. Third, we further propose to utilize multiple visual-policies as optimization goals, including visual rarity levels, visual aesthetic scores, and CLIP-based text-image relevances. This design ensures that our proposed Diffusion-MVP is capable of minting NFT images with high visual quality and market value. To facilitate this research, we have collected the largest publicly available NFT image dataset to date, consisting of 1.5 million high-quality images with corresponding texts and market values. Extensive experiments including objective evaluations and user studies demonstrate that our framework can generate NFT images showing more visually engaging elements and higher market value, compared with state-of-the-art approaches. Huiguo He, Tianfu Wang 0002, Huan Yang 0005, Jianlong Fu, Nicholas Jing Yuan, Jian Yin 0001, Hongyang Chao, Qi Zhang 0066 |
ACM Multimedia | 7 |
| 2023 | A Low Rank Promoting Prior for Unsupervised Contrastive LearningabstractUnsupervised learning is just at a tipping point where it could really take off. Among these approaches, contrastive learning has led to state-of-the-art performance. In this paper, we construct a novel probabilistic graphical model that effectively incorporates the low rank promoting prior into the framework of contrastive learning, referred to as LORAC. In contrast to the existing conventional self-supervised approaches that only considers independent learning, our hypothesis explicitly requires that all the samples belonging to the same instance class lie on the same subspace with small dimension. This heuristic poses particular joint learning constraints to reduce the degree of freedom of the problem during the search of the optimal network parameterization. Most importantly, we argue that the low rank prior employed here is not unique, and many different priors can be invoked in a similar probabilistic way, corresponding to different hypotheses about underlying truth behind the contrastive features. Empirical evidences show that the proposed algorithm clearly surpasses the state-of-the-art approaches on multiple benchmarks, including image classification, object detection, instance segmentation and keypoint detection. Code is available: https://github.com/ssl-codelab/lorac. Yu Wang 0060, Yingwei Pan, Ting Yao 0003, Hongyang Chao, Tao Mei 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | Boosting Vision-and-Language Navigation with Direction Guiding and BacktracingabstractVision-and-Language Navigation (VLN) has been an emerging and fast-developing research topic, where an embodied agent is required to navigate in a real-world environment based on natural language instructions. In this article, we present a Direction-guided Navigator Agent (DNA) that novelly integrates direction clues derived from instructions into the essential encoder-decoder navigation framework. Particularly, DNA couples the standard instruction encoder with an additional direction branch which sequentially encodes the direction clues in the instructions to boost navigation. Furthermore, an Instruction Flipping mechanism is uniquely devised to enable fast data augmentation as well as a follow-up backtracing for navigating the agent in a backward direction. Such a way naturally amplifies the grounding of instruction in the local visual scenes along both forward and backward directions, and thus strengthens the alignment between instruction and action sequence. Extensive experiments conducted on Room to Room (R2R) dataset validate our proposal and demonstrate quantitatively compelling results. Jingwen Chen 0001, Jianjie Luo, Yingwei Pan, Yehao Li, Ting Yao 0003, Hongyang Chao, Tao Mei 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2023 | Retrieval Augmented Convolutional Encoder-decoder Networks for Video CaptioningabstractVideo captioning has been an emerging research topic in computer vision, which aims to generate a natural sentence to correctly reflect the visual content of a video. The well-established way of doing so is to rely on encoder-decoder paradigm by learning to encode the input video and decode the variable-length output sentence in a sequence-to-sequence manner. Nevertheless, these approaches often fail to produce complex and descriptive sentences as natural as those from human being, since the models are incapable of memorizing all visual contents and syntactic structures in the human-annotated video-sentence pairs. In this article, we uniquely introduce a Retrieval Augmentation Mechanism (RAM) that enables the explicit reference to existing video-sentence pairs within any encoder-decoder captioning model. Specifically, for each query video, a video-sentence retrieval model is first utilized to fetch semantically relevant sentences from the training sentence pool, coupled with the corresponding training videos. RAM then writes the relevant video-sentence pairs into memory and reads the memorized visual contents/syntactic structures in video-sentence pairs from memory to facilitate the word prediction at each timestep. Furthermore, we present Retrieval Augmented Convolutional Encoder-Decoder Network (R-ConvED), which novelly integrates RAM into convolutional encoder-decoder structure to boost video captioning. Extensive experiments on MSVD, MSR-VTT, Activity Net Captions, and VATEX datasets validate the superiority of our proposals and demonstrate quantitatively compelling results. Jingwen Chen 0001, Yingwei Pan, Yehao Li, Ting Yao 0003, Hongyang Chao, Tao Mei 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2023 | Aggregated Contextual Transformations for High-Resolution Image InpaintingabstractImage inpainting that completes large free-form missing regions in images is a promising yet challenging task. State-of-the-art approaches have achieved significant progress by taking advantage of generative adversarial networks (GAN). However, these approaches can suffer from generating distorted structures and blurry textures in high-resolution images (e.g., 512×512). The challenges mainly drive from (1) image content reasoning from distant contexts, and (2) fine-grained texture synthesis for a large missing region. To overcome these two challenges, we propose an enhanced GAN-based model, named Aggregated COntextual-Transformation GAN (AOT-GAN), for high-resolution image inpainting. Specifically, to enhance context reasoning, we construct the generator of AOT-GAN by stacking multiple layers of a proposed AOT block. The AOT blocks aggregate contextual transformations from various receptive fields, allowing to capture both informative distant image contexts and rich patterns of interest for context reasoning. For improving texture synthesis, we enhance the discriminator of AOT-GAN by training it with a tailored mask-prediction task. Such a training objective forces the discriminator to distinguish the detailed appearances of real and synthesized patches, and in turn facilitates the generator to synthesize clear textures. Extensive comparisons on Places2, the most challenging benchmark with 1.8 million high-resolution images of 365 complex scenes, show that our model outperforms the state-of-the-art. A user study including more than 30 subjects further validates the superiority of AOT-GAN. We further evaluate the proposed AOT-GAN in practical applications, e.g., logo removal, face editing, and object removal. Results show that our model achieves promising completions in the real world. We release codes and models in https://github.com/researchmm/AOT-GAN-for-Inpainting. Yanhong Zeng, Jianlong Fu, Hongyang Chao, Baining Guo |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2022 | Compression loss-based spatial-temporal attention module for compressed video quality enhancement
Huiguo He, Hongyang Chao, Jian Yin 0001 |
Neurocomputing | 2 |
| 2021 | Rethinking and Improving Relative Position Encoding for Vision TransformerabstractRelative position encoding (RPE) is important for transformer to capture sequence ordering of input tokens. General efficacy has been proven in natural language processing. However, in computer vision, its efficacy is not well studied and even remains controversial, e.g., whether relative position encoding can work equally well as absolute position? In order to clarify this, we first review existing relative position encoding methods and analyze their pros and cons when applied in vision transformers. We then propose new relative position encoding methods dedicated to 2D images, called image RPE (iRPE). Our methods consider directional relative distance modeling as well as the interactions between queries and relative position embeddings in self-attention mechanism. The proposed iRPE methods are simple and lightweight. They can be easily plugged into transformer blocks. Experiments demonstrate that solely due to the proposed encoding methods, DeiT [21] and DETR [1] obtain up to 1.5% (top-1 Acc) and 1.3% (mAP) stable improvements over their original versions on ImageNet and COCO respectively, without tuning any extra hyperparameters such as learning rate and weight decay. Our ablation and analysis also yield interesting findings, some of which run counter to previous understanding. Code and models are open-sourced at https://github.com/microsoft/Cream/tree/main/iRPE. Houwen Peng, Jianlong Fu, Hongyang Chao |
ICCV | 5 |
| 2021 | Core-Text: Improving Scene Text Detection with Contrastive Relational ReasoningabstractLocalizing text instances in natural scenes is regarded as a fundamental challenge in computer vision. Nevertheless, owing to the extremely varied aspect ratios and scales of text instances in real scenes, most conventional text detectors suffer from the sub-text problem that only localizes the fragments of text instance (i.e., sub-texts). In this work, we quantitatively analyze the sub-text problem and present a simple yet effective design, COntrastive RElation (CORE) module, to mitigate that issue. CORE first leverages a vanilla relation block to model the relations among all text proposals (sub-texts of multiple text instances) and further enhances relational reasoning via instance-level sub-text discrimination in a contrastive manner. Such way naturally learns instance-aware representations of text proposals and thus facilitates scene text detection. We integrate the CORE module into a two-stage text detector of Mask R-CNN and devise our text detector CORE-Text. Extensive experiments on four benchmarks demonstrate the superiority of CORE-Text. Yingwei Pan, Rongfeng Lai, Xuehang Yang, Hongyang Chao, Ting Yao 0003 |
ICME | 5 |
| 2021 | CoCo-BERT: Improving Video-Language Pre-training with Contrastive Cross-modal Matching and DenoisingabstractBERT-type structure has led to the revolution of vision-language pre-training and the achievement of state-of-the-art results on numerous vision-language downstream tasks. Existing solutions dominantly capitalize on the multi-modal inputs with mask tokens to trigger mask-based proxy pre-training tasks (e.g., masked language modeling and masked object/frame prediction). In this work, we argue that such masked inputs would inevitably introduce noise for cross-modal matching proxy task, and thus leave the inherent vision-language association under-explored. As an alternative, we derive a particular form of cross-modal proxy objective for video-language pre-training, i.e., Contrastive Cross-modal matching and denoising (CoCo). By viewing the masked frame/word sequences as the noisy augmentation of primary unmasked ones, CoCo strengthens video-language association by simultaneously pursuing inter-modal matching and intra-modal denoising between masked and unmasked inputs in a contrastive manner. Our CoCo proxy objective can be further integrated into any BERT-type encoder-decoder structure for video-language pre-training, named as Contrastive Cross-modal BERT (CoCo-BERT). We pre-train CoCo-BERT on TV dataset and a newly collected large-scale GIF video dataset (ACTION). Through extensive experiments over a wide range of downstream tasks (e.g., cross-modal retrieval, video question answering, and video captioning), we demonstrate the superiority of CoCo-BERT as a pre-trained structure. Jianjie Luo, Yehao Li, Yingwei Pan, Ting Yao 0003, Hongyang Chao, Tao Mei 0001 |
ACM Multimedia | 5 |
| 2021 | Searching the Search Space of Vision TransformerabstractVision Transformer has shown great visual representation power in substantial vision tasks such as recognition and detection, and thus been attracting fast-growing efforts on manually designing more effective architectures. In this paper, we propose to use neural architecture search to automate this process, by searching not only the architecture but also the search space. The central idea is to gradually evolve different search dimensions guided by their E-T Error computed using a weight-sharing supernet. Moreover, we provide design guidelines of general vision transformers with extensive analysis according to the space searching process, which could promote the understanding of vision transformer. Remarkably, the searched models, named S3 (short for Searching the Search Space), from the searched space achieve superior performance to recently proposed models, such as Swin, DeiT and ViT, when evaluated on ImageNet. The effectiveness of S3 is also illustrated on object detection, semantic segmentation and visual question answering, demonstrating its generality to downstream vision and vision-language tasks. Code and models will be available at https://github.com/microsoft/Cream. Bolin Ni, Houwen Peng, Bei Liu 0001, Jianlong Fu, Hongyang Chao, Haibin Ling |
NeurIPS | 7 |
| 2021 | Improving Visual Quality of Image Synthesis by A Token-based Generator with TransformersabstractWe present a new perspective of achieving image synthesis by viewing this task as a visual token generation problem. Different from existing paradigms that directly synthesize a full image from a single input (e.g., a latent code), the new formulation enables a flexible local manipulation for different image regions, which makes it possible to learn content-aware and fine-grained style control for image synthesis. Specifically, it takes as input a sequence of latent tokens to predict the visual tokens for synthesizing an image. Under this perspective, we propose a token-based generator (i.e., TokenGAN). Particularly, the TokenGAN inputs two semantically different visual tokens, i.e., the learned constant content tokens and the style tokens from the latent space. Given a sequence of style tokens, the TokenGAN is able to control the image synthesis by assigning the styles to the content tokens by attention mechanism with a Transformer. We conduct extensive experiments and show that the proposed TokenGAN has achieved state-of-the-art results on several widely-used image synthesis benchmarks, including FFHQ and LSUN CHURCH with different resolutions. In particular, the generator is able to synthesize high-fidelity images with (1024x1024) size, dispensing with convolutions entirely. Yanhong Zeng, Huan Yang 0005, Hongyang Chao, Jianlong Fu |
NeurIPS | 3 |
| 2021 | oComm: Overlapping Community Detection in Multi-View Brain NetworkabstractMany efforts have been made on developing multi-view network community detection approaches. However, most of them can only reveal non-overlapping community structure. In this paper, we propose a novel approach for Overlapping Community Detection in Multi-view Brain Network (oComm). For modeling the overlapping community structure, a community membership strength vector is introduced for each node in each view, based on which a network generative model is designed to measure the within-view community quality. For measuring the consistency of overlapping community structures across different views, the Jaccard similarity is adopted to measure the first-order structural consistency of one node across different views, based on which a cross-view community consistency model is established. One objective function is defined by integrating the above two components. By solving the objective function via the alternative coordinate gradient ascent method, the optimal community membership strength vectors are generated, from which the multi-view overlapping community structure is obtained. Additionally, this study collects a set of EEG data of 147 subjects from Department of Otolaryngology of Sun Yat-sen Memorial Hospital, Sun Yat-sen University, based on which three multi-view brain networks are constructed. Comparison results with several existing approaches have confirmed the effectiveness of the proposed method. Ling Huang 0002, Chang-Dong Wang 0001, Hongyang Chao |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2021 | Reference-Based Defect Detection NetworkabstractThe defect detection task can be regarded as a realistic scenario of object detection in the computer vision field and it is widely used in the industrial field. Directly applying vanilla object detector to defect detection task can achieve promising results, while there still exists challenging issues that have not been solved. The first issue is the texture shift which means a trained defect detector model will be easily affected by unseen texture, and the second issue is partial visual confusion which indicates that a partial defect box is visually similar with a complete box. To tackle these two problems, we propose a Reference-based Defect Detection Network (RDDN). Specifically, we introduce template reference and context reference to against those two problems, respectively. Template reference can reduce the texture shift from image, feature or region levels, and encourage the detectors to focus more on the defective area as a result. We can use either well-aligned template images or the outputs of a pseudo template generator as template references in this work, and they are jointly trained with detectors by the supervision of normal samples. To solve the partial visual confusion issue, we propose to leverage the carried context information of context reference, which is the concentric bigger box of each region proposal, to perform more accurate region classification and regression. Experiments on two defect detection datasets demonstrate the effectiveness of our proposed approach. Zhaoyang Zeng, Bei Liu 0001, Jianlong Fu, Hongyang Chao |
IEEE Trans. Image Process. | 4 |
| 2021 | HM-Modularity: A Harmonic Motif Modularity Approach for Multi-Layer Network Community DetectionabstractMulti-layer network community detection has drawn an increasing amount of attention recently. Despite success, the existing methods mainly focus on the lower-order connectivity structure at the level of individual nodes and edges. And the higher-order connectivity structure has been largely ignored, which contains better signature of community compared with edges. The main challenges in utilizing higher-order structure for multi-layer network community detection are that the most representative higher-order structure may vary from one layer to another and the connectivity structure formed by the same node subset may exhibit different higher-order connectivity patterns in different layers. To this end, this paper proposes a novel higher-order structure, termed harmonic motif, which is a dense subgraph having on average the largest statistical significance in each layer. Based on the harmonic motif, a primary layer is constructed by integrating higher-order structural information from all layers. Additionally, the higher-order structural information of each individual layer is taken as the auxiliary information. A coupling is established between the primary layer and each auxiliary layer. Accordingly, a harmonic motif modularity is designed to generate the community structure. Extensive experiments on eleven real-world multi-layer network datasets have been conducted to confirm the effectiveness of the proposed method. Ling Huang 0002, Chang-Dong Wang 0001, Hongyang Chao |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2020 | MuMod: A Micro-Unit Connection Approach for Hybrid-Order Community DetectionabstractIn the past few years, higher-order community detection has drawn an increasing amount of attention. Compared with the lower-order approaches that rely on the connectivity pattern of individual nodes and edges, the higher-order approaches discover communities by leveraging the higher-order connectivity pattern via constructing a motif-based hypergraph. Despite success in capturing the building blocks of complex networks, recent study has shown that the higher-order approaches unavoidably suffer from the hypergraph fragmentation issue. Although an edge enhancement strategy has been designed previously to address this issue, adding additional edges may corrupt the original lower-order connectivity pattern. To this end, this paper defines a new problem of community detection, namely hybrid-order community detection, which aims to discover communities by simultaneously leveraging the lower-order connectivity pattern and the higherorder connectivity pattern. For addressing this new problem, a new Micro-unit Modularity (MuMod) approach is designed. The basic idea lies in constructing a micro-unit connection network, where both of the lower-order connectivity pattern and the higher-order connectivity pattern are utilized. And then a new micro-unit modularity model is proposed for generating the micro-unit groups, from which the overlapping community structure of the original network can be derived. Extensive experiments are conducted on five real-world networks. Comparison results with twelve existing approaches confirm the effectiveness of the proposed method. Ling Huang 0002, Hongyang Chao, Guangqiang Xie |
AAAI | 2 |
| 2020 | Learning Joint Spatial-Temporal Transformations for Video Inpainting
Yanhong Zeng, Jianlong Fu, Hongyang Chao |
ECCV (16) | 3 |
| 2020 | Endowing Deep 3d Models With Rotation Invariance Based On Principal Component AnalysisabstractIn this paper, we propose to endow deep 3D models with rotation invariance by expressing the coordinates in an intrinsic frame determined by the object shape itself. Key to our approach is to find such an intrinsic frame which should be unique to the identical object shape and consistent across different instances of the same category. Interestingly, the principal component analysis exactly provides an effective way to define such a frame, i.e. setting the principal components as the frame axes. As the principal components have direction ambiguity, there exist several intrinsic frames for each object. To achieve absolute rotation invariance for a deep model, we adopt the coordinates expressed in all intrinsic frames as inputs to obtain multiple output features, which will be aggregated as a final feature via a self-attention module. Comprehensive experiments demonstrate that our approach can achieve state-of-the-art performance on rotated 3D object classification and retrieval tasks. Zelin Xiao, Hongxin Lin, Lishuai Geng, Hongyang Chao, Shengyong Ding |
ICME | 5 |
| 2020 | VideoTRM: Pre-training for Video Captioning Challenge 2020abstractThe Pre-training for Video Captioning Challenge 2020 mainly focuses on developing video captioning systems by pre-training on the newly released large-scale Auto-captions on GIF dataset and further transferring the pre-trained model to MSR-VTT benchmark. As a part of the submission to this challenge, we propose a Transformer based framework named VideoTRM, which consists of four modules: a textual encoder for encoding the linguistic relationship among words in the input sentence, a visual encoder for capturing the temporal dynamics in the input video, a cross-modal encoder for modeling the interactions between the two modalities (i.e., textual and visual) and a decoder for sentence generation conditioned on the input video and words generated previously. Additionally, we extend the decoder in our VideoTRM with mesh-like connections and gate fusion mechanism in multi-head attention during fine-tuning to take advantage of multi-level visual features and bypass less informative attention results, respectively. In the evaluation on test server, our VideoTRM achieves superior performances and ranks the second place on the leadboard finally. Jingwen Chen 0001, Hongyang Chao |
ACM Multimedia | 2 |
| 2020 | Deep Metric Learning With Density AdaptivityabstractThe problem of distance metric learning is mostly considered from the perspective of learning an embedding space, where the distances between pairs of examples are in correspondence with a similarity metric. With the rise and success of Convolutional Neural Networks (CNN), deep metric learning (DML) involves training a network to learn a nonlinear transformation to the embedding space. Existing DML approaches often express the supervision through maximizing inter-class distance and minimizing intra-class variation. However, the results can suffer from overfitting problem, especially when the training examples of each class are embedded together tightly and the density of each class is very high. In this paper, we integrate density, i.e., the measure of data concentration in the representation, into the optimization of DML frameworks to adaptively balance inter-class similarity and intra-class variation by training the architecture in an end-to-end manner. Technically, the knowledge of density is employed as a regularizer, which is pluggable to any DML architecture with different objective functions such as contrastive loss, N-pair loss and triplet loss. Extensive experiments on three public datasets consistently demonstrate clear improvements by amending three types of embedding with the density adaptivity. More remarkably, our proposal increases Recall@1 from 67.95% to 77.62%, from 52.01% to 55.64% and from 68.20% to 70.56% on Cars196, CUB-200-2011 and Stanford Online Products dataset, respectively. Yehao Li, Ting Yao 0003, Yingwei Pan, Hongyang Chao, Tao Mei 0001 |
IEEE Trans. Multim. | 4 |
| 2020 | The Interpretable Fast Multi-Scale Deep Decoder for the Standard HEVC BitstreamsabstractIt is a research hotspot to restore decoded videos with existing bitstreams by applying deep neural network to improve compression efficiency at decoder-end. Existing research has verified that the utilization of redundancy at decoder-end, which is underused by the encoder, can bring an increase of compression efficiency. However, most existing research neglects the abundant multi-scale information among video frames as a typical type of such redundancy. It remains an interesting yet challenging topic how to build an effective, interpretable and fast deep neural network for the purpose of using the multi-scale similarity at decoder-end and further enhancing compression efficiency. To this end, this paper considers the use of underused inter multi-scale information and proposes the Fast Multi-Scale Deep Decoder (Fast MSDD) for the state-of-the-art video coding standard HEVC. The advantages of Fast MSDD are three-fold. First, it achieves a higher coding efficiency without modifying any encoding algorithm. Second, Fast MSDD is interpretable based on the framework of using the underused redundancy. Third, it guarantees the model's inference speed while fully using the multi-scale similarity among video frames. Extensive experimental results verify Fast MSDD's effectiveness, interpretability, and computational efficiency. Fast MSDD obtains averagely 14.3%, 10.8%, 8.5% and 7.6% BD gains for AI, LP, LB and RA respectively. Compared with our previous work MSDD, Fast MSDD achieves increases of 59.3%, 49.1%, 61.0% and 29.3%. Meanwhile, 16.9%, 11.2%, 9.2% and 8.3% BD gains are observed on videos with scale changes, which validate the interpretability of the proposed method. Furthermore, Fast MSDD can save at most 56.3% time compared to MSDD. Wenhui Xiao, Huiguo He, Tingting Wang 0004, Hongyang Chao |
IEEE Trans. Multim. | 4 |
| 2020 | MVStream: Multiview Data Stream ClusteringabstractThis article studies a new problem of data stream clustering, namely, multiview data stream (MVStream) clustering. Although many data stream clustering algorithms have been developed, they are restricted to the single-view streaming data, and clustering MVStreams still remains largely unsolved. In addition to the many issues encountered by the conventional single-view data stream clustering, such as capturing cluster evolution and discovering clusters of arbitrary shapes under the limited computational resources, the main challenge of MVStream clustering lies in integrating information from multiple views in a streaming manner and abstracting summary statistics from the integrated features simultaneously. In this article, we propose a novel MVStream clustering algorithm for the first time. The main idea is to design a multiview support vector domain description (MVSVDD) model, by which the information from multiple insufficient views can be integrated, and the outputting support vectors (SVs) are utilized to abstract the summary statistics of the historical multiview data objects. Based on the MVSVDD model, a new multiview cluster labeling method is designed, whereby clusters of arbitrary shapes can be discovered for each view. By tracking the cluster labels of SVs in each view, the cluster evolution associated with concept drift can be captured. Since the SVs occupy only a small portion of data objects, the proposed MVStream algorithm is quite efficient with the limited computational resources. Extensive experiments are conducted to demonstrate the effectiveness and efficiency of the proposed method. Ling Huang 0002, Chang-Dong Wang 0001, Hongyang Chao, Philip S. Yu |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2019 | Temporal Deformable Convolutional Encoder-Decoder Networks for Video CaptioningabstractIt is well believed that video captioning is a fundamental but challenging task in both computer vision and artificial intelligence fields. The prevalent approach is to map an input video to a variable-length output sentence in a sequence to sequence manner via Recurrent Neural Network (RNN). Nevertheless, the training of RNN still suffers to some degree from vanishing/exploding gradient problem, making the optimization difficult. Moreover, the inherently recurrent dependency in RNN prevents parallelization within a sequence during training and therefore limits the computations. In this paper, we present a novel design — Temporal Deformable Convolutional Encoder-Decoder Networks (dubbed as TDConvED) that fully employ convolutions in both encoder and decoder networks for video captioning. Technically, we exploit convolutional block structures that compute intermediate states of a fixed number of inputs and stack several blocks to capture long-term relationships. The structure in encoder is further equipped with temporal deformable convolution to enable free-form deformation of temporal sampling. Our model also capitalizes on temporal attention mechanism for sentence generation. Extensive experiments are conducted on both MSVD and MSR-VTT video captioning datasets, and superior results are reported when comparing to conventional RNN-based encoder-decoder techniques. More remarkably, TDConvED increases CIDEr-D performance from 58.8% to 67.2% on MSVD. Jingwen Chen 0001, Yingwei Pan, Yehao Li, Ting Yao 0003, Hongyang Chao, Tao Mei 0001 |
AAAI | 5 |
| 2019 | Higher-Order Multi-Layer Community DetectionabstractIn this paper, we define a new problem of multi-layer network community detection, namely higher-order multi-layer community detection. A multi-layer motif (M-Motif) approach is proposed, which discovers communities with good intralayer higher-order community quality while preserving interlayer higher-order community consistency. Experimental results have confirmed the superiority of the proposed method. Ling Huang 0002, Chang-Dong Wang 0001, Hongyang Chao |
AAAI | 3 |
| 2019 | Pointing Novel Objects in Image CaptioningabstractImage captioning has received significant attention with remarkable improvements in recent advances. Nevertheless, images in the wild encapsulate rich knowledge and cannot be sufficiently described with models built on image-caption pairs containing only in-domain objects. In this paper, we propose to address the problem by augmenting standard deep captioning architectures with object learners. Specifically, we present Long Short-Term Memory with Pointing (LSTM-P) --- a new architecture that facilitates vocabulary expansion and produces novel objects via pointing mechanism. Technically, object learners are initially pre-trained on available object recognition data. Pointing in LSTM-P then balances the probability between generating a word through LSTM and copying a word from the recognized objects at each time step in decoder stage. Furthermore, our captioning encourages global coverage of objects in the sentence. Extensive experiments are conducted on both held-out COCO image captioning and ImageNet datasets for describing novel objects, and superior results are reported when comparing to state-of-the-art approaches. More remarkably, we obtain an average of 60.9% in F1 score on held-out COCO dataset. Yehao Li, Ting Yao 0003, Yingwei Pan, Hongyang Chao, Tao Mei 0001 |
CVPR | 4 |
| 2019 | Learning Pyramid-Context Encoder Network for High-Quality Image InpaintingabstractHigh-quality image inpainting requires filling missing regions in a damaged image with plausible content. Existing works either fill the regions by copying high-resolution patches or generating semantically-coherent patches from region context, while neglecting the fact that both visual and semantic plausibility are highly-demanded. In this paper, we propose a Pyramid-context Encoder Network (denoted as PEN-Net) for image inpainting by deep generative models. The proposed PEN-Net is built upon a U-Net structure with three tailored components, ie., a pyramid-context encoder, a multi-scale decoder, and an adversarial training loss. First, we adopt a U-Net as backbone which can encode the context of an image from high-resolution pixels into high-level semantic features, and decode the features reversely. Second, we propose a pyramid-context encoder, which progressively learns region affinity by attention from a high-level semantic feature map, and transfers the learned attention to its adjacent high-resolution feature map. As the missing content can be filled by attention transfer from deep to shallow in a pyramid fashion, both visual and semantic coherence for image inpainting can be ensured. Third, we further propose a multi-scale decoder with deeply-supervised pyramid losses and an adversarial loss. Such a design not only results in fast convergence in training, but more realistic results in testing. Extensive experiments on a broad range of datasets shows the superior performance of the proposed network. Yanhong Zeng, Jianlong Fu, Hongyang Chao, Baining Guo |
CVPR | 3 |
| 2019 | TCN: Transferable Coupled Network for Cross-Resolution Face Recognition*abstractCross-resolution face recognition (CRFR) aims to learn the matching of a low-resolution (LR) probe image with a database of high-resolution (HR) gallery images. Existing methods including super resolution and projection-based algorithms are not recognition-oriented and computationally expensive, or ignore the inter-class associations across resolutions. To address the issues, we propose a novel end-to-end Transferable Coupled Network (TCN) for CRFR. Specifically, the TCN consists of two networks for the HR and LR domains, respectively. To reduce the resolution mismatch, a transferrable triple loss (TTL) is introduced to pull together cross-resolution positive pairs (intra-class) and also enforce margins towards negative ones (inter-class) from both domains. Besides, to keep stability and faster convergence, a novel online triplet selection method is proposed. Empirically, the proposed TCN model consistently outperforms the state-of-the-art methods among various low resolutions and architectures on public LFW and SCFace benchmarks. Juan Zha, Hongyang Chao |
ICASSP | 2 |
| 2019 | WSOD2: Learning Bottom-Up and Top-Down Objectness Distillation for Weakly-Supervised Object DetectionabstractWe study on weakly-supervised object detection (WSOD) which plays a vital role in relieving human involvement from object-level annotations. Predominant works integrate region proposal mechanisms with convolutional neural networks (CNN). Although CNN is proficient in extracting discriminative local features, grand challenges still exist to measure the likelihood of a bounding box containing a complete object (i.e., “objectness”). In this paper, we propose a novel WSOD framework with Objectness Distillation (i.e., WSOD2) by designing a tailored training mechanism for weakly-supervised object detection. Multiple regression targets are specifically determined by jointly considering bottom-up (BU) and top-down (TD) objectness from low-level measurement and CNN confidences with an adaptive linear combination. As bounding box regression can facilitate a region proposal learning to approach its regression target with high objectness during training, deep objectness representation learned from bottom-up evidences can be gradually distilled into CNN by optimization. We explore different adaptive training curves for BU/TD objectness, and show that the proposed WSOD2 can achieve state-of-the-art results. Zhaoyang Zeng, Bei Liu 0001, Jianlong Fu, Hongyang Chao, Lei Zhang 0001 |
ICCV | 4 |
| 2019 | Justlookup: One Millisecond Deep Feature Extraction for Point Clouds By Lookup TablesabstractDeep models are capable of fitting complex high dimensional functions while usually yielding large computation load. There is no way to speed up the inference process by classical lookup tables due to the high-dimensional input and limited memory size. Recently, a novel architecture (PointNet) for point clouds has demonstrated that it is possible to obtain a complicated deep function from a set of 3-variable functions. In this paper, we exploit this property and apply a lookup table to encode these 3-variable functions. This method ensures that the inference time is only determined by the memory access no matter how complicated the deep function is. We conduct extensive experiments on ModelNet and ShapeNet datasets and demonstrate that we can complete the inference process in 1.5 ms on an Intel i7-8700 CPU (single core mode), 32× speedup over the PointNet architecture without any performance degradation. Hongxin Lin, Zelin Xiao, Hongyang Chao, Shengyong Ding |
ICME | 4 |
| 2019 | LSCD: Low-rank and sparse cross-domain recommendation
Ling Huang 0002, Zhi-Lin Zhao 0001, Chang-Dong Wang 0001, Dong Huang 0001, Hongyang Chao |
Neurocomputing | 5 |
| 2019 | Item orientated recommendation by multi-view intact space learning with overlapping
Qi-Ying Hu, Ling Huang 0002, Chang-Dong Wang 0001, Hongyang Chao |
Knowl. Based Syst. | 4 |
| 2019 | See and chat: automatically generating viewer-level comments on images
Jingwen Chen 0001, Ting Yao 0003, Hongyang Chao |
Multim. Tools Appl. | 3 |
| 2019 | Enhancing Face Recognition from Massive Weakly Labeled Data of New Domains
Wei Xu 0014, Junyu Wu, Shengyong Ding, Linggan Lian, Hongyang Chao |
Neural Process. Lett. | 5 |
| 2019 | Multi-view intact space clustering
Ling Huang 0002, Hongyang Chao, Chang-Dong Wang 0001 |
Pattern Recognit. | 2 |
| 2019 | Learning Click-Based Deep Structure-Preserving Embeddings with Visual AttentionabstractOne fundamental problem in image search is to learn the ranking functions (i.e., the similarity between query and image). Recent progress on this topic has evolved through two paradigms: the text-based model and image ranker learning. The former relies on image surrounding texts, making the similarity sensitive to the quality of textual descriptions. The latter may suffer from the robustness problem when human-labeled query-image pairs cannot represent user search intent precisely. We demonstrate in this article that the preceding two limitations can be well mitigated by learning a cross-view embedding that leverages click data. Specifically, a novel click-based Deep Structure-Preserving Embeddings with visual Attention (DSPEA) model is presented, which consists of two components: deep convolutional neural networks followed by image embedding layers for learning visual embedding, and a deep neural networks for generating query semantic embedding. Meanwhile, visual attention is incorporated at the top of the convolutional neural network to reflect the relevant regions of the image to the query. Furthermore, considering the high dimension of the query space, a new click-based representation on a query set is proposed for alleviating this sparsity problem. The whole network is end-to-end trained by optimizing a large margin objective that combines cross-view ranking constraints with in-view neighborhood structure preservation constraints. On a large-scale click-based image dataset with 11.7 million queries and 1 million images, our model is shown to be powerful for keyword-based image search with superior performance over several state-of-the-art methods and achieves, to date, the best reported NDCG@25 of 52.21%. Yehao Li, Yingwei Pan, Ting Yao 0003, Hongyang Chao, Yong Rui, Tao Mei 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2018 | Overlapping Community Detection in Multi-view Brain Network
Ling Huang 0002, Chang-Dong Wang 0001, Hongyang Chao |
BIBM | 3 |
| 2018 | Image Blind Denoising With Generative Adversarial Network Based Noise ModelingabstractIn this paper, we consider a typical image blind denoising problem, which is to remove unknown noise from noisy images. As we all know, discriminative learning based methods, such as DnCNN, can achieve state-of-the-art denoising results, but they are not applicable to this problem due to the lack of paired training data. To tackle the barrier, we propose a novel two-step framework. First, a Generative Adversarial Network (GAN) is trained to estimate the noise distribution over the input noisy images and to generate noise samples. Second, the noise patches sampled from the first step are utilized to construct a paired training dataset, which is used, in turn, to train a deep Convolutional Neural Network (CNN) for denoising. Extensive experiments have been done to demonstrate the superiority of our approach in image blind denoising. Jingwen Chen 0001, Hongyang Chao, Ming Yang 0039 |
CVPR | 3 |
| 2018 | Jointly Localizing and Describing Events for Dense Video CaptioningabstractAutomatically describing a video with natural language is regarded as a fundamental challenge in computer vision. The problem nevertheless is not trivial especially when a video contains multiple events to be worthy of mention, which often happens in real videos. A valid question is how to temporally localize and then describe events, which is known as "dense video captioning." In this paper, we present a novel framework for dense video captioning that unifies the localization of temporal event proposals and sentence generation of each proposal, by jointly training them in an end-to-end manner. To combine these two worlds, we integrate a new design, namely descriptiveness regression, into a single shot detection structure to infer the descriptive complexity of each detected proposal via sentence generation. This in turn adjusts the temporal locations of each event proposal. Our model differs from existing dense video captioning methods since we propose a joint and global optimization of detection and captioning, and the framework uniquely capitalizes on an attribute-augmented video captioning architecture. Extensive experiments are conducted on ActivityNet Captions dataset and our framework shows clear improvements when compared to the state-of-the-art techniques. More remarkably, we obtain a new record: METEOR of 12.96% on ActivityNet Captions official test set. Yehao Li, Ting Yao 0003, Yingwei Pan, Hongyang Chao, Tao Mei 0001 |
CVPR | 4 |
| 2018 | Multi-view Proximity Learning for Clustering
Kun-Yu Lin, Ling Huang 0002, Chang-Dong Wang 0001, Hongyang Chao |
DASFAA (2) | 4 |
| 2018 | The Multi-Scale Deep Decoder for the Standard HEVC BitstreamsabstractAs we all know, there is strong multi-scale similarity among video frames. However, almost none of the current video coding standards takes this similarity into consideration. There exist two major problems when utilizing the multi-scale information at encoder-end: one is the extra motion models and the overheads brought by new motion parameters; the other is the extreme increment of the encoding algorithms' complexity. Is it possible to employ the multi-scale similarity only at the decoder-end to improve the decoded videos' quality, i.e., to further boost the coding efficiency? This paper mainly studies how to answer this question by proposing a novel Multi-Scale Deep Decoder (MSDD) for HEVC. Benefiting from the efficiency of deep learning technology (Convolutional Neural Network and Long Short-Term Memory network), MSDD achieves a higher coding efficiency only at the decoder-end without changing any encoding algorithms. Extensive experiments validate the feasibility and effectiveness of MSDD. MSDD leads to on averagely 6.5%, 8.0%, 6.4%, and 6.7% BD-rate reduction compared to HEVC anchor, for AI, LP, LB and RA coding configurations respectively. Especially for the videos with multi-scale similarity, the proposed approach obviously improves the coding efficiency indeed. Tingting Wang 0004, Wenhui Xiao, Mingjin Chen, Hongyang Chao |
DCC | 4 |
| 2018 | Fast H.264/AVC to HEVC Transcoding Based on Compressed Domain InformationabstractOne of the most important issue in applying High Efficiency Video Coding (HEVC) is how to fast transcode data from H.264/AVC to HEVC, because there are a large amount of H.264/AVC devices and videos. In this paper, unlike the traditional transcoding algorithms and other existing methods, we present a novel fast transcoding approach by exploiting the temporal and spatial information in the compressed domain which comes from both the input H.264/AVC and the previously-encoded HEVC bitstream. To solve the key issue of enhancing the relevance between compressed domain and the decisions making process in HEVC, the basic features, such as motion vectors, dividing depths and bit allocations, are further processed and composed into complex features. In the re-encoding process, an approach based on Bayesian rule is utilized to help fast decide the coding unit size and partition mode with above features. Experimental results show that this proposed transcoding approach can significantly reduce the computational complexity while retaining high coding efficiency. Juan Zha, Hongyang Chao |
DCC | 3 |
| 2018 | A Harmonic Motif Modularity Approach for Multi-layer Network Community DetectionabstractDuring the past several years, multi-layer network community detection has drawn an increasing amount of attention and many approaches have been developed from different perspectives. Despite the success, they mainly rely on the lower-order connectivity structure at the level of individual nodes and edges. However, the higher-order connectivity structure plays the essential role as the building block for multiplex networks, which may contain better signature of community than edge. The main challenge in utilizing higher-order structure for multi-layer network community detection is that the most representative higher-order structure may vary from one layer to another. In this paper, we propose a higher-order structural approach for multi-layer network community detection, termed harmonic motif modularity (HM-Modularity). The key idea is to design a novel higher-order structure, termed harmonic motif, which is able to integrate higher-order structural information from multiple layers to construct a primary layer. The higher-order structural information of each individual layer is also extracted, which is taken as the auxiliary information for discovering the multi-layer community structure. A coupling is established between the primary layer and each auxiliary layer. Finally, a harmonic motif modularity is designed to generate the community structure. By solving the optimization problem of the harmonic motif modularity, the community labels of the primary layer can be obtained to reveal the community structure of the original multi-layer network. Experiments have been conducted to show the effectiveness of the proposed method. Ling Huang 0002, Chang-Dong Wang 0001, Hongyang Chao |
ICDM | 3 |
| 2018 | Matrix completion with capped nuclear norm via majorized proximal minimization
Shenfen Kuang, Hongyang Chao, Qia Li |
Neurocomputing | 2 |
| 2017 | Building an End-to-End Spatial-Temporal Convolutional Network for Video Super-ResolutionabstractWe propose an end-to-end deep network for video super-resolution. Our network is composed of a spatial component that encodes intra-frame visual patterns, a temporal component that discovers inter-frame relations, and a reconstruction component that aggregates information to predict details. We make the spatial component deep, so that it can better leverage spatial redundancies for rebuilding high-frequency structures. We organize the temporal component in a bidirectional and multi-scale fashion, to better capture how frames change across time. The effectiveness of the proposed approach is highlighted on two datasets, where we observe substantial improvements relative to the state of the arts. Jun Guo 0024, Hongyang Chao |
AAAI | 2 |
| 2017 | One-To-Many Network for Visually Pleasing Compression Artifacts ReductionabstractWe consider the compression artifacts reduction problem, where a compressed image is transformed into an artifact-free image. Recent approaches for this problem typically train a one-to-one mapping using a per-pixel L2loss between the outputs and the ground-truths. We point out that these approaches used to produce overly smooth results, and PSNR doesn't reflect their real performance. In this paper, we propose a one-to-many network, which measures output quality using a perceptual loss, a naturalness loss, and a JPEG loss. We also avoid grid-like artifacts during deconvolution using a shift-and-average strategy. Extensive experimental results demonstrate the dramatic visual improvement of our approach over the state of the arts. Jun Guo 0024, Hongyang Chao |
CVPR | 2 |
| 2017 | A Novel Deep Learning-Based Method of Improving Coding Efficiency from the Decoder-End for HEVCabstractImproving the coding efficiency is the eternal theme in video coding field. The traditional way for this purpose is to reduce the redundancies inside videos by adding numerous coding options at the encoder side. However, no matter what we have done, it is still hard to guarantee the optimal coding efficiency. On the other hand, the decoded video can be treated as a certain compressive sampling of the original video. According to the compressive sensing theory, it might be possible to further enhance the quality of the decoded video by some restoration methods. Different from the traditional methods, without changing the encoding algorithm, this paper focuses on an approach to improve the video's quality at the decoder end, which equals to further boosting the coding efficiency. Furthermore, we propose a very deep convolutional neural network to automatically remove the artifacts and enhance the details of HEVC-compressed videos, by utilizing that underused information left in the bit-streams and external images. Benefit from the prowess and efficiency of the fully end-to-end feed forward architecture, our approach can be treated as a better decoder to efficiently obtain the decoded frames with higher quality. Extensive experiments indicate our approach can further improve the coding efficiency post the deblocking and SAO in current HEVC decoder, averagely 5.0%, 6.4%, 5.3%, 5.5% BD-rate reduction for all intra, lowdelay P, lowdelay B and random access configurations respectively. This method can aslo be extended to any video coding standards. Tingting Wang 0004, Mingjin Chen, Hongyang Chao |
DCC | 3 |
| 2017 | An Optimally Scalable and Cost-Effective Algorithm for 1/8-Pixel Motion Estimation for HEVCabstractIn the state-of-the-art H.265/HEVC video coding standard, the highest resolution of motion vectors is 1/4-pixel. However, the 1/8-pixel motion vector performs better in some specific scenes but brings more time cost. In this paper, we propose an optimally scalable and cost-effective algorithm for fractional-pixel motion estimation with 1/8-pixel motion vector resolution, which fits well to varying and constrained computing resources. In proposed method, each of the search points is assigned with a cost-effective priority, which determines the order in a refinement search. Also, we demonstrate a new search pattern, which makes it more likely to check the optimal motion vector earlier. Extensive experiments indicate that our method is capable of maximizing the R-D gains of each search point. As a side product, our approach can also serve as a fast algorithm. It can enhance time performance greatly with nearly no loss in R-D gains. Even compared with full fractional pixels search with 1/4-pixel resolution, the proposed method shows averagely 9.8% improvement in terms of time performance, with almost 1.3% increase in coding efficiency. Wenhui Xiao, Tingting Wang 0004, Hongyang Chao |
DCC | 4 |
| 2017 | Abnormal Event Detection in Surveillance Video: A Compressed Domain Approach for HEVCabstractRecently, detecting abnormal events in surveillance videos has become one of the most important tasks of video analysis. There is a huge demand in developing fast and accurate abnormal event detection approach. However, traditional pixel-domain approaches are time-consuming and require fully decoding of the bit streams. On the other hand, the compression format may provide useful information to solve the challenge: compared to raw pixels, the advantage of the compression format is that it already contains some valuable clues for video analysis. In this work, we focus on anomaly detection in traffic video with the information provided in the HEVC compressed domain. There are mainly two contributions. The first is that we propose a novel feature namely motion energy intensity (MEI) to represent the motion intensity within coding unit, based on the MV fields, bit allocations and partitions of coding units in the HEVC compressed domain, which can provide energy and motion information of the LCU. The second contribution is that we have proposed an anomaly detection model based on the MEI. Furthermore, this approach can also serve as preprocess for pixel domain detection algorithms, by labeling suspicious part of the video which are likely to be abnormities. Figure 1 shows the flowchart of the proposed algorithm. Experimental results show that the proposed approach can obtain rather high detecting accuracy with very little overhead. The processing speed for detection is as fast as 1200 fps on 480x360 video. When served as a preprocess approach, the pixel-domain methods may also benefit with up to 50% time reduction. Hongyang Chao |
DCC | 2 |
| 2017 | A fast intra-prediction decision algorithm in inter-frame based on a novel feature of HEVCabstractWith the quad-tree based coding structure and more flexible intramodes, the coding efficiency provided by intra-technique in interframes of HEVC is much higher than the preceding standard H.264/AVC. However, the computing complexity is also significantly increased. Although only a few CUs are encoded as intra-mode in inter-frames finally, almost all CUs need to be checked all the intra-options to obtain the optimal mode, which may be in fact unnecessary. Unlike most existing fast algorithms, which focus on accelerating intra-prediction itself, we propose a novel feature to quantitatively estimate the possibilities of doing intra-prediction for CUs. Based on this feature, we further design a fast intra-prediction decision algorithm to determine whether a CU needs to be checked the intra-prediction. The experimental results show that our algorithm can reduce more than 60% computing resources with a negligible loss in rate-distortion performance for intra-module in all inter-frames. Besides, our method can be compatible with all other existing fast algorithms. Tingting Wang 0004, Yangyang Men, Hongyang Chao |
ICASSP | 4 |
| 2017 | Searching Personal Photos on the Phone with Instant Visual Query Suggestion and Joint Text-Image HashingabstractThe ubiquitous mobile devices have led to the unprecedented growing of personal photo collections on the phone. One significant pain point of today's mobile users is instantly finding specific photos of what they want. Existing applications (e.g., Google Photo and OneDrive) have predominantly focused on cloud-based solutions, while leaving the client-side challenges (e.g., query formulation, photo tagging and search, etc.) unsolved. This considerably hinders user experience on the phone. In this paper, we present an innovative personal photo search system on the phone, which enables instant and accurate photo search by visual query suggestion and joint text-image hashing. Specifically, the system is characterized by several distinctive properties: 1) visual query suggestion (VQS) to facilitate the formulation of queries in a joint text-image form, 2) light-weight convolutional and sequential deep neural networks to extract representations for both photos and queries, and 3) joint text-image hashing (with compact binary codes) to facilitate binary image search and VQS. It is worth noting that all the components run on the phone with client optimization by deep learning techniques. We have collected 270 photo albums taken by 30 mobile users (corresponding to 37,000 personal photos) and conducted a series of field studies. We show that our system significantly outperforms the existing client-based solutions by 10 x in terms of search efficiency, and 92.3% precision in terms of search accuracy, leading to a remarkably better user experience of photo discovery on the phone. Zhaoyang Zeng, Jianlong Fu, Hongyang Chao, Tao Mei 0001 |
ACM Multimedia | 3 |
| 2017 | Efficient l q norm based sparse subspace clustering via smooth IRLS and ADMM
Shenfen Kuang, Hongyang Chao, Jun Yang 0051 |
Multim. Tools Appl. | 2 |
| 2016 | Joint Multiview Segmentation and Localization of RGB-D Images Using Depth-Induced Silhouette ConsistencyabstractIn this paper, we propose an RGB-D camera localization approach which takes an effective geometry constraint, i.e. silhouette consistency, into consideration. Unlike existing approaches which usually assume the silhouettes are provided, we consider more practical scenarios and generate the silhouettes for multiple views on the fly. To obtain a set of accurate silhouettes, precise camera poses are required to propagate segmentation cues across views. To perform better localization, accurate silhouettes are needed to constrain camera poses. Therefore the two problems are intertwined with each other and require a joint treatment. Facilitated by the available depth, we introduce a simple but effective silhouette consistency energy term that binds traditional appearance-based multiview segmentation cost and RGB-D frame-to-frame matching cost together. Optimization of the problem w.r.t. binary segmentation masks and camera poses naturally fits in the graph cut minimization framework and the Gauss-Newton non-linear least-squares method respectively. Experiments show that the proposed approach achieves state-of-the-arts performance on both tasks of image segmentation and camera localization. Chi Zhang 0069, Zhiwei Li 0006, Rui Cai 0002, Hongyang Chao, Yong Rui |
CVPR | 4 |
| 2016 | A Framework of Complexity Optimally Scalable Algorithms for HEVCabstractDifferent from conventional profiles in the state-of-the-art video coding standard HEVC and related optimization methods, we focus on building the optimally scalable algorithms under constrained and varying computational capacity to take full advantages of HEVC as far as possible in order to meet the growing demands of computational capacity adaptive applications such as real-time video communication and video coding on different mobile devices. We propose a video coding framework based on priority order for a special profile and give the general thoughts of designing algorithms by utilizing cost-performance as priority. For the framework, we invent a feasible solution by introducing a three-level coding structure and some novel features that express the relationship between video contents and their priorities. Experimental results partially prove our framework may nearly achieve the optimal coding efficiency under the continuously changeable computing limitations at every time with negligible extra time consuming. Tingting Wang 0004, Hongyang Chao, Feng Wu 0001 |
DCC | 4 |
| 2016 | Building Dual-Domain Representations for Compression Artifacts Reduction
Jun Guo 0024, Hongyang Chao |
ECCV (1) | 2 |
| 2016 | Share-and-Chat: Achieving Human-Level Video Commenting by Search and Multi-View EmbeddingabstractVideo has become a predominant social media for the booming live interactions. Automatic generation of emotional comments to a video has great potential to significantly increase user engagement in many socio-video applications (e.g., chat bot). Nevertheless, the problem of video commenting has been overlooked by the research community. The major challenges are that the generated comments are to be not only as natural as those from human beings, but also relevant to the video content. We present in this paper a novel two-stage deep learning-based approach to automatic video commenting. Our approach consists of two components. The first component, similar video search, efficiently finds the visually similar videos w.r.t. a given video using approximate nearest-neighbor search based on the learned deep video representations, while the second dynamic ranking effectively ranks the comments associated with the searched similar videos by learning a deep multi-view embedding space. For modeling the emotional view of videos, we incorporate visual sentiment, video content, and text comments into the learning of the embedding space. On a newly collected dataset with over 102K videos and 10.6M comments, we demonstrate that our approach outperforms several state-of-the-art methods and achieves human-level video commenting. Yehao Li, Ting Yao 0003, Tao Mei 0001, Hongyang Chao, Yong Rui |
ACM Multimedia | 4 |
| 2016 | Real-Time Logo Recognition from Live Video Streams Using an Elastic Cloud Platform
Jianbing Ding, Hongyang Chao, Mansheng Yang |
WAIM (2) | 2 |
| 2016 | A versatile sparse representation based post-processing method for improving image super-resolution
Jun Yang 0051, Jun Guo 0024, Hongyang Chao |
Neurocomputing | 3 |
| 2016 | High-quality image restoration from partial mixed adaptive-random measurements
Jun Yang 0051, Wei E. I. Sha, Hongyang Chao, Zhu Jin |
Multim. Tools Appl. | 3 |
| 2016 | Building Hierarchical Representations for Oracle Character and Sketch RecognitionabstractIn this paper, we study oracle character recognition and general sketch recognition. First, a data set of oracle characters, which are the oldest hieroglyphs in China yet remain a part of modern Chinese characters, is collected for analysis. Second, typical visual representations in shape- and sketch-related works are evaluated. We analyze the problems suffered when addressing these representations and determine several representation design criteria. Based on the analysis, we propose a novel hierarchical representation that combines a Gabor-related low-level representation and a sparse-encoder-related mid-level representation. Extensive experiments show the effectiveness of the proposed representation in both oracle character recognition and general sketch recognition. The proposed representation is also complementary to convolutional neural network (CNN)-based models. We introduce a solution to combine the proposed representation with CNN-based models, and achieve better performances over both approaches. This solution has beaten humans at recognizing general sketches. Jun Guo 0024, Changhu Wang, Edgar Roman-Rangel, Hongyang Chao, Yong Rui |
IEEE Trans. Image Process. | 4 |
| 2016 | Effective Clipart Image Vectorization through Direct Optimization of BezigonsabstractBezigons, i.e., closed paths composed of Bézier curves, have been widely employed to describe shapes in image vectorization results. However, most existing vectorization techniques infer the bezigons by simply approximating an intermediate vector representation (such as polygons). Consequently, the resultant bezigons are sometimes imperfect due to accumulated errors, fitting ambiguities, and a lack of curve priors, especially for low-resolution images. In this paper, we describe a novel method for vectorizing clipart images. In contrast to previous methods, we directly optimize the bezigons rather than using other intermediate representations; therefore, the resultant bezigons are not only of higher fidelity compared with the original raster image but also more reasonable because they were traced by a proficient expert. To enable such optimization, we have overcome several challenges and have devised a differentiable data energy as well as several curve-based prior terms. To improve the efficiency of the optimization, we also take advantage of the local control property of bezigons and adopt an overlapped piecewise optimization strategy. The experimental results show that our method outperforms both the current state-of-the-art method and commonly used commercial software in terms of bezigon quality. Ming Yang 0039, Hongyang Chao, Chi Zhang 0026, Jun Guo 0024, Lu Yuan 0001, Jian Sun 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2015 | Building Effective Representations for Sketch RecognitionabstractAs the popularity of touch-screen devices, understanding a user's hand-drawn sketch has become an increasingly important research topic in artificial intelligence and computer vision. However, different from natural images, the hand-drawn sketches are often highly abstract, with sparse visual information and large intra-class variance, making the problem more challenging. In this work, we study how to build effective representations for sketch recognition. First, to capture saliency patterns of different scales and spatial arrangements, a Gabor-based low-level representation is proposed. Then, based on this representation, to discovery more complex patterns in a sketch, a Hybrid Multilayer Sparse Coding (HMSC) model is proposed to learn mid-level representations. An improved dictionary learning algorithm is also leveraged in HMSC to reduce overfitting to common but trivial patterns. Extensive experiments show that the proposed representations are highly discriminative and lead to large improvements over the state of the arts. Jun Guo 0024, Changhu Wang, Hongyang Chao |
AAAI | 3 |
| 2015 | MeshStereo: A Global Stereo Model with Mesh Alignment Regularization for View InterpolationabstractWe present a novel global stereo model designed for view interpolation. Unlike existing stereo models which only output a disparity map, our model is able to output a 3D triangular mesh, which can be directly used for view interpolation. To this aim, we partition the input stereo images into 2D triangles with shared vertices. Lifting the 2D triangulation to 3D naturally generates a corresponding mesh. A technical difficulty is to properly split vertices to multiple copies when they appear at depth discontinuous boundaries. To deal with this problem, we formulate our objective as a two-layer MRF, with the upper layer modeling the splitting properties of the vertices and the lower layer optimizing a region-based stereo matching. Experiments on the Middlebury and the Herodion datasets demonstrate that our model is able to synthesize visually coherent new view angles with high PSNR, as well as outputting high quality disparity maps which rank at the first place on the new challenging high resolution Middlebury 3.0 benchmark. Chi Zhang 0069, Zhiwei Li 0006, Yanhua Cheng, Rui Cai 0002, Hongyang Chao, Yong Rui |
ICCV | 5 |
| 2015 | Inverse halftoning with grouping singular value decompositionabstractThe objective of inverse halftoning refers to reconstruct a high quality gray scale image from bi-level halftone image. However, reconstructing continuous-tone images from their halftoned versions is highly underdetermined, making this technique very difficult. In this paper, we present the Grouping Singular Value Decomposition (G-SVD), a novel approach which first groups similar image patches as input and then characterizes lower-dimensional regions in input space where the data density is peaked. By adding a constraint formulated via G-SVD into inverse halftoning, noises are separated from meaningful contents and similarity of nonlocal image patches is promoted. Our experiments shown that the proposed approach could improve the visual quality of reconstructed results and outperformed the state of the arts in terms of both objective and subjective measurements. Jun Yang 0051, Jun Guo 0024, Hongyang Chao |
ICIP | 3 |
| 2015 | ECISER: Efficient Clip-art Image SEgmentation by Re-rasterization
Ming Yang 0039, Hongyang Chao |
Comput. Aided Des. | 2 |
| 2015 | Deep feature learning with relative distance comparison for person re-identification
Shengyong Ding, Liang Lin 0004, Guangrun Wang, Hongyang Chao |
Pattern Recognit. | 4 |
| 2014 | As-Rigid-As-Possible Stereo under Second Order Smoothness Priors
Chi Zhang 0069, Zhiwei Li 0006, Rui Cai 0002, Hongyang Chao, Yong Rui |
ECCV (2) | 4 |
| 2014 | A rapid abnormal event detection method for surveillance video based on a novel feature in compressed domain of HEVCabstractEvent detection plays an essential role in video content analysis. On the other hand, according to our analysis, the coding structures in new video coding standard High Efficient Video Coding (HEVC) have a high correlation with video contents. Hence there is large potential to identify events by reusing coding structures in HEVC, which can save a huge amount of computational resources. In this paper, we proposed a new compressed-domain feature for abnormal event detection, namely Motion Intensity Count (MIC), which makes use of motion vectors, coding unit and prediction unit modes in HEVC with little computational cost. MIC can well predict the normal paths of moving objects, which enables us to identify motions in unexpected locations where abnormal events are likely to happen. Our experiments show that MIC can correctly detect abnormal events at about 1250 fps. Ming Yang 0039, Yangyang Men, Hongyang Chao |
ICME | 5 |
| 2013 | An optimally scalable and cost-effective fractional-pixel motion estimation algorithm for HEVCabstractFractional-pixel motion compensation is still one of the most time-consuming parts in the upcoming High Efficiency Video Coding (HEVC) standard. In this paper, we propose an optimally scalable and cost-effective fractional-pixel motion estimation (FPME) algorithm to optimally fit to different and varying constrains of computing resources. Our main contribution include two aspects. Firstly, an optimally scalable and cost-effective FPME algorithm based on a cost-benefit analysis is proposed, where we present a improved fractional-pixel MV prediction method and a new cost-effective priority for each search point in HEVC. Secondly, a complexity adjustment strategy is delivered to enable the ability for the FPME to adjust its complexity to match different given constraints on time. Experiments show that the proposed algorithm can achieve best R-D performance while optimally adjust its complexity, based on any given time constraints. As a side product, the proposed algorithm can also serve as the best fast algorithm which has already reduced computing complexity by a factor 74% with almost no loss on PSNR and bitrates. Hongyang Chao |
ICASSP | 3 |
| 2013 | An optimally complexity scalable multi-mode decision algorithm for HEVCabstractQuad-tree based Coding Unit structure in HEVC provides more motion compensation sizes to improve rate-distortion performance at the cost of greatly increased computational complexity. Different from other researches on fast algorithms, we develop an optimally complexity scalable multi-mode decision algorithm (OCSMD) for HEVC. There are two major contributions in this paper. The first one is a novel feature proposed to describe the relationship between MV field and CU depth. The second is that we build a cost-performance priority predicting model in frame level based on the feature with negligible overhead as well as no conflict with the standard. Our method may allocate computational resources to the MD of all the CUs in frame level under arbitrary complexity constraints, while obtaining nearly optimal coding performance. The experimental result shows that our algorithm can adjust complexity under varying computing capacity while achieving near-optimal R-D performance. Shichao Huang, Hongyang Chao |
ICIP | 4 |
| 2013 | Discovering Video Shot Categories by Unsupervised Stochastic Graph PartitionabstractVideo shots are often treated as the basic elements for retrieving information from videos. In recent years, video shot categorization has received increasing attention, but most of the methods involve a procedure of supervised learning, i.e., training a multi-class predictor (classifier) on the labeled data. In this paper, we study a general framework to unsupervisedly discover video shot categories. The contributions are three-fold in feature, representation, and inference: (1) A new feature is proposed to capture local information in videos, defined with small video patches (e.g.,$11 \times 11 \times 5$pixels). A dictionary of video words can be thus clustered off-line, characterizing both appearance and motion dynamics. (2) We pose the problem of categorization as an automated graph partition task, in that each graph vertex represents a video shot, and a partitioned sub-graph consisting of connected graph vertices represents a clustered category. The model of each video shot category can be analytically calculated by a projection pursuit type of learning process. (3) An MCMC-based cluster sampling algorithm, namely Swendsen-Wang cuts, is adopted to efficiently solve the graph partition. Unlike traditional graph partition techniques, this algorithm is able to explore the nearly global optimal solution and eliminate the need for good initialization. We apply our method on a wide variety of 1600 video shots collected from Internet as well as a subset of TRECVID 2010 data, and two benchmark metrics, i.e., Purity and Conditional Entropy, are adopted for evaluating performance. The experimental results demonstrate superior performance of our method over other popular state-of-the-art methods. Xiaohua Duan, Liang Lin 0004, Hongyang Chao |
IEEE Trans. Multim. | 3 |
| 2012 | Object categorization with sketch representation and generalized samples
Liang Lin 0004, Xiaobai Liu, Shaowu Peng, Hongyang Chao, Yongtian Wang, Bo Jiang 0002 |
Pattern Recognit. | 4 |
| 2011 | Group Crumb: Sharing Web Navigation by Visualizing Group Traces on the Web
Qing Wang 0018, Gaoqiang Zheng, HuiYou Chang, Hongyang Chao |
ECSCW | 5 |
| 2011 | On Combining Fractional-Pixel Interpolation and Motion Estimation: A Cost-Effective ApproachabstractThe additional complexity of the adoption of fractional-pixel motion compensation technology arises from two aspects: fractional-pixel interpolation (FPI) and fractional-pixel motion estimation (FPME). Different from current fast algorithms, we use the internal link between FPME and FPI as a factor in considering optimization by integrally manipulating them rather than attempting to speed them up separately. In this paper, a refinement search order for FPME is proposed to satisfy the criteria of cost/performance efficiency. And then, some strategies, i.e., FPME skipping, early termination and search pattern pruning, are also given for reducing the number of search positions with negligible coding loss. We also propose a FPI algorithm to save redundant interpolation as well as reduce duplicate calculation. Experimental results show that our integrated algorithm significantly improves the overall speed of FPME and FPI. Compared with the FFPS+XFPI and CBFPS+XFPI, the proposed algorithm has already reduced the speed by a factor of 65% and 32%. Additionally, our FPI algorithm can be used to cooperate with any fast FPME algorithms to greatly reduce the computational time of FPI. Jiyuan Lu, Peizhao Zhang, Hongyang Chao, Paul S. Fisher |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2010 | An Integrated Algorithm for Fractional Pixel Interpolation and Motion Estimation of H.264abstractOur work has two main contributions. The first is to give a cost performance search order which checks fractional SPs by not only maximizing the expectation for the R-D gains but also minimizing the computational cost at the same time. We assign 48 SPs to one of six categories with different cost performance priorities. Each category has a different probability of being the optimal fractional MV (performance), and it will have a different computational complexity for interpolation (cost). This algorithm checks fractional SPs in the order of the likelihood for being the optimal position in consideration with the simplicity of interpolation. After that we propose a strategy to further reduce the number of SPs for FPME. Jiyuan Lu, Peizhao Zhang, Hongyang Chao, Paul S. Fisher |
DCC | 3 |
| 2010 | Learning Shape Detector by Quantizing Curve Segments with Multiple Distance Metrics
Ping Luo 0002, Liang Lin 0004, Hongyang Chao |
ECCV (3) | 3 |
| 2010 | Annotating and navigating tourist videosabstractDue to the rapid increase in video capture technology, more and more tourist videos are captured every day, creating a challenge for organization and association with metadata. In this paper, we present a novel system for annotating and navigating tourist videos. Placing annotations in a video is difficult because of the need to track the movement of the camera. Navigation of a regular video is also challenging due to the sequential nature of the media. To overcome these challenges, we introduce a system for registering videos to geo-referenced 3D models and analyzing the video contents. We also introduce a novel scheduling algorithm for showing annotations in video. We show results in automatically annotated videos and in a map-based application for browsing videos. Our user study indicates the system is very useful. Qinlin Li, Hongyang Chao, Billy Chen, Eyal Ofek, Ying-Qing Xu |
GIS | 3 |
| 2010 | An adaptive real-time descreening method based on SVM and improved SUSAN filterabstractScanned halftone images are degraded for the presence of screen patterns. It's a challenge to automatically detect the halftone images and remove the noises on the fly. This paper proposes a novel adaptive real-time descreening method based on Support Vector Machine (SVM) and modified Smoothing over Univalue Segment Assimilating Nucleus (SUSAN) filter for imaging devices, including scanners and multifunction printers. The proposed algorithm contains two major steps: image classification and adaptive descreening. The image classification uses SVM methods to accurately classify the scanned images into three categories: continuous tone, amplitude modulation (AM) halftone or frequency modulation (FM) halftone. The halftone images are needed to descreen. The proposed descreening method is based on modified SUSAN filter. It considers screen cell size to choose the optimal filter parameters which can preserve more high frequency image detail. The experiment results show that the algorithm is effective and fully automatism, and maintains higher image quality. Xiaohua Duan, Guifeng Zheng, Hongyang Chao |
ICASSP | 3 |
| 2010 | A split and merge algorithm for inter-mode decision in extended macroblocksabstractExtended macroblock size (EMS) improves the highest possible coding efficiency for the current video coding standards, but a large number of additional candidate coding modes results in high computational expense. In this paper, the correlation among different modes is exploited to drastically reduce the number of candidate modes in EMS. And a fast inter-mode decision algorithm for EMS is proposed based on above facts. Furthermore, a split and merge technology is used to find the optimal partitioning with computational efficiency. Experimental results show that the proposed fast inter-mode decision algorithm reduces the computational complexity significantly with negligible coding lost. Jiyuan Lu, Peizhao Zhang, Hongyang Chao |
ICIP | 3 |
| 2010 | Semantics-driven portrait cartoon stylizationabstractThis paper proposes an efficient framework for transforming an input human portrait image into an artistic cartoon style. Compared to the previous work of non-photorealistic rendering (NPR), our method exploits the portrait semantics for enriching and manipulating the cartooning style, based on a semantic grammar model. The proposed framework consists of two phases: a portrait parsing phase to localize and recognize facial components in a hierarchic manner, and further calculate the portrait saliency with the facial components; a cartoon stylizing phase to abstract and cartoonize the portrait according to the parsed semantics and saliency, in which the regions and structure (edges/boundaries) of the portrait are rendered in two layers. In the experiments, we test our method with different types of human portraits: daily photos, identification photos, and studio photos, and find satisfactory results; a quantitative evaluation of subjective preference is presented as well. Ming Yang 0039, Ping Luo 0002, Liang Lin 0004, Hongyang Chao |
ICIP | 5 |
| 2010 | Inter-mode decision with varied computational complexityabstractVariable block size motion compensation significantly improves the rate-distortion performance of video coding at the cost of high computational complexity. Currently, fast inter-mode decision algorithms only improve the speed of intermode decision (MD) but do not provide a flexible computational complexity control to adapt to different hardware platforms with optimized rate-distortion (R-D) performance. Instead of speeding up inter-mode decision merely, we attempt to propose a complexity adjustable inter-mode decision algorithm to attain optimized coding performance under different computational complexity constraints herein. Our algorithm predicts the Lagrangian cost and complexity slope (J-C slope) of MD for each macroblock (MB) by exploiting their temporal and spatial correlations. MD is applied on the MBs with larger J-C slope to provide the better prediction performance with less computational cost. Adjustable complexity is obtained by preset the number of MBs on which MD is applied for each frame. According to our experiments, the algorithm can both freely adjust the computational complexity and provides an improved R-D performance under different computational constraints. Jiyuan Lu, Peizhao Zhang, Hongyang Chao, Paul S. Fisher |
VCIP | 3 |
| 2009 | Hierarchical 3D perception from a single imageabstractInspirited by the human vision mechanism, this paper discusses a hierarchical grammar model for 3D inference of man-made object from a single image. This model decomposes an object with two layers: (i) 3D parts (primitives) with 3D spatial relationship and (ii) 2D aspects with prediction (production) rules. Thus each object is represented by a set of co-related 3D primitives that are generated by a set of 2D aspects. The 3D relationships can be learned for each object category specifically by a discriminative boosting method, and the 2D production rules are defined according to the human visual experience. With this representation, the inference follows a data-driven Markov Chain Monte Carlo computing method in the Bayesian framework. In the experiments, we demonstrate the 3D inference results on 8 object categories and also propose a psychology analysis to evaluate our work. Ping Luo 0002, Liang Lin 0004, Hongyang Chao |
ICIP | 4 |
| 2007 | An Optimization Method for Real-Time Natural Phenomena Simulation on WinCE PlatformabstractWe present a new optimized algorithm of particle system for simulating flame, so that it can be implemented on WinCE platform which has only limited resources. Our approach mainly intends to store some traces first and then load them randomly during the simulation to reduce the computation complexity of the particle system. We have tested this approach in the simulation of a flame, and our method shows a significant performance gain with little loss in the visual appearance of the simulation, this indicates the potential to extend our approach to particle simulation involved with other natural phenomena for resource limited platforms. Hongyang Chao |
COMPSAC (2) | 4 |
| 2006 | A High Accurate Predictor Based Fractional Pixel Search for H.264abstractIn this paper, we proposed a new fractional pixel motion estimation method that effectively extends predictor based algorithm for integer pixel search to fractional pixel search. It generates very precise predictors and makes use of simple refinement patterns in successive search. According to our experiments, the proposed method not only has superior speed compared with other methods, but also produces the same PSNR as full fractional pixel search (FFPS) for all block modes of H.264. However, current fast fractional pixel search methods adopted by H.264 JM do not guarantee the same search quality as FFPS in different block mode. Hongyang Chao, Jiyuan Lu |
ICIP | 1 |
| 2002 | Rate Scalable Video Compression Based on Flexible Block Wavelet Coding TechniqueabstractSummary form only given. We present a new wavelet coding technique, flexible block wavelet coding (FBWC). It is specially designed for delta frame compression. FBWC not only makes delta frame compression more efficient but also keeps the rate scalability and other scalabilities of wavelet based scalable video compression techniques. Based on this new delta frame compression technique, we have implemented a highly efficient rate scalable video codec. Its overall performance surpasses the video codec based on SPIHT and is comparable to that of H.263. Using FBWC, delta frames are compressed in three steps: First, we use motion compensation to eliminate the temporal redundancy between adjacent frames and the predicted error frames (PEF) are generated. Next, the PEF are transformed into the wavelet domain using the flexible block wavelet transform (FBWT). Finally, the zerotree coding mechanism and entropy coder are used to actually compress the FBWT transformed wavelet coefficients to the target data rate. Hongyang Chao |
DCC | 2 |
| 2001 | Rate scalable video compression based on flexible block wavelet coding techniqueabstractRate scalable video compression techniques are very attractive, because they are able to encode the video in embedded way that makes the decoding process more flexible to the bandwidth changes. Embedded wavelet coding technology plays an important role in this area. Traditional wavelet transforms are globally optimized. It is very efficient to code the natural images where sudden transitions rarely exist. On the other hand, video compression mostly deals with the differences between adjacent frames, also known as delta frames. Unfortunately, commonly used wavelet transforms are not so efficient on delta frame coding. We present a new wavelet coding technique - flexible block wavelet coding (FBWC). It is specially designed for delta frame compression. FBWC not only makes delta frame compression more efficient but also keeps the rate scalability of wavelet based rate scalable video compression techniques, and other scalabilities as well. Based on this new delta frame compression technique, we implement a highly efficient rate scalable video codec. Its overall performance surpasses the video codec based on SPIHT and is comparable to that of H.263. Hongyang Chao |
MMSP | 1 |