Xin Jin 0015

dblp:68/3340-15 · DBLP profile ↗
← Back
64ranked-venue papers
31as first author
30since 2021 · last 2026
0000-0003-3873-1653ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 45 · 22 first-author · 18 since 2021Artificial intelligence and machine learning · 11 · 6 first-author · 4 since 2021Computer networks · 7 · 4 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Video Echoed in Music: Semantic, Temporal, and Rhythmic Alignment for Video-to-Music Generation
abstract
Video-to-Music generation seeks to generate musically appropriate background music that enhances audiovisual immersion for videos. However, current approaches suffer from two critical limitations: 1) incomplete representation of video details, leading to weak alignment, and 2) inadequate temporal and rhythmic correspondence, particularly in achieving precise beat synchronization. To address the challenges, we propose Video Echoed in Music (VeM), a latent music diffusion that generates high-quality soundtracks with semantic, temporal, and rhythmic alignment for input videos. To capture video details comprehensively, VeM employs a hierarchical video parsing that acts as a music conductor, orchestrating multi-level information across modalities. Modality-specific encoders, coupled with a storyboard-guided cross-attention mechanism (SG-CAtt), integrate semantic cues while maintaining temporal coherence through position and duration encoding. For rhythmic precision, the frame-level transition-beat aligner and adapter (TB-As) dynamically synchronize visual scene transitions with music beats. We further contribute a novel video-music paired dataset sourced from e-commerce advertisements and video-sharing platforms, which imposes stricter transition-beat synchronization requirements. Meanwhile, we introduce novel metrics tailored to the task. Experimental results demonstrate superiority, particularly in semantic relevance and rhythmic precision.
Xinyi Tong 0001, Yiran Zhu, Jishang Chen, Chunru Zhan, Tianle Wang 0007, Sirui Zhang, Nian Liu 0003, Tiezheng Ge, Duo Xu 0004, Xin Jin 0015, Feng Yu 0032, Song-Chun Zhu
AAAI10
2026 Fine-grained aesthetic multi-attribute captioning with aligned vision-language representations
Yehui Liu, Minzheng Jia, Yongqiang Kong, Xin Jin 0015, Ping Shi 0001
J. Vis. Commun. Image Represent.6
2025 KLMN: Knowledge distillation based lightweight multi-clue image forgery detection and localization
abstract
Current image forensics methods often utilize image features from various frequency domains. However, the effective use of these features frequently depends on complex network architectures and a large number of parameters. In this paper, we introduce a lightweight Multi-Clue image forgery detection and localization network (KLMN) along with a novel extractor-adapter framework based on multi-target knowledge distillation. Our extractor-adapter framework effectively integrates RGB features with low-level features, enhancing resilience against various forms of forgery while reducing the number of parameters. We also propose a feature-aggregated decoder that reconstructs the predicted mask by remixing features. The multi-target knowledge distillation process includes intermediate feature distillation, predicted map distillation from the teacher networks, and supervised training using ground truths. Experimental results on multiple datasets demonstrate that our approach can accurately and efficiently detect and localize manipulated regions while meeting the requirements for fewer parameters and reduced memory usage.
Heng Huang 0002, Xin Jin 0015
ICASSP3
2025 InteractMove: Text-Controlled Human-Object Interaction Generation in 3D Scenes with Movable Objects
Xinhao Cai, Minghang Zheng, Xin Jin 0015, Yang Liu 0105
ACM Multimedia3
2025 VADB: A Large-Scale Video Aesthetic Database with Professional and Multi-Dimensional Annotations
abstract
Video aesthetic assessment, a vital area in multimedia computing, integrates computer vision with human cognition. Its progress is limited by the lack of standardized datasets and robust models, as the temporal dynamics of video and multimodal fusion challenges hinder direct application of image-based methods. This study introduces VADB, the largest video aesthetic database with 10,490 diverse videos annotated by 37 professionals across multiple aesthetic dimensions, including overall and attribute-specific aesthetic scores, rich language comments and objective tags. We propose VADB-Net, a dual-modal pre-training framework with a two-stage training strategy, which outperforms existing video quality assessment models in scoring tasks and supports downstream video aesthetic assessment tasks. The dataset and source code are available at https://github.com/BestiVictory/VADB.
Qianqian Qiao, Yihang Bo, Bao Peng, Heng Huang 0002, Longteng Jiang, Huaye Wang, Jingdong Chen, Xin Jin 0015
NeurIPS10
2025 MusicAOG: An Energy-Based Model for Learning and Sampling a Hierarchical Representation of Symbolic Music
abstract
In addressing the challenge of interpretability and generalizability of artificial music intelligence, this article introduces a novel symbolic representation that amalgamates both explicit and implicit musical information across diverse traditions and granularities. Utilizing a hierarchical and-or graph representation, the model employs nodes and edges to encapsulate a broad spectrum of musical elements, including structures, textures, rhythms, and harmonies. This hierarchical approach expands the representability across various scales of music. This representation serves as the foundation for an energy-based model, uniquely tailored to learn musical concepts through a flexible algorithm framework relying on the minimax entropy principle. Utilizing an adapted Metropolis–Hastings sampling technique, the model enables fine-grained control over music generation. Through a comprehensive empirical evaluation, this novel approach demonstrates significant improvements in interpretability and controllability compared to existing methodologies. This study marks a substantial contribution to the fields of music analysis, composition, and computational musicology.
Yikai Qian, Tianle Wang 0007, Jishang Chen, Peiyang Yu, Duo Xu 0004, Xin Jin 0015, Feng Yu 0032, Song-Chun Zhu
IEEE Trans. Comput. Soc. Syst.6
2024 Paintings and Drawings Aesthetics Assessment with Rich Attributes for Various Artistic Categories
Xin Jin 0015, Qianqian Qiao, Huaye Wang, Heng Huang 0002, Guangdong Li
IJCAI1
2024 Reproducibility Companion Paper: Aesthetics-Driven Virtual Time-Lapse Photography Generation
abstract
In this paper, we replicate the experimental results from our previous work titled "Aesthetics-Driven Virtual Time-Lapse Photography Generation", which was presented at ACM Multimedia 2023. Our primary objective is to confirm the validity of our earlier findings and to provide a more comprehensive understanding of our software framework. We provide the necessary artifacts to reproduce the results from our prior research. This paper details the technical aspects of our package, including dataset preparation, source code structure, and the experimental environment. By utilizing these artifacts, we demonstrate the reproducibility of our results. We encourage others to use our software framework for purposes beyond reproducibility.
Xin Jin 0015, Longteng Jiang, Yihao Zhang 0010, Xiaobo Gao, Boyan Dong
ACM Multimedia1
2024 Unlimited Vision: Professional Composition by Yourself
abstract
This paper introduces a novel method for enhancing image composition guidance in photography. It utilizes advanced composition rules to guide a Real-Time Detection Transformer (RT-DETR) model in predicting aesthetically pleasing compositions for photographs. Unlike traditional methods constrained by original image boundaries, our approach allows the predicted framing to extend beyond these limits, offering dynamic, real-time guidance for image composition in photography. The system integrates multi-label composition classification and compositional element annotation, using YOLOv8 for key object detection and an enhanced Deep Hough Transform for compositional lines to guide photographers. It provides photographers with real-time guidance for optimal camera adjustments, transforming traditional post-processing tasks into an intuitive, interactive process. This method significantly enhances photographers' flexibility and effectiveness in capturing visually superior photographs.
Xin Jin 0015, Liaoruxing Zhang, Longteng Jiang
ACM Multimedia1
2024 APDDv2: Aesthetics of Paintings and Drawings Dataset with Artist Labeled Scores and Comments
abstract
Datasets play a pivotal role in training visual models, facilitating the development of abstract understandings of visual features through diverse image samples and multidimensional attributes. However, in the realm of aesthetic evaluation of artistic images, datasets remain relatively scarce. Existing painting datasets are often characterized by limited scoring dimensions and insufficient annotations, thereby constraining the advancement and application of automatic aesthetic evaluation methods in the domain of painting.To bridge this gap, we introduce the Aesthetics Paintings and Drawings Dataset (APDD), the first comprehensive collection of paintings encompassing 24 distinct artistic categories and 10 aesthetic attributes. Building upon the initial release of APDDv1, our ongoing research has identified opportunities for enhancement in data scale and annotation precision. Consequently, APDDv2 boasts an expanded image corpus and improved annotation quality, featuring detailed language comments to better cater to the needs of both researchers and practitioners seeking high-quality painting datasets.Furthermore, we present an updated version of the Art Assessment Network for Specific Painting Styles, denoted as ArtCLIP. Experimental validation demonstrates the superior performance of this revised model in the realm of aesthetic evaluation, surpassing its predecessor in accuracy and efficacy.The dataset and model are available at https://github.com/BestiVictory/APDDv2.git.
Xin Jin 0015, Qianqian Qiao, Huaye Wang, Heng Huang 0002
NeurIPS1
2024 GPU Accelerated Full Homomorphic Encryption Cryptosystem, Library, and Applications for IoT Systems
abstract
Deep learning, such as convolutional neural networks (CNNs), has been utilized in a number of cloud-based Internet of Things (IoT) applications. Security and privacy are two key considerations in any commercial deployment. Fully homomorphic encryption (FHE) is a popular privacy protection approach, and there have been attempts to integrate FHE with CNNs. However, a simple integration may lead to inefficiency in single-user services and fail to support many of the requirements in real-time applications. In this article, we propose a novel confused modulo projection-based FHE algorithm (CMP-FHE) that is designed to support floating-point operations. Then, we developed a parallelized runtime library based on CMP-FHE and compared it with the widely employed FHE library. Our results show that our library achieves faster speeds. Furthermore, we compared it with the state-of-the-art confused modulo projection-based library and the results demonstrated a speed improvement of 841.67 to 3056.25 times faster. Additionally, we construct a real-time homomorphic CNN (RT-HCNN) under the ciphertext-based framework using CMP-FHE, as well as using graphics processing units (GPUs) to facilitate acceleration. To demonstrate utility, we evaluate the proposed approach on the MNIST data set. Findings demonstrate that our proposed approach achieves a high accuracy rate of 99.13%. Using GPUs acceleration for ciphertext prediction results in us achieving a single prediction time of 79.5 ms. This represents the first homomorphic CNN capable of supporting real-time application and is approximately 58 times faster than Microsoft’s Lola scheme.
Xiaodong Li 0013, Hehe Gao, Shuya Yang, Xin Jin 0015, Kim-Kwang Raymond Choo
IEEE Internet Things J.5
2023 An Order-Complexity Model for Aesthetic Quality Assessment of Symbolic Homophony Music Scores
abstract
Computational aesthetics evaluation has made great achievements in the field of visual arts, but the research work on music still needs to be explored. Although the existing work of music generation is very substantial, the quality of music score generated by AI is relatively poor compared with that created by human composers. The music scores created by AI are usually monotonous and devoid of emotion. Based on Birkhoff’s aesthetic measure, this paper proposes an objective quantitative evaluation method for homophony music score aesthetic quality assessment. The main contributions of our work are as follows: first, we put forward a homophony music score aesthetic model to objectively evaluate the quality of music score as a baseline model; second, we put forward eight basic music features and four music aesthetic features.
Xin Jin 0015, Wu Zhou 0006, Jinyu Wang 0007, Duo Xu 0004, Yiqing Rong, Shuai Cui
ICME1
2023 An Order-Complexity Aesthetic Assessment Model for Aesthetic-aware Music Recommendation
abstract
Computational aesthetic evaluation has made remarkable contribution to visual art works, but its application to music is still rare. Currently, subjective evaluation is still the most effective form of evaluating artistic works. However, subjective evaluation of artistic works will consume a lot of human and material resources. The popular AI generated content (AIGC) tasks nowadays have flooded all industries, and music is no exception. While compared to music produced by humans, AI generated music still sounds mechanical, monotonous, and lacks aesthetic appeal. Due to the lack of music datasets with rating annotations, we have to choose traditional aesthetic equations to objectively measure the beauty of music. In order to improve the quality of AI music generation and further guide computer music production, synthesis, recommendation and other tasks, we use Birkhoff's aesthetic measure to design a aesthetic model, objectively measuring the aesthetic beauty of music, and form a recommendation list according to the aesthetic feeling of music. Experiments show that our objective aesthetic model and recommendation method are effective.
Xin Jin 0015, Wu Zhou 0006, Jinyu Wang 0007, Duo Xu 0004, Yongsen Zheng
ACM Multimedia1
2023 Aesthetics-Driven Virtual Time-Lapse Photography Generation
abstract
Time-lapse videos can visualize the temporal change of dynamic scenes and present wonderful sights with drastic variance in color appearance and rapid movement that interests people. We propose an aesthetics-driven virtual time-lapse photography framework to explore the automatic generation of time-lapse videos in the virtual world, which has potential applications like artistic creation and entertainment in the virtual space. We first define shooting parameters to parameterize the time-lapse photography process and accordingly propose image, video, and time-lapse aesthetic assessments to optimize these parameters, enabling the process to be autonomous and adaptive. We also build an interactive interface to visualize the shooting process and help users conduct virtual time-lapse photography by personalizing shooting parameters according to their aesthetic preferences. Finally, we present a two-stream time-lapse aesthetic model and a time-lapse aesthetic dataset, which can evaluate the aesthetic quality of time-lapse videos. Experimental results demonstrate our method can automatically generate time-lapse videos comparable to those of professional photographers and is more efficient.
Hui Wei 0005, Xin Jin 0015, Yihao Zhang 0010, Boyan Dong, Longteng Jiang, Xiaohui Zhang 0017, Ruyang Li, Yaqian Zhao
ACM Multimedia3
2023 TBFormer: Two-Branch Transformer for Image Forgery Localization
abstract
Image forgery localization aims to identify forged regions by capturing subtle traces from high-quality discriminative features. In this paper, we propose a Transformer-style network with two feature extraction branches for image forgery localization, and it is named as Two-Branch Transformer (TBFormer). Firstly, two feature extraction branches are elaborately designed, taking advantage of the discriminative stacked Transformer layers, for both RGB and noise domain features. Secondly, an Attention-aware Hierarchical-feature Fusion Module (AHFM) is proposed to effectively fuse hierarchical features from two different domains. Although the two feature extraction branches have the same architecture, their features have significant differences since they are extracted from different domains. We adopt position attention to embed them into a unified feature domain for hierarchical feature investigation. Finally, a Transformer decoder is constructed for feature reconstruction to generate the predicted mask. Extensive experiments on publicly available datasets demonstrate the effectiveness of the proposed model.
Binbin Lv, Xin Jin 0015, Xiaokun Zhang 0002
IEEE Signal Process. Lett.3
2022 Reproducibility Companion Paper: Focusing on Persons: Colorizing Old Images Learning from Modern Historical Movies
abstract
In this paper we reproduce experimental results presented in our earlier work titled "Focusing on Persons: Colorizing Old Images Learning from Modern Historical Movies" that was presented in the course of the 29th ACM International Conference on Multimedia. The paper aims at verifying the soundness of our prior results and helping others understand our software framework. We present artifacts that help reproduce results that were included in our earlier work. Specifically, this paper contains the technical details of the package, including dataset preparation, source code structure and experimental environment. Using the artifacts we show that our results are reproducible. We invite everyone to use our software framework going beyond reproducibility efforts.
Xin Jin 0015, Dongqing Zou, Zhonglan Li, Heng Huang 0002, Vajira Thambawita
ACM Multimedia1
2022 Attribute Controllable Beautiful Caucasian Face Generation by Aesthetics Driven Reinforcement Learning
abstract
In recent years, image generation has made great strides in improving the quality of images, producing high-fidelity ones. Also, quite recently, there are architecture designs, which enable GAN to unsupervisedly learn the semantic attributes represented in different layers. However, there is still a lack of research on generating face images more consistent with human aesthetics. Based on EigenGAN [He et al., ICCV 2021], we build the techniques of reinforcement learning into the generator of EigenGAN. The agent tries to figure out how to alter the semantic attributes of the generated human faces towards more preferable ones. To accomplish this, we trained an aesthetics scoring model that can conduct facial beauty prediction. We also can utilize this scoring model to analyze the correlation between face attributes and aesthetics scores. Empirically, using off-the-shelf techniques from reinforcement learning would not work well. So instead, we present a new variant incorporating the ingredients emerging in the reinforcement learning communities in recent years. Compared to the original generated images, the adjusted ones show clear distinctions concerning various attributes. Experimental results using the MindSpore, show the effectiveness of the proposed method. Altered facial images are commonly more attractive, with significantly improved aesthetic levels.
Xin Jin 0015, Le Zhang 0015, Qiang Deng, Chaoen Xiao
ACM Multimedia1
2022 FOV Recognizer: Telling the Field of View of Movie Shots
Xin Jin 0015, Chenyu Fan, Yihang Bo, Xinzhe Pan, Zihan Jia, Ya Zhuo, Runqi Zhang, Shuai Cui
PRCV (3)1
2022 Pyramid Copy-move Forgery Detection Using Adversarial Optimized Self Deep Matching Network
abstract
In this paper, we propose a pyramid copy-move forgery detection framework based on multiscale Conditional Random Field (CRF) using an adversarial optimized self deep matching network. We present a patch-level adversarial optimization scheme to optimize a pre-trained self deep matching network, in order to approximate the distribution of ground truth masks more accurately. The optimized network is adopted as the basic self deep matching network in pyramid copy-move forgery detection. Coarse-to-fine images in pyramid form are put into the adversarial optimized self deep matching network to generate a set of score maps which contain rich multiscale information. Multiscale CRF is designed based on average score unary potentials and pairwise potentials with multiscale information kernels, to adequately explore multiscale information. Extensive experiments on publicly available datasets demonstrate the state-of-the-art performance of the proposed method.
Qiang Cai 0001, Xin Jin 0015
TrustCom4
2022 Pseudo-Labeling and Meta Reweighting Learning for Image Aesthetic Quality Assessment
abstract
In the tasks of image aesthetic quality assessment, it is difficult to reach both the high score area and low score area due to the normal distribution of aesthetic datasets. To reduce the error in labelling and solve the problem of normal data distribution, we propose a new aesthetic mixed dataset with classification and regression called AMD-CR, and we train a meta reweighting network to reweight the loss of training data differently. In addition, we provide a training strategy according to different stages based on pseudo labels, and then we use it for aesthetic training according to different stages in classification and regression tasks. In the construction of the network structure, we construct an aesthetic adaptive block (AAB) structure that can adapt to any size of the input images. Besides, we also use the efficient channel attention (ECA) to strengthen the feature extracting ability of each task. The experimental result shows that our method improves 0.1112 compared with the conventional method in SROCC. The method can also help to find best aesthetic path planning for unmanned aerial vehicles (UAV) and vehicles.
Xin Jin 0015, Hao Lou, Heng Huang 0002, Xinning Li, Xiaodong Li 0013, Shuai Cui, Xiaokun Zhang 0002, Xiqiao Li
IEEE Trans. Intell. Transp. Syst.1
2022 Aesthetic Attribute Assessment of Images Numerically on Mixed Multi-attribute Datasets
abstract
With the continuous development of social software and multimedia technology, images have become a kind of important carrier for spreading information and socializing. How to evaluate an image comprehensively has become the focus of recent researches. The traditional image aesthetic assessment methods often adopt single numerical overall assessment scores, which has certain subjectivity and can no longer meet the higher aesthetic requirements. In this article, we construct an new image attribute dataset called aesthetic mixed dataset with attributes (AMD-A) and design external attribute features for fusion. Besides, we propose an efficient method for image aesthetic attribute assessment on mixed multi-attribute dataset and construct a multitasking network architecture by using the EfficientNet-B0 as the backbone network. Our model can achieve aesthetic classification, overall scoring, and attribute scoring. In each sub-network, we improve the feature extraction through ECA channel attention module. As for the final overall scoring, we adopt the idea of the teacher-student network and use the classification sub-network to guide the aesthetic overall fine-grain regression. Experimental results, using the MindSpore, show that our proposed method can effectively improve the performance of the aesthetic overall and attribute assessment.
Xin Jin 0015, Xinning Li, Hao Lou, Chenyu Fan, Qiang Deng, Chaoen Xiao, Shuai Cui, Amit Kumar Singh 0001
ACM Trans. Multim. Comput. Commun. Appl.1
2022 Double-Blinded Finder: a two-side secure children face recognition system
Xin Jin 0015, Jicheng Lei, Shiming Ge, Chenggen Song, Chuanqiang Wu
Wirel. Networks1
2021 Focusing on Persons: Colorizing Old Images Learning from Modern Historical Movies
abstract
In industry, there exist plenty of scenarios where old gray photos need to be automatically colored, such as video sites and archives. In this paper, we present the HistoryNet focusing on historical person's diverse high fidelity clothing colorization based on fine grained semantic understanding and prior. Colorization of historical persons is realistic and practical, however, existing methods do not perform well in the regards. In this paper, a HistoryNet including three parts, namely, classification, fine grained semantic parsing and colorization, is proposed. Classification sub-module supplies classifying of images according to the eras, nationalities and garment types; Parsing sub-network supplies the semantic for person contours, clothing and background in the image to achieve more accurate colorization of clothes and persons and prevent color overflow. In the training process, we integrate classification and semantic parsing features into the coloring generation network to improve colorization. Through the design of classification and parsing subnetwork, the accuracy of image colorization can be improved and the boundary of each part of image can be more clearly. Moreover, we also propose a novel Modern Historical Movies Dataset (MHMD) containing 1,353,166 images and 42 labels of eras, nationalities, and garment types for automatic colorization from 147 historical movies or TV series made in modern time. Various quantitative and qualitative comparisons demonstrate that our method outperforms the state-of-the-art colorization methods, especially on military uniforms, which has correct colors according to the historical literatures.
Xin Jin 0015, Zhonglan Li, Dongqing Zou, Xiaodong Li 0013, Xingfan Zhu, Ziyin Zhou, Qilong Sun
ACM Multimedia1
2021 Aesthetic Evaluation and Guidance for Mobile Photography
abstract
Nowadays, almost everyone can shoot photos using smart phones. However, not everyone can take good photos. We propose to use computational aesthetics to automatically teach people without photography training to take excellent photos. We present Aesthetic Dashboard: a system of rich aesthetic evaluation and guidance for mobile photography. We take 2 most used types of photos: landscapes and portraits into consideration. When people take photos in the preview mode, for landscapes, we show the overall aesthetic score and scores of 3 basic attributes: light, composition and color usage. Meanwhile, the matching scores of the 3 basic attributes of current preview to typical templates are shown, which can help users to adjust 3 basic attributes accordingly. For portraits, besides the above basic attributes, the facial appearance, the guidance of face light, body pose and the garment color are also shown to the users. This is the first system that can teach mobile users to shoot good photos in the form of aesthetic dashboard, through which, users can adjust several aesthetic attributes to take good photos easily.
Hao Lou, Heng Huang 0002, Chaoen Xiao, Xin Jin 0015
ACM Multimedia4
2021 AR CAPTCHA: Recognizing robot by augmented reality
abstract
Summary We propose a novel AR CAPTCHA that first uses Augmented Reality to design CAPTCHA. Users should use their cameras on mobile devices to capture a marker in the 3D physical world or PC screens to find the appropriate angle to recognize each 3D character rendered on the marker in a 3D registration manner, which is very hard for machines or robots to do the same thing. Besides, we add many random 2D characters on the marker to make the state‐of‐the‐art scene text recognition methods fail to recognize the target 3D characters. To ensure availability, we design dual‐channel scenario setting (mobile phone and physical world) that enhances the security level of CAPTCHA and single‐channel scenario setting (mobile phone only) that the gyroscope could be used for our CAPTCHA. The experimental results reveal that the recognition accuracy of machines is nearly 0%, while a human could recognize AR CAPTCHA in only a few seconds.
Xin Jin 0015, Yuwei Duan, Xiaodong Li 0013
Concurr. Comput. Pract. Exp.1
2021 Confused-Modulo-Projection-Based Somewhat Homomorphic Encryption - Cryptosystem, Library, and Applications on Secure Smart Cities
abstract
With the development of cloud computing, the storage and processing of massive visual media data has gradually transferred to the cloud server. For example, if the intelligent video monitoring system cannot process a large amount of data locally, the data will be uploaded to the cloud. Therefore, how to process data in the cloud without exposing the original data has become an important research topic. We propose a single-server version of somewhat homomorphic encryption cryptosystem based on confused modulo projection theorem named CMP-SWHE, which allows the server to complete blind data processing withoutseeingthe effective information of user data. On the client side, the original data is encrypted by amplification, randomization, and setting confusing redundancy. Operating on the encrypted data on the server side is equivalent to operating on the original data. As an extension, we designed and implemented a blind computing scheme of accelerated version based on batch processing technology to improve efficiency. To make this algorithm easy to use, we also designed and implemented an efficient general blind computing library based on CMP-SWHE. We have applied this library to foreground extraction, optical flow tracking, and object detection with satisfactory results, which are helpful for building smart cities. We also discuss how to extend the algorithm to deep learning applications. Compared with other homomorphic encryption cryptosystems and libraries, the results show that our method has obvious advantages in computing efficiency. Although our algorithm has some tiny errors ($10^{-6}$) when the data is too large, it is very efficient and practical, especially suitable for blind image and video processing.
Xin Jin 0015, Xiaodong Li 0013, Beisheng Liu, Shujiang Xie, Amit Kumar Singh 0001, Yujie Li 0001
IEEE Internet Things J.1
2021 RSANet: Towards Real-Time Object Detection with Residual Semantic-Guided Attention Feature Pyramid Network
Quan Zhou 0004, Jie Wang 0024, Shenghua Li, Weihua Ou, Xin Jin 0015
Mob. Networks Appl.6
2021 SCCGAN: Style and Characters Inpainting Based on CGAN
Xiangshang Wang, Huimin Lu 0001, Shanxi Li, Xin Jin 0015
Mob. Networks Appl.7
2021 Random base image representation for efficient blind vision
Xin Jin 0015, Xiaodong Li 0013, Mingxue Yu
Multim. Tools Appl.1
2021 Fast Search of Lightweight Block Cipher Primitives via Swarm-like Metaheuristics for Cyber Security
abstract
With the construction and improvement of 5G infrastructure, more devices choose to access the Internet to achieve some functions. People are paying more attention to information security in the use of network devices. This makes lightweight block ciphers become a hotspot. A lightweight block cipher with superior performance can ensure the security of information while reducing the consumption of device resources. Traditional optimization tools, such as brute force or random search, are often used to solve the design of Symmetric-Key primitives. The metaheuristic algorithm was first used to solve the design of Symmetric-Key primitives of SKINNY. The genetic algorithm and the simulated annealing algorithm are used to increase the number of active S-boxes in SKINNY, thus improving the security of SKINNY. Based on this, to improve search efficiency and optimize search results, we design a novel metaheuristic algorithm, named particle swarm-like normal optimization algorithm (PSNO) to design the Symmetric-Key primitives of SKINNY. With our algorithm, one or better algorithm components can be obtained more quickly. The results in the experiments show that our search results are better than those of the genetic algorithm and the simulated annealing algorithm. The search efficiency is significantly improved. The algorithm we proposed can be generalized to the design of Symmetric-Key primitives of other lightweight block ciphers with clear evaluation indicators, where the corresponding indicators can be used as the objective functions.
Xin Jin 0015, Yuwei Duan, Mengdong Li, Ming Mao, Amit Kumar Singh 0001, Yujie Li 0001
ACM Trans. Internet Techn.1
2020 Look One and More: Distilling Hybrid Order Relational Knowledge for Cross-Resolution Image Recognition
abstract
In spite of great success in many image recognition tasks achieved by recent deep models, directly applying them to recognize low-resolution images may suffer from low accuracy due to the missing of informative details during resolution degradation. However, these images are still recognizable for subjects who are familiar with the corresponding high-resolution ones. Inspired by that, we propose a teacher-student learning approach to facilitate low-resolution image recognition via hybrid order relational knowledge distillation. The approach refers to three streams: the teacher stream is pretrained to recognize high-resolution images in high accuracy, the student stream is learned to identify low-resolution images by mimicking the teacher's behaviors, and the extra assistant stream is introduced as bridge to help knowledge transfer across the teacher to the student. To extract sufficient knowledge for reducing the loss in accuracy, the learning of student is supervised with multiple losses, which preserves the similarities in various order relational structures. In this way, the capability of recovering missing details of familiar low-resolution images can be effectively enhanced, leading to a better knowledge transfer. Extensive experiments on metric learning, low-resolution image classification and low-resolution face recognition tasks show the effectiveness of our approach, while taking reduced models.
Shiming Ge, Kangkai Zhang, Yingying Hua, Shengwei Zhao, Xin Jin 0015
AAAI6
2020 Efficient blind face recognition in the cloud
Xin Jin 0015, Xiaodong Li 0013, Chuanqiang Wu
Multim. Tools Appl.1
2020 Complex object relighting via split-then-composition by semantics and materials
Xin Jin 0015, Xiaodong Li 0013, Xiaokun Zhang 0002
Multim. Tools Appl.1
2020 IDEA: A new dataset for image aesthetic scoring
Xin Jin 0015, Geng Zhao 0001, Xinghui Zhou, Xiaokun Zhang 0002, Xiaodong Li 0013
Multim. Tools Appl.1
2020 Blind background extraction from videos in the cloud
Xin Jin 0015, Xiaodong Li 0013
Multim. Tools Appl.1
2020 Video encryption based on hyperchaotic system
Xiaodong Li 0013, Xin Jin 0015
Multim. Tools Appl.4
2020 Painting completion with generative translation models
Shanxi Li, Yuqian Shi, Xin Jin 0015
Multim. Tools Appl.5
2020 Deep Multimodality Learning for UAV Video Aesthetic Quality Assessment
abstract
Despite the growing number of unmanned aerial vehicles (UAVs) and aerial videos, there is a paucity of studies focusing on the aesthetics of aerial videos that can provide valuable information for improving the aesthetic quality of aerial photography. In this article, we present a method of deep multimodality learning for UAV video aesthetic quality assessment. More specifically, a multistream framework is designed to exploit aesthetic attributes from multiple modalities, including spatial appearance, drone camera motion, and scene structure. A novel specially designed motion stream network is proposed for this new multistream framework. We construct a dataset with 6,000 UAV video shots captured by drone cameras. Our model can judge whether a UAV video was shot by professional photographers or amateurs together with the scene type classification. The experimental results reveal that our method outperforms the video classification methods and traditional SVM-based methods for video aesthetics. In addition, we present three application examples of UAV video grading, professional segment detection and aesthetic-based UAV path planning using the proposed method.
Qi Kuang, Xin Jin 0015, Qinping Zhao
IEEE Trans. Multim.2
2019 Defending Against Adversarial Examples via Soft Decision Trees Embedding
abstract
Convolutional neural networks (CNNs) have shown vulnerable to adversarial examples which contain imperceptible perturbations. In this paper, we propose an approach to defend against adversarial examples with soft decision trees embedding. Firstly, we extract the semantic features of adversarial examples with a feature extraction network. Then, a specific soft decision tree is trained and embedded to select the key semantic features for each feature map from convolutional layers and the selected features are fed to a light-weight classification network. To this end, we use the probability distributions of each tree node to quantify the semantic features. In this way, some small perturbations can be effectively removed and the selected features are more discriminative in identifying adversarial examples. Moreover, the influence of adversarial perturbations on classification can be reduced by migrating the interpretability of soft decision trees into the black-box neural networks. We conduct experiments to defend the state-of-the-art adversarial attacks. The experimental results demonstrate that our proposed approach can effectively defend against these attacks and improve the robustness of deep neural networks.
Yingying Hua, Shiming Ge, Xindi Gao, Xin Jin 0015, Dan Zeng 0001
ACM Multimedia4
2019 Aesthetic Attributes Assessment of Images
abstract
Image aesthetic quality assessment has been a relatively hot topic during the last decade. Most recently, comments type assessment (aesthetic captions) has been proposed to describe the general aesthetic impression of an image using text. In this paper, we propose Aesthetic Attributes Assessment of Images, which means the aesthetic attributes captioning. This is a new formula of image aesthetic assessment, which predicts aesthetic attributes captions together with the aesthetic score of each attribute. We introduce a new dataset named DPC-Captions which contains comments of up to 5 aesthetic attributes of one image through knowledge transfer from a full-annotated small-scale dataset. Then, we propose Aesthetic Multi-Attribute Network (AMAN), which is trained on a mixture of fully-annotated small-scale PCCD dataset and weakly-annotated large-scale DPC-Captions dataset. Our AMAN makes full use of transfer learning and attention model in a single framework. The experimental results on our DPC-Captions and PCCD dataset reveal that our method can predict captions of 5 aesthetic attributes together with numerical score assessment of each attribute. We use the evaluation criteria used in image captions to prove that our specially designed AMAN model outperforms traditional CNN-LSTM model and modern SCA-CNN model of image captions.
Xin Jin 0015, Geng Zhao 0001, Xiaodong Li 0013, Xiaokun Zhang 0002, Shiming Ge, Dongqing Zou, Xinghui Zhou
ACM Multimedia1
2019 ESNet: An Efficient Symmetric Network for Real-Time Semantic Segmentation
Yu Wang 0109, Quan Zhou 0004, Jian Xiong 0005, Xiaofu Wu, Xin Jin 0015
PRCV (2)5
2019 An open-source project for real-time image semantic segmentation
Quan Zhou 0004, Yu Wang 0109, Xin Jin 0015, Longin Jan Latecki
Sci. China Inf. Sci.4
2019 Two open-source projects for image aesthetic quality assessment
Xin Jin 0015, Geng Zhao 0001, Xinghui Zhou
Sci. China Inf. Sci.2
2019 ILGNet: inception modules with connected local and global features for efficient image aesthetic quality classification using domain adaptation
abstract
In this study, the authors address a challenging problem of aesthetic image classification, which is to label an input image as high‐ or low‐aesthetic quality. We take both the local and global features of images into consideration. A novel deep convolutional neural network named ILGNet is proposed, which combines both the inception modules and a connected layer of both local and global features. The ILGnet is based on GoogLeNet. Thus, it is easy to use a pre‐trained GoogLeNet for large‐scale image classification problem and fine tune their connected layers on a large‐scale database of aesthetic‐related images: AVA, i.e. domain adaptation . The experiments reveal that their model achieves the state of the arts in AVA database. Both the training and testing speeds of their model are higher than those of the original GoogLeNet.
Xin Jin 0015, Xiaodong Li 0013, Xiaokun Zhang 0002, Jingying Chi, Siwei Peng, Shiming Ge, Geng Zhao 0001
IET Comput. Vis.1
2019 Secure face retrieval for group mobile users
Xin Jin 0015, Yujie Li 0001, Shiming Ge, Chenggen Song, Xinghui Zhou
Soft Comput.1
2018 Predicting Aesthetic Score Distribution Through Cumulative Jensen-Shannon Divergence
abstract
Aesthetic quality prediction is a challenging task in the computer vision community because of the complex interplay with semantic contents and photographic technologies. Recent studies on the powerful deep learning based aesthetic quality assessment usually use a binary high-low label or a numerical score to represent the aesthetic quality. However the scalar representation cannot describe well the underlying varieties of the human perception of aesthetics. In this work, we propose to predict the aesthetic score distribution (i.e., a score distribution vector of the ordinal basic human ratings) using Deep Convolutional Neural Network (DCNN). Conventional DCNNs which aim to minimize the difference between the predicted scalar numbers or vectors and the ground truth cannot be directly used for the ordinal basic rating distribution. Thus, a novel CNN based on the Cumulative distribution with Jensen-Shannon divergence (CJS-CNN) is presented to predict the aesthetic score distribution of human ratings, with a new reliability-sensitive learning method based on the kurtosis of the score distribution, which eliminates the requirement of the original full data of human ratings (without normalization). Experimental results on large scale aesthetic dataset demonstrate the effectiveness of our introduced CJS-CNN in this task.
Xin Jin 0015, Xiaodong Li 0013, Siwei Peng, Jingying Chi, Shiming Ge, Chenggen Song, Geng Zhao 0001
AAAI1
2018 Predicting Aesthetic Radar Map Using a Hierarchical Multi-task Network
Xin Jin 0015, Xinghui Zhou, Geng Zhao 0001, Xiaokun Zhang 0002, Xiaodong Li 0013, Shiming Ge
PRCV (2)1
2018 Image editing by object-aware optimal boundary searching and mixed-domain composition
abstract
When combining very different images which often contain complex objects and backgrounds, producing consistent compositions is a challenging problem requiring seamless image editing. In this paper, we propose a general approach, called object-aware image editing , to obtain consistency in structure, color, and texture in a unified way. Our approach improves upon previous gradient-domain composition in three ways. Firstly, we introduce an iterative optimization algorithm to minimize mismatches on the boundaries when the target region contains multiple objects of interest. Secondly, we propose a mixed-domain consistency metric for measuring gradients and colors, and formulate composition as a unified minimization problem that can be solved with a sparse linear system. In particular, we encode texture consistency using a patch-based approach without searching and matching. Thirdly, we adopt an object-aware approach to separately manipulate the guidance gradient fields for objects of interest and backgrounds of interest, which facilitates a variety of seamless image editing applications. Our unified method outperforms previous state-of-the-art methods in preserving global texture consistency in addition to local structure continuity.
Shiming Ge, Xin Jin 0015, Qiting Ye, Zhao Luo, Qiang Li 0007
Comput. Vis. Media2
2018 Color image encryption in non-RGB color spaces
Xin Jin 0015, Sui Yin, Ningning Liu, Xiaodong Li 0013, Geng Zhao 0001, Shiming Ge
Multim. Tools Appl.1
2017 Compressing deep neural networks for efficient visual inference
abstract
The deployments of deep neural network models on mobile or embedded devices have been challenged due to two main reasons: 1) the large model size for storage, and 2) the large memory bandwidth for inference. To address these issues, this paper develops a deep neural network compression framework to reduce the resource usage for efficient visual inference. By reviewing the trained deep model, we propose a hybrid model compression algorithm via four major modules. Approximation module reduces the number of weights in each fully connected layer with low rank approximation. Then, quantization module analyzes weight distribution in each layer and represents them with low precision fixed point, which reduces the bits for storing each weight. After that, pruning module suppresses small weights to further reduce the number of parameters. Finally, coding module joint optimizes the representation and encoding of the sparse structure of the pruned weights with relative index by Huffman coding. Beyond the compression of model size, we propose an adaptive fixed point memory allocation algorithm to reduce memory footprint in inference. The proposed framework, along with the model compression and memory allocation algorithms, can provide 20-30x compression rate with negligible accuracy loss. We conduct an evaluation on two representative models, AlexNet and VGG-16, for object recognition and face verification tasks, which demonstrate the effectiveness of our proposed compression framework.
Shiming Ge, Zhao Luo, Shengwei Zhao, Xin Jin 0015, Xiaoyu Zhang 0002
ICME4
2017 Efficient privacy preserving Viola-Jones type object detection via random base image representation
abstract
A cloud server spent a lot of time, energy and money to train a Viola-Jones type object detector [1] with high accuracy. Clients can upload their photos to the cloud server to find objects. However, the client does not want the leakage of the content of his/her photos. In the meanwhile, the cloud server is also reluctant to leak any parameters of the trained object detectors. 10 years ago, Avidan & Butman introduced Blind Vision, which is a method for securely evaluating a ViolaJones type object detector. Blind Vision uses standard cryptographic tools and is painfully slow to compute, taking a couple of hours to scan a single image. The purpose of this work is to explore an efficient method that can speed up the process. We propose the Random Base Image (RBI) Representation. The original image is divided into random base images. Only the base images are submitted randomly to the cloud server. Thus, the content of the image can not be leaked. In the meanwhile, a random vector and the secure Millionaire protocol are leveraged to protect the parameters of the trained object detector. The RBI makes the integral-image enable again for the great acceleration. The experimental results reveal that our method can retain the detection accuracy of that of the plain vision algorithm and is significantly faster than the traditional blind vision, with only a very low probability of the information leakage theoretically.
Xin Jin 0015, Xiaodong Li 0013, Chenggen Song, Shiming Ge, Geng Zhao 0001, Yingya Chen
ICME1
2017 3D textured model encryption via 3D Lu chaotic mapping
Xin Jin 0015, Shuyun Zhu, Chaoen Xiao, Xiaodong Li 0013, Geng Zhao 0001, Shiming Ge
Sci. China Inf. Sci.1
2016 Private Video Foreground Extraction Through Chaotic Mapping Based Encryption in the Cloud
Xin Jin 0015, Kui Guo, Chenggen Song, Xiaodong Li 0013, Geng Zhao 0001, Yingya Chen, Huaichao Wang
MMM (1)1
2016 Chaos-based image encryption scheme combining DNA coding and entropy
Ping Zhen, Geng Zhao 0001, Lequan Min, Xin Jin 0015
Multim. Tools Appl.4
2015 Context-aware local abnormality detection in crowded scene
Xiaobin Zhu 0001, Xin Jin 0015, Xiaoyu Zhang 0002, Fugang He, Lei Wang 0101
Sci. China Inf. Sci.2
2015 Learning Templates for Artistic Portrait Lighting Analysis
abstract
Lighting is a key factor in creating impressive artistic portraits. In this paper, we propose to analyze portrait lighting by learning templates of lighting styles. Inspired by the experience of artists, we first define several novel features that describe the local contrasts in various face regions. The most informative features are then selected with a stepwise feature pursuit algorithm to derive the templates of various lighting styles. After that, the matching scores that measure the similarity between a testing portrait and those templates are calculated for lighting style classification. Furthermore, we train a regression model by the subjective scores and the feature responses of a template to predict the score of a portrait lighting quality. Based on the templates, a novel face illumination descriptor is defined to measure the difference between two portrait lightings. Experimental results show that the learned templates can well describe the lighting styles, whereas the proposed approach can assess the lighting quality of artistic portraits as human being does.
Xiaowu Chen 0001, Xin Jin 0015, Qinping Zhao
IEEE Trans. Image Process.2
2014 Lighting virtual objects in a single image via coarse scene understanding
Xiaowu Chen 0001, Xin Jin 0015
Sci. China Inf. Sci.2
2014 Geodesic Propagation for Semantic Labeling
abstract
This paper presents a semantic labeling framework with geodesic propagation (GP). Under the same framework, three algorithms are proposed, including GP, supervised GP (SGP) for image, and hybrid GP (HGP) for video. In these algorithms, we resort to the recognition proposal map and select confident pixels with maximum probability as the initial propagation seeds. From these seeds, the GP algorithm iteratively updates the weights of geodesic distances until the semantic labels are propagated to all pixels. On the contrary, the SGP algorithm further exploits the contextual information to guide the direction of propagation, leading to better performance but higher computational complexity than the GP. For video labeling, we further propose the HGP algorithm, in which the geodesic metric is used in both spatial and temporal spaces. Experiments on four public data sets show that our algorithms outperform several state-of-the-art methods. With the GP framework, convincing results for both image and video semantic labeling can be obtained.
Xiaowu Chen 0001, Yafei Song 0002, Yu Zhang 0035, Xin Jin 0015, Qinping Zhao
IEEE Trans. Image Process.5
2013 Face Illumination Manipulation Using a Single Reference Image by Adaptive Layer Decomposition
abstract
This paper proposes a novel image-based framework to manipulate the illumination of human face through adaptive layer decomposition. According to our framework, only a single reference image, without any knowledge of the 3D geometry or material information of the input face, is needed. To transfer the illumination effects of a reference face image to a normal lighting face, we first decompose the lightness layers of the reference and the input images into large-scale and detail layers through weighted least squares (WLS) filter with adaptive smoothing parameters according to the gradient values of the face images. The large-scale layer of the reference image is filtered with the guidance of the input image by guided filter with adaptive smoothing parameters according to the face structures. The relit result is obtained by replacing the largescale layer of the input image with that of the reference image. To normalize the illumination effects of a non-normal lighting face (i.e., face delighting), we introduce similar reflectance prior to the layer decomposition stage by WLS filter, which make the normalized result less affected by the high contrast light and shadow effects of the input face. Through these two procedures, we can change the illumination effects of a non-normal lighting face by first normalizing the illumination and then transferring the illumination of another reference face to it. We acquire convincing relit results of both face relighting and delighting on numerous input and reference face images with various illumination effects and genders. Comparisons with previous papers show that our framework is less affected by geometry differences and can preserve better the identification structure and skin color of the input face.
Xiaowu Chen 0001, Xin Jin 0015, Qinping Zhao
IEEE Trans. Image Process.3
2012 Supervised Geodesic Propagation for Semantic Label Transfer
Xiaowu Chen 0001, Yafei Song 0002, Xin Jin 0015, Qinping Zhao
ECCV (3)4
2012 Artistic Illumination Transfer for Portraits
abstract
Abstract Relighting a portrait in a single image is still a challenging problem, particularly when only a single artistic reference photograph or painting is provided. In this paper, we propose an artistic illumination transfer system for portraits based on a database of portrait images (photographs and paintings) associated with hand‐drawn illumination templates (276) by artists. Users can select a reference portrait image in the database, and the corresponding illumination template is transferred to an input portrait using image warping. Users can also provide reference portrait images those are not in the database. Based on the Face Illumination Descriptor (FID), the system selects from the database the reference image with the closest illumination to that of the user‐provided reference image and adjusts the corresponding illumination template to match the contrast of the user‐provided reference image. Experiments on not only paintings but also photographs, paper‐cuts and sketches demonstrate that convincing illumination transferred results can be rendered by our system.
Xiaowu Chen 0001, Xin Jin 0015, Qinping Zhao
Comput. Graph. Forum2
2011 Single Image Based Illumination Estimation for Lighting Virtual Object in Real Scene
abstract
Rendering virtual objects into real scenes with real illumination can greatly increase the realism of virtual objects and the consistency between the virtual and the real. The main challenge lies in illumination estimation from a single image. This article proposes a novel method of single image based illumination estimation for lighting virtual object in real scene. Only a single image, without any knowledge of the 3D geometry or reflectance, is needed, which greatly increases the applicability of the method. We first estimate coarse scene geometry and intrinsic components including shading image and reflectance image. Then the sparse radiance map of the scene is inferred based on the scene geometry and intrinsic components. Finally, the virtual objects are illuminated by the estimated sparse radiance map. Some experimental results show that this method can convincingly light virtual objects into a single real image, without any pre-recorded 3D geometry and reflectance, illumination acquisition equipments or imaging information of the image.
Xiaowu Chen 0001, Xin Jin 0015
CAD/Graphics3
2011 Face illumination transfer through edge-preserving filters
abstract
This article proposes a novel image-based method to transfer illumination from a reference face image to a target face image through edge-preserving filters. According to our method, only a single reference image, without any knowledge of the 3D geometry or material information of the target face, is needed. We first decompose the lightness layers of the reference and the target images into large-scale and detail layers through weighted least square (WLS) filter after face alignment. The large-scale layer of the reference image is filtered with the guidance of the target image. Adaptive parameter selection schemes for the edge-preserving filters is proposed in the above two filtering steps. The final relit result is obtained by replacing the large-scale layer of the target image with that of the reference image. We acquire convincing relit result on numerous target and reference face images with different lighting effects and genders. Comparisons with previous work show that our method is less affected by geometry differences and can preserve better the identification structure and skin color of the target face.
Xiaowu Chen 0001, Xin Jin 0015, Qinping Zhao
CVPR3
2010 Learning Artistic Lighting Template from Portrait Photographs
Xin Jin 0015, Mingtian Zhao, Xiaowu Chen 0001, Qinping Zhao, Song-Chun Zhu
ECCV (4)1