Bin Jiang 0006

dblp:18/4625-6 · DBLP profile ↗
← Back
42ranked-venue papers
13as first author
29since 2021 · last 2026
0000-0002-5840-9664ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 32 · 9 first-author · 26 since 2021Artificial intelligence and machine learning · 9 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CREAM: Collaborative Representation with Self-supervised Alignment for Multimedia Recommendation
abstract
Recent works in multimedia recommendation combine interaction data with rich multimedia content to generate personalized recommendations, which has attracted considerable attention. Despite their effectiveness, the existing methods still suffer from three limitations: (1) Incomplete user preference modeling due to the late fusion strategy which ignores modality interactivity and user-specific modality preference; (2) Insufficient multimodal item representation caused by conventional homogeneous graph structure that faces a message-passing bottleneck and struggles to capture two-hop node similarity; (3) Ineffective modality alignment based on entity-specific fine-grained mutual information may undermine modality specificity and semantic richness, while user-item behavior alignment neglects the usefulness from negative feedback. To this end, we propose a novel Collaborative Representation with SElf-supervised Alignment for Multimedia Recommendation (CREAM). Specifically, we design a collaborative attention mechanism for late fusion to model user preferences, capturing both cross-modal interactivity and specific modality preferences. Secondly, we propose dual graph-view modeling that combines a traditional pair-wise graph with a derived group-wise hypergraph, improving message-passing efficiency and capturing two-hop node similarity. Finally, we design a self-supervised alignment based on Cauchy-Schwarz divergence, achieving coarse-grained cross-modal alignment. We also propose a user-item behavior alignment that explicitly distinguishes users’ positive and negative preferences through contrastive learning. Extensive experiments on three public datasets demonstrate the effectiveness of CREAM. Our code is available at https://github.com/GarethJellingham/CREAM.
Junhao Gao, Chao Yang 0015, Bin Jiang 0006
ICMR3
2025 Dynamic Spectral Graph Anomaly Detection
abstract
Graph anomaly detection is crucial for identifying anomalous nodes within graphs and addressing applications like financial fraud detection and social spam detection. Recent spectral graph neural network methods advance graph anomaly detection by focusing on anomalies that notably affect the distribution of graph spectral energy. Such spectrum-based methods rely on two steps: graph wavelet extraction and feature fusion. However, both steps are hand-designed, capturing incomprehensive anomaly information of wavelet-specific features and resulting in their inconsistent feature fusion. To address these problems, we propose a dynamic spectral graph anomaly detection framework DSGAD to adaptively capture comprehensive anomaly information and perform consistent feature fusion. DSGAD introduces dynamic wavelets, consisting of trainable wavelets to adaptively learn anomalous patterns and capture wavelet-specific features with comprehensive anomaly information. Furthermore, the consistent fusion of wavelet-specific features achieves dynamic fusion by combining wavelet-specific feature extraction with energy difference and channel convolution fusion using location correlation. Experimental results on four datasets substantiate the efficacy of our DSGAD method, surpassing state-of-the-art methods in both homogeneous and heterogeneous graphs.
Jianbo Zheng, Chao Yang 0015, Tairui Zhang, Longbing Cao, Bin Jiang 0006, Xuhui Fan 0001, Xiao-Ming Wu 0002, Xianxun Zhu
AAAI5
2025 Subdomain Uncertainty Optimization for Cross-Speed Fault Diagnosis
abstract
Cross-speed bearing fault diagnosis based on unsupervised domain adaptation can handle data distribution differences across various operating speeds, supporting intelligent maintenance of equipment like wind turbines with variable operating speeds. Existing methods focus on aligning sample distributions between source and target domains through global or subdomain correlations. However, these methods overlook essential relationships, such as possible high sample similarity between target subdomains and discrepancies in decision boundaries between source and target domains, leading to sub-stantial class confusion issues. To address class confusion, this paper proposes a subdomain uncertainty optimization method by using these relationships. Class uncertainty is proposed to quantify the degree of classification ambiguity among target domain samples, facilitating the differentiation of high-similarity samples. Boundary optimization is introduced to refine decision boundaries learned from the source domain, alleviating the adverse effects of boundary discrepancies between domains. Additionally, the CL-CNN network is adopted and adjusted to collaborate with the class uncertainty term and boundary optimization term, thus achieving optimal cross-speed fault diagnosis. Extensive experiments conducted across 18 cross-speed tasks demonstrate the superiority of the proposed method, which achieves a stable average accuracy of 99.86%. All code will be released on https://github.com/IWantBe/SUO.
Jianbo Zheng, Lida Huang, Tairui Zhang, Bin Jiang 0006, Chao Yang 0015
ICASSP4
2025 Multi-Passage Retrieval-Augmented Multimodal Language Generation Model for Knowledge-Based Visual Question Answering
abstract
Knowledge-based Visual Question Answering is a challenging task that requires a VQA system to utilize external knowledge to answer questions about a given image. Most current retrieval-augmented methods suffer from two notable limitations. Firstly, these methods tend to convert images into plain text and subsequently discard the original image input, resulting in a loss of valuable visual information. Secondly, they often fail to utilize multiple retrieved passages effectively, typically depending on just one passage to generate answers or directly concatenating multiple passages, which leads to excessively long sequences of knowledge input. To overcome the limitations of previous methods, we introduce an innovative retriever-generator framework for knowledge-based VQA. This framework comprises a Multimodal Queries-oriented Knowledge Retriever (MQ-KR) and a Multi-Passage Retrieval-Augmented Generator (MP-RAG). We utilize the original image directly for both knowledge retrieval and answer generation. The generator is a multimodal language generation model capable of producing accurate answers by effectively aggregating evidence from multiple passages. Our proposed method achieves a VQA Score of 59.89% on the OK-VQA dataset, surpassing the SOTA retrieval-augmented method by a large margin (+4.37%).
Siyu Cheng, Chao Yang 0015, Bin Jiang 0006
ICME3
2025 Rethinking DeNoising Training for DETR-based Object Detection
abstract
DETR-based methods have shown impressive performance in object detection tasks. The original DETR employs one-to-one sparse supervision, resulting in poor supervision capability. Denoising training methods introduce additional noisy queries and train the decoder to reconstruct the ground truth boxes, thereby providing extra supervision and improving training efficiency. However, existing denoising training methods generate additional noisy queries based on random distributions, which are inconsistent with the noise distribution of the one-to-one queries, thus limiting the effectiveness of their supervision. To alleviate this, we propose a novel Sampling DeNoising method (SDN) that includes two key components: Noise Collection Builder and Sampling Query Generator. In the Noise Collection Builder, we continuously collect one-to-one query noise and iteratively update the initial query noise collection, ensuring that its noise distribution is similar to that of one-to-one queries, thereby alleviating the inconsistency. In the Sampling Query Generator, we randomly sample from the Noise Collection Builder and introduce random fluctuations to enhance the diversity of the noise samples. Meanwhile, we inject hard noise samples into the Sampling Query Generator to improve the decoder’s bounding box refinement ability. Experiments on two datasets demonstrate that the SDN method leads to a significant performance improvement. Our code is available at https://github.com/yimingfy/SDNDETR.
Bin Jiang 0006, Chao Yang 0015, Chenglong Lei, Ruiqi Hu
ICME1
2025 Advancing Multi-Hop Question Answering via Alternating Retrieval and Reasoning over Multi-view Knowledge Integration
abstract
Open-Domain Multi-Hop Question Answering aims to retrieve knowledge based on a given complex question and generate an answer through reasoning. Existing methods typically rely solely on structured or unstructured knowledge, neglecting the potential benefits of their integration. This limitation leads to issues such as incomplete retrieval for complex questions or inefficient utilization and aggregation of unstructured information from raw text. To address these issues, we leverage the complementary nature of structured and unstructured knowledge by performing parallel retrieval of both types and enhancing semantic information by integrating them to guide reasoning. Furthermore, to better assist the LLM in generating appropriate query, we adopt a data augmentation strategy to construct a triple-query dataset based on the 2WikiMultiHopQA dataset, which is used to fine-tuning the LLM. Extensive experiments conducted on three open-domain multi-hop QA datasets demonstrate that our approach consistently outperforms existing methods across various benchmarks. Our code is available at https://github.com/liumc14/multiview-knowledge-MHQA.git.
Mengchao Liu, Chao Yang 0015, Bin Jiang 0006, Chenglong Lei
ICME3
2025 Multi-Grained Alignment with Knowledge Distillation for Partially Relevant Video Retrieval
abstract
Partially Relevant Video Retrieval (PRVR) aims to accurately retrieve the most relevant video in response to a query from untrimmed videos. The analysis of video content can be done at three different granularities: frame-level, clip-level, and video-level. Previous methods have focused on one or two of these levels for alignment, limiting the exploration of the video semantics. Moreover, some methods use video-level alignment and apply a self-attention mechanism to generate video-level features, but this may not be ideal as the entire video may not be relevant to the query. We propose a M ulti- G rained A lignment framework with K nowledge D istillation (MGAKD), which purifies the cross-modal alignment knowledge from the Contrastive Language-Image Pre-training (CLIP) model and achieves multi-grained alignment. It extracts cross-modal alignment knowledge from CLIP and imparts this knowledge to the designed student model. For the student model, two branches are designed: an inheritance branch and an exploration branch. The inheritance branch absorbs the knowledge of cross-modal alignment from the CLIP. The exploration branch explores visual features at three granularities: frame-level, clip-level, and video-level. Specifically, we directly align the extracted frame features of the video with the query features to achieve frame-level alignment. In clip-level alignment, the use of Gaussian masks allows for the representation of the beginning, climax, and end of an event. By employing Gaussian masks, we are able to implicitly model clip-level features, resulting in clip features that contain a richer set of contextual information. To further enhance video-level feature exploration, we apply clip-guided attention to generate diverse video-level features based on different queries. This strategy effectively prevents irrelevant video moments from affecting the alignment of videos and queries. We conduct extensive experiments on two publicly available datasets, and the experimental results have surpassed those of the state-of-the-art method, showcasing the superior performance of the proposed method.
Chao Yang 0015, Bin Jiang 0006
ACM Trans. Multim. Comput. Commun. Appl.3
2024 ScribbleEditor: Guided Photo-realistic and Identity-preserving Image Editing with Interactive Scribble
abstract
Free-form scribbles are a convenient way for users to describe their intentions for image editing. However, prevalent issues arise in existing methods. First, the rich color of scribbles may contain information that deviates significantly from the original image distribution, causing common blending methods to yield unrealistic results due to inadequate interaction between the scribble and the image. Secondly, inputting extensive scribble areas may obscure crucial image components, resulting in identity loss post-editing. To address these issues, we propose a two-stage scribble editing method that achieves photo-realistic and identity-preserving editing. Our proposed method employs a two-channel multimodal interaction module, facilitating deep interaction between scribbles and images in channel and position domains, achieving semantic alignment and enhancement. In the second stage, we interact the editing offset with scribble area image content via the texture supplement module to achieve texture detail supplementation. Our method adeptly handles complex scribbles, even across extensive areas, demonstrating superior photo-realism and identity preservation in experimental results.
Haotian Hu, Bin Jiang 0006, Chao Yang 0015, Xinjiao Zhou, Xiaofei Huo
ICME2
2024 DE-GAN: Text-to-image synthesis with dual and efficient fusion model
Bin Jiang 0006, Weiyuan Zeng, Chao Yang 0015, Renjun Wang
Multim. Tools Appl.1
2024 A Dual Reinforcement Learning Framework for Weakly Supervised Phrase Grounding
abstract
Weakly-supervised phrase grounding aims to localize a specific region in an image that corresponds to the given textual phrase, where the mapping between noun phrases and image regions is not available in the training stage. Previous methods typically exploit an additional proxy task (e.g., phrase reconstruction or image-phrase alignment) to provide supervision for training, since the lack of region-level annotations in the weakly-supervised setting. However, there exists a significant gap in optimization objectives between the proxy tasks and the target grounding task, which may result in low-efficient optimization for the target model. Therefore, in this paper, we propose a novel dual reinforcement learning framework to directly optimize the phrase grounding model. Specifically, we consider the duality of phrase grounding and phrase generation tasks. These two tasks form a closed loop that can provide quality feedback signals to measure the performance of each other. In this way, we can measure the correctness of the localized regions and thus be able to optimize the grounding model directly. We design two reward functions to quantify the feedback signals and train the models via reinforcement learning. In addition, to relieve the training difficulty of our framework, we present a heuristic algorithm to generate pseudo region-phrase pairs to warm-start our models. We perform experiments on two popular phrase grounding datasets: ReferItGame and Flickr30K Entities, and the results demonstrate that our method outperforms the previous methods by a large margin.
Chao Yang 0015, Bin Jiang 0006, Junsong Yuan 0001
IEEE Trans. Multim.3
2024 FBGAN: multi-scale feature aggregation combined with boosting strategy for low-light image enhancement
Bin Jiang 0006, Renjun Wang, Jiawu Dai, Weiyuan Zeng
Vis. Comput.1
2023 Dynamic Dense-Sparse Representations for Real-Time Question Answering
abstract
Existing real-time question answering models have shown speed benefits on open-domain tasks. However, they possess limited phrase representations and are susceptible to information loss, which leads to low accuracy. In this paper, we propose modified contextualized sparse and dense encoders to improve the context embedding quality. For sparse encoding, we propose the JM-Sparse, which utilizes joint multi-head attention to focus on crucial information in different context locations and subsequently learn sparse vectors within an n-gram vocabulary space. Moreover, we leverage the similarity-enhanced dense(SE-Dense) vector to obtain rich contextual dense representations. To effectively combine dense and sparse features, we train the weights of dense and sparse vectors dynamically. Extensive experiments on standard benchmarks demonstrate the effectiveness of the proposed method compared with other query-agnostic models.
Minyu Sun, Bin Jiang 0006, Chao Yang 0015
ICME2
2023 DSP-Net: Diverse Structure Prior Network for Image Inpainting
abstract
The latest deep learning-based approaches have advanced diverse image inpainting task. However, existing methods limit to be aware of the structure information well, which constricts the performance of diverse generations. The intuitive representation of diversity generation is the structure change since the structure is the basis of the image. In this paper, we make full use of the structure information and propose the diverse structure prior network (DSP-Net). Specifically, there are two stages in DSP-Net to generate the diverse structure first and refine the texture next. For the diverse structure generation, we prompt the structural distribution to be similar to the Gaussian distribution to sample the diverse structural prior. With these priors, we refine the texture with a proposed propagation attention module. Meanwhile, we propose a structure diversity loss to enhance the ability of diverse structure generation further. Experiments on benchmark datasets including CelebA-HQ and Places2 indicate that DSP-Net is effective for diverse and visually realistic image restoration.
Chao Yang 0015, Bin Jiang 0006
ICME3
2023 DF-CLIP: Towards Disentangled and Fine-grained Image Editing from Text
abstract
Inspired by CLIP’s excellent image/text representation capability and StyleGAN’s disentangled latent space, text-guide image editing techniques make significant progress. However, as CLIP cannot perform local fine-grained image/text alignment, existing methods suffer from entanglement problems. Moreover, there lacks a deep interaction between textual tokens and visual features, which may lead to unfaithful editing results. In this paper, we propose DF-CLIP for Disentangled and Fine-grained text-guide image editing. Specifically, we design a novel dual-branch LatentMask module to generate more accurate editing directions in StyleGAN’s latent space, which can avoid changes in text-unrelated areas. Furthermore, we present a Multi-modal Interaction module to associate the text embedding with the image embedding and perform a deep interaction between them, which greatly enhance the guidance of text in image editing process and accelerate the training convergence. Extensive experiments show that our models perform more disentangled and natural editing results with a shorter training time.
Xinjiao Zhou, Bin Jiang 0006, Chao Yang 0015, Haotian Hu, Xiaofei Huo
ICME2
2023 Multi-scale dual-modal generative adversarial networks for text-to-image synthesis
Bin Jiang 0006, Chao Yang 0015, Fangqiang Xu
Multim. Tools Appl.1
2023 Image inpainting based on cross-hierarchy global and local aware network
Bin Jiang 0006, Chao Yang 0015
Multim. Tools Appl.1
2023 Effective low-light image enhancement with multiscale and context learning network
Bin Jiang 0006, Xiaochen Bo, Chao Yang 0015
Multim. Tools Appl.2
2022 Stacked Multi-Scale Attention Network for Image Colorization
abstract
Deep convolutional networks (CNNs) show their potential in image colorization for producing plausible results. Recently, the attention mechanism further boosts the performances of CNNs by constructing channel and spatial interactions. However, existing attention methods are performed in a single-scale manner, which is hard to capture multi-scale information interactions in a limited computational cost. This can limit the performance of the network to reconstruct color channels. In this paper, we propose a stacked multi-scale attention network (SMSANet) for image colorization. The core idea is to perform the attentions of a feature map in a multi-scale manner so that sufficient interactions are conducted to efficiently and adaptively capture multi-scale and long-range dependencies. By stacking the multi-scale attention layers in different convolutional layers, the SMSANet can focus on more discriminative features to reconstruct color channels. Moreover, a salient loss is designed to further refine the generated image at both pixel-level and object-level. Extensive experiments on the ImageNet dataset have demonstrated that SMSANet outperforms the state-of-the-art automatic colorization methods.
Bin Jiang 0006, Fangqiang Xu, Chao Yang 0015
ICASSP1
2022 Diverse Image Colorizaiton via Deep Color Prior and Triplet Latent Regularization
abstract
Automatic image colorization has made tremendous progress in recent years. However, previous methods rarely consider the inherent diversity of natural images, and existing diverse colorization approaches still suffer from the following two limitations. i) The diversity of generated images is limited. ii) Some inevitable artifacts substantially degrade the quality of colorization results. To cope with such problems, we propose a novel diverse image colorization network based on vanilla GAN, which extracts the deep color priors through an initial colorization network first, and then utilizes these priors to modulate the diverse generation to achieve high-fidelity colorization outputs. In addition, a novel triplet latent regularization is proposed to constrain the correspondence between latent codes and generated images, which efficiently alleviates mode collapse and encourages more diverse colorization results. Extensive experiments on three benchmark datasets demonstrate the superiority of our method over existing diverse image colorization models.
Bin Jiang 0006, Jiawu Dai, Weiyuan Zeng, Renjun Wang
ICME1
2022 LPGN: Language-Guided Proposal Generation Network for Referring Expression Comprehension
abstract
Referring expression comprehension(REC) aims to ground the referring expression in an image. Many mainstream frameworks are implemented in a two-stage process: the first-stage model generates candidate region proposals, and the second-stage model locates the referent among the proposals. Existing proposal generators create proposals entirely based on images, so there is a gap between generated proposals and referring expression, which leads to a bottleneck limiting the performance of the whole model. In order to break the bottle-neck, we introduce a novel language-guided proposal generation network: LPGN. Moreover, we introduce an uncertainty-aware proposal generation strategy to tackle the vagueness of language, so as to improve the training effectiveness. LPGN is convenient to integrate it into the existing two-stage REC models, because it is agnostic to the second stage model. Through extensive experiments on benchmark datasets, we demonstrate that our LPGN can generate proposals of higher quality than existing proposal generators and effectively alleviate the proposal bottleneck of the existing two-stage REC model.
Chao Yang 0015, Su Feng, Bin Jiang 0006
ICME4
2022 A Rolling Bearing Fault Diagnosis Method Using Multi-Sensor Data and Periodic Sampling
abstract
In recent years, bearing fault diagnosis based on deep learning has gradually become the mainstream. However, the existing studies still have some defects, such as unreasonable sampling and incomplete utilization of bearing data, limiting the further improvement of the performance of the fault diagnosis model. This paper proposes a fault diagnosis method using multi-sensor data and periodic sampling to solve the problems above. First, the vibration data of different bearing positions are fused into multi-channel fusion data to improve the defect of insufficient data utilization. Second, based on the sampling length and sampling stride, periodic sampling is carried out for the fusion data to solve the problem of unreasonable sampling. Third, the traditional convolutional neural network is adjusted to extract more detailed fault features and obtain the best recognition effect. Finally, the experimental results verify the effectiveness of the proposed method.
Jianbo Zheng, Chao Yang 0015, Fangrong Zheng, Bin Jiang 0006
ICME4
2022 Dual-Channel Localization Networks for Moment Retrieval with Natural Language
abstract
According to the given natural language query, moment retrieval aims to localize the most relevant moment in an untrimmed video. The existing solutions for this problem can be roughly divided into two categories based on whether candidate moments are generated: i) Moment-based approach: It pre-cuts the video into a set of candidate moments, performs multimodal fusion, and evaluates matching scores with the query. ii) Clip-based approach: It directly aligns video clips and query with predicting matching scores without generating candidate moments. Both frameworks have respective shortcomings: the moment-based models suffer from heavy computations, while the performance of clip-based models is familiarly inferior to moment-based counterparts. To this end, we design an intuitive and efficient Dual-Channel Localization Network (DCLN) to balance computational cost and retrieval performance. For reducing computational cost, we capture the temporal relations of only a few video moments with the same start or end boundary in the proposed dual-channel structure. The start or end channel map index represents the corresponding video moment's start or end time boundary. For improving model performance, we apply the proposed dual-channel localization network to efficiently encode the temporal relations on the dual-channel map and learn discriminative features to distinguish the matching degree between natural language query and video moments. The extensive experiments on two standard benchmarks demonstrate the effectiveness of our proposed method.
Bin Jiang 0006, Chao Yang 0015, Liang Pang 0001
ICMR2
2022 Video Moment Retrieval with Hierarchical Contrastive Learning
abstract
This paper explores the task of video moment retrieval (VMR), which aims to localize the temporal boundary of a specific moment from an untrimmed video by a sentence query. Previous methods either extract pre-defined candidate moment features and select the moment that best matches the query by ranking, or directly align the boundary clips of a target moment with the query and predict matching scores. Despite their effectiveness, these methods mostly focus only on aligning the query and single-level clip or moment features, and ignore the different granularities involved in the video itself, such as clip, moment, or video, resulting in insufficient cross-modal interaction. To this end, we propose a Temporal Localization Network with Hierarchical Contrastive Learning (HCLNet) for the VMR task. Specifically, we introduce a hierarchical contrastive learning method to better align the query and video by maximizing the mutual information (MI) between query and three different granularities of video to learn informative representations. Meanwhile, we introduce a self-supervised cycle-consistency loss to enforce the further semantic alignment between fine-grained video clips and query words. Experiments on three standard benchmarks show the effectiveness of our proposed method.
Chao Yang 0015, Bin Jiang 0006, Xiaokang Zhou
ACM Multimedia3
2022 Recognizing Very Small Face Images Using Convolution Neural Networks
abstract
Face recognition can be installed in a surveillance system so that it can be used for monitoring, tracking and access control. An excellent, intelligent surveillance system should be sensitive to the objects far away from the camera. Unfortunately, due to the long-distance, objects like human faces captured by the camera are too small to identify. As to enhance the subtle color differences in the face image, in this paper we first improve the resolution of the captured image using deep convolution neural networks (DCNNs). Then the efficient features are extracted and used to do classification. As for verifying the effectiveness of the proposed method, we used three databases including AR face database, Georgia Tech face database (GT) database, and Labelled Faces in the Wild (LFW) database, altogether, to conduct the training and testing. Compared to the existing approaches, experimental results show that the identification accuracy of the proposed method outperforms any existing approaches.
Shi-Jinn Horng, Julian Supardi, Wanlei Zhou 0001, Chin-Teng Lin, Bin Jiang 0006
IEEE Trans. Intell. Transp. Syst.5
2021 ParaPindel: a scalable coordinated parallel detection framework for human genome-wide structural variation
abstract
Detecting the existence of variation from massive human genome data, and determining the breakpoints and types of variations is essential for analyzing structural variation. Pindel, an accurate detection tool based on pattern growth approach, is commonly used for discovering indels and other types of structural variations from next-generation sequencing data. The explosive growth of sequencing data poses new challenges to current implementation of Pindel. Here, we proposed ParaPindel, an optimized version of Pindel that utilizes distributed multiprocess, for efficient large-scale detection of structural variation for human whole-genome sequencing data. ParaPindel divides the chromosome into multiple small windows with a fixed-length window size, so as to realize the parallel detection between different windows and different chromosomes. A crosswindow with a smaller length is introduced to cope with possible structural variations at the edge of the window. The experimental results show that ParaPindel shortens the time to detect an individual’s genome-wide structural variation from 186 hours to 33 minutes under the premise that the detection results are basically consistent. Employing 256 processes on 128 nodes on the TH-IHN supercomputer, the speedup ratio has reached 163 times, and the parallel efficiency has reached 69.74%.
Yaning Yang, Chao Yang 0015, Bin Jiang 0006, Shaoliang Peng
BIBM5
2021 DeFLOCNet: Deep Image Editing via Flexible Low-Level Controls
abstract
User-intended visual content fills the hole regions of an input image in the image editing scenario. The coarse low- level inputs, which typically consist of sparse sketch lines and color dots, convey user intentions for content creation (i.e., free-form editing). While existing methods combine an input image and these low-level controls for CNN inputs, the corresponding feature representations are not sufficient to convey user intentions, leading to unfaithfully generated content. In this paper, we propose DeFLOCNet which relies on a deep encoder-decoder CNN to retain the guidance of these controls in the deep feature representations. In each skip-connection layer, we design a structure generation block. Instead of attaching low-level controls to an input image, we inject these controls directly into each structure generation block for sketch line refinement and color propagation in the CNN feature space. We then concatenate the modulated features with the original decoder features for structure generation. Meanwhile, DeFLOCNet involves another decoder branch for texture generation and detail enhancement. Both structures and textures are rendered in the decoder, leading to user-intended editing results. Experiments on benchmarks demonstrate that DeFLOCNet effectively transforms different user intentions to create visually pleasing content.
Ziyu Wan, Yibing Song, Xintong Han, Jing Liao 0001, Bin Jiang 0006, Wei Liu 0005
CVPR7
2021 Learning Content and Context with Language Bias for Visual Question Answering
abstract
Visual Question Answering (VQA) is a challenging multi-modal task to answer questions about an image. Many works concentrate on how to reduce language bias which makes models answer questions ignoring visual content and language context. However, reducing language bias also weakens the ability of VQA models to learn context prior. To address this issue, we propose a novel learning strategy named CCB, which forces VQA models to answer questions relying on Content and Context with language Bias. Specifically, CCB establishes Content and Context branches on top of a base VQA model and forces them to focus on local key content and global effective context respectively. Moreover, a joint loss function is proposed to reduce the importance of biased samples and retain their beneficial influence on answering questions. Experiments show that CCB outperforms the state-of-the-art methods on VQA-CP v2.
Chao Yang 0015, Su Feng, Dongsheng Li 0002, Huawei Shen, Bin Jiang 0006
ICME6
2021 Multi-Modality Image Manipulation Detection
abstract
State-of-the-art multi-stream methods for image manipulation detection suffer from gaps between features. Moreover, the rapid development of GANs makes it an emerging method of image tampering. However, existing natural scene image tampered datasets are limited to manual tampering. In this paper, we address these two issues. Firstly, we propose a novel two-stream multi-modality image manipulation detection model (MM-net) that abandons the way of fusion. The main idea of the proposed model is to exploit one stream to guide the learning of the other stream through attention mechanism. Therefore, our approach enjoys the benefits of the multi-stream methods while avoiding the semantic gaps caused by bridging gaps between different streams. Secondly, we build the first tampered dataset of natural images based on GANs, pushing manipulation detection toward more realistic and challenging scenarios. Extensive experimental results demonstrate that our model outperforms state-of-the-art approaches on both manually and GANs tampered images.
Chao Yang 0015, Huawei Shen, Huizhou Li, Bin Jiang 0006
ICME5
2021 Accurate and Explainable Recommendation via Hierarchical Attention Network Oriented Towards Crowd Intelligence
Chao Yang 0015, Weixin Zhou, Bin Jiang 0006, Dongsheng Li 0002, Huawei Shen
Knowl. Based Syst.4
2020 METNet: A Mutual Enhanced Transformation Network for Aspect-based Sentiment Analysis
abstract
Aspect-based sentiment analysis (ABSA) aims to determine the sentiment polarity of each specific aspect in a given sentence.Existing researches have realized the importance of the aspect for the ABSA task and have derived many interactive learning methods that model context based on specific aspect.However, current interaction mechanisms are ill-equipped to learn complex sentences with multiple aspects, and these methods underestimate the representation learning of the aspect.In order to solve the two problems, we propose a mutual enhanced transformation network (METNet) for the ABSA task.First, the aspect enhancement module in METNet improves the representation learning of the aspect with contextual semantic features, which gives the aspect more abundant information.Second, METNet designs and implements a hierarchical structure, which enhances the representations of aspect and context iteratively.Experimental results on SemEval 2014 Datasets demonstrate the effectiveness of METNet, and we further prove that METNet is outstanding in multi-aspect scenarios.
Bin Jiang 0006, Wanyue Zhou, Chao Yang 0015, Shihan Wang 0001, Liang Pang 0001
COLING1
2020 PEDNet: A Persona Enhanced Dual Alternating Learning Network for Conversational Response Generation
abstract
Endowing a chatbot with a personality is essential to deliver more realistic conversations.Various persona-based dialogue models have been proposed to generate personalized and diverse responses by utilizing predefined persona information.However, generating personalized responses is still a challenging task since the leverage of predefined persona information is often insufficient.To alleviate this problem, we propose a novel Persona Enhanced Dual Alternating Learning Network (PEDNet) aiming at producing more personalized responses in various opendomain conversation scenarios.PEDNet consists of a Context-Dominated Network (CDNet) and a Persona-Dominated Network (PDNet), which are built upon a common encoder-decoder backbone.CDNet learns to select a proper persona as well as ensure the contextual relevance of the predicted response, while PDNet learns to enhance the utilization of persona information when generating the response by weakening the disturbance of specific content in the conversation context.CDNet and PDNet are trained alternately using a multi-task training approach to equip PEDNet with the both capabilities they have learned.Both automatic and human evaluations on a newly released dialogue dataset Persona-chat demonstrate that our method could deliver more personalized responses than baseline methods.
Bin Jiang 0006, Wanyue Zhou, Jingxu Yang, Chao Yang 0015, Shihan Wang 0001, Liang Pang 0001
COLING1
2020 Rethinking Image Inpainting via a Mutual Encoder-Decoder with Feature Equalizations
Bin Jiang 0006, Yibing Song, Chao Yang 0015
ECCV (2)2
2020 Predicting functional elements and variants effects in non-coding regions based on deep learning
abstract
Accurate recognition and annotation of the important functional elements in the genome is an important prerequisite to understand the coding mode of complex regulatory networks in the one-dimensional genome. Despite rapid advances in sequencing and recognition technologies, accurately calling non-coding variant effects from large-scale sequence reads remains challenging. Here we present a deep neural network-based algorithmic framework, DeepMSA, which directly learns a regulatory sequence code from large-scale chromatin-profiling data,enabling to evaluate chromatin effects caused by SNP(single nucleotide polymorphism).
Yunhao Liu 0001, Shaoliang Peng, Wenjie Shu 0001, Bin Jiang 0006, Chao Yang 0015, Kun Xie 0001
HealthCom4
2020 Constrained R-Cnn: A General Image Manipulation Detection Model
abstract
Recently, deep learning-based models have exhibited remarkable performance for image manipulation detection. However, most of them suffer from poor universality of handcrafted or predetermined features. Meanwhile, they only focus on manipulation localization and overlook manipulation classification. To address these issues, we propose a coarse-to-fine architecture named Constrained R-CNN for complete and accurate image forensics. First, the learnable manipulation feature extractor learns a unified feature representation directly from data. Second, the attention region proposal network effectively discriminates manipulated regions for the next manipulation classification and coarse localization. Then, the skip structure fuses low-level and high-level information to refine the global manipulation features. Finally, the coarse localization information guides the model to further learn the finer local features and segment out the tampered region. Experimental results show that our model achieves state-of-the-art performance. Especially, the F1score is increased by 28.4%, 73.2%, 13.3% on the NIST16, COVERAGE, and Columbia dataset.
Chao Yang 0015, Huizhou Li, Fangting Lin, Bin Jiang 0006
ICME4
2020 Hybrid Dilated Convolution Network Using Attentive Kernels for Real-Time Semantic Segmentation
Jiankai He, Bin Jiang 0006, Chao Yang 0015, Wenxuan Tu
PRCV (1)2
2020 Knowledge Augmented Dialogue Generation with Divergent Facts Selection
Bin Jiang 0006, Jingxu Yang, Chao Yang 0015, Wanyue Zhou, Liang Pang 0001, Xiaokang Zhou
Knowl. Based Syst.1
2020 Gated and attentive neural collaborative filtering for user generated list recommendation
Chao Yang 0015, Lianhai Miao, Bin Jiang 0006, Dongsheng Li 0002, Da Cao
Knowl. Based Syst.3
2020 Context-Integrated and Feature-Refined Network for Lightweight Object Parsing
abstract
Semantic segmentation for lightweight object parsing is a very challenging task, because both accuracy and efficiency (e.g., execution speed, memory footprint or computational complexity) should all be taken into account. However, most previous works pay too much attention to one-sided perspective, either accuracy or speed, and ignore others, which poses a great limitation to actual demands of intelligent devices. To tackle this dilemma, we propose a novel lightweight architecture named Context-Integrated and Feature-Refined Network (CIFReNet). The core components of CIFReNet are the Long-skip Refinement Module (LRM) and the Multi-scale Context Integration Module (MCIM). The LRM is designed to ease the propagation of spatial information between low-level and high-level stages. Furthermore, channel attention mechanism is introduced into the process of long-skip learning to boost the quality of low-level feature refinement. Meanwhile, the MCIM consists of three cascaded Dense Semantic Pyramid (DSP) blocks with image-level features, which is presented to encode multiple context information and enlarge the field of view. Specifically, the proposed DSP block exploits a dense feature sampling strategy to enhance the information representations without significantly increasing the computation cost. Comprehensive experiments are conducted on three benchmark datasets for object parsing including Cityscapes, CamVid, and Helen. As indicated, the proposed method reaches a better trade-off between accuracy and efficiency compared with the other state-of-the-art methods.
Bin Jiang 0006, Wenxuan Tu, Chao Yang 0015, Junsong Yuan 0001
IEEE Trans. Image Process.1
2019 Coherent Semantic Attention for Image Inpainting
abstract
The latest deep learning-based approaches have shown promising results for the challenging task of inpainting missing regions of an image. However, the existing methods often generate contents with blurry textures and distorted structures due to the discontinuity of the local pixels. From a semantic-level perspective, the local pixel discontinuity is mainly because these methods ignore the semantic relevance and feature continuity of hole regions. To handle this problem, we investigate the human behavior in repairing pictures and propose a fined deep generative model-based approach with a novel coherent semantic attention (CSA) layer, which can not only preserve contextual structure but also make more effective predictions of missing parts by modeling the semantic relevance between the holes features. The task is divided into rough, refinement as two steps and we model each step with a neural network under the U-Net architecture, where the CSA layer is embedded into the encoder of refinement step. Meanwhile, we further propose consistency loss and feature patch discriminator to stabilize the network training process and improve the details. The experiments on CelebA, Places2, and Paris StreetView datasets have validated the effectiveness of our proposed methods in image inpainting tasks and can obtain images with a higher quality as compared with the existing state-of-the-art approaches. The codes and pre-trained models will be available at https://github.com/KumapowerLIU/CSA-inpainting.
Bin Jiang 0006, Chao Yang 0015
ICCV2
2019 Cross-Modal Video Moment Retrieval with Spatial and Language-Temporal Attention
abstract
Given an untrimmed video and a description query, temporal moment retrieval aims to localize the temporal segment within the video that best describes the textual query. Existing studies predominantly employ coarse frame-level features as the visual representation, obfuscating the specific details which may provide critical cues for localizing the desired moment. We propose a SLTA (short for "Spatial and Language-Temporal Attention") method to address the detail missing issue. Specifically, the SLTA method takes advantage of object-level local features and attends to the most relevant local features (e.g., the local features "girl", "cup") by spatial attention. Then we encode the sequence of local features on consecutive frames to capture the interaction information among these objects (e.g., the interaction "pour" involving these two objects). Meanwhile, a language-temporal attention is utilized to emphasize the keywords based on moment context information. Therefore, our proposed two attention sub-networks can recognize the most relevant objects and interactions in the video, and simultaneously highlight the keywords in the query. Extensive experiments on TACOS, Charades-STA and DiDeMo datasets demonstrate the effectiveness of our model as compared to state-of-the-art methods.
Bin Jiang 0006, Chao Yang 0015, Junsong Yuan 0001
ICMR1
2019 SLTFNet: A spatial and language-temporal tensor fusion network for video moment retrieval
Bin Jiang 0006, Chao Yang 0015, Junsong Yuan 0001
Inf. Process. Manag.1
2019 Aspect-based sentiment analysis with alternating coattention networks
Chao Yang 0015, Hefeng Zhang, Bin Jiang 0006, Keqin Li 0001
Inf. Process. Manag.3