Cheng Jin 0001

dblp:84/4201-1 · DBLP profile ↗
← Back
74ranked-venue papers
7as first author
44since 2021 · last 2026
0000-0003-3063-1957ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 56 · 5 first-author · 37 since 2021Artificial intelligence and machine learning · 39 · 1 first-author · 23 since 2021Databases, data management, data science and information retrieval · 7 · 2 first-authorHuman-computer interaction and ubiquitous computing · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-authorComputer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Explore and Establish Synergistic Effects Between Weight Pruning and Coreset Selection in Neural Network Training
abstract
Modern deep neural networks rely heavily on massive model weights and training samples, incurring substantial computational costs. Weight pruning and coreset selection are two emerging paradigms proposed to improve computational efficiency. In this paper, we first explore the interplay between redundant weights and training samples through a transparent analysis: redundant samples, particularly noisy ones, cause model weights to become unnecessarily overtuned to fit them, complicating the identification of irrelevant weights during pruning; conversely, irrelevant weights tend to overfit noisy data, undermining coreset selection effectiveness. To further investigate and harness this interplay in deep learning, we develop a Simultaneous Weight and Sample Tailoring mechanism (SWaST) that alternately performs weight pruning and coreset selection to establish a synergistic effect in training. During this investigation, we observe that when simultaneously removing a large number of weights and samples, a phenomenon we term critical double-loss can occur, where important weights and their supportive samples are mistakenly eliminated at the same time, leading to model instability and nearly irreversible degradation that cannot be recovered in subsequent training. Unlike classic machine learning models, this issue can arise in deep learning due to the lack of theoretical guarantees on the correctness of weight pruning and coreset selection, which explains why these paradigms are often developed independently. We mitigate this by integrating a state preservation mechanism into SWaST, enabling stable joint optimization. Extensive experiments reveal a strong synergy between pruning and coreset selection across varying prune rates and coreset sizes, delivering accuracy boosts of up to 17.83% alongside 10% to 90% FLOPs reductions.
Weilin Wan 0002, Cheng Jin 0001
AAAI5
2026 Advanced Black-Box Tuning of Large Language Models with Limited API Calls
abstract
Black-box tuning is an emerging paradigm for adapting large language models (LLMs) to better achieve desired behaviors, particularly when direct access to model parameters is unavailable. Current strategies, however, often present a dilemma of suboptimal extremes: either separately train a small proxy model and then use it to shift the predictions of the foundation model, offering notable efficiency but often yielding limited improvement; or making API calls in each tuning iteration to the foundation model, which entails prohibitive computational costs. In this paper, we argue that a more reasonable way for black-box tuning is to train the proxy model with limited API calls. The underlying intuition is based on two key observations: first, the training samples may exhibit correlations and redundancies, suggesting that the foundation model’s predictions can be estimated from previous calls; second, foundation models frequently demonstrate low accuracy on downstream tasks. Therefore, we propose a novel advanced black-box tuning method for LLMs with limited API calls. Our core strategy involves training a Gaussian Process (GP) surrogate model with "LogitMap Pairs" derived from querying the foundation model on a minimal but highly informative training subset. This surrogate can approximate the outputs of the foundation model to guide the training of the proxy model, thereby effectively reducing the need for direct queries to the foundation model. Extensive experiments verify that our approach elevates pre-trained language model accuracy from 55.92% to 86.85%, reducing the frequency of API queries to merely 1.38%. This significantly outperforms offline approaches that operate entirely without API access. Notably, our method also achieves comparable or superior accuracy to query-intensive approaches, while significantly reducing API costs. This offers a robust and high-efficiency paradigm for language model adaptation.
Zhikang Xie, Weilin Wan 0002, Peizhu Gong, Cheng Jin 0001
AAAI5
2025 Optimized Gradient Clipping for Noisy Label Learning
abstract
Previous research has shown that constraining the gradient of loss function w.r.t. model-predicted probabilities can enhance the model robustness against noisy labels. These methods typically specify a fixed optimal threshold for gradient clipping through validation data to obtain the desired robustness against noise. However, this common practice overlooks the dynamic distribution of gradients from both clean and noisy-labeled samples at different stages of training, significantly limiting the model capability to adapt to the variable nature of gradients throughout the training process. To address this issue, we propose a simple yet effective approach called Optimized Gradient Clipping (OGC), which dynamically adjusts the clipping threshold based on the ratio of noise gradients to clean gradients after clipping, estimated by modeling the distributions of clean and noisy samples. This approach allows us to modify the clipping threshold at each training step, effectively controlling the influence of noise gradients. Additionally, we provide statistical analysis to certify the noise-tolerance ability of OGC. Our extensive experiments across various types of label noise, including symmetric, asymmetric, instance-dependent, and real-world noise, demonstrate the effectiveness of our approach.
Xichen Ye, Yifan Wu 0011, Xiaoqiang Li 0002, Yifan Chen 0004, Cheng Jin 0001
AAAI6
2025 An Exemplar-based Framework for Chinese Text Recognition
abstract
This paper introduces a novel exemplar-based framework for reading Chinese texts in natural scene or document images. We present the Deep Exemplar-based Chinese Text Recognizer, which is structured to first identify candidate characters as exemplars from each text-line, and subsequently recognize them by retrieving analogous exemplars from a database. With text-line level annotations, we design the exemplar discovery network to simultaneously recognize texts and capture individual character positions in a weak-supervision manner. The exemplar retrieval module is then crafted to identify the most similar exemplar and propagate the corresponding character label. This enables us to effectively rectify the misrecognized characters and boost the performance of scene text recognition. Experiments on four scenarios of Chinese texts demonstrate the effectiveness of our proposed framework.
Zhao Zhou, Xiangcheng Du, Yingbin Zheng, Xingjiao Wu, Cheng Jin 0001
AAAI5
2025 Achieving Ensemble-Like Performance in a Single Model: A Feature Diversification Framework for Image-Text Matching
abstract
Model ensembling is a widely used technique that enhances performance in image-text matching tasks by combining multiple models, each trained with different initializations. However, the inefficiencies associated with training several models and generating outputs from them constrain their practical applicability. In this paper, we argue that while the parameters of two randomly initialized models can differ significantly, their feature distributions can be similar at certain stages. By employing a proposed technique called cross-modal realignment, we demonstrate that features derived from differently initialized models maintain similarity at the feature extraction stage and can be effectively transformed by fine-tuning a small number of parameters. These findings provide an efficient way to achieve ensemble-like performance within a single model. Specifically, we propose a Feature Diversification Framework (FDF) that emulates the outputs of multiple model initializations to generate diverse features from a common shared feature. Firstly, we introduce feature conversion methods to transform shared features into a set of distinct features. Next, a realignment training strategy is presented to optimize negative pairs for realigning these transformed features, thereby enhancing their diversification to resemble the outputs of different models. Additionally, we propose a reweighting module that assigns weights to these features, enabling a weighted fusion approach for robust feature representation. Extensive experiments on the Flickr30K and MS-COCO datasets demonstrate the effectiveness and generalizability of our framework.
Zhao Zhou, Yingbin Zheng, Xiangcheng Du, Cheng Jin 0001
AAAI6
2025 Expanding the Scope of Negatives: Boosting Image-Text Matching with Negatives Distribution Guided Learning
abstract
Image-text matching is a crucial task that bridges visual and linguistic modalities. Recent research typically formulates it into the problem of maximizing the margin with the truly hardest negatives to enhance the learning efficiency and avoid the poor local optima. We argue that such formulation can lead to a serious limitation, i.e., under this formulation, conventional trainers would confine their horizon within the hardest negative examples, while other negative examples offer a range of semantic differences not present in the hardest negatives. In this paper, we propose an efficient negative distribution guided training framework for image-text matching to unlock the substantial promotion space left by the above limitation. Rather than simply incorporating additional negative examples into the training objective, which could diminish both the leading role of the hardest negatives in training and the effect of a large margin learning in producing a robust matching model, our central idea is to supply the objective with distributional information on the entire set of negative examples. To be precise, we first construct the sample similarity matrix based on several pretrained models to extract the distributional information of the entire negative sample dataset. Then we encode it into a margin regularization module to smooth the similarities differences of all negatives. This enhancement facilitates the capture of fine-grained semantic differences and guides the main learning process by maximizing the margin with hard negative examples. Furthermore, we propose a hardest negative rectification module to address the instability in hardest negative selection based on predicted similarity and to correct erroneous hardest negatives. We evaluate our method in combination with several state-of-the-art image-text matching methods, and our quantitative and qualitative experiments demonstrate its significant generalizability and effectiveness.
Zhao Zhou, Xiangcheng Du, Yingbin Zheng, Cheng Jin 0001
AAAI5
2025 Population Normalization for Federated Learning
abstract
Batch normalization (BN) is widely recognized as an essential method in training deep neural networks, facilitating convergence and enhancing model stability. However, in Federated Learning (FL) contexts, where training data are typically heterogeneous and clients often face resource constraints, the effectiveness of BN is considerably limited for two primary reasons. First, the population statistics, specifically the mean and variance of these heterogeneous datasets, vary substantially, resulting in inconsistent BN layers across client models, which ultimately drives these models to diverge further. Second, estimating statistics from a mini-batch is often imprecise since the batch size has to be small in resource-limited clients. This paper introduces Population Normalization, a novel technique for FL, in which the statistics are learned as trainable parameters rather than calculated from mini-batches as in BN. Thus, our normalization layers are homogeneous among the clients and the adverse impact of small batch size is eliminated as the model can be well-trained even when the batch size equals to one. To enhance the flexibility of our method in practical applications, we investigate the role of stochastic uncertainty in BN’s statistical estimation. When larger batch sizes are available, we demonstrate that injecting simple artificial noise can effectively mimic this stochastic uncertainty and improve the model’s generalization capability. Experimental results validate the efficacy of our approach across various FL tasks.
Peizhu Gong, Caitou He, Cheng Jin 0001
CVPR5
2025 DreamText: High Fidelity Scene Text Synthesis
abstract
Scene text synthesis involves rendering specified texts onto arbitrary images. Current methods typically formulate this task in an end-to-end manner but lack effective character-level guidance during training. Besides, their text encoders, pre-trained on a single font type, struggle to adapt to the diverse font styles encountered in practical applications. Consequently, these methods suffer from character distortion, repetition, and absence, particularly in polystylistic scenarios. To this end, this paper proposes DreamText for high-fidelity scene text synthesis. Our key idea is to reconstruct the diffusion training process, introducing more refined guidance tailored to this task, to expose and rectify the model’s attention at the character level and strengthen its learning of text regions. This transformation poses a hybrid optimization challenge, involving both discrete and continuous variables. To effectively tackle this challenge, we employ a heuristic alternate optimization strategy. Meanwhile, we jointly train the text encoder and generator to comprehensively learn and utilize the diverse font present in the training dataset. This joint training is seamlessly integrated into the alternate optimization process, fostering a synergistic relationship between learning character embedding and re-estimating character attention. Specifically, in each step, we first encode potential character-generated position information from cross-attention maps into latent character masks. These masks are then utilized to update the representation of specific characters in the current step, which, in turn, enables the generator to correct the character’s attention in the subsequent steps. Both qualitative and quantitative results demonstrate the superiority of our method to the state of the art. Our project page is here.
Honghui Xu 0002, Cheng Jin 0001
CVPR4
2025 Towards Robust Influence Functions with Flat Validation Minima
abstract
The Influence Function (IF) is a widely used technique for assessing the impact of individual training samples on model predictions. However, existing IF methods often fail to provide reliable influence estimates in deep neural networks, particularly when applied to noisy training data. This issue does not stem from inaccuracies in parameter change estimation, which has been the primary focus of prior research, but rather from deficiencies in loss change estimation, specifically due to the sharpness of validation risk. In this work, we establish a theoretical connection between influence estimation error, validation set risk, and its sharpness, underscoring the importance of flat validation minima for accurate influence estimation. Furthermore, we introduce a novel estimation form of Influence Function specifically designed for flat validation minima. Experimental results across various tasks validate the superiority of our approach.
Xichen Ye, Yifan Wu 0011, Cheng Jin 0001, Yifan Chen 0004
ICML4
2025 Unleashing the Semantic Adaptability of Controlled Diffusion Model for Image Colorization
abstract
Recent data-driven image colorization methods have leveraged pre-trained Text-to-Image (T2I) diffusion models as generative prior, while still suffering from unsatisfactory and inaccurate semantic-level color control. To address these issues, we propose a Semantic Adaptation method (SeAda) that enhances the prior while considering the semantic discrepancy between color and grayscale image pairs. The SeAda employs a semantic adapter to produce refined semantic embeddings and a controlled T2I diffusion model to create reasonably colored images. Specifically, the semantic adapter transfers the embedding from grayscale to color domain, while the diffusion model utilizes the refined embedding and prior knowledge to achieve realistic and diverse results. We also design a three-staged training strategy to improve semantic comprehension and prior integration for further performance improvement. Extensive experiments on public datasets demonstrate that our method outperforms existing state-of-the-art techniques, yielding superior performance in image colorization.
Xiangcheng Du, Zhao Zhou, Yingbin Zheng, Xingjiao Wu, Peizhu Gong, Cheng Jin 0001
IJCAI7
2025 McGE '25: The 3rd International Workshop on Multimedia Content Generation and Evaluation: New Methods and Practice
abstract
This workshop addresses next-generation methods in multimedia research, with a focus on content generation, quality assessment, and dataset development. These three areas are foundational for advancing multimedia technologies and applications. Emerging approaches in multimedia content generation, powered by generative AI and multimodal learning, are reshaping domains such as entertainment, advertising, education, and healthcare. At the same time, robust quality assessment is essential to ensure that generated content achieves high standards of perceptual fidelity, semantic consistency, and user satisfaction, thereby determining the real-world impact of multimedia systems. Datasets remain indispensable for training and evaluating algorithms, and innovative strategies in dataset construction-ranging from augmentation and annotation to addressing issues of bias and small-sample imbalance-are driving the development of more reliable and ethical multimedia applications. By convening leading researchers and practitioners, this workshop provides a platform to explore state-of-the-art methods, share best practices, and discuss open challenges in next-generation multimedia research. The goal is to foster interdisciplinary collaboration and inspire innovative solutions that advance the creation, evaluation, and application of multimedia content, setting new benchmarks for the field and shaping the future of multimedia technologies.
Cheng Jin 0001, Mingli Song, Rui Wang 0032, Xingjiao Wu
ACM Multimedia1
2025 Computational Budget Should Be Considered in Data Selection
abstract
Data selection improves computational efficiency by choosing informative subsets of training samples. However, existing methods ignore the compute budget, treating data selection and importance evaluation independently of compute budget constraints. Yet empirical studies show no algorithm can consistently outperform others (or even random selection) across varying budgets. We therefore argue that compute budget must be integral to data-selection strategies, since different budgets impose distinct requirements on data quantity, quality, and distribution for effective training. To this end, we propose a novel Computational budget-Aware Data Selection (CADS) method and naturally formulate it into a bilevel optimization framework, where the inner loop trains the model within the constraints of the computational budget on some selected subset of training data, while the outer loop optimizes data selection based on model evaluation. Our technical contributions lie in addressing two main challenges in solving this bilevel optimization problem: the expensive Hessian matrix estimation for outer-loop gradients and the computational burden of achieving inner-loop optimality during iterations. To solve the first issue, we propose a probabilistic reparameterization strategy and compute the gradient using a Hessian-free policy gradient estimator. To address the second challenge, we transform the inner optimization problem into a penalty term in the outer objective, further discovering that we only need to estimate the minimum of a one-dimensional loss to calculate the gradient, significantly improving efficiency. To accommodate different data selection granularities, we present two complementary CADS variants: an example-level version (CADS-E) offering fine-grained control and a source-level version (CADS-S) aggregating samples into source groups for scalable, efficient selection without sacrificing effectiveness. Extensive experiments show that our method achieves performance gains of up to 14.42\% over baselines in vision and language benchmarks. Additionally, CADS achieves a 3-20× speedup compared to conventional bilevel implementations, with acceleration correlating positively with compute budget size.
Weilin Wan 0002, Cheng Jin 0001
NeurIPS3
2024 Point Cloud Part Editing: Segmentation, Generation, Assembly, and Selection
abstract
Ideal part editing should guarantee the diversity of edited parts, the fidelity to the remaining parts, and the quality of the results. However, previous methods do not disentangle each part completely, which means the edited parts will affect the others, resulting in poor diversity and fidelity. In addition, some methods lack constraints between parts, which need manual selections of edited results to ensure quality. Therefore, we propose a four-stage process for point cloud part editing: Segmentation, Generation, Assembly, and Selection. Based on this process, we introduce SGAS, a model for part editing that employs two strategies: feature disentanglement and constraint. By independently fitting part-level feature distributions, we realize the feature disentanglement. By explicitly modeling the transformation from object-level distribution to part-level distributions, we realize the feature constraint. Considerable experiments on different datasets demonstrate the efficiency and effectiveness of SGAS on point cloud part editing. In addition, SGAS can be pruned to realize unsupervised part-aware point cloud generation and achieves state-of-the-art results.
Kaiyi Zhang 0002, Ximing Yang, Cheng Jin 0001
AAAI5
2024 FusionFormer: A Concise Unified Feature Fusion Transformer for 3D Pose Estimation
abstract
Depth uncertainty is a core challenge in 3D human pose estimation, especially when the camera parameters are unknown. Previous methods try to reduce the impact of depth uncertainty by multi-view and/or multi-frame feature fusion to utilize more spatial and temporal information. However, they generally lead to marginal improvements and their performance still cannot match the camera-parameter-required methods. The reason is that their handcrafted fusion schemes cannot fuse the features flexibly, e.g., the multi-view and/or multi-frame features are fused separately. Moreover, the diverse and complicated fusion schemes make the principle for developing effective fusion schemes unclear and also raises an open problem that whether there exist more simple and elegant fusion schemes. To address these issues, this paper proposes an extremely concise unified feature fusion transformer (FusionFormer) with minimized handcrafted design for 3D pose estimation. FusionFormer fuses both the multi-view and multi-frame features in a unified fusion scheme, in which all the features are accessible to each other and thus can be fused flexibly. Experimental results on several mainstream datasets demonstrate that FusionFormer achieves state-of-the-art performance. To our best knowledge, this is the first camera-parameter-free method to outperform the existing camera-parameter-required methods, revealing the tremendous potential of camera-parameter-free models. These impressive experimental results together with our concise feature fusion scheme resolve the above open problem. Another appealing feature of FusionFormer we observe is that benefiting from its effective fusion scheme, we can achieve impressive performance with smaller model size and less FLOPs.
Yanlu Cai, Yuan Wu 0004, Cheng Jin 0001
AAAI4
2024 PoseIRM: Enhance 3D Human Pose Estimation on Unseen Camera Settings via Invariant Risk Minimization
abstract
Camera-parameter-free multi-view pose estimation is an emerging technique for 3D human pose estimation (HPE). They can infer the camera settings implicitly or explicitly to mitigate the depth uncertainty impact, showcasing significant potential in real applications. However, due to the limited camera setting diversity in the available datasets, the inferred camera parameters are always simply hard-coded into the model during training and not adaptable to the input in inference, making the learned models cannot generalize well under unseen camera settings. A natural solution is to artificially synthesize some samples, i.e., 2D-3D pose pairs, under massive new camera settings. Un-fortunately, to prevent over-fitting the existing camera setting, the number of synthesized samples for each new camera setting should be comparable with that for the existing one, which multiplies the scale of training and even makes it computationally prohibitive. In this paper, we propose a novel HPE approach under the invariant risk minimization (IRM) paradigm. Precisely, we first synthesize 2D poses from myriad camera settings. We then train our model under the IRM paradigm, which targets at learning a common optimal model across all camera settings and thus enforces the model to automatically learn the camera parameters based on the input data. This allows the model to accurately infer 3D poses on unseen data by training on only a hand-ful of samples from each synthesized setting and thus avoid the unbearable training cost increment. Another appealing feature of our method is that benefited from the capability of IRM in identifying the invariant features, its performance on the seen camera settings is enhanced as well. Compre-hensive experiments verify the superiority of our approach.
Yanlu Cai, Yuan Wu 0004, Cheng Jin 0001
CVPR4
2024 High-fidelity Person-centric Subject-to-Image Synthesis
abstract
Current subject-driven image generation methods en-counter significant challenges in person-centric image generation. The reason is that they learn the semantic scene and person generation by fine-tuning a common pre-trained diffusion, which involves an irreconcilable training imbalance. Precisely, to generate realistic persons, they need to sufficiently tune the pre-trained model, which inevitably causes the model to forget the rich semantic scene prior and makes scene generation over-fit to the training data. Moreover, even with sufficient fine-tuning, these methods can still not generate high-fidelity persons since joint learning of the scene and person generation also lead to quality compromise. In this paper, we propose Face-diffuser, an effective collaborative generation pipeline to eliminate the above training imbal-ance and quality compromise. Specifically, we first develop two specialized pre-trained diffusion models, i.e., Text-driven Diffusion Model (TDM) and Subject-augmented Diffusion Model (SDM), for scene and person generation, respectively. The sampling process is divided into three sequential stages, i.e., semantic scene construction, subject-scene fusion, and subject enhancement. The first and last stages are performed by TDM and SDM respectively. The subject-scene fusion stage, that is the collaboration achieved through a novel and highly effective mechanism, Saliency-adaptive Noise Fusion (SNF). Specifically, it is based on our key observation that there exists a robust link between classifier-free guidance responses and the saliency of generated images. In each time step, SNF leverages the unique strengths of each model and allows for the spatial blending of predicted noises from both models automatically in a saliency-aware manner, all of which can be seamlessly integrated into the DDIM sampling process. Extensive experiments confirm the impressive effectiveness and robustness of the Face-diffuser in gener-ating high-fidelity person images depicting multiple unseen persons with varying contexts. Code is available at https://github.com/CodeGoat24/Face-diffuser.
Jianwei Zheng 0001, Cheng Jin 0001
CVPR4
2024 Target Optimization Direction Guided Transfer Learning for Image Classification
abstract
At present, deep learning has made impressive achievements in various fields; however, effectively training deep neural networks on small data sets remains a significant challenge. Transfer learning, as a method of efficient training across multiple tasks, has been widely used to solve this problem. However, when the domain gap or the data volume difference between the two tasks is too large, the transfer learning may not perform well, and other optimization methods will be required to improve the performance. In this paper, we propose a new transfer learning method guided by the direction of objective optimization from the perspective of gradient. This method guides the gradient direction of the source task towards the gradient direction of the target task. In several similar and conflicting tasks, this method has achieved good results in efficiency and performance. In comparison with other transfer learning methods, the results shown by this method are generally better.
Kelvin Ting Zuo Han, Shengxuming Zhang, Gerard Marcos Freixas, Zunlei Feng, Cheng Jin 0001
ICASSP5
2024 Efficient Scene Text Image Super-Resolution with Semantic Guidance
abstract
Scene text image super-resolution has significantly improved the accuracy of scene text recognition. However, many existing methods emphasize performance over efficiency and ignore the practical need for lightweight solutions in deployment scenarios. Faced with the issues, our work proposes an efficient framework called SGENet to facilitate deployment on resource-limited platforms. SGENet contains two branches: super-resolution branch and semantic guidance branch. We apply a lightweight pre-trained recognizer as a semantic extractor to enhance the understanding of text information. Meanwhile, we design the visual-semantic alignment module to achieve bidirectional alignment between image features and semantics, resulting in the generation of high-quality prior guidance. We conduct extensive experiments on benchmark dataset, and the proposed SGENet achieves excellent performance with fewer computational costs.
LeoWu TomyEnrique, Xiangcheng Du, Kangliang Liu, Zhao Zhou, Cheng Jin 0001
ICASSP6
2024 Fine-Grained Scene Image Classification with Modality-Agnostic Adapter
abstract
When dealing with the task of fine-grained scene image classification, most previous works lay much emphasis on global visual features when doing multi-modal feature fusion. In other words, models are deliberately designed based on prior intuitions about the importance of different modalities. In this paper, we present a new multi-modal feature fusion approach named MAA (Modality-Agnostic Adapter), trying to make the model learn the importance of different modalities in different cases adaptively, without giving a prior setting in the model architecture. More specifically, we eliminate the modal differences in distribution and then use a modality-agnostic Transformer encoder for a semantic-level feature fusion. Our experiments demonstrate that MAA achieves state-of-the-art results on benchmarks by applying the same modalities with previous methods. Besides, it is worth mentioning that new modalities can be easily added when using MAA and further boost the performance.
Zhao Zhou, Xiangcheng Du, Xingjiao Wu, Yingbin Zheng, Cheng Jin 0001
ICME6
2024 Minutes to Seconds: Speeded-up DDPM-based Image Inpainting with Coarse-to-Fine Sampling
abstract
For image inpainting, the existing Denoising Diffusion Probabilistic Model (DDPM) based method i.e. RePaint can produce high-quality images for any inpainting form. It utilizes a pre-trained DDPM as a prior and generates inpainting results by conditioning on the reverse diffusion process, namely denoising process. However, this process is significantly time-consuming. In this paper, we propose an efficient DDPM-based image inpainting method which includes three speed-up strategies. First, we utilize a pre-trained Light-Weight Diffusion Model (LWDM) to reduce the number of parameters. Second, we introduce a skip-step sampling scheme of Denoising Diffusion Implicit Models (DDIM) for the denoising process. Finally, we propose Coarse-to-Fine Sampling (CFS), which speeds up inference by reducing image resolution in the coarse stage and decreasing denoising timesteps in the refinement stage. We conduct extensive experiments on both faces and general-purpose image inpainting tasks, and our method achieves competitive performance with approximately 60 times speedup.
Xiangcheng Du, LeoWu TomyEnrique, Yingbin Zheng, Cheng Jin 0001
ICME6
2024 ELiTe: Efficient Image-to-LiDAR Knowledge Transfer for Semantic Segmentation
abstract
Cross-modal knowledge transfer enhances point cloud representation learning in LiDAR semantic segmentation. Despite its potential, the weak teacher challenge arises due to repetitive and non-diverse car camera images and sparse, inaccurate ground truth labels. To address this, we propose the Efficient Image-to-LiDAR Knowledge Transfer (ELiTe) paradigm. ELiTe introduces Patch-to-Point Multi-Stage Knowledge Distillation, transferring comprehensive knowledge from the Vision Foundation Model (VFM), extensively trained on diverse open-world images. This enables effective knowledge transfer to a lightweight student model across modalities. ELiTe employs Parameter-Efficient Fine-Tuning to strengthen the VFM teacher and expedite large-scale model training with minimal costs. Additionally, we introduce the Segment Anything Model based Pseudo-Label Generation approach to enhance low-quality image labels, facilitating robust semantic representations. Efficient knowledge transfer in ELiTe yields state-of-the-art results on the SemanticKITTI benchmark, outperforming real-time inference models. Our approach achieves this with significantly fewer parameters, confirming its effectiveness and efficiency.
Ximing Yang, Cheng Jin 0001
ICME4
2024 MultiColor: Image Colorization by Learning from Multiple Color Spaces
Xiangcheng Du, Zhao Zhou, Xingjiao Wu, Yingbin Zheng, Cheng Jin 0001
ACM Multimedia7
2024 PrimeComposer: Faster Progressively Combined Diffusion for Image Composition with Attention Steering
abstract
an oil painting of an eiffel tower in the distance, Van Gogh Style an oil painting of a shopping mall in the distance, Van Gogh Style a pencil drawing of a car and a willow, black and white painting a pencil drawing of buildings in the distance, black and white painting a cartoon animation of an elephant in the forest a cartoon animation of a cabinet a professional photograph of a wet puppy in a pool, ultra realistic a professional photograph of a castle in the distance, ultra realistic Figure 1: Displayed are the results generated using PrimeComposer, showcasing its prowess across various domains: oil painting, sketching, cartoon animation, and photorealism.
Jianwei Zheng 0001, Cheng Jin 0001
ACM Multimedia4
2024 Future Motion Dynamic Modeling via Hybrid Supervision for Multi-Person Motion Prediction Uncertainty Reduction
abstract
Multi-person motion prediction remains a challenging problem due to the intricate motion dynamics and complex interpersonal interactions, where uncertainty escalates rapidly across the forecasting horizon. Existing approaches always overlook the motion dynamic modeling among the prediction frames to reduce the uncertainty, but leave it entirely up to the deep neural networks, which lacks a dynamic inductive bias, leading to suboptimal performance. This paper addresses this limitation by proposing an effective multi-person motion prediction method named Hybrid Supervision Transformer (HSFormer), which formulates the dynamic modeling within the prediction horizon as a novel hybrid supervision task. To be precise, our method performs a rolling predicting process equipped with a hybrid supervision mechanism, which enforces the model to be able to predict the pose in the next frames based on the (typically error-contained) earlier predictions. Addition to the standard supervision loss, two self and auxiliary supervision mechanisms, which minimize the distance of the predictions with error-contained inputs and the predictions with error-free inputs (ground truth) and guide the model to make accurate predictions based on the ground truth, are introduced to improve the robustness of our model to the input deviation in inference and stabilize the training process, respectively. The optimization techniques, such as stop-gradient, are extended to our model to improve the training efficiency. Furthermore, we develop a fine-grained spatio-temporal correlation capture module to assist the feature learning and reduce the uncertainties arising from the intricate and varying interactions among the individuals. Our approach achieves state-of-the-art results on multiple multi-person datasets in both short- and long-term prediction.
Yan Zhuang 0008, Yanlu Cai, Cheng Jin 0001
ACM Multimedia4
2024 Cross-domain document layout analysis using document style guide
Xingjiao Wu, Luwei Xiao, Xiangcheng Du, Yingbin Zheng, Xin Li 0110, Tianlong Ma, Cheng Jin 0001, Liang He 0001
Expert Syst. Appl.7
2024 S2CL-Leaf Net: Recognizing Leaf Images Like Human Botanists
abstract
Automatically classifying plant leaves is a challenging fine-grained classification task because of the diversity in leaf morphology, including size, texture, shape, and venation. Although powerful deep learning-based methods have achieved great improvement in leaf classification, these methods still require a large number of well-labeled samples for supervised training, which is difficult to get. In contrast, relying on the specific coarse-to-fine classification strategy, human botanists only require a small number of samples for accurate leaf recognition. Inspired by the classification strategy of human botanists, we propose a novel S 2 CL-Leaf Net , which exploits multi-granularity clues with a hierarchical attention mechanism and boosts the learning ability with the supervised sampling contrastive learning with limited training samples to classify plant leaves as human botanists do. Specifically, to fully explore and exploit the subtle details of the leaves, a novel sampling transformation mechanism is combined with the supervised contrastive learning to enhance the network’s perception of details by amplifying the discriminative regions with a weighted sampling of different regions. Furthermore, we construct the hierarchical attention mechanism to produce attention maps of different granularity, which helps to discover details in leaves that are important for classification. Experiments are conducted on the open-access leaf datasets, including Flavia, Swedish, and LeafSnap, which prove the effectiveness of the proposed S 2 CL-Leaf Net .
Cong Zou, Rui Wang 0032, Cheng Jin 0001, Sanyi Zhang, Xin Wang 0019
ACM Trans. Multim. Comput. Commun. Appl.3
2023 DDT: Dual-branch Deformable Transformer for Image Denoising
abstract
Transformer is beneficial for image denoising tasks since it can model long-range dependencies to overcome the limitations presented by inductive convolutional biases. However, directly applying the transformer structure to remove noise is challenging because its complexity grows quadratically with the spatial resolution. In this paper, we propose an efficient Dual-branch Deformable Transformer (DDT) denoising network which captures both local and global interactions in parallel. We divide features with a fixed patch size and a fixed number of patches in local and global branches, respectively. In addition, we apply deformable attention operation in both branches, which helps the network focus on more important regions and further reduces computational complexity. We conduct extensive experiments on real-world and synthetic denoising tasks, and the proposed DDT achieves state-of-the-art performance with significantly fewer computational costs.
Kangliang Liu, Xiangcheng Du, Yingbin Zheng, Xingjiao Wu, Cheng Jin 0001
ICME6
2023 Image Layer Modeling for Complex Document Layout Generation
abstract
Document layout analysis (DLA) plays an essential role in information extraction and document understanding. At present, DLA has reached the milestone achievement; however, DLA of non-Manhattan is still challenging because of annotation data limitations. In this paper, we propose an image layer modeling method to mitigate this issue. The image layer modeling method generates document images of non-Manhattan layouts by superimposing images under pre-defined aesthetic rules. Due to the lack of evaluation benchmark for non-Manhattan layout, we have constructed a manually-labeled non-Manhattan layout fine-grained segmentation dataset. To the best of our knowledge, this is the first manually-labeled non-Manhattan layout fine-grained segmentation dataset. Extensive experimental results verify that our proposed image layer modeling method can better deal with the fine-grained segmented document of the non-Manhattan layout.
Tianlong Ma, Xingjiao Wu, Xiangcheng Du, Cheng Jin 0001
ICME5
2023 Weakly-supervised Temporal Action Localization with Adaptive Clustering and Refining Network
abstract
Weakly-supervised temporal action localization task aims to localize temporal boundaries of action instances by using only video-level labels. Existing methods primarily adopt Multi-Instance-Learning (MIL) scheme to handle this task. The effectiveness of MIL scheme depends heavily on the selection of top-k action snippets, which is unstable and requires manual tuning. To address these deficiencies, we propose an Adaptive Clustering and Refining Network (ACRNet). Specifically, we present an action-aware clustering strategy that is adaptable and requires no manual tuning to separate action and background snippets of diverse videos based on intra-class activation distribution. And a cluster refining step is included to eliminate false action snippets by considering inter-class activation distribution, which greatly improves robustness and localization accuracy. Extensive experiments on THUMOS14, ActivityNet 1.2&1.3 benchmarks show that our method achieves state-of-the-art performance.
Hao Ren 0002, Wu Ran, Xingson Liu, Hong Lu 0001, Cheng Jin 0001
ICME7
2023 Pose-Motion Video Anomaly Detection via Memory-Augmented Reconstruction and Conditional Variational Prediction
abstract
Video anomaly detection (VAD) is a challenging computer vision problem. Due to the scarcity of anomalous events in training, the models learned by existing methods would mistakenly fit the ubiquitous non-causal or even spurious correlations, leading to failure in inference. In this paper, we propose a new two-phase Pose-Motion Video Anomaly Detection (PoMo) approach by jointly exploiting the informative features including the poses and optical flows that have rich causal correlations with abnormality. PoMo can effectively prevent the non-causal features from leaking in by either encoding only the essential information, i.e., the poses and optical flows, with our normalized autoencoder (phase one), or separately modeling the knowledge learned in phase one using our causal-conditioned autoencoder (phase two). The difference between normal and abnormal events can be amplified through these two phases. Thus the generalization ability can be reinforced. Extensive experimental results demonstrate the superiority of our approach over the existing methods and the improvements in AUC-ROC can be up to 1.5%.
Weilin Wan 0002, Cheng Jin 0001
ICME3
2023 McGE '23: 1st International Workshop on Multimedia Content Generation and Evaluation: New Methods and Practice
abstract
The proposed workshop's topics, focusing on multimedia content generation, quality assessment, datasets and construction, is crucial due to its direct impact on the growth and success of the multimedia field. Multimedia content generation is essential for various applications, such as entertainment, advertising, and education. Quality assessment ensures the overall value and effectiveness of multimedia content, directly influencing user satisfaction and application success. Datasets are indispensable for training and evaluating multimedia algorithms, driving innovation, and fostering progress in the field. Finally, effective dataset construction methods set new benchmarks for the research community, stimulating innovation and unlocking new opportunities for leveraging multimedia data in various applications. The goal of this workshop is to bring together leading researchers in the field in a joint forum for advancing multimedia content generation and evaluation.
Cheng Jin 0001, Liang He 0001, Mingli Song, Rui Wang 0032
ACM Multimedia1
2023 Weakly-Supervised Temporal Action Localization with Regional Similarity Consistency
Hao Ren 0002, Hong Lu 0001, Cheng Jin 0001
MMM (1)4
2023 Modeling Stroke Mask for End-to-End Text Erasing
abstract
Scene text erasing aims to wipe text regions in scene images with reasonable background. Most previous approaches employ scene text detectors to assist localization of the text regions. However, detected text boxes contain both text strokes and background clutters, and directly in-painting on the whole boxes may remain text artifacts and make regions unnatural. In this paper, we present an end-to-end network that focuses on modeling text stroke masks that provide more accurate locations to compute erased images. The network consists of two stages, i.e., a basic network with stroke generation and a refinement network with stroke awareness. The basic network predicts the text stroke masks and initial erasing results simultaneously. The refinement network receives the masks as supervision to generate natural erased results. Experiments on both synthetic and real-world scene images demonstrate the effectiveness of our framework in producing high quality erasing results.
Xiangcheng Du, Zhao Zhou, Yingbin Zheng, Tianlong Ma, Xingjiao Wu, Cheng Jin 0001
WACV6
2023 ETR: An Efficient Transformer for Re-ranking in Visual Place Recognition
abstract
Visual place recognition is to estimate the geographical location of a given image, which is usually addressed by recognizing its similar reference images from a database. The reference images are usually retrieved via similarity search using global descriptor, and the local descriptors are used to re-rank the initial retrieved candidates. The local descriptors re-ranking can significantly improve the accuracy of global retrieval but comes at a high computational cost. To achieve a good trade-off between accuracy and efficiency, we propose an Efficient Transformer for Re-ranking (ETR), utilizing both global and local descriptors to re-rank the top candidates in a single shot. In contrast to traditional re-ranking methods, we leverage self-attention to capture relationships between local descriptors in a single image and cross-attention to explore the similarity of the image pairs. We show that the proposed model can be regarded as a general re-ranking algorithm for significantly boosting the performance of other global-only retrieval methods. Extensive experimental results show that our method outperforms state-of-the-arts and is orders of magnitude faster in terms of computational efficiency.
Heming Jing, Yingbin Zheng, Yuan Wu 0004, Cheng Jin 0001
WACV6
2023 Progressive scene text erasing with self-supervision
Xiangcheng Du, Zhao Zhou, Yingbin Zheng, Xingjiao Wu, Tianlong Ma, Cheng Jin 0001
Comput. Vis. Image Underst.6
2023 Reading Scene Text with Aggregated Temporal Convolutional Encoder
abstract
Reading scene text in the natural image is of fundamental importance in many real-world problems. Text recognition has a profound effect on information processing by enabling automated extraction and interpretation. Recent scene text recognition methods employ the encoder-decoder framework, which constructs the encoder by obtaining the visual representations based on the last layer of the backbone network and then feeding them into a sequence model. In this article, we propose a novel encoder structure that performs the feature extractor and the sequence modeling within a unified framework. The introduced Aggregated Temporal Convolutional Encoder (ATCE) first incorporates the temporal convolutional layers to consider the long-term temporal relationship in the encoder stage. The aggregation of these temporal convolution modules is designed to utilize visual features from different levels, by augmenting the standard architecture with deeper aggregation to better fuse information across modules. We also study the impact of different attention modules in convolutional blocks for learning accurate text representations. We conduct comparisons on several scene text recognition benchmarks for both Chinese and English; the experiments demonstrate the complementary ability with different decoder variants and the effectiveness of our proposed approach.
Tianlong Ma, Xiangcheng Du, Xingjiao Wu, Zhao Zhou, Yingbin Zheng, Cheng Jin 0001
ACM Trans. Asian Low Resour. Lang. Inf. Process.6
2022 Safe Distillation Box
abstract
Knowledge distillation (KD) has recently emerged as a powerful strategy to transfer knowledge from a pre-trained teacher model to a lightweight student, and has demonstrated its unprecedented success over a wide spectrum of applications. In spite of the encouraging results, the KD process \emph{per se} poses a potential threat to network ownership protection, since the knowledge contained in network can be effortlessly distilled and hence exposed to a malicious user. In this paper, we propose a novel framework, termed as Safe Distillation Box~(SDB), that allows us to wrap a pre-trained model in a virtual box for intellectual property protection. Specifically, SDB preserves the inference capability of the wrapped model to all users, but precludes KD from unauthorized users. For authorized users, on the other hand, SDB carries out a knowledge augmentation scheme to strengthen the KD performances and the results of the student model. In other words, all users may employ a model in SDB for inference, but only authorized users get access to KD from the model. The proposed SDB imposes no constraints over the model architecture, and may readily serve as a plug-and-play solution to protect the ownership of a pre-trained network. Experiments across various datasets and architectures demonstrate that, with SDB, the performance of an unauthorized KD drops significantly while that of an authorized gets enhanced, demonstrating the effectiveness of SDB.
Jingwen Ye, Yining Mao, Jie Song 0011, Xinchao Wang, Cheng Jin 0001, Mingli Song
AAAI5
2022 Attention-Based Transformation from Latent Features to Point Clouds
abstract
In point cloud generation and completion, previous methods for transforming latent features to point clouds are generally based on fully connected layers (FC-based) or folding operations (Folding-based). However, point clouds generated by FC-based methods are usually troubled by outliers and rough surfaces. For folding-based methods, their data flow is large, convergence speed is slow, and they are also hard to handle the generation of non-smooth surfaces. In this work, we propose AXform, an attention-based method to transform latent features to point clouds. AXform first generates points in an interim space, using a fully connected layer. These interim points are then aggregated to generate the target point cloud. AXform takes both parameter sharing and data flow into account, which makes it has fewer outliers, fewer network parameters, and a faster convergence speed. The points generated by AXform do not have the strong 2-manifold constraint, which improves the generation of non-smooth surfaces. When AXform is expanded to multiple branches for local generations, the centripetal constraint makes it has properties of self-clustering and space consistency, which further enables unsupervised semantic segmentation. We also adopt this scheme and design AXformNet for point cloud completion. Considerable experiments on different datasets show that our methods achieve state-of-the-art results.
Kaiyi Zhang 0002, Ximing Yang, Yuan Wu 0004, Cheng Jin 0001
AAAI4
2022 ST2PE: Spatial and Temporal Transformer for Pose Estimation
Yuan Wu 0004, Yanlu Cai, Rui Feng 0001, Cheng Jin 0001
ICANN (2)4
2022 GLTA-GCN: Global-Local Temporal Attention Graph Convolutional Network for Unsupervised Skeleton-Based Action Recognition
abstract
Unsupervised skeleton-based action recognition has attracted increasing attention. Existing methods have several limitations: (1) Many actions are highly related to local joints, which is often neglected. (2) Most methods directly employ joint coordinates as frame feature and do not utilize skeleton graph, e.g., topological information. (3) Long-range dependency is not captured well. In this work, a novel unsupervised method called Global-Local Temporal Attention Graph Convolutional Network (GLTA-GCN) is proposed to alleviate the above problems. The network consists of two branches, local and global branches. Each one utilizes graph convolution units and self-attention mechanism to better extract spatio-temporal features. Furthermore, two loss functions are designed to constrain the model to extract more essential local joint feature and maintain intrinsic structural information. Extensive experiments demonstrate that GLTA-GCN achieves state-of-the-art performance. Our code is released on https://github.com/HaoyueQiu/GLTA-GCN.
Haoyue Qiu, Yuan Wu 0004, Mengmeng Duan, Cheng Jin 0001
ICME4
2022 Point Cloud Completion via Multi-Scale Edge Convolution and Attention
abstract
Point cloud completion aims to recover a complete shape of a 3D object from its partial observation. Existing methods usually predict complete shapes from global representations, consequently, local geometric details may be ignored. Furthermore, they tend to overlook relations among different local regions, which are valuable during shape inference. To solve these problems, we propose a novel point cloud completion network based on multi-scale edge convolution and attention mechanism, named MEAPCN. We represent a point cloud as a set of embedded points, each of which contains geometric information of local patches around it. Firstly, we devise an encoder to extract multi-scale local features of the input point cloud and produce partial embedded points. Then, we generate coarse complete embedded points to represent the overall shape. In order to enrich features of complete embedded points, attention mechanism is utilized to selectively aggregate local informative features of partial ones. Lastly, we recover a fine-grained point cloud with highly detailed geometries using folding-based strategy. To better reflect real-world occlusion scenarios, we contribute a more challenging dataset, which consists of view-occluded partial point clouds. Experimental results on various benchmarks demonstrate that our method achieves a superior completion performance with much smaller model size and much lower computation cost.
Kaiyi Zhang 0002, Ximing Yang, Cheng Jin 0001
ACM Multimedia5
2022 Memory Enhanced Spatial-Temporal Graph Convolutional Autoencoder for Human-Related Video Anomaly Detection
Sibo Luo, Shangshang Wang, Yuan Wu 0004, Cheng Jin 0001
PRCV (3)4
2022 Weakly-Supervised Temporal Action Localization with Multi-Head Cross-Modal Attention
Hao Ren 0002, Wu Ran, Hong Lu 0001, Cheng Jin 0001
PRICAI (3)5
2021 CPCGAN: A Controllable 3D Point Cloud Generative Adversarial Network with Semantic Label Generating
abstract
Generative Adversarial Networks (GAN) are good at generating variant samples of complex data distributions. Generating a sample with certain properties is one of the major tasks in the real-world application of GANs. In this paper, we propose a novel generative adversarial network to generate 3D point clouds from random latent codes, named Controllable Point Cloud Generative Adversarial Network(CPCGAN). A two-stage GAN framework is utilized in CPCGAN and a sparse point cloud containing major structural information is extracted as the middle-level information between the two stages. With their help, CPCGAN has the ability to control the generated structure and generate 3D point clouds with semantic labels for points. Experimental results demonstrate that the proposed CPCGAN outperforms state-of-the-art point cloud GANs.
Ximing Yang, Yuan Wu 0004, Kaiyi Zhang 0002, Cheng Jin 0001
AAAI4
2020 Look into Multi-Person: A New Benchmark for Pose Estimation and Human Parsing
abstract
Human parsing and pose estimation, regarded as two fundamental tasks to analyze human in the wild, are the basis of upper-level tasks, such as human action recognition and person re-identification. The lack of a comprehensive multi-person dataset, which contains the annotations of both human part labels and skeleton keypoint labels, makes that most of the joint learning work for pose estimation and human parsing can only focus on the single-person scene. To fill this gap, we proposed a comprehensive multi-person dataset Look into Multi-Person (LIMP) with 10,082 multi-person images. And we adopt data augmentation strategies to enrich the dataset's diversity which contains more texture, more occlusion, and more in-the-wild images than CIHP. To the best of our knowledge, this is the first comprehensive multi-person dataset for pose estimation and human parsing.
Yanlu Cai, Runyu Peng, Yipei Xu, Chenzhe Jin, Cheng Jin 0001
IEEE BigData6
2020 One-sample Guided Object Representation Disassembling
abstract
The ability to disassemble the features of objects and background is crucial for many machine learning tasks, including image classification, image editing, visual concepts learning, and so on. However, existing (semi-)supervised methods all need a large amount of annotated samples, while unsupervised methods can't handle real-world images with complicated backgrounds. In this paper, we introduce the One-sample Guided Object Representation Disassembling (One-GORD) method, which only requires one annotated sample for each object category to learn disassembled object representation from unannotated images. For the annotated one-sample, we first adopt some data augmentation strategies to generate some synthetic samples, which can guide the disassembling of the object features and background features. For the unannotated images, two self-supervised mechanisms: dual-swapping and fuzzy classification are introduced to disassemble object features from the background with the guidance of annotated one-sample. What's more, we devise two metrics to evaluate the disassembling performance from the perspective of representation and image, respectively. Experiments demonstrate that the One-GORD achieves competitive dissembling performance and can handle natural scenes with complicated backgrounds.
Zunlei Feng, Yongming He, Xinchao Wang, Xin Gao 0032, Jie Lei 0002, Cheng Jin 0001, Mingli Song
NeurIPS6
2019 Fully Convolutional Video Captioning with Coarse-to-Fine and Inherited Attention
abstract
Automatically generating natural language description for video is an extremely complicated and challenging task. To tackle the obstacles of traditional LSTM-based model for video captioning, we propose a novel architecture to generate the optimal descriptions for videos, which focuses on constructing a new network structure that can generate sentences superior to the basic model with LSTM, and establishing special attention mechanisms that can provide more useful visual information for caption generation. This scheme discards the traditional LSTM, and exploits the fully convolutional network with coarse-to-fine and inherited attention designed according to the characteristics of fully convolutional structure. Our model cannot only outperform the basic LSTM-based model, but also achieve the comparable performance with those of state-of-the-art methods
Kuncheng Fang, Lian Zhou, Cheng Jin 0001, Yuejie Zhang, Kangnian Weng, Tao Zhang 0022, Weiguo Fan
AAAI3
2018 Sketch-based image retrieval with deep visual semantic descriptor
Cheng Jin 0001, Yuejie Zhang, Kangnian Weng, Tao Zhang 0022, Weiguo Fan
Pattern Recognit.2
2017 Towards sketch-based image retrieval with deep cross-modal correlation learning
abstract
A novel scheme with deep cross-modal correlation learning is developed in this paper to facilitate more effective Sketch-based Image Retrieval (SBIR) for large-scale annotated images. It integrates the deep multimodal feature generation, deep cross-modal correlation learning and similarity search optimization through mining all the beneficial multimodal information sources in sketches and images, which can be treated as an inter-related correlation distribution over deep representations of sketches and images. Very positive results were obtained in our experiments using a large quantity of public data.
Cheng Jin 0001, Yuejie Zhang, Tao Zhang 0022
ICME2
2017 A Hierarchical Multimodal Attention-based Neural Network for Image Captioning
abstract
A novel hierarchical multimodal attention-based model is developed in this paper to generate more accurate and descriptive captions for images. Our model is an "end-to-end" neural network which contains three related sub-networks: a deep convolutional neural network to encode image contents, a recurrent neural network to identify the objects in images sequentially, and a multimodal attention-based recurrent neural network to generate image captions. The main contribution of our work is that the hierarchical structure and multimodal attention mechanism is both applied, thus each caption word can be generated with the multimodal attention on the intermediate semantic objects and the global visual content. Our experiments on two benchmark datasets have obtained very positive results.
Lian Zhou, Cheng Jin 0001, Yuejie Zhang, Tao Zhang 0022
SIGIR4
2017 Deep Multimodal Embedding Model for Fine-grained Sketch-based Image Retrieval
abstract
Fine-grained Sketch-based Image Retrieval (Fine-grained SBIR), which uses hand-drawn sketches to search the target object images, has been an emerging topic over the last few years. The difficulties of this task not only come from the ambiguous and abstract characteristics of sketches with less useful information, but also the cross-modal gap at both visual and semantic level. However, images on the web are always exhibited with multimodal contents. In this paper, we consider Fine-grained SBIR as a cross-modal retrieval problem and propose a deep multimodal embedding model that exploits all the beneficial multimodal information sources in sketches and images. In our experiment with large quantity of public data, we show that the proposed method outperforms the state-of-the-art methods for Fine-grained SBIR.
Cheng Jin 0001, Yuejie Zhang, Tao Zhang 0022
SIGIR3
2016 A Novel Cross-Modal Topic Correlation Model for Cross-Media Retrieval
abstract
A novel cross-modal topic correlation model CMTCM is developed in this paper to facilitate more effective cross-modal analysis and cross-media retrieval for large-scale multimodal document collections. It can be modeled as a cross-modal topic correlation model which explores the inter-related correlation distribution over the deep representations of multimodal documents. It integrates the deep multimodal document representation, relational topic correlation modeling, and cross-modal topic correlation learning, which aims to characterize the correlations between the heterogeneous topic distributions of inter-related visual images and semantic texts, and measure their association degree more precisely. Very positive results were obtained in our experiments using a large quantity of public data.
Cheng Jin 0001, Yuejie Zhang, Tao Zhang 0022
ECAI3
2016 Enhancing Sketch-Based Image Retrieval via Deep Discriminative Representation
abstract
In this paper we aim to employ deep learning to enhance SBIR via deep discriminative representation. Our main contributions focus on: 1) The deep discriminative representation is established to bridge both the visual appearance gap and the semantic gap between sketches and images; 2) The deep learning pattern is applied to our SBIR model through training on our transformed sketch-like images to overcome the rarity of training sketches. Our experiments on a large number of public sketch and image data have obtained very positive results.
Cheng Jin 0001, Yuejie Zhang, Tao Zhang 0022
ECAI3
2016 Sketch-Based Image Retrieval with a Novel BoVW Representation
Cheng Jin 0001, Chenjie Li, Zheming Wang, Yuejie Zhang, Tao Zhang 0022
MMM (1)1
2015 Cross-Modal Image Clustering via Canonical Correlation Analysis
abstract
A new algorithm via Canonical Correlation Analysis (CCA) is developed in this paper to support more effective cross-modal image clustering for large-scale annotated image collections. It can be treated as a bi-media multimodal mapping problem and modeled as a correlation distribution over multimodal feature representations. It integrates the multimodal feature generation with the Locality Linear Coding (LLC) and co-occurrence association network, multimodal feature fusion with CCA, and accelerated hierarchical k-means clustering, which aims to characterize the correlations between the inter-related visual features in images and semantic features in captions, and measure their association degree more precisely. Very positive results were obtained in our experiments using a large quantity of public data.
Cheng Jin 0001, Wenhui Mao, Yuejie Zhang, Xiangyang Xue 0001
AAAI1
2015 People News Search via Name-Face Association Analysis
abstract
By integrating multimodal information in multimodal news, a novel scheme is developed in this paper for facilitating more effective people news search via name-face association analysis. It is treated as a problem of bi-media multimodal semantic mapping on multimodal news, and modeled as an inter-related correlation distribution over multimodal semantic representations of name-face associations. Very positive results have been obtained in our experiments using a large quantity of public multimodal news data.
Cheng Jin 0001, Yuejie Zhang, Tao Zhang 0022
ICMR4
2015 A Novel Visual-Region-Descriptor-based Approach to Sketch-based Image Retrieval
abstract
A novel Visual-Region-Descriptor-based approach is developed in this paper to facilitate more effective Sketch-based Image Retrieval (SBIR), which can be treated as a problem of bilateral visual mapping and modeled as an inter-related correlation distribution over visual semantic representations of sketches and images. For crossing the matching barrier between binary query sketches and full color natural images, we focus on constructing a visual pre-analysis via the sketch-like representation transformation to improve the general sketch-image resemblance, creating a special visual region descriptor to obtain better visual feature generation for sketches and images, and a dynamic sketch-image matching scheme to achieve more precise characterization of the correlations between sketches and images. Such a visual-region-descriptor-based SBIR pattern can not only enable users to present whatever they imagine in their mind on the sketch query panel but also return the most similar images to the picture in users' mind. Very positive results were obtained in our experiments using a large quantity of public data.
Cheng Jin 0001, Zheming Wang, Qinen Zhu, Yuejie Zhang
ICMR1
2015 Cross-Modal Image-Tag Relevance Learning for Social Images
abstract
A new algorithm is developed in this paper to support more effective cross-modal image-tag relevance learning for large-scale social images, which integrates the multimodal feature representation, multimodal relevance measurement, and cross- modal relevance fusion. The main contribution of our work is that we provide a more reasonable base to learn cross-modal relevance among social images, which can be acquired from integrating multimodal image and tag relevance with multiple features in different modalities. Very positive results were obtained in our experiments using a large quantity of public social image data.
Zhengxiang Cai, Rui Feng 0001, Cheng Jin 0001, Yuejie Zhang, Tao Zhang 0022
ACM Multimedia4
2014 Sketch-Based Image Retrieval via Adaptive Weighting
abstract
As touch devices become more and more popular these days, it would be convenient if the user could draw a sketch and then use the sketch as the input for an image retrieval system. Although Sketch-Based Image Retrieval (SBIR) had been studied since 1990s, how to measure the similarity between a sketch and an image with high precision is still a challenging problem. In this paper, a novel adaptive weighting method is proposed for the matching process of SBIR. We integrate a cost aggregation step into the matching process, both the neighborhood and multi-scale information are taken into account. The experiments on the public image dataset show that our method can yield promising results.
Cheng Jin 0001, Yuejie Zhang
ICMR2
2013 Automatic Name-Face Alignment to Enable Cross-Media News Retrieval
Yuejie Zhang, Cheng Jin 0001, Xiangyang Xue 0001, Jianping Fan 0001
IJCAI4
2012 Learning attention map from images
abstract
While bottom-up and top-down processes have shown effectiveness during predicting attention and eye fixation maps on images, in this paper, inspired by the perceptual organization mechanism before attention selection, we propose to utilize figure-ground maps for the purpose. So as to take both pixel-wise and region-wise interactions into consideration when predicting label probabilities for each pixel, we develop a context-aware model based on multiple segmentation to obtain final results. The MIT attention dataset [14] is applied finally to evaluate both new features and model. Quantitative experiments demonstrate that figure-ground cues are valid in predicting attention selection, and our proposed model produces improvements over baseline method.
Yao Lu 0028, Wei Zhang 0016, Cheng Jin 0001, Xiangyang Xue 0001
CVPR3
2012 Groupwise Constrained Reconstruction for Subspace Clustering
Ruijiang Li, Bin Li 0015, Cheng Jin 0001, Xiangyang Xue 0001
ICML3
2012 A simplified multi-class support vector machine with reduced dual optimization
Xisheng He, Zhe Wang 0002, Cheng Jin 0001, Yingbin Zheng, Xiangyang Xue 0001
Pattern Recognit. Lett.3
2011 Tracking User-Preference Varying Speed in Collaborative Filtering
abstract
In real-world recommender systems, some users are easily influenced by new products and whereas others are unwilling to change their minds. So the preference varying speeds for users are different. Based on this observation, we propose a dynamic nonlinear matrix factorization model for collaborative filtering, aimed to improve the rating prediction performance as well as track the preference varying speeds for different users. We assume that user-preference changes smoothly over time, and the preference varying speeds for users are different. These two assumptions are incorporated into the proposed model as prior knowledge on user feature vectors, which can be learned efficiently by MAP estimation. The experimental results show that our method not only achieves state-of-the-art performance in the rating prediction task, but also provides an effective way to track user-preference varying speed.
Ruijiang Li, Bin Li 0015, Cheng Jin 0001, Xiangyang Xue 0001, Xingquan Zhu 0001
AAAI3
2011 Learning Inter-Related Statistical Query Translation Models for English-Chinese Bi-Directional CLIR
abstract
To support more precise query translation for English-Chinese Bi-Directional Cross-Language Information Retrieval (CLIR), we have developed a novel framework by integrating a semantic network to characterize the correlations between multiple inter-related text terms of interest and learn their inter-related statistical query translation models. First, a semantic network is automatically generated from large-scale English-Chinese bilingual parallel corpora to characterize the correlations between a large number of text terms of interest. Second, the semantic network is exploited to learn the statistical query translation models for such text terms of interest. Finally, these inter-related query translation models are used to translate the queries more precisely and achieve more effective CLIR. Our experiments on a large number of official public data have obtained very positive results.
Yuejie Zhang, Lei Cen, Cheng Jin 0001, Xiangyang Xue 0001, Jianping Fan 0001
IJCAI3
2011 Fusion of Multiple Features and Supervised Learning for Chinese OOV Term Detection and POS Guessing
abstract
In this paper, to support more precise Chinese Out-of-Vocabulary (OOV) term detection and Part-of-Speech (POS) guessing, a unified mechanism is proposed and formulated based on the fusion of multiple features and supervised learning. Besides all the traditional features, the new features for statistical information and global contexts are introduced, as well as some constraints and heuristic rules, which reveal the relationships among OOV term candidates. Our experiments on the Chinese corpora from both People’s Daily and SIGHAN 2005 have achieved the consistent results, which are better than those acquired by pure rule-based or statistics-based models. From the experimental results for combining our model with Chinese monolingual retrieval on the data sets of TREC-9, it is found that the obvious improvement for the retrieval performance can also be obtained.
Yuejie Zhang, Lei Cen, Cheng Jin 0001, Xiangyang Xue 0001
IJCAI4
2011 Integrating hierarchical feature selection and classifier training for multi-label image annotation
abstract
It is well accepted that using high-dimensional multi-modal visual features for image content representation and classifier training may achieve more sufficient characterization of the diverse visual properties of the images and further result in higher discrimination power of the classifiers. However, training the classifiers in a high-dimensional multi-modal feature space requires a large number of labeled training images, which will further result in the problem of curse of dimensionality. To tackle this problem, a hierarchical feature subset selection algorithm is proposed to enable more accurate image classification, where the processes for feature selection and classifier training are seamlessly integrated in a single framework. First, a feature hierarchy (i.e., concept tree for automatic feature space partition and organization) is used to automatically partition high-dimensional heterogeneous multi-modal visual features into multiple low-dimensional homogeneous single-modal feature subsets according to their certain physical meanings and each of them is used to characterize one certain type of the diverse visual properties of the images. Second, principal component analysis (PCA) is performed on each homogeneous singlemodal feature subset to select the most representative feature dimensions and a weak classifier is learned simultaneously. After the weak classifiers and their representative feature dimensions are available for all these homogeneous single-modal feature subsets, they are combined to generate an ensemble image classifier and achieve hierarchical feature subset selection. Our experiments on a specific domain of natural images have also obtained very positive results.
Cheng Jin 0001, Chunlei Yang
SIGIR1
2010 Quick matting: A matting method based on pixel spread and propagation
abstract
The problem of matting is always solved by finding the alpha value for each pixel in the image. Many recent methods combine color sampling and affinity definition in different steps, leading to large computational cost. In the proposed method, when the alpha value of a pixel Piis calculated, the pixel is regarded as a foreground pixel to help calculate its adjacent pixels' alpha values, resulted in a faster solution. This spreading way of traversal also ensures local continuity of foreground object and improves the visual result. Experiments show our Quick Matting can achieve comparable alpha mattes as Robust Matting, while the speed is enhanced by about 25 times.
Yiyang Gu, Cheng Jin 0001, Xiangyang Xue 0001
ICIP2
2010 How context helps: A discriminative codeword selection method for object detection
abstract
We first propose in this paper to localize objects in images based on the models learned from the weakly labeled images. This task is termed as region of interest (ROI) detection. Local features such as SIFT or HOG are extracted and the discriminative words from clustered codewords based on SIFT and HOG are selected to model the objects. Then how to find the discriminative words to model the object is important. Existing ROI detection methods consider the information from the foreground objects by selecting the words appearing more in the images belonging to one specific image class. Considering the information from background/context is also helpful for object detection and classification, we propose to select the discriminative words which appear more in the foreground/object and less in the background/context. Second, another task is to give the class label (object in this setting) for a given image and also give the position of the object appearing in the image. This task is termed as objection detection. A normal way for this task after ROI is to extract features from the detected regions and not from the whole image. Since the discriminative words extracted during ROI detection has good discriminative ability, we propose to use these words for object detection. Experimental results on PASCAL VOC 2006 dataset and a larger dataset containing 29 classes demonstrate the effectiveness of the proposed method.
Renzhong Wei, Hong Lu 0001, Yingbin Zheng, Lei Cen, Cheng Jin 0001, Xiangyang Xue 0001, Weiguo Wu
ICIP5
2010 Bilingual query translation and expansion for supporting more effective cross-language image retrieval
abstract
To support more effective Cross-Language Image Retrieval (ImageCLIR), a novel algorithm is developed by integrating a bilingual semantic network to achieve more precise bilingual query translation and expansion. An English-Chinese bilingual parallel corpus is used to construct the bilingual semantic network for determining more meaningful text terms and characterizing the inter-term correlations and similarity contexts between multiple inter-related text terms more precisely. Our experiments on CWMT2009 and CLEF have provided very promising results.
Yuejie Zhang, Lei Cen, Cheng Jin 0001, Xiangyang Xue 0001
ACM Multimedia3
2009 Incorporating Spatial Correlogram into Bag-of-Features Model for Scene Categorization
Yingbin Zheng, Hong Lu 0001, Cheng Jin 0001, Xiangyang Xue 0001
ACCV (1)3
2008 An Improved Generalized Discriminant Analysis for Large-Scale Data Set
abstract
In order to overcome the computation and storage problem for large-scale data set, an efficient iterative method of generalized discriminant analysis is proposed. Because sample vectors cannot explicitly be denoted in kernel space, some mathematical tricks are firstly used to transform the kernel matrix. Then, the columns of transformed matrix are used for iterative algorithm to extract nonlinear discriminant vectors. The proposed method reduces space complexity from O(m2) to O(m) and its effectiveness is validated from experimental results.
Weiya Shi, Yue-Fei Guo, Cheng Jin 0001, Xiangyang Xue 0001
ICMLA3
2005 Sketch Based Facial Expression Recognition Using Graphics Hardware
Jiajun Bu, Mingli Song, Qi Wu 0014, Chun Chen 0001, Cheng Jin 0001
ACII5
2003 A new approach of blind image restoration
abstract
Blurring is inevitable in many image acquisition systems. Blind image restoration is one of deblur methods, and because it does not need a priori point spread function (PSF), it becomes more and more popular. In this paper, a novel method of blind image restoration is presented which can be applied to various kinds of images. The method has no limitation on blurred images, and does not need any extra database supports. Gaussian Blur is assumed to be able to present every other blur, and by successfully dealing with Gaussian Blur, our method can be suitable to deal with other blur methods. In order to reduce the noise, a pre-processing step is introduced. With edge detection, point spread function can be reconstructed automatically. Then an iterative algorithm is used to restore the image. Experimental results show the good performance of this method.
Cheng Jin 0001, Chun Chen 0001, Jiajun Bu
SMC1