EDBT 2026 Demo / reviewers in the wild / expert
Jianlong Chang
dblp:92/2332
· DBLP profile ↗
38ranked-venue papers
6as first author
26since 2021 · last 2026
0000-0002-0610-907XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 28 · 6 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 1 first-author · 17 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LookFlow: Training-Free and Efficient High-Resolution Image Synthesis via Dynamic Lookahead Guidance FlowabstractRectification flow Transformers (RFTs) have shown promising performance in diffusion-based image synthesis but are typically confined to lower-resolution scenarios, limiting their ability to generate high-resolution images. Existing resolution extrapolation approaches often suffer from excessive computational overhead, resulting in prolonged inference times. We propose LookFlow, a training-free high-resolution synthesis framework that accelerates inference while preserving visual quality. Building on pretrained text-to-image RFTs, LookFlow employs a dynamic lookahead guidance flow mechanism to refine high-resolution velocity predictions by leveraging multi-timestep lookahead information extracted from a low-resolution flow. Additionally, reusing temporally similar features across consecutive timesteps drastically reduces computation and significantly decreases inference time overhead. Extensive experiments on COCO demonstrate that LookFlow robustly scales resolutions from 4× to 25×, achieving up to a maximum speedup of 2.01× while maintaining competitive visual fidelity. Jianlong Chang, Ying Wang 0008, Kun Ding 0001, Shiming Xiang |
AAAI | 3 |
| 2026 | Beyond Counting: Evaluating Abstract and Emotional Reasoning in Vision-Language ModelsabstractDespite the rapid progress of Vision Language Models (VLMs), existing benchmarks still concentrate on coarse-grained object recognition or simple relational reasoning, leaving the fine-grained and higher-order reasoning abilities of these systems largely unexamined. To bridge this critical evaluation gap, we introduce EmojiGrid, a novel diagnostic benchmark specifically designed to probe these fine-grained and higher-order skills. Leveraging the universal and semantically rich nature of emojis, we synthesize a grid‑based visual dataset paired with 29,000+ QA pairs. Each pair is explicitly anchored in a three-level cognitive taxonomy comprising (i) Perception and Information Extraction, (ii) Relational and Structural Reasoning, and (iii) Abstraction and Advanced Cognition. These dimensions further decompose into nine categories covering a broad range of cognitive skills, including counting, spatial relations, compositional logic, semantic sentiment, and related higher-order reasoning tasks. Our extensive evaluation of 25 state-of-the-art open-source and proprietary VLMs reveals a significant performance gap between foundational perceptual tasks and higher-level cognitive abilities, particularly in abstraction and advanced emotional reasoning. Notably, all models struggle with compositional logic, spatial consistency, and especially emotional and semantic understanding. EmojiGrid provides a quantifiable, fine-grained benchmark to diagnose VLM limitations and guides future progress toward models that can truly perceive, reason about, and interpret complex, symbol-rich visual scenes. Jianlong Chang, Ying Wang 0008, Kun Ding 0001, Shiming Xiang |
AAAI | 3 |
| 2025 | Make "V" and "Q" Inseparable: Deliberately Dual-Channel Adversarial Learning for Robust Visual Question AnsweringabstractVisual Question Answering (VQA) is a challenging task due to the vision-language biases which restrict the model to sufficiently learn the multi-modal knowledge from visual image and natural language simultaneously. Several recent works attempt to alleviate this problem via weakening language prior but ignore vision prior, hindering further performance improvement. In this paper, we propose a novel Deliberately Dual-Channel Adversarial Learning (DCAL) to make "V" and "Q" inseparable, which aims to weaken prior from both vision and language. Specifically, DCAL introduces in-batch random negative sampling to force the model to be wrong when given the wrong questions or images. DCAL maximizes the likelihood of correct answers for the original question-image pairs and minimizes it for random negative samples. In order to solve the problem of false negatives, DCAL exploits a deliberate training strategy to utilize the sampled question-image pairs. The proposed DCAL is model-agnostic and can be applied to various VQA models. Experiments demonstrate that our proposed DCAL method improves the performance of existing robust VQA models on the sensitive VQA-CP dataset while performing robustly on the balanced VQA v2 dataset. Hanxiao Wu, Zhaowen Li, Liquan Hu, Huaixuan Cao, Jinqiao Wang, Jianlong Chang |
IJCNN | 10 |
| 2025 | Practical incremental learning: Striving for better performance-efficiency trade-off
Shixiong Xu, Bolin Ni, Xing Nie, Fei Zhu 0004, Jianlong Chang, Gaofeng Meng |
Neurocomputing | 6 |
| 2024 | LION: Implicit Vision Prompt TuningabstractDespite recent promising performances across a range of vision tasks, vision Transformers still have an issue of high computational costs. Recently, vision prompt learning has provided an economical solution to this problem without fine-tuning the whole large-scale model. However, the efficiency and effectiveness of existing models are still far from satisfactory due to the parameter cost of extensive prompt blocks and tricky prompt framework designs. In this paper, we propose a light-weight prompt framework named impLicit vIsion prOmpt tuNing (LION), which is motivated by deep implicit models with stable low memory costs for various complex tasks. In particular, we merely insect two equilibrium implicit layers in two ends of the pre-trained backbone with parameters frozen. Moreover, according to the lottery hypothesis, we further prune the parameters to relieve the computation burden in implicit layers. Various experiments have validated that our LION obtains promising performances on a wide range of datasets. Most importantly, LION reduces up to 11.5 % of training parameter numbers while obtaining higher performance than the state-of-the-art VPT, especially under challenging scenes. Furthermore, we find that our proposed LION has an excellent generalization performance, making it an easy way to boost transfer learning in the future. Haixin Wang 0003, Jianlong Chang, Yihang Zhai, Xiao Luo 0001, Jinan Sun, Zhouchen Lin, Qi Tian 0001 |
AAAI | 2 |
| 2024 | Accurate Fine-Grained Object Recognition with Structure-Driven Relation Graph Networks
Shijie Wang 0003, Zhihui Wang 0001, Jianlong Chang, Wanli Ouyang, Qi Tian 0001 |
Int. J. Comput. Vis. | 4 |
| 2024 | Content-Aware Rectified Activation for Zero-Shot Fine-Grained Image RetrievalabstractFine-grained image retrieval mainly focuses on learning salient features from the seen subcategories as discriminative embedding while neglecting the problems behind zero-shot settings. We argue that retrieving fine-grained objects from unseen subcategories may rely on more diverse clues, which are easily restrained by the salient features learnt from seen subcategories. To address this issue, we propose a novel Content-aware Rectified Activation model, which enables this model to suppress the activation on salient regions while preserving their discrimination, and spread activation to adjacent non-salient regions, thus mining more diverse discriminative features for retrieving unseen subcategories. Specifically, we construct a content-aware rectified prototype (CARP) by perceiving semantics of salient regions. CARP acts as a channel-wise non-destructive activation upper bound and can be selectively used to suppress salient regions for obtaining the rectified features. Moreover, two regularizations are proposed: 1) a semantic coherency constraint that imposes a restriction on semantic coherency of CARP and salient regions, aiming at propagating the discriminative ability of salient regions to CARP, 2) a feature-navigated constraint to further guide the model to adaptively balance the discrimination power of rectified features and the suppression power of salient features. Experimental results on fine-grained and product retrieval benchmarks demonstrate that our method consistently outperforms the state-of-the-art methods. Shijie Wang 0003, Jianlong Chang, Zhihui Wang 0001, Wanli Ouyang, Qi Tian 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Pro-Tuning: Unified Prompt Tuning for Vision TasksabstractIn computer vision, fine-tuning is the de-facto approach to leverage pre-trained vision models to perform downstream tasks. However, deploying it in practice is quite challenging, due to adopting parameter inefficient global update and heavily relying on high-quality downstream data. Recently, prompt-based learning, which adds the task-relevant prompt to adapt the pre-trained models to downstream tasks, has drastically boosted the performance of many natural language downstream tasks. In this work, we extend this notable transfer ability benefited from prompt into vision models as an alternative to fine-tuning. To this end, we propose parameter-efficient Prompt tuning (Pro-tuning) to adapt diverse frozen pre-trained models to a wide variety of downstream vision tasks. The key to Pro-tuning is prompt-based tuning, i.e., learning task-specific vision prompts for downstream input images with the pre-trained model frozen. By only training a small number of additional parameters, Pro-tuning can generate compact and robust downstream models both for CNN-based and transformer-based network architectures. Comprehensive experiments evidence that the proposed Pro-tuning outperforms fine-tuning on a broad range of vision tasks and scenarios, including image classification (under generic objects, class imbalance, image corruption, natural adversarial examples, and out-of-distribution generalization), and dense prediction tasks such as object detection and semantic segmentation. Xing Nie, Bolin Ni, Jianlong Chang, Gaofeng Meng, Chunlei Huo, Shiming Xiang, Qi Tian 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | KNLConv: Kernel-Space Non-Local Convolution for Hyperspectral Image Super-ResolutionabstractPixel-level adaptive convolution, which overcomes the deficiency of the spatial-invariance of standard convolution, is always limited to performing feature extraction from local patches and ignores the latent long-range dependencies imperceptible in the feature space, which are more significant in pixel-level tasks such as hyperspectral image super-resolution (HSISR). To handle such limitations, we propose kernel-space non-local convolution (KNLConv), which explores non-local dependencies in the generated kernel space, to leverage these global information to guide the network to extract image features more flexibly. Technically, the proposed KNLConv first decomposes the convolutional kernel space into spatial and channel dimensions, and designs a depth-wise non-local expansion convolution (NLEC) in the spatial dimension of the kernel-space to explore underlying global correlations. Then introduce an adaptive point-wise convolution (APC), generalizing the NLEC to the pixel-level while integrating features in the channel dimension. In addition, applying KNLConv, we design an effective network architecture for hyperspectral image super-resolution. Extensive experiments demonstrate that our approach performs favorably against current state-of-the-art HSISR methods, both on quantitative indicators and visual quality. Ran Ran 0001, Liang-Jian Deng, Tianjing Zhang, Jianlong Chang, Qi Tian 0001 |
IEEE Trans. Multim. | 4 |
| 2024 | Structure Aware Multi-Graph Network for Multi-Modal Emotion Recognition in ConversationsabstractMulti-Modal Emotion Recognition in Conversations (MMERC) is an increasingly active research field that leverages multi-modal signals to understand the feelings behind each utterance. Modeling contextual interactions and multi-modal fusion lie at the heart of this field, with graph-based models recently being widely used for MMERC to capture global multi-modal contextual information. However, these models generally mix all modality representations in a single graph, and utterances in each modality are fully connected, potentially ignoring three problems: 1) the heterogeneity of the multi-modal context, 2) the redundancy of contextual information, and 3) over-smoothing of the graph networks. To address these problems, we propose a Structure Aware Multi-Graph Network (SAMGN) for MMERC. Specifically, we construct multiple modality-specific graphs to model the heterogeneity of the multi-modal context. Instead of fully connecting the utterances in each modality, we design a structure learning module that determines whether edges exist between the utterances. This module reduces redundancy by forcing each utterance to focus on the contextual ones that contribute to its emotion recognition, acting like a message propagating reducer to alleviate over-smoothing. Then, we develop the SAMGN via Dual-Stream Propagation (DSP), which contains two propagation streams, i.e., intra- and inter-modal, performed in parallel to aggregate the heterogeneous modality information from multi-graphs. DSP also contains a gating unit that adaptively integrates the co-occurrence information from the above two propagations for emotion recognition. Experiments on two popular MMERC datasets demonstrate that SAMGN achieves new State-Of-The-Art (SOTA) results. Duzhen Zhang, Jianlong Chang, Xiuyi Chen, Qi Tian 0001 |
IEEE Trans. Multim. | 3 |
| 2023 | Fine-Grained Retrieval Prompt TuningabstractFine-grained object retrieval aims to learn discriminative representation to retrieve visually similar objects. However, existing top-performing works usually impose pairwise similarities on the semantic embedding spaces or design a localization sub-network to continually fine-tune the entire model in limited data scenarios, thus resulting in convergence to suboptimal solutions. In this paper, we develop Fine-grained Retrieval Prompt Tuning (FRPT), which steers a frozen pre-trained model to perform the fine-grained retrieval task from the perspectives of sample prompting and feature adaptation. Specifically, FRPT only needs to learn fewer parameters in the prompt and adaptation instead of fine-tuning the entire model, thus solving the issue of convergence to suboptimal solutions caused by fine-tuning the entire model. Technically, a discriminative perturbation prompt (DPP) is introduced and deemed as a sample prompting process, which amplifies and even exaggerates some discriminative elements contributing to category prediction via a content-aware inhomogeneous sampling operation. In this way, DPP can make the fine-grained retrieval task aided by the perturbation prompts close to the solved task during the original pre-training. Thereby, it preserves the generalization and discrimination of representation extracted from input samples. Besides, a category-specific awareness head is proposed and regarded as feature adaptation, which removes the species discrepancies in features extracted by the pre-trained model using category-guided instance normalization. And thus, it makes the optimized features only include the discrepancies among subcategories. Extensive experiments demonstrate that our FRPT with fewer learnable parameters achieves the state-of-the-art performance on three widely-used fine-grained datasets. Shijie Wang 0003, Jianlong Chang, Zhihui Wang 0001, Wanli Ouyang, Qi Tian 0001 |
AAAI | 2 |
| 2023 | Distilling Vision-Language Pre-Training to Collaborate with Weakly-Supervised Temporal Action LocalizationabstractWeakly-supervised temporal action localization (WTAL) learns to detect and classify action instances with only category labels. Most methods widely adopt the off-the-shelf Classification-Based Pre-training (CBP) to generate video features for action localization. However, the different optimization objectives between classification and localization, make temporally localized results suffer from the serious incomplete issue. To tackle this issue without additional annotations, this paper considers to distill free action knowledge from Vision-Language Pre-training (VLP), as we surprisingly observe that the localization results of vanilla VLP have an over-complete issue, which is just complementary to the CBP results. To fuse such complementarity, we propose a novel distillation-collaboration framework with two branches acting as CBP and VLP respectively. The framework is optimized through a dual-branch alternate training strategy. Specifically, during the B step, we distill the confident background pseudo-labels from the CBP branch; while during the F step, the confident foreground pseudo-labels are distilled from the VLP branch. As a result, the dualbranch complementarity is effectively fused to promote one strong alliance. Extensive experiments and ablation studies on THUMOS14 and ActivityNet1.2 reveal that our method significantly outperforms state-of-the-art methods. Chen Ju, Kunhao Zheng, Jinxiang Liu, Peisen Zhao, Ya Zhang 0002, Jianlong Chang, Qi Tian 0001, Yanfeng Wang 0001 |
CVPR | 6 |
| 2023 | Being Comes from Not-Being: Open-Vocabulary Text-to-Motion Generation with Wordless TrainingabstractText-to-motion generation is an emerging and challenging problem, which aims to synthesize motion with the same semantics as the input text. However, due to the lack of diverse labeled training data, most approaches either limit to specific types of text annotations or require online optimizations to cater to the texts during inference at the cost of efficiency and stability. In this paper, we investigate offline open-vocabulary text-to-motion generation in a zero-shot learning manner that neither requires paired training data nor extra online optimization to adapt for unseen texts. Inspired by the prompt learning in NLP, we pretrain a motion generator that learns to reconstruct the full motion from the masked motion. During inference, instead of changing the motion generator, our method reformulates the input text into a masked motion as the prompt for the motion generator to “reconstruct” the motion. In constructing the prompt, the unmasked poses of the prompt are synthesized by a text-to-pose generator. To supervise the optimization of the text-to-pose generator, we propose the first text-pose alignment model for measuring the alignment between texts and 3D poses. And to prevent the pose generator from over-fitting to limited training texts, we further propose a novel wordless training mechanism that optimizes the text-to-pose generator without any training texts. The comprehensive experimental results show that our method obtains a significant improvement against the baseline methods. The code is available at https://github.com/junfanlin/oohmg. Junfan Lin, Jianlong Chang, Lingbo Liu, Guanbin Li, Liang Lin 0004, Qi Tian 0001, Chang Wen Chen |
CVPR | 2 |
| 2023 | Open-Set Fine-Grained Retrieval via Prompting Vision-Language EvaluatorabstractOpen-set fine-grained retrieval is an emerging challenge that requires an extra capability to retrieve unknown subcategories during evaluation. However, current works focus on close-set visual concepts, where all the subcategories are pre-defined, and make it hard to capture discriminative knowledge from unknown subcategories, consequently failing to handle unknown subcategories in open-world scenarios. In this work, we propose a novel Prompting vision-Language Evaluator (PLEor) framework based on the recently introduced contrastive language-image pretraining (CLIP) model, for open-set fine-grained retrieval. PLEor could leverage pre-trained CLIP model to infer the discrepancies encompassing both pre-defined and unknown subcategories, called category-specific discrepancies, and transfer them to the backbone network trained in the close-set scenarios. To make pre-trained CLIP model sensitive to category-specific discrepancies, we design a dual prompt scheme to learn a vision prompt specifying the categoryspecific discrepancies, and turn random vectors with category names in a text prompt into category-specific discrepancy descriptions. Moreover, a vision-language evaluator is proposed to semantically align the vision and text prompts based on CLIP model, and reinforce each other. In addition, we propose an open-set knowledge transfer to transfer the category-specific discrepancies into the backbone network using knowledge distillation mechanism. Quantitative and qualitative experiments show that our PLEor achieves promising performance on open-set fine-grained datasets. Shijie Wang 0003, Jianlong Chang, Zhihui Wang 0001, Wanli Ouyang, Qi Tian 0001 |
CVPR | 2 |
| 2023 | Parameter-efficient Tuning of Large-scale Multimodal Foundation ModelabstractDriven by the progress of large-scale pre-training, parameter-efficient transfer learning has gained immense popularity across different subfields of Artificial Intelligence. The core is to adapt the model to downstream tasks with only a small set of parameters. Recently, researchers have leveraged such proven techniques in multimodal tasks and achieve promising results. However, two critical issues remain unresolved: how to further reduce the complexity with lightweight design and how to boost alignment between modalities under extremely low parameters. In this paper, we propose A gracefUl pRompt framewOrk for cRoss-modal trAnsfer (AURORA) to overcome these challenges. Considering the redundancy in existing architectures, we first utilize the mode approximation to generate 0.1M trainable parameters to implement the multimodal parameter-efficient tuning, which explores the low intrinsic dimension with only 0.04% parameters of the pre-trained model. Then, for better modality alignment, we propose the Informative Context Enhancement and Gated Query Transformation module under extremely few parameters scenes. A thorough evaluation on six cross-modal benchmarks shows that it not only outperforms the state-of-the-art but even outperforms the full fine-tuning approach. Our code is available at: https://github.com/WillDreamer/Aurora. Haixin Wang 0003, Xinlong Yang, Jianlong Chang, Dian Jin 0004, Jinan Sun, Shikun Zhang, Xiao Luo 0001, Qi Tian 0001 |
NeurIPS | 3 |
| 2023 | Learning to Parameterize Visual Attributes for Open-set Fine-grained RetrievalabstractOpen-set fine-grained retrieval is an emerging challenging task that allows to retrieve unknown categories beyond the training set.
The best solution for handling unknown categories is to represent them using a set of visual attributes learnt from known categories, as widely used in zero-shot learning. Though important, attribute modeling usually requires significant manual annotations and thus is labor-intensive. Therefore, it is worth to investigate how to transform retrieval models trained by image-level supervision from category semantic extraction to attribute modeling. To this end, we propose a novel Visual Attribute Parameterization Network (VAPNet) to learn visual attributes from known categories and parameterize them into the retrieval model, without the involvement of any attribute annotations.
In this way, VAPNet could utilize its parameters to parse a set of visual attributes from unknown categories and precisely represent them.
Technically, VAPNet explicitly attains some semantics with rich details via making use of local image patches and distills the visual attributes from these discovered semantics. Additionally, it integrates the online refinement of these visual attributes into the training process to iteratively enhance their quality. Simultaneously, VAPNet treats these attributes as supervisory signals to tune the retrieval models, thereby achieving attribute parameterization. Extensive experiments on open-set fine-grained retrieval datasets validate the superior performance of our VAPNet over existing solutions. Shijie Wang 0003, Jianlong Chang, Zhihui Wang 0001, Wanli Ouyang, Qi Tian 0001 |
NeurIPS | 2 |
| 2023 | Semantic-Guided Information Alignment Network for Fine-Grained Image RecognitionabstractExisting fine-grained image recognition works have attempted to dig into low-level details for emphasizing subtle discrepancies among sub-categories. However, a potential limitation of these methods is that they integrate the low-level details and high-level semantics directly, and neglect their content complementarity and spatial corresponding correlation. To handle this limitation, we propose an end-to-end Semantic-guided Information Alignment Network (SIA-Net) to dynamically pick out the low-level details under the guidance of accurate semantics to make selected details spatially corresponding to high-level semantics and complementary in content. Technically, SIA-Net consists of an Accurate Semantic Calibration (ASC) module for providing accurate semantics and a Discriminative Feature Alignment (DFA) module for aggregating low-level details and high-level semantics using accurate semantics generated by ASC. ASC learns the pixel-level feature shifting caused by convolutional operations, which is utilized for replacing the incorrectly highlighted semantics by shifting discriminative semantics or background features. After obtaining the accurate semantic features, DFA digs into the complementary details and simultaneously makes the selected details spatially corresponding via applying the guidance of accurate semantics to obtain the reassembly features. Finally, the reassembly features, which serve as discriminative cues, are used for more accurate discriminative region localization. Extensive experiments verify that our proposed method yields the best performance under the same settings with the most competitive approaches on CUB-birds, Stanford-Cars, and FGVC Aircraft datasets. Shijie Wang 0003, Zhihui Wang 0001, Jianlong Chang, Wanli Ouyang, Qi Tian 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Efficient and Scalable Implicit Graph Neural Networks with Virtual EquilibriumabstractOn large-scale graphs, many graph neural networks are problematic in capturing long-range dependencies due to the oversmoothing problem. Recently, Graph Equilibrium Models (GEQs) arise as a promising solution to this issue. Their output is the equilibrium of a fixed-point equation, which can be seen as the result of iterating a GNN layer for infinite times, so that they inherently have global receptive fields. However, to find the equilibrium, GEQs require running costly full-batch root-finding algorithms from scratch during each model update, which leads to severe efficiency and scalability issues that prevent them from scaling to large graphs. To address these limitations, we propose VEQ, an efficient learning method to scale GEQs to large graphs. Instead of initializing the equilibrium from scratch in full-batch training, VEQ uses the latest equilibrium of in-batch nodes and their 1-hop neighbors (dubbed Virtual Equilibrium) to accelerate and calibrate the root-finding process in mini-batch training. With virtual equilibrium as an informative prior, VEQ is able to reach the equilibrium in fewer steps while still capturing global dependencies. Theoretically, we provide convergence analysis for the forward and backward pass of VEQ. Empirically, VEQ significantly outperforms existing GEQs by a large margin (more than 1.5%) on all benchmark datasets, with much less training time and memory. Also, VEQ achieves competitive and even superior performance to many highly engineered explicit GNNs on large-scale benchmark datasets like ogbn-arxiv and ogbn-products. VEQ shows that after we resolve the efficiency and scalability issues, GEQs are indeed favorable on large graphs due to their advantage of capturing long-range dependencies. Yifei Wang 0001, Yisen Wang 0001, Jianlong Chang, Qi Tian 0001, Jiansheng Yang, Zhouchen Lin |
IEEE Big Data | 4 |
| 2022 | AME: Attention and Memory Enhancement in Hyper-Parameter OptimizationabstractTraining Deep Neural Networks (DNNs) is inherently subject to sensitive hyper-parameters and untimely feedbacks of performance evaluation. To solve these two difficulties, an efficient parallel hyper-parameter optimization model is proposed under the framework of Deep Reinforcement Learning (DRL). Technically, we develop Attention and Memory Enhancement (AME), that includes multi-head attention and memory mechanism to enhance the ability to capture both the short-term and long-term relationships between different hyper-parameter configurations, yielding an attentive sampling mechanism for searching high-performance configurations embedded into a huge search space. During the optimization of transformer-structured configuration searcher, a conceptually intuitive yet powerful strategy is applied to solve the problem of insufficient number of samples due to the untimely feedback. Experiments on three visual tasks, including image classification, object detection, semantic segmentation, demonstrate the effectiveness of AME. Nuo Xu 0006, Jianlong Chang, Xing Nie, Chunlei Huo, Shiming Xiang, Chunhong Pan |
CVPR | 2 |
| 2022 | SSL++: Improving Self-Supervised Learning by Mitigating the Proxy Task-Specificity ProblemabstractThe success of deep convolutional networks (ConvNets) generally relies on a massive amount of well-labeled data, which is labor-intensive and time-consuming to collect and annotate in many scenarios. To eliminate such limitation, self-supervised learning (SSL) is recently proposed. Specifically, by solving a pre-designed proxy task, SSL is capable of capturing general-purpose features without requiring human supervision. Existing efforts focus obsessively on designing a particular proxy task but ignore the semanticity of samples that are advantageous to downstream tasks, resulting in the inherent limitation that the learned features are specific to the proxy task, namely the proxy task-specificity of features. In this work, to improve the generalizability of features learned by existing SSL methods, we present a novel self-supervised framework SSL++ to incorporate the proxy task-independent semanticity of samples into the representation learning process. Technically, SSL++ aims to leverage the complementarity, between the low-level generic features learned by a proxy task and the high-level semantic features newly learned by the generated semantic pseudo-labels, to mitigate the task-specificity and improve the generalizability of features. Extensive experiments show that SSL++ performs favorably against the state-of-the-art approaches on the established and latest SSL benchmarks. Jing-Hao Xue, Jianlong Chang, Jianzhong Zhang 0003, Jufeng Yang, Qi Tian 0001 |
IEEE Trans. Image Process. | 3 |
| 2022 | MS-Net: Multi-Source Spatio-Temporal Network for Traffic Flow PredictionabstractPredicting urban traffic flow is a challenging task, due to the complicated spatio-temporal dependencies on traffic networks. Urban traffic flow usually has both short-term neighboring and long-term periodic temporal dependencies. It is also noticed that the spatial correlations over different traffic nodes are both local and non-local. What’s more, the traffic flow is affected by various external factors. To capture the non-local spatial correlations, we propose a Dilated Attentional Graph Convolution (DAGC). The DAGC utilizes a dilated graph convolution kernel to expand the nodes’ receptive field and exploit multi-order neighborhood. Technically, the lower-order neighborhood corresponds to local spatial dependencies, while the higher-order neighborhood corresponds to non-local spatial dependencies between nodes. Based on DAGC, a Multi-Source Spatio-Temporal Network (MS-Net) is designed, which suffices to integrate long-range historical traffic data as well as multi-modal external information. MS-Net consists of four components: a spatial feature extraction module, a temporal feature fusion module, an external factors embedding module, and a multi-source data fusion module. Extensive experiments on three real traffic datasets demonstrates that the proposed model performs well on both the public transportation networks, road networks, and can handle large-scale traffic networks in particular the Beijing bus network which has more than 4,000 traffic nodes. Shen Fang, Véronique Prinet, Jianlong Chang, Michael Werman, Chunxia Zhang 0001, Shiming Xiang, Chunhong Pan |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | Decoupled Representation Learning for Character Glyph SynthesisabstractCharacter glyph synthesis is still an open challenging problem, which involves two related aspects,i.e., font style transfer and content consistency. In this paper, we propose a novel model named FontGAN, which integrates the character structure stylization, de-stylization and texture transfer into a unified framework. Specifically, we decouple character images into style representation and content representation, which offers fine-grained control of these two types of variables, thus improving the quality of the generated results. To effectively capture the style information, a style consistency module (SCM) is introduced. Technically, SCM exploits category-guided Kullback-Leibler divergence to explicitly model the style representation into different prior distributions. In this way, our model is capable of implementing transformations between multiple domains in one framework. In addition, we propose content prior module (CPM) to provide content prior for the model to guide the content encoding process and alleviates the problem of stroke deficiency during structure de-stylization. Benefiting from the idea of decoupling and regrouping, our FontGAN suffices to achieve many-to-many translation tasks for glyph structure. Experimental results demonstrate that the proposed FontGAN achieves the state-of-the-art performance in character glyph synthesis. Xiyan Liu, Gaofeng Meng, Jianlong Chang, Ruiguang Hu, Shiming Xiang, Chunhong Pan |
IEEE Trans. Multim. | 3 |
| 2021 | Camera-Space Hand Mesh Recovery via Semantic Aggregation and Adaptive 2D-1D RegistrationabstractRecent years have witnessed significant progress in 3D hand mesh recovery. Nevertheless, because of the intrinsic 2D-to-3D ambiguity, recovering camera-space 3D information from a single RGB image remains challenging. To tackle this problem, we divide camera-space mesh recovery into two sub-tasks, i.e., root-relative mesh recovery and root recovery. First, joint landmarks and silhouette are extracted from a single input image to provide 2D cues for the 3D tasks. In the root-relative mesh recovery task, we exploit semantic relations among joints to generate a 3D mesh from the extracted 2D cues. Such generated 3D mesh coordinates are expressed relative to a root position, i.e., wrist of the hand. In the root recovery task, the root position is registered to the camera space by aligning the generated 3D mesh back to 2D cues, thereby completing camera-space 3D mesh recovery. Our pipeline is novel in that (1) it explicitly makes use of known semantic relations among joints and (2) it exploits 1D projections of the silhouette and mesh to achieve robust registration. Extensive experiments on popular datasets such as FreiHAND, RHD, and Human3.6M demonstrate that our approach achieves state-of-the-art performance on both root-relative mesh recovery and root recovery. Our code is publicly available at https://github.com/SeanChenxy/HandMesh. Chongyang Ma, Jianlong Chang, Huayan Wang, Pengfei Wan 0001 |
CVPR | 4 |
| 2021 | Differentiable Convolution Search for Point Cloud ProcessingabstractExploiting convolutional neural networks for point cloud processing is quite challenging, due to the inherent irregular distribution and discrete shape representation of point clouds. To address these problems, many handcrafted convolution variants have sprung up in recent years. Though with elaborate design, these variants could be far from optimal in sufficiently capturing diverse shapes formed by discrete points. In this paper, we propose PointSeaConv, i.e., a novel differential convolution search paradigm on point clouds. It can work in a purely data-driven manner and thus is capable of auto-creating a group of suitable convolutions for geometric shape modeling. We also propose a joint optimization framework for simultaneous search of internal convolution and external architecture, and introduce epsilon-greedy algorithm to alleviate the effect of discretization error. As a result, PointSeaNet, a deep network that is sufficient to capture geometric shapes at both convolution level and architecture level, can be searched out for point cloud processing. Extensive experiments strongly evidence that our proposed PointSeaNet surpasses current handcrafted deep models on challenging benchmarks across multiple tasks with remarkable margins. Xing Nie, Yongcheng Liu, Shaohong Chen, Jianlong Chang, Chunlei Huo, Gaofeng Meng, Qi Tian 0001, Chunhong Pan |
ICCV | 4 |
| 2021 | DATA: Differentiable ArchiTecture Approximation With Distribution Guided SamplingabstractNeural architecture search (NAS) is inherently subject to the gap of architectures during searching and validating. To bridge this gap effectively, we develop Differentiable ArchiTecture Approximation (DATA) with Ensemble Gumbel-Softmax (EGS) estimator and Architecture Distribution Constraint (ADC) to automatically approximate architectures during searching and validating in a differentiable manner. Technically, the EGS estimator consists of a group of Gumbel-Softmax estimators, which is capable of converting probability vectors to binary codes and passing gradients reversely, reducing the estimation bias in a differentiable way. To narrow the distribution gap between sampled architectures and supernet, further, the ADC is introduced to reduce the variance of sampling during searching. Benefiting from such modeling, architecture probabilities and network weights in the NAS model can be jointly optimized with the standard back-propagation, yielding an end-to-end learning mechanism for searching deep neural architectures in an extended search space. Conclusively, in the validating process, a high-performance architecture that approaches to the learned one during searching is readily built. Extensive experiments on various tasks including image classification, few-shot learning, unsupervised clustering, semantic segmentation and language modeling strongly demonstrate that DATA is capable of discovering high-performance architectures while guaranteeing the required efficiency. Code is available at https://github.com/XinbangZhang/DATA-NAS. Xinbang Zhang, Jianlong Chang, Yiwen Guo, Gaofeng Meng, Shiming Xiang, Zhouchen Lin, Chunhong Pan |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | BiSPL: Bidirectional Self-Paced Learning for Recognition From Web DataabstractDeep learning (DL) is inherently subject to the requirement of a large amount of well-labeled data, which is expensive and time-consuming to obtain manually. In order to broaden the reach of DL, leveraging free web data becomes an attractive strategy to alleviate the issue of data scarcity. However, directly utilizing collected web data to train a deep model is ineffective because of the mixed noisy data. To address such problems, we develop a novel bidirectional self-paced learning (BiSPL) framework which reduces the effect of noise by learning from web data in a meaningful order. Technically, the BiSPL framework consists of two essential steps. Relying on distances defined between web samples and labeled source samples, first, the web samples with short distances are sampled and combined to form a new training set. Second, based on the new training set, both easy and hard samples are initially employed to train deep models for higher stability, and hard samples are gradually dropped to reduce the noise as the training progresses. By iteratively alternating such steps, deep models converge to a better solution. We mainly focus on the fine-grained visual classification (FGVC) tasks because their corresponding datasets are generally small and therefore face a more significant data scarcity problem. Experiments conducted on six public FGVC tasks demonstrate that our proposed method outperforms the state-of-the-art approaches. Especially, BiSPL suffices to achieve the highest stable performance when the scale of the well-labeled training set decreases dramatically. Jianlong Chang, Yukun Lai, Jufeng Yang, Qi Tian 0001 |
IEEE Trans. Image Process. | 2 |
| 2020 | Cross-Modality Paired-Images Generation for RGB-Infrared Person Re-IdentificationabstractRGB-Infrared (IR) person re-identification is very challenging due to the large cross-modality variations between RGB and IR images. The key solution is to learn aligned features to the bridge RGB and IR modalities. However, due to the lack of correspondence labels between every pair of RGB and IR images, most methods try to alleviate the variations with set-level alignment by reducing the distance between the entire RGB and IR sets. However, this set-level alignment may lead to misalignment of some instances, which limits the performance for RGB-IR Re-ID. Different from existing methods, in this paper, we propose to generate cross-modality paired-images and perform both global set-level and fine-grained instance-level alignments. Our proposed method enjoys several merits. First, our method can perform set-level alignment by disentangling modality-specific and modality-invariant features. Compared with conventional methods, ours can explicitly remove the modality-specific features and the modality variation can be better reduced. Second, given cross-modality unpaired-images of a person, our method can generate cross-modality paired images from exchanged images. With them, we can directly perform instance-level alignment by minimizing distances of every pair of images. Extensive experimental results on two standard benchmarks demonstrate that the proposed model favourably against state-of-the-art methods. Especially, on SYSU-MM01 dataset, our model can achieve a gain of 9.2% and 7.7% in terms of Rank-1 and mAP. Code is available at https://github.com/wangguanan/JSIA-ReID. Guan'an Wang, Tianzhu Zhang 0001, Yang Yang 0062, Jian Cheng 0001, Jianlong Chang, Zeng-Guang Hou |
AAAI | 5 |
| 2020 | Spatio-Temporal Graph Structure Learning for Traffic ForecastingabstractAs an indispensable part in Intelligent Traffic System (ITS), the task of traffic forecasting inherently subjects to the following three challenging aspects. First, traffic data are physically associated with road networks, and thus should be formatted as traffic graphs rather than regular grid-like tensors. Second, traffic data render strong spatial dependence, which implies that the nodes in the traffic graphs usually have complex and dynamic relationships between each other. Third, traffic data demonstrate strong temporal dependence, which is crucial for traffic time series modeling. To address these issues, we propose a novel framework named Structure Learning Convolution (SLC) that enables to extend the traditional convolutional neural network (CNN) to graph domains and learn the graph structure for traffic forecasting. Technically, SLC explicitly models the structure information into the convolutional operation. Under this framework, various non-Euclidean CNN methods can be considered as particular instances of our formulation, yielding a flexible mechanism for learning on the graph. Along this technical line, two SLC modules are proposed to capture the global and local structures respectively and they are integrated to construct an end-to-end network for traffic forecasting. Additionally, in this process, Pseudo three Dimensional convolution (P3D) networks are combined with SLC to capture the temporal dependencies in traffic data. Extensively comparative experiments on six real-world datasets demonstrate our proposed approach significantly outperforms the state-of-the-art ones. Jianlong Chang, Gaofeng Meng, Shiming Xiang, Chunhong Pan |
AAAI | 2 |
| 2020 | Deep Self-Evolution ClusteringabstractClustering is a crucial but challenging task in pattern analysis and machine learning. Existing methods often ignore the combination between representation learning and clustering. To tackle this problem, we reconsider the clustering task from its definition to develop Deep Self-Evolution Clustering (DSEC) to jointly learn representations and cluster data. For this purpose, the clustering task is recast as a binary pairwise-classification problem to estimate whether pairwise patterns are similar. Specifically, similarities between pairwise patterns are defined by the dot product between indicator features which are generated by a deep neural network (DNN). To learn informative representations for clustering, clustering constraints are imposed on the indicator features to represent specific concepts with specific representations. Since the ground-truth similarities are unavailable in clustering, an alternating iterative algorithm called Self-Evolution Clustering Training (SECT) is presented to select similar and dissimilar pairwise patterns and to train the DNN alternately. Consequently, the indicator features tend to be one-hot vectors and the patterns can be clustered by locating the largest response of the learned indicator features. Extensive experiments strongly evidence that DSEC outperforms current models on twelve popular image, text and audio datasets consistently. Jianlong Chang, Gaofeng Meng, Lingfeng Wang 0002, Shiming Xiang, Chunhong Pan |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2020 | Local-Aggregation Graph NetworksabstractConvolutional neural networks (CNNs) provide a dramatically powerful class of models, but are subject to traditional convolution that can merely aggregate permutation-ordered and dimension-equal local inputs. It causes that CNNs are allowed to only manage signals on Euclidean or grid-like domains (e.g., images), not ones on non-Euclidean or graph domains (e.g., traffic networks). To eliminate this limitation, we develop a local-aggregation function, a sharable nonlinear operation, to aggregate permutation-unordered and dimension-unequal local inputs on non-Euclidean domains. In the context of the function approximation theory, the local-aggregation function is parameterized with a group of orthonormal polynomials in an effective and efficient manner. By replacing the traditional convolution in CNNs with the parameterized local-aggregation function, Local-Aggregation Graph Networks (LAGNs) are readily established, which enable to fit nonlinear functions without activation functions and can be expediently trained with the standard back-propagation. Extensive experiments on various datasets strongly demonstrate the effectiveness and efficiency of LAGNs, leading to superior performance on numerous pattern recognition and machine learning tasks, including text categorization, molecular activity detection, taxi flow prediction, and image classification. Jianlong Chang, Lingfeng Wang 0002, Gaofeng Meng, Shiming Xiang, Chunhong Pan |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2019 | DATA: Differentiable ArchiTecture ApproximationabstractNeural architecture search (NAS) is inherently subject to the gap of architectures during searching and validating. To bridge this gap, we develop Differentiable ArchiTecture Approximation (DATA) with an Ensemble Gumbel-Softmax (EGS) estimator to automatically approximate architectures during searching and validating in a differentiable manner. Technically, the EGS estimator consists of a group of Gumbel-Softmax estimators, which is capable of converting probability vectors to binary codes and passing gradients from binary codes to probability vectors. Benefiting from such modeling, in searching, architecture parameters and network weights in the NAS model can be jointly optimized with the standard back-propagation, yielding an end-to-end learning mechanism for searching deep models in a large enough search space. Conclusively, during validating, a high-performance architecture that approaches to the learned one during searching is readily built. Extensive experiments on a variety of popular datasets strongly evidence that our method is capable of discovering high-performance architectures for image classification, language modeling and semantic segmentation, while guaranteeing the requisite efficiency during searching. Jianlong Chang, Xinbang Zhang, Yiwen Guo, Gaofeng Meng, Shiming Xiang, Chunhong Pan |
NeurIPS | 1 |
| 2019 | Learning graph structure via graph convolutional networks
Jianlong Chang, Gaofeng Meng, Shibiao Xu, Shiming Xiang, Chunhong Pan |
Pattern Recognit. | 2 |
| 2018 | Kernel-Weighted Graph Convolutional Network: A Deep Learning Approach for Traffic ForecastingabstractTraffic forecasting is of great significance and has many applications in Intelligent Traffic System (ITS). In spite of many thoughtful attempts in the past decades, this task still remains far from being solved, due to the diversity, complexity and nonlinearity of traffic situations. Technically, it can be cast on the framework of regressions with spatial-template data. Typically, one may consider to employ the Convolutional Neural Network (CNN) to achieve this goal. Unfortunately, the traditional CNN is developed for grid data. By contrast, here we are facing with non-grid traffic data points that are observed spatially at locations of interest. To this end, this paper proposes a novel Kernel-Weighted Graph Convolutional Network (KW-GCN) for traffic forecasting, which learns simultaneously a group of convolutional kernels and their linear combination weights for each of the nodes in the graph. This yields a mechanism that is able to learn the features locally and exploit the structure information of traffic road-network globally. By introducing additional parameters, our KW-GCN can relax the restriction of weight sharing in classical CNN to better handle the traffic data of non-stationarity. Furthermore, it has been illustrated that the proposed linear weighting of kernels can be viewed as the low-rank decomposition of the well-known locally-connected networks, and thus it avoids over-fitting to some degree. We apply our approach to the real-world GPS data set of about 30,000 taxis in seven months in Beijing. Experiments on both taxi-flow forecasting and road-speed forecasting demonstrate that our method significantly outperforms the state-of-the-art ones. Qizhao Jin, Jianlong Chang, Shiming Xiang, Chunhong Pan |
ICPR | 3 |
| 2018 | Structure-Aware Convolutional Neural NetworksabstractConvolutional neural networks (CNNs) are inherently subject to invariable filters that can only aggregate local inputs with the same topological structures. It causes that CNNs are allowed to manage data with Euclidean or grid-like structures (e.g., images), not ones with non-Euclidean or graph structures (e.g., traffic networks). To broaden the reach of CNNs, we develop structure-aware convolution to eliminate the invariance, yielding a unified mechanism of dealing with both Euclidean and non-Euclidean structured data. Technically, filters in the structure-aware convolution are generalized to univariate functions, which are capable of aggregating local inputs with diverse topological structures. Since infinite parameters are required to determine a univariate function, we parameterize these filters with numbered learnable parameters in the context of the function approximation theory. By replacing the classical convolution in CNNs with the structure-aware convolution, Structure-Aware Convolutional Neural Networks (SACNNs) are readily established. Extensive experiments on eleven datasets strongly evidence that SACNNs outperform current models on various machine learning tasks, including image classification and clustering, text categorization, skeleton-based action recognition, molecular activity detection, and taxi flow prediction. Jianlong Chang, Jie Gu 0002, Lingfeng Wang 0002, Gaofeng Meng, Shiming Xiang, Chunhong Pan |
NeurIPS | 1 |
| 2018 | Deep unsupervised learning with consistent inference of latent representations
Jianlong Chang, Lingfeng Wang 0002, Gaofeng Meng, Shiming Xiang, Chunhong Pan |
Pattern Recognit. | 1 |
| 2017 | Deep Adaptive Image ClusteringabstractImage clustering is a crucial but challenging task in machine learning and computer vision. Existing methods often ignore the combination between feature learning and clustering. To tackle this problem, we propose Deep Adaptive Clustering (DAC) that recasts the clustering problem into a binary pairwise-classification framework to judge whether pairs of images belong to the same clusters. In DAC, the similarities are calculated as the cosine distance between label features of images which are generated by a deep convolutional network (ConvNet). By introducing a constraint into DAC, the learned label features tend to be one-hot vectors that can be utilized for clustering images. The main challenge is that the ground-truth similarities are unknown in image clustering. We handle this issue by presenting an alternating iterative Adaptive Learning algorithm where each iteration alternately selects labeled samples and trains the ConvNet. Conclusively, images are automatically clustered based on the label features. Experimental results show that DAC achieves state-of-the-art performance on five popular datasets, e.g., yielding 97.75% clustering accuracy on MNIST, 52.18% on CIFAR-10 and 46.99% on STL-10. Jianlong Chang, Lingfeng Wang 0002, Gaofeng Meng, Shiming Xiang, Chunhong Pan |
ICCV | 1 |
| 2014 | On efficiently generating realistic social media timeline structuresabstractA framework of synthetic data generator to generate social media timeline structures is proposed in this paper, which is useful for benchmarking query processing over social media data, and validating hypothesis over users' behavior. It is flexible to generate synthetic data with different distributions. With the help of its asynchronized parallel processing model and delayed update strategy, it is efficient to feed out timeline structure with high throughput. We show in experiments that our method can generate realistic social media timeline structures efficiently. Chengcheng Yu, Weining Qian, Aoying Zhou, Jianlong Chang |
SSDBM | 5 |
| 2008 | SMART: A System for Online Monitoring Large Volumes of Network TrafficabstractNetwork traffic monitoring have been gaining attentions due to its importance in telecom industry. However, the monitoring systems deployed in telecom operators are usually too slow because of their disk-based processing approach. To address this problem, an online network traffic monitoring system, named SMART, is designed and developed. The system converts different formats of raw Netflow data (Netflow IPv5, IPv7 and IPv9) to user-defined control flows through combination and filtering. It can compute top-k frequent flows with sliding window, detects burst on arbitrary attributes, and presents results visually to users. The system could be used to replace the traditional offline monitoring system used in Shanghai Telecom. In its daily operation, it is shown that the processing speed achieves 30,000 flows per second. The basis of advanced streaming algorithms and design of robust system architecture enable SMART to achieve good performance. Aoying Zhou, Ying Yan 0002, Xueqing Gong, Jianlong Chang, Dai Dai |
ICDE | 4 |