EDBT 2026 Demo / reviewers in the wild / expert
Jiangtong Li
dblp:220/0990
· DBLP profile ↗
30ranked-venue papers
10as first author
27since 2021 · last 2026
0000-0003-3873-4053ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 7 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 5 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Targeting Borderline Fraudsters: Multi-View Hypergraph Fraud Detection with LLM-Guided Contrastive LearningabstractGraph fraud detection (GFD) on transaction networks is crucial for safeguarding financial systems. However, due to the limited perspective of existing graph neural networks (GNNs) in the single transaction view, sophisticated fraudsters can disguise themselves to exhibit weak fraud signals, appearing as borderline fraudsters. To address this challenge, we propose MH-LGC, a multi-view hypergraph fraud detection model with large language model (LLM) guided contrastive learning. MH-LGC tackles two key limitations of existing GNN-based GFD methods: (1) Due to the local aggregation mechanism, existing methods struggle to capture high-order trading patterns among distant fraudsters. MH-LGC introduces two temporal hyper-views as complements to the transaction view and employs a Temporal Hypergraph Attention Network (THAN) to integrate the three views. (2) Most GFD methods overlook the rich semantic cues embedded in transaction data. Although some general graph learning studies have explored LLM integration, the high computational overhead and task-specific fine-tuning make them impractical for GFD tasks. MH-LGC introduces a semantic view through a fine-tuning-free LLM-Guided Contrastive learning (LGC), adopting a novel paradigm for integrating GNN and LLM to reduce the computational overhead of LLM. Extensive experiments on three real-world datasets demonstrate that MH-LGC outperforms twelve state-of-the-art baselines, with AUC improvements ranging from 1.10% to 5.70%. Rui Ou, Kun Zhu 0024, Jiangtong Li, Chaochao Chen 0001, Yuhua Xu 0011, Changjun Jiang 0002 |
AAAI | 4 |
| 2026 | LEGO: Supporting LLM-Enhanced Games with One Gaming GPUabstractArtificial intelligence (AI) has been increasingly applied to gaming, with large language models (LLMs) playing a key role in character control. However, efficiently co-locating game rendering and LLM inference on one GPU presents challenges due to resource constraints, diverse latency requirements, and fine-grained task scheduling. We propose LEGO, an algorithm-system co-design that enables the efficient co-location of LLM inference and game rendering tasks. Algorithmwise, LEGO features a resource-oriented layer-skipping adaptor, which distills knowledge from skipped layers to reduce computational demand while maintaining inference accuracy. System-wise, LEGO proposes a headroom-maximizing LLM scheduler, which dynamically partitions inference tasks to utilize available rendering headroom. Evaluations on an Nvidia RTX 4090 show that LEGO meets latency targets in all scenarios, improves rendering headroom utilization by up to 28.6 %, and reduces LLM inference accuracy loss by up to 86.3 % compared to current layer-skipping approaches. Han Zhao 0005, Weihao Cui, Zeshen Zhang, Jiangtong Li, Quan Chen 0002, Pu Pang, Zijun Li 0001, Zhenhua Han, Yuqing Yang 0001, Minyi Guo |
HPCA | 5 |
| 2026 | Bridging Visual Dynamics and Narrative Reasoning: Multimodal Large Language Models for Short Drama Quality Assessment
Qingyang Liu 0008, Jiangtong Li, Zelin Peng, Shaobo Wang 0001, Zhaohe Liao, Shuochen Chang, Bingjie Gao, Mu Liu, Jidong Jiang, Li Niu 0002 |
WWW | 2 |
| 2026 | STG-DGR: Fraud Detection on Streaming Transaction Graphs with Diffusion-based Generative ReplayabstractFraud detection on streaming transaction graphs (STGs) faces challenges on the catastrophic forgetting of previously learned fraud patterns when adapting to evolving patterns. Although some Graph Continual Learning (GCL) approaches mitigate this issue by storing and revisiting historical samples, practical storage constraints prevent them from fully preserving previous patterns. In this work, we propose STG-DGR, a streaming GNN model with diffusion-based generative replay that generates synthetic samples to retain previously learned patterns without storing real samples. The generation of replay samples for STGs faces two key challenges: (1) Heterogeneity challenge of generating STG samples with discrete adjacency table, user features, transaction features, and transaction timestamps. (2) Dependency challenge of capturing bottom-up dependencies across layers in STG samples. To address these challenges, STG-DGR integrates two novel components: (1) a Computational Subgraph Processor (CSP) that transforms heterogeneous STG samples into well-organized hierarchical subgraphs, and (2) a Diffusion-based Subgraph Generator (DSG) that captures the bottom-up dependencies using a novel Transformer-based Hierarchical Denoising Network (THDN), and generates synthetic replay samples that preserve these dependencies. Extensive experiments on four streaming fraud detection datasets demonstrate STG-DGR's superiority in reducing forgetting and improving accuracy over nineteen state-of-the-art baselines. Rui Ou, Kun Zhu 0024, Jiangtong Li, Chaochao Chen 0001, Changjun Jiang 0002 |
WWW | 4 |
| 2026 | Category-Specific Trigger Backdoor Attacks on Graph Neural NetworksabstractGraph Neural Networks (GNNs) have achieved remarkable success in various applications, while still exhibiting high vulnerability to backdoor attacks when applied to node classification. Existing single-category attack methods typically rely on adaptive triggers that force victim nodes to be misclassified into a fixed target label, but they often neglect the inherent structural and feature priors associated with the target category. In this work, we propose a novel and effective backdoor attack framework Category-Specific Trigger Backdoor Attacks (CSTBA), employing category-specific information to generate more natural and unnoticeable triggers. Specifically, we introduce a Category-Specific Subgraph Triggers Pool (CS-STP) to capture representative patterns of the target category, along with a Match-and-Attach Strategy (MAS) to unnoticeably attach triggers to victim nodes, thereby ensuring that the modifications remain effective and unnoticeable within the graph. Extensive experiments on multiple benchmark datasets demonstrate that our method significantly enhances both the attack success rate (ASRs) and the unnoticeability of the attack compared with existing single-category approaches, highlighting the critical importance of category-aware trigger design in GNN backdoor attacks. Yiwen Jiang, Wensi Liu, Linbo Shao, Dongyi Liu, Jiangtong Li |
Int. J. Pattern Recognit. Artif. Intell. | 5 |
| 2026 | QMATE: Mitigating gradient conflict with query-enhanced multi-task adaptive transformer for unifying EEG representation
Junming Lin, Jiangtong Li, Jie Li 0049, Changjun Jiang 0002 |
Knowl. Based Syst. | 3 |
| 2026 | Parse, Align and Aggregate: Graph-Driven Compositional Reasoning for Video Question AnsweringabstractVideo Question-Answering (VideoQA) enables machines to interpret and respond to complex video content, advancing human-computer interaction. However, existing multimodal large language models (MLLMs) often provide incomplete or opaque explanations and existing benchmarks mainly focus on the correction of final answers, limiting insight into their reasoning processes and hindering both transparency and verifiability. To address this gap, we propose the Question Parsing, Video Alignment and Answer Aggregation framework (QPVA$^{3}$3), which leverages a compositional graph to drive visual and logical reasoning in VideoQA. Specifically, QPVA$^{3}$3 consists of three core components, the planner, executor, and reasoner to generate the compositional graph and conduct graph-driven reasoning. For the original question, the planner parses it into the compositional graph, capturing the underlying reasoning logic and structuring it into a series of interconnected questions. For each question in compositional graph, the executor aligns the video by selecting relevant video clips and generates answers, ensuring accurate, context-specific responses. For each question with its first-order descents, the reasoner aggregates answers by integrating reasoning logic with visual evidence, resolving conflicts to produce a coherent and accurate response. Moreover, to assess the performance of existing MLLMs in the reasoning processes of VideoQA, we introduce novel compositional consistency metrics and construct a VideoQA benchmark (QPVA$^{3}$3 Bench) with 3,492 question-video tuples, each annotated with detailed compositional graphs and fine-grained answers. We evaluate the QPVA$^{3}$3 framework on QPVA$^{3}$3 Bench and 5 other VideoQA benchmarks. Experimental results demonstrate that our framework improves both consistency and accuracy compared to baselines, leading to a more transparent and verifiable VideoQA system. This approach has the potential to advance the field, as supported by our comprehensive evaluation and benchmarking efforts. Jiangtong Li, Zhaohe Liao, Fengshun Xiao, Tianjiao Li 0001, Qiang Zhang 0055, Haohua Zhao 0001, Li Niu 0002, Guang Chen 0001, Liqing Zhang 0001, Changjun Jiang 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | InspireDebate: Multi-Dimensional Subjective-Objective Evaluation-Guided Reasoning and Optimization for DebatingabstractWith the rapid advancements in large language models (LLMs), debating tasks, such as argument quality assessment and debate process simulation, have made significant progress. However, existing LLM-based debating systems focus on responding to specific arguments while neglecting objective assessments such as authenticity and logical validity. Furthermore, these systems lack a structured approach to optimize across various dimensions—including evaluation metrics, chain-of-thought (CoT) reasoning, and multi-turn debate refinement—thereby limiting their effectiveness. To address these interconnected challenges, we propose a dual-component framework: (1) InspireScore, a novel evaluation system that establishes a multi-dimensional assessment architecture incorporating four subjective criteria (emotional appeal, argument clarity, argument arrangement, and topic relevance) alongside two objective metrics (fact authenticity and logical validity); and (2) InspireDebate, an optimized debating framework employing a phased optimization approach through CoT reasoning enhancement, multi-dimensional Direct Preference Optimization (DPO), and real-time knowledge grounding via web-based Retrieval Augmented Generation (Web-RAG). Empirical evaluations demonstrate that InspireScore achieves 44% higher correlation with expert judgments compared to existing methods, while InspireDebate shows significant improvements, outperforming baseline models by 57%. Source code is available at https://github.com/fywang12/InspireDebate. Fuyu Wang 0004, Jiangtong Li, Kun Zhu 0036, Changjun Jiang 0002 |
ACL (1) | 2 |
| 2025 | Rethinking Classifier Re-Training in Long-Tailed Recognition: Label Over-Smooth Can BalanceabstractIn the field of long-tailed recognition, the Decoupled Training paradigm has shown exceptional promise by dividing training into two stages: representation learning and classifier re-training. While previous work has tried to improve both stages simultaneously, this complicates isolating the effect of classifier re-training. Recent studies reveal that simple regularization can produce strong feature representations, highlighting the need to reassess classifier re-training methods. In this study, we revisit classifier re-training methods based on a unified feature representation and re-evaluate their performances.
We propose two new metrics, Logits Magnitude and Regularized Standard Deviation, to compare the differences and similarities between various methods.
Using these two newly proposed metrics, we demonstrate that when the Logits Magnitude across classes is nearly balanced, further reducing its overall value can effectively decrease errors and disturbances during training, leading to better model performance.
Based on our analysis using these metrics, we observe that adjusting the logits could improve model performance, leading us to develop a simple label over-smoothing approach to adjust the logits without requiring prior knowledge of class distribution.
This method softens the original one-hot labels by assigning a probability slightly higher than $\frac{1}{K}$ to the true class and slightly lower than $\frac{1}{K}$ to the other classes, where $K$ is the number of classes.
Our method achieves state-of-the-art performance on various imbalanced datasets, including CIFAR100-LT, ImageNet-LT, and iNaturalist2018. Han Lu 0004, Jiangtong Li, Yichen Xie 0002, Tianjiao Li 0001, Xiaokang Yang 0001, Liqing Zhang 0001, Junchi Yan |
ICLR | 3 |
| 2025 | Divide and Conquer: Exploring Language-centric Tree Reasoning for Video Question-AnsweringabstractVideo Question-Answering (VideoQA) remains challenging in achieving advanced cognitive reasoning due to the uncontrollable and opaque reasoning processes in existing Multimodal Large Language Models (MLLMs). To address this issue, we propose a novel Language-centric Tree Reasoning (LTR) framework that targets on enhancing the reasoning ability of models. In detail, it recursively divides the original question into logically manageable parts and conquers them piece by piece, enhancing the reasoning capabilities and interpretability of existing MLLMs. Specifically, in the first stage, the LTR focuses on language to recursively generate a language-centric logical tree, which gradually breaks down the complex cognitive question into simple perceptual ones and plans the reasoning path through a RAG-based few-shot approach. In the second stage, with the aid of video content, the LTR performs bottom-up logical reasoning within the tree to derive the final answer along with the traceable reasoning path. Experiments across 11 VideoQA benchmarks demonstrate that our LTR framework significantly improves both accuracy and interpretability compared to state-of-the-art MLLMs. To our knowledge, this is the first work to implement a language-centric logical tree to guide MLLM reasoning in VideoQA, paving the way for language-centric video understanding from perception to cognition. Zhaohe Liao, Jiangtong Li, Qingyang Liu 0002, Fengshun Xiao, Tianjiao Li 0001, Qiang Zhang 0055, Guang Chen 0001, Li Niu 0002, Changjun Jiang 0002, Liqing Zhang 0001 |
ICML | 2 |
| 2025 | Attack by Yourself: Effective and Unnoticeable Multi-Category Graph Backdoor Attacks with Subgraph Triggers PoolabstractGraph Neural Networks (GNNs) have achieved significant success in various real-world applications, including social networks, finance systems, and traffic management. Recent researches highlight their vulnerability to backdoor attacks in node classification, where GNNs trained on a poisoned graph misclassify a test node only when specific triggers are attached. These studies typically focus on single attack categories and use adaptive trigger generators to create node-specific triggers. However, adaptive trigger generators typically have a simple structure, limited parameters, and lack category-aware graph knowledge, which makes them struggle to handle backdoor attacks across multiple categories as the number of target categories increases. We address this gap by proposing a novel approach for Effective and Unnoticeable Multi-Category (EUMC) graph backdoor attacks, leveraging subgraph from the attacked graph as category-aware triggers to precisely control the target category. To ensure the effectiveness of our method, we construct a Multi-Category Subgraph Triggers Pool (MC-STP) using the subgraphs of the attacked graph as triggers. We then exploit the attachment probability shifts of each subgraph trigger as category-aware priors for target category determination. Moreover, we develop a ``select then attach'' strategy that connects suitable category-aware trigger to attacked nodes for unnoticeability. Extensive experiments across different real-world datasets confirm the efficacy of our method in conducting multi-category graph backdoor attacks on various GNN models and defense strategies. Jiangtong Li, Dongyi Liu, Kun Zhu 0036, Dawei Cheng, Changjun Jiang 0002 |
NeurIPS | 1 |
| 2025 | Multimodal LiDAR-Camera Novel View Synthesis with Unified Pose-free Neural FieldsabstractPose-free Neural Radiance Field (NeRF) aims at novel view synthesis (NVS) without relying on accurate poses, exhibiting significant practical value. Image and LiDAR point cloud are two pivotal modalities in autonomous driving scenarios. While demonstrating impressive performance, single-modality pose-free NeRFs often suffer from local optima due to the limited geometric information provided by dense image textures or the sparse, textureless nature of point clouds. Although prior methods have explored the complementary strengths of both modalities, they have only leveraged inherently sparse point clouds for discrete, non-pixel-wise depth supervision, and are limited to NVS of images. As a result, a Multimodal Unified Pose-free framework remains notably absent. In light of this, we propose MUP, a pose-free framework for LiDAR-Camera joint NVS in large-scale scenes. This unified framework enables continuous depth supervision for image reconstruction using LiDAR-Fields rather than discrete point clouds. By leveraging multimodal inputs, pose optimization receives gradients from the rendering loss of point cloud geometry and image texture, thereby alleviating the issue of local optima commonly encountered in single-modality pose-free tasks. Moreover, to further guide pose optimization of NeRF, we propose a multimodal geometric optimizer that leverages geometric relations from point clouds and photometric regularization from adjacent image frames. Besides, to alleviate the domain gap between modalities, we propose a multimodal-specific coarse-to-fine training approach for unified, compact reconstruction. Extensive experiments on KITTI-360 and NuScenes datasets demonstrate MUP's superiority in accomplishing geometry-aware, modality-consistent, and pose-free 3D reconstruction. Weiyi Xue, Fan Lu 0001, Yunwei Zhu, Zehan Zheng, Sanqing Qu, Jiangtong Li, Haiyun Wei, Guang Chen 0001 |
NeurIPS | 6 |
| 2025 | HFTCRNet: Hierarchical Fusion Transformer for Interbank Credit Rating and Risk AssessmentabstractAs a prominent application of deep neural networks in financial literature, bank credit ratings play a pivotal role in safeguarding global economic stability and preventing crises. In the contemporary financial system, interconnectivity among banks has reached unprecedented levels. However, many existing credit risk models continue to assess each bank independently, resulting in inevitable suboptimal performance. Thus, developing advanced neural networks to model intricate temporal dynamics and interconnected relationships in the banking system is essential for an effective credit rating and risk assessment learning system. To this end, we propose a novel hierarchical fusion transformer for interbank credit rating and risk assessment (HFTCRNet), which includes the long-term temporal transformer (LT3) module, short-term cross-graph transformer (STCGT) module, attentive risk contagion transformer (ARCT) module, and hierarchical fusion transformer (HFT) module to capture the long-term growth trajectories of banks, the short-term interbank network variance, the potential propagation of risks within interbank network, and integrate these information hierarchically. We further develop an interbank credit rating dataset, encompassing quarterly financial data, interbank lending networks, and key indicators such as credit ratings and systemic risk (SRISK) for 4548 banks from 2016Q1 to 2023Q1. Notably, we also adapt the minimum density algorithm to stabilize the interbank loan network over time, aiding in the analysis of long-term and short-term network effects. Our learning system uses semi-supervised learning to handle labels of varying sparsity, integrating credit ratings and SRISK for a comprehensive assessment of individual bank creditworthiness and systemic interbank risk. Extensive experimental results on our interbank dataset show that HFTCRNet not only outperforms all the baselines in terms of credit rating accuracy but also can evaluate the systemic risk within the interbank network. Code will be available at: https://github.com/AI4Risk/HFTCRNet. Jiangtong Li, Ziyuan Zhou 0004, Jingkai Zhang, Dawei Cheng, Changjun Jiang 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Align and Aggregate: Compositional Reasoning with Video Alignment and Answer Aggregation for Video Question-AnsweringabstractDespite the recent progress made in Video Question-Answering (VideoQA), these methods typically function as black-boxes, making it difficult to understand their reasoning processes and perform consistent compositional reasoning. To address these challenges, we propose a model-agnostic Video Alignment and Answer Aggregation (VA3) framework, which is capable of enhancing both compositional consistency and accuracy of existing VidQA methods by integrating video aligner and answer aggregator modules. The video aligner hierarchically selects the relevant video clips based on the question, while the answer ag-gregator deduces the answer to the question based on its sub-questions, with compositional consistency ensured by the information flow along question decomposition graph and the contrastive learning strategy. We evaluate our framework on three settings of the AGQA-Decomp dataset with three baseline methods, and propose new metrics to measure the compositional consistency of VidQA methods more comprehensively. Moreover, we propose a large language model (LLM) based automatic question decomposition pipeline to apply our framework to any VidQA dataset. We extend MSVD and NExT-QA datasets with it to evaluate our VA3framework on broader scenarios. Extensive experiments show that our framework improves both compositional consistency and accuracy of existing methods, leading to more interpretable real-world VidQA models. Zhaohe Liao, Jiangtong Li, Li Niu 0002, Liqing Zhang 0001 |
CVPR | 2 |
| 2024 | COIN-Matting: Confounder Intervention for Image Matting
Zhaohe Liao, Jiangtong Li, Jun Lan 0001, Huijia Zhu, Weiqiang Wang 0002, Li Niu 0002, Liqing Zhang 0001 |
ECCV (19) | 2 |
| 2024 | Multi-Patch Prediction: Adapting Language Models for Time Series Representation LearningabstractIn this study, we present $\text{aL\small{LM}4T\small{S}}$, an innovative framework that adapts Large Language Models (LLMs) for time-series representation learning. Central to our approach is that we reconceive time-series forecasting as a self-supervised, multi-patch prediction task, which, compared to traditional mask-and-reconstruction methods, captures temporal dynamics in patch representations more effectively. Our strategy encompasses two-stage training: (i). a causal continual pre-training phase on various time-series datasets, anchored on next patch prediction, effectively syncing LLM capabilities with the intricacies of time-series data; (ii). fine-tuning for multi-patch prediction in the targeted time-series context. A distinctive element of our framework is the patch-wise decoding layer, which departs from previous methods reliant on sequence-level decoding. Such a design directly transposes individual patches into temporal sequences, thereby significantly bolstering the model’s proficiency in mastering temporal patch-based representations. $\text{aL\small{LM}4T\small{S}}$ demonstrates superior performance in several downstream tasks, proving its effectiveness in deriving temporal representations with enhanced transferability and marking a pivotal advancement in the adaptation of LLMs for time-series analysis. Yuxuan Bian, Xuan Ju, Jiangtong Li, Dawei Cheng, Qiang Xu 0001 |
ICML | 3 |
| 2024 | RA-CFGPT: Chinese financial assistant with retrieval-augmented large language model
Jiangtong Li, Yuxuan Bian, Dawei Cheng, Zhijun Ding, Changjun Jiang 0002 |
Frontiers Comput. Sci. | 1 |
| 2023 | Knowledge Proxy Intervention for Deconfounded Video Question AnsweringabstractRecently, Video Question-Answering (VideoQA) has drawn more and more attention from both the industry and the research community. Despite all the success achieved by recent works, dataset bias always harmfully misleads current methods focusing on spurious correlations in training data. To analyze the effects of dataset bias, we frame the VideoQA pipeline into a causal graph, which shows the causalities among video, question, aligned feature between video and question, answer, and underlying confounder. Through the causal graph, we prove that the confounder and the backdoor path lead to spurious causality. To tackle the challenge that the confounder in VideoQA is unobserved and non-enumerable in general, we propose a model-agnostic framework called Knowledge Proxy Intervention (KPI), which introduces an extra knowledge proxy variable in the causal graph to cut the backdoor path and remove the effect of confounder. Our KPI framework exploits the front-door adjustment, which requires no prior knowledge about the confounder. The effectiveness of our KPI framework is corroborated by three baseline methods on five benchmark datasets, including MSVD-QA, MSRVTT-QA, TGIF-QA, NExT-QA, and Causal-VidQA. Jiangtong Li, Li Niu 0002, Liqing Zhang 0001 |
ICCV | 1 |
| 2023 | Painterly Image Harmonization using Diffusion ModelabstractPainterly image harmonization aims to insert photographic objects into paintings and obtain artistically coherent composite images. Previous methods for this task mainly rely on inference optimization or generative adversarial network, but they are either very time-consuming or struggling at fine control of the foreground objects (e.g., texture and content details). To address these issues, we propose a novel Painterly Harmonization stable Diffusion model (PHDiffusion), which includes a lightweight adaptive encoder and a Dual Encoder Fusion (DEF) module. Specifically, the adaptive encoder and the DEF module first stylize foreground features within each encoder. Then, the stylized foreground features from both encoders are combined to guide the harmonization process. During training, besides the noise loss in diffusion model, we additionally employ content loss and two style losses, i.e., AdaIN style loss and contrastive style loss, aiming to balance the trade-off between style migration and content preservation. Compared with the state-of-the-art models from related fields, our PHDiffusion can stylize the foreground more sufficiently and simultaneously retain finer content. Our code and model are available at https://github.com/bcmi/PHDiffusion-Painterly-Image-Harmonization Lingxiao Lu, Jiangtong Li, Junyan Cao, Li Niu 0002, Liqing Zhang 0001 |
ACM Multimedia | 2 |
| 2023 | Deep Image Harmonization in Dual Color SpacesabstractImage harmonization is an essential step in image composition that adjusts the appearance of composite foreground to address the inconsistency between foreground and background. Existing methods primarily operate in correlated RGB color space, leading to entangled features and limited representation ability. In contrast, decorrelated color space (e.g., Lab) has decorrelated channels that provide disentangled color and illumination statistics. In this paper, we explore image harmonization in dual color spaces, which supplements entangled RGB features with disentangled L, a, b features to alleviate the workload in harmonization process. The network comprises a RGB harmonization backbone, an Lab encoding module, and an Lab control module. The backbone is a U-Net network translating composite image to harmonized image. Three encoders in Lab encoding module extract three control codes independently from L, a, b channels, which are used to manipulate the decoder features in harmonization backbone via Lab control module. Our code and model are available at https://github.com/bcmi/DucoNet-Image-Harmonization. Linfeng Tan, Jiangtong Li, Li Niu 0002, Liqing Zhang 0001 |
ACM Multimedia | 2 |
| 2022 | Action-Aware Embedding Enhancement for Image-Text RetrievalabstractImage-text retrieval plays a central role in bridging vision and language, which aims to reduce the semantic discrepancy between images and texts. Most of existing works rely on refined words and objects representation through the data-oriented method to capture the word-object cooccurrence. Such approaches are prone to ignore the asymmetric action relation between images and texts, that is, the text has explicit action representation (i.e., verb phrase) while the image only contains implicit action information. In this paper, we propose Action-aware Memory-Enhanced embedding (AME) method for image-text retrieval, which aims to emphasize the action information when mapping the images and texts into a shared embedding space. Specifically, we integrate action prediction along with an action-aware memory bank to enrich the image and text features with action-similar text features. The effectiveness of our proposed AME method is verified by comprehensive experimental results on two benchmark datasets. Jiangtong Li, Li Niu 0002, Liqing Zhang 0001 |
AAAI | 1 |
| 2022 | From Representation to Reasoning: Towards both Evidence and Commonsense Reasoning for Video Question-AnsweringabstractVideo understanding has achieved great success in representation learning, such as video caption, video object grounding, and video descriptive question-answer. However, current methods still struggle on video reasoning, including evidence reasoning and commonsense reasoning. To facilitate deeper video understanding towards video reasoning, we present the task of Causal-VidQA, which includes four types of questions ranging from scene description (description) to evidence reasoning (explanation) and commonsense reasoning (prediction and counterfactual). For commonsense reasoning, we set up a two-step solution by answering the question and providing a proper reason. Through extensive experiments on existing VideoQA methods, we find that the state-of-the-art methods are strong in descriptions but weak in reasoning. We hope that Causal-VidQA can guide the research of video understanding from representation learning to deeper reasoning. The dataset and related resources are available at https://github.com/bcmi/Causal-VidQA.git. Jiangtong Li, Li Niu 0002, Liqing Zhang 0001 |
CVPR | 1 |
| 2022 | Multi-Level Region Matching for Fine-Grained Sketch-Based Image RetrievalabstractFine-Grained Sketch-Based Image Retrieval (FG-SBIR) is to use free-hand sketches as queries to perform instance-level retrieval in an image gallery. Existing works usually leverage only high-level information and perform matching in a single region. However, both low-level and high-level information are helpful to establish fine-grained correspondence. Besides, we argue that matching different regions between each sketch-image pair can further boost model robustness. Therefore, we propose Multi-Level Region Matching (MLRM) for FG-SBIR, which consists of two modules: a Discriminative Region Extraction module (DRE) and a Region and Level Attention module (RLA). In DRE, we propose Light-weighted Attention Map Augmentation (LAMA) to extract local feature from different regions. In RLA, we propose a transformer-based attentive matching module to learn attention weights to explore different importance from different image/sketch regions and feature levels. Furthermore, to ensure that the geometrical and semantic distinctiveness is well modeled, we also explore a novel LAMA overlapping penalty and a local region-negative triplet loss in our proposed MLRM method. Comprehensive experiments conducted on five datasets (i.e., Sketchy, QMUL-ChairV2, QMUL-ShoeV2, QMUL-Chair, QMUL-Shoe) demonstrate effectiveness of our method. Zhixin Ling, Jiangtong Li, Li Niu 0002 |
ACM Multimedia | 3 |
| 2022 | Zero-shot sketch-based image retrieval with structure-aware asymmetric disentanglement
Jiangtong Li, Zhixin Ling, Li Niu 0002, Liqing Zhang 0001 |
Comput. Vis. Image Underst. | 1 |
| 2021 | Activity Image-to-Video Retrieval by Disentangling Appearance and MotionabstractWith the rapid emergence of video data, image-to-video retrieval has attracted much attention. There are two types of image-to-video retrieval: instance-based and activity-based. The former task aims to retrieve videos containing the same main objects as the query image, while the latter focuses on finding the similar activity. Since dynamic information plays a significant role in the video, we pay attention to the latter task to explore the motion relation between images and videos. In this paper, we propose a Motion-assisted Activity Proposal-based Image-to-Video Retrieval (MAP-IVR) approach to disentangle the video features into motion features and appearance features and obtain appearance features from the images. Then, we perform image-to-video translation to improve the disentanglement quality. The retrieval is performed in both appearance and video feature spaces. Extensive experiments demonstrate that our MAP-IVR approach remarkably outperforms the state-of-the-art approaches on two benchmark activity-based video datasets. Liu Liu 0022, Jiangtong Li, Li Niu 0002, Ruicong Xu, Liqing Zhang 0001 |
AAAI | 2 |
| 2021 | Video Semantic Segmentation via Sparse Temporal TransformerabstractCurrently, video semantic segmentation mainly faces two challenges: 1) the demand of temporal consistency; 2) the balance between segmentation accuracy and inference efficiency. For the first challenge, existing methods usually use optical flow to capture the temporal relation in consecutive frames and maintain the temporal consistency, but the low inference speed by means of optical flow limits the real-time applications. For the second challenge, flow based key frame warping is one mainstream solution. However, the unbalanced inference latency of flow-based key frame warping makes it unsatisfactory for real-time applications. Considering the segmentation accuracy and inference efficiency, we propose a novel Sparse Temporal Transformer (STT) to bridge temporal relation among video frames adaptively, which is also equipped with query selection and key selection. The key selection and query selection strategies are separately applied to filter out temporal and spatial redundancy in our temporal transformer. Specifically, our STT can reduce the time complexity of temporal transformer by a large margin without harming the segmentation accuracy and temporal consistency. Experiments on two benchmark datasets, Cityscapes and Camvid, demonstrate that our method achieves the state-of-the-art segmentation accuracy and temporal consistency with comparable inference speed. Jiangtong Li, Wentao Wang 0009, Junjie Chen 0008, Li Niu 0002, Jianlou Si, Chen Qian 0006, Liqing Zhang 0001 |
ACM Multimedia | 1 |
| 2021 | Memorize, Associate and Match: Embedding Enhancement via Fine-Grained Alignment for Image-Text RetrievalabstractImage-text retrieval aims to capture the semantic correlation between images and texts. Existing image-text retrieval methods can be roughly categorized into embedding learning paradigm and pair-wise learning paradigm. The former paradigm fails to capture the fine-grained correspondence between images and texts. The latter paradigm achieves fine-grained alignment between regions and words, but the high cost of pair-wise computation leads to slow retrieval speed. In this paper, we propose a novel method named MEMBER by using Memory-based EMBedding Enhancement for image-text Retrieval (MEMBER), which introduces global memory banks to enable fine-grained alignment and fusion in embedding learning paradigm. Specifically, we enrich image (resp., text) features with relevant text (resp., image) features stored in the text (resp., image) memory bank. In this way, our model not only accomplishes mutual embedding enhancement across two modalities, but also maintains the retrieval efficiency. Extensive experiments demonstrate that our MEMBER remarkably outperforms state-of-the-art approaches on two large-scale benchmark datasets. Jiangtong Li, Liu Liu 0022, Li Niu 0002, Liqing Zhang 0001 |
IEEE Trans. Image Process. | 1 |
| 2019 | Lattice-Based Transformer Encoder for Neural Machine TranslationabstractNeural machine translation (NMT) takes deterministic sequences for source representations.However, either wordlevel or subword-level segmentations have multiple choices to split a source sequence with different word segmentors or different subword vocabulary sizes.We hypothesize that the diversity in segmentations may affect the NMT performance.To integrate different segmentations with the state-of-the-art NMT model, Transformer, we propose lattice-based encoders to explore effective word or subword representation in an automatic way during training.We propose two methods: 1) lattice positional encoding and 2) lattice-aware self-attention.These two methods can be used together and show complementary to each other to further improve translation performance.Experiment results show superiorities of lattice-based encoders in word-level and subword-level representations over conventional Transformer encoder. Fengshun Xiao, Jiangtong Li, Hai Zhao 0001, Rui Wang 0015, Kehai Chen |
ACL (1) | 2 |
| 2019 | Effective Subword Segmentation for Text ComprehensionabstractRepresentation learning is the foundation of machine reading comprehension and inference. In state-of-the-art models, character-level representations have been broadly adopted to alleviate the problem of effectively representing rare or complex words. However, character itself is not a natural minimal linguistic unit for representation or word embedding composing due to ignoring the linguistic coherence of consecutive characters inside word. This paper presents a general subword-augmented embedding framework for learning and composing computationally derived subword-level representations. We survey a series of unsupervised segmentation methods for subword acquisition and different subword-augmented strategies for text understanding, showing that subword-augmented embedding significantly improves our baselines in various types of text understanding tasks on both English and Chinese benchmarks. Zhuosheng Zhang 0001, Hai Zhao 0001, Kangwei Ling, Jiangtong Li, Zuchao Li, Shexia He, Guohong Fu |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2018 | Modeling Multi-turn Conversation with Deep Utterance AggregationabstractMulti-turn conversation understanding is a major challenge for building intelligent dialogue systems. This work focuses on retrieval-based response matching for multi-turn conversation whose related work simply concatenates the conversation utterances, ignoring the interactions among previous utterances for context modeling. In this paper, we formulate previous utterances into context using a proposed deep utterance aggregation model to form a fine-grained context representation. In detail, a self-matching attention is first introduced to route the vital information in each utterance. Then the model matches a response with each refined utterance and the final matching score is obtained after attentive turns aggregation. Experimental results show our model outperforms the state-of-the-art methods on three multi-turn conversation benchmarks, including a newly introduced e-commerce dialogue corpus. Zhuosheng Zhang 0001, Jiangtong Li, Pengfei Zhu 0003, Hai Zhao 0001, Gongshen Liu |
COLING | 2 |