EDBT 2026 Demo / reviewers in the wild / expert
Jidong Tian
dblp:230/4307
· DBLP profile ↗
24ranked-venue papers
3as first author
21since 2021 · last 2026
0000-0003-1936-4595ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 3 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 2 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Consistent 3D Human Reconstruction From Monocular Video: Learning Correctable Appearance and Temporal Motion PriorsabstractRecent advancements in rendering dynamic humans using NeRF and 3D Gaussian splatting have made significant progress, leveraging implicit geometry learning and image appearance rendering to create digital humans. However, in monocular video rendering, there are still challenges in rendering subtle and complex motion from different viewpoints and states, primarily due to the imbalance of viewpoints. Additionally, ensuring continuity between adjacent frames when rendering from novel and free viewpoints remains a difficult task. To address these challenges, we first propose a pixel-level motion correction module that adjusts the errors in the learned representation between different viewpoints. We also introduce a temporal information-based model to improve motion continuity by leveraging adjacent frames. Experimental results on dynamic human rendering, using the NeuMan, ZJU-Mocap, and People-Snapshot datasets, demonstrate that our method outperforms state-of-the-art techniques both quantitatively and qualitatively. Cheng Shang, Liang An 0001, Jiajun Zhang 0012, Yuxiang Zhang 0006, Jidong Tian, Yebin Liu, Xubo Yang |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2025 | ChemActor: Enhancing Automated Extraction of Chemical Synthesis Actions with LLM-Generated DataabstractYu Zhang, Ruijie Yu, Jidong Tian, Feng Zhu, Jiapeng Liu, Xiaokang Yang, Yaohui Jin, Yanyan Xu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Ruijie Yu, Jidong Tian, Feng Zhu 0006, Xiaokang Yang 0001, Yaohui Jin, Yanyan Xu 0002 |
ACL (1) | 3 |
| 2025 | Look Before You Leap: Problem Elaboration Prompting Improves Mathematical Reasoning in Large Language ModelsabstractLarge language models (LLMs) still grapple with complex tasks like mathematical reasoning. Despite significant efforts invested in improving prefix prompts or reasoning process, the crucial role of problem context might have been neglected. Accurate recognition of inputs is fundamental for solving mathematical tasks, as ill-formed problems could potentially mislead LLM’s reasoning. In this study, we propose a new approach named Problem Elaboration Prompting (PEP) to enhance the mathematical capacities of LLMs. Specifically, PEP decomposes and elucidates the problem context before reasoning, therefore enhancing the context modeling and parsing efficiency. Experiments across datasets and models demonstrate promising performances: (1) PEP demonstrates an overall enhancement in various situation. (2) PEP can be easily implemented and integrated with other prompting methods. (3) PEP shows particular strength in handling distraction problems. Haoran Liao, Jidong Tian, Shaohua Hu, Hao He 0007, Yaohui Jin |
ICASSP | 2 |
| 2025 | KinFormer: Generalizable Dynamical Symbolic Regression for Catalytic Organic Reaction KineticsabstractModeling kinetic equations is essential for understanding the mechanisms of chemical reactions, yet a complex and time-consuming task. Kinetic equation prediction is formulated as a problem of dynamical symbolic regression (DSR) subject to physical chemistry constraints. Deep learning (DL) holds the potential to capture reaction patterns and predict kinetic equations from data of chemical species, effectively avoiding empirical bias and improving efficiency compared with traditional analytical methods. Despite numerous studies focusing on DSR and the introduction of Transformers to predict ordinary differential equations, the corresponding models lack generalization abilities across diverse categories of reactions. In this study, we propose KinFormer, a generalizable kinetic equation prediction model. KinFormer utilizes a conditional Transformer to model DSR under physical constraints and employs Monte Carlo Tree Search to apply the model to new types of reactions. Experimental results on 20 types of organic reactions demonstrate that KinFormer not only outperforms classical baselines, but also exceeds Transformer baselines in out-of-domain evaluations, thereby proving its generalization ability. Jindou Chen, Jidong Tian, ChenXinWei, Xiaokang Yang 0001, Yaohui Jin, Yanyan Xu 0002 |
ICLR | 2 |
| 2024 | Self-Hint Prompting Improves Zero-shot Reasoning in Large Language Models via Reflective Cycle
Jindou Chen, Jidong Tian, Yaohui Jin |
CogSci | 2 |
| 2024 | Comparable Demonstrations Are Important In In-Context Learning: A Novel Perspective On Demonstration SelectionabstractIn-Context Learning (ICL) is an important paradigm for adapting Large Language Models (LLMs) to downstream tasks through a few demonstrations. Despite the great success of ICL, the limitation of the demonstration number may lead to demonstration bias, i.e. the input-label mapping induced by LLMs misunderstands the task’s essence. Inspired by human experience, we attempt to mitigate such bias through the perspective of the inter-demonstration relationship. Specifically, we construct Comparable Demonstrations (CDs) by minimally editing the texts to flip the corresponding labels, in order to highlight the task’s essence and eliminate potential spurious correlations through the inter-demonstration comparison. Through a series of experiments on CDs, we find that (1) demonstration bias does exist in LLMs, and CDs can significantly reduce such bias; (2) CDs exhibit good performance in ICL, especially in out-of-distribution scenarios. In summary, this study explores the ICL mechanisms from a novel perspective, providing a deeper insight into the demonstration selection strategy for ICL. Caoyun Fan, Jidong Tian, Hao He 0007, Yaohui Jin |
ICASSP | 2 |
| 2024 | Free-view Rendering of Dynamic Human from Monocular Video Via Modeling Temporal Information Globally and Locally among Adjacent FramesabstractRecent research developments on rendering dynamic humans using neural radiance fields are remarkable. These methods often utilize learning implicit geometry and image appearance rendering for digital humans. However, keeping the complex and fast motions in detail, such as fingers, clothes, and faces, remains a challenge. Inspired by temporal information from human motion, we propose an architecture among adjacent frames by constructing a model on global and local levels. For the global level, we propose a hidden Markov model (HMM)based method to capture the global similarity among adjacent frames. At the local level, we introduce a module composed of a multi-head attention mechanism on a triplet canonical space structure for patch-level local temporal information. Experiments on two public datasets of dynamic human rendering (ZJU-MoCap and the People-Snapshot dataset) demonstrate that the proposed method outperforms advanced methods quantitatively and qualitatively. Cheng Shang, Jidong Tian, Jiannan Ye, Xubo Yang |
ICME | 2 |
| 2024 | Unlock the Potential of Counterfactually-Augmented Data in Out-Of-Distribution Generalization
Caoyun Fan, Wenqing Chen, Jidong Tian, Hao He 0007, Yaohui Jin |
Expert Syst. Appl. | 3 |
| 2023 | Preference-Controlled Multi-Objective Reinforcement Learning for Conditional Text GenerationabstractConditional text generation is to generate text sequences conditioning on linguistic or non-linguistic data. The main line of existing work proposed deterministic models to improve the fidelity of the generated text but often ignored the diversity. Another line relied on conditional variational auto-encoders (CVAEs), which increased the diversity over their deterministic backbones. However, CVAEs regard diversity as an implicit objective and may not be optimal. In this paper, we raise two questions: i) Can diversity be further improved with an explicit objective? ii) Since fidelity and diversity are two conflicting objectives, how can we obtain different multi-objective optimal solutions according to user preferences? To answer question i), we propose a multi-objective reinforcement learning (MORL) method which explicitly takes CIDEr and Self-CIDEr scores as the fidelity-oriented and diversity-oriented rewards respectively. To answer question ii), we propose a preference-controlled MORL method, which can obtain infinite multi-objective optimal solutions by tuning the preference variable. We conduct extensive experiments on paraphrasing and image captioning tasks, which show that in the fidelity-diversity trade-off space, our model outperforms both deterministic and CVAE-based baselines. Wenqing Chen, Jidong Tian, Caoyun Fan, Hao He 0007, Yaohui Jin |
AAAI | 2 |
| 2023 | Latent Constraints on Unsupervised Text-Graph Alignment with Information AsymmetryabstractUnsupervised text-graph alignment (UTGA) is a fundamental task that bidirectionally generates texts and graphs without parallel data. Most available models of UTGA suffer from information asymmetry, a common phenomenon that texts and graphs include additional information invisible to each other. On the one hand, these models fail to supplement asymmetric information effectively due to the lack of ground truths. On the other hand, it is challenging to indicate asymmetric information with explicit indicators because it cannot be decoupled from the data directly. To address the challenge posed by information asymmetry, we propose the assumption that asymmetric information is encoded in unobservable latent variables and only affects the one-way generation processes. These latent variables corresponding to asymmetric information should obey prior distributions recovered approximately from original data. Therefore, we first propose a taxonomy of the latent variable that classifies the latent variable into transferrable (TV) and non-transferable (NTV) variables and further distinguish NTV as the dependent variable (DV) and the independent variable (IV). Next, we propose three latent VAE-based regularizations on TV, DV, and IV to constrain their distributions to well-designed prior distributions to introduce asymmetric information into models and enhance the preservation of shared contents. Finally, we impose the three proposed constraints on a cycle-consistent learning framework, back-translation (BT), named ConstrainedBT. Experimental results on three UTGA tasks demonstrate the effectiveness of ConstrainedBT on the information-asymmetric challenge. Jidong Tian, Wenqing Chen, Caoyun Fan, Hao He 0007, Yaohui Jin |
AAAI | 1 |
| 2023 | Task-Level Thinking Steps Help Large Language Models for Challenging Classification TaskabstractLarge language models (LLMs) have shown incredible performance on many tasks such as dialogue generation, commonsense reasoning and question answering.In-context learning (ICL) is an important paradigm for adapting LLMs to the downstream tasks by prompting few demonstrations.However, the distribution of demonstrations can severely affect the performance, especially for challenging classification tasks.In this paper, we propose the concept of task-level thinking steps that can eliminate bias introduced by demonstrations.Further, to help LLMs distinguish confusing classes, we design a progressive revision framework, which can improve the thinking steps by correcting hard demonstrations.Experimental results prove the superiority of our proposed method, achieving best performance on three kinds of challenging classification tasks in the zero-shot and few-shot settings.Besides, with task-level thinking steps, automatically generated chain-of-thoughts (CoTs) bring more competitive performance. Jidong Tian, Haoran Liao, Jindou Chen, Hao He 0007, Yaohui Jin |
EMNLP | 2 |
| 2023 | Chain-of-Thought Tuning: Masked Language Models can also Think Step By Step in Natural Language UnderstandingabstractChain-of-Thought (CoT) is a technique that guides Large Language Models (LLMs) to decompose complex tasks into multi-step reasoning through intermediate steps in natural language form.Briefly, CoT enables LLMs to think step by step.However, although many Natural Language Understanding (NLU) tasks also require thinking step by step, LLMs perform less well than small-scale Masked Language Models (MLMs).To migrate CoT from LLMs to MLMs, we propose Chain-of-Thought Tuning (CoTT), a two-step reasoning framework based on prompt tuning, to implement step-by-step thinking for MLMs on NLU tasks.From the perspective of CoT, CoTT's two-step framework enables MLMs to implement task decomposition; CoTT's prompt tuning allows intermediate steps to be used in natural language form.Thereby, the success of CoT can be extended to NLU tasks through MLMs.To verify the effectiveness of CoTT, we conduct experiments on two NLU tasks: hierarchical classification and relation extraction, and the results show that CoTT outperforms baselines and achieves state-of-the-art performance. Caoyun Fan, Jidong Tian, Wenqing Chen, Hao He 0007, Yaohui Jin |
EMNLP | 2 |
| 2023 | Improving the out-of-Distribution Generalization Capability of Language Models: Counterfactually-Augmented Data is not EnoughabstractCounterfactually-Augmented Data (CAD) has the potential to improve language models’ Out-Of-Distribution (OOD) generalization capability, as CAD induces language models to exploit causal features and exclude spurious correlations. However, the empirical results of OOD generalization on CAD are not as efficient as expected. In this paper, we attribute the inefficiency to Myopia Phenomenon caused by CAD: language models only focus on causal features that are edited in the augmentation and exclude other non-edited causal features. As a result, the potential of CAD is not fully exploited. Based on the structural properties of CAD, we design two additional constraints to help language models extract more complete causal features contained in CAD, thus improving the OOD generalization capability. We evaluate our method on two tasks: Sentiment Analysis and Natural Language Inference, and the experimental results demonstrate that our method could unlock CAD’s potential and improve language models’ OOD generalization capability. Caoyun Fan, Wenqing Chen, Jidong Tian, Hao He 0007, Yaohui Jin |
ICASSP | 3 |
| 2023 | Accurate use of label dependency in multi-label text classification through the lens of causality
Caoyun Fan, Wenqing Chen, Jidong Tian, Hao He 0007, Yaohui Jin |
Appl. Intell. | 3 |
| 2022 | Weakly Supervised Neural Symbolic Learning for Cognitive TasksabstractDespite the recent success of end-to-end deep neural networks, there are growing concerns about their lack of logical reasoning abilities, especially on cognitive tasks with perception and reasoning processes. A solution is the neural symbolic learning (NeSyL) method that can effectively utilize pre-defined logic rules to constrain the neural architecture making it perform better on cognitive tasks. However, it is challenging to apply NeSyL to these cognitive tasks because of the lack of supervision, the non-differentiable manner of the symbolic system, and the difficulty to probabilistically constrain the neural network. In this paper, we propose WS-NeSyL, a weakly supervised neural symbolic learning model for cognitive tasks with logical reasoning. First, WS-NeSyL employs a novel back search algorithm to sample the possible reasoning process through logic rules. This sampled process can supervise the neural network as the pseudo label. Based on this algorithm, we can backpropagate gradients to the neural network of WS-NeSyL in a weakly supervised manner. Second, we introduce a probabilistic logic regularization into WS-NeSyL to help the neural network learn probabilistic logic. To evaluate WS-NeSyL, we have conducted experiments on three cognitive datasets, including temporal reasoning, handwritten formula recognition, and relational reasoning datasets. Experimental results show that WS-NeSyL not only outperforms the end-to-end neural model but also beats the state-of-the-art neural symbolic learning models. Jidong Tian, Wenqing Chen, Liqiang Xiao, Hao He 0007, Yaohui Jin |
AAAI | 1 |
| 2022 | MaxGNR: A Dynamic Weight Strategy via Maximizing Gradient-to-Noise Ratio for Multi-task Learning
Caoyun Fan, Wenqing Chen, Jidong Tian, Hao He 0007, Yaohui Jin |
ACCV (1) | 3 |
| 2022 | To What Extent Do Natural Language Understanding Datasets Correlate to Logical Reasoning? A Method for Diagnosing Logical ReasoningabstractReasoning and knowledge-related skills are considered as two fundamental skills for natural language understanding (NLU) tasks such as machine reading comprehension (MRC) and natural language inference (NLI). However, it is not clear to what extent an NLU task defined on a dataset correlates to a specific NLU skill. On the one hand, evaluating the correlation requires an understanding of the significance of the NLU skill in a dataset. Significance judges whether a dataset includes sufficient material to help the model master this skill. On the other hand, it is also necessary to evaluate the dependence of the task on the NLU skill. Dependence is a measure of how much the task defined on a dataset depends on the skill. In this paper, we propose a systematic method to diagnose the correlations between an NLU dataset and a specific skill, and then take a fundamental reasoning skill, logical reasoning, as an example for analysis. The method adopts a qualitative indicator to indicate the significance while adopting a quantitative indicator to measure the dependence. We perform diagnosis on 8 MRC datasets (including two types) and 3 NLI datasets and acquire intuitively reasonable results. We then perform the analysis to further understand the results and the proposed indicators. Based on the analysis, although the diagnostic method has some limitations, it is still an effective method to perform a basic diagnosis of the correlation between the dataset and logical reasoning skill, which also can be generalized to other NLU skills. Jidong Tian, Wenqing Chen, Caoyun Fan, Hao He 0007, Yaohui Jin |
COLING | 2 |
| 2022 | CALM: Commen-Sense Knowledge Augmentation for Document Image UnderstandingabstractPerformance of document image understanding has been significantly fueled by encoding multi-modal information in recent years. However, existing works heavily rely on the superficial appearance of the observed data, resulting in counter-intuitive model behavior in many critical cases. To overcome this issue, this paper proposes a common-sense knowledge augmented model CALM for document image understanding tasks. It firstly produces purified representations of document contents to extract key information and learn common-sense augmented representation for inputs. Then, relevant common-sense knowledge is extracted from the external ConceptNet knowledge base, and a derived knowledge graph is built to enhance the common-sense reasoning capability of CALM jointly. In order to further highlight the importance of common-sense knowledge in document image understanding, we propose the first question-answering dataset, CS-DVQA, focused on common-sense reasoning for document images, in which questions are answered by taking both document contents and common-sense knowledge into consideration. Through extensive evaluation, the proposed CALM approach outperforms the state-of-the-art models in three document image understanding tasks, including key information extraction(from 85.37 to 86.52), document image classification(from 96.08 to 96.17), document visual question answering(from 86.72 to 88.03). Qinyi Du, Keqian Li, Jidong Tian, Liqiang Xiao, Yaohui Jin |
ACM Multimedia | 4 |
| 2021 | De-Confounded Variational Encoder-Decoder for Logical Table-to-Text GenerationabstractWenqing Chen, Jidong Tian, Yitian Li, Hao He, Yaohui Jin. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Wenqing Chen, Jidong Tian, Hao He 0007, Yaohui Jin |
ACL/IJCNLP (1) | 2 |
| 2021 | Diagnosing the First-Order Logical Reasoning Ability Through LogicNLIabstractRecently, language models (LMs) have achieved significant performance on many NLU tasks, which has spurred widespread interest for their possible applications in the scientific and social area.However, LMs have faced much criticism of whether they are truly capable of reasoning in NLU.In this work, we propose a diagnostic method for first-order logic (FOL) reasoning with a new proposed benchmark, LogicNLI.LogicNLI is an NLI-style dataset that effectively disentangles the target FOL reasoning from commonsense inference and can be used to diagnose LMs from four perspectives: accuracy, robustness, generalization, and traceability.Experiments on BERT, RoBERTa, and XLNet, have uncovered the weaknesses of these LMs on FOL reasoning, which motivates future exploration to enhance the reasoning ability. Jidong Tian, Wenqing Chen, Liqiang Xiao, Hao He 0007, Yaohui Jin |
EMNLP (1) | 1 |
| 2021 | Dependent Multi-Task Learning with Causal Intervention for Image CaptioningabstractRecent work for image captioning mainly followed an extract-then-generate paradigm, pre-extracting a sequence of object-based features and then formulating image captioning as a single sequence-to-sequence task. Although promising, we observed two problems in generated captions: 1) content inconsistency where models would generate contradicting facts; 2) not informative enough where models would miss parts of important information. From a causal perspective, the reason is that models have captured spurious statistical correlations between visual features and certain expressions (e.g., visual features of "long hair" and "woman"). In this paper, we propose a dependent multi-task learning framework with the causal intervention (DMTCI). Firstly, we involve an intermediate task, bag-of-categories generation, before the final task, image captioning. The intermediate task would help the model better understand the visual features and thus alleviate the content inconsistency problem. Secondly, we apply Pearl's do-calculus on the model, cutting off the link between the visual features and possible confounders and thus letting models focus on the causal visual features. Specifically, the high-frequency concept set is considered as the proxy confounders where the real confounders are inferred in the continuous space. Finally, we use a multi-agent reinforcement learning (MARL) strategy to enable end-to-end training and reduce the inter-task error accumulations. The extensive experiments show that our model outperforms the baseline models and achieves competitive performance with state-of-the-art models. Wenqing Chen, Jidong Tian, Caoyun Fan, Hao He 0007, Yaohui Jin |
IJCAI | 2 |
| 2020 | A Semantically Consistent and Syntactically Variational Encoder-Decoder Framework for Paraphrase GenerationabstractParaphrase generation aims to generate semantically consistent sentences with different syntactic realizations.Most of the recent studies rely on the typical encoder-decoder framework where the generation process is deterministic.However, in practice, the ability to generate multiple syntactically different paraphrases is important.Recent work proposed to cooperate variational inference on a target-related latent variable to introduce the diversity.But the latent variable may be contaminated by the semantic information of other unrelated sentences, and in turn, change the conveyed meaning of generated paraphrases.In this paper, we propose a semantically consistent and syntactically variational encoder-decoder framework, which uses adversarial learning to ensure the syntactic latent variable be semantic-free.Moreover, we adopt another discriminator to improve the word-level and sentence-level semantic consistency.So the proposed framework can generate multiple semantically consistent and syntactically different paraphrases.The experiments show that our model outperforms the baseline models on the metrics based on both n-gram matching and semantic similarity, and our model can generate multiple different paraphrases by assembling different syntactic variables. Wenqing Chen, Jidong Tian, Liqiang Xiao, Hao He 0007, Yaohui Jin |
COLING | 2 |
| 2020 | Exploring Logically Dependent Multi-task Learning with Causal InferenceabstractPrevious studies have shown that hierarchical multi-task learning (MTL) can utilize task dependencies by stacking encoders and outperform democratic MTL.However, stacking encoders only considers the dependencies of feature representations and ignores the label dependencies in logically dependent tasks.Furthermore, how to properly utilize the labels remains an issue due to the cascading errors between tasks.In this paper, we view logically dependent MTL from the perspective of causal inference and suggest a mediation assumption instead of the confounding assumption in conventional MTL models.We propose a model including two key mechanisms: label transfer (LT) for each task to utilize the labels of all its lower-level tasks, and Gumbel sampling (GS) to deal with cascading errors.In the field of causal inference, GS in our model is essentially a counterfactual reasoning process, trying to estimate the causal effect between tasks and utilize it to improve MTL.We conduct experiments on two English datasets and one Chinese dataset.Experiment results show that our model achieves state-of-the-art on six out of seven subtasks and improves predictions' consistency. Wenqing Chen, Jidong Tian, Liqiang Xiao, Hao He 0007, Yaohui Jin |
EMNLP (1) | 2 |
| 2019 | TransMS: Knowledge Graph Embedding for Complex Relations by Multidirectional SemanticsabstractKnowledge graph embedding, which projects the symbolic relations and entities onto low-dimension continuous spaces, is essential to knowledge graph completion. Recently, translation-based embedding models (e.g. TransE) have aroused increasing attention for their simplicity and effectiveness. These models attempt to translate semantics from head entities to tail entities with the relations and infer richer facts outside the knowledge graph. In this paper, we propose a novel knowledge graph embedding method named TransMS, which translates and transmits multidirectional semantics: i) the semantics of head/tail entities and relations to tail/head entities with nonlinear functions and ii) the semantics from entities to relations with linear bias vectors. Our model has merely one additional parameter α than TransE for each triplet, which results in its better scalability in large-scale knowledge graph. Experiments show that TransMS achieves substantial improvements against state-of-the-art baselines, especially the Hit@10s of head entity prediction for N-1 relations and tail entity prediction for 1-N relations improved by about 27.1% and 24.8% on FB15K database respectively. Shihui Yang 0001, Jidong Tian, Honglun Zhang, Junchi Yan, Hao He 0007, Yaohui Jin |
IJCAI | 2 |