Yaohui Jin

dblp:27/7040 · DBLP profile ↗
← Back
83ranked-venue papers
0as first author
46since 2021 · last 2026
0000-0001-6158-6277ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 46 · 32 since 2021Graphics, computer vision, multimedia, augmented reality and games · 23 · 17 since 2021Computer networks · 18 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 7 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 3Systems, architecture and hardware · 2
YearPublicationVenuePosition
2026 ChemReason-Bench: Benchmarking Large Language Models for Procedural Reasoning in Experimental Chemistry
abstract
Jinwei Zhang, Xucheng Liang, Yu Zhang, Ruijie Yu, Xiaokang Yang, Yaohui Jin, Yanyan Xu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Xucheng Liang, Ruijie Yu, Xiaokang Yang 0001, Yaohui Jin, Yanyan Xu 0002
ACL (1)6
2025 ChemActor: Enhancing Automated Extraction of Chemical Synthesis Actions with LLM-Generated Data
abstract
Yu Zhang, Ruijie Yu, Jidong Tian, Feng Zhu, Jiapeng Liu, Xiaokang Yang, Yaohui Jin, Yanyan Xu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Ruijie Yu, Jidong Tian, Feng Zhu 0006, Xiaokang Yang 0001, Yaohui Jin, Yanyan Xu 0002
ACL (1)7
2025 Enhancing Objectivity in LLM-as-a-Judge through Perturbation Injection
Haoran Liao, Yaohui Jin
CogSci3
2025 FD-Bench: Fine-Grained Evaluating the Decision-Making Capability of LLM Agents in Dynamic Scenarios
Yaohui Jin
CogSci3
2025 TopoRefine: Iterative Refinement with Reasoning Topology as High-Level Feedback
abstract
By leveraging effective signals to refine their outputs, large language models (LLMs) can achieve superior performance compared to single-pass outputs. However, internal signals often suffer from accumulated hallucinations and a lack of confidence, while external signals are typically difficult to obtain and apply, hindering the development of the self-refinement paradigm. In this work, we propose a novel method named TopoRefine, which integrates the reasoning topology within the model’s outputs as reliable high-level feedback. Specifically, TopoRefine encourages LLMs to explore reasoning paths without assuming a predefined direction for improvement, while using a consistency mechanism to prevent performance degradation. Reasoning topology can provide higher-level information as feedback, offering more precise semantics and facilitating self-analysis. Experiments on mathematical datasets and with various LLMs demonstrate significant improvements using our method. We also provide a detailed analysis of TopoRefine’s efficiency.
Haoran Liao, Shaohua Hu, Hao He 0007, Yaohui Jin
ICASSP5
2025 Look Before You Leap: Problem Elaboration Prompting Improves Mathematical Reasoning in Large Language Models
abstract
Large language models (LLMs) still grapple with complex tasks like mathematical reasoning. Despite significant efforts invested in improving prefix prompts or reasoning process, the crucial role of problem context might have been neglected. Accurate recognition of inputs is fundamental for solving mathematical tasks, as ill-formed problems could potentially mislead LLM’s reasoning. In this study, we propose a new approach named Problem Elaboration Prompting (PEP) to enhance the mathematical capacities of LLMs. Specifically, PEP decomposes and elucidates the problem context before reasoning, therefore enhancing the context modeling and parsing efficiency. Experiments across datasets and models demonstrate promising performances: (1) PEP demonstrates an overall enhancement in various situation. (2) PEP can be easily implemented and integrated with other prompting methods. (3) PEP shows particular strength in handling distraction problems.
Haoran Liao, Jidong Tian, Shaohua Hu, Hao He 0007, Yaohui Jin
ICASSP6
2025 Faithful Self-Refinement in Mathematical Reasoning via Progressive Back-Translation
abstract
Large language models (LLMs) can achieve superior results through iterative refinement based on internal or external signals, compared to the unstable outputs from a single pass. However, the reliability of existing internal signals is questionable due to their susceptibility to intrinsic hallucinations, while external signals are only useful in limited scenarios. In this paper, we introduce a novel framework called Progressive Back-Translation refinement (PBT). Specifically, PBT prompts LLMs to extract and reconstruct the question from the answer, avoiding unreliable inferences, critiques, or judgments on intermediate results. We then derive and provide accurate, fine-grained feedback by identifying discrepancies between the back-translated and original questions. Experiments across various large language models and challenging mathematical datasets demonstrate consistent improvements. We also provide a detailed analysis to confirm the effectiveness of the proposed method.
Haoran Liao, Shaohua Hu, Hao He 0007, Yaohui Jin
ICASSP5
2025 Context Guided Transformer Entropy Modeling for Video Compression
abstract
Conditional entropy models effectively leverage spatio-temporal contexts to reduce video redundancy. However, incorporating temporal context often introduces additional model complexity and increases computational cost. In parallel, many existing spatial context models lack explicit modeling the ordering of spatial dependencies, which may limit the availability of relevant context during decoding. To address these issues, we propose the Context Guided Transformer (CGT) entropy model, which estimates probability mass functions of the current frame conditioned on resampled temporal context and dependency-weighted spatial context. A temporal context resampler learns predefined latent queries to extract critical temporal information using transformer encoders, reducing downstream computational overhead. Meanwhile, a teacher-student network is designed as dependency-weighted spatial context assigner to explicitly model the dependency of spatial context order. The teacher generates an attention map to represent token importance and an entropy map to reflect prediction certainty from randomly masked inputs, guiding the student to select the weighted top-k tokens with the highest spatial dependency. During inference, only the student is used to predict undecoded tokens based on high-dependency context. Experimental results demonstrate that our CGT model reduces entropy modeling time by approximately 65% and achieves an 11% BD-Rate reduction compared to the previous state-of-the-art conditional entropy model.
Junlong Tong, Wei Zhang 0185, Yaohui Jin, Xiaoyu Shen 0001
ICCV3
2025 PASG: A Closed-Loop Framework for Automated Geometric Primitive Extraction and Semantic Anchoring in Robotic Manipulation
abstract
The fragmentation between high-level task semantics and low-level geometric features remains a persistent challenge in robotic manipulation. While vision-language models (VLMs) have shown promise in generating affordance-aware visual representations, the lack of semantic grounding in canonical spaces and reliance on manual annotations severely limit their ability to capture dynamic semantic-affordance relationships. To address these, we propose Primitive-Aware Semantic Grounding (PASG), a closed-loop framework that introduces: (1) Automatic primitive extraction through geometric feature aggregation, enabling cross-category detection of keypoints and axes; (2) VLM-driven semantic anchoring that dynamically couples geometric primitives with functional affordances and task-relevant description; (3) A spatial-semantic reasoning benchmark and a fine-tuned VLM (Qwen2.5VL-PA). We demonstrate PASG's effectiveness in practical robotic manipulation tasks across diverse scenarios, achieving performance comparable to manual annotations. PASG achieves a finer-grained semantic-affordance understanding of objects, establishing a unified paradigm for bridging geometric primitives with task semantics in robotic manipulation.
Yaohui Jin, Yao Mu 0001
ICCV4
2025 KinFormer: Generalizable Dynamical Symbolic Regression for Catalytic Organic Reaction Kinetics
abstract
Modeling kinetic equations is essential for understanding the mechanisms of chemical reactions, yet a complex and time-consuming task. Kinetic equation prediction is formulated as a problem of dynamical symbolic regression (DSR) subject to physical chemistry constraints. Deep learning (DL) holds the potential to capture reaction patterns and predict kinetic equations from data of chemical species, effectively avoiding empirical bias and improving efficiency compared with traditional analytical methods. Despite numerous studies focusing on DSR and the introduction of Transformers to predict ordinary differential equations, the corresponding models lack generalization abilities across diverse categories of reactions. In this study, we propose KinFormer, a generalizable kinetic equation prediction model. KinFormer utilizes a conditional Transformer to model DSR under physical constraints and employs Monte Carlo Tree Search to apply the model to new types of reactions. Experimental results on 20 types of organic reactions demonstrate that KinFormer not only outperforms classical baselines, but also exceeds Transformer baselines in out-of-domain evaluations, thereby proving its generalization ability.
Jindou Chen, Jidong Tian, ChenXinWei, Xiaokang Yang 0001, Yaohui Jin, Yanyan Xu 0002
ICLR6
2025 Forest for the Trees: Overarching Prompting Evokes High-Level Reasoning in Large Language Models
abstract
Haoran Liao, Shaohua Hu, Zhihao Zhu, Hao He, Yaohui Jin. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Haoran Liao, Shaohua Hu, Hao He 0007, Yaohui Jin
NAACL (Long Papers)5
2025 Unified and Generalizable Reinforcement Learning for Facility Location Problems on Graphs
abstract
Facility location problems on graphs are ubiquitous in the real world and hold significant importance, yet their resolution is often impeded by NP-hardness. MIP solvers can find the optimal solutions but fail to handle large instances, while algorithm efficiency has a higher priority in cases of emergency. Recently, machine learning methods have been proposed to tackle such classical problems with fast inference, but they are limited to the myopic constructive pattern and only consider simple cases in Euclidean space. This paper introduces a unified and generalizable approach to tackle facility location problems on weighted graphs with deep reinforcement learning, demonstrating a keen awareness of complex graph structures. Striking a harmonious balance between solution quality and running time, our method stands out with superior efficiency and steady performance. Our model trained on small graphs is highly scalable and consistently generates high-quality solutions, achieving a speedup of more than 2000 times to Gurobi on instances with 1000 nodes. The experiments on Shanghai road networks further demonstrate its practical value in solving real-world problems. The source codes are available at https://github.com/AryaGuo/PPO-swap.
Runzhong Wang, Yanyan Xu 0002, Yaohui Jin
WWW4
2025 Optimizing Resource Allocation and Energy Efficiency in Vehicle Mobile-Edge Computing With Blockchain Integration
abstract
The availability of conventional mobile edge computing (MEC) for vehicles is often hindered by signal interference and attenuation, limiting its efficiency in supporting computationally intensive and latency-sensitive applications. To address these challenges, we propose a novel blockchainenabled vehicular mobile edge computing (VMEC) system that enhances resource sharing and energy efficiency in electric vehicle (EV)-centric services. The system employs an improved RAFT-based consensus mechanism (mRAFT), which dynamically evaluates the reputation of access point (AP) nodes based on their available resources, ensuring fair leader election and enhancing consensus reliability and efficiency. Furthermore, a probabilistic model is introduced to describe AP behaviors, improving the security of the consensus process. To minimize overall energy consumption, we develop a decentralized optimization framework using the Alternating Direction Method of Multipliers (ADMM). This framework jointly optimizes AP clustering, computation resource allocation, and bandwidth scheduling to achieve energy-efficient task offloading and consensus. Simulation results demonstrate that the proposed VMEC system reduces latency by 29.53 and energy consumption by 43.43 schemes, showcasing its effectiveness in delivering low-latency, energy-efficient services for advanced vehicular applications.
Yongsheng Cao, Caiping Zhao, Yihong Zhang 0002, Yaohui Jin
IEEE Internet Things J.4
2025 TrajGEOS: Trajectory Graph Enhanced Orientation-Based Sequential Network for Mobility Prediction
abstract
Human mobility studies how people move to access their needed resources and plays a significant role in urban planning and location-based services. As a paramount task of human mobility modeling, next location prediction is challenging because of the diversity of users’ historical trajectories that gives rise to complex mobility patterns and various contexts. Deep sequential models have been widely used to predict the next location by leveraging the inherent sequentiality of trajectory data. However, they do not fully leverage the relationship between locations and fail to capture users’ multilevel preferences. This work constructs a trajectory graph from users’ historical traces and proposes a Trajectory Graph Enhanced Orientation-based Sequential network (TrajGEOS) for next-location prediction tasks. TrajGEOS introduces hierarchical graph convolution to capture location and user embeddings. Such embeddings consider not only the contextual feature of locations but also the relation between them, and serve as additional features in downstream modules. Specifically, TrajGEOS models short-term preferences from recent trajectory sequences via GRU, mid-term preferences from historical trajectories over weeks through attention mechanisms, and long-term preferences from the global trajectory graph. Extensive experiments on three real-world LBSN datasets corroborate the value of graph and orientation-based modules and demonstrate that TrajGEOS outperforms the state-of-the-art methods on the next location prediction task.
Zhaoping Hu, Zongyuan Huang, Yaohui Jin, Yanyan Xu 0002
IEEE Trans. Comput. Soc. Syst.6
2024 Can Large Language Models Serve as Rational Players in Game Theory? A Systematic Analysis
abstract
Game theory, as an analytical tool, is frequently utilized to analyze human behavior in social science research. With the high alignment between the behavior of Large Language Models (LLMs) and humans, a promising research direction is to employ LLMs as substitutes for humans in game experiments, enabling social science research. However, despite numerous empirical researches on the combination of LLMs and game theory, the capability boundaries of LLMs in game theory remain unclear. In this research, we endeavor to systematically analyze LLMs in the context of game theory. Specifically, rationality, as the fundamental principle of game theory, serves as the metric for evaluating players' behavior --- building a clear desire, refining belief about uncertainty, and taking optimal actions. Accordingly, we select three classical games (dictator game, Rock-Paper-Scissors, and ring-network game) to analyze to what extent LLMs can achieve rationality in these three aspects. The experimental results indicate that even the current state-of-the-art LLM (GPT-4) exhibits substantial disparities compared to humans in game theory. For instance, LLMs struggle to build desires based on uncommon preferences, fail to refine belief from many simple patterns, and may overlook or modify refined belief when taking actions. Therefore, we consider that introducing LLMs into game experiments in the field of social science should be approached with greater caution.
Caoyun Fan, Jindou Chen, Yaohui Jin, Hao He 0007
AAAI3
2024 Self-Hint Prompting Improves Zero-shot Reasoning in Large Language Models via Reflective Cycle
Jindou Chen, Jidong Tian, Yaohui Jin
CogSci3
2024 Comparable Demonstrations Are Important In In-Context Learning: A Novel Perspective On Demonstration Selection
abstract
In-Context Learning (ICL) is an important paradigm for adapting Large Language Models (LLMs) to downstream tasks through a few demonstrations. Despite the great success of ICL, the limitation of the demonstration number may lead to demonstration bias, i.e. the input-label mapping induced by LLMs misunderstands the task’s essence. Inspired by human experience, we attempt to mitigate such bias through the perspective of the inter-demonstration relationship. Specifically, we construct Comparable Demonstrations (CDs) by minimally editing the texts to flip the corresponding labels, in order to highlight the task’s essence and eliminate potential spurious correlations through the inter-demonstration comparison. Through a series of experiments on CDs, we find that (1) demonstration bias does exist in LLMs, and CDs can significantly reduce such bias; (2) CDs exhibit good performance in ICL, especially in out-of-distribution scenarios. In summary, this study explores the ICL mechanisms from a novel perspective, providing a deeper insight into the demonstration selection strategy for ICL.
Caoyun Fan, Jidong Tian, Hao He 0007, Yaohui Jin
ICASSP5
2024 AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents
abstract
Evaluating large language models (LLMs) as general-purpose agents is essential for understanding their capabilities and facilitating their integration into practical applications. However, the evaluation process presents substantial challenges. A primary obstacle is the benchmarking of agent performance across diverse scenarios within a unified framework, especially in maintaining partially-observable environments and ensuring multi-round interactions. Moreover, current evaluation frameworks mostly focus on the final success rate, revealing few insights during the process and failing to provide a deep understanding of the model abilities. To address these challenges, we introduce AgentBoard, a pioneering comprehensive benchmark and accompanied open-source evaluation framework tailored to analytical evaluation of LLM agents. AgentBoard offers a fine-grained progress rate metric that captures incremental advancements as well as a comprehensive evaluation toolkit that features easy assessment of agents for multi-faceted analysis through interactive visualization. This not only sheds light on the capabilities and limitations of LLM agents but also propels the interpretability of their performance to the forefront. Ultimately, AgentBoard serves as a significant step towards demystifying agent behaviors and accelerating the development of stronger LLM agents.
Junlei Zhang, Cheng Yang 0007, Yujiu Yang 0001, Yaohui Jin, Zhen-Zhong Lan, Lingpeng Kong, Junxian He
NeurIPS6
2024 User re-identification via human mobility trajectories with siamese transformer networks
Bin Wang 0052, Yaohui Jin, Yanyan Xu 0002
Appl. Intell.5
2024 Unlock the Potential of Counterfactually-Augmented Data in Out-Of-Distribution Generalization
Caoyun Fan, Wenqing Chen, Jidong Tian, Hao He 0007, Yaohui Jin
Expert Syst. Appl.6
2024 PG² Net: Personalized and Group Preferences Guided Network for Next Place Prediction
abstract
Predicting the next destination is a key in human mobility behavior modeling, which is significant in various fields, such as epidemic control, urban planning, traffic management and recommendation. To achieve this, one typical solution is designing modules based on RNN to capture their preferences to various locations. Although these RNN-based methods can effectively learn individual’s hidden personalized preferences to her visited places, the interactions among users can only be weakly learned through the representations of locations. Targeting this, we propose an end-to-end framework named personalized and group preference guided network (PG2Net), considering the users’ preferences to various places at both individual and collective levels. Specifically, PG2Net concatenates Bi-LSTM and attention mechanism to capture each user’s long-term mobility tendency. To learn population’s group preferences, we utilize spatial and temporal information of the visitations to construct a spatial-temporal dependency module. We adopt a graph embedding method to map users’ trajectory into a hidden space, capturing their sequential relation. In addition, we devise an auxiliary loss to learn the vectorial representation of her next location. Experimental results on two Foursquare check-in datasets and one mobile phone dataset indicate the advantages of our model compared to the state-of-the-art baselines. Source code is available at https://github.com/urbanmobility/PG2Net.
Bin Wang 0052, Yaohui Jin, Yanyan Xu 0002
IEEE Trans. Intell. Transp. Syst.5
2024 Mitigating Bus Bunching via Hierarchical Multi-Agent Reinforcement Learning
abstract
Bus bunching is harmful to the efficiency and stability of bus transit systems, consequently delaying the arrival time of passengers and lowering the public transportation’s adoption rate. Traditional solutions adjust the additional holding time of buses at certain stations to mitigate this phenomenon. These methods sacrifice the system efficiency in exchange for even headway between neighboring buses. Little work focuses on optimizing multiple strategies when a single bus line not only has a bus bay to increase bus dwell time but also owns several dedicated bus lanes to accelerate. In this work, we develop a hierarchical multi-agent reinforcement learning (HMARL) framework to combine these two strategies. Speeding up certain buses via dedicated lanes can counteract the negative influence of additional holding time. Next, to support the two strategies, we devise a two-layer policy scheme, one for high-level policy deciding holding or accelerating and the other for low-level policy determining the specific dwell time or increase of speed. Besides, to handle the issue that the controlling actions of agents are asynchronous and temporally extended, we establish a duration-critic module based on the Recurrent Neural Networks (RNN) mechanism to model other agents’ impact during the period between two consecutive control. We evaluate the proposed framework on a simulated bus line with a quasi-real-world pattern to compare the performance of both traditional headway-based control methods and existing MARL methods. Results show that our method outperforms other baselines, not only stabilizing a strongly unstable bus line but also shortening the traveling times of passengers.
Mengdi Yu, Yaohui Jin, Yanyan Xu 0002
IEEE Trans. Intell. Transp. Syst.4
2024 Enhancing Unsupervised Anomaly Detection With Score-Guided Network
abstract
Anomaly detection plays a crucial role in various real-world applications, including healthcare and finance systems. Owing to the limited number of anomaly labels in these complex systems, unsupervised anomaly detection methods have attracted great attention in recent years. Two major challenges faced by the existing unsupervised methods are as follows: 1) distinguishing between normal and abnormal data when they are highly mixed together and 2) defining an effective metric to maximize the gap between normal and abnormal data in a hypothesis space, which is built by a representation learner. To that end, this work proposes a novel scoring network with a score-guided regularization to learn and enlarge the anomaly score disparities between normal and abnormal data, enhancing the capability of anomaly detection. With such score-guided strategy, the representation learner can gradually learn more informative representation during the model training stage, especially for the samples in the transition field. Moreover, the scoring network can be incorporated into most of the deep unsupervised representation learning (URL)-based anomaly detection models and enhances them as a plug-in component. We next integrate the scoring network into an autoencoder (AE) and four state-of-the-art models to demonstrate the effectiveness and transferability of the design. These score-guided models are collectively called SG-Models. Extensive experiments on both synthetic and real-world datasets confirm the state-of-the-art performance of SG-Models.
Zongyuan Huang, Longyuan Li, Yanyan Xu 0002, Yaohui Jin
IEEE Trans. Neural Networks Learn. Syst.6
2023 Preference-Controlled Multi-Objective Reinforcement Learning for Conditional Text Generation
abstract
Conditional text generation is to generate text sequences conditioning on linguistic or non-linguistic data. The main line of existing work proposed deterministic models to improve the fidelity of the generated text but often ignored the diversity. Another line relied on conditional variational auto-encoders (CVAEs), which increased the diversity over their deterministic backbones. However, CVAEs regard diversity as an implicit objective and may not be optimal. In this paper, we raise two questions: i) Can diversity be further improved with an explicit objective? ii) Since fidelity and diversity are two conflicting objectives, how can we obtain different multi-objective optimal solutions according to user preferences? To answer question i), we propose a multi-objective reinforcement learning (MORL) method which explicitly takes CIDEr and Self-CIDEr scores as the fidelity-oriented and diversity-oriented rewards respectively. To answer question ii), we propose a preference-controlled MORL method, which can obtain infinite multi-objective optimal solutions by tuning the preference variable. We conduct extensive experiments on paraphrasing and image captioning tasks, which show that in the fidelity-diversity trade-off space, our model outperforms both deterministic and CVAE-based baselines.
Wenqing Chen, Jidong Tian, Caoyun Fan, Hao He 0007, Yaohui Jin
AAAI6
2023 Latent Constraints on Unsupervised Text-Graph Alignment with Information Asymmetry
abstract
Unsupervised text-graph alignment (UTGA) is a fundamental task that bidirectionally generates texts and graphs without parallel data. Most available models of UTGA suffer from information asymmetry, a common phenomenon that texts and graphs include additional information invisible to each other. On the one hand, these models fail to supplement asymmetric information effectively due to the lack of ground truths. On the other hand, it is challenging to indicate asymmetric information with explicit indicators because it cannot be decoupled from the data directly. To address the challenge posed by information asymmetry, we propose the assumption that asymmetric information is encoded in unobservable latent variables and only affects the one-way generation processes. These latent variables corresponding to asymmetric information should obey prior distributions recovered approximately from original data. Therefore, we first propose a taxonomy of the latent variable that classifies the latent variable into transferrable (TV) and non-transferable (NTV) variables and further distinguish NTV as the dependent variable (DV) and the independent variable (IV). Next, we propose three latent VAE-based regularizations on TV, DV, and IV to constrain their distributions to well-designed prior distributions to introduce asymmetric information into models and enhance the preservation of shared contents. Finally, we impose the three proposed constraints on a cycle-consistent learning framework, back-translation (BT), named ConstrainedBT. Experimental results on three UTGA tasks demonstrate the effectiveness of ConstrainedBT on the information-asymmetric challenge.
Jidong Tian, Wenqing Chen, Caoyun Fan, Hao He 0007, Yaohui Jin
AAAI6
2023 Task-Level Thinking Steps Help Large Language Models for Challenging Classification Task
abstract
Large language models (LLMs) have shown incredible performance on many tasks such as dialogue generation, commonsense reasoning and question answering.In-context learning (ICL) is an important paradigm for adapting LLMs to the downstream tasks by prompting few demonstrations.However, the distribution of demonstrations can severely affect the performance, especially for challenging classification tasks.In this paper, we propose the concept of task-level thinking steps that can eliminate bias introduced by demonstrations.Further, to help LLMs distinguish confusing classes, we design a progressive revision framework, which can improve the thinking steps by correcting hard demonstrations.Experimental results prove the superiority of our proposed method, achieving best performance on three kinds of challenging classification tasks in the zero-shot and few-shot settings.Besides, with task-level thinking steps, automatically generated chain-of-thoughts (CoTs) bring more competitive performance.
Jidong Tian, Haoran Liao, Jindou Chen, Hao He 0007, Yaohui Jin
EMNLP6
2023 Chain-of-Thought Tuning: Masked Language Models can also Think Step By Step in Natural Language Understanding
abstract
Chain-of-Thought (CoT) is a technique that guides Large Language Models (LLMs) to decompose complex tasks into multi-step reasoning through intermediate steps in natural language form.Briefly, CoT enables LLMs to think step by step.However, although many Natural Language Understanding (NLU) tasks also require thinking step by step, LLMs perform less well than small-scale Masked Language Models (MLMs).To migrate CoT from LLMs to MLMs, we propose Chain-of-Thought Tuning (CoTT), a two-step reasoning framework based on prompt tuning, to implement step-by-step thinking for MLMs on NLU tasks.From the perspective of CoT, CoTT's two-step framework enables MLMs to implement task decomposition; CoTT's prompt tuning allows intermediate steps to be used in natural language form.Thereby, the success of CoT can be extended to NLU tasks through MLMs.To verify the effectiveness of CoTT, we conduct experiments on two NLU tasks: hierarchical classification and relation extraction, and the results show that CoTT outperforms baselines and achieves state-of-the-art performance.
Caoyun Fan, Jidong Tian, Wenqing Chen, Hao He 0007, Yaohui Jin
EMNLP6
2023 Improving the out-of-Distribution Generalization Capability of Language Models: Counterfactually-Augmented Data is not Enough
abstract
Counterfactually-Augmented Data (CAD) has the potential to improve language models’ Out-Of-Distribution (OOD) generalization capability, as CAD induces language models to exploit causal features and exclude spurious correlations. However, the empirical results of OOD generalization on CAD are not as efficient as expected. In this paper, we attribute the inefficiency to Myopia Phenomenon caused by CAD: language models only focus on causal features that are edited in the augmentation and exclude other non-edited causal features. As a result, the potential of CAD is not fully exploited. Based on the structural properties of CAD, we design two additional constraints to help language models extract more complete causal features contained in CAD, thus improving the OOD generalization capability. We evaluate our method on two tasks: Sentiment Analysis and Natural Language Inference, and the experimental results demonstrate that our method could unlock CAD’s potential and improve language models’ OOD generalization capability.
Caoyun Fan, Wenqing Chen, Jidong Tian, Hao He 0007, Yaohui Jin
ICASSP6
2023 Contrast with major classifier vectors for federated medical relation extraction with heterogeneous label distribution
Hao He 0007, Yaohui Jin
Appl. Intell.3
2023 Accurate use of label dependency in multi-label text classification through the lens of causality
Caoyun Fan, Wenqing Chen, Jidong Tian, Hao He 0007, Yaohui Jin
Appl. Intell.6
2023 Learning Robust Deep State Space for Unsupervised Anomaly Detection in Contaminated Time-Series
abstract
Anomalies are ubiquitous in real-world time-series data which call for effective and timely detection, especially in an unsupervised setting for labeling cost saving. In this paper, we develop an unsupervised density reconstruction model for multi-dimensional time-series anomaly detection. In particular, it directly handles an important realistic setting that the detection is achieved towards raw time-series contaminated with noise for training, in contrast to most existing anomaly detection works that assume the training data is in general clean i.e. not contaminated with anomaly. It extends recent advancements in deep generative models and state space models to achieve robust anomaly detection. Our approach comprises of a novel state space based generative model, a filtering based inference model, together with a carefully-designated emission model based on robust statistics theory. Extensive experimental results are conducted to show that our approach can adapt to complex patterns even given severely contaminated training data. We also develop visualization techniques to help better understand the behavior of the anomaly detection models. Empirical results show that our method outperforms state-of-the-arts on both synthetic and real-world datasets.
Longyuan Li, Junchi Yan, Qingsong Wen, Yaohui Jin, Xiaokang Yang 0001
IEEE Trans. Knowl. Data Eng.4
2023 Learning Generative RNN-ODE for Collaborative Time-Series and Event Sequence Forecasting
abstract
Time-series and event sequences are widely collected data types in real-world applications. Modeling and forecasting of such temporal data play an important role in an informed decision-making process. A major limitation of previous methods is that they either focus on time-series or events, rather than the combination of the two worlds. In fact, the two types of data often provide complementary information, emphasizing the necessity of jointly modeling the both. In this paper, we propose the RNN-ODE collaborative model for joint modeling and forecasting of heterogeneous time-series and event sequence data, which combines several useful techniques from both Bayesian and deep learning for its interpretability. Specifically, we devise a tailored encoder to combine the advances in deep temporal point processes models and variational recurrent neural networks. To predict the probability of event occurrence over an arbitrary continuous-time horizon, we base our model on the mathematical foundation of Neural Ordinary Differential Equations (NODE). Extensive experimental results on simulations and real data sets show that compared with existing methods, our integrated approach can achieve more competitive forecasting performance of both time-series and event sequences.
Longyuan Li, Junchi Yan, Jihai Zhang 0002, Yaohui Jin, Xiaokang Yang 0001
IEEE Trans. Knowl. Data Eng.6
2022 Weakly Supervised Neural Symbolic Learning for Cognitive Tasks
abstract
Despite the recent success of end-to-end deep neural networks, there are growing concerns about their lack of logical reasoning abilities, especially on cognitive tasks with perception and reasoning processes. A solution is the neural symbolic learning (NeSyL) method that can effectively utilize pre-defined logic rules to constrain the neural architecture making it perform better on cognitive tasks. However, it is challenging to apply NeSyL to these cognitive tasks because of the lack of supervision, the non-differentiable manner of the symbolic system, and the difficulty to probabilistically constrain the neural network. In this paper, we propose WS-NeSyL, a weakly supervised neural symbolic learning model for cognitive tasks with logical reasoning. First, WS-NeSyL employs a novel back search algorithm to sample the possible reasoning process through logic rules. This sampled process can supervise the neural network as the pseudo label. Based on this algorithm, we can backpropagate gradients to the neural network of WS-NeSyL in a weakly supervised manner. Second, we introduce a probabilistic logic regularization into WS-NeSyL to help the neural network learn probabilistic logic. To evaluate WS-NeSyL, we have conducted experiments on three cognitive datasets, including temporal reasoning, handwritten formula recognition, and relational reasoning datasets. Experimental results show that WS-NeSyL not only outperforms the end-to-end neural model but also beats the state-of-the-art neural symbolic learning models.
Jidong Tian, Wenqing Chen, Liqiang Xiao, Hao He 0007, Yaohui Jin
AAAI6
2022 MaxGNR: A Dynamic Weight Strategy via Maximizing Gradient-to-Noise Ratio for Multi-task Learning
Caoyun Fan, Wenqing Chen, Jidong Tian, Hao He 0007, Yaohui Jin
ACCV (1)6
2022 To What Extent Do Natural Language Understanding Datasets Correlate to Logical Reasoning? A Method for Diagnosing Logical Reasoning
abstract
Reasoning and knowledge-related skills are considered as two fundamental skills for natural language understanding (NLU) tasks such as machine reading comprehension (MRC) and natural language inference (NLI). However, it is not clear to what extent an NLU task defined on a dataset correlates to a specific NLU skill. On the one hand, evaluating the correlation requires an understanding of the significance of the NLU skill in a dataset. Significance judges whether a dataset includes sufficient material to help the model master this skill. On the other hand, it is also necessary to evaluate the dependence of the task on the NLU skill. Dependence is a measure of how much the task defined on a dataset depends on the skill. In this paper, we propose a systematic method to diagnose the correlations between an NLU dataset and a specific skill, and then take a fundamental reasoning skill, logical reasoning, as an example for analysis. The method adopts a qualitative indicator to indicate the significance while adopting a quantitative indicator to measure the dependence. We perform diagnosis on 8 MRC datasets (including two types) and 3 NLI datasets and acquire intuitively reasonable results. We then perform the analysis to further understand the results and the proposed indicators. Based on the analysis, although the diagnostic method has some limitations, it is still an effective method to perform a basic diagnosis of the correlation between the dataset and logical reasoning skill, which also can be generalized to other NLU skills.
Jidong Tian, Wenqing Chen, Caoyun Fan, Hao He 0007, Yaohui Jin
COLING6
2022 Open Source MagicData-RAMC: A Rich Annotated Mandarin Conversational(RAMC) Speech Dataset
abstract
This paper introduces a high-quality rich annotated Mandarin conversational (RAMC) speech dataset called MagicData-RAMC.The MagicData-RAMC corpus contains 180 hours of conversational speech data recorded from native speakers of Mandarin Chinese over mobile phones with a sampling rate of 16 kHz.The dialogs in MagicData-RAMC are classified into 15 diversified domains and tagged with topic labels, ranging from science and technology to ordinary life.Accurate transcription and precise speaker voice activity timestamps are manually labeled for each sample.Speakers' detailed information is also provided.As a Mandarin speech dataset designed for dialog scenarios with high quality and rich annotations, MagicData-RAMC enriches the data diversity in the Mandarin speech community and allows extensive research on a series of speechrelated tasks, including automatic speech recognition, speaker diarization, topic detection, keyword search, text-to-speech, etc.We also conduct several relevant tasks and provide experimental results to help evaluate the dataset.
Zehui Yang, Runyan Yang, Lingxuan Ye, Gaofeng Cheng, Yaohui Jin, Pengyuan Zhang, Lei Xie 0001, Yonghong Yan 0002
INTERSPEECH8
2022 CALM: Commen-Sense Knowledge Augmentation for Document Image Understanding
abstract
Performance of document image understanding has been significantly fueled by encoding multi-modal information in recent years. However, existing works heavily rely on the superficial appearance of the observed data, resulting in counter-intuitive model behavior in many critical cases. To overcome this issue, this paper proposes a common-sense knowledge augmented model CALM for document image understanding tasks. It firstly produces purified representations of document contents to extract key information and learn common-sense augmented representation for inputs. Then, relevant common-sense knowledge is extracted from the external ConceptNet knowledge base, and a derived knowledge graph is built to enhance the common-sense reasoning capability of CALM jointly. In order to further highlight the importance of common-sense knowledge in document image understanding, we propose the first question-answering dataset, CS-DVQA, focused on common-sense reasoning for document images, in which questions are answered by taking both document contents and common-sense knowledge into consideration. Through extensive evaluation, the proposed CALM approach outperforms the state-of-the-art models in three document image understanding tasks, including key information extraction(from 85.37 to 86.52), document image classification(from 96.08 to 96.17), document visual question answering(from 86.72 to 88.03).
Qinyi Du, Keqian Li, Jidong Tian, Liqiang Xiao, Yaohui Jin
ACM Multimedia6
2022 FusionSum: Abstractive summarization with sentence fusion and cooperative reinforcement learning
Liqiang Xiao, Hao He 0007, Yaohui Jin
Knowl. Based Syst.3
2021 Synergetic Learning of Heterogeneous Temporal Sequences for Multi-Horizon Probabilistic Forecasting
abstract
Time-series is ubiquitous across applications, such as transportation, finance and healthcare. Time-series is often influenced by external factors, especially in the form of asynchronous events, making forecasting difficult. However, existing models are mainly designated for either synchronous time-series or asynchronous event sequence, and can hardly provide a synthetic way to capture the relation between them. We propose Variational Synergetic Multi-Horizon Network (VSMHN), a novel deep conditional generative model. To learn complex correlations across heterogeneous sequences, a tailored encoder is devised to combine the advances in deep point processes models and variational recurrent neural networks. In addition, an aligned time coding and an auxiliary transition scheme are carefully devised for batched training on unaligned sequences. Our model can be trained effectively using stochastic variational inference and generates probabilistic predictions with Monte-Carlo simulation. Furthermore, our model produces accurate, sharp and more realistic probabilistic forecasts. We also show that modeling asynchronous event sequences is crucial for multi-horizon time-series forecasting.
Longyuan Li, Jihai Zhang 0002, Junchi Yan, Yaohui Jin, Yanjie Duan, Guangjian Tian
AAAI4
2021 De-Confounded Variational Encoder-Decoder for Logical Table-to-Text Generation
abstract
Wenqing Chen, Jidong Tian, Yitian Li, Hao He, Yaohui Jin. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Wenqing Chen, Jidong Tian, Hao He 0007, Yaohui Jin
ACL/IJCNLP (1)5
2021 Diagnosing the First-Order Logical Reasoning Ability Through LogicNLI
abstract
Recently, language models (LMs) have achieved significant performance on many NLU tasks, which has spurred widespread interest for their possible applications in the scientific and social area.However, LMs have faced much criticism of whether they are truly capable of reasoning in NLU.In this work, we propose a diagnostic method for first-order logic (FOL) reasoning with a new proposed benchmark, LogicNLI.LogicNLI is an NLI-style dataset that effectively disentangles the target FOL reasoning from commonsense inference and can be used to diagnose LMs from four perspectives: accuracy, robustness, generalization, and traceability.Experiments on BERT, RoBERTa, and XLNet, have uncovered the weaknesses of these LMs on FOL reasoning, which motivates future exploration to enhance the reasoning ability.
Jidong Tian, Wenqing Chen, Liqiang Xiao, Hao He 0007, Yaohui Jin
EMNLP (1)6
2021 End-to-End Conversational Search for Online Shopping with Utterance Transfer
abstract
Liqiang Xiao, Jun Ma, Xin Luna Dong, Pascual Martínez-Gómez, Nasser Zalmout, Wei Chen, Tong Zhao, Hao He, Yaohui Jin. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021.
Liqiang Xiao, Jun Ma 0029, Xin Dong 0001, Pascual Martínez-Gómez, Nasser Zalmout, Tong Zhao 0002, Hao He 0007, Yaohui Jin
EMNLP (1)9
2021 Learning Hierarchical Reasoning for Text-Based Visual Question Answering
Caiyuan Li, Qinyi Du, Yaohui Jin
ICANN (3)4
2021 Dependent Multi-Task Learning with Causal Intervention for Image Captioning
abstract
Recent work for image captioning mainly followed an extract-then-generate paradigm, pre-extracting a sequence of object-based features and then formulating image captioning as a single sequence-to-sequence task. Although promising, we observed two problems in generated captions: 1) content inconsistency where models would generate contradicting facts; 2) not informative enough where models would miss parts of important information. From a causal perspective, the reason is that models have captured spurious statistical correlations between visual features and certain expressions (e.g., visual features of "long hair" and "woman"). In this paper, we propose a dependent multi-task learning framework with the causal intervention (DMTCI). Firstly, we involve an intermediate task, bag-of-categories generation, before the final task, image captioning. The intermediate task would help the model better understand the visual features and thus alleviate the content inconsistency problem. Secondly, we apply Pearl's do-calculus on the model, cutting off the link between the visual features and possible confounders and thus letting models focus on the causal visual features. Specifically, the high-frequency concept set is considered as the proxy confounders where the real confounders are inferred in the continuous space. Finally, we use a multi-agent reinforcement learning (MARL) strategy to enable end-to-end training and reduce the inter-task error accumulations. The extensive experiments show that our model outperforms the baseline models and achieves competitive performance with state-of-the-art models.
Wenqing Chen, Jidong Tian, Caoyun Fan, Hao He 0007, Yaohui Jin
IJCAI5
2021 Towards Reasoning Ability in Scene Text Visual Question Answering
abstract
Works on scene text visual question answering (TextVQA) always emphasize the importance of reasoning questions and image contents. However, we find current TextVQA models lack reasoning ability and tend to answer questions by exploiting dataset bias and language priors. Moreover, our observations indicate that recent accuracy improvement in TextVQA is mainly contributed by stronger OCR engines, better pre-training strategies and more Transformer layers, instead of newly proposed networks. In this work, towards the reasoning ability, we 1) conduct module-wise contribution analysis to quantitatively investigate how existing works improve accuracies in TextVQA; 2) design a gradient-based explainability method to explore why TextVQA models answer what they answer and find evidence for their predictions; 3) perform qualitative experiments to visually analyze models reasoning ability and explore potential reasons behind such a poor ability.
Liqiang Xiao, Yue Lu 0001, Yaohui Jin, Hao He 0007
ACM Multimedia4
2021 Anomaly Detection of Time Series With Smoothness-Inducing Sequential Variational Auto-Encoder
abstract
Deep generative models have demonstrated their effectiveness in learning latent representation and modeling complex dependencies of time series. In this article, we present a smoothness-inducing sequential variational auto-encoder (VAE) (SISVAE) model for the robust estimation and anomaly detection of multidimensional time series. Our model is based on VAE, and its backbone is fulfilled by a recurrent neural network to capture latent temporal structures of time series for both the generative model and the inference model. Specifically, our model parameterizes mean and variance for each time-stamp with flexible neural networks, resulting in a nonstationary model that can work without the assumption of constant noise as commonly made by existing Markov models. However, such flexibility may cause the model fragile to anomalies. To achieve robust density estimation which can also benefit detection tasks, we propose a smoothness-inducing prior over possible estimations. The proposed prior works as a regularizer that places penalty at nonsmooth reconstructions. Our model is learned efficiently with a novel stochastic gradient variational Bayes estimator. In particular, we study two decision criteria for anomaly detection: reconstruction probability and reconstruction error. We show the effectiveness of our model on both synthetic data sets and public real-world benchmarks.
Longyuan Li, Junchi Yan, Yaohui Jin
IEEE Trans. Neural Networks Learn. Syst.4
2020 Copy or Rewrite: Hybrid Summarization with Hierarchical Reinforcement Learning
abstract
Jointly using the extractive and abstractive summarization methods can combine their complementary advantages, generating both informative and concise summary. Existing methods that adopt an extract-then-abstract strategy have achieved impressive results, yet they suffer from the information loss in the abstraction step because they compress all the selected sentences without distinguish. Especially when the whole sentence is summary-worthy, salient content would be lost by compression. To address this problem, we propose HySum, a hybrid framework for summarization that can flexibly switch between copying sentence and rewriting sentence according to the degree of redundancy. In this way, our approach can effectively combine the advantages of two branches of summarization, juggling informativity and conciseness. Moreover, we based on Hierarchical Reinforcement Learning, propose an end-to-end reinforcing method to bridge together the extraction module and rewriting module, which can enhance the cooperation between them. Automatic evaluation shows that our approach significantly outperforms the state-of-the-arts on the CNN/DailyMail corpus. Human evaluation also demonstrates that our generated summaries are more informative and concise than popular models.
Liqiang Xiao, Lu Wang 0008, Hao He 0007, Yaohui Jin
AAAI4
2020 A Semantically Consistent and Syntactically Variational Encoder-Decoder Framework for Paraphrase Generation
abstract
Paraphrase generation aims to generate semantically consistent sentences with different syntactic realizations.Most of the recent studies rely on the typical encoder-decoder framework where the generation process is deterministic.However, in practice, the ability to generate multiple syntactically different paraphrases is important.Recent work proposed to cooperate variational inference on a target-related latent variable to introduce the diversity.But the latent variable may be contaminated by the semantic information of other unrelated sentences, and in turn, change the conveyed meaning of generated paraphrases.In this paper, we propose a semantically consistent and syntactically variational encoder-decoder framework, which uses adversarial learning to ensure the syntactic latent variable be semantic-free.Moreover, we adopt another discriminator to improve the word-level and sentence-level semantic consistency.So the proposed framework can generate multiple semantically consistent and syntactically different paraphrases.The experiments show that our model outperforms the baseline models on the metrics based on both n-gram matching and semantic similarity, and our model can generate multiple different paraphrases by assembling different syntactic variables.
Wenqing Chen, Jidong Tian, Liqiang Xiao, Hao He 0007, Yaohui Jin
COLING5
2020 Exploring Logically Dependent Multi-task Learning with Causal Inference
abstract
Previous studies have shown that hierarchical multi-task learning (MTL) can utilize task dependencies by stacking encoders and outperform democratic MTL.However, stacking encoders only considers the dependencies of feature representations and ignores the label dependencies in logically dependent tasks.Furthermore, how to properly utilize the labels remains an issue due to the cascading errors between tasks.In this paper, we view logically dependent MTL from the perspective of causal inference and suggest a mediation assumption instead of the confounding assumption in conventional MTL models.We propose a model including two key mechanisms: label transfer (LT) for each task to utilize the labels of all its lower-level tasks, and Gumbel sampling (GS) to deal with cascading errors.In the field of causal inference, GS in our model is essentially a counterfactual reasoning process, trying to estimate the causal effect between tasks and utilize it to improve MTL.We conduct experiments on two English datasets and one Chinese dataset.Experiment results show that our model achieves state-of-the-art on six out of seven subtasks and improves predictions' consistency.
Wenqing Chen, Jidong Tian, Liqiang Xiao, Hao He 0007, Yaohui Jin
EMNLP (1)5
2020 Modeling Content Importance for Summarization with Pre-trained Language Models
abstract
Modeling content importance is an essential yet challenging task for summarization.Previous work is mostly based on statistical methods that estimate word-level salience, which does not consider semantics and larger context when quantifying importance.It is thus hard for these methods to generalize to semantic units of longer text spans.In this work, we apply information theory on top of pretrained language models and define the concept of importance from the perspective of information amount.It considers both the semantics and context when evaluating the importance of each semantic unit.With the help of pre-trained language models, it can easily generalize to different kinds of semantic units (n-grams or sentences).Experiments on CNN/Daily Mail and New York Times datasets demonstrate that our method can better model the importance of content than prior work based on F1 and ROUGE scores.
Liqiang Xiao, Lu Wang 0008, Hao He 0007, Yaohui Jin
EMNLP (1)4
2019 Learning Interpretable Deep State Space Model for Probabilistic Time Series Forecasting
abstract
Probabilistic time series forecasting involves estimating the distribution of future based on its history, which is essential for risk management in downstream decision-making. We propose a deep state space model for probabilistic time series forecasting whereby the non-linear emission model and transition model are parameterized by networks and the dependency is modeled by recurrent neural nets. We take the automatic relevance determination (ARD) view and devise a network to exploit the exogenous variables in addition to time series. In particular, our ARD network can incorporate the uncertainty of the exogenous variables and eventually helps identify useful exogenous variables and suppress those irrelevant for forecasting. The distribution of multi-step ahead forecasts are approximated by Monte Carlo simulation. We show in experiments that our model produces accurate and sharp probabilistic forecasts. The estimated uncertainty of our forecasting also realistically increases over time, in a spontaneous manner.
Longyuan Li, Junchi Yan, Xiaokang Yang 0001, Yaohui Jin
IJCAI4
2019 TransMS: Knowledge Graph Embedding for Complex Relations by Multidirectional Semantics
abstract
Knowledge graph embedding, which projects the symbolic relations and entities onto low-dimension continuous spaces, is essential to knowledge graph completion. Recently, translation-based embedding models (e.g. TransE) have aroused increasing attention for their simplicity and effectiveness. These models attempt to translate semantics from head entities to tail entities with the relations and infer richer facts outside the knowledge graph. In this paper, we propose a novel knowledge graph embedding method named TransMS, which translates and transmits multidirectional semantics: i) the semantics of head/tail entities and relations to tail/head entities with nonlinear functions and ii) the semantics from entities to relations with linear bias vectors. Our model has merely one additional parameter α than TransE for each triplet, which results in its better scalability in large-scale knowledge graph. Experiments show that TransMS achieves substantial improvements against state-of-the-art baselines, especially the Hit@10s of head entity prediction for N-1 relations and tail entity prediction for 1-N relations improved by about 27.1% and 24.8% on FB15K database respectively.
Shihui Yang 0001, Jidong Tian, Honglun Zhang, Junchi Yan, Hao He 0007, Yaohui Jin
IJCAI6
2019 Differential Privacy with Variant-Noise for Gaussian Processes Classification
Zhili Xiong, Longyuan Li, Junchi Yan, Hao He 0007, Yaohui Jin
PRICAI (3)6
2019 Online detection of abnormal passenger out-flow in urban metro system
Longyuan Li, Pingjun Pan, Yongkun Wang, Yaohui Jin
Neurocomputing5
2018 Learning What to Share: Leaky Multi-Task Network for Text Classification
abstract
Neural network based multi-task learning has achieved great success on many NLP problems, which focuses on sharing knowledge among tasks by linking some layers to enhance the performance. However, most existing approaches suffer from the interference between tasks because they lack of selection mechanism for feature sharing. In this way, the feature spaces of tasks may be easily contaminated by helpless features borrowed from others, which will confuse the models for making correct prediction. In this paper, we propose a multi-task convolutional neural network with the Leaky Unit, which has memory and forgetting mechanism to filter the feature flows between tasks. Experiments on five different datasets for text classification validate the benefits of our approach.
Liqiang Xiao, Honglun Zhang, Wenqing Chen, Yongkun Wang, Yaohui Jin
COLING5
2018 MCapsNet: Capsule Network for Text with Multi-Task Learning
abstract
Multi-task learning has an ability to share the knowledge among related tasks and implicitly increase the training data.However, it has long been frustrated by the interference among tasks.This paper investigates the performance of capsule network for text, and proposes a capsule-based multi-task learning architecture, which is unified, simple and effective.With the advantages of capsules for feature clustering, proposed task routing algorithm can cluster the features for each task in the network, which helps reduce the interference among tasks.Experiments on six text classification datasets demonstrate the effectiveness of our models and their characteristics for feature clustering.
Liqiang Xiao, Honglun Zhang, Wenqing Chen, Yongkun Wang, Yaohui Jin
EMNLP5
2018 Multi-Task Label Embedding for Text Classification
abstract
Multi-task learning in text classification leverages implicit correlations among related tasks to extract common features and yield performance gains.However, a large body of previous work treats labels of each task as independent and meaningless one-hot vectors, which cause a loss of potential label information.In this paper, we propose Multi-Task Label Embedding to convert labels in text classification into semantic vectors, thereby turning the original tasks into vector matching tasks.Our model utilizes semantic correlations among tasks and makes it convenient to scale or transfer when new tasks are involved.Extensive experiments on five benchmark datasets for text classification show that our model can effectively improve the performances of related tasks with semantic representations of labels and additional information from each other.
Honglun Zhang, Liqiang Xiao, Wenqing Chen, Yongkun Wang, Yaohui Jin
EMNLP5
2018 Transformable Convolutional Neural Network for Text Classification
abstract
Convolutional neural networks (CNNs) have shown their promising performance for natural language processing tasks, which extract n-grams as features to represent the input. However, n-gram based CNNs are inherently limited to fixed geometric structure and cannot proactively adapt to the transformations of features. In this paper, we propose two modules to provide CNNs with the flexibility for complex features and the adaptability for transformation, namely, transformable convolution and transformable pooling. Our method fuses dynamic and static deviations to redistribute the sampling locations, which can capture both current and global transformations. Our modules can be easily integrated by other models to generate new transformable networks. We test proposed modules on two state-of-the-art models, and the results demonstrate that our modules can effectively adapt to the feature transformation in text classification.
Liqiang Xiao, Honglun Zhang, Wenqing Chen, Yongkun Wang, Yaohui Jin
IJCAI5
2018 Generative Warfare Nets: Ensemble via Adversaries and Collaborators
abstract
Generative Adversarial Nets are a powerful method for training generative models of complex data, where a Generator and a Discriminator confront with each other and get optimized in a two-player minmax manner. In this paper, we propose the Generative Warfare Nets (GWN) that involve multiple generators and multiple discriminators from two sides to exploit the advantages of Ensemble Learning. We maintain the authorities for the generators and the discriminators to enhance inter-side interactions, and utilize the mechanisms of imitation and innovation to model intra-side interactions among the generators, where they can not only learn from but also compete with each other. Extensive experiments on three natural image datasets show that GWN can achieve state-of-the-art Inception scores and produce diverse high-quality synthetic results.
Honglun Zhang, Liqiang Xiao, Wenqing Chen, Yongkun Wang, Yaohui Jin
IJCAI5
2018 Troubleshooting Data Plane With Rule Verification in Software-Defined Networks
abstract
Data plane network issues, caused by software bugs or hardware failures inside network devices, usually manifest themselves as failed rules, which can be verified by the comparison between actual and desired network behavior. Previous efforts implement this leveraging end-to-end active probing, which falls short on timely locating the exact failure points and identifying responsible rules. With the help of the out-band channel in software-defined network (SDN), per hop active probing can be employed to enable per device or even per rule network behavior comparison to verify the effectiveness of all the rules. Therefore, we present SERVE, an SDN-enabled rule verification framework that automatically identifies data plane network issues with periodic per rule network behavior comparison. By modeling network devices as stateful multi-rooted trees with respect to pipeline processing, a small number of probes can be generated from a per device perspective in a timely manner. We evaluate the performance of SERVE's probe generation and validate SERVE's effectiveness on a small deployment with typical use cases.
Yusu Zhao, Yongkun Wang, Yaohui Jin
IEEE Trans. Netw. Serv. Manag.4
2017 A Generalized Recurrent Neural Architecture for Text Classification with Multi-Task Learning
abstract
Multi-task learning leverages potential correlations among related tasks to extract common features and yield performance gains. However, most previous works only consider simple or weak interactions, thereby failing to model complex correlations among three or more tasks. In this paper, we propose a multi-task learning architecture with four types of recurrent neural layers to fuse information across multiple related tasks. The architecture is structurally flexible and considers various interactions among tasks, which can be regarded as a generalized case of many previous works. Extensive experiments on five benchmark datasets for text classification show that our model can significantly improve performances of related tasks with additional information from others.
Honglun Zhang, Liqiang Xiao, Yongkun Wang, Yaohui Jin
IJCAI4
2017 SDN enhanced tomography for performance profiling in cloud network
abstract
For cloud network performance profiling, network tomography is useful for deducing the network performance based on end‐to‐end measurement. However, most tomography problems are under‐constrained, thus requires additional assumptions in order to be solvable, which sacrifices the accuracy. On the other hand, packet traces from switches could provide accurate and direct performance measurement, but it is hard to cover the whole network with packet trace analysis per link and flow. In this study, the authors propose ScoutFlow, a method combining software‐defined networking (SDN) flow measurement and end‐to‐end performance tomography, to achieve accurate performance profiling for cloud network while keeping low monitoring overhead. In ScoutFlow, they mirror the flow packet trace using SDN, to solve the under‐constrained problem in tomography. ScoutFlow only requires a small amount of flow mirror traces for the measurement, which leads to much lower overhead of flow mirroring than that of traditional packet‐level monitoring methods. The proposed methodology is evaluated with simulation and testbed experiments, which demonstrates ScoutFlow's scalability and accuracy.
Yusu Zhao, Yongkun Wang, Yaohui Jin
IET Commun.4
2017 Discovering and modeling meta-structures in human behavior from city-scale cellular data
Xiaming Chen, Siwei Qiang, Yongkun Wang, Yaohui Jin
Pervasive Mob. Comput.5
2016 Netography: Troubleshoot your network with packet behavior in SDN
abstract
Network troubleshooting is always a tough and daunting task for network operators to struggle with, due to difficultly observed network state, large network size and limited tools such as ping and traceroute. Software-defined networking (SDN) brings us the centralized network control and the customized network management over the entire network, which enables new ways of network troubleshooting. Previous efforts focus on static checking, passive monitoring and active probing, which rely on scraping rules from either controller or network devices. Since those rules describe the actions supposed to be performed on packets, the actions having been performed on packets can be different. We propose the concept of packet behavior to describe the real changes of packets and highlight its importance towards network troubleshooting. Based on the novel approach of exporting packet behavior and flow rules via copies triggered by probes being actively sent, we present the design of Netography system and illustrate the procedures of troubleshooting tasks regarding forwarding errors as well as performance degradation caused by non-tenant-contention reasons. We implement a prototype and verify our system on a small deployment with three typical use cases.
Yusu Zhao, Yaohui Jin
NOMS3
2016 Passive profiling of mobile engaging behaviours via user-end application performance assessment
Xiaming Chen, Siwei Qiang, Jianwen Wei, Kaida Jiang, Yaohui Jin
Pervasive Mob. Comput.5
2015 Analyzing and modeling spatio-temporal dependence of cellular traffic at city scale
abstract
Traffic characteristics over space and time constitute an important aspect of cellular networks in consideration of resource provision, traffic engineering and system optimization. Despite recent progress in revealing temporal dynamics and spatial inhomogeneity of cellular traffic, limited knowledge about traffic dependence is gained. One of challenges comes from the absence of sustained observations at a network-wide scale. In this paper, we make an analysis on the week-long traffic generated by a large population of users in a city of China, and model traffic dependence along both space and time dimensions. The evaluation results suggest connections between spatio-temporal dependence of cellular traffic and the organization of human lives. Region differences are observed to impact traffic dependence to a great extent. Additionally, interactive knowledge between space and time enhances traffic prediction with a decrease in root-mean-square error of 2.8%~25.2%. We believe that these achievements will benefit multiple research and development areas such as network deploying and simulation researches.
Xiaming Chen, Yaohui Jin, Siwei Qiang, Weisheng Hu, Kaida Jiang
ICC2
2015 FT-INDEX: A distributed indexing scheme for switch-centric cloud storage system
abstract
Nowadays, cloud storage systems may contain tens of thousands of servers and large scale data sets, which significantly require efficient data management scheme and query processing mechanism. To fulfill these requirements in modern data centers, the infrastructure of cloud systems, we propose FT-Index, a secondary indexing scheme for cloud system with switch-centric topology. FT-Index has a two-layer design. The upper-layer index, called global index, is distributed across different hosts in the system, while the lower-layer index, named local index, is a B+-tree for local query. We further adopt the Interval tree to reorganize the global index and propose two versions of FT-Index with different publishing methods to lower the rate of false positives and reduce the cost of forwarding queries. We provide detailed theoretical analysis on the upper bound of false positives, physical hops per query, and the relationship between them. We also conduct abundant experiments to validate the efficiency of FT-Index.
Xiaofeng Gao 0001, Binjie Li, Zongchen Chen, Maofan Yin, Guihai Chen, Yaohui Jin
ICC6
2015 Intelligent timeout master: Dynamic timeout for SDN-based data centers
abstract
Software Defined Networks (SDN) such as OpenFlow provides better network management for data center by decoupling control plane from data plane. Current OpenFlow controllers install flow rules with a fixed timeout after which the switch automatically removes the rules from its flow table. However, this fixed timeout has shown many disadvantages. For flows with short packet interval, the timeout may be too large so that flow rules stay in the flow table for too long time and result in unnecessary occupation of flow table; for flows with long packet interval or periodic flows, the timeout may be too short, hence producing too many packet-in events and causing overload on the controller. In this paper, we propose the Intelligent Timeout Master, which can assign suitable timeout to different flows according to their characteristics, as well as conduct a feedback control to adjust the max timeout value according to the current flow table occupation, in an effort to avoid flow table overflow. In our experiments, we use real traffic trace and the result confirms that our Intelligent Timeout Master performs quite well in reducing the number of packet-in events as well as flow table occupation.
Huikang Zhu, Hongbo Fan, Yaohui Jin
IM4
2014 NICE: Network-aware VM Consolidation scheme for Energy Conservation in Data Centers
abstract
Energy conservation and network performance have become two of the most important issues in data center as the scale of cloud services continues growing. Recent researches usually consider these two issues separately. Energy conservation mainly deals with hosts, which reduces total energy consumption by consolidating virtual machine(VM)s to fewer hosts, and network performance mainly deals with network scalability and energy efficiency, which improves data center network(DCN) scalability by applying new network topologies or routing schemes and improves DCN energy efficiency by consolidating trafile. In this paper, we jointly consider these two issues and define Combined VM Consolidation (CVC) problem. We prove that CVC is NP-complete and is inapproximable by a factor of 3/2 ε unless P = NP. Next, we propose NICE: Network-aware VM Consolidation scheme for Energy Conservation in Data CEnter to solve CVC. Instead of taking the unrealistic hypothesis that migration cost is negligible, a common assumption in most literatures, we precisely analyze VM migration cost according to real-trace experiments in a 6-server testbed via VMware. Massive simulations validate the efficiency of NICE, In all, to the best of our knowledge, we arc the first work to combine VM consolidation with network optimization and migration cost.
Xiaofeng Gao 0001, Guihai Chen, Yaohui Jin
ICPADS4
2014 Characterizing home network traffic: an inside view
Kuai Xu, Feng Wang 0002, Lin Gu 0001, Yaohui Jin
Pers. Ubiquitous Comput.5
2013 ReSurf: Reconstructing web-surfing activity from network traffic
Guowu Xie, Marios Iliofotou, Thomas Karagiannis, Michalis Faloutsos, Yaohui Jin
Networking5
2012 Minimizing mean packet delay in EPONs through Integrated Grant Scheduling
abstract
In Ethernet Passive Optical Networks (EPONs) with offline Dynamic Bandwidth Allocation (DBA) framework, the Optical Line Terminal (OLT) will first collect bandwidth requests from all Optical Network Units (ONUs), and then make bandwidth allocation and scheduling decisions for the shared upstream channel. Due to varying Round-Trip Time (RTT) and grant window sizes, the transmission order of ONUs will greatly affect the mean packet delay. In this paper, we address this grant scheduling problem aiming at minimizing the mean packet delay. We prove several theorems which could determine the ONU to transmit first which will minimize the mean packet delay. And then, by iteratively using these theorems, we propose an Integrated Grant Scheduling (IGS) algorithm to determine the optimal transmission order to minimize the mean packet delay. We then conduct simulations to evaluate the performance of our algorithm, and find that the performance of our algorithm is better compared with other algorithms under various conditions.
Qingqi Shi, Wei Guo 0003, Zhe Liu 0003, Yaohui Jin, Weiqiang Sun, Weisheng Hu
ICC4
2012 Characterizing Home Network Traffic: An Inside View
Kuai Xu, Feng Wang 0002, Lin Gu 0001, Yaohui Jin
WASA5
2011 Combining sFlow and Tracker traffic analysis: A novel estimation approach for network-wide BitTorrent distribution
abstract
BitTorrent(BT) has emerged as one of the most popular protocols for content sharing in recent years. Most BT applications are network-oblivious which brings great challenges to traffic engineering. As basic input information, network-wide BT traffic distribution is of vital importance for service providers (SP) or carriers to enforce appropriate control policies. Direct online measurement of network-wide BT traffic faces scalability issue, for it requires deploying application identification deep packet inspection (DPI) device on every subject link. In this paper, we propose a novel approach to estimate network-wide BT traffic distribution through combining sFlow and Tracker traffic analysis. Owing to a noticeable increase of BT traffic using random port numbers, sFlow alone is no longer sufficient for BT recognition. We tackle this problem by combining sFlow with information collected from BT tracker's traffic. With assistance of this information, BT samples can be accurately distinguished from other sFlow samples. The efficacy of our approach is evaluated through experiments on backbone links of campus network. Evaluation results show that our approach can reach a high accuracy in network-wide BT traffic estimation.
Wei Ye 0010, Kaida Jiang, Yaohui Jin
APNOMS6
2011 Routing for Deadline-Constrained Bulk Data Transfers Based on Transfer Failure Probability
abstract
Bulk data transfers, which require reliable and efficient transfer of terabits or even petabits of data for data-intensive applications, have been extensively studied. In these applications, usually, a specified amount of data needs to be transferred within strict deadline. Previous researches mainly focus on routing deadline-constrained bulk data transfers to improve the utilization of network resources for guarantying deadline requirements, without considering network failures. This paper proposes novel routing algorithm for deadline-constrained bulk data transfers based on Transfer Failure Probability using continuous-time Markov model. The goal of the novel algorithm is to guarantee the deadline requirement by decreasing the failure probability of data transfer request under network failure settings. Simulation evaluates the performance and demonstrates the effectiveness of our algorithm in terms of deadline violation ratio and service blocking ratio under network failures.
Yaoquan Zhong, Wei Guo 0003, Yaohui Jin, Weiqiang Sun, Weisheng Hu
ICC3
2010 Performances of Random IPTV Channel Change with Finite Duration Multi-Channel Delivery
abstract
Multi-channel delivery is one competitive solution to reduce channel change time in IPTV. We presented the performance analysis on finite duration multi-channel delivery method in our previous work, when channel changes are sequential and adjacent. In this paper, we develop mathematical models to evaluate the performance of this method when channel change happens randomly. This in fact generalizes the previous model. We find that, in typical setup, a delivery duration of 20 seconds is enough to reduce the channel change time by 47.5%, yet the peak bandwidth increase on carrier's uplink is 66.7%. This indicates that the finite duration multi-channel delivery method can improve viewers' experience effectively when channel changes are random.
Bing Xun, Weiqiang Sun, Yaohui Jin, Wei Guo 0003, Weisheng Hu, Kan Lin
ICC3
2009 On the Efficiency of Inter-Domain State Advertising in Multi-Domain Networks
abstract
In this paper, we study the efficiency of inter-domain state advertising in multi-domain networks, in which the intra-domain topology is substituted with a full-mesh topology of abstract links interconnecting the border nodes of the domain. The advertising trigger rate of individual abstract link is taken as the metric of efficiency. We present an analytical model that reveals the contributions of different physical links to the advertising trigger rate of individual abstract links. According to the model, we propose a state based advertising mechanism, called partial link based advertising (PLA) which can discern and advertise the state variation of individual abstract link by monitoring only a portion of physical links. We formulate the problem of how to optimally selecting physical links to monitor, and provide a heuristic algorithm for link selection. Results show that the PLA performs more efficiently than existing mechanisms in terms of advertising trigger rate and blocking probability.
Ying Di Yu, Yaohui Jin, Weiqiang Sun, Wei Guo 0003, Weisheng Hu
GLOBECOM2
2008 Fault-Tolerant Policy for Optical Network Based Distributed Computing System
abstract
The optical network based distributed computing system has been thought as a promising technology to support large-scale data-intensive distributed applications. For such a system with so many heterogeneous resources and middlewares involved, faults seem to be inevitable. However, for those applications that need to be finished before the given deadline, a fault in the system will lead to the failure of the application. Therefore, fault-tolerant policy is necessary to improve the performance of the system when faults could happen. In this paper, we address to the fault-tolerant problem for the optical network based distributed computing system. We first propose an overlay approach which applies the existing fault-tolerant policies for distributed computing and optical network. Then we present a joint fault-tolerant policy which takes into account the fault tolerance for computing resource and network resource in the same time. We compare the performances of different polices by simulation. The simulation results show that the joint fault-tolerant policy achieves much better performances compared to overlay approaches.
Wei Guo 0003, Yaohui Jin, Weiqiang Sun, Weisheng Hu
CCGRID3
2008 Nonblocking Multicast-Capable Optical Cross Connects Based on the 4-Stage Multicast Network
abstract
In this paper, we investigated designing multicast-capable optical cross-connects (MC-OXCs) on a basis of the 4- stage multicast network. Firstly, we derive the sufficient wide- sense nonblocking (WSNB) and rearrangeable nonblocking (RNB) conditions for the 4-stage multicast network with only two (the second and output) stages being multicast-capable. Both WSNB and RNB 4-stage multicast networks need 0(N3/2) crosspoints. Then MC-OXCs empoying the WSNB and RNB 4-stage multicast network are proposed and proven to be power efficient in reducing the total power loss caused by light splitting.
Fangfang Yan, Weisheng Hu, Weiqiang Sun, Wei Guo 0003, Yaohui Jin
GLOBECOM5
2008 Per-Flow Re-Sequencing in Load-Balanced Switches by Using Dynamic Mailbox Sharing
abstract
Load-balanced switches have received much attention because they are more scalable than other switch architectures. However, a load-balanced switch has the problem of packet mis-sequencing. In this paper, we propose a dynamic mailbox sharing (DMS) scheme to eliminate the mis-sequencing problem of load-balanced switches only at the cost of a very small increase of delay. The key idea is to keep packets of the same flow in order in the load-balanced switch. The DMS scheme is based on two statistical facts in operational networks: the number of simultaneous active flows in the router buffer is far less than that of in-progress flows, and most of the intra-flow packet intervals are longer than the packet delay in the high speed router. In DMS, the packet sequence of the same flow arrived in the input ports is recorded in the mailbox maintained in the output ports. Then, packets of the same flow are delivered according to the order of their arrivals. The mailbox becomes the bottleneck in order to accommodate a large number of flows. We thus propose a dynamic sharing scheme to alleviate the bottleneck and greatly enhance the scalability of the mailbox. By simulations using the real internet traffic traces, we show that with a simple flow splitter mechanism restraining mis-sequencing, the average packet delay using DMS is considerably lower than that of other schemes including uniform frame spreading, padded frame and the CR switch, and it is close to the ideal case without re-sequencing even when the load is very high. The results also demonstrate that the size of mailbox is in the hundreds.
Yaohui Jin, Ying Di Yu, Weisheng Hu, Nirwan Ansari
ICC2
2007 Multicast Flow Aggregation in IP over Optical Networks
abstract
It is widely believed that IP over optical networks will be a major component of the next generation Internet However, it is not efficient to map a single multicast IP flow into one light-tree, since the bandwidth of an IP flow required is usually much less than that of a light-tree. In this paper, we study the problem of multicast flow aggregation (MFA) in the IP over optical two-layered networks under the overlay model, which can be defined as follows: given a set of head ends (i.e. optical multicasting sources), each of which can provide a set of contents (i.e. multicast IP flows) with different required transmission bandwidth, and a set of requested content at the access routers (i.e. optical multicasting destinations), find a set of light-trees as well as the optimal aggregation of multicast IP flows in each light-tree. We model MFA by a tri-partite graph with multiple criteria and show that the problem is NP-complete. Optimal solutions are designed by exploiting MFA to formulate an integer linear programming (ILP), with two parameters: the multicast receiving index alpha and the redundant transmitting index beta. We also propose a heuristic algorithm. Finally, we compare the performance of MFA for different combination of alpha and beta via experiments and show our heuristic algorithm is effective for large-scale network in numerical results
Yaohui Jin, Weiqiang Sun, Wei Guo 0003, Weisheng Hu, Wen-De Zhong, Min-You Wu
IEEE J. Sel. Areas Commun.2
2005 On topology-independent IP group aggregation in multicast capable optical networks
abstract
To support large-scale IP-TV signal delivering, a new multicast-IP over light-tree network model was proposed in W. Sun et al. In this paper, we study the problem of topology-independent IP group aggregation in the multicast capable optical network. We first use a tri-partite graph to describe the problem, and then formulate the problem by a mixed integer linear programming (MILP) model. Finally, we propose a pre-estimation for the number of light-trees to reduce the searching space, which significantly saves solving time of MILP
Yaohui Jin, Weiqiang Sun, Wei Guo 0003, Weisheng Hu
GLOBECOM2
2004 On-line integrated routing in dynamic multifiber IP/WDM networks
abstract
This paper focuses on dynamic integrated routing in multifiber Internet protocol/wavelength-division multiplexing (IP/WDM) networks, which can be implemented through either one-step routing (OSR) or two-step routing (TSR) approach. Based on an extended layered-graph, two resource assignment strategies, termed channel-level balance (CLB) and link-level balance (LLB), are proposed to balance the traffic in the network at different levels. To further improve the performance, a parameter K is introduced to make a dynamic tradeoff between the logical-layer links and the optical-layer links. Simulation studies are carried out for various topologies. The results show that LLB is better than CLB in most cases, and LLB combined with OSR has the optimal performance. Also, we find that the routing approach and the resource assignment strategy individually play different roles with different values of r/sub l/ that is introduced to indicate the resource richness of the network. As a multifiber network is functionally equivalent to a single-fiber network with limited wavelength conversion, we investigate the effects of wavelength conversion by studying the multifiber IP/WDM networks. The analysis shows that, when the granularity of each connection request is much smaller than the wavelength granularity, wavelength conversion may increase the request blocking probability in the network.
Tong Ye 0002, Qingji Zeng, Yikai Su, Lufeng Leng, Wei Wei 0010, Wei Guo 0003, Yaohui Jin
IEEE J. Sel. Areas Commun.8