EDBT 2026 Demo / reviewers in the wild / expert
Jingjing Jiang
dblp:41/8612
· DBLP profile ↗
36ranked-venue papers
15as first author
31since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 7 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 first-author · 8 since 2021Systems, architecture and hardware · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Improving modality alignment via multi-scale and selective feature learning in reasoning segmentation
Zhongye Liu, Weiliang Zuo, Liguo Liu, Jingjing Jiang |
Knowl. Based Syst. | 4 |
| 2026 | MonoA2: Adaptive depth with augmented head for monocular 3D object detection
Jinpeng Dong, Sanping Zhou, Jingjing Jiang, Weiliang Zuo, Shi-tao Chen, Nanning Zheng 0001 |
Pattern Recognit. | 5 |
| 2026 | $k$-Step Look-Ahead Active Concurrent Learning-Based Dual Control of Exploration and Exploitation for Auto-OptimizationabstractThis study introduces a $k$ -step look-ahead active concurrent learning-based dual control of exploration and exploitation (KSLCL-DCEE) framework designed to address the challenges of auto-optimization in systems with unknown references and environments, inherently balancing parameter estimation and optimal reference tracking. The KSLCL-DCEE algorithm incorporates two loops that employ future gradients of the cost function to generate the subsequent control command by looking ahead $k$ -steps: the inner loop generates $k$ -step look-ahead gradients (i.e., estimated reference trajectory), while the outer loop utilizes the gradient at the $k$ th step to generate the dual control commands which act on a general linear system. Active concurrent learning with a modified learning rate in the initial period is introduced to relax the reliance on the condition of persistent excitation and achieve faster convergence. A comprehensive stability analysis of KSLCL-DCEE is provided. The effectiveness and performance of KSLCL-DCEE are demonstrated through numerical studies and applications on photovoltaic (PV) arrays. Yalei Yu, Jingjing Jiang, Wen-Hua Chen 0001, Yuefei Zuo |
IEEE Trans. Cybern. | 2 |
| 2025 | Sequence Knowledge Enhancement Distillation Framework for Ultra-Fast Image DerainingabstractTraditional knowledge distillation techniques are aimed at compressing models and speeding up inference, but they often fail to maintain the superior capabilities of complex models in simpler ones. To address this issue, this paper focus on the deraining task and introduces the Sequential Knowledge-Enhanced Distillation Framework (SKEDF). SKEDF, as a two-stage strategy, comprises a Knowledge Completion Stage (KCS) and a Knowledge Enhancement Stage (KES). The KCS employs feature sequence to enhance the student network’s ability to understand and learn superior deraining capabilities from various teacher networks. The KES independently trains the student network to further refine its deraining abilities. Moreover, the framework incorporates a Contrastive Structural Similarity Regularization (CSSR) loss, ingeniously integrating contrastive learning with the Structural Similarity Index (SSIM) to enhance training efficacy and achieve nuanced model improvements. Quantitative and qualitative results demonstrate that SKEDF not only achieves a breakthrough improvement in model efficiency but also delivers more promising deraining performance compared to other SOTA solutions. Pengyu Fu, Jingjing Jiang, Yuanjian Zhang 0001 |
ICASSP | 5 |
| 2025 | Corvid: Improving Multimodal Large Language Models Towards Chain-of-Thought ReasoningabstractRecent advancements in multimodal large language models (MLLMs) have demonstrated exceptional performance in multimodal perception and understanding. However, leading open-source MLLMs exhibit significant limitations in complex and structured reasoning, particularly in tasks requiring deep reasoning for decision-making and problem-solving. In this work, we present Corvid, an MLLM with enhanced chain-of-thought (CoT) reasoning capabilities. Architecturally, Corvid incorporates a hybrid vision encoder for informative visual representation and a meticulously designed connector (GateMixer) to facilitate cross-modal alignment. To enhance Corvid's CoT reasoning capabilities, we introduce MCoT-Instruct-287K, a high-quality multimodal CoT instruction-following dataset, refined and standardized from diverse public reasoning sources. Leveraging this dataset, we fine-tune Corvid with a two-stage CoT-formatted training approach to progressively enhance its step-by-step reasoning abilities. Furthermore, we propose an effective inference-time scaling strategy that enables Corvid to mitigate over-reasoning and under-reasoning through self-verification. Extensive experiments demonstrate that Corvid outperforms existing o1-like MLLMs and state-of-the-art MLLMs with similar parameter scales, with notable strengths in mathematical reasoning and science problem-solving. Project page: https://mm-vl.github.io/corvid. Jingjing Jiang, Xurui Song, Hanwang Zhang, Jun Luo 0001 |
ICCV | 1 |
| 2025 | Towards a Japanese Full-duplex Spoken Dialogue System
Atsumoto Ohashi, Shinya Iizuka, Jingjing Jiang, Ryuichiro Higashinaka |
INTERSPEECH | 3 |
| 2025 | Co-Reinforcement Learning for Unified Multimodal Understanding and GenerationabstractThis paper presents a pioneering exploration of reinforcement learning (RL) via group relative policy optimization for unified multimodal large language models (ULMs), aimed at simultaneously reinforcing generation and understanding capabilities. Through systematic pilot studies, we uncover the significant potential of ULMs to enable the synergistic co-evolution of dual capabilities within a shared policy optimization framework. Building on this insight, we introduce \textbf{CoRL}, a \textbf{Co}-\textbf{R}einforcement \textbf{L}earning framework comprising a unified RL stage for joint optimization and a refined RL stage for task-specific enhancement. With the proposed CoRL, our resulting model, \textbf{ULM-R1}, achieves average improvements of 7\% on three text-to-image generation datasets and 23\% on nine multimodal understanding benchmarks. These results demonstrate the effectiveness of CoRL and highlight the substantial benefits of reinforcement learning in facilitating cross-task synergy and optimization for ULMs. Code is available at \url{https://github.com/mm-vl/ULM-R1}. Jingjing Jiang, Chongjie Si, Jun Luo 0001, Hanwang Zhang |
NeurIPS | 1 |
| 2025 | Integrating Physiological, Speech, and Textual Information Toward Real-Time Recognition of Emotional Valence in DialogueabstractAccurately estimating users’ emotional states in real time is crucial for enabling dialogue systems to respond adaptively. While existing approaches primarily rely on verbal information, such as text and speech, these modalities are often unavailable in non-speaking situations. In such cases, non-verbal information, particularly physiological signals, becomes essential for understanding users’ emotional states. In this study, we aimed to develop a model for real-time recognition of users’ binary emotional valence (high-valence vs. low-valence) during conversations. Specifically, we utilized an existing Japanese multimodal dialogue dataset, which includes various physiological signals, namely electrodermal activity (EDA), blood volume pulse (BVP), photoplethysmography (PPG), and pupil diameter, along with speech and textual data. We classify the emotional valence of every 15-second segment of dialogue interaction by integrating such multimodal inputs. To this end, time-series embeddings of physiological signals are extracted using a self-supervised encoder, while speech and textual features are obtained from pre-trained Japanese HuBERT and BERT models, respectively. The modality-specific embeddings are integrated using a feature fusion mechanism for emotional valence recognition. Experimental results show that while each modality individually contributes to emotion recognition, the inclusion of physiological signals leads to a notable performance improvement, particularly in non-speaking or minimally verbal situations. These findings underscore the importance of physiological information for enhancing real-time valence recognition in dialogue systems, especially when verbal information is limited. Jingjing Jiang, Ryuichiro Higashinaka |
SIGDIAL | 1 |
| 2025 | Detecting question relatedness in programming Q&A communities via bimodal feature fusion
Qirong Bu, Xiangqiang Guo, Jingjing Jiang, Xiaodi Zhao, Wang Zou, Xuxin Wang, Jianqiang Yan |
Autom. Softw. Eng. | 4 |
| 2025 | MIFS: A low overhead and efficient mixture file index management method in flash file system
Jingjing Jiang, Mengfei Yang, Lei Qiao 0002, Tingyu Wang 0003 |
J. Syst. Archit. | 1 |
| 2025 | COSDA: Covariance regularized semantic data augmentation for self-supervised visual representation learning
Hui Chen 0036, Jingjing Jiang, Nanning Zheng 0001 |
Knowl. Based Syst. | 3 |
| 2025 | High-Level Decision Making in a Hierarchical Control Framework: Integrating HMDP and MPC for Autonomous SystemsabstractThis article addresses challenges of autonomous decisions making influenced by discrete system states, underlying continuous dynamics, and evolving operational environments. A comprehensive framework is proposed, encompassing new modeling, problem formulation, control design, and stability analysis. The framework integrates continuous system dynamics, used for low-level control, with discrete Markov decision processes (MDP) for high-level decision making. To capture the interactions between these domains, the decision-making system is modeled as a hybrid system consisting of a controlled MDP and autonomous (uncontrolled) continuous dynamics, collectively referred to as the hybrid Markov decision process (HMDP). The design focuses on ensuring safety and optimality by accounting for both discrete and continuous state variables across different levels. With the help of the model predictive control (MPC) concept, a decision-making scheme is developed for the hybrid model, with guarantees for recursive feasibility and stability. The proposed framework is applied to the autonomous lane changing system for intelligent vehicles, and simulation shows its capability to handle diverse behaviors in dynamic and complex environments. Xuefang Wang 0001, Jingjing Jiang, Wen-Hua Chen 0001 |
IEEE Trans. Cybern. | 2 |
| 2025 | Ultra-Fast Deraining Plugin for Vision-Based Perception of Autonomous DrivingabstractRain deviates the distribution of rainy images and the clean, rain-free data typically used during perception model training, this kind of out-of-distribution (OOD) issue making it difficult for models to generalize effectively in rainy scenarios, leading the performance degrade of autonomous perception systems in visual tasks such as lane detection and depth estimation, posing serious safety risks. To address this issue, we propose the Ultra-Fast Deraining Plugin (UFDP), a model-efficient deraining solution specifically designed to realign the distribution of rainy images and their rain-free counterparts. UFDP not only effectively removes rain from images but also seamlessly integrates into existing visual perception models, significantly enhancing their robustness and stability under rainy conditions. Through a detailed analysis of single-image color histograms and dataset-level distribution, we demonstrate how UFDP improves the similarity between rainy and non-rainy image distributions. Additionally, qualitative and quantitative results highlight UFDP’s superiority over state-of-the-art (SOTA) methods, showing a 5.4% improvement in SSIM and 8.1% in PSNR. UFDP also excels in terms of efficiency, achieving 7 times higher FPS than the slowest method, reducing FLOPs by 53.7 times, and using 28.8 times fewer MACs, with 6.2 times fewer parameters. This makes UFDP an ideal solution for ensuring reliable performance in autonomous driving visual perception systems, particularly in challenging rainy environments. Pengyu Fu, Jun Yang 0011, Jingjing Jiang, Yuanjian Zhang 0001 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2024 | Estimating the Emotional Valence of Interlocutors Using Heterogeneous Sensors in Human-Human DialogueabstractDialogue systems need to accurately understand the user's mental state to generate appropriate responses, but accurately discerning such states solely from text or speech can be challenging.To determine which information is necessary, we first collected human-human multimodal dialogues using heterogeneous sensors, resulting in a dataset containing various types of information including speech, video, physiological signals, gaze, and body movement.Additionally, for each time step of the data, users provided subjective evaluations of their emotional valence while reviewing the dialogue videos.Using this dataset and focusing on physiological signals, we analyzed the relationship between the signals and the subjective evaluations through Granger causality analysis.We also investigated how sensor signals differ depending on the polarity of the valence.Our findings revealed several physiological signals related to the user's emotional valence. Jingjing Jiang, Ryuichiro Higashinaka |
SIGDIAL | 1 |
| 2024 | An Adaptive Real-Time Garbage Collection Method Based on File Write Prediction
Jingjing Jiang, Mengfei Yang, Lei Qiao 0002, Tingyu Wang 0003, Shenghui Zhu |
TASE | 1 |
| 2024 | Correlation Information Bottleneck: Towards Adapting Pretrained Multimodal Models for Robust Visual Question Answering
Jingjing Jiang, Ziyi Liu 0001, Nanning Zheng 0001 |
Int. J. Comput. Vis. | 1 |
| 2024 | Segmentation from localization: a weakly supervised semantic segmentation method for resegmenting CAM
Jingjing Jiang, Jiali Wu |
Multim. Tools Appl. | 1 |
| 2024 | FFINet: Future Feedback Interaction Network for Motion ForecastingabstractMotion forecasting plays a crucial role in autonomous driving, with the aim of predicting the future reasonable motions of traffic agents. Most existing methods mainly model the historical interactions between agents and the environment, and predict multi-modal trajectories in a feedforward process, ignoring potential trajectory changes caused by future interactions between agents. In this paper, we propose a novel Future Feedback Interaction Network (FFINet) to aggregate the current, observations and potential future interaction features for trajectory prediction. Firstly, we employ different spatial-temporal encoders to embed the decomposed position vectors and the current position of each scene, providing rich features for the subsequent cross-temporal aggregation. Secondly, the relative interaction and cross-temporal aggregation strategies are sequentially adopted to integrate features in the current fusion module, observation interaction module, future feedback module and global fusion module, in which the future feedback module can enable the understanding of pre-action by feeding the influence of preview information to feedforward prediction. Thirdly, the comprehensive interaction features are further fed into final predictor to generate the joint predicted trajectories of multiple agents. Extensive experimental results show that our FFINet achieves the state-of-the-art performance on Argoverse 1 and Argoverse 2 motion forecasting benchmarks. Miao Kang, Shengqi Wang, Sanping Zhou, Ke Ye, Jingjing Jiang, Nanning Zheng 0001 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2023 | MixPHM: Redundancy-Aware Parameter-Efficient Tuning for Low-Resource Visual Question AnsweringabstractRecently, finetuning pretrained vision-language models (VLMs) has been a prevailing paradigm for achieving state-of-the-art performance in VQA. However, as VLMs scale, it becomes computationally expensive, storage inefficient, and prone to overfitting when tuning full model parameters for a specific task in low-resource settings. Although current parameter-efficient tuning methods dramatically reduce the number of tunable parameters, there still exists a significant performance gap with full finetuning. In this paper, we propose MixPHM, a redundancy-aware parameter-efficient tuning method that outperforms full finetuning in low-resource VQA. Specifically, MixPHM is a lightweight module implemented by multiple PHM-experts in a mixture-of-experts manner. To reduce parameter redundancy, we reparameterize expert weights in a low-rank subspace and share part of the weights inside and across MixPHM. Moreover, based on our quantitative analysis of representation redundancy, we propose Redundancy Regularization, which facilitates MixPHM to reduce task-irrelevant redundancy while promoting task-relevant correlation. Experiments conducted on VQA v2, GQA, and OK-VQA with different low-resource settings show that our MixPHM outperforms state-of-the-art parameter-efficient methods and is the only one consistently surpassing full finetuning. Jingjing Jiang, Nanning Zheng 0001 |
CVPR | 1 |
| 2023 | A congestion-aware path planning method considering crowd spatial-temporal anomalies for long-term autonomy of mobile robotsabstractA congestion-aware path planning method is pre-sented for mobile robots during long-term deployment in human occupied environments. With known spatial-temporal crowd patterns, the robot will navigate to its destination via less congested areas. Traditional traffic-aware routing methods do not consider spatial-temporal anomalies of macroscopic crowd behaviour that can deviate from the predicted crowd spatial distribution. The proposed method improves long-term path planning adaptivity by integrating a partially updated memory (PUM) model that utilizes observed anomalies to generate a multi-layer crowd density map to improve estimation accuracy. Using this map, we are able to generate a path that has less chance to encounter the crowded areas. Simulation results show that our method outperforms the benchmark congestion-aware routing method in terms of reducing the probability of robot's proximity to dense crowds. Zijian Ge, Jingjing Jiang, Matthew Coombes |
ICRA | 2 |
| 2023 | Learning to Infer Unseen Single-/ Multi-Attribute-Object Compositions With Graph NetworksabstractInferring the unseen attribute-object composition is critical to make machines learn to decompose and compose complex concepts like people. Most existing methods are limited to the composition recognition of single-attribute-object, and can hardly learn relations between the attributes and objects. In this paper, we propose an attribute-object semantic association graph model to learn the complex relations and enable knowledge transfer between primitives. With nodes representing attributes and objects, the graph can be constructed flexibly, which realizes both single- and multi-attribute-object composition recognition. In order to reduce mis-classifications of similar compositions (e.g., scratched screen and broken screen), driven by the contrastive loss, the anchor image feature is pulled closer to the corresponding label feature and pushed away from other negative label features. Specifically, a novel balance loss is proposed to alleviate the domain bias, where a model prefers to predict seen compositions. In addition, we build a large-scale Multi-Attribute Dataset (MAD) with 116,099 images and 8,030 label categories for inferring unseen multi-attribute-object compositions. Along with MAD, we propose two novel metrics Hard and Soft to give a comprehensive evaluation in the multi-attribute setting. Experiments on MAD and two other single-attribute-object benchmarks (MIT-States and UT-Zappos50K) demonstrate the effectiveness of our approach. Hui Chen 0036, Jingjing Jiang, Nanning Zheng 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | LiVLR: A Lightweight Visual-Linguistic Reasoning Framework for Video Question AnsweringabstractVideo Question Answering (VideoQA), aiming to correctly answer a given question based on understanding multimodal video content, is challenging due to the richness of the video content. From the perspective of video understanding, a complete VideoQA framework needs to understand the video content at different semantic levels and flexibly integrate diverse video content to distill question-related content. To this end, we propose a Lightweight Visual-Linguistic Reasoning framework named$\text{LiVLR}$. Specifically,$\text{LiVLR}$first utilizes graph-based visual and linguistic encoders to obtain multi-grained visual and linguistic representations, respectively. Subsequently, the obtained representations are integrated with the devised Diversity-aware Visual-Linguistic Reasoning module ($\text{DaVL}$).$\text{DaVL}$distinguishes different types of representations with the learnable index embedding in graph embedding. Therefore,$\text{DaVL}$can flexibly adjust the importance of different representations when generating the question-related joint representation. The proposed$\text{LiVLR}$is lightweight and shows its performance advantage on three VideoQA benchmarks, MRSVTT-QA, KnowIT VQA, and TVQA. Extensive ablation studies demonstrate the effectiveness of the key components of$\text{LiVLR}$. Jingjing Jiang, Ziyi Liu 0001, Nanning Zheng 0001 |
IEEE Trans. Multim. | 1 |
| 2023 | Reinforcement Learning-Based Fixed-Time Trajectory Tracking Control for Uncertain Robotic Manipulators With Input SaturationabstractA fixed-time trajectory tracking control method for uncertain robotic manipulators with input saturation based on reinforcement learning (RL) is studied. The designed RL control algorithm is implemented by a radial basis function (RBF) neural network (NN), in which the actor NN is used to generate the control strategy and the critic NN is used to evaluate the execution cost. A new nonsingular fast terminal sliding mode technique is used to ensure the convergence of tracking error in fixed time, and the upper bound of convergence time is estimated. To solve the saturation problem of an actuator, a nonlinear antiwindup compensator is designed to compensate for the saturation effect of the joint torque actuator in real time. Finally, the stability of the closed-loop system based on the Lyapunov candidate is analyzed, and the timing convergence of the closed-loop system is proven. Simulation and experimental results show the effectiveness and superiority of the proposed control law. Shengjie Cao, Liang Sun 0004, Jingjing Jiang, Zongyu Zuo |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | Robust Hierarchical Adaptive Fuzzy Relative Motion Coordination for Feature Points of Two Rigid Bodies With Input and Output ConstraintsabstractIn this article, a model-based six-degrees-of-freedom relative motion coordinated control approach is developed for the relative position tracking and attitude synchronization between feature points of two rigid bodies subject to control input constraints, output constraints, and model uncertainties. In the designing framework of adaptive backstepping control technique, the control input saturation is compensated by the nonlinear antiwindup compensator and the output constraints are handled by the barrier Lyapunov function-based backstepping design. The unknown misalignment vector of the feature point with respect to the center of the mass for the chaser is estimated by the element-wise adaptive law, while the model uncertainties and unknown dynamical couplings are compensated by the adaptive hierarchical fuzzy logic system to decrease the computational burden with respect to the traditional adaptive fuzzy system. The ultimately uniformly bounded convergence of the relative pose and relative velocities is analyzed in the Lyapunov framework and the effectiveness of the proposed approach is validated by the numerical simulations. Liang Sun 0004, Jingjing Jiang |
IEEE Trans. Fuzzy Syst. | 2 |
| 2022 | Robust Adaptive Learning-Based Path Tracking Control of Autonomous Vehicles Under Uncertain Driving EnvironmentsabstractThis paper investigates the path tracking control problem of autonomous vehicles subject to modelling uncertainties and external disturbances. The problem is approached by employing a 2-degree of freedom vehicle model, which is reformulated into a newly defined parametric form with the system uncertainties being lumped into an unknown parametric vector. On top of the parametric system representation, a novel robust adaptive learning control (RALC) approach is then developed, which estimates the system uncertainties through iterative learning while treating the external disturbances by adopting a robust term. It is shown that the proposed approach is able to improve the lateral tracking performance gradually through learning from previous control experiences, despite only partial knowledge of the vehicle dynamics being available. It is noteworthy that a novel technique targeting at the non-square input distribution matrix is employed so as to deal with the under-actuation property of the vehicle dynamics, which extends the adaptive learning control theory from square systems to non-square systems. Moreover, the convergence properties of the RALC algorithm are analysed under the framework of Lyapunov-like theory by virtue of the composite energy function and the$\lambda $-norm. The effectiveness of the proposed control scheme is verified by representative simulation examples and comparisons with existing methods. Boli Chen, Jingjing Jiang |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2021 | Personalized Recommendation Algorithm Based on Fuzzy Semantics in Big Data EnvironmentabstractWith the continuous expansion of the scale of e-government, personalized recommendation technology has been widely used. However, the traditional recommendation system has been unable to meet the needs of the current data processing, and good big data processing ability has become the basic requirement of the new personalized recommendation system, Combined with the application background of big data: firstly, this paper puts forward the construction method of personalized information service system with portal as the core, including the construction object, mode, goal and principle system. After information integration, in order to improve the security of the system, the authority management and access control are introduced. Secondly, a unified model is established to describe the integration of information resources and user files. Thirdly, the data mining method is used to discover the user's interest in integrated information. Combined with Hadoop cloud computing platform, the distributed construction of the model is realized, and the stability of the classic collaborative filtering algorithm is improved. Finally, a fuzzy semantic model of personalized recommendation system is proposed and implemented with a fuzzy description logic language. Jingjing Jiang, Mengxuan Wu |
IWCMC | 1 |
| 2021 | Research on the Optimization Algorithm of Big Data Computing SystemabstractWith the social progress and development, the scale of data continues to expand, in order to realize the processing and analysis of large-scale data, graph computing system came into being. At present, with the continuous maturity of graph computing system, graph computing has been widely used in various fields, such as social field, Internet of things field and neural network field. In recent years, different graph computing models have emerged, and some typical distributed graph computing models show good expansibility in the formulation of graph data for big data processing. However, in order to further expand the expansibility, many graph calculation models are studied by algorithms. At present, the SFA algorithm is mostly used in the graph calculation system. However, with the continuous development of graph calculation, many inadaptability of the SFA algorithm appear which restricts the further development of graph calculation. Therefore, it is an urgent problem to optimize the algorithm of graph computing system. On the basis of scholars' research, this paper firstly gives a simple overview of graph calculation and graph calculation model. On this basis, it analyzes the specific formula and significance of SFA algorithm, puts forward the specific scheme of algorithm optimization, and carries out experimental detection of optimization algorithm. Mengxuan Wu, Jingjing Jiang |
IWCMC | 2 |
| 2021 | X-GGM: Graph Generative Modeling for Out-of-distribution Generalization in Visual Question AnsweringabstractEncouraging progress has been made towards Visual Question Answering (VQA) in recent years, but it is still challenging to enable VQA models to adaptively generalize to out-of-distribution (OOD) samples. Intuitively, recompositions of existing visual concepts (i.e., attributes and objects) can generate unseen compositions in the training set, which will promote VQA models to generalize to OOD samples. In this paper, we formulate OOD generalization in VQA as a compositional generalization problem and propose a graph generative modeling-based training scheme (X-GGM) to handle the problem implicitly. X-GGM leverages graph generative modeling to iteratively generate a relation matrix and node representations for the predefined graph that utilizes attribute-object pairs as nodes. Furthermore, to alleviate the unstable training issue in graph generative modeling, we propose a gradient distribution consistency loss to constrain the data distribution with adversarial perturbations and the generated distribution. The baseline VQA model (LXMERT) trained with the X-GGM scheme achieves state-of-the-art OOD performance on two standard VQA OOD benchmarks, i.e., VQA-CP v2 and GQA-OOD. Extensive ablation studies demonstrate the effectiveness of X-GGM components. Jingjing Jiang, Ziyi Liu 0001, Yifan Liu 0014, Zhixiong Nan, Nanning Zheng 0001 |
ACM Multimedia | 1 |
| 2021 | Predicting short-term next-active-object through visual attention and hand position
Jingjing Jiang, Zhixiong Nan, Hui Chen 0036, Shi-tao Chen, Nanning Zheng 0001 |
Neurocomputing | 1 |
| 2021 | A joint object detection and semantic segmentation model with cross-attention and inner-attention mechanisms
Zhixiong Nan, Jizhi Peng, Jingjing Jiang, Hui Chen 0036, Ben Yang, Jingmin Xin, Nanning Zheng 0001 |
Neurocomputing | 3 |
| 2021 | Predicting Task-Driven Attention via Integrating Bottom-Up Stimulus and Top-Down GuidanceabstractTask-free attention has gained intensive interest in the computer vision community while relatively few works focus on task-driven attention (TDAttention). Thus this paper handles the problem of TDAttention prediction in daily scenarios where a human is doing a task. Motivated by the cognition mechanism that human attention allocation is jointly controlled by the top-down guidance and bottom-up stimulus, this paper proposes a cognitively-explanatory deep neural network model to predict TDAttention. Given an image sequence, bottom-up features, such as human pose and motion, are firstly extracted. At the same time, the coarse-grained task information and fine-grained task information are embedded as a top-down feature. The bottom-up features are then fused with the top-down feature to guide the model to predict TDAttention. Two public datasets are re-annotated to make them qualified for TDAttention prediction, and our model is widely compared with other models on the two datasets. In addition, some ablation studies are conducted to evaluate the individual modules in our model. Experiment results demonstrate the effectiveness of our model. Zhixiong Nan, Jingjing Jiang, Xiaofeng Gao 0002, Sanping Zhou, Weiliang Zuo, Ping Wei 0001, Nanning Zheng 0001 |
IEEE Trans. Image Process. | 2 |
| 2020 | Automatic Lane Change Maneuver in Dynamic Environment Using Model Predictive Control MethodabstractThe lane change maneuver is one of the typical maneuvers in various driving situations. Therefore the automatic lane change function is one of the key functions for autonomous vehicles. Many researches have been conducted in this field. Most existing work focused on the solutions for the static environment and assume that the surrounding vehicles are running at constant speeds. However, in reality, if not all the vehicles on the road are fully autonomous, the situation could be much more complicated and the ego vehicle has to deal with the dynamic environment. This paper proposes a Model Predictive Control (MPC)-based method to achieve automatic lane change in a dynamic environment. A two-wheel dynamic bicycle model, which combines the longitudinal and lateral motion of the ego vehicle, together with a utility function, which helps to automatically determine the target lane have been used in the algorithm. The simulation results have demonstrated the capability of the proposed algorithm in a dynamic environment. Zhaolun Li, Jingjing Jiang, Wen-Hua Chen 0001 |
IROS | 2 |
| 2020 | Logistics industry monitoring system based on wireless sensor network platform
Jingjing Jiang, Haiwen Wang, Xiangwei Mu, Sheng Guan |
Comput. Commun. | 1 |
| 2018 | An effective pattern-based Bayesian classifier for evolving data stream
Jidong Yuan, Yange Sun, Wei Zhang 0180, Jingjing Jiang |
Neurocomputing | 5 |
| 2017 | A Pattern-Based Bayesian Classifier for Data Stream
Jidong Yuan, Yange Sun, Wei Zhang 0180, Jingjing Jiang |
ICONIP (4) | 5 |
| 2017 | Shared-Control for a Rear-Wheel Drive Car: Dynamic Environments and Disturbance RejectionabstractThis paper studies the shared-control problem for the kinematic model of a group of rear-wheel drive cars in a (possibly) dynamic (i.e., time-varying) environment. The design of the shared-controller is based on measurements of distances to obstacles, angle differences, and the human input. The shared-controller is used to guarantee the safety of the car when the driver behaves “dangerously.” Formal properties of the closed-loop system with the shared-controller are presented through a Lyapunov-like analysis. In addition, we consider uncertainties in the dynamics and prove that the shared-controller is able to help the driver drive the car safely even in the presence of disturbances. Finally, the effectiveness of the controller is verified by two case studies: traffic at a junction and at a roundabout. Jingjing Jiang, Alessandro Astolfi |
IEEE Trans. Hum. Mach. Syst. | 1 |