EDBT 2026 Demo / reviewers in the wild / expert
Jing Zhao 0015
dblp:69/5882-15
· DBLP profile ↗
52ranked-venue papers
10as first author
35since 2021 · last 2026
0000-0003-0158-5330ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 35 · 7 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 1 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Temporal and Spatial Representation Learning for Multimodal Low-Beam 3D Object DetectionabstractTo facilitate the large-scale deployment of autonomous driving in real-world scenarios, developing low-cost and high-performance 3D object detection systems has become a critical technical challenge. Although high-beam LiDARs provide denser point cloud data, their prohibitive hardware cost and high power consumption limit their practicality. In contrast, low-beam LiDARs offer advantages in terms of affordability and energy efficiency, but often suffer from inadequate perception accuracy due to their sparser point cloud data. This paper focuses on the task of multimodal 3D object detection with low-beam LiDARs, and proposes a novel approach that integrates temporal and spatial representation learning to enhance detection accuracy under sparser sensor conditions. Specifically, our approach comprises: (1) a Temporal Feature Prediction Learning (TFPL) module, which predicts the current BEV representation based on a sequence of historical BEV features; (2) a Spatial Feature Observation Learning (SFOL) module, which aligns BEV (Bird's-Eye-View) features from high-beam and low-beam LiDAR to enforce the low-beam features to approximate high-beam representations; (3) an Uncertainty-Aware Fusion (UAF) strategy, which performs feature-wise weighting between the predicted and observed BEV features by leveraging channel-wise variances, effectively mitigating perturbations in the learned BEV representations. Extensive experiments on the KITTI and nuScenes 3D object detection datasets demonstrate that the proposed approach significantly improves detection performance under low-beam LiDAR configurations. Lin Wang 0053, Shiliang Sun, Jing Zhao 0015 |
AAAI | 3 |
| 2026 | Hate Speech Detection in Blockchain Transaction Metadata via Structured Reasoning and Uncertainty-Aware Retrieval Augmentation
Jing Zhao 0015 |
ICIC | 2 |
| 2026 | RoLiC: A Robust LiDAR-Camera Fusion Framework for 3D Object DetectionabstractIn the 3D object detection task of autonomous driving systems, LiDAR and camera are the most crucial sensors, and current methods primarily focus on fusion strategies for these two complementary modalities. However, in real-world driving scenarios, potential sensor failures pose critical risks, potentially undermining the effectiveness of fusion-based detection approaches. This work presents RoLiC, a robust LiDAR-camera fusion framework designed to handle three challenging deployment scenarios within a unified model: LiDAR failure, camera failure, and simultaneous LiDAR-camera failure. To mitigate cross-modal dependency and recover missing information, RoLiC introduces two cross-modality feature transformers (L2C and C2L) that bidirectionally complete features between modalities when partial data is available. To further enhance feature reliability, we design a Sparse Similarity Loss (SSL) that constrains feature learning within high-probability object regions. Moreover, RoLiC integrates a task-aware two-stage feature knowledge distillation strategy (MS1 and MS2), where MS1 captures cross-modality complementarities, and MS2 distills knowledge between the fused modality and complete modality settings. Extensive experiments on the nuScenes and KITTI benchmarks demonstrate that RoLiC consistently outperforms state-of-the-art methods across all sensor-failure conditions. Code is available at https://github.com/JLIN77/RoLiC. Lin Wang 0053, Shiliang Sun, Jing Zhao 0015 |
IEEE Trans. Image Process. | 3 |
| 2026 | Visual Context and Commonsense-Guided Causal Chain-of-Thoughts for Visual Commonsense ReasoningabstractHumans are capable of inferring dynamic context from a still image and, with the provision of additional commonsense knowledge, can accurately complete visual commonsense reasoning tasks. Nevertheless, this remains a highly challenging cognitive-level task for current vision-language models. Previous work has primarily focused on utilizing models fine-tuned for specific downstream tasks and introduces external world knowledge to tackle these challenging tasks, while neglecting the importance of accurate context and the key role of commonsense knowledge in reasoning. In this paper, we propose a novel framework to enhance visual commonsense reasoning by incorporating context and commonsense knowledge. We decompose the visual commonsense reasoning problem into four distinct but interrelated sub-problems and combine visual language models with a large language model to enable zero-shot reasoning. The uniqueness of this work lies in the proposed commonsense knowledge filtering module, which filters out relevant commonsense knowledge through the causal strength of visual context. This process constructs Visual Context and Commonsense-guided Causal Chain-of-Thought ($\mathrm{VC^{3}}$-CoT) reasoning paths, thereby providing double robustness to visual commonsense reasoning by incorporating weighted majority voting strategy. Extensive experiments on several downstream tasks demonstrate that the proposed method significantly improves performance compared to baseline models and the state-of-the-art method, and confirm the effectiveness of the proposed components. Jing Zhao 0015, Tongquan Wei, Shiliang Sun |
IEEE Trans. Multim. | 2 |
| 2026 | DVD: A Debiased Visual Dialog Model via Disentangling Knowledge Features
Chenyu Lu, Jing Zhao 0015, Shiliang Sun |
IEEE Trans. Multim. | 2 |
| 2025 | CiGA: A Cross-Layer Fine-Grained Attention Correction Method for Large Language ModelabstractFine-grained text processing is a significant domain in Natural Language Processing (NLP), including tasks such as long-document question answering, aspect-based sentiment analysis, and document summarization. Although Large Language Models (LLMs) perform excellent on many NLP tasks, they often exhibit hallucination, such as detail loss or inaccuracies in tasks that require handling fine-grained content. This shortcoming arises because LLMs’ final layers tend to lose attention to details compared to the middle layers. Existing optimization methods for LLMs lack a focus on attention mechanisms for fine-grained information. To address this issue, we propose a novel Cross-Layer Fine-Grained Attention Correction method (CiGA). CiGA includes two correction terms that integrate detail-oriented attention from middle layers into the final layers. Experimental results demonstrate that CiGA significantly improves LLMs’ performance on fine-grained text processing tasks. Jing Zhao 0015 |
ICASSP | 2 |
| 2025 | ITJP: Image and Text Joint Prompts for Few-Shot Whole Slide Image ClassificationabstractMultiple instance learning has achieved remarkable results in whole slide image (WSI) classification. However, constrained by patient privacy and the scarcity of cancer data, the insufficient quantity of data poses a challenge in training models with the extensive parameters required for large-pixel WSIs, giving rise to the requirement of few-shot WSI classification. Recently, pre-trained vision-language models (preVLM) have demonstrated great superiority in few-shot WSI classification tasks due to their good transfer-learning and few-shot capabilities. Therefore, we propose an Image and Text Joint Prompts for few-shot whole slide image classification method, ITJP, to construct a few-shot WSI classification model under the multiple instance learning framework. ITJP utilizes the pre-VLM to extract instance features under the unsupervised setting, constructs image prompts to guide the aggregation of instance features into bag features, and subsequently guides the classification of bag features by text prompts. Specifically, we propose an image prototype guided aggregation strategy, where image prototype is obtained by clustering patches extracted from representative WSIs. The image prototype guides the aggregation of instance features and is directly compared with the image patches to achieve more accurate similarity, thereby enhancing the focus on classification-relevant features. Furthermore, low-rank linear transformation is designed to obtain variable image prototype, enhancing adaptability in specific WSI classification task and enabling effective adaptation to the few-shot scenario. We conduct extensive experiments on two WSI datasets to demonstrate the significant performance of the ITJP for few-shot WSI classification. Xinzhu Zhang, Zhikang Zhao, Jing Zhao 0015 |
ICME | 4 |
| 2025 | RPMIL: Rethinking Uncertainty-Aware Probabilistic Multiple Instance Learning for Whole Slide Pathology DiagnosisabstractWhole slide images (WSIs) are gigapixel digital scans of traditional pathology slides, offering substantial support for cancer diagnosis. Current multiple instance learning (MIL) methods for WSIs typically extract instance features and aggregate these into a single bag feature for prediction. We observe that these MIL methods rely on point estimation, where each bag is mapped to a deterministic embedding. Such MIL methods based on point estimation fail to capture the full spectrum of data variability due to the reliance on fixed embedding, especially when the number of trainable bags is limited. In this paper, we rethink probabilistic modeling in MIL and propose RPMIL, an uncertainty-aware probabilistic MIL method for whole slide pathology diagnosis. RPMIL learns a probabilistic aggregator to consolidate instance features into dynamic bag feature distributions instead of a deterministic bag feature. Specifically, we employ a variational autoencoder approach to compress multiple instance features into a low-dimension space with probabilistic representation and obtain the bag feature distribution formulated by the mean and variance. Furthermore, we drive the prediction by jointly leveraging the instance feature distribution and bag feature distribution. We evaluate the WSI classification performance on two public datasets: Camelyon16 and TCGA-NSCLC. Extensive experiments demonstrate that our method surpasses point estimation methods in MIL, achieving state-of-the-art levels. Zhikang Zhao, Kaitao Chen, Jing Zhao 0015 |
IJCAI | 3 |
| 2025 | MedCDA: Counterfactual Data Augmentation for Medical Image Analysis
Kexin Yao, Jing Zhao 0015 |
PRCV (14) | 2 |
| 2025 | Multiple instance learning with hierarchical discrimination and smoothing attention for histopathological diagnosis
Jing Zhao 0015, Zhikang Zhao, Xueru Song, Shiliang Sun |
Appl. Intell. | 1 |
| 2025 | Taming vision transformers for clinical laryngoscopy assessment
Xinzhu Zhang, Jing Zhao 0015, Daoming Zong, Henglei Ren, Chunli Gao |
J. Biomed. Informatics | 2 |
| 2025 | Few-Shot Knowledge Graph Completion With Star and Ring Topology Information AggregationabstractFew-shot knowledge graph completion (FKGC) addresses the long-tail problem of relations by leveraging a few observed support entity pairs to infer unknown facts for tail-located relations. Learning the relation representation of entity pairs and evaluating the match of query and support entity pairs are the two key steps of FKGC. Existing methods learn the representation of entity pairs by either aggregating neighbors of entities or integrating relation representations in the connected paths from head to tail. However, in few-shot scenarios, the limited number of support entity pairs and insufficient structural information with a single neighborhood topology will lead to matching failure. To this end, we consider the star and ring topological information for a given entity pair: (1) Entity neighborhood, which captures multi-hop neighbors of entities; (2) Relational path, which characterizes compound relation forms. Furthermore, to effectively fuse the two kinds of heterogeneous topological information, we design the multi-aggregator and the fine-grained path correlation matching algorithm to obtain more delicate and balanced matching. Based on the proposed relational path correlation matching module, we propose the relation adaptive network to solve the few-shot temporal knowledge graph completion problem. The experimental results show that our method continuously outperforms the state-of-the-art methods. Jing Zhao 0015, Xinzhu Zhang, Shiliang Sun |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2025 | De-Pessimism Offline Reinforcement Learning via Value CompensationabstractOffline reinforcement learning (RL) has been widely used in practice due to its efficient data utilization, but it still faces the challenge of training vulnerability caused by policy deviation. Existing offline RL methods that add policy constraints or perform conservative Q-value estimation are pessimistic, making the learned policy suboptimal. In this article, we address the pessimism problem by focusing on accurate Q-value estimation. We propose the de-pessimism (DEP) operator to estimate Q values using the optimal Bellman operator or the compensation operator according to whether the actions are in the behavior support set. The compensation operator qualitatively determines the positive or negative nature of out-of-distribution (OOD) actions based on their performance compared with the behavior actions. It leverages differences in state values to compensate for the Q value of positive OOD actions, thereby alleviating pessimism. We theoretically demonstrate the convergence of DEP and its effectiveness in policy improvement. To further advance the practical application, we integrate DEP into the soft actor-critic (SAC) algorithm, yielding the value-compensated de-pessimism offline RL (DoRL-VC). Experimentally, DoRL-VC achieves state-of-the-art (SOTA) performance across mujoco locomotion, Maze 2-D, and challenging Adroit tasks, illustrating the efficacy of DEP in mitigating pessimism. Zhenbo Huang, Jing Zhao 0015, Shiliang Sun |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | VB-Adapter: Variational Bayesian Adapter for Cross-Domain Speech Representation LearningabstractTo leverage the abundant speech data available for pretraining, current models excel in generalization across diverse tasks. Nevertheless, real-world challenges emerge when addressing unfamiliar speech scenarios far from the pretrained speech, owing to the domain shift between pretraining (source) and fine-tuning (target) data. To overcome this barrier, we propose a variational Bayesian adapter (VB-Adapter) for cross-domain speech representation learning during fine-tuning. First, we establish a latent variable model to construct a desired posterior distribution after incorporating domain-specific knowledge to bridge the gap between the source and target domains. Then, an adaptive objective is presented to maximize the mutual information of the latent variables with and without domain-specific knowledge to facilitate model adaptation. Finally, we introduce contrastive learning on samples to optimize the lower bound of the above adaptive objective. Our experiments apply the VB-Adapter on transformers for dysarthric speech recognition (DSR) and the integration of Whisper-encoder and Llama for Mandarin speech recognition (MSR). The results reveal the effectiveness of VB-Adapter in modeling the uncertainties arising from domain shift and enhancing the robustness of speech representations. Jing Zhao 0015, Qimin Huang, Shanhu Wang, Shiliang Sun |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | CaMIL: Causal Multiple Instance Learning for Whole Slide Image ClassificationabstractWhole slide image (WSI) classification is a crucial component in automated pathology analysis. Due to the inherent challenges of high-resolution WSIs and the absence of patch-level labels, most of the proposed methods follow the multiple instance learning (MIL) formulation. While MIL has been equipped with excellent instance feature extractors and aggregators, it is prone to learn spurious associations that undermine the performance of the model. For example, relying solely on color features may lead to erroneous diagnoses due to spurious associations between the disease and the color of patches. To address this issue, we develop a causal MIL framework for WSI classification, effectively distinguishing between causal and spurious associations. Specifically, we use the expectation of the intervention P(Y | do(X)) for bag prediction rather than the traditional likelihood P(Y | X). By applying the front-door adjustment, the spurious association is effectively blocked, where the intervened mediator is aggregated from patch-level features. We evaluate our proposed method on two publicly available WSI datasets, Camelyon16 and TCGA-NSCLC. Our causal MIL framework shows outstanding performance and is plug-and-play, seamlessly integrating with various feature extractors and aggregators. Kaitao Chen, Shiliang Sun, Jing Zhao 0015 |
AAAI | 3 |
| 2024 | Reward-free offline reinforcement learning: Optimizing behavior policy via action exploration
Zhenbo Huang, Shiliang Sun, Jing Zhao 0015 |
Knowl. Based Syst. | 3 |
| 2024 | HWLane: HW-Transformer for Lane DetectionabstractLane detection is one of the most fundamental tasks in autonomous driving perception, but it still faces many challenges in some special driving scenarios. For example, in dazzling light, crowded roads, etc., lane detection is very dependent on surrounding visual cues. Previous segmentation-based lane detection methods have not paid enough attention to the surrounding visual range, resulting in poor performance. In this paper, we design a novel lane detection network namely HW-Transformer, which is based on row and column multi-head self-attention. It restricts the attention only to their respective rows and columns, and transfers information across rows and columns by intersection features. In this way, the attention to the visual range around the lane is greatly expanded, and the communication of global information can be achieved through intersecting features. In addition, we further propose a self-attention knowledge distillation (SAKD) method for the Transformer model, where higher-level attention guides lower-level attention to learn. SAKD not only helps to improve the performance of lane detection, but also has universality in better learning semantic features from general images. Extensive experiments on BDD100K, TuSimple, CULane, and VIL100 datasets demonstrate that our method outperforms the state-of-the-art segmentation-based lane detection methods. We also apply the proposed SAKD to DeiT-tiny, and it achieves 1.51 Top-1 accuracy improvement on ImageNet-1K dataset. Our code will be available at https://github.com/Cuibaby/HWLane. Jing Zhao 0015, Zengyu Qiu, Huiqin Hu, Shiliang Sun |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2024 | VirPNet: A Multimodal Virtual Point Generation Network for 3D Object DetectionabstractLiDAR and camera are the most common used sensors to percept the road scenes in autonomous driving. Current methods tried to fuse the two complementary information to boost 3D object detection. However, there are still two burning problems for multi-modality 3D object detection. One is the detection problem for the objects with sparse point clouds. The other is the misalignment of different sensors caused by the fixed physical locations. Therefore, this paper argues that explicitly fusing information from the two modalities with the physical misalignment is suboptimal for multi-modality 3D object detection. This paper presents a novel virtual point generation network, VirPNet, to overcome the multi-modality fusion challenges. On one hand, it completes sparse point cloud objects from image source and improves the final detection accuracy. On the other hand, it directly detects 3D targets from raw point clouds to avoid the physical misalignment between LiDAR and camera sensors. Different from previous point cloud completion methods, VirPNet fully utilizes the geometric information of pixels and point clouds and simplifies 3D point cloud regression into a 2D distance regression problem through a virtual plane. Experimental results on KITTI 3D object detection dataset and nuScenes dataset demonstrate that VirPNet improves the detection accuracy with the help of the generated virtual points. Lin Wang 0053, Shiliang Sun, Jing Zhao 0015 |
IEEE Trans. Multim. | 3 |
| 2023 | Effective Domain Adaptation for Robust Dysarthric Speech Recognition
Shanhu Wang, Jing Zhao 0015, Shiliang Sun |
ICONIP (10) | 2 |
| 2023 | Knowledge enhanced zero-resource machine translation using image-pivoting
Jing Zhao 0015, Shilinag Sun, Yichu Lin |
Appl. Intell. | 2 |
| 2023 | Multi-view Collaborative Gaussian Process Dynamical SystemsabstractGaussian process dynamical systems (GPDSs) have shown their effectiveness in many tasks of machine learning. However, when they address multi-view data, current GPDSs do not explicitly model the dependence between private and shared latent variables. Instead, they introduce structurally and intrinsically discrete segmentation in the latent space. In this paper, we propose the multi-view collaborative Gaussian process dynamical systems (McGPDSs) model, which assumes that the private latent variable for each view is controlled by its dynamical prior and the shared latent variable. The relevance between private and shared latent variables can be automatically learned by optimization in the Bayesian framework. The model is capable of learning an effective latent representation and generating novel data of one view given data of the other view. We evaluate our model on two-view data sets, and our model obtains better performance compared with the state-of-the-art multi-view GPDSs. Shiliang Sun, Jingjing Fei, Jing Zhao 0015 |
J. Mach. Learn. Res. | 3 |
| 2023 | Concept Drift Adaptation for Time Series Anomaly Detection via Transformer
Chaoyue Ding, Jing Zhao 0015, Shiliang Sun |
Neural Process. Lett. | 2 |
| 2023 | Fair Transfer Learning with Factor Variational Auto-Encoder
Shaofan Liu, Shiliang Sun, Jing Zhao 0015 |
Neural Process. Lett. | 3 |
| 2022 | Enhancing Unsupervised Domain Adaptation via Semantic Similarity Constraint for Medical Image SegmentationabstractThis work proposes a novel unsupervised cross-modality adaptive segmentation method for medical images to tackle the performance degradation caused by the severe domain shift when neural networks are being deployed to unseen modalities. The proposed method is an end-2-end framework, which conducts appearance transformation via a domain-shared shallow content encoder and two domain-specific decoders. The feature extracted from the encoder is enhanced to be more domain-invariant by a similarity learning task using the proposed Semantic Similarity Mining (SSM) module which has a strong help of domain adaptation. The domain-invariant latent feature is then fused into the target domain segmentation sub-network, trained using the original target domain images and the translated target images from the source domain in the framework of adversarial training. The adversarial training is effective to narrow the remaining gap between domains in semantic space after appearance alignment. Experimental results on two challenging datasets demonstrate that our method outperforms the state-of-the-art approaches. Shiliang Sun, Jing Zhao 0015, Dongyu Shi |
IJCAI | 3 |
| 2022 | TiRGN: Time-Guided Recurrent Graph Network with Local-Global Historical Patterns for Temporal Knowledge Graph ReasoningabstractTemporal knowledge graphs (TKGs) have been widely used in various fields that model the dynamics of facts along the timeline. In the extrapolation setting of TKG reasoning, since facts happening in the future are entirely unknowable, insight into history is the key to predicting future facts. However, it is still a great challenge for existing models as they hardly learn the characteristics of historical events adequately. From the perspective of historical development laws, comprehensively considering the sequential, repetitive, and cyclical patterns of historical facts is conducive to predicting future facts. To this end, we propose a novel representation learning model for TKG reasoning, namely TiRGN, a time-guided recurrent graph network with local-global historical patterns. Specifically, TiRGN uses a local recurrent graph encoder network to model the historical dependency of events at adjacent timestamps and uses the global history encoder network to collect repeated historical facts. After the trade-off between the two encoders, the final inference is performed by a decoder with periodicity. We use six benchmark datasets to evaluate the proposed method. The experimental results show that TiRGN outperforms the state-of-the-art TKG reasoning methods in most cases. Shiliang Sun, Jing Zhao 0015 |
IJCAI | 3 |
| 2022 | Multi-Modal Adversarial Example Detection with TransformerabstractAlthough deep neural networks have shown great potential for many tasks, they are vulnerable to adversarial examples, which are generated by adding small perturbations to natural examples. Recently, many studies have proved that making full use of different modalities can effectively enhance the representational ability of deep neural networks. We propose a multi-modal deep fusion Transformer, termed MDFT. First, the audio feature and the rich semantic text features are extracted by audio encoders and text encoders, respectively. Then, multi-modal attention mechanisms are established to capture the high-level interactions between the audio and linguistic domains to obtain joint multi-modal representation. Finally, the representation is propagated to a dense layer to generate the detection result. The accuracy of this model compared with its unimodal variant on WiAd dataset and BlAd dataset are improved by 0.12 % and 0.19 %, respectively. Experimental results on the two datasets show that MDFT outperforms its unimodal variant model. Chaoyue Ding, Shiliang Sun, Jing Zhao 0015 |
IJCNN | 3 |
| 2022 | Robust Cross-Modal Retrieval by Adversarial TrainingabstractCross-modal retrieval is usually implemented based on cross-modal representation learning, which is used to extract semantic information from cross-modal data. Recent work shows that cross-modal representation learning is vulnerable to adversarial attacks, even using large-scale pre-trained networks. By attacking the representation, it can be simple to attack the downstream tasks, especially for cross-modal retrieval tasks. Adversarial attacks on any modality will easily lead to obvious retrieval errors, which brings the challenge to improve the adversarial robustness of cross-modal retrieval. In this paper, we propose a robust cross-modal retrieval method (RoCMR), which generates adversarial examples for both the query modality and candidate modality and performs adversarial training for cross-modal retrieval. Specifically, we generate adversarial examples for both image and text modalities and train the model with benign and adversarial examples in the framework of contrastive learning. We evaluate the proposed RoCMR on two datasets and show its effectiveness in defending against gradient-based attacks. Shiliang Sun, Jing Zhao 0015 |
IJCNN | 3 |
| 2022 | Diverse Machine Translation with Translation MemoryabstractThe challenge of diverse machine translation is to ensure both diversity and quality, in which diversity desires the distinction between multiple hypotheses in the syntactic or word level, while quality requires hypotheses to be consistent with certain references. Existing work on diverse machine translation, most of which boosts translation diversity but compromises trans-lation quality, either resorts to special search strategies to realize word-level diversity or trains several decoders simultaneously to enable possible syntactic multiplicity. In this work, we propose to improve the translation diversity without a severe drop in quality through constructing a novel translation memory based NMT and designing a global diverse beam search strategy. Specifically, the translation memory retrieved from target-size corpus act as experts to interact with standard NMT, which not only generates various hypotheses, but also enhances the quality to a great degree. Moreover, we exploit a novel diverse beam search to further avoid token reuse across different hypotheses and improve local diversity. Experiments on two benchmarks (JRC-Acquis and WMT) demonstrate that our approaches achieve a compelling promotion both of translation quality and diversity compared with other diverse approaches. Jing Zhao 0015, Shiliang Sun |
IJCNN | 2 |
| 2022 | MFIALane: Multiscale Feature Information Aggregator Network for Lane DetectionabstractLane detection differs from general object detection in that lane lines are usually long and narrow in the road image, and more attention to image features at different scales is required to reason about lane lines under occlusion, degradation, and bad weather. However, most existing semantic segmentation-based lane detection methods focus on solving the convolutional receptive field through aggregating information vertically and horizontally in the same feature map, which may ignore important information contained in multi-scale features. Besides, the high-level semantic information of whether the lane exists is not fully utilized, as they often add a module at the final stage of the network output to determine whether the lane exists, which is a dispensable for their network. Based on the above analysis, we design a novel lane detection network based on semantic segmentation which consists of a Multi-scale Feature Information Aggregator (MFIA) module and a Channel Attention (CA) module. Many experiments on the TRLane dataset, the generated Lane dataset, BDD100K dataset, TuSimple dataset, VIL-100 dataset and CULane dataset show that our approach can achieve the state-of-the-art performance (our code will be available athttps://github.com/Cuibaby/MFIALane). In addition, considering that different perceptual tasks in autonomous driving are able to share the feature extraction network, we also conduct the experiment for drivable area segmentation on BDD100K dataset. Our approach also achieves good results compared to many existing methods, showing that our proposed model is capable of simultaneously handling multiple perceptual tasks in autonomous driving scenarios. Zengyu Qiu, Jing Zhao 0015, Shiliang Sun |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2022 | Conditional Random Fields for Multiview Sequential Data ModelingabstractRecently, multiview learning has been increasingly focused on machine learning. However, most existing multiview learning methods cannot directly deal with multiview sequential data, in which the inherent dynamical structure is often ignored. Especially, most traditional multiview machine learning methods assume that the items at different time slices within a sequence are independent of each other. In order to solve this problem, we propose a new multiview discriminant model based on conditional random fields (CRFs) to model multiview sequential data, called multiview CRF. It inherits the advantages of CRFs that build a relationship between items in each sequence. Moreover, by introducing specific features designed on the CRFs for multiview data, the multiview CRF not only considers the relationship among different views but also captures the correlation between the features from the same view. Particularly, some features can be reused or divided into different views to build an appropriate size of feature space. This helps to avoid underfitting problems caused by too small feature space or overfitting problems caused by too large feature space. In order to handle large-scale data, we use the stochastic gradient method to speed up our model. The experimental results on the text and video data illustrate the superiority of the proposed model. Shiliang Sun, Ziang Dong, Jing Zhao 0015 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2021 | ASHF-Net: Adaptive Sampling and Hierarchical Folding Network for Robust Point Cloud CompletionabstractEstimating the complete 3D point cloud from an incomplete one lies at the core of many vision and robotics applications. Existing methods typically predict the complete point cloud based on the global shape representation extracted from the incomplete input. Although they could predict the overall shape of 3D objects, they are incapable of generating structure details of objects. Moreover, the partial input point sets obtained from range scans are often sparse, noisy and non-uniform, which largely hinder shape completion. In this paper, we propose an adaptive sampling and hierarchical folding network (ASHF-Net) for robust 3D point cloud completion. Our main contributions are two-fold. First, we propose a denoising auto-encoder with an adaptive sampling module, aiming at learning robust local region features that are insensitive to noise. Second, we propose a hierarchical folding decoder with the gated skip-attention and multi-resolution completion goal to effectively exploit the local structure details of partial inputs. We also design a KL regularization term to evenly distribute the generated points. Extensive experiments demonstrate that our method outperforms existing state-of-the-art methods on multiple 3D point cloud completion benchmarks. Daoming Zong, Shiliang Sun, Jing Zhao 0015 |
AAAI | 3 |
| 2021 | Multi-Task Transformer with Input Feature Reconstruction for Dysarthric Speech RecognitionabstractDysarthria is a motor speech disorder caused by damage to the part of the nervous system that controls the physical production of speech. It poses great challenges in building robust dysarthric speech recognition (DSR) due to the high inter- and intra-speaker variability. To this end, we propose a multi-task Transformer with input feature reconstruction as an auxiliary task, where the main task of DSR and the auxiliary reconstruction task share the same encoder network. The auxiliary task aims to reconstruct clear speech features from corrupted speech of healthy speakers (intra-domain) or dysarthric speakers (cross-domain). Further, to alleviate the imbalanced distribution of dysarthria data sets, we devise an adaptive rebalance sampling scheme to improve the utterance sampling frequency of dysarthric speech. Experimental results show that the proposed model considerably outperforms other baselines across speakers with varying severity of dysarthria. Chaoyue Ding, Shiliang Sun, Jing Zhao 0015 |
ICASSP | 3 |
| 2021 | A Sequential Contrastive Learning Framework for Robust Dysarthric Speech RecognitionabstractDysarthria is a manifestation of disruption in the neuromuscular physiology resulting in uneven, slow, slurred, harsh, or quiet speech. Despite the remarkable progress of automatic speech recognition (ASR), it poses great challenges in developing stable ASR for dysarthric individuals due to the high intra- and inter-speaker variations and data deficiency. In this paper, we propose a contrastive learning framework for robust dysarthric speech recognition (DSR) by capturing the dysarthric speech variability. Several speech data augmentation strategies are explored to form two branches of the framework, meanwhile alleviating the scarcity of dysarthria data. We also develop an efficient projection head acting on a sequence of learned hidden representations for defining contrastive loss. Experiment results on DSR demonstrate that the model is better than or comparable to the supervised baseline. Lidan Wu, Daoming Zong, Shiliang Sun, Jing Zhao 0015 |
ICASSP | 4 |
| 2021 | Resilient Abstractive Summarization Model with Adaptively Weighted Training LossabstractRecently, abstractive summarization models are preferred over extractive summarization models as they can generate words that do not exist in the original text, whose summary descriptions are more flexible and natural. Neural network-based models learn the pattern of summary generation from the training data by modeling the relationship between the original text and the reference summary, which is very dependent on the reference summary. Although we intuitively feel that summary with higher Abstraction Degree (quantified by the number of words in the summary which do not appear in the original text) will be more general, manually generated summaries with high Abstraction Degree are most likely subliminally written with additional knowledge. It's difficult to learn the generation pattern of such reference summary using limited training data. What's more, such reference summaries can even harm the model performance. To this end, we design a learning method that can adaptively weighted difference training samples based on their Abstraction Degree, so that the model will pay less attention to the samples with higher Abstraction Degree. Experiments of LCSTS and CNN-DM dataset show that our method greatly improves the performance of the summarization model and is resilient in the face of training data containing low quality reference summaries. Shiqi Guo, Jing Zhao 0015, Shiliang Sun |
IJCNN | 2 |
| 2021 | Stick-Breaking Dependent Beta Processes with Variational Inference
Zehui Cao, Jing Zhao 0015, Shiliang Sun |
Neural Process. Lett. | 2 |
| 2020 | Multi-View Deep Attention Network for Reinforcement Learning (Student Abstract)abstractThe representation approximated by a single deep network is usually limited for reinforcement learning agents. We propose a novel multi-view deep attention network (MvDAN), which introduces multi-view representation learning into the reinforcement learning task for the first time. The proposed model approximates a set of strategies from multiple representations and combines these strategies based on attention mechanisms to provide a comprehensive strategy for a single-agent. Experimental results on eight Atari video games show that the MvDAN has effective competitive performance than single-view reinforcement learning methods. Yueyue Hu, Shiliang Sun, Xin Xu 0001, Jing Zhao 0015 |
AAAI | 4 |
| 2020 | Bayesian Adversarial Attack on Graph Neural Networks (Student Abstract)abstractAdversarial attack on graph neural network (GNN) is distinctive as it often jointly trains the available nodes to generate a graph as an adversarial example. Existing attacking approaches usually consider the case that all the training set is available which may be impractical. In this paper, we propose a novel Bayesian adversarial attack approach based on projected gradient descent optimization, called Bayesian PGD attack, which gets more general attack examples than deterministic attack approaches. The generated adversarial examples by our approach using the same partial dataset as deterministic attack approaches would make the GNN have higher misclassification rate on graph node classification. Specifically, in our approach, the edge perturbation Z is used for generating adversarial examples, which is viewed as a random variable with scale constraint, and the optimization target of the edge perturbation is to maximize the KL divergence between its true posterior distribution p(Z|D) and its approximate variational distribution qθ(Z). We experimentally find that the attack performance will decrease with the reduction of available nodes, and the effect of attack using different nodes varies greatly especially when the number of nodes is small. Through experimental comparison with the state-of-the-art attack approaches on GNNs, our approach is demonstrated to have better and robust attack performance. Jing Zhao 0015, Shiliang Sun |
AAAI | 2 |
| 2020 | Multi-view Deep Gaussian Process with a Pre-training Acceleration Technique
Han Zhu 0009, Jing Zhao 0015, Shiliang Sun |
PAKDD (2) | 2 |
| 2020 | Probabilistic inference of Bayesian neural networks with generalized expectation propagation
Jing Zhao 0015, Shaojie He, Shiliang Sun |
Neurocomputing | 1 |
| 2020 | Promoting active learning with mixtures of Gaussian processes
Jing Zhao 0015, Shiliang Sun, Zehui Cao |
Knowl. Based Syst. | 1 |
| 2020 | A Survey of Optimization Methods From a Machine Learning PerspectiveabstractMachine learning develops rapidly, which has made many theoretical breakthroughs and is widely applied in various fields. Optimization, as an important part of machine learning, has attracted much attention of researchers. With the exponential growth of data amount and the increase of model complexity, optimization methods in machine learning face more and more challenges. A lot of work on solving optimization problems or improving optimization methods in machine learning has been proposed successively. The systematic retrospect and summary of the optimization methods from the perspective of machine learning are of great significance, which can offer guidance for both developments of optimization and machine learning research. In this article, we first describe the optimization problems in machine learning. Then, we introduce the principles and progresses of commonly used optimization methods. Finally, we explore and give some challenges and open problems for the optimization in machine learning. Shiliang Sun, Zehui Cao, Han Zhu 0009, Jing Zhao 0015 |
IEEE Trans. Cybern. | 4 |
| 2019 | A Conditional Random Fields Based Framework for Multiview Sequential Data Modeling
Ziang Dong, Jing Zhao 0015, Shiliang Sun |
ICONIP (5) | 2 |
| 2018 | Active Learning Methods with Deep Gaussian Processes
Jingjing Fei, Jing Zhao 0015, Shiliang Sun |
ICONIP (3) | 2 |
| 2018 | Intelligent Educational Data Analysis with Gaussian Processes
Jiachun Wang, Jing Zhao 0015, Shiliang Sun, Dongyu Shi |
ICONIP (6) | 2 |
| 2018 | Multi-label Active Learning with Conditional Bernoulli Mixtures
Shiliang Sun, Jing Zhao 0015 |
PRICAI (1) | 3 |
| 2017 | Educational and Non-educational Text Classification Based on Deep Gaussian Processes
Jing Zhao 0015, Zeheng Tang, Shiliang Sun |
ICONIP (1) | 2 |
| 2016 | Key Course Selection for Academic Early Warning Based on Gaussian Processes
Min Yin, Jing Zhao 0015, Shiliang Sun |
IDEAL | 2 |
| 2016 | Variational Dependent Multi-output Gaussian Process Dynamical SystemsabstractThis paper presents a dependent multi-output Gaussian process (GP) for modeling complex dynamical systems. The outputs are dependent in this model, which is largely different from previous GP dynamical systems. We adopt convolved multi-output GPs to model the outputs, which are provided with a flexible multi-output covariance function. We adapt the variational inference method with inducing points for learning the model. Conjugate gradient based optimization is used to solve parameters involved by maximizing the variational lower bound of the marginal likelihood. The proposed model has superiority on modeling dynamical systems under the more reasonable assumption and the fully Bayesian learning framework. Further, it can be flexibly extended to handle regression problems. We evaluate the model on both synthetic and real-world data including motion capture data, traffic flow data and robot inverse dynamics data. Various evaluation methods are taken on the experiments to demonstrate the effectiveness of our model, and encouraging results are observed. Jing Zhao 0015, Shiliang Sun |
J. Mach. Learn. Res. | 1 |
| 2016 | High-Order Gaussian Process Dynamical Models for Traffic Flow PredictionabstractTraffic flow prediction, which predicts the future flow using historic flows, is an important task in intelligent transportation systems (ITS). Efficient and accurate models for traffic flow prediction greatly contribute to the development of ITS. In this paper, we adopt the Gaussian process dynamical model (GPDM) to a fourth-order GPDM, which is more suitable for modeling traffic flow data. Specifically, the latent variables in the fourth-order GPDM is a fourth-order Markov Gaussian process, and the weighted k-NN is incorporated in the model to predict latent variables for efficient prediction. After training the model, the future flow is estimated by the average of the results predicted by the fourth-order GPDM and k-NN. Compared with other popular methods, the proposed method performs best and yields significant improvements of prediction performance. Jing Zhao 0015, Shiliang Sun |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2015 | Revisiting Gaussian Process Dynamical Models
Jing Zhao 0015, Shiliang Sun |
IJCAI | 1 |
| 2015 | Modeling and recognizing human trajectories with beta process hidden Markov models
Shiliang Sun, Jing Zhao 0015, Qingbin Gao |
Pattern Recognit. | 2 |
| 2014 | Variational Dependent Multi-output Gaussian Process Dynamical Systems
Jing Zhao 0015, Shiliang Sun |
Discovery Science | 1 |