EDBT 2026 Demo / reviewers in the wild / expert
Yuntian Chen
dblp:97/10115
· DBLP profile ↗
23ranked-venue papers
4as first author
21since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 1 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Latent Knowledge-Guided Video Diffusion for Scientific Phenomena Generation from a Single Initial FrameabstractVideo diffusion models have achieved impressive results in natural scene generation, yet they struggle to generalize to scientific phenomena such as fluid simulations and meteorological processes, where underlying dynamics are governed by scientific laws. These tasks pose unique challenges, including severe domain gaps, limited training data, and the lack of descriptive language annotations. To handle this dilemma, we extracted the latent scientific phenomena knowledge and further proposed a fresh framework that teaches video diffusion models to generate scientific phenomena from a single initial frame. Particularly, static knowledge is extracted via pre-trained masked autoencoders, while dynamic knowledge is derived from pre-trained optical flow prediction. Subsequently, based on the aligned spatial relations between the CLIP vision and language encoders, the visual embeddings of scientific phenomena, guided by latent scientific phenomena knowledge, are projected to generate the pseudo-language prompt embeddings in both spatial and frequency domains. By incorporating these prompts and fine-tuning the video diffusion model, we enable the generation of videos that better adhere to scientific laws. Extensive experiments on both computational fluid dynamics simulations and real-world typhoon observations demonstrate the effectiveness of our approach, achieving superior fidelity and consistency across diverse scientific scenarios. Qinglong Cao, Chao Ma 0004, Yuntian Chen, Xiaokang Yang 0001 |
AAAI | 5 |
| 2026 | SuperEar: Eavesdropping on Mobile Voice Calls via Stealthy Acoustic Metamaterials
Zhiyuan Ning 0003, Zhanyong Tang, Juan He 0007, Weizhi Meng 0001, Yuntian Chen, Jie Zhang 0028, Zheng Wang 0001 |
WWW | 5 |
| 2025 | Auto-Regressive Moving Diffusion Models for Time Series ForecastingabstractTime series forecasting (TSF) is essential in various domains, and recent advancements in diffusion-based TSF models have shown considerable promise. However, these models typically adopt traditional diffusion patterns, treating TSF as a noise-based conditional generation task. This approach neglects the inherent continuous sequential nature of time series, leading to a fundamental misalignment between diffusion mechanisms and the TSF objective, thereby severely impairing performance. To bridge this misalignment, and inspired by the classic Auto-Regressive Moving Average (ARMA) theory, which views time series as continuous sequential progressions evolving from previous data points, we propose a novel Auto-Regressive Moving Diffusion (ARMD) model to first achieve the continuous sequential diffusion-based TSF. Unlike previous methods that start from white Gaussian noise, our model employs chain-based diffusion with priors, accurately modeling the evolution of time series and leveraging intermediate state information to improve forecasting accuracy and stability. Specifically, our approach reinterprets the diffusion process by considering future series as the initial state and historical series as the final state, with intermediate series generated using a sliding-based technique during the forward process. This design aligns the diffusion model's sampling procedure with the forecasting objective, resulting in an unconditional, continuous sequential diffusion TSF model. Extensive experiments conducted on seven widely used datasets demonstrate that our model achieves state-of-the-art performance, significantly outperforming existing diffusion-based TSF models. Qinglong Cao, Yuntian Chen |
AAAI | 3 |
| 2025 | Context-Alignment: Activating and Enhancing LLMs Capabilities in Time SeriesabstractRecently, leveraging pre-trained Large Language Models (LLMs) for time series (TS) tasks has gained increasing attention, which involves activating and enhancing LLMs' capabilities. Many methods aim to activate LLMs' capabilities based on token-level alignment, but overlook LLMs' inherent strength in natural language processing — their deep understanding of linguistic logic and structure rather than superficial embedding processing. We propose Context-Alignment (CA), a new paradigm that aligns TS with a linguistic component in the language environments familiar to LLMs to enable LLMs to contextualize and comprehend TS data, thereby activating their capabilities. Specifically, such context-level alignment comprises structural alignment and logical alignment, which is achieved by Dual-Scale Context-Alignment GNNs (DSCA-GNNs) applied to TS-language multimodal inputs. Structural alignment utilizes dual-scale nodes to describe hierarchical structure in TS-language, enabling LLMs to treat long TS data as a whole linguistic component while preserving intrinsic token features. Logical alignment uses directed edges to guide logical relationships, ensuring coherence in the contextual semantics. Following the DSCA-GNNs framework, we propose an instantiation method of CA, termed Few-Shot prompting Context-Alignment (FSCA), to enhance the capabilities of pre-trained LLMs in handling TS tasks. FSCA can be flexibly and repeatedly integrated into various layers of pre-trained LLMs to improve awareness of logic and structure, thereby enhancing performance. Extensive experiments show the effectiveness of FSCA and the importance of Context-Alignment across tasks, particularly in few-shot and zero-shot forecasting, confirming that Context-Alignment provides powerful prior knowledge on context. The code is open-sourced at https://github.com/tokaka22/ICLR25-FSCA. Yuxiao Hu 0003, Dongxiao Zhang, Jinyue Yan, Yuntian Chen |
ICLR | 5 |
| 2025 | DragSolver: A Multi-Scale Transformer for Real-World Automotive Drag Coefficient EstimationabstractAutomotive drag coefficient ($C_d$) is pivotal to energy efficiency, fuel consumption, and aerodynamic performance. However, costly computational fluid dynamics (CFD) simulations and wind tunnel tests struggle to meet the rapid-iteration demands of automotive design. We present DragSolver, a Transformer-based framework for rapid and accurate $C_d$ estimation from large-scale, diverse 3D vehicle models.
DragSolver tackles four key real-world challenges:
(1) multi-scale feature extraction to capture both global shape and fine local geometry;
(2) heterogeneous scale normalization to handle meshes with varying sizes and densities;
(3) surface-guided gating to suppress internal structures irrelevant to external aerodynamics;
and (4) epistemic uncertainty estimation via Monte Carlo dropout for risk-aware design.
Extensive evaluations on three industrial-scale datasets (DrivaerNet, DrivaerNet++, and DrivaerML) show that DragSolver outperforms existing approaches in accuracy and generalization, achieving an average reduction of relative $L_2$ error by 58.7% across real-world datasets. Crucially, DragSolver is the first to achieve reliable, real-time $C_d$ inference on production-level automotive geometries. Yuntian Chen |
ICML | 2 |
| 2025 | Domain Prompt Learning with Quaternion Networks (Extended Abstract)abstractFoundational vision-language models (VLMs) like CLIP have revolutionized image recognition, but adapting them to specialized domains with limited data remains challenging. We propose Domain Prompt Learning with Quaternion Networks (DPLQ), which leverages domain-specific foundation models and quaternion-based prompt tuning to effectively transfer recognition capabilities. Our method achieves state-of-the-art results in remote sensing and medical imaging tasks. This extended abstract highlights the key contributions and performance of DPLQ. Qinglong Cao, Zhengqin Xu, Yuntian Chen, Chao Ma 0004, Xiaokang Yang 0001 |
IJCAI | 3 |
| 2025 | MMT-fgMDI: Metapath-driven multimodal transformer framework for predicting fine-grained metabolite-drug interactions
Wenzhi Liu, Pengli Lu, Yuntian Chen |
Knowl. Based Syst. | 3 |
| 2025 | Open-Vocabulary High-Resolution Remote Sensing Image Semantic SegmentationabstractOpen-vocabulary image semantic segmentation (OVS) seeks to segment images into semantic regions across an open set of categories. Existing OVS methods commonly depend on foundational vision-language models and utilize similarity computation to tackle OVS tasks. However, these approaches are predominantly tailored to natural images and struggle with the unique characteristics of high-resolution remote sensing images, such as rapidly changing orientations and significant scale variations. To tackle this dilemma, we propose the first OVS framework specifically designed for high-resolution remote sensing imagery, introducing a rotation-aggregative similarity computation module to enhance segmentation across varying orientations and a multi-scale feature integration strategy to generate scale-aware semantic masks. Additionally, we establish the first open-sourced OVS benchmark for remote sensing, comprising four public datasets. Experiments demonstrate that our framework effectively addresses orientation and scale challenges, achieving state-of-the-art performance. All codes and datasets are available at https://github.com/caoql98/OVRS. Qinglong Cao, Yuntian Chen, Chao Ma 0004, Xiaokang Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | HSFormer: Multiscale Hybrid Sparse Transformer for Uncertainty-Aware Cloud and Shadow RemovalabstractClouds and their shadows hinder accurate analysis of optical remote sensing imagery, making cloud removal an indispensable preprocessing step in remote sensing. However, existing methods lack efficient long-range modeling capabilities and overlook the impact of cloud irregularities and uncertainties on cloud removal. Additionally, resolving spectral confusion between cloud shadows and surface information remains a significant challenge. To tackle this issue, this study introduces an innovative cloud removal algorithm termed the multiscale hybrid sparse transformer (HSFormer), which adaptively removes clouds and shadows while reconstructing land surface semantics. HSFormer leverages pixel correlation explicit sparsity and uncertainty-driven implicit sparsity to maximize attention gains, enabling efficient cloud recognition and removal. The global pixel correlation based on attention relations enhances the semantic integrity of reconstructed images and avoids information loss across frequency domains. Furthermore, the uncertainty-guided adaptive receptive field enhances the model’s ability to resolve complex cloud-covered spatial relationships and reduces the spatial uncertainty of the reconstructed image. Experiments on simulated cloud shadow, real RICE, and full-band WHUS2-CRv datasets demonstrate HSFormer’s superiority over existing methods, improving PSNR and SSIM by 0.49% and 0.62%, respectively, in average evaluations across all bands of the WHUS2-CRv dataset and effectively addresses spectral aliasing between cloud shadows and dark surfaces. Changqi Sun, Yuntian Chen, Qinglong Cao, Longfeng Nie, Zhenzhong Zeng, Dongxiao Zhang |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Accelerating Private Large Transformers Inference Through Fine-Grained Collaborative ComputationabstractHomomorphic encryption (HE) and secret sharing (SS) enable computations on encrypted data, providing significant privacy benefits for large transformer-based models (TBM) in sensitive sectors like medicine and finance. However, private TBM inference incurs significant costs due to the coarse-grained application of HE and SS. We present FASTLMPI, a new approach to accelerate private TBM inference through fine-grained computation optimization. Specifically, through the fine-grained co-design of homomorphic encryption and secret sharing, FASTLMPI achieves efficient protocols for matrix multiplication, SoftMax, LayerNorm, and GeLU. In addition, FASTLMPI introduces a precise segmented approximation technique for differentiable non-linear functions, improving its fitting accuracy while maintaining a low polynomial degree. Compared to solution BOLT (S&P’24), FASTLMPI shows a remarkable 25.1% to 55.3% decrease in runtime and an impressive 39.0% reduction in communication costs. Yuntian Chen, Zhanyong Tang, Tianpei Lu, Bingsheng Zhang, Zhiying Shi, Zheng Wang 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2024 | Domain-Controlled Prompt LearningabstractLarge pre-trained vision-language models, such as CLIP, have shown remarkable generalization capabilities across various tasks when appropriate text prompts are provided. However, adapting these models to specific domains, like remote sensing images (RSIs), medical images, etc, remains unexplored and challenging. Existing prompt learning methods often lack domain-awareness or domain-transfer mechanisms, leading to suboptimal performance due to the misinterpretation of specific images in natural image patterns. To tackle this dilemma, we proposed a Domain-Controlled Prompt Learning for the specific domains. Specifically, the large-scale specific domain foundation model (LSDM) is first introduced to provide essential specific domain knowledge. Using lightweight neural networks, we transfer this knowledge into domain biases, which control both the visual and language branches to obtain domain-adaptive prompts in a directly incorporating manner. Simultaneously, to overcome the existing overfitting challenge, we propose a novel noisy-adding strategy, without extra trainable parameters, to help the model escape the suboptimal solution in a global domain oscillation manner. Experimental results show our method achieves state-of-the-art performance in specific domain image recognition datasets. Our code is available at https://github.com/caoql98/DCPL. Qinglong Cao, Zhengqin Xu, Yuntian Chen, Chao Ma 0004, Xiaokang Yang 0001 |
AAAI | 3 |
| 2024 | Domain Prompt Learning with Quaternion NetworksabstractPrompt learning has emerged as a potent and resource-efficient technique in large Vision-Language Models (VLMs). However, its application in adapting VLMs to specialized domains like remote sensing and medical imaging, termed domain prompt learning, remains relatively unexplored. Although large-scale domain-specific foundation models offer a potential solution, their focus on a singular vision level presents challenges in prompting both vision and language modalities. To address this limitation, we propose leveraging domain-specific knowledge from these foundation models to transfer the robust recognition abilities of VLMs from generalized to specialized domains, employing quaternion networks. Our method entails utilizing domain-specific vision features from domain-specific foundation models to guide the transformation of generalized contextual embeddings from the language branch into a specialized space within quaternion networks. Furthermore, we introduce a hierarchical approach that derives vision prompt features by analyzing intermodal relationships between hierarchical language prompt features and domain-specific vision features. Through this mechanism, quaternion networks can effectively explore intermodal relationships in specific domains, facilitating domain-specific vision-language contrastive learning. Extensive experiments conducted on domain-specific datasets demonstrate that our proposed method achieves new state-of-the-art results in prompt learning. Codes are available at https://github.com/caoq198/DPLQ. Qinglong Cao, Zhengqin Xu, Yuntian Chen, M. Chao, Xiaokang Yang 0001 |
CVPR | 3 |
| 2024 | Focus on Hiders: Exploring Hidden Threats for Enhancing Adversarial TrainingabstractAdversarial training is often formulated as a min-max problem, however, concentrating only on the worst adversarial examples causes alternating repetitive confusion of the model, i.e., previously defended or correctly classified samples are not defensible or accurately classifiable in subsequent adversarial training. We characterize such non-ignorable samples as “hiders”, which reveal the hidden high-risk regions within the secure area obtained through adversarial training and prevent the model from finding the real worst cases. We demand the model to prevent hiders when defending against adversarial examples for improving accuracy and robustness simultaneously. By rethinking and redefining the min-max optimization problem for adversarial training, we propose a generalized adversarial training algorithm called Hider-Focused Adversarial Training (HFAT). HFAT introduces the iterative evolution optimization strategy to simplify the optimization problem and employs an auxiliary model to reveal hiders, effectively combining the optimization directions of standard adversarial training and prevention hiders. Furthermore, we introduce an adaptive weighting mechanism that facilitates the model in adaptively adjusting its focus between adversarial examples and hiders during different training periods. We demonstrate the effectiveness of our method based on extensive experiments, and ensure that HFAT can provide higher robustness and accuracy. Yuxiao Hu 0003, Yinpeng Dong, Dongxiao Zhang, Yuntian Chen |
CVPR | 5 |
| 2024 | SecureTLM: Private inference for transformer-based large model with MPCabstractTransformer-based Large Models (TLM), such as generative pre-trained models (GPT), have become increasingly popular for practical applications through Deep Learning as a Service (DLaaS). They have been extensively used in natural language processing and computer vision. However, concerns regarding potential private data leakage arise with this type of inference service. While some private inference techniques can protect privacy, they often introduce high latency and approximate replacements in the design protocols, resulting in changes to the model structure and decreased accuracy. In this research, we present SecureTLM, a private inference method based on secure multi-party computation (MPC) that does not require modifications to the underlying model structure. SecureTLM offers protocols for crucial computations in TLM, such as Multiplication, Softmax, GeLU, and LayerNorm, without altering the model structure. Experimental results demonstrate that SecureTLM ensures data privacy, maintains correctness, and achieves efficiency in private inference tasks. Yuntian Chen, Xianjia Meng, Zhiying Shi, Jingzhi Lin |
Inf. Sci. | 1 |
| 2024 | Break the Bias: Delving Semantic Transform Invariance for Few-Shot SegmentationabstractFew-shot semantic segmentation (FSS) aims to segment objects of unseen classes in query images with only a few annotated support images. Existing FSS algorithms typically focus on mining category representations from the single-view support to match semantic objects of the single-view query. However, the limited annotated samples render the single-view matching struggle to perceive the varying characteristics of novel objects, which results in a restricted learning space for novel categories and further induces a biased segmentation with demoted parsing performance. To address this challenge, inspired by the semantic transform invariance, this paper proposes a fresh few-shot segmentation framework to break the bias and perform invariant segmentation in a multi-view matching manner. Specifically, original and transform support features from different perspectives with the same semantics are learnable fused to obtain the transform invariance prototype with a stronger category representation ability. Simultaneously, aiming at providing better parsing guidance, the Transform Invariance Guidance Mask Generation (TIGM) module is proposed to integrate prior knowledge from different perspectives. Finally, segmentation predictions from varying views are complementarily merged in the Transform Invariance Semantic Prediction (TISP) module to decide the uncertain area and yield precise segmentation predictions. Extensive experiments on both PASCAL-5i and COCO-20i datasets demonstrate the effectiveness of our approach and show that our method could achieve state-of-the-art performance. Code is available at https://github.com/caoql98/BBD. Qinglong Cao, Yuntian Chen, Chao Ma 0004, Xiaokang Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Few-Shot Rotation-Invariant Aerial Image Semantic SegmentationabstractFew-shot aerial image semantic segmentation is a challenging task that requires precisely parsing unseen-category objects in query aerial images with limited annotated support aerial images. Formally, category prototypes would be extracted from support samples to segment query images in a pixel-to–pixel matching manner. However, aerial objects in aerial images are often distributed with arbitrary orientations, and varying orientations could cause a dramatic feature change. This unique property of aerial images renders conventional matching manner without consideration of orientations fails to activate same-category objects with different orientations. Furthermore, the oscillation of the confidence scores in existing rotation-insensitive algorithms, engendered by the striking changes of object orientations, often leads to false recognition of lower-scored rotated semantic objects. To tackle these challenges, inspired by the intrinsic rotation invariance in aerial images, we propose a novel few-shot rotation-invariant aerial semantic segmentation network (FRINet) to efficiently segment aerial semantic objects with diverse orientations. Specifically, through extracting orientation-varying yet category-consistent support information, FRINet provides rotation-adaptive matching for each query feature in a feature-aggregation manner. Meanwhile, to encourage consistent predictions for aerial objects with arbitrary orientations, segmentation predictions from different orientations are supervised by the same label and further fused to obtain the final rotation-invariant prediction in a complementary manner. Moreover, aiming at providing a better solution searching space, the backbones are newly pre-trained in the base category to basically boost the segmentation performance. Extensive experiments on the few-shot aerial image semantic segmentation benchmark demonstrate that the proposed FRINet achieves a new state-of-the-art performance. The code is available at https://github.com/caoql98/FRINet. Qinglong Cao, Yuntian Chen, Chao Ma 0004, Xiaokang Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | A Phone-Based Distributed Ambient Temperature Measurement System With an Efficient Label-Free Automated Training StrategyabstractEnhancing the energy efficiency of buildings significantly relies on monitoring indoor ambient temperature. The potential limitations of conventional temperature measurement techniques, together with the omnipresence of smartphones, have redirected researchers' attention towards the exploration of phone-based ambient temperature estimation methods. However, existing phone-based methods face challenges such as insufficient privacy protection, difficulty in adapting models to various phones, and hurdles in obtaining enough labeled training data. In this study, we propose a distributed phone-based ambient temperature estimation system which enables collaboration among multiple phones to accurately measure the ambient temperature in different areas of an indoor space. This system also provides an efficient, cost-effective approach with a few-shot meta-learning module and an automated label generation module. It shows that with just 5 new training data points, the temperature estimation model can adapt to a new phone and reach a good performance. Moreover, the system uses crowdsourcing to generate accurate labels for all newly collected training data, significantly reducing costs. Additionally, we highlight the potential of incorporating federated learning into our system to enhance privacy protection. We believe this study can advance the practical application of phone-based ambient temperature measurement, facilitating energy-saving efforts in buildings. Dayin Chen, Xiaodan Shi, Haoran Zhang 0002, Xuan Song 0001, Dongxiao Zhang, Yuntian Chen, Jinyue Yan |
IEEE Trans. Mob. Comput. | 6 |
| 2023 | Discrete Point-Wise Attack is Not Enough: Generalized Manifold Adversarial Attack for Face RecognitionabstractClassical adversarial attacks for Face Recognition (FR) models typically generate discrete examples for target identity with a single state image. However, such paradigm of point-wise attack exhibits poor generalization against numerous unknown states of identity and can be easily defended. In this paper, by rethinking the inherent relationship between the face of target identity and its variants, we introduce a new pipeline of Generalized Manifold Adversarial Attack (GMAA)11https://github.com/tokaka22/GMAA to achieve a better attack performance by expanding the attack range. Specifically, this expansion lies on two aspects - GMAA not only expands the target to be attacked from one to many to encourage a good generalization ability for the generated adversarial examples, but it also expands the latter from discrete points to manifold by leveraging the domain knowledge that face expression change can be continuous, which enhances the attack effect as a data augmentation mechanism did. Moreover, we further design a dual supervision with local and global constraints as a minor contribution to improve the visual quality of the generated adversarial examples. We demonstrate the effectiveness of our method based on extensive experiments, and reveal that GMAA promises a semantic continuous adversarial space with a higher generalization ability and visual quality. Yuxiao Hu 0003, Dongxiao Zhang, Xin Jin 0002, Yuntian Chen |
CVPR | 6 |
| 2023 | Physics-Guided Discovery of Highly Nonlinear Parametric Partial Differential EquationsabstractPartial differential equations (PDEs) that fit scientific data can represent physical laws with explainable mechanisms for various mathematically-oriented subjects, such as physics and finance. The data-driven discovery of PDEs from scientific data thrives as a new attempt to model complex phenomena in nature, but the effectiveness of current practice is typically limited by the scarcity of data and the complexity of phenomena. Especially, the discovery of PDEs with highly nonlinear coefficients from low-quality data remains largely under-addressed. To deal with this challenge, we propose a novel physics-guided learning method, which can not only encode observation knowledge such as initial and boundary conditions but also incorporate the basic physical principles and laws to guide the model optimization. We theoretically show that our proposed method strictly reduces the coefficient estimation error of existing baselines, and is also robust against noise. Extensive experiments show that the proposed method is more robust against data noise, and can reduce the estimation error by a large margin. Moreover, all the PDEs in the experiments are correctly discovered, and for the first time we are able to discover three-dimensional PDEs with highly nonlinear coefficients. Yingtao Luo, Qiang Liu 0006, Yuntian Chen, Wenbo Hu 0001, Tian Tian 0001, Jun Zhu 0001 |
KDD | 3 |
| 2023 | TaG-Net: Topology-Aware Graph Network for Centerline-Based Vessel LabelingabstractAnatomical labeling of head and neck vessels is a vital step for cerebrovascular disease diagnosis. However, it remains challenging to automatically and accurately label vessels in computed tomography angiography (CTA) since head and neck vessels are tortuous, branched, and often spatially close to nearby vasculature. To address these challenges, we propose a novel topology-aware graph network (TaG-Net) for vessel labeling. It combines the advantages of volumetric image segmentation in the voxel space and centerline labeling in the line space, wherein the voxel space provides detailed local appearance information, and line space offers high-level anatomical and topological information of vessels through the vascular graph constructed from centerlines. First, we extract centerlines from the initial vessel segmentation and construct a vascular graph from them. Then, we conduct vascular graph labeling using TaG-Net, in which techniques of topology-preserving sampling, topology-aware feature grouping, and multi-scale vascular graph are designed. After that, the labeled vascular graph is utilized to improve volumetric segmentation via vessel completion. Finally, the head and neck vessels of 18 segments are labeled by assigning centerline labels to the refined segmentation. We have conducted experiments on CTA images of 401 subjects, and experimental results show superior vessel segmentation and labeling of our method compared to other state-of-the-art methods. Linlin Yao, Feng Shi 0001, Sheng Wang 0014, Xiao Zhang 0028, Zhong Xue, Xiaohuan Cao, Yiqiang Zhan, Lizhou Chen, Yuntian Chen, Bin Song 0002, Qian Wang 0001, Dinggang Shen |
IEEE Trans. Medical Imaging | 9 |
| 2021 | Mining Cross Features for Financial Credit Risk AssessmentabstractFor reliability, machine learning models in some areas, e.g., finance and healthcare, require to be both accurate and globally interpretable. Among them, credit risk assessment is a major application of machine learning for financial institutions to evaluate credit of users and detect default or fraud. Simple white-box models, such as Logistic Regression (LR), are usually used for credit risk assessment, but not powerful enough to model complex nonlinear interactions among features. In contrast, complex black-box models are powerful at modeling, but lack of interpretability, especially global interpretability. Fortunately, automatic feature crossing is a promising way to find cross features to make simple classifiers to be more accurate without heavy handcrafted feature engineering. However, existing automatic feature crossing methods have problems in efficiency on credit risk assessment, for corresponding data usually contains hundreds of feature fields. Qiang Liu 0006, Zhaocheng Liu, Haoli Zhang, Yuntian Chen, Jun Zhu 0001 |
CIKM | 4 |
| 2020 | Physics-Constrained Deep Learning of Geomechanical LogsabstractGeomechanical logs are of ultimate importance for subsurface description and evaluation, as well as for the exploration of underground resources, such as oil and gas, groundwater, minerals, and geothermal energy. Together with geological and hydrological properties, low-cost and high-accuracy models can be generated based on geomechanical parameters. However, it is challenging to directly measure geomechanical parameters, and they are usually estimated based on other measured quantities. For example, geomechanical logs may be obtained with certain empirical models from sonic logs together with prior information such as rock types, which are not readily available. Finding a way to directly estimate geomechanical logs based on easily available conventional well logs can result in significant cost savings and increased efficiency. In this article, we showed that deep learning via the long short-term memory network (LSTM) is effective in constructing an end-to-end model that takes the spatial dependence in well logs into consideration. We further proposed a physics-constrained LSTM, in which the physical mechanism behind the geomechanical parameters is utilized as a priori information. This state-of-the-art model is capable to directly estimate geomechanical logs based on easily available data, and it achieves higher prediction accuracy since the domain knowledge of the problem is considered. Yuntian Chen, Dongxiao Zhang |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2019 | Ensemble Neural Networks (ENN): A gradient-free stochastic methodabstractIn this study, an efficient stochastic gradient-free method, the ensemble neural networks (ENN), is developed. In the ENN, the optimization process relies on covariance matrices rather than derivatives. The covariance matrices are calculated by the ensemble randomized maximum likelihood algorithm (EnRML), which is an inverse modeling method. The ENN is able to simultaneously provide estimations and perform uncertainty quantification since it is built under the Bayesian framework. The ENN is also robust to small training data size because the ensemble of stochastic realizations essentially enlarges the training dataset. This constitutes a desirable characteristic, especially for real-world engineering applications. In addition, the ENN does not require the calculation of gradients, which enables the use of complicated neuron models and loss functions in neural networks. We experimentally demonstrate benefits of the proposed model, in particular showing that the ENN performs much better than the traditional Bayesian neural networks (BNN). The EnRML in ENN is a substitution of gradient-based optimization algorithms, which means that it can be directly combined with the feed-forward process in other existing (deep) neural networks, such as convolutional neural networks (CNN) and recurrent neural networks (RNN), broadening future applications of the ENN. Yuntian Chen, Haibin Chang, Dongxiao Zhang |
Neural Networks | 1 |