Jiankai Li

dblp:193/6273 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Beyond benchmarks of IUGC: Rethinking requirements of deep learning method for intrapartum ultrasound biometry from fetal ultrasound videos
Jieyun Bai, Yitong Tang, Zhuonan Liang, Jianan Fan, Lisa Mcguire, Jillian Clarke, Tom Weidong Cai, Jacqueline Spurway, Yubo Tang, Shiye Wang, Wenda Shen, Wangwang Yu, Philippe Zhang, Weili Jiang, Salem Muhsin Ali Binqahal Al Nasim, Arsen Abzhanov, Numan Saeed, Mohammad Yaqub, Zunhui Xia, Hongxing Li 0001, Libin Lan, Jayroop Ramesh, Valentin Bacher, Mark Eid, Hoda Kalabizadeh, Christian Rupprecht 0001, Ana I. L. Namburete, Pak-Hei Yeung, Madeleine K. Wyburd, Nicola K. Dinsdale, Assanali Serikbey, Jiankai Li, Sung-Liang Chen, Zicheng Hu, Nana Liu, Yian Deng, Wenfeng Zhang, Mai Tuyet Nhi, Gregor Koehler, Rapheal Stock, Klaus H. Maier-Hein, Marawan Elbatel, Xiaomeng Li 0001, Saad Slimani, Victor M. Campello, Benard Ohene Botwe, Isaac Khobo, Zhenyan Han, Hongying Hou, Di Qiu, Gongning Luo, Dong Ni 0001, Yaosheng Lu, Karim Lekadir, Shuo Li 0001
Medical Image Anal.36
2024 Initno: Boosting Text-to-Image Diffusion Models via Initial Noise Optimization
abstract
Recent strides in the development of diffusion models, ex-emplified by advancements such as Stable Diffusion, have underscored their remarkable prowess in generating visu-ally compelling images. However, the imperative of achieving a seamless alignment between the generated image and the provided prompt persists as a formidable challenge. This paper traces the root of these difficulties to invalid initial noise, and proposes a solution in the form of Initial Noise Optimization (INITNO), a paradigm that refines this noise. Considering text prompts, not all random noises are effective in synthesizing semantically-faithful images. We design the cross-attention response score and the selfattention conflict score to evaluate the initial noise, bifurcating the initial latent space into valid and invalid sectors. A strategically crafted noise optimization pipeline is developed to guide the initial noise towards valid regions. Our method, validated through rigorous experimentation, shows a commendable proficiency in generating images in strict accordance with text prompts. Our code is available at https://github.com/xiefan-guo/initno.
Xiefan Guo, Jinlin Liu, Miaomiao Cui, Jiankai Li, Hongyu Yang 0001, Di Huang 0001
CVPR4
2024 Leveraging Predicate and Triplet Learning for Scene Graph Generation
abstract
Scene Graph Generation (SGG) aims to identify entities and predict the relationship tripletsin visual scenes. Given the prevalence of large visual variations of subject-object pairs even in the same predicate, it can be quite challenging to model and refine predicate representations directly across such pairs, which is however a common strategy adopted by most existing SGG methods. We observe that visual variations within the identical triplet are relatively small and certain relation cues are shared in the same type of triplet, which can potentially facilitate the relation learning in SGG. Moreover, for the long-tail problem widely studied in SGG task, it is also crucial to deal with the limited types and quantity of triplets in tail predicates. Accordingly, in this paper, we propose a Dual-granularity Relation Modeling (DRM) network to leverage fine-grained triplet cues besides the coarse-grained predicate ones. DRM utilizes contexts and semantics of predicate and triplet with Dual-granularity Constraints, generating compact and balanced representations from two perspectives to facilitate relation recognition. Furthermore, a Dual-granularity Knowledge Transfer (DKT) strategy is introduced to transfer variation from head predicates/triplets to tail ones, aiming to enrich the pattern diversity of tail classes to alleviate the long-tail problem. Extensive experiments demonstrate the effectiveness of our method, which establishes new state-of-the-art performance on Visual Genome, Open Image, and GQA datasets. Our code is available at https://github.com/jkli1998/DRM
Jiankai Li, Yunhong Wang 0001, Xiefan Guo, Ruijie Yang, Weixin Li 0001
CVPR1
2024 Three-Dimensional Transient Electromagnetic Forward Modeling for Simulating Arbitrary Source Waveform and e, db/dt, b Responses Using Rational Krylov Subspace Method
abstract
The rational Krylov subspace methods can improve the computational speed compared to conventional time-stepping approaches for calculating 3-D transient electromagnetic (TEM) method forward modeling. However, the rational Krylov subspace method simulates only the step-off response. Because primary source waveforms have nonnegligible effects on the induced responses, it is crucial to model the response induced by any given source waveform. The electric field (e) and the time derivative of the magnetic induction ($\mathrm {d} {\mathbf { b}} / \mathrm {d}t$) are commonly measured TEM responses. Case studies also show the magnetic induction ($\bf b$) response measured by magnetometers has a good resolution for exploring conductive mineral deposits. Therefore, modern TEM forward modeling algorithms should be able to simulate different types of responses. We present a new algorithm for TEM modeling using the rational Krylov subspace method. The following improvements are implemented in our approach: 1) the algorithm can efficiently compute the e and$\mathrm {d} {\mathbf { b}} / \mathrm {d}t$responses, and especially the$\bf b$response, which was less considered in other 3-D TEM studies; 2) a convolution approach is employed that allows the Krylov subspace method to simulate the source waveform effects on all three types of responses; and 3) we present the approach for computing the initial condition of b in cases of using galvanic sources. This work extends the flexibility of existing 3-D TEM modeling algorithms. Numerical examples demonstrate that the new algorithm is accurate and computationally efficient.
Jingyu Gao, Jiankai Li, Ling Huang 0007, Ji Cai, Maxim Smirnov, Thorkild Maack Rasmussen, Xiaojun Liu 0004, Guangyou Fang
IEEE Trans. Geosci. Remote. Sens.2
2024 MHRN: A Multimodal Hierarchical Reasoning Network for Topic Detection
abstract
Multimodal topic detection is an important social media analysis task with a wide variety of real-world applications. However, modeling data jointly, and inferring their topics, is challenging due to the semantic gaps between different modalities. Our insights are from the psychological findings pretaining to the hierarchical structure in humans? inherent perception of images and texts. In this paper, we propose a Multimodal Hierarchical Reasoning Network (MHRN) to perform multimodal inference for topic detection. The images and texts are represented in a hierarchical model named the Multimodal Part-whole Aware Graph (MPAG). MHRN then performs reasoning for topic inference based on three modules, which include a Bottom-Up Aggregation (BUA) module for encoding the hierarchical connections and sibling relations in MPAG, a Top-Down Guidance (TDG) module for enriching features of the nodes in MPAG guided by their parents, and a Bottom-Up Cross Aggregation (BUCA) module for capturing and aggregating the cross-modality cues to achieve effective multimodal reasoning. Extensive experiments are conducted on two benchmarks, and the results demonstrate the superiority of our approach.
Jiankai Li, Yunhong Wang 0001, Weixin Li 0001
IEEE Trans. Multim.1
2024 Zero-shot Scene Graph Generation via Triplet Calibration and Reduction
abstract
Scene Graph Generation (SGG) plays a pivotal role in downstream vision-language tasks. Existing SGG methods typically suffer from poor compositional generalizations on unseen triplets. They are generally trained on incompletely annotated scene graphs that contain dominant triplets and tend to bias toward these seen triplets during inference. To address this issue, we propose a Triplet Calibration and Reduction (T-CAR) framework in this article. In our framework, a triplet calibration loss is first presented to regularize the representations of diverse triplets and to simultaneously excavate the unseen triplets in incompletely annotated training scene graphs. Moreover, the unseen space of scene graphs is usually several times larger than the seen space, since it contains a huge number of unrealistic compositions. Thus, we propose an unseen space reduction loss to shift the attention of excavation to reasonable unseen compositions to facilitate the model training. Finally, we propose a contextual encoder to improve the compositional generalizations of unseen triplets by explicitly modeling the relative spatial relations between subjects and objects. Extensive experiments show that our approach achieves consistent improvements for zero-shot SGG over state-of-the-art methods. The code is available at https://github.com/jkli1998/T-CAR .
Jiankai Li, Yunhong Wang 0001, Weixin Li 0001
ACM Trans. Multim. Comput. Commun. Appl.1
2023 Three-Dimensional Transient Electromagnetic Forward Modeling for Simulating Arbitrary Source Waveform Using Convolution Approach
abstract
The transient electromagnetic (TEM) method utilizes artificial transmitters and measures electromagnetic (EM) responses to reveal the resistivity information of the subsurface. The current waveform of transmitters has nonnegligible effects on induced fields. Therefore, 3-D TEM forward modeling algorithms need the capability of simulating arbitrary waveforms to obtain accurate responses. In time-stepping-based 3-D TEM forward modeling, the source term (ST) approach is frequently used, which employs the source current density to model the waveform variation during time-stepping. The ST approach, however, requires fine-time discretization to describe complex waveforms, which could significantly raise the computational cost. We present a robust convolution (Conv) approach that computes the convolution between the time derivative of the waveform and the step-off response to incorporate the waveform effects in 3-D TEM modeling. The Conv approach does not discretize the waveform using time steps. Hence, it is advantageous when modeling full-waveform cases. The developed algorithm is based on the finite-element (FE) method using unstructured grids and the implicit backward Euler approach. Both galvanic and inductive transmitters are incorporated. Ground and airborne TEM surveys are tested using an actual airborne TEM waveform, a full waveform of the$2^{(n)}$-sequence pseudorandom signal, and various synthetic waveforms. Accuracy is validated against the 1-D and 3-D solutions of published studies. The ST and Conv approaches are compared. Synthetic examples show that the latter approach simplifies the waveform incorporation in TEM modeling and substantially improves time-stepping efficiency without sacrificing accuracy.
Jingyu Gao, Xiaojun Liu 0004, Wanhua Zhu, Maxim Smirnov, Thorkild Maack Rasmussen, Ling Huang 0007, Jiankai Li, Guangyou Fang
IEEE Trans. Geosci. Remote. Sens.7
2022 MGMP: Multimodal Graph Message Propagation Network for Event Detection
Jiankai Li, Yunhong Wang 0001, Weixin Li 0001
MMM (1)1
2021 Entity Relation Fusion for Real-Time One-Stage Referring Expression Comprehension
abstract
Referring Expression Comprehension (REC) is the task of grounding object which is referred by the language expression. Previous one-stage REC methods usually use one single language feature vector to represent the whole query for grounding and no reasoning between different objects is performed despite the rich relation cues of objects contained in the language expression, which depresses their grounding accuracy. Additionally, these methods mostly use the feature pyramid networks for multi-scale visual object feature extraction but ground on different feature layers separately, neglecting the connections between objects with different scales. To address these problems, we propose a novel one-stage REC method, i.e. the Entity Relation Fusion Network (ERFN) to locate referred object by relation guided reasoning on different objects. In ERFN, instead of grounding objects at each layer separately, we propose a Language Guided Multi-Scale Fusion (LGMSF) model to utilize language to guide the fusion of representations of objects with different scales into one feature map.For modeling connections between different objects, we design a Relation Guided Feature Fusion (RGFF) model that extracts entities in the language expression to enhance the referred entity feature in the visual object feature map, and further extracts relations to guide object feature fusion based on the self-attention mechanism. Experimental results show that our method is competitive with the state-of-the-art one-stage and two-stage REC methods, and can also keep inferring in real time.
Weixin Li 0001, Jiankai Li
MMAsia3
2021 3-D Marine CSEM Forward Modeling With General Anisotropy Using an Adaptive Finite-Element Method
abstract
To investigate the effect of azimuthal anisotropy on frequency-domain marine controlled-source electromagnetic (CSEM) responses, an adaptive edge-based finite-element (FE) modeling algorithm is presented in this letter. The 3-D algorithm is capable of dealing with generally anisotropic conductive media. It is implemented on unstructured tetrahedral grids, which allow for complex model geometries. The accuracy of the FE solution is controlled through adaptive mesh refinement, which is performed iteratively until the solution converges to the desired accuracy tolerance. The algorithm is validated against the quasi-analytic solutions for a 1-D layered model with anisotropy. We then simulate the marine CSEM responses over a set of 3-D anisotropic models and illustrate that the azimuthal anisotropy has a considerable influence on both the inline and broadside marine CSEM responses but to different extents.
Jiankai Li, Yuguo Li, Klaus Spitzer 0002
IEEE Geosci. Remote. Sens. Lett.1