VLDB 2026 Research / reviewers in the wild / expert
Jiahui Wei
dblp:223/5846
· DBLP profile ↗
18ranked-venue papers
8as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Theory of computation · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | On the Rényi Rate-Distortion-Perception Function and Functional RepresentationsabstractWe extend the Rate-Distortion-Perception (RDP) framework to the Rényi information-theoretic regime, utilizing Sibson's $α$-mutual information to characterize the fundamental limits under distortion and perception constraints. For scalar Gaussian sources, we derive closed-form expressions for the Rényi RDP function, showing that the perception constraint induces a feasible interval for the reproduction variance. Furthermore, we establish a Rényi-generalized version of the Strong Functional Representation Lemma. Our analysis reveals a phase transition in the complexity of optimal functional representations: for $0.5<α< 1$, the coding cost is bounded by the $α$-divergence of order $α+1$, necessitating a codebook with heavy-tailed polynomial decay; conversely, for $α> 1$, the representation collapses to one with finite support, offering new insights into the compression of shared randomness under generalized notions of mutual information. Jiahui Wei, Marios Kountouris |
ISIT | 1 |
| 2026 | GAA-TSO: Geometry-Aware-Assisted Depth Completion for Transparent and Specular ObjectsabstractTransparent and specular objects are frequently encountered in daily life, factories, and laboratories. However, due to the unique optical properties, the depth information on these objects is usually incomplete and inaccurate, which poses significant challenges for downstream robotics tasks. Therefore, it is crucial to accurately restore the depth information of transparent and specular objects. Previous depth completion methods for these objects usually generate structure-less or ambiguous depth predictions. To address these issues, we propose a geometry-aware assisted depth completion method for transparent and specular objects, which focuses on exploring the 3D structural cues of the scene. Specifically, besides extracting 2D features from RGB-D input, we back-project the input depth to a point cloud and build the 3D branch to extract hierarchical scene-level 3D structural features. To exploit 3D geometric information, we design several gated cross-modal fusion modules to effectively propagate multi-level 3D geometric features to the image branch. In addition, we propose an adaptive correlation aggregation strategy to appropriately assign 3D features to the corresponding 2D features. Extensive experiments on ClearGrasp, OOD, TransCG, and STD datasets show that our method outperforms other state-of-the-art methods. We further demonstrate that our method significantly enhances the performance of downstream robotic grasping tasks. The code will be available at: https://github.com/lyz3356/GAA-TSO. Yizhe Liu, Tong Jia 0001, Jiahui Wei, Hao Wang 0073, Dongyue Chen 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Non-Asymptotic Achievable Rate-Distortion Region for Indirect Wyner-Ziv Source CodingabstractIn the Wyner-Ziv source coding problem, a source X has to be encoded while the decoder has access to side information Y. This paper investigates the indirect setup, in which a latent source S, unobserved by both the encoder and the decoder, must also be reconstructed at the decoder. This scenario is increasingly relevant in the context of goal-oriented communications, where S can represent semantic information obtained from X. This paper derives the indirect Wyner-Ziv rate-distortion function in asymptotic regime and provides an achievable region in finite block-length. Furthermore, a Blahut-Arimoto algorithm tailored for the indirect Wyner-Ziv setup, is proposed. This algorithm is then used to give a numerical evaluation of the achievable indirect rate-distortion region when S is treated as a classification label. Jiahui Wei, Philippe Mary, Elsa Dupraz |
ITW | 1 |
| 2025 | Dynamic window sampling strategy for image captioning
Zhixin Li 0001, Jiahui Wei, Tiantao Xian, Canlong Zhang, Huifang Ma |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | Fusing grid and adaptive region features for image captioning
Jiahui Wei, Zhixin Li 0001, Canlong Zhang, Huifang Ma |
Image Vis. Comput. | 1 |
| 2024 | Practical Coding Schemes based on LDPC Codes for Distributed Parametric RegressionabstractIn the framework of goal-oriented communications, this paper investigates parametric regression over coded data. For this problem, information-theoretic bounds are provided in terms of rate versus regression generalization error, by considering quantize and binning achievability schemes. Alternatively, this paper focuses on practical implementations by proposing a coding scheme that combines a scalar quantizer with a non-binary LDPC code for the binning part. Given that the LDPC decoder requires prior knowledge of the regression parameters, the paper introduces a novel method to estimate these parameters directly over the LDPC-coded syndrome, without the need for prior decoding. This technique allows to both address the regression task and initialize the LDPC decoder for further data reconstruction. Monte-Carlo simulations show the efficiency of the proposed approach in terms of regression generalization error. Jiahui Wei, Elsa Dupraz, Philippe Mary |
ITW | 1 |
| 2024 | Mining core information by evaluating semantic importance for unpaired image captioning
Jiahui Wei, Zhixin Li 0001, Canlong Zhang, Huifang Ma |
Neural Networks | 1 |
| 2023 | Enhance understanding and reasoning ability for image captioning
Jiahui Wei, Zhixin Li 0001, Jianwei Zhu, Huifang Ma |
Appl. Intell. | 1 |
| 2023 | Modeling graph-structured contexts for image captioning
Zhixin Li 0001, Jiahui Wei, Feicheng Huang, Huifang Ma |
Image Vis. Comput. | 2 |
| 2023 | Fuzzy Adaptive Control for Vehicular Platoons With Constraints and Unknown Dead-Zone InputabstractIn this paper, an adaptive fuzzy control problem is studied for a connected automated vehicles platoon subject to unknown dead-zone input and constraints. To better handle the unknown nonlinear dynamical functions and disturbances, the nonlinear dynamics model is transformed to a new model. Then, the fuzzy logic system (FLS) is used to identify the unknown nonlinear functions. A dead-zone inverse technique is introduced to eliminate the negative effects of the unknown dead-zone input nonlinearity. In the framework of backstepping, the tangent barrier Lyapunov function (BLF) is introduced in this paper, and a distributed adaptive fuzzy control scheme is designed so that the position, velocity and acceleration of the vehicle platoon do not violate the given constrained boundaries. Finally, based on the Lyapunov stability theory, it is noted that all signals in the closed-loop system are bounded and the tracking errors converge to a small neighborhood of the origin. The effectiveness of the proposed approach is validated by simulation results. Jiahui Wei, Yan-Jun Liu 0003, Hao Chen 0099, Lei Liu 0006 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2022 | Mixed Knowledge Relation Transformer for Image CaptioningabstractInternal relationship of image objects has contributed significantly to the development of image captioning, especially when combined with Transformer architecture. Most of these methods only calculate the relationship between entities and ignore the information between entities and background. Besides, the way of exploring the relational information inside the image can also be extended. In this paper, we continually explore the relationship between objects from both internal and external perspectives, and embed the vital image global information into the internal relationship module. To validate the effectiveness of our model, we conduct extensive experiments on the most popular MSCOCO dataset, and achieve state-of-the-art performance on both online and offline test sets. Zhixin Li 0001, Jiahui Wei, Tiantao Xian |
ICASSP | 3 |
| 2022 | Image-Text Matching with Fine-Grained Relational Dependency and Bidirectional Attention-Based Generative NetworksabstractGenerally, most existing cross-modal retrieval methods only consider global or local semantic embeddings, lacking fine-grained dependencies between objects. At the same time, it is usually ignored that the mutual transformation between modalities also facilitates the embedding of modalities. Given these problems, we propose a method called BiKA (Bidirectional Knowledge-assisted embedding and Attention-based generation). The model uses a bidirectional graph convolutional neural network to establish dependencies between objects. In addition, it employs a bidirectional attention-based generative network to achieve the mutual transformation between modalities. Specifically, the knowledge graph is used for local matching to constrain the local expression of the modalities, in which the generative network is used for mutual transformation to constrain the global expression of the modalities. In addition, we also propose a new position relation embedding network to embed position relation information between objects. The experiments on two public datasets show that the performance of our method has been dramatically improved compared to many state-of-the-art models. Jianwei Zhu, Zhixin Li 0001, Yufei Zeng, Jiahui Wei, Huifang Ma |
ACM Multimedia | 4 |
| 2022 | Fine-Grained Bidirectional Attention-Based Generative Networks for Image-Text Matching
Zhixin Li 0001, Jianwei Zhu, Jiahui Wei, Yufei Zeng |
ECML/PKDD (3) | 3 |
| 2022 | Flexible Image Captioning via Internal Understanding and External ReasoningabstractImage captioning aims to generate a grammatically correct and semantically accurate natural language description of a given image. In order to capture the more complex information contained in the image and expand the relevant external knowledge outside the image to generate better image caption, this paper proposes an end-to-end image captioning framework Flexible Image Captioning via Internal Understanding and External Reasoning (IUER) based on the Transformer model. IUER enhances visual understanding ability and caption reasoning ability to improve image captioning performance. To achieve this goal, we use the semantic features of the core objects detected from the image to guide the visual feature, where the visual feature incorporate the spatial positional relationship information between the objects, then we introduce external knowledge network to obtain information other than the intuitive content from the image. In this way, a high-quality image caption sentence about the given image is generated. Experiments prove that our method is superior to the baseline model and comparable to other state-of-the-art methods. Jiahui Wei, Zhixin Li 0001, Jianwei Zhu, Huifang Ma |
SDM | 1 |
| 2022 | Fine-grained bidirectional attentional generation and knowledge-assisted networks for cross-modal retrieval
Jianwei Zhu, Zhixin Li 0001, Jiahui Wei, Yufei Zeng, Huifang Ma |
Image Vis. Comput. | 3 |
| 2022 | PBGN: Phased Bidirectional Generation Network in Text-to-Image Synthesis
Jianwei Zhu, Zhixin Li 0001, Jiahui Wei, Huifang Ma |
Neural Process. Lett. | 3 |
| 2018 | Automatic Semantic Content Removal by Learning to Neglect
Siyang Qin, Jiahui Wei, Roberto Manduchi |
BMVC | 2 |
| 2018 | An Overlapping Microblog Community Detection Method Using New Partition Criterion
Huifang Ma, Meng Xie, Jiahui Wei, Tingnian He |
KSEM (2) | 3 |