VLDB 2026 Research / reviewers in the wild / expert
Jinchao Xu
dblp:13/1124
· DBLP profile ↗
8ranked-venue papers
1as first author
6since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 5 since 2021Security and privacy · 1 · 1 first-authorTheory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Second Language (Arabic) Acquisition of LLMs via Progressive Vocabulary ExpansionabstractJianqing Zhu, Huang Huang, Zhihang Lin, Juhao Liang, Zhengyang Tang, Khalid Almubarak, Mosen Alharthi, Bang An, Juncai He, Xiangbo Wu, Fei Yu, Junying Chen, Ma Zhuoheng, Yuhao Du, He Zhang, Saied Alshahrani, Emad A. Alghamdi, Lian Zhang, Ruoyu Sun, Haizhou Li, Benyou Wang, Jinchao Xu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Jianqing Zhu, Zhihang Lin, Juhao Liang, Zhengyang Tang, Khalid Almubarak, Mosen Alharthi, Bang An 0004, Juncai He 0001, Xiangbo Wu, Fei Yu 0017, Zhuoheng Ma, Saied Alshahrani, Emad A. Alghamdi, Ruoyu Sun 0001, Haizhou Li 0001, Benyou Wang, Jinchao Xu |
ACL (1) | 22 |
| 2024 | MgNO: Efficient Parameterization of Linear Operators via MultigridabstractIn this work, we propose a concise neural operator architecture for operator learning. Drawing an analogy with a conventional fully connected neural network, we define the neural operator as follows: the output of the $i$-th neuron in a nonlinear operator layer is defined by $\mathcal O_i(u) = \sigma\left( \sum_j \mathcal W_{ij} u + \mathcal B_{ij}\right)$. Here, $\mathcal W_{ij}$ denotes the bounded linear operator connecting $j$-th input neuron to $i$-th output neuron, and the bias $\mathcal B_{ij}$ takes the form of a function rather than a scalar. Given its new universal approximation property, the efficient parameterization of the bounded linear operators between two neurons (Banach spaces) plays a critical role. As a result, we introduce MgNO, utilizing multigrid structures to parameterize these linear operators between neurons. This approach offers both mathematical rigor and practical expressivity. Additionally, MgNO obviates the need for conventional lifting and projecting operators typically required in previous neural operators. Moreover, it seamlessly accommodates diverse boundary conditions. Our empirical observations reveal that MgNO exhibits superior ease of training compared to CNN-based models, while also displaying a reduced susceptibility to overfitting when contrasted with spectral-type neural operators. We demonstrate the efficiency and accuracy of our method with consistently state-of-the-art performance on different types of partial differential equations (PDEs). Juncai He 0001, Jinchao Xu |
ICLR | 3 |
| 2024 | AceGPT, Localizing Large Language Models in ArabicabstractHuang Huang, Fei Yu, Jianqing Zhu, Xuening Sun, Hao Cheng, Song Dingjie, Zhihong Chen, Mosen Alharthi, Bang An, Juncai He, Ziche Liu, Junying Chen, Jianquan Li, Benyou Wang, Lian Zhang, Ruoyu Sun, Xiang Wan, Haizhou Li, Jinchao Xu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Fei Yu 0017, Jianqing Zhu, Xuening Sun, Dingjie Song, Mosen Alharthi, Bang An 0004, Juncai He 0001, Ziche Liu, Benyou Wang, Ruoyu Sun 0001, Haizhou Li 0001, Jinchao Xu |
NAACL-HLT | 19 |
| 2024 | Alignment at Pre-training! Towards Native Alignment for Arabic LLMsabstractThe alignment of large language models (LLMs) is critical for developing effective and safe language models. Traditional approaches focus on aligning models during the instruction tuning or reinforcement learning stages, referred to in this paper as `\textit{post alignment}'. We argue that alignment during the pre-training phase, which we term 'native alignment', warrants investigation. Native alignment aims to prevent unaligned content from the beginning, rather than relying on post-hoc processing. This approach leverages extensively aligned pre-training data to enhance the effectiveness and usability of pre-trained models. Our study specifically explores the application of native alignment in the context of Arabic LLMs. We conduct comprehensive experiments and ablation studies to evaluate the impact of native alignment on model performance and alignment stability. Additionally, we release open-source Arabic LLMs that demonstrate state-of-the-art performance on various benchmarks, providing significant benefits to the Arabic LLM community. Juhao Liang, Zhenyang Cai, Jianqing Zhu, Kewei Zong, Bang An 0004, Mosen Alharthi, Juncai He 0001, Haizhou Li 0001, Benyou Wang, Jinchao Xu |
NeurIPS | 12 |
| 2023 | An interpretive constrained linear model for ResNet and MgNet
Juncai He 0001, Jinchao Xu, Jianqing Zhu |
Neural Networks | 2 |
| 2022 | Optimal Convergence Rates for the Orthogonal Greedy AlgorithmabstractWe analyze the orthogonal greedy algorithm when applied to dictionaries$\mathbb {D}$whose convex hull has small entropy. We show that if the metric entropy of the convex hull of$\mathbb {D}$decays at a rate of$O\left({n^{-\frac {1}{2}-\alpha }}\right)$for$\alpha > 0$, then the orthogonal greedy algorithm converges at the same rate on the variation space of$\mathbb {D}$. This improves upon the well-known$O\left({n^{-\frac {1}{2}}}\right)$convergence rate of the orthogonal greedy algorithm in many cases, most notably for dictionaries corresponding to shallow neural networks. These results hold under no additional assumptions on the dictionary beyond the decay rate of the entropy of its convex hull. In addition, they are robust to noise in the target function and can be extended to convergence rates on the interpolation spaces of the variation norm. We show empirically that the predicted rates are obtained for the dictionary corresponding to shallow neural networks with Heaviside activation function in two dimensions. Finally, we show that these improved rates are sharp and prove a negative result showing that the iterates generated by the orthogonal greedy algorithm cannot in general be bounded in the variation norm of$\mathbb {D}$. Jonathan W. Siegel, Jinchao Xu |
IEEE Trans. Inf. Theory | 2 |
| 2020 | Approximation rates for neural networks with general activation functions
Jonathan W. Siegel, Jinchao Xu |
Neural Networks | 2 |
| 2010 | A Software Watermarking Algorithm Based on Stack-State Transition GraphabstractIn the Internet age, Software security and piracy becomes a more and more important issue. In order to prevent software from piracy and unauthorized modification, various techniques have been developed. Among them is software watermarking which protects software through embedding some secret information into software as an identifier of the ownership of copyright for this software. This paper gives an new algorithm based on stack-state transition graph, watermarks is embed by adding additional code in the executable file, and extracted by recognizing the relationship of stack-state which processed in runtime. Analysis proves that our algorithm is more reliable. Jinchao Xu, Guosun Zeng |
NSS | 1 |