Jinchao Xu

dblp:13/1124 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
6since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 5 since 2021Security and privacy · 1 · 1 first-authorTheory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Second Language (Arabic) Acquisition of LLMs via Progressive Vocabulary Expansion
abstract
Jianqing Zhu, Huang Huang, Zhihang Lin, Juhao Liang, Zhengyang Tang, Khalid Almubarak, Mosen Alharthi, Bang An, Juncai He, Xiangbo Wu, Fei Yu, Junying Chen, Ma Zhuoheng, Yuhao Du, He Zhang, Saied Alshahrani, Emad A. Alghamdi, Lian Zhang, Ruoyu Sun, Haizhou Li, Benyou Wang, Jinchao Xu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Jianqing Zhu, Zhihang Lin, Juhao Liang, Zhengyang Tang, Khalid Almubarak, Mosen Alharthi, Bang An 0004, Juncai He 0001, Xiangbo Wu, Fei Yu 0017, Zhuoheng Ma, Saied Alshahrani, Emad A. Alghamdi, Ruoyu Sun 0001, Haizhou Li 0001, Benyou Wang, Jinchao Xu
ACL (1)22
2024 MgNO: Efficient Parameterization of Linear Operators via Multigrid
abstract
In this work, we propose a concise neural operator architecture for operator learning. Drawing an analogy with a conventional fully connected neural network, we define the neural operator as follows: the output of the $i$-th neuron in a nonlinear operator layer is defined by $\mathcal O_i(u) = \sigma\left( \sum_j \mathcal W_{ij} u + \mathcal B_{ij}\right)$. Here, $\mathcal W_{ij}$ denotes the bounded linear operator connecting $j$-th input neuron to $i$-th output neuron, and the bias $\mathcal B_{ij}$ takes the form of a function rather than a scalar. Given its new universal approximation property, the efficient parameterization of the bounded linear operators between two neurons (Banach spaces) plays a critical role. As a result, we introduce MgNO, utilizing multigrid structures to parameterize these linear operators between neurons. This approach offers both mathematical rigor and practical expressivity. Additionally, MgNO obviates the need for conventional lifting and projecting operators typically required in previous neural operators. Moreover, it seamlessly accommodates diverse boundary conditions. Our empirical observations reveal that MgNO exhibits superior ease of training compared to CNN-based models, while also displaying a reduced susceptibility to overfitting when contrasted with spectral-type neural operators. We demonstrate the efficiency and accuracy of our method with consistently state-of-the-art performance on different types of partial differential equations (PDEs).
Juncai He 0001, Jinchao Xu
ICLR3
2024 AceGPT, Localizing Large Language Models in Arabic
abstract
Huang Huang, Fei Yu, Jianqing Zhu, Xuening Sun, Hao Cheng, Song Dingjie, Zhihong Chen, Mosen Alharthi, Bang An, Juncai He, Ziche Liu, Junying Chen, Jianquan Li, Benyou Wang, Lian Zhang, Ruoyu Sun, Xiang Wan, Haizhou Li, Jinchao Xu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Fei Yu 0017, Jianqing Zhu, Xuening Sun, Dingjie Song, Mosen Alharthi, Bang An 0004, Juncai He 0001, Ziche Liu, Benyou Wang, Ruoyu Sun 0001, Haizhou Li 0001, Jinchao Xu
NAACL-HLT19
2024 Alignment at Pre-training! Towards Native Alignment for Arabic LLMs
abstract
The alignment of large language models (LLMs) is critical for developing effective and safe language models. Traditional approaches focus on aligning models during the instruction tuning or reinforcement learning stages, referred to in this paper as `\textit{post alignment}'. We argue that alignment during the pre-training phase, which we term 'native alignment', warrants investigation. Native alignment aims to prevent unaligned content from the beginning, rather than relying on post-hoc processing. This approach leverages extensively aligned pre-training data to enhance the effectiveness and usability of pre-trained models. Our study specifically explores the application of native alignment in the context of Arabic LLMs. We conduct comprehensive experiments and ablation studies to evaluate the impact of native alignment on model performance and alignment stability. Additionally, we release open-source Arabic LLMs that demonstrate state-of-the-art performance on various benchmarks, providing significant benefits to the Arabic LLM community.
Juhao Liang, Zhenyang Cai, Jianqing Zhu, Kewei Zong, Bang An 0004, Mosen Alharthi, Juncai He 0001, Haizhou Li 0001, Benyou Wang, Jinchao Xu
NeurIPS12
2023 An interpretive constrained linear model for ResNet and MgNet
Juncai He 0001, Jinchao Xu, Jianqing Zhu
Neural Networks2
2022 Optimal Convergence Rates for the Orthogonal Greedy Algorithm
abstract
We analyze the orthogonal greedy algorithm when applied to dictionaries$\mathbb {D}$whose convex hull has small entropy. We show that if the metric entropy of the convex hull of$\mathbb {D}$decays at a rate of$O\left({n^{-\frac {1}{2}-\alpha }}\right)$for$\alpha > 0$, then the orthogonal greedy algorithm converges at the same rate on the variation space of$\mathbb {D}$. This improves upon the well-known$O\left({n^{-\frac {1}{2}}}\right)$convergence rate of the orthogonal greedy algorithm in many cases, most notably for dictionaries corresponding to shallow neural networks. These results hold under no additional assumptions on the dictionary beyond the decay rate of the entropy of its convex hull. In addition, they are robust to noise in the target function and can be extended to convergence rates on the interpolation spaces of the variation norm. We show empirically that the predicted rates are obtained for the dictionary corresponding to shallow neural networks with Heaviside activation function in two dimensions. Finally, we show that these improved rates are sharp and prove a negative result showing that the iterates generated by the orthogonal greedy algorithm cannot in general be bounded in the variation norm of$\mathbb {D}$.
Jonathan W. Siegel, Jinchao Xu
IEEE Trans. Inf. Theory2
2020 Approximation rates for neural networks with general activation functions
Jonathan W. Siegel, Jinchao Xu
Neural Networks2
2010 A Software Watermarking Algorithm Based on Stack-State Transition Graph
abstract
In the Internet age, Software security and piracy becomes a more and more important issue. In order to prevent software from piracy and unauthorized modification, various techniques have been developed. Among them is software watermarking which protects software through embedding some secret information into software as an identifier of the ownership of copyright for this software. This paper gives an new algorithm based on stack-state transition graph, watermarks is embed by adding additional code in the executable file, and extracted by recognizing the relationship of stack-state which processed in runtime. Analysis proves that our algorithm is more reliable.
Jinchao Xu, Guosun Zeng
NSS1