Xinjie Chen

dblp:157/6842 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
6since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Machine translation · 58% Language models and text generation · 25% Efficient and distributed learning · 16%

Topics — the 4 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Machine translation › simultaneous machine translation
simultaneous speech translation
1.822026
Efficient and Adaptive Simultaneous Speech Translation with Fully Unidirectional Architecture · AAAI 2026
Divergence-Guided Simultaneous Speech Translation · AAAI 2024
Natural language and speech › Machine translation
speech translation
1.822026
Efficient and Adaptive Simultaneous Speech Translation with Fully Unidirectional Architecture · AAAI 2026
Divergence-Guided Simultaneous Speech Translation · AAAI 2024
Natural language and speech › Language models and text generation
large language model reasoning
1.012026
From Data-Centric to Sample-Centric: Enhancing LLM Reasoning via Progressive Optimization · ACL (1) 2026
Natural language and speech › Language models and text generation
large language model
0.312026
Efficient and Adaptive Simultaneous Speech Translation with Fully Unidirectional Architecture · AAAI 2026

Methods — techniques the papers use, named apart from their topics

unidirectional architecture · 1.0progressive optimization · 1.0policy head · 1.0multi-stage training · 1.0data curation · 1.0prefix-based training · 0.8divergence-guided policy · 0.8adaptive read/write policy · 0.8
YearPublicationVenuePosition
2026 Efficient and Adaptive Simultaneous Speech Translation with Fully Unidirectional Architecture
abstract
Simultaneous speech translation (SimulST) produces translations incrementally while processing partial speech input. Although large language models (LLMs) have shown strong capabilities in offline translation tasks, applying them to SimulST poses notable challenges. Existing LLM-based SimulST approaches either incur significant computational overhead due to repeated encoding of bidirectional speech encoder, or they depend on a fixed read/write policy, limiting the efficiency and performance. In this work, we introduce Efficient and Adaptive Simultaneous Speech Translation (EASiST) with fully unidirectional architecture, including both speech encoder and LLM. EASiST includes a multi-latency data curation strategy to generate semantically aligned SimulST training samples and redefines SimulST as an interleaved generation task with explicit read/write tokens. To facilitate adaptive inference, we incorporate a lightweight policy head that dynamically predicts read/write actions. Additionally, we employ a multi-stage training strategy to align speech-text modalities and optimize both translation and policy behavior. Experiments on both in-domain (MuST-C) and out-of-domain (Europarl-ST) En-De and En-Es datasets demonstrate that EASiST offers superior latency-quality trade-offs compared to several strong baselines.
Biao Fu, Donglei Yu, Minpeng Liao, Chengxi Li 0014, Xinjie Chen, Yidong Chen 0001, Kai Fan 0002, Xiaodong Shi
AAAI5
2026 From Data-Centric to Sample-Centric: Enhancing LLM Reasoning via Progressive Optimization
abstract
Xinjie Chen, Minpeng Liao, Guoxin Chen, Chengxi Li, Biao Fu, Kai Fan, Xinggao Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Xinjie Chen, Minpeng Liao, Guoxin Chen, Chengxi Li 0014, Biao Fu, Kai Fan 0002, Xinggao Liu
ACL (1)1
2025 SheepDA-YOLO: Cross-Domain Adaptive Mean Teacher with Dual-Path Decoupling for Sheep Behavior Recognition
abstract
With the rapid advancement of smart farming towards large-scale livestock operations, the demand for model generalization in cross-pen behavior recognition has significantly increased. Traditional deep learning models suffer from substantial performance degradation due to variations in illumination and structure across different sheep pens, often necessitating the re-annotation of tens of thousands of frames for each new environment to mitigate domain shift issues. This severely limits the deployment of models in large-scale sheep farms. To achieve the goal of ’annotate once, generalize across pens,’ we propose the SheepDA-YOLO framework, which innovatively integrates contrastive image translation and feature decoupling to address cross-domain adaptation challenges in agriculture. The core of our method consists of four parts: generating bidirectional pseudo-images for source and target domains based on CUT method to reduce image-level domain discrepancies through mixed training sets; employing a Mean Teacher architecture combined with a quadruple loss function to ensure stable knowledge transfer; proposing DP-DMAF module, which suppresses illumination interference and feature confusion through dual-path feature decoupling and separable large-kernel attention, complemented by a high-resolution detection layer to enhance small-target recognition accuracy. Experimental results demonstrate that SheepDA-YOLO achieves 89.7% mAP in cross-domain testing on target sheep pens, outperforming state-of-the-art methods by 3.4% and significantly reducing annotation costs. The study is the first to validate the feasibility of cross-pen adaptation, providing an efficient solution for the scalable implementation of smart livestock farming.
Xinjie Chen, Yongyuan Qiao
IROS1
2024 Divergence-Guided Simultaneous Speech Translation
abstract
To achieve high-quality translation with low latency, a Simultaneous Speech Translation (SimulST) system relies on a policy module to decide whether to translate immediately or wait for additional streaming input, along with a translation model capable of effectively handling partial speech input. Prior research has tackled these components separately, either using ``wait-k'' policies based on fixed-length segments or detected word boundaries, or dynamic policies based on different strategies (e.g., meaningful units), while employing offline models for prefix-to-prefix translation. In this paper, we propose Divergence-Guided Simultaneous Speech Translation (DiG-SST), a tightly integrated approach focusing on both translation quality and latency for streaming input. Specifically, we introduce a simple yet effective prefix-based strategy for training translation models with partial speech input, and develop an adaptive policy that makes read/write decisions for the translation model based on the expected divergence in translation distributions resulting from future input. Our experiments on multiple translation directions of the MuST-C benchmark demonstrate that our approach achieves a better trade-off between translation quality and latency compared to existing methods.
Xinjie Chen, Kai Fan 0002, Xinggao Liu, Zhongqiang Huang
AAAI1
2024 Integrating regular expressions into neural networks for relation extraction
Zhaoran Liu, Xinjie Chen, Hao Wang 0049, Xinggao Liu
Expert Syst. Appl.2
2024 Hidformer: Hierarchical dual-tower transformer using multi-scale mergence for long-term time series forecasting
Zhaoran Liu, Yizhi Cao, Hu Xu 0007, Yuxin Huang 0007, Qunshan He, Xinjie Chen, Xinggao Liu
Expert Syst. Appl.6