Conghui Zhu

dblp:28/2613 · DBLP profile ↗
← Back
30ranked-venue papers
2as first author
15since 2021 · last 2026
0000-0003-3132-3059ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 28 · 2 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Long-form RewardBench: Evaluating Reward Models for Long-form Generation
abstract
The widespread adoption of reinforcement learning-based alignment highlights the growing importance of reward models. Various benchmarks have been built to evaluate reward models in various domains and scenarios. However, a significant gap remains in assessing reward models for long-form generation, despite its critical role in real-world applications. To bridge this, we introduce Long-form RewardBench, the first reward modeling testbed specifically designed for long-form generation. Our benchmark encompasses five key subtasks: QA, RAG, Chat, Writing, and Reasoning. We collected instruction and preference data through a meticulously designed multi-stage data collection process, and conducted extensive experiments on 20+ mainstream reward models, including both classifiers and generative models. Our findings reveal that current models still lack long-form reward modeling capabilities. Furthermore, we designed a novel Long-form Needle-in-a-Haystack Test, which revealed a correlation between reward modeling performance and the error's position within a response, as well as the overall response length, with distinct characteristics observed between classification and generative models. Finally, we demonstrate that classifier exhibit better generalizability compared to generative models trained on the same data. As the first benchmark for long-form reward modeling, this work aims to offer a robust platform for visualizing progress in this crucial area.
Hui Huang 0021, Yancheng He, Muyun Yang, Kehai Chen, Conghui Zhu, Hailong Cao, Tiejun Zhao
AAAI8
2026 Lost in Benchmarks? Rethinking Large Language Model Benchmarking with Item Response Theory
abstract
The evaluation of large language models (LLMs) via benchmarks is widespread, yet inconsistencies between different leaderboards and poor separability among top models raise concerns about their ability to accurately reflect authentic model capabilities. This paper provides a critical analysis of benchmark effectiveness, examining mainstream prominent LLM benchmarks using results from diverse models. We first propose Pseudo-Siamese Network for Item Response Theory (PSN-IRT), an enhanced Item Response Theory framework that incorporates a rich set of item parameters within an IRT-grounded architecture. PSN-IRT can be utilized for accurate and reliable estimations of item characteristics and model abilities. Based on PSN-IRT, we conduct extensive analysis on 11 LLM benchmarks comprising 41,871 items, revealing significant and varied shortcomings in their measurement quality. Furthermore, we demonstrate that leveraging PSN-IRT is able to construct smaller benchmarks while maintaining stronger alignment with human preference.
Hongli Zhou 0001, Hui Huang 0021, Ziqing Zhao, Lvyuan Han, Huicheng Wang, Kehai Chen, Muyun Yang, Conghui Zhu, Hailong Cao, Tiejun Zhao
AAAI11
2025 MuSC: Improving Complex Instruction Following with Multi-granularity Self-Contrastive Training
abstract
Hui Huang, Jiaheng Liu, Yancheng He, Shilong Li, Bing Xu, Conghui Zhu, Muyun Yang, Tiejun Zhao. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Hui Huang 0021, Yancheng He, Conghui Zhu, Muyun Yang, Tiejun Zhao
ACL (1)6
2025 A Chain-of-Task Framework for Instruction Tuning of LLMs Based on Chinese Grammatical Error Correction
abstract
Over-correction is a critical issue for large language models (LLMs) to address Grammatical Error Correction (GEC) task, esp. for Chinese. This paper proposes a Chain-of-Task (CoTask) framework to reduce over-correction. The CoTask framework is applied as multi-task instruction tuning of LLMs by decomposing the process of grammatical error analysis to design auxiliary tasks and adjusting the types and combinations of training tasks. A supervised fine-tuning (SFT) strategy is also presented to enhance the performance of LLMs, together with an algorithm for automatic dataset annotation to avoid additional manual costs. Experimental results demonstrate that our method achieves new state-of-the-art results on both FCGEC (in-domain) and NaCGEC (out-of-domain) test sets.
Xinpeng Liu 0008, Muyun Yang, Hailong Cao, Conghui Zhu, Tiejun Zhao, Wenpeng Lu
COLING5
2025 LoRA-drop: Efficient LoRA Parameter Pruning based on Output Evaluation
abstract
Low-Rank Adaptation (LoRA) is currently the most commonly used Parameter-efficient fine-tuning (PEFT) method. However, it still faces high computational and storage costs to models with billions of parameters. Most previous studies have tackled this issue by using pruning techniques. Nonetheless, these efforts only analyze LoRA parameter features to evaluate their importance, such as parameter count, size, and gradient. In fact, the output of LoRA directly impacts the fine-tuned model. Preliminary experiments indicate that a fraction of LoRA possesses significantly high output values, substantially influencing the layer output. Motivated by the observation, we propose LoRA-drop. Concretely, LoRA-drop evaluates the importance of LoRA based on the LoRA output. Then we retain LoRA for important layers and the other layers share the same LoRA. We conduct abundant experiments with models of different scales on NLU and NLG tasks. Results demonstrate that LoRA-drop can achieve performance comparable to full fine-tuning and LoRA while retaining 50% of the LoRA parameters on average.
Hongyun Zhou, Conghui Zhu, Tiejun Zhao, Muyun Yang
COLING4
2025 Self-Relevance-Based Multimodal In-Context Learning for Multimodal Named Entity Recognition
abstract
Recently, Multimodal Named Entity Recognition (MNER) has attracted significant attention. Although MNER utilizing in-context learning has shown improved performance, modality retrieval bias often diminishes the relevance of in-context examples. To address this issue, we propose a self-relevance-based multimodal in-context learning method to mitigate modality retrieval bias by dynamically adjusting the weight of each modality. Specifically, we first measure the self-relevance of the query by calculating the similarity between textual and visual modalities, which helps to assess how much visual information contributes to the textual context. Then, we rank the similarity of different modalities, adjust the image rankings based on self-relevance to reduce modality retrieval bias, and integrate them to select the k most relevant examples. Finally, we use task definition and retrieved examples as effective guidance provided to the Multimodal Large Language Models to obtain feedback. Experimental results demonstrate that our method achieves SOTA performance on two benchmark datasets.
Muyun Yang, Hailong Cao, Conghui Zhu, Wenpeng Lu, Tiejun Zhao
ICME5
2025 DuplexMamba: Enhancing Real-Time Speech Conversations with Duplex and Streaming Capabilities
Hongyun Zhou, Conghui Zhu, Tiejun Zhao, Muyun Yang
NLPCC (2)6
2025 Thoughts Behind Attack: Enhancing Security Against Jailbreak Attacks Using Chain-of-Thought
Zhe Tao, Muyun Yang, Hongjiao Guan, Wenpeng Lu, Hailong Cao, Conghui Zhu, Tiejun Zhao
NLPCC (4)7
2024 An efficient confusing choices decoupling framework for multi-choice tasks over texts
Yingyao Wang, Junwei Bao 0001, Chaoqun Duan, Youzheng Wu, Xiaodong He 0001, Conghui Zhu, Tiejun Zhao
Neural Comput. Appl.6
2023 Unsupervised Clustering for Negative Sampling to Optimize Open-Domain Question Answering Retrieval
Feiqing Zhuang, Conghui Zhu, Tiejun Zhao
NLPCC (2)2
2023 Dual Word Embedding for Robust Unsupervised Bilingual Lexicon Induction
abstract
The word embedding models such as Word2vec and FastText simultaneously learn dual representations of input vectors and output vectors. In contrast, almost all existing unsupervised bilingual lexicon induction (UBLI) methods use only input vectors without utilizing output vectors. In this paper, we propose a novel approach to making full use of both input and output vectors for more robust and strong UBLI. We discover the Common Difference Property that one orthogonal transformation can connect not only the input vectors of two languages but also the output vectors. Therefore, we can learn just one transformation to induce two different dictionaries from the input and output vectors, respectively. Between these two quite different dictionaries, a more accurate lexicon with less noise can be induced by taking the intersection of them in UBLI procedure. Extensive experiments show that our method achieves much more robust and strong results than state-of-the-art methods in distant language pairs, while reserving comparable performances in similar language pairs.
Hailong Cao, Liguo Li, Conghui Zhu, Muyun Yang, Tiejun Zhao
IEEE ACM Trans. Audio Speech Lang. Process.3
2022 Toward automatic support for leading court debates: a novel task proposal & effective approach of judicial question generation
Changzhen Ji, Xiaozhong Liu 0001, Adam Jatowt, Sourav S. Bhowmick, Changlong Sun, Conghui Zhu, Tiejun Zhao
Neural Comput. Appl.7
2021 A Neural Conversation Generation Model via Equivalent Shared Memory Investigation
abstract
Conversation generation as a challenging task in Natural Language Generation (NLG) has been increasingly attracting attention over the last years. A number of recent works adopted sequence-to-sequence structures along with external knowledge, which successfully enhanced the quality of generated conversations. Nevertheless, few works utilized the knowledge extracted from similar conversations for utterance generation. Taking conversations in customer service and court debate domains as examples, it is evident that essential entities/phrases, as well as their associated logic and inter-relationships, can be extracted and borrowed from similar conversation instances. Such information could provide useful signals for improving conversation generation. In this paper, we propose a novel reading and memory framework called Deep Reading Memory Network (DRMN) which is capable of remembering useful information of similar conversations for improving utterance generation. We apply our model to two large-scale conversation datasets of justice and e-commerce fields. Experiments prove that the proposed model outperforms the state-of-the-art approaches.
Changzhen Ji, Xiaozhong Liu 0001, Adam Jatowt, Changlong Sun, Conghui Zhu, Tiejun Zhao
CIKM6
2021 Modeling Future Cost for Neural Machine Translation
abstract
Existing neural machine translation (NMT) systems utilize sequence-to-sequence neural networks to generate target translation word by word, and then make the generated word at each time-step and the counterpart in the references as consistent as possible. However, the trained translation model tends to focus on ensuring the accuracy of the generated target word at the current time-step and does not consider its future cost which means the expected cost of generating the subsequent target translation (i.e., the next target word). To respond to this issue, in this article, we propose a simple and effective method to model the future cost of each target word for NMT systems. In detail, a future cost representation is learned based on the current generated target word and its contextual information to compute an additional loss to guide the training of the NMT model. Furthermore, the learned future cost representation at the current time-step is used to help the generation of the next target word in the decoding. Experimental results on three widely-used translation datasets, including the WMT14 English-to-German, WMT14 English-to-French, and WMT17 Chinese-to-English, show that the proposed approach achieves significant improvements over strong Transformer-based NMT baseline.
Chaoqun Duan, Kehai Chen, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita, Conghui Zhu, Tiejun Zhao
IEEE ACM Trans. Audio Speech Lang. Process.6
2021 Detecting Source Contextual Barriers for Understanding Neural Machine Translation
abstract
In machine translation evaluation, the traditional wisdom measures model's generalization ability in an average sense, for example by using corpus BLEU. However, the statistics of corpus BLEU cannot provide comprehensive understanding and fine-grained analysis on model's generalization ability. As a remedy, this paper attempts to understand NMT at fine-grained level, by detecting contextual barriers within an unseen input sentence that \textit{cause} the degradation in model's translation quality. It proposes a principled definition of source contextual barriers as well as its modified version which is tractable in computation and operates at word-level. Based on the modified one, three simple methods are proposed for barrier detection by search-aware risk estimation through counterfactual generation. Extensive analyses are conducted on those detected contextual barrier words on both Zh$\Leftrightarrow$En NIST benchmarks. Potential usages motivated from barrier words are also discussed.
Lemao Liu, Conghui Zhu, Rui Wang 0015, Tiejun Zhao, Shuming Shi 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2020 Demographics Should Not Be the Reason of Toxicity: Mitigating Discrimination in Text Classifications with Instance Weighting
abstract
With the recent proliferation of the use of text classifications, researchers have found that there are certain unintended biases in text classification datasets.For example, texts containing some demographic identity-terms (e.g., "gay", "black") are more likely to be abusive in existing abusive language detection datasets.As a result, models trained with these datasets may consider sentences like "She makes me happy to be gay" as abusive simply because of the word "gay."In this paper, we formalize the unintended biases in text classification datasets as a kind of selection bias from the non-discrimination distribution to the discrimination distribution.Based on this formalization, we further propose a model-agnostic debiasing training framework by recovering the non-discrimination distribution using instance weighting, which does not require any extra resources or annotations apart from a pre-defined set of demographic identity-terms.Experiments demonstrate that our method can effectively alleviate the impacts of the unintended biases without significantly hurting models' generalization ability.
Conghui Zhu, Tiejun Zhao
ACL5
2020 Multimodal Matching Transformer for Live Commenting
abstract
Automatic live commenting aims to provide real-time comments on videos for viewers. It encourages users engagement on online video sites, and is also a good benchmark for video-to-text generation. Recent work on this task adopts encoder-decoder models to generate comments. However, these methods do not model the interaction between videos and comments explicitly, so they tend to generate popular comments that are often irrelevant to the videos. In this work, we aim to improve the relevance between live comments and videos by modeling the cross-modal interactions among different modalities. To this end, we propose a multimodal matching transformer to capture the relationships among comments, vision, and audio. The proposed model is based on the transformer framework and can iteratively learn the attention-aware representations for each modality. We evaluate the model on a publicly available live commenting dataset. Experiments show that the multimodal matching transformer model outperforms the state-of-the-art methods.
Chaoqun Duan, Lei Cui 0001, Shuming Ma, Furu Wei, Conghui Zhu, Tiejun Zhao
ECAI5
2020 Cross Copy Network for Dialogue Generation
abstract
07 2 8 22 0 2 02 !7" 2 #$% &' () (* + ' + ,+ -./0-12(.3 .456#$%&' (672' ($ 83 ' &$&$9% .,:6#$(4;2.,672' ($ ) (<' $($=(' >-% * ' + 5?3 ..@' (4+ .(6=A8BCDEFEGHIJGKIILMBINOPQREBMCSORTURTUMCVGWHTKEXTXTYEUBMBIN KEJZ[\HEU]ETUTMQ]JO BFTU^KIU^M_BKHGTIXTIMBINOPBIU^FJEOGDCFTIWHFEGMQ]JMBU `a3 8 1b8 ) (+ 2-:$* +/ -c5-$% * 6$,<' -(1-*/ % .@<'/ / -% -(+ d-3 <*c' + (-* *+ 2-$12' ->-@-(+ *./* -e,-(1-f + .f* -e,-(1-@.<-3*g -h 4h 6iA0jk$+ + -(+ ' .(6l.' (+ -%9-(-% $+ .%m-+c.% n*$(<0% $(* / .%@-% o + .-(2$(1-<'$3 .4,-1.(+ -(+4-(-% $+ ' .(hp2' 3 -1.(+-(+q,-(15$(<$11,% $15./ + -(* -% >-$*+ 2-@$r .%'(<' 1$+ .%*/ .%@.<-3+ % $' (' (46<' $3 .4,-3.4'1* 61$% % 5' (41% ' + ' 1$3' (/ .%@$+ ' .(/.%*.@-:$%+ ' 1,3 $%<.@$' (* 6$% -./ + -(' 4(.% - ' 1-$(<1.,%+<-&$+ -<' $3 .4,-$*-s$@:3 -* 61.@:$+ ' &3 -3 .4'1*1$(&-.&*-% >-< $1% .**<' / / -% -(+<' $3 .4,-'(* + $(1-* 6$(<+ 2' *' (f / .%@$+ ' .(1$(:%.>'<->' + $3->' <-(1-/ .%,++ -% f $(1-4-(-% $+ ' .(h)(+ 2' *:$:-% 6c-:% .:.* -$ (.>-3(-+ c.% n$% 12' + -1+ ,% -f7% .**7.:5m-+ f c.% n*g 006o+ .-s:3.%-+ 2-1,% % -(+<' $3 .41.(+ -s+$(<* ' @' 3 $%<' $3 .4,-'(* + $(1-* t3 .4'1$3 * + % ,1+ ,% -* ' @,3 + $(-.,* 3 5h us:-% ' @-(+ *c' + 2 + c.+ $* n* 61.,% +<-&$+ -$(<1,* + .@-%*-% >' 1-1.(+ -(+4-(-% $+ ' .(6:%.>-<+2$++ 2-:% .:.* -< $3 4.% ' + 2@ ' ** ,:-% ' .%+.-s'* + ' (4* + $+ -f ./f $% + 1.(+ -(+4-(-% $+ ' .(@.<-3 * h v w 8 12 b8 2 8*$(' @:.% + $(++ $* n' (m$+ ,% $3i$(4,$4-9-(-% $f + ' .(x6y 6<' $3 .4,-4-(-%$+ ' .(-@:.c-% *$c' <-* :-1+ % ,@./$::3 ' 1$+ ' .(*6* ,12$*12$+ &.+$(<1,* f + .@-%*-% >' 1-$,+ .@$+' .(h)(+ 2-:$* +/ -c5-$% * 6 &% -$n+ 2% .,42*'(<' $3 .4,-4-(-%$+ ' .(+-12(.3 .45/ .1,*-<.($* -% ' -*./* -e,-(1-f + .f* -e,-(1-@. -%-+$3 h 6z{|}o hj.% -% -1-(+ 3 56-sf + -% ($3n(.c3-<4-' *-@:3 .5-<+.-(2$(1-@.<-3:-% / .%@$(1-hp,-+$3 hg z{|~o i' ,-+$3 hg z{|o 1$($* * ' * +<' $3 .4,-4-(-%$+ ' .(&5,*' (4n(.c3-<4-+ % ' :3 -* hA' @' 3 $% 3 56i'-+$3 hg z{|~o $r :,% n$%-+$3 h g z{|o #,$(4-+$3 hg z{|o -<<5-+$3 hg z{|~o -s:3 .%-<.1,@-(+$*n(.c3-<4-<' * 1.>-% 5/ .%<'f $3 .4,-4-(-%$+ ' .(6$(<'$-+$3 hg z{|o --+$3 h g z{z{o 92$;>' (' (-r $<-+$3 hg z{|o l$% + 2$* $% $+ 2' $( -% 6,($/ / .%<$&3 -n(.c3 -<4-1.(*+ % ,1f + ' .($(<<-/-1+ ' >-<.@$' ($<$:+ $+ ' .(%-* + % ' 1++ 2-' % ,+ ' 3 ' ;$+ ' .(h7.:5f &$* -<4-(-% $+ ' .(@.<-3 *g ' (5$3 *-+$3 h 6 z{|9,-+$3 h 6z{|o2$>-&--(c' <-3 5$<.:+ -< ' (1.(+ -(+4-(-% $+ ' .(+$* n*$(<* 2.c&-+ + -%% -* ,3 + * 1.@:$% -<+ .*-e,-(1-f + .f* -e,-(1-@.<-3*c2-( / $1- .1$&,3 $% 5:% .&3-@h02$(n*+ .+ 2-' %($+ ,% -./3 ->-% $4' (4>.1$&,3 $% 5$(<1.(+-s+ <' * + % ' &,+ ' .(*/.%1.(+ -(+1.:56'+-($&3 -*+ .1.:5e,' % ' -*/ % .@+2-1,* + .@-%*c' 3 34-+* ' @' 3 $%% -* :.(* -*/ % .@+2-* + $/ / h ) +@.+ ' >$+ -*,*+ .&,' 3 <$@.<-3+2$+1$((.+.(3 5 1.:5+ 2-1.(+ -(+c' + 2' (+ 2-,::-%1.(+-s+./+ 2-+ $% 4-+<' $3 .4,-'(* + $(1-6&,+$3 * .3-$% (+ 2-* ' @' 3 $% :$+ + -% (*$1% .**<' / / -% -(+* ' @' 3 $%1$* -*./+ 2-+ $% 4-+ ' (*+ $(1-hA,12-s+ -% ($31.:51$(&-1%' + ' 1$3' (*.@-* 1-($% ' .*h 8** 2.c(' (' 4,% -h |6c-:% .:.* -+ c.<' / / -% f -(+n' (<*./1.:5@-12$(' * @*' (+ 2' ** + ,<571 8 bb2451.(+-s+ f <-:-(<-(+' (/ .%@$+ ' .(c'+ 2' ( + 2-+ $% 4-+<' $3 .4,-'(* + $(1-6$(<21 28 b245 3 .4'1f <-:-(<-(+1.(+-(+$1% .**<' / / -% -(+t A' @' 3 $% 7$* -* tx 0y h02' */ % $@-c.%n' *3 $&-3 -<$*7% .** f 7.:5m-+ c.% n*x 006y h8*-s-@:3 $%<' $3 .4,-<-:'1+ -<6r ,<4-*@$5% -:-$+g 2.% ' ;.(+ $31.:' - $3' <$+ -+ 2-:% .:.* -<@.<-3 6c--@f :3 .5+c.<' / / -% -(+<' $3 .4,-<$+$* -+ */ % .@+c..% f + 2.4.($3<.@$'(*f $(<
Changzhen Ji, Xiaozhong Liu 0001, Changlong Sun, Conghui Zhu, Tiejun Zhao
EMNLP (1)6
2019 Selection Bias Explorations and Debias Methods for Natural Language Sentence Matching Datasets
abstract
Natural Language Sentence Matching (NLSM) has gained substantial attention from both academics and the industry, and rich public datasets contribute a lot to this process.However, biased datasets can also hurt the generalization performance of trained models and give untrustworthy evaluation results.For many NLSM datasets, the providers select some pairs of sentences into the datasets, and this sampling procedure can easily bring unintended pattern, i.e., selection bias.One example is the QuoraQP dataset, where some content-independent naïve features are unreasonably predictive.Such features are the reflection of the selection bias and termed as the "leakage features."In this paper, we investigate the problem of selection bias on six NLSM datasets and find that four out of them are significantly biased.We further propose a training and evaluation framework to alleviate the bias.Experimental results on QuoraQP suggest that the proposed framework can improve the generalization ability of trained models, and give more trustworthy evaluation results for real-world adoptions.
Jian Liang 0002, Shiyu Chang, Mo Yu, Conghui Zhu, Tiejun Zhao
ACL (1)7
2019 Understanding Data Augmentation in Neural Machine Translation: Two Perspectives towards Generalization
abstract
Guanlin Li, Lemao Liu, Guoping Huang, Conghui Zhu, Tiejun Zhao. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Lemao Liu, Guoping Huang, Conghui Zhu, Tiejun Zhao
EMNLP/IJCNLP (1)4
2018 Attention-Fused Deep Matching Network for Natural Language Inference
abstract
Natural language inference aims to predict whether a premise sentence can infer another hypothesis sentence. Recent progress on this task only relies on a shallow interaction between sentence pairs, which is insufficient for modeling complex relations. In this paper, we present an attention-fused deep matching network (AF-DMN) for natural language inference. Unlike existing models, AF-DMN takes two sentences as input and iteratively learns the attention-aware representations for each side by multi-level interactions. Moreover, we add a self-attention mechanism to fully exploit local context information within each sentence. Experiment results show that AF-DMN achieves state-of-the-art performance and outperforms strong baselines on Stanford natural language inference (SNLI), multi-genre natural language inference (MultiNLI), and Quora duplicate questions datasets.
Chaoqun Duan, Lei Cui 0001, Xinchi Chen, Furu Wei, Conghui Zhu, Tiejun Zhao
IJCAI5
2014 Improving Pivot-Based Statistical Machine Translation by Pivoting the Co-occurrence Count of Phrase Pairs
abstract
To overcome the scarceness of bilingual corpora for some language pairs in machine translation, pivot-based SMT uses pivot language as a "bridge" to generate source-target translation from sourcepivot and pivot-target translation.One of the key issues is to estimate the probabilities for the generated phrase pairs.In this paper, we present a novel approach to calculate the translation probability by pivoting the co-occurrence count of source-pivot and pivot-target phrase pairs.Experimental results on Europarl data and web data show that our method leads to significant improvements over the baseline systems.
Zhongjun He, Hua Wu 0003, Conghui Zhu, Haifeng Wang 0001, Tiejun Zhao
EMNLP4
2014 Discriminative Training for Log-Linear Based SMT: Global or Local Methods
abstract
In statistical machine translation, the standard methods such as MERT tune a single weight with regard to a given development data. However, these methods suffer from two problems due to the diversity and uneven distribution of source sentences. First, their performance is highly dependent on the choice of a development set, which may lead to an unstable performance for testing. Second, the sentence level translation quality is not assured since tuning is performed on the document level rather than on sentence level. In contrast with the standard global training in which a single weight is learned, we propose novel local training methods to address these two problems. We perform training and testing in one step by locally learning the sentence-wise weight for each input sentence. Since the time of each tuning step is unnegligible and learning sentence-wise weights for the entire test set means many passes of tuning, it is a great challenge for the efficiency of local training. We propose an efficient two-phase method to put the local training into practice by employing the ultraconservative update. On NIST Chinese-to-English translation tasks with both medium and large scales of training data, our local training methods significantly outperform standard methods with the maximal improvements up to 2.0 BLEU points, meanwhile their efficiency is comparable to that of the standard methods.
Lemao Liu, Tiejun Zhao, Taro Watanabe, Hailong Cao, Conghui Zhu
ACM Trans. Asian Lang. Inf. Process.5
2013 Hierarchical Phrase Table Combination for Machine Translation
Conghui Zhu, Taro Watanabe, Eiichiro Sumita, Tiejun Zhao
ACL (1)1
2013 Improving Pivot-Based Statistical Machine Translation Using Random Walk
abstract
This paper proposes a novel approach that utilizes a machine learning method to improve pivot-based statistical machine translation (SMT).For language pairs with few bilingual data, a possible solution in pivot-based SMT using another language as a "bridge" to generate source-target translation.However, one of the weaknesses is that some useful sourcetarget translations cannot be generated if the corresponding source phrase and target phrase connect to different pivot phrases.To alleviate the problem, we utilize Markov random walks to connect possible translation phrases between source and target language.Experimental results on European Parliament data, spoken language data and web data show that our method leads to significant improvements on all the tasks over the baseline system.
Zhongjun He, Hua Wu 0003, Haifeng Wang 0001, Conghui Zhu, Tiejun Zhao
EMNLP5
2013 Phrase Table Combination Deficiency Analyses in Pivot-Based SMT
Yiming Cui 0001, Conghui Zhu, Tiejun Zhao, Dequan Zheng
NLDB2
2013 Image Classification Based on the Combination of Text Features and Visual Features
abstract
With more and more text-image co-occurrence data becoming available on the Web, we are interested in how text especially Chinese context around images can aid image classification. The goal is to construct a classification system for images, and we used the context of the images to improve the classification system. First, we extracted three kinds of features, including global visual features, local visual features, and text features using both the image content and context. Then, we tried various feature combination methods and train classifiers for each kind of feature vector. Finally, we used a classifier fusion strategy based on weight learning, combining classifier outputs together, and we obtained the category of unlabeled images. In our experiments on the data set extracted from Google Image Search, we demonstrated the benefit of using context to help image classification. By comparing different feature combination methods on our feature set, we adopted the most effective one. Meanwhile, the classifier fusion approach improves the classification accuracy.
Lexiao Tian, Dequan Zheng, Conghui Zhu
Int. J. Intell. Syst.3
2012 Locally Training the Log-Linear Model for SMT
Lemao Liu, Hailong Cao, Taro Watanabe, Tiejun Zhao, Mo Yu, Conghui Zhu
EMNLP-CoNLL6
2010 Chinese Named Entity Recognition with a Sequence Labeling Approach: Based on Characters, or Based on Words?
Zhangxun Liu, Conghui Zhu, Tiejun Zhao
ICIC (2)2
2007 A Unified Tagging Approach to Text Normalization
Conghui Zhu, Jie Tang 0001, Hang Li 0001, Hwee Tou Ng, Tiejun Zhao
ACL1