Authors
Ido Dagan, Kenneth Church, Willian Gale
Publication date
1999
Journal
Natural Language Processing Using Very Large Corpora
Pages
209-224
Publisher
Springer Netherlands
Description
We have developed a new program called word_align for aligning parallel text, text such as the Canadian Hansards that are available in two or more languages. The program takes the output of char_align (Church, 1993), a robust alternative to sentence-based alignment programs, and applies word-level constraints using a version of Brown et al.’s Model 2 (Brown et al., 1993), modified and extended to deal with robustness issues. Word_align was tested on a subset of Canadian Hansards supplied by Simard (Simard et al., 1992). The combination of word_align plus char_align reduces the variance (average square error) by a factor of 5 over char_align alone. More importantly, because word_align and char_align were designed to work robustly on texts that are smaller and more noisy than the Hansards, it has been possible to successfully deploy the programs at AT&T Language Line Services, a …
Total citations
19931994199519961997199819992000200120022003200420052006200720082009201020112012201320142015201620172018201920202021202220232024210111317169231317211813510125121010515332152331
Scholar articles
I Dagan, K Church, W Gale - Natural Language Processing Using Very Large …, 1999