Automatic Bilingual Lexical Knowledge Acquisition from a Parallel Chinese-English Corpus
Date Issued
2000-09-30
Date
2000-09-30
Author(s)
DOI
892411H002043
Abstract
In this research, we combine
statistics-based and dictionary-based
algorithms in deriving and augmenting a
translation lexicon at word and phrase
level. We first reimplement Fung and
Church's (1994) K-vec algorithm. K-vec
is a simple algorithm to find word
correspondences from parallel texts. The
basic idea is that a true word pair should
have similar distributions in terms of the
position of its occurrence in the text. To
estimate the similarity of co-occurrence,
the parallel texts are split into the same
number of segments (K) and the
distributions of each word are
represented in a 1… K binary vector.
Two statistical measures, mutual
information (MI) and t-score, are then
employed to calculate the similarity of
distributions of any Chinese-English
word pair. Although K-vec is a quick
and easy method to derive an initial
translation lexicon, its precision is
subject to the constraint of genres,
frequency, and several other factors. We
test the K-vec algorithm against the
Sinorama Chinese-English corpus and
find that the precision ranges from 30%
to 70%. We also experiment on
augmenting translation lexicon based on
existing machine-readable dictionaries.
The result also turns out to be very poor,
the main reason being that translations
are seldom word for word. Furthermore,
exact string matching between
dictionary listing and words in the
parallel corpus can only identify a small
portion of correct word correspondences,
whereas partial matching of dictionary
lookup can find many word
correspondences, most of which are
incorrect. We propose two methods to
address these problems. The first
method, based on proximity of word
pairs and the similarities of word order
between Chinese and English, can
extract phrasal translations and reliable
word correspondences simultaneously
from an unaligned parallel corpus. The
second method utilizes the cooccurrence
information of word pairs in
different documents to filter out spurious
word correspondences. This method can
achieve high-precision word and phrasal
correspondences which can then be used
to derive sentence correspondences. We
process the aligned Chinese-English sentences with Chinese and English
part-of-speech taggers. We also use
Abney’s parser to derive the syntactic
structures of English. After identifying
the subject noun phrase and verb phrase
of each aligned Chinese-English
sentence, we derive phrasal
correspondences based on the syntactic
structures and clues of dictionary lookup.
The basic assumption is that the
structures of aligned sentences are very
likely to correspond to each other if it is
further supported by evidence of partial
match with dictionary lookup. This
hybrid approach can construct a highprecision
translation lexicon at word and
phrase level which can be fruitfully
applied to computational lexicography,
machine translation, and cross-lingual
information retrieval.
Subjects
automatic bilingual lexical knowledge acquisition
parallel Chinese-
English corpus
English corpus
automatic extraction of
translation equivalents
translation equivalents
Publisher
臺北市:國立臺灣大學外國語文學系暨研究所
Type
report
File(s)![Thumbnail Image]()
Loading...
Name
892411H002043.pdf
Size
53.14 KB
Format
Adobe PDF
Checksum
(MD5):7fad4ca0a9f4f3605add082030e9afef
