Moby Project
From Wikipedia, the free encyclopedia
The Moby Project is a collection of public domain lexical resources. It was created by Grady Ward. The resources were dedicated to the public domain, and are now mirrored at Project Gutenberg. As of 2007, it contains the largest free phonetic database, with 177,267 words and corresponding pronunciation.
Contents |
[edit] Hyphenator
The Moby Hyphenator II contains 187,175 hyphenated words, with 9,752 indicating that they should not be hyphenated. Hyphenation is indicated by the ASCII 165 character.
[edit] Language
Moby Language II contains wordlists of five languages - French, German, Italian, Japanese, and Spanish:
Language | Words | Size (in bytes) |
---|---|---|
French | 138,257 | 1,524,757 |
German | 159,809 | 2,055,986 |
Italian | 60,453 | 561,981 |
Japanese | 115,523 | 934,783 |
Spanish | 86,059 | 850,523 |
Total | 560,101 | 5,928,030 |
However, it should be noted that some of the lists are contaminated, for example the Japanese list contains English words such as abnormal and non-words such as abcdefgh and m,./.
[edit] Part-of-Speech
Moby Part-of-Speech contains 233,356 words fully described by part(s) of speech, listed in priority order. The format of the file is word\parts-of-speech, with the following parts of speech being identified:
Part-of-speech | Code |
---|---|
Noun | N |
Plural | p |
Noun phrase | h |
Verb (usually participle) | V |
Transitive verb | t |
Intransitive verb | i |
Adjective | A |
Adverb | v |
Conjunction | C |
Preposition | P |
Interjection | ! |
Pronoun | r |
Definite article | D |
Indefinite article | I |
Nominative | o |
[edit] Pronunciator
The Moby Pronunciator II contains 177,267 words with corresponding pronunciation. The Project Gutenberg distribution also contains a copy of the cmudict v0.3. The file follows the format word[/part-of-speech] pronunciation. The part-of-speech field is used to disambiguate 770 of the words which have differing pronunciations depending on their part-of-speech. For example for the words spelled close, the verb has the pronunciation IPA: /ˈkloʊz/, whereas the adjective is /ˈkloʊs/. The parts-of-speech have been assigned the following codes:
Part-of-speech | Code |
---|---|
Noun | n |
Verb | v |
Adjective | aj |
Adverb | av |
Interjection | interj |
Following this is the pronunciation. Several special symbols are present:
Symbol | Meaning |
---|---|
/ | Used to separate phonemes |
_ | Used to separate words |
' | Primary stress on the following syllable |
, | Secondary stress on the following syllable |
The rest of the symbols are used to represent IPA characters, according to the following table:
Symbol | IPA (GA) |
---|---|
& | æ |
(@)r | ɛɹ |
- | ə |
@ | ʌ |
@r | ɚ |
A | ɑ |
aI | aɪ |
Ar | ɑɹ |
AU | aʊ |
b | b |
d | d |
D | ð |
dZ | ʤ |
E | ɛ |
eI | eɪ |
f | f |
g | g |
h | h |
hw | ʍ |
i | i |
I | ɪ |
ir | ɪɹ |
j | j |
k | k |
l | l |
m | m |
n | n |
N | ŋ |
O | ɔ |
Oi | ɔɪ |
oU | oʊ |
oUr | ɔɹ |
p | p |
R | ɻ |
r | ɹ |
S | ʃ |
s | s |
T | θ |
t | t |
tS | ʧ |
U | ʊ |
u | u |
Ur | ʊɹ |
v | v |
w | w |
x | x |
Y | u |
y | y |
Z | ʒ |
z | z |
[@]r | ɝ |
[edit] Shakespeare
Moby Shakespeare contains the complete unabridged works of Shakespeare. This specific resource is not available from Project Gutenberg.
[edit] Thesaurus
The Moby Thesaurus II contains 30,260 root words, with 2,520,264 synonyms and related terms - an average of 83.3 per root word. Each line consists of a list of comma-separated values, with the first term being the root word, and all following words being related terms.
Grady Ward placed this thesaurus in the public domain in 1996. It is also available as a Debian package.
[edit] Words
Moby Words II is the largest wordlist in the world[1]. The distribution consists of the following 16 files:
Filename | Words | Description |
---|---|---|
ACRONYMS.TXT | 6,213 | Common acronyms and abbreviations |
COMMON.TXT | 74,550 | Common words present in two or more published dictionaries |
COMPOUND.TXT | 256,772 | Phrases, proper nouns, and acronyms not included in the common words file |
CROSSWD.TXT | 113,809 | Words included in the first edition of the Official Scrabble Players Dictionary |
CRSWD-D.TXT | 4,160 | Additions to the Official Scrabble Players Dictionary in the second edition |
FICTION.TXT | 467 | A list of the most commonly occurring substrings in the book The Joy Luck Club |
FREQ.TXT | 1,000 | Most frequently occurring words in the English language, listed in descending order |
FREQ-INT.TXT | 1,000 | Most frequently occurring words on Usenet in 1992, listed with corresponding percentage in decreasing order |
KJVFREQ.TXT | 1,185 | Most frequently occurring substrings in the King James Version of the Bible, listed in descending order |
NAMES.TXT | 21,986 | Most common names used in the USA and Great Britain |
NAMES-F.TXT | 4,946 | Common English female names |
NAMES-M.TXT | 3,897 | Common English male names |
OFTENMIS.TXT | 366 | Most common misspelled English words |
PLACES.TXT | 10,196 | Place names in the USA |
SINGLE.TXT | 354,984 | Single words excluding proper nouns, acronyms, compound words and phrases, but including archaic words and significant variant spellings |
USACONST.TXT | 7,618 | United States Constitution including all amendments current to 1993 |
Total | 863,149 |