Consonant clusters in Proto-Japanese


In existing reconstructions, Proto-Japanese (PJ)[1] was thought to possess only one kind of consonant clusters, nasal–obstruent combinations (*NC), which developed into prenasalized consonants in all varieties including the well-attested Western Old Japanese (WOJ)[2]. However, a careful observation of problematic correspondences among different Japanese–Ryukyuan varieties suggests otherwise: it is actually able to show there were clusters consisting of a nasal and an approximant, two nasals, or even three consonants (*C₁NC₂) in Proto-Japanese. In this article, we present cases where such clusters should be reconstructed, along with the phonological processes closely related to the history of such clusters.

1 Word-initial clusters

1.1 The case of “cat”

It has been widely assumed that the Japanese word for “cat”, neko, has an onomatopoetic origin (e.g. Martin 1977: 495), i.e. *nia “meow” followed by a diminutive suffix etymologically equivalent to Western Old Japanese  “child”, and comes from a different etymon than the Ryukyuan forms such as Shuri Okinawan /majaː/ and /majuː/[3], from which we can reconstruct Proto-Ryukyuan (PR) *maja and *maju/o.

However, it is also possible that the first part of neko came from *mjə/a < *majə/a, and the fossilized Japanese dialectal form *meko in Horobetsu and Raichishka Ainu mekó “cat” (Hattori 1964: 185) supports this idea. Since there is at least one Ryukyuan form with the same contraction (Tarama Miyako /nika/ “cat” < *ne-ka, appearing in Shimoji & Shimoji 2010: 215), it is unable to say that PR did not undergo the contraction; instead, it seems that the contraction took place in words with four or more syllables in PJ. In this scenario, there was already a diminutive form for “cat” in PJ, which was four syllables long and underwent contraction while the unsuffixed form remained unaffected.

Aside from *meko, there is one more evidence that shows *majV is not a PR innovation. Ainu cápe~cappe “cat” (?< *Car-pe “to scatter-NMLZ”; cf. cátcari “to scatter” < *Car-Car-[4]) seems to be a calque of *majV or vice versa (cf. WOJ mayôp- “to be frayed, to be bewildered” < PJ *maju-ap-). This would be hard to explain if *majV did not exist in PJ.

Since there are some words in Japanese that shows word-initial m~n alternation, the suggestion of first-syllable vowel loss along with the change *mj > n should not be considered too far-fetched. Actually, we can find an additional couple of examples that allows us to establish a change *V > ∅ / #m_C in PJ, which we are going to call the Late Proto-Japanese Syncope (LPJS).

1.2 The Late Proto-Japanese Syncope

Attempts to etymologize Japanese namida “tear” have been not so successful so far. The traditional Altaic etymology is that the na part means “eye” and the mida “water”; while the second half of the argument makes sense Japanese-internally (cf. WOJ mîdu /mʲintu/ “water”), the first half is based entirely on comparative evidence and the Altaic hypothesis is highly controversial at best. Vovin (2010) proposed that the PJ form for “tear” should be reconstructed as *namVta = *na “water” + *mVta “eye”, but for some reasons we disagree with his etymology.

Vovin recovers *na by analyzing the WOJ bound form mîna- /mʲina-/ for “water” as /mʲi-na/ “HON-water” rather than /mʲi-na-/ “water-GEN” (the traditional analysis), and *mVtV by analyzing Early Middle Japanese (EMJ) matuge /matuᵑge/ “eyelashes” as coming from *matu-n-kay “eye-GEN-hair” (Vovin 2009: 18–20). However, we think that we should follow the traditional analysis for /mʲina-/ (§2.2), and also that EMJ matuge and PR *matuge are simply corrupted forms of earlier matukï /matuki/ (< *ma “eye” + *tuk- “to attach” + *-ui, perhaps a fossilized PJ verb ending), a WOJ hapax legomenon whose existence was noticed by Vovin, presumably under the influence of mayuge “eyebrows”.

Our etymology for namida derives it from *mnaminta < *ma-na-minta “eye-GEN-water”, with the change *V > ∅ / #m_C it shares with “cat” (§1.1). In addition to the perfect semantic fit, the reconstruction of *mnaminta is a good choice since it can explain the existence of WOJ namîta /namʲita/ along with expected namîda /namʲinta/: it is already well known that WOJ tends to avoid two /NC/ clusters in a single word (i.e. Lyman’s law).

Another word that might have undergone LPJS is doro “mud”. Since this word appears in Ryukyuan (e.g. Shuri Okinawan /duru/), perhaps we should be able to reconstruct the word at the PJ level; however, words that begin with d < PJ *Nt are rare and require special attention. With LPJS, we can reconstruct post-LPJS PJ *mtɨrɨ~*mtərə, which might be cognate to PJ *mita “dirt”, a root that is already well-established (Martin 1977: 481). Here, note that we use the seven-vowel reconstruction of PJ as proposed by Frellesvig & Whitman (2004) where WOJ ö /ə/ has two PJ sources, *ɨ and *ə, which are different in that *ɨi > WOJ ï /i/ but *əi > WOJ ë /e/. This fact will be important in §2.2.

2 Clusters of three consonants

2.1 Word-final consonants in Proto-Japanese

Traditional explanation for the “intrusive” /s/ in compounds such as WOJ parusamë “spring rain” = paru “spring” + amë “rain” is that either the /s/ is a fossilized genitive marker or amë originally started with *z which was dropped word-initially but preserved word-medially as /s/ (Martin 1977: 36). However, perhaps we should reconstruct PJ *parus “spring”; while the evidence for the genitive marker *s or the PJ consonant *z is highly limited, we have at least some evidence supporting PJ word-final *s, since it is able to analyze WOJ adjectives as having stems ending with *s, and then it would be natural to assume that adjectival nouns such as aka “red” was originally *akas.

In our reconstruction of PJ adjective conjugations, the difference between ku-type and siku-type is due to the third-syllable vowel being lost in the former. The case of uresi- “to be happy”, traditionally explained as ura “inside” + yö- “to be good”, seems to suggest that an adjective’s membership to either type is determined by its stem length (i.e. PJ disyllabic stems becoming ku-type and trisyllabic ones siku-type), and also a tendency of third-syllable vowel loss in PJ.

Indicative (終止形)

Attributive (連体形)

Adverbial (連用形)

Traditional reconstruction

“to be bright/red”

(ku-type)

akasi

< PJ *aka-si

akakî

< PJ *aka-ki

akaku

< PJ *aka-ku

“to be sad”

(siku-type)

kanasi

< PJ *kanasi-si

kanasikî

< PJ *kanasi-ki

kanasiku

< PJ *kanasi-ku

This article

“to be bright/red”

(ku-type)

akasi

< PJ *akas-i

akakî < *akas-k-i

< PJ *akas-ik-i

akaku < *akas-k-u

< PJ *akas-ik-u

“to be sad”

(siku-type)

kanasi

< PJ *kanas-i

kanasikî

< PJ *kanas-ik-i

kanasiku

< PJ *kanas-ik-u

Note that the ending *-i for indicative is also used by ar- “to exist”. Our analysis eliminates the need to introduce three new endings *-si, *-ki, and *-ku; we need only one suffix, *-ik, and having an extra suffix in attributive and adverbial forms (but not in the indicative) finds a parallel in PJ verb conjugations (§2.3).

2.2 Compensatory lengthening in Proto-Japanese

The existence of word-final consonants in PJ implies that compounds such as WOJ akatöki /aka-təkʲi/ “dawn”, assuming it already existed in PJ, should come from earlier *akas-tɨki or *akas-təki, in which the cluster *st occurs. However, in case of akatöki as well as most other compounds, it would be hard to show that such cluster actually existed during a certain period. Fortunately, we observe a regular compensatory lengthening of type VC > Vː / _NC in PJ:

  1. *miC-n-tu > *miːntu “water”
  2. *itɨC-n-pi > *itɨːnpi “strawberry”
  3. *jəC-n-ta > *jəːnta “branch”
  4. *mijɨC-n-si > *mjɨːnsi “rainbow”

The first-syllable vowel of “water” is often reconstructed as *e on the basis that it is never lost in Ryukyuan languages, but such reconstruction is not without any problem: in Okinawan, the vowel corresponding to /i/ in WOJ bound form mîna- “water” is lost (Shuri /nnatu/, Nakijin /naːtˀu/ “port”; cf. WOJ mînatô “id.”), suggesting the reconstruction of PR *i < PJ *i there. Then, it follows that WOJ mîdu and its Ryukyuan cognates come from a etymon different from that of WOJ mîna-, which is quite counterintuitive. Vovin (2010) was aware of this problem and suggested an alternative etymology for mîna- (§1.2); instead, we can posit a distinction based on vowel length: *miC-n-tu > *miːntu but *miC-na- > *mina-. It seems that length was preserved only in the first syllable in PR, since “tear” should have been *mnamiːnta (cf. *miːntu) with a long third-syllable vowel but the vowel was invariably lost in every Ryukyuan variety (e.g. Shuri Okinawan /nada/).

In the case of (2) and (3), the long mid vowel breaks into *ɨi > /i/ or *əi > /e/ in some varieties, while in others (including PR) the length is simply lost to yield *ə > /o/ or PR *o. Therefore, WOJ itibîkô /itinpʲi-ko/ but Shuri Okinawan /ʔitɕubi/ and Shodon Amami /ʔitɕʰup/ (Martin 1970: 120), suggesting PR *itobi; WOJ yeda /jenta/ but Shuri Okinawan /juda/ < PR *joda. The same vowel correspondence can be seen in verbs “to flee” (WOJ nigë- but Shuri Okinawan /nugi-/ < *noge-) and “to let flee”, but there are non-Ryukyuan (i.e. “Mainland Japanese”) varieties with /o/ in the first syllable, suggesting that the breaking of long mid vowels (“strawberry breaking”) was a rather late innovation in pre-WOJ.

In (4), the first-syllable vowel was lost due to LPJS (§1.2), and *mj became /n/ in most dialects; thus we observe Early Middle Japanese /nizi/ (suggesting WOJ *nizi) and Shuri Okinawan /nuːdʑi/ (< PR *nozi). We also have Eastern Old Japanese (EOJ) 努自 ?/nozi/ in Man’yōshū 14.3414, but it is hard to determine whether that /o/ is from pre-EOJ *ɨ or *ɨi since EOJ often drops the *i in *Vi. There is also a modern dialectal form /mjoːzi/ for “rainbow” (Martin 1977: 498–499), which seems to have escaped the effect of LPJS (i.e. *mijɨC-n-si > *mijɨːnsi > *mijozi > /mjoːzi/); however, the first-syllable vowel might have been reinserted by analogy with the unsuffixed form *mijɨC which did not undergo LPJS, which is our reason for positing *i in the first syllable. Figure 259 of Kokuritsu kokugo kenkyūjo (1974) shows that the form /mjoːzi/ and its likes are widespread.

2.3 Identification of more word-final consonants

While the regular compensatory lengthening (“strawberry lengthening”) as described in §2.2 can be a piece of evidence that there was three-consonant clusters of type *C₁NC₂ in PJ, it tells us nothing about what *C₁ was. However, since the first morpheme in (3) also appears in WOJ as standalone ye /je/, it is able to identify *C in *jəC-n-ta as *r, following the argument for apophonic noun kamu-~kamï presented by Whitman (2016: 29–30).

Positing a word-final *r in apophonic nouns would require a sound change *r > *j / _# if one tries to explain the phenomenon in a purely phonological way, but evidence for such sound change is limited. Instead, we can reconstruct *jər-i for ye “branch”, *kamur-i for kamï “deity”, and so on, since the WOJ verb conjugations seem to support *r > ∅ / _i. Note that the following reconstruction of PJ verb conjugations is structurally similar to that of adjective conjugations given in §2.1, strengthening the case for it and, in turn, word-final *s in adjectival stems.

Indicative (終止形)

Attributive (連体形)

Infinitive (連用形)

Quadrigrade

*-u > WOJ -u

*-u > WOJ -u

*-i

Upper monograde

*-u

*-ir-u

*-ir-i

Upper bigrade

*-u > WOJ -u

*-ur-u > WOJ -uru

*-ur-i > WOJ 

Middle bigrade

*-u > WOJ -u

*-ɨr-u >> WOJ -uru

*-ɨr-i > WOJ 

Lower bigrade

*-u > WOJ -u

*-ar-u >> WOJ -uru

*-ar-i > WOJ 

Here, we suppose there were *-Vr suffixes, attached to attributive and infinitive forms of certain verbs, which later became known as monograde and bigrade verbs. WOJ verbs such as tömar- “to stop”, whose stem violates Arisaka’s law, might support the existence of *-Vr suffixes. Double angle brackets denote analogical changes.

Additional explanation would be needed for the quadrigrade and upper monograde conjugations. First, for the quadrigrade, we expect *nar-i > ne for “cry-INF”, but in reality, ne is only attested as a deverbal noun meaning “sound”. Instead, for the verb form, WOJ shows nari, which can be either due to the influence of other forms (i.e. *nar-u > naru “cry-IND” and *nar-u > naru “cry-ATTR”), or to the stem–ending boundary which blocked the application of the rule *r > ∅ / _i. Since y /j/ and w /w/ never occur stem-finally in quadrigrade verbs, it is possible that there was indeed such a boundary lost in monograde and bigrade verbs (presumably due to confusion; i.e. in forms which acquired a *-Vr suffix, should the boundary be placed before the suffix or after it?), but preserved in quadrigrade verbs, and the absence of quadrigrade stems ending in /j/ or /w/ is related to the boundary. For the upper monograde, assuming the PJ stem *mir- for “to see” gives the correct WOJ forms (*mir-u > mîru “see-IND”, *mir-ir-u > mîru “see-ATTR”, *mir-ir-i >  “see-INF”) but it does not work for “to turn” (*mɨr-ir-u > mïru “turn-ATTR”, *mɨr-ir-i >  “turn-INF” but *mɨr-u >> mïru “turn-IND”), and we are forced to posit an analogical change there.

3 Conclusion and implications

In this article, we showed that there existed a wider variety of consonant clusters in PJ than was previously thought, and suggested new etymologies and modified PJ reconstructions for a number of words (“cat”, “tear”, “mud”, “spring rain”, “water”, “strawberry”, “branch”, and “rainbow”). Furthermore, we proposed a new theory concerning the origin of WOJ apophonic nouns and verb and adjective conjugations.

Our reconstruction of PJ implies that long vowels reconstructed for PJ based on Ryukyuan evidence are secondary, and perhaps two PJ mid vowels should be reconstructed, contra Whitman (2016: 23–24): we have both *ɨː > *ɨi > WOJ ï and *əː > *əi > WOJ ë after PJ *j.

        References

Frellesvig, Bjarke & Whitman, John. 2004. The Vowels of Proto-Japanese. Japanese Language and Literature 38(2): 281–299.

Hattori, Shirō (ed.). 1964. Ainu-go hōgen jiten. Tokyo: Iwanami shoten.

Kokuritsu kokugo kenkyūjo (ed.). 1963. Okinawa-go jiten. Tokyo: Ōkurashō insatsukyoku.

Kokuritsu kokugo kenkyūjo (ed.). 1974. Nihon gengo chizu (vol. 6). Tokyo: Ōkurashō insatsukyoku.

Martin, Samuel E. 1970. Shodon: A dialect of Northern Ryukyus. Journal of the American Oriental Society 90(1): 97–139.

Martin, Samuel E. 1977. The Japanese language through time. New Haven and London: Yale University Press.

Nakasone, Seizen. 1983. Okinawa Nakijin hōgen jiten. Tokyo: Kadokawa shoten.

Shimoji, Kayoko & Shimoji, Seijun. 2010. Taramajima hōgen no goi shiryō (1): ‘Taramajima hōgen jiten’ sakusei no tame no. Ryūkyū no hōgen 34: 209–239.

Vovin, Alexander. 1993. A reconstruction of Proto-Ainu. Leiden: Brill.

Vovin, Alexander. 2009. Ryūkyū-go, jōdai nihongo to shūhen no sho-gengo: Saikō to setten no sho-mondai. Nihon kenkyū 39: 11–27.

Vovin, Alexander. 2010. Jōdai nihongo to kodai/chūsei kankokugo no ‘mizu’ to ‘namida’. In Komazawa daigaku gengo kenkyū sentā (ed.) Nikkan gengogakusha kaigi: Kankokugo wo tsūjita nikkan ryōgo no sōgo rikai to kyōsei. Tokyo: Komazawa University, pp. 115–120.

Whitman, John. 2016. Nichiryū sogo no on’in taikei rentaikei/izenkei no kigen. In Takubo, Yukinori, Whitman, John & Tatsuya, Hirako (eds.) Ryūkyū shogo to kodai nihongo: Nichiryū sogo no saiken ni mukete. Tokyo: Kuroshio shuppan, pp. 21–38.

[1] In this article, we will use the term “Proto-Japanese” to refer to the language ancestral to all varieties of Japanese–Ryukyuan, which is natural since we think that Japanese–Ryukyuan is a multifurcate family, “(Mainland) Japanese” is a paraphyletic designation, and there was no Proto-Mainland Japanese.

[2] For the sake of convenience, we will treat WOJ prenasalized obstruents as /NC/ clusters. Regarding the phonemic inventory of Old Japanese, it should also be noted that we will attribute the so-called kō–otsu distinction between îê and ïë to phonemic palatalization of consonants, making WOJ a language with 12 consonants and 6 vowels.

[3] All Okinawan forms cited in this article are from Kokuritsu kokugo kenkyūjo (1963) for the Shuri dialect and Nakasone (1983) for the Nakijin dialect.

[4] Ainu c is secondary, and has various sources including *tj and *tiʔ (Vovin 1993). Here, the symbol *C is used to cover all possible sources of c.