The Persian Language: A History of Old, Middle, and New Persian

A philological problem opens this article. The Behistun inscription of Darius I, carved about 100 metres up a cliff face in western Iran, was written in three cuneiform scripts: Old Persian, Elamite, and Babylonian. The Old Persian version, which is the shortest, was the key to the entire cuneiform tradition. Henry Rawlinson’s copies of the 1830s and 1840s, made by dangling on ropes from the cliff above, gave scholars their first sustained access to a working Old Persian text, and the comparative method applied to the three versions gave us the cuneiform syllabaries and the bulk of the Mesopotamian textual tradition.

What is striking, on rereading the Old Persian text, is how narrow the language is. The Old Persian cuneiform corpus, which is the entire surviving body of the language, runs to about 1,200 distinct words. The grammar is deliberately simplified — a true alphabet for the vowel /a/, a small set of syllabograms, and a chancery style that keeps the morphology to a minimum. The Old Persian of the inscriptions is an artificial language: a chancery dialect, designed for monumental inscription, that flattens the dialectal variation of the spoken language into a single written form. The spoken language of the Achaemenid court was almost certainly broader than the written form suggests, and the relationship between the written Old Persian and the spoken Median and the other Iranian dialects is one of the harder problems in the field.

I want to write about the long arc. It runs from the Old Persian of the Achaemenid chancery (c. 550–330 BCE) through the Middle Persian Pahlavi of the Sasanian state (c. 224–651 CE) to the New Persian of the Islamic and modern periods (c. 950 CE – present), and the technical core of the article is the question of how the language changed across those three stages. The standard works are Roland Kent, Old Persian: Grammar, Texts, Lexicon (1950, revised 1989) for Old Persian; D. N. MacKenzie, A Concise Pahlavi Dictionary (1971) for Middle Persian; and Gilbert Lazard, La langue des plus anciens monuments de la prose persane (1963) and the later Lazard survey for early New Persian. The philological work of Prods Oktor Skjærvø on the Avestan and Old Persian boundary is the most important recent contribution to the older end. The wider context is in the Persian Culture and Society pillar.

The Indo-Iranian context

Persian is a member of the Indo-Iranian branch of the Indo-European family. Its closest relatives are Sanskrit, Avestan (the language of the Zoroastrian scripture), Kurdish, Balochi, and Pashto. The Indo-Iranian languages share a number of distinctive features that distinguish them from the other Indo-European branches: the merger of the original voiced and voiceless stop series into a single voiced series, the reduction of the complicated Indo-European verb system to a simpler two-tense system, and the preservation of a rich system of nominal case and verbal mood. The split between Iranian and Indo-Aryan, the two main branches of Indo-Iranian, is conventionally dated to around 2000–1500 BCE, on the evidence of the distinct phonological and morphological developments that differentiate the two branches.

The earliest Iranian texts are the Old Avestan hymns of the Zoroastrian scripture, conventionally dated to the second millennium BCE and preserved in the Yasna, the Yashts, and the Visperad. Old Avestan is related to the language of the Rig Veda, though the chronological relationship is disputed: Michael Witzel and others have argued that the Vedic material postdates the Gathas, which would invert the older assumption that the Gathas are simply “older sister” of the Vedas. I find the chronological relationship genuinely unresolved; the Gathas could be second millennium or early first millennium, and the older consensus of a high second-millennium date is being slowly eroded by the recent work.

Old Persian (c. 550–330 BCE)

Old Persian is the earliest attested stage of the Persian language proper. The corpus of Old Persian inscriptions, beginning with the inscription of Cyrus the Great at Pasargadae (c. 540 BCE) and culminating in the long inscriptions of Darius I at Behistun (c. 520 BCE), Naqsh-e Rostam, and Persepolis, is one of the most important documentary sources for the Achaemenid period. The Old Persian cuneiform, designed in the sixth century BCE, was a remarkable adaptation of the Mesopotamian cuneiform to the phonology of an Indo-European language — a mix of logograms, syllabograms, and (uniquely among cuneiform scripts) a true alphabet for the vowel /a/.

The script was used for only about two hundred years, after which the Achaemenid chancery switched to Aramaic for most administrative purposes. The Aramaic script, derived from the Phoenician alphabet, was the working language of the imperial administration, and the Achaemenid chancery used it for the bulk of the surviving administrative documents. The Old Persian inscriptions, however, were preserved in the cuneiform tradition and continued to be copied and studied long after the fall of the Achaemenid empire.

The Old Persian language itself is preserved in a relatively small corpus of about 1,200 distinct words in the royal inscriptions, but the related Median dialect — the language of the Medes and the western Iranian peoples — is attested only in a handful of loanwords in Old Persian, in the names of persons and places, and in a small set of Aramaic transcriptions. The relationship between Old Persian and Median is comparable to that between classical Latin and Oscan or Umbrian, with Old Persian serving as the prestige language of the Achaemenid court and Median as the spoken language of the western Iranian heartland.

Middle Persian or Pahlavi (c. 224–651 CE)

After the fall of the Achaemenid empire and the long Hellenistic and Parthian periods, the Iranian languages entered a new stage. The Arsacid Parthian state, which ruled Iran from c. 247 BCE to 224 CE, used a north-eastern Iranian dialect as its chancery language, and the Sasanian Empire, which succeeded the Parthians, used a south-western Iranian dialect conventionally called Middle Persian or Pahlavi. The Arsacid state also used Aramaic, Greek, and other languages for its inscriptions and coin legends, and the Parthian Empire was, in linguistic terms, more multilingual than the Achaemenid.

The Sasanian state standardised Middle Persian as the official language of the chancery, the law, the Zoroastrian church, and the literary tradition, and the Middle Persian corpus — preserved in the inscriptions, the religious texts, the historical chronicles, and the legal documents of the period — is the most extensive in any pre-Islamic Iranian language. The Sasanian Middle Persian inscriptions, such as the great inscription of Shapur I at the Ka’ba-ye Zartosht, are written in a form of the script that combines the Aramaic-derived Pahlavi writing system with a number of older orthographic conventions. The Sasanian use of Middle Persian was a marker of its distinctive Iranian identity, and the linguistic policy of the Sasanians was a major factor in the survival of Persian identity through the long Hellenistic and Parthian periods.

The Pahlavi writing system was a complex mixture of inherited Aramaic words and signs and Middle Persian pronunciation. Many Pahlavi words were written in their Aramaic form, with the Middle Persian reading supplied by the reader — a convention that has made the decipherment of Pahlavi texts one of the most demanding problems in Iranian philology. The major Pahlavi texts — the Bundahishn, the Denkard, the Arda Viraf Namag, the Zand-i Wahman Yasn, the Khwaday-namag — were the principal vehicles of the Sasanian religious, philosophical, and historical tradition, and their translation into the modern languages is the foundation of the modern study of pre-Islamic Iranian religion and culture. The great Pahlavi scholar of the twentieth century was H. S. Nyberg, whose Hilfsbuch des Pehlevi (1928–1931) opened the field; the work of MacKenzie and his successors built on Nyberg’s foundations.

The Arabic-script revolution

The third and current stage of the Persian language, New Persian, emerged in the eastern Iranian lands, in the Samanid domains of Khorasan and Transoxiana, in the ninth and tenth centuries CE. The earliest New Persian texts are poems in the fahlawiyat, the dialect poetry of the eastern Iranian countryside, but the language was soon adopted as the literary medium of the Islamic Iranian world, replacing both the older Middle Persian and the Arabic that had been the dominant literary language of the early Islamic period. The New Persian of the classical period, in the works of Rudaki, Ferdowsi, Rumi, Hafez, and Saadi, is sometimes called classical Persian or Farsi-ye Dari, the language of the court.

The most striking development in New Persian was the adoption of the Arabic script, which occurred gradually during the ninth and tenth centuries. The Arabic script, a consonantal alphabet with a small set of vowel diacritics, was well adapted to the phonology of Arabic, but the phonology of Persian included several sounds not present in Arabic, including the consonants /p/, /ch/, /g/, and /zh/, which were written using the Arabic letters pe, che, gaf, and zhe, with the dots added to the existing Arabic forms. The result is a script that can be read by any speaker of Arabic but that represents, with a few additional conventions, the full phonology of Persian. The naskh and nasta’liq styles are the most common in Persian, and the nasta’liq style, developed in the fifteenth and sixteenth centuries, is particularly associated with Persian poetry.

New Persian borrowed extensively from Arabic, especially in the domains of religion, philosophy, science, and administration, and the resulting vocabulary includes an estimated 40%–60% of Arabic loanwords. The grammatical structure, however, remained Indo-Iranian, and the literary tradition preserved the use of Persian words and constructions in poetic and elevated contexts, with Arabic words used in technical and learned contexts. The interplay between the two lexicons — the farsi and the arabi — became one of the great themes of Persian literary criticism, and the farsi-e sarf (pure Persian) movement of the twentieth century tried, without complete success, to replace the Arabic borrowings with native Persian equivalents.

Farsi, Dari, Tajik

The modern descendants of classical Persian are conventionally divided into three major varieties: Farsi, the language of Iran; Dari, the language of Afghanistan; and Tajik, the language of Tajikistan and Uzbekistan. The three varieties are mutually intelligible to a very high degree — speakers of Farsi, Dari, and Tajik can communicate with little or no difficulty — and they are conventionally considered dialects of a single language, Persian, rather than separate languages. The differences are largely phonological, lexical, and in the case of Tajik, orthographic: Tajik is written in a Cyrillic alphabet, in part a legacy of the Soviet period, although there has been a movement to restore the Arabic-based script.

The split between Farsi and Dari is older than is sometimes thought. The Khwarazmian and the Khorasani dialects of the Samanid period, from which the New Persian of the classical poets came, were already differentiated from the dialects of the Iranian plateau, and the Ghaznavid and Timurid courts, based in Afghanistan and eastern Iran, cultivated a recognisably distinct Dari — the “language of the court” — that was the prestige form of Persian in the east. The Mughal court of India, descended from the Timurid, continued to use a recognisably Dari form, and the modern Afghan Dari is the direct descendant of that eastern tradition.

Tajik, the language of Tajikistan, is the easternmost variety of Persian, and the speakers of Tajik are descended from the Sogdians, Bactrians, and other eastern Iranian peoples who, like the western Iranians, adopted the New Persian language of the Samanid period. The Tajik literary tradition, in classical and modern times, includes the poetry of Rudaki, Ferdowsi (whose family was from the Samanid domains of Khorasan and Transoxiana), the Timurid poets of Herat, and the modern poets of Dushanbe and Samarkand. The Tajik ASSR was created in 1924, becoming the Tajik SSR in 1929, and the cultural policy of the Soviet period, which introduced the Cyrillic script and promoted the study of Russian, left a deep mark on the language.

What we don’t know

The biggest open question, in my view, is the question of the spoken language of the Achaemenid court. The Old Persian cuneiform corpus is chancery, not vernacular, and the spoken form of the language is recoverable only by inference from the Aramaic transcriptions and from the later Middle Persian forms. The Median language, which was probably the spoken language of the western Iranian heartland, is even more thinly attested, and the question of how the two languages coexisted — Old Persian as a chancery language, Median as a spoken language, with the imperial court moving between the two — is, I think, not answerable on the present evidence. The recent work of Skjærvø, Cantera, and others on the Avestan corpus is, indirectly, helping to clarify the older end of the Iranian dialect continuum, and we may, in another generation, have a clearer picture of the spoken language of the Achaemenid period.