Сравнивайте с английским: нажмите на абзац — оригинал откроется в окне. Кнопка EN под абзацем показывает его прямо в тексте.
Содержание
Введение
Самая маленькая функциональная единица письма
Smallest functional written unit
В лингвистике графема — это наименьшая смыслоразличительная единица системы письма. Слово «графема» образовано от основы и суффикса *eme* по аналогии с фонемой и другими терминами для эмических единиц. Изучение графем называется графемикой. Понятие графемы абстрактно и сопоставимо с понятием символа в информатике. В отличие от этого, конкретная графическая форма, представляющая ту или иную графему в определенном шрифте, называется глифом.
In linguistics, a grapheme is the smallest functional unit of a writing system. The word grapheme is derived and the suffix eme by analogy with phoneme and other names of emic units. The study of graphemes is called graphemics. The concept of graphemes is abstract and similar to the notion in computing of a character. By comparison, a specific shape that represents any particular grapheme in a given typeface is called a glyph.
Концептуализация
Существуют две основные противоположные концепции графемы. В так называемой референтной концепции графемы интерпретируются как наименьшие единицы письма, соответствующие звукам (точнее, фонемам). В рамках этой концепции сочетание sh в английском слове shake является графемой, поскольку оно представляет фонему /ʃ/. Эта референтная концепция связана с гипотезой зависимости, утверждающей, что письмо лишь отображает речь. В отличие от неё, аналогическая концепция определяет графемы аналогично фонемам, то есть посредством письменных минимальных пар, таких как shake и snake. В этом примере h и n являются графемами, поскольку они различают два слова. Эта аналогическая концепция связана с гипотезой автономии, согласно которой письмо представляет собой самостоятельную систему и должно изучаться независимо от речи. Обе концепции имеют свои недостатки. Некоторые модели придерживаются обеих концепций одновременно, выделяя две отдельные единицы, которым даются такие названия, как графема, определяемая по аналогическому принципу (h в слове shake), и графема, соответствующая референтной концепции (sh в слове shake). В более новых концепциях, где графема интерпретируется семиотически как диадический лингвистический знак, она определяется как минимальная единица письма, которая одновременно является лексически различимой и соответствует языковой единице (фонеме, слогу или морфеме).
There are two main opposing grapheme concepts. In the so called referential conception, graphemes are interpreted as the smallest units of writing that correspond with sounds (more accurately phonemes). In this concept, the sh in the written English word shake would be a grapheme because it represents the phoneme /ʃ/. This referential concept is linked to the dependency hypothesis that claims that writing merely depicts speech. By contrast, the analogical concept defines graphemes analogously to phonemes, i. e. via written minimal pairs such as shake vs. snake. In this example, h and n are graphemes because they distinguish two words. This analogical concept is associated with the autonomy hypothesis which holds that writing is a system in its own right and should be studied independently from speech. Both concepts have weaknesses. Some models adhere to both concepts simultaneously by including two individual units, which are given names such as graphemic grapheme for the grapheme according to the analogical conception (h in shake), and phonological fit grapheme for the grapheme according to the referential concept (sh in shake). In newer concepts, in which the grapheme is interpreted semiotically as a dyadic linguistic sign, it is defined as a minimal unit of writing that is both lexically distinctive and corresponds with a linguistic unit (phoneme, syllable, or morpheme).
Обозначение
Графемы часто обозначаются в угловых скобках: < >, < >, и т. д. Это аналогично как обозначению с косой чертой (/ /) для фонем, так и обозначению в квадратных скобках ([ ]) для фонетической транскрипции.
Graphemes are often notated within angle brackets: , , etc. This is analogous to both the slash notation (, ) used for phonemes and to the square bracket notation used for phonetic transcriptions (, ).
Глифы
Точно так же, как поверхностные формы фонем – это звуки речи или телефоны (а разные телефоны, представляющие одну и ту же фонему, называются аллофонами), поверхностные формы графем – это глифы (иногда графики), то есть конкретные письменные представления символов (а разные глифы, представляющие одну и ту же графем, называются аллографами). Таким образом, графему можно рассматривать как абстракцию набора глифов, которые все функционально эквивалентны. Например, в письменном английском (или других языках, использующих латинский алфавит), существуют два различных физических представления строчной латинской буквы "a": "a" и "ɑ". Однако, поскольку замена одного из них другим не может изменить значение слова, они считаются аллографами одной и той же графемы, которую можно записать аналогичным образом. Графема, соответствующая "арабскому нулю", имеет уникальную семантическую идентичность и значение Unicode, но проявляется в виде перечеркнутого нуля. Курсивные и полужирные формы также являются аллографическими, как и вариации, наблюдаемые в шрифтах с засечками (например, Times New Roman) и без засечек (например, Helvetica). Существуют разногласия относительно того, являются ли прописные и строчные буквы аллографами или отдельными графемами. Прописные буквы обычно встречаются в определенных контекстах, не меняющих значение слова: например, в именах собственных или в начале предложения, или во всей заглавной печати в газетном заголовке. В других контекстах капитализация может определять значение: сравните, например, Polish и polish: первое – это язык, второе – для начистки обуви. Некоторые лингвисты считают диграфы, такие как «sh» в слове ship, отдельными графемами, но обычно они анализируются как последовательности графем. Однако нестилистические лигатуры, такие как «æ», являются отдельными графемами, как и различные буквы с отличительными диакритическими знаками, такие как .
In the same way that the surface forms of phonemes are speech sounds or phones (and different phones representing the same phoneme are called allophones), the surface forms of graphemes are glyphs (sometimes graphs), namely concrete written representations of symbols (and different glyphs representing the same grapheme are called allographs). Thus, a grapheme can be regarded as an abstraction of a collection of glyphs that are all functionally equivalent. For example, in written English (or other languages using the Latin alphabet), there are two different physical representations of the lowercase Latin letter "a": "a" and "ɑ". Since, however, the substitution of either of them for the other cannot change the meaning of a word, they are considered to be allographs of the same grapheme, which can be written Similarly, the grapheme corresponding to "Arabic numeral zero" has a unique semantic identity and Unicode value but exhibits variation in the form of slashed zero. Italic and bold face forms are also allographic, as is the variation seen in serif (as in Times New Roman) versus sans serif (as in Helvetica) forms. There is some disagreement as to whether capital and lower case letters are allographs or distinct graphemes. Capitals are generally found in certain triggering contexts that do not change the meaning of a word: a proper name, for example, or at the beginning of a sentence, or all caps in a newspaper headline. In other contexts, capitalization can determine meaning: compare, for example Polish and polish: the former is a language, the latter is for shining shoes. Some linguists consider digraphs like the in ship to be distinct graphemes, but these are generally analyzed as sequences of graphemes. Non stylistic ligatures, however, such as , are distinct graphemes, as are various letters with distinctive diacritics, such as
Identical glyphs may not always represent the same grapheme. For example, the three letters , and appear identical but each has a different meaning: in order, they are the Latin letter A, the Cyrillic letter Azǔ/Азъ and the Greek letter Alpha. Each has its own code point in Unicode: , and .
Идентичные глифы не всегда представляют одну и ту же графему. Например, три буквы «A», «А» и «Α» выглядят идентично, но каждая имеет разное значение: в порядке, это латинская буква A, кириллическая буква Азъ/Аз и греческая буква Альфа. У каждой есть свой код в Unicode: U+0041, U+0041 и U+0391.
In the same way that the surface forms of phonemes are speech sounds or phones (and different phones representing the same phoneme are called allophones), the surface forms of graphemes are glyphs (sometimes graphs), namely concrete written representations of symbols (and different glyphs representing the same grapheme are called allographs). Thus, a grapheme can be regarded as an abstraction of a collection of glyphs that are all functionally equivalent. For example, in written English (or other languages using the Latin alphabet), there are two different physical representations of the lowercase Latin letter "a": "a" and "ɑ". Since, however, the substitution of either of them for the other cannot change the meaning of a word, they are considered to be allographs of the same grapheme, which can be written Similarly, the grapheme corresponding to "Arabic numeral zero" has a unique semantic identity and Unicode value but exhibits variation in the form of slashed zero. Italic and bold face forms are also allographic, as is the variation seen in serif (as in Times New Roman) versus sans serif (as in Helvetica) forms. There is some disagreement as to whether capital and lower case letters are allographs or distinct graphemes. Capitals are generally found in certain triggering contexts that do not change the meaning of a word: a proper name, for example, or at the beginning of a sentence, or all caps in a newspaper headline. In other contexts, capitalization can determine meaning: compare, for example Polish and polish: the former is a language, the latter is for shining shoes. Some linguists consider digraphs like the in ship to be distinct graphemes, but these are generally analyzed as sequences of graphemes. Non stylistic ligatures, however, such as , are distinct graphemes, as are various letters with distinctive diacritics, such as
Identical glyphs may not always represent the same grapheme. For example, the three letters , and appear identical but each has a different meaning: in order, they are the Latin letter A, the Cyrillic letter Azǔ/Азъ and the Greek letter Alpha. Each has its own code point in Unicode: , and .
Типы графемы
Основными типами графем являются логограммы (точнее – морфограммы), которые представляют слова или морфемы (например, китайские иероглифы, амперсанд "&", обозначающий слово "и", арабские цифры); слоговые знаки, представляющие слоги (как в японской кана); и алфавитные буквы, приблизительно соответствующие фонемам (см. следующий раздел). Для более подробного обсуждения различных типов см. В письменности также используются дополнительные графемические элементы, такие как знаки препинания, математические символы, разделители слов, например пробел, и другие типографские символы. В древних логографических системах письма часто применялись немая детерминативы для уточнения значения соседнего (не немая) слова.
The principal types of graphemes are logograms (more accurately termed morphograms), which represent words or morphemes (for example Chinese characters, the ampersand "&" representing the word and, Arabic numerals); syllabic characters, representing syllables (as in Japanese kana); and alphabetic letters, corresponding roughly to phonemes (see next section). For a full discussion of the different types, see
There are additional graphemic components used in writing, such as punctuation marks, mathematical symbols, word dividers such as the space, and other typographic symbols. Ancient logographic scripts often used silent determinatives to disambiguate the meaning of a neighboring (non silent) word.
Связь с фонемами
Как упоминалось в предыдущем разделе, в языках, использующих алфавитные системы письма, многие графемы в принципе соответствуют фонемам (значимым звукам) языка. Однако на практике орфография таких языков в той или иной степени отклоняется от идеала точного соответствия между графемами и фонемами. Фонема может быть представлена мультиграфом (последовательностью из более чем одной графемы), как, например, диграф *sh* представляет один звук в английском языке (а иногда одна графема может представлять более одной фонемы, как в случае с русской буквой я или испанской *c*). Некоторые графемы могут вообще не соответствовать какому-либо звуку (например, *b* в английском слове *debt* или *h* во всех испанских словах, содержащих эту букву), и часто правила соответствия между графемами и фонемами становятся сложными или нерегулярными, особенно в результате исторических звуковых изменений, которые не всегда отражаются в написании. "Поверхностные" орфографии, такие как стандартные испанская и финская, имеют относительно регулярное (хотя и не всегда однозначное) соответствие между графемами и фонемами, в то время как орфографии французского и английского языков характеризуются гораздо менее регулярным соответствием и называются "глубокими". Мультиграфы, представляющие одну фонему, обычно рассматриваются как сочетания отдельных букв, а не как самостоятельные графемы. Однако в некоторых языках мультиграф может рассматриваться как единый элемент при сортировке; например, в чешском словаре раздел для слов, начинающихся с , следует за разделом для слов, начинающихся с .
As mentioned in the previous section, in languages that use alphabetic writing systems, many of the graphemes stand in principle for the phonemes (significant sounds) of the language. In practice, however, the orthographies of such languages entail at least a certain amount of deviation from the ideal of exact grapheme–phoneme correspondence. A phoneme may be represented by a multigraph (sequence of more than one grapheme), as the digraph sh represents a single sound in English (and sometimes a single grapheme may represent more than one phoneme, as with the Russian letter я or the Spanish c). Some graphemes may not represent any sound at all (like the b in English debt or the h in all Spanish words containing the said letter), and often the rules of correspondence between graphemes and phonemes become complex or irregular, particularly as a result of historical sound changes that are not necessarily reflected in spelling. "Shallow" orthographies such as those of standard Spanish and Finnish have relatively regular (though not always one to one) correspondence between graphemes and phonemes, while those of French and English have much less regular correspondence, and are known as deep orthographies. Multigraphs representing a single phoneme are normally treated as combinations of separate letters, not as graphemes in their own right. However, in some languages a multigraph may be treated as a single unit for the purposes of collation; for example, in a Czech dictionary, the section for words that start with comes after that for For more examples, see .