Ағылшыншамен салыстырыңыз: абзацты басыңыз — түпнұсқа терезеде ашылады. Абзац астындағы EN түймесі оны мәтін ішінде көрсетеді.
Мазмұны
Кіріспе
Машиналық аударманың түрі
Type of machine translation
Тіл аралық машиналық аударма – машиналық аударманың классикалық тәсілдерінің бірі. Бұл тәсілде бастапқы тіл, яғни аударылатын мәтін, тілге тәуелсіз абстрактілі бейнелеу – интерлингваға түрлендіріледі. Содан кейін мақсатты тіл интерлингвадан жасалады. Ережелерге негізделген машиналық аударма парадигмасы шеңберінде, тіл аралық тәсіл тікелей және трансферлік тәсілдерге балама болып табылады. Тікелей тәсілде сөздер қосымша бейнелеуден өтпей тікелей аударылады. Трансферлік тәсілде бастапқы тіл абстрактілі, тілге тән емес бейнелеуге түрлендіріледі. Тіл жұптарына тән тілдік ережелер бастапқы тілдің бейнелеуін абстрактілі мақсатты тілдің бейнелеуіне айналдырады, содан мақсатты сөйлем құрылады. Тіл аралық машиналық аударманың артықшылықтары мен кемшіліктері бар. Оның артықшылықтары: әр бастапқы тілді әр мақсатты тілмен байланыстыру үшін аз компоненттер қажет, жаңа тілді қосу үшін де аз компоненттер жеткілікті, кіріс мәтінді түпнұсқа тілінде қайта формулиреуге мүмкіндік береді, талдаушылар мен генераторларды бір тілді жүйе әзірлеушілер жасай алады, сондай-ақ бір-бірінен өте ерекшеленетін тілдерді (мысалы, ағылшын және араб тілдерін) өңдей алады. Керісінше, айқын кемшілігі – кең ауқымда интерлингваны анықтау қиын, тіпті мүмкін емес болуы. Сондықтан, тіл аралық машиналық аударма үшін ең қолайлы жағдай – нақты бір салаға қатысты көп тілді машиналық аударма. Мысалы, Интерлингва халықаралық конференцияларда делдал тіл ретінде қолданылып келеді және Еуропалық Одақ үшін делдал тіл ретінде ұсынылған.
Interlingual machine translation is one of the classic approaches to machine translation. In this approach, the source language, i. e. the text to be translated is transformed into an interlingua, i. e., an abstract language independent representation. The target language is then generated from the interlingua. Within the rule based machine translation paradigm, the interlingual approach is an alternative to the direct approach and the transfer approach. In the direct approach, words are translated directly without passing through an additional representation. In the transfer approach the source language is transformed into an abstract, less language specific representation. Linguistic rules which are specific to the language pair then transform the source language representation into an abstract target language representation and from this the target sentence is generated. The interlingual approach to machine translation has advantages and disadvantages. The advantages are that it requires fewer components in order to relate each source language to each target language, it takes fewer components to add a new language, it supports paraphrases of the input in the original language, it allows both the analysers and generators to be written by monolingual system developers, and it handles languages that are very different from each other (e. g. English and Arabic). The obvious disadvantage is that the definition of an interlingua is difficult and maybe even impossible for a wider domain. The ideal context for interlingual machine translation is thus multilingual machine translation in a very specific domain. For example, Interlingua has been used as a pivot language in international conferences and has been proposed as a pivot language for the European Union.
Тарих
Тіл аралық машиналық аударма туралы алғашқы идеялар 17 ғасырда Декарт пен Лейбницпен пайда болды, олар қазіргі үлкен тілдік модельдер қолданатын сандық таңбалардан айырмашылығы жоқ, әмбебап сандық кодтарды пайдалана отырып сөздіктерді қалай жасау керектігі туралы теориялар ұсынды. Басқалар, мысалы, Кев Бек, Афанасий Кирхер және Иоганн Йоахим Бехер логика мен иконография принциптеріне негізделген түсінікті әмбебап тілді әзірлеуде жұмыс істеді. 1668 жылы Джон Уилкинс өзінің интерлингвасын "Нағыз кейіпкер мен философиялық тілге арналған эссесінде" сипаттады. 18-19 ғасырларда "әмбебап" халықаралық тілдер туралы көптеген ұсыныстар жасалды, ең танымал – Эсперанто. Алайда, әмбебап тіл идеясын машиналық аудармаға қолдану алғашқы маңызды тәсілдердің ешқайсысында көрінбеді. Оның орнына, тілдер жұптары бойынша жұмыс басталды. Бірақ 1950 және 60-жылдары Маргарет Мастерман басқаратын Кембридждегі, Николай Андреев басқаратын Ленинградтағы және Сильвио Чекато басқаратын Миландағы зерттеушілер осы салада жұмыс істей бастады. Бұл идеяны 1969 жылы Израиль философы Йехошуа Бар Хиллель кеңінен талқылады. 1970 жылдары Гренобльде физика және математикалық мәтіндерді орыс тілінен француз тіліне аударуға тырысқан зерттеушілер, ал Техаста орыс тілінен ағылшын тіліне ұқсас жоба (METAL) жүргізілді. Алғашқы тіл аралық МТ жүйелері 1970 жылдары Стэнфордта Роджер Шэнк пен Йорк Уилкс құрды; біріншісі қаржы аударудың коммерциялық жүйесінің негізіне айналды, ал екіншісінің коды Бостондағы Компьютер мұражайында алғашқы тіл аралық машиналық аударма жүйесі ретінде сақталған. 1980 жылдары машиналық аудармаға интерлингвалық және білімге негізделген тәсілдерге жаңа мән берілді, бұл салада көптеген зерттеулер жүргізілді. Бұл зерттеулерді біріктіретін фактор – жоғары сапалы аударма мәтінді толық түсінуді талап ету идеясынан бас тартуды қажет етеді. Керісінше, аударма тілдік білімге және жүйені қолданатын нақты салаға негізделуі керек. Бұл дәуірдегі ең маңызды зерттеулер Эсперантоның өңделген нұсқасымен жұмыс істейтін Утрехттегі үлестірілген тілдік аударма (DLT) және Жапониядағы Фуджитсу жүйесімен жүргізілді.
The first ideas about interlingual machine translation appeared in the 17th century with Descartes and Leibniz, who came up with theories of how to create dictionaries using universal numerical codes, not unlike numerical tokens used by large language models nowadays. Others, such as Cave Beck, Athanasius Kircher and Johann Joachim Becher worked on developing an unambiguous universal language based on the principles of logic and iconographs. In 1668, John Wilkins described his interlingua in his "Essay towards a Real Character and a Philosophical Language". In the 18th and 19th centuries many proposals for "universal" international languages were developed, the most well known being Esperanto. That said, applying the idea of a universal language to machine translation did not appear in any of the first significant approaches. Instead, work started on pairs of languages. However, during the 1950s and 60s, researchers in Cambridge headed by Margaret Masterman, in Leningrad headed by Nikolai Andreev and in Milan by Silvio Ceccato started work in this area. The idea was discussed extensively by the Israeli philosopher Yehoshua Bar Hillel in 1969. During the 1970s, noteworthy research was done in Grenoble by researchers attempting to translate physics and mathematical texts from Russian to French, and in Texas a similar project (METAL) was ongoing for Russian to English. Early interlingual MT systems were also built at Stanford in the 1970s by Roger Schank and Yorick Wilks; the former became the basis of a commercial system for the transfer of funds, and the latter's code is preserved at The Computer Museum at Boston as the first interlingual machine translation system. In the 1980s, renewed relevance was given to interlingua based, and knowledge based approaches to machine translation in general, with much research going on in the field. The uniting factor in this research was that high quality translation required abandoning the idea of requiring total comprehension of the text. Instead, the translation should be based on linguistic knowledge and the specific domain in which the system would be used. The most important research of this era was done in distributed language translation (DLT) in Utrecht, which worked with a modified version of Esperanto, and the Fujitsu system in Japan.
Контурлық
Аударманың осы әдісінде интерлингва – бастапқы тілде жазылған мәтіннің талдауын сипаттаудың бір жолы ретінде қарастырылады, осы арқылы оның морфологиялық, синтаксистік, семантикалық (тіпті прагматикалық) ерекшеліктерін, яғни "мағынасын" мақсатты тілге түрлендіруге болады. Бұл интерлингва бір тілден екінші тілге аударма жасаудың орнына, аударма жасалатын барлық тілдердің барлық ерекшеліктерін сипаттауға қабілетті. Кейде аудармада екі интерлингва қолданылады. Мүмкін, олардың біреуі бастапқы тілдің ерекшеліктерін көбірек қамтиды, ал екіншісі – мақсатты тілдің ерекшеліктерін. Аударма екі кезеңде жүзеге асырылады: бастапқы тілдегі сөйлемдерді екінші интерлингва арқылы мақсатты тілге жақын сөйлемдерге түрлендіру. Жүйені сондай етіп құруға болады, екінші интерлингва мақсатты тілге жақын немесе оған сәйкес келетін нақты сөздіктерді пайдаланады, бұл аударма сапасын арттыруға мүмкіндік береді. Аталған жүйе бір түпнұсқалық талдаудан көптеген құрылымдық жағынан ұқсас тілдерге аударма сапасын жақсарту үшін тілдік жақындықты пайдалану идеясына негізделген. Бұл принцип тіркесімді машиналық аудармада да қолданылады, онда бір табиғи тіл екі алыс тілдің арасындағы "көпір" ретінде қолданылады. Мысалы, украин тілінен ағылшын тіліне аударғанда орыс тілі аралық тіл ретінде пайдаланылуы мүмкін.
In this method of translation, the interlingua can be thought of as a way of describing the analysis of a text written in a source language such that it is possible to convert its morphological, syntactic, semantic (and even pragmatic) characteristics, that is "meaning" into a target language. This interlingua is able to describe all of the characteristics of all of the languages which are to be translated, instead of simply translating from one language to another. Sometimes two interlinguas are used in translation. It is possible that one of the two covers more of the characteristics of the source language, and the other possess more of the characteristics of the target language. The translation then proceeds by converting sentences from the first language into sentences closer to the target language through two stages. The system may also be set up such that the second interlingua uses a more specific vocabulary that is closer, or more aligned with the target language, and this could improve the translation quality. The above mentioned system is based on the idea of using linguistic proximity to improve the translation quality from a text in one original language to many other structurally similar languages from only one original analysis. This principle is also used in pivot machine translation, where a natural language is used as a "bridge" between two more distant languages. For example, in the case of translating to English from Ukrainian using Russian as an intermediate language.
Тиімділік
Бұл стратегияның басты артықшылығы – көп тілді аударма жүйелерін жасаудың тиімді жолын ұсынады. Интерлингва қолданылғанда, жүйедегі әр тіл жұбы үшін жеке аударма жасау қажет болмайды. Яғни, жүйедегі тілдер саны болса, тіл жұптарын жасаудың орнына, тек тілдер мен интерлингва арасында жұптар жасау жеткілікті. Бұл стратегияның басты кемшілігі – тиімді интерлингва құрудың қиындығы. Ол бастапқы және мақсатты тілдерге тәуелсіз, сонымен қатар абстрактілі болуы керек. Аударма жүйесіне қаншама тіл қосылса, және олар бір-бірінен қаншама өзгеше болса, интерлингваның барлық мүмкін аударма бағыттарын жеткізуге соншалықты күшті болуы қажет. Тағы бір мәселе – түпнұсқа тілдердегі мәтіндерден аралық өрнекті жасау үшін мағынаны анықтау қиын.
One of the main advantages of this strategy is that it provides an economical way to make multilingual translation systems. With an interlingua it becomes unnecessary to make a translation pair between each pair of languages in the system. So instead of creating language pairs, where is the number of languages in the system, it is only necessary to make pairs between the languages and the interlingua. The main disadvantage of this strategy is the difficulty of creating an adequate interlingua. It should be both abstract and independent of the source and target languages. The more languages added to the translation system, and the more different they are, the more potent the interlingua must be to express all possible translation directions. Another problem is that it is difficult to extract meaning from texts in the original languages to create the intermediate representation.