Ағылшыншамен салыстырыңыз: абзацты басыңыз — түпнұсқа терезеде ашылады. Абзац астындағы EN түймесі оны мәтін ішінде көрсетеді.
Мазмұны
Кіріспе
Машиналық аударма - мәтінді немесе сөйлеуді бір табиғи тілден екіншісіне аудару үшін бағдарламалық жасақтаманы қолдануды зерттейтін есептеу лингвистикасының кіші саласы. 1950 жылдары машиналық аударма зерттеулерде нақты нәрсеге айналды, дегенмен бұл тақырыпқа сілтемелер 17 ғасырдың басында кездеседі. 1954 жылы алтымыстан астам орыс сөйлемдерін ағылшын тіліне толық автоматты түрде аударуды қамтитын Джорджтаун эксперименті ең алғашқы жобалардың бірі болды. Джорджтаун экспериментіне қатысқан зерттеушілер машиналық аударма мәселесі бірнеше жылдың ішінде шешіледі деген сенімде. Көп ұзамай Совет Одағында да осыған ұқсас тәжірибелер жүргізілді. Нәтижесінде, эксперименттің сәттілігі Америка Құрама Штаттарында машиналық аударма зерттеулеріне елеулі қаржы бөлу дәуірін бастады. Жетілген прогресс күтілгеннен әлдеқайда баяу болды; 1966 жылы ALPAC баяндамасы он жылдық зерттеу Джорджтаун эксперименті күткенді орындамағандығын және қаржыландырудың күрт азаюына әкелгенін анықтады. 1980 жылдары қол жетімді есептеу қуатының артуымен машиналық аударма үшін статистикалық модельдерге қызығушылық артып, олар кең таралған және арзан болды. "Текстті шектеусіз толық автоматты түрде жоғары сапалы аударма" автономды жүйесі болмаса да, қатаң шектеулер шеңберінде пайдалы нәтижелер беру мүмкіндігі бар көптеген бағдарламалар қазір бар. Бұл бағдарламалардың бірнешеуі онлайн режимінде қол жетімді, мысалы Google Translate және AltaVista-ның BabelFish-ін (ол 2012 жылдың мамырында Microsoft Bing аудармашысымен ауыстырылды) қамтамасыз ететін SYSTRAN жүйесі.
Machine translation is a sub field of computational linguistics that investigates the use of software to translate text or speech from one natural language to another. In the 1950s, machine translation became a reality in research, although references to the subject can be found as early as the 17th century. The Georgetown experiment, which involved successful fully automatic translation of more than sixty Russian sentences into English in 1954, was one of the earliest recorded projects. Researchers of the Georgetown experiment asserted their belief that machine translation would be a solved problem within a few years. In the Soviet Union, similar experiments were performed shortly after. Consequently, the success of the experiment ushered in an era of significant funding for machine translation research in the United States. The achieved progress was much slower than expected; in 1966, the ALPAC report found that ten years of research had not fulfilled the expectations of the Georgetown experiment and resulted in dramatically reduced funding. Interest grew in statistical models for machine translation, which became more common and also less expensive in the 1980s as available computational power increased. Although there exists no autonomous system of "fully automatic high quality translation of unrestricted text," there are many programs now available that are capable of providing useful output within strict constraints. Several of these programs are available online, such as Google Translate and the SYSTRAN system that powers AltaVista's BabelFish (which was replaced by Microsoft Bing translator in May 2012).
Бастауы
Машиналық аударманың басталуы 9-ғасырдағы араб криптографы Аль Киндидің еңбегіне байланысты болуы мүмкін, ол жүйелі тіл аудармасының әдістерін, соның ішінде криптоанализ, жиілікті талдау, ықтималдық және статистиканы дамытты, олар қазіргі заманғы машиналық аудармада қолданылады. Машиналық аударма идеясы 17 ғасырда пайда болды. 1629 жылы Рене Декарт әртүрлі тілдерде бірдей идеялар бір белгімен бөлісетін әмбебап тілді ұсынды. 1930 жылдардың ортасында "аударма машиналарына" алғашқы патенттерді Жорж Арцруни қағаз таспасын пайдаланатын автоматты екі тілді сөздік үшін қолданды. Орыс Пётр Троянский екі тілді сөздік пен эсперантоның грамматикалық жүйесіне негізделген тілдер арасындағы грамматикалық рөлдерді шешу әдісін қамтитын егжей-тегжейлі ұсынысты ұсынды. Бұл жүйе үш кезеңге бөлінді: бірінші кезеңге сөздерді логикалық формаларына реттеу және синтаксистік функцияларды жүзеге асыру үшін бастапқы тілдегі ана тілімен сөйлейтін редактор кірді; екінші кезеңге машина осы формаларды мақсатты тілге "аудармалауды" қажет етті; үшінші кезеңге мақсатты тілдегі ана тілімен сөйлейтін редактор осы шығысты қалыпқа келтірді. Троянскийдің ұсынысы 1950 жылдардың аяғына дейін белгісіз болды, сол кезде компьютерлер жақсы танымал болды және пайдаланылды.
The origins of machine translation can be traced back to the work of Al Kindi, a 9th century Arabic cryptographer who developed techniques for systemic language translation, including cryptanalysis, frequency analysis, and probability and statistics, which are used in modern machine translation. The idea of machine translation later appeared in the 17th century. In 1629, René Descartes proposed a universal language, with equivalent ideas in different tongues sharing one symbol. In the mid 1930s the first patents for "translating machines" were applied for by Georges Artsrouni, for an automatic bilingual dictionary using paper tape. Russian Peter Troyanskii submitted a more detailed proposal that included both the bilingual dictionary and a method for dealing with grammatical roles between languages, based on the grammatical system of Esperanto. This system was separated into three stages: stage one consisted of a native speaking editor in the source language to organize the words into their logical forms and to exercise the syntactic functions; stage two required the machine to "translate" these forms into the target language; and stage three required a native speaking editor in the target language to normalize this output. Troyanskii's proposal remained unknown until the late 1950s, by which time computers were well known and utilized.
Алғашқы жылдары
Компьютерлік машиналық аударма үшін ұсыныстардың алғашқы топтамасын 1949 жылы Рокфеллер қорының зерттеушісі Уоррен Уивер ұсынды, "Аударма меморандумы". Бұл ұсыныстар ақпараттық теорияға, Екінші дүниежүзілік соғыс кезінде кодтарды бұзудағы табыстарға және табиғи тілдің негізі болып табылатын әмбебап принциптер туралы теорияларға негізделген. Уивер өз ұсыныстарын ұсынғаннан кейін бірнеше жылдан кейін АҚШ-тың көптеген университеттерінде зерттеулер басталды. 1954 жылдың 7 қаңтарында Нью-Йоркте IBM бас кеңсесінде Джорджтаун IBM эксперименті өтті. Бұл машиналық аударма жүйесін алғаш рет көпшілік алдында көрсету болды. Бұл демонстрация туралы газеттерде кеңінен жарияланып, көпшілік қызығушылық танытты. Алайда, жүйенің өзі "ойыншық" жүйе ғана болды. Ол тек 250 сөзден тұратын және 49 ресейлік сөйлемді ағылшын тіліне мұқият аударған, негізінен химия саласында. Дегенмен, бұл машиналық аударманың жақын арада болатыны туралы ойды көтермеледі және АҚШ-та ғана емес, бүкіл әлемде зерттеулерді қаржыландыруды ынталандырды. 1960 жылдары жүргізілген зерттеулер шектеулі тіл жұптары мен кіріске бағытталған болса, 1970 жылдары сұраныс техникалық және коммерциялық құжаттардың аудармаларын жасай алатын арзан жүйелерге қатысты болды. Бұл сұранысқа жаһанданудың өсуі және Канада, Еуропа және Жапонияда аударма сұранысының артуы себеп болды.
The first set of proposals for computer based machine translation was presented in 1949 by Warren Weaver, a researcher at the Rockefeller Foundation, "Translation memorandum". These proposals were based on information theory, successes in code breaking during the Second World War, and theories about the universal principles underlying natural language. A few years after Weaver submitted his proposals, research began in earnest at many universities in the United States. On 7 January 1954 the Georgetown–IBM experiment was held in New York at the head office of IBM. This was the first public demonstration of a machine translation system. The demonstration was widely reported in the newspapers and garnered public interest. The system itself, however, was no more than a "toy" system. It had only 250 words and translated 49 carefully selected Russian sentences into English – mainly in the field of chemistry. Nevertheless, it encouraged the idea that machine translation was imminent and stimulated the financing of the research, not only in the US but worldwide. While research in the 1960s concentrated on limited language pairs and input, demand in the 1970s was for low cost systems that could translate a range of technical and commercial documents. This demand was spurred by the increase of globalisation and the demand for translation in Canada, Europe, and Japan.
1980 жылдар мен 1990 жылдардың басында
1980 жылдары автоматты аударма үшін орнатылған жүйелердің саны да, әртүрлілігі де өсті. SYSTRAN, Logos, Ariane G5 және Metal сияқты мейнфрейм технологиясына негізделген бірқатар жүйелер қолданылды. Микрокомпьютерлердің қол жетімділігінің жақсаруы нәтижесінде төменгі деңгейдегі машиналық аударма жүйелері нарығы пайда болды. Көптеген компаниялар Еуропа, Жапония және АҚШ-та осы мүмкіндікті пайдаланды. Сонымен қатар, Қытай, Шығыс Еуропа, Корея және Кеңес Одағында да жүйелер нарыққа шығарылды. 1980 жылдары ТМ-де әсіресе Жапонияда көп іс-әрекет болды. Бесінші буын компьютерімен Жапония компьютерлік аппараттық және бағдарламалық қамтамасыз етудегі бәсекелестерінен асып түсуді көздеді, ал көптеген ірі жапондық электроника фирмаларының қатысуымен ағылшын тіліне және ағылшын тіліне аудару бағдарламалық қамтамасыз етуін жасау болды (Fujitsu, Toshiba, NTT, Brother, Catena, Matsushita, Mitsubishi, Sharp, Sanyo, Hitachi, NEC, Panasonic, Kodensha, Nova, Oki). 1980 жылдардағы зерттеулер әдетте морфологиялық, синтаксистік және семантикалық талдауды қамтитын аралық тілдік бейнелеудің кейбір түрлері арқылы аудармаға сүйенді. 1980 жылдардың соңында машиналық аударма үшін бірқатар жаңа әдістердің пайда болуы болды. IBM-де статистикалық әдістерге негізделген бір жүйе жасалды. Макото Нагао мен оның тобы көптеген аударма үлгілеріне негізделген әдістерді қолданды, бұл әдіс қазір үлгіге негізделген машиналық аударма деп аталады. Бұл екі тәсілдің де ерекшелігі синтаксистік және семантикалық ережелерді елемеу және үлкен мәтіндік корпустарды манипуляциялауға сүйену болды. 1990 жылдары сөйлеуді тану мен сөйлеу синтезі саласындағы жетістіктерден жігер алған сөйлеуді аудару бойынша зерттеулер неміс Verbmobil жобасын әзірлеу арқылы басталды. Forward Area Language Converter (FALCon) жүйесі - Армияның ғылыми-зерттеу зертханасы әзірлеген машиналық аударма технологиясы. Ол 1997 жылы Босниядағы солдаттар үшін құжаттарды аудару үшін қолданылды. Автоматты аудармалардың қолданылуы арзан және қуатты компьютерлердің пайда болуы нәтижесінде айтарлықтай өсті. Бұл 1990-шы жылдардың басында машиналық аударма үлкен мейнфрейм компьютерлерден дербес компьютерлер мен жұмыс станцияларына көшуді бастады. PC нарығын бір уақыт бойы басқаратын екі компания Globalink және MicroTac болды, содан кейін екі компанияның бірігуі (1994 жылдың желтоқсанында) екеуінің де корпоративтік мүддесіне сай деп табылды. Осы уақытта Intergraph және Systran компьютерлік нұсқаларын ұсына бастады. Сондай-ақ, Интернетте AltaVista's Babel Fish (Systran технологиясын пайдалана отырып) және Google Language Tools (алғашқыда Systran технологиясын ғана пайдалана отырып) сияқты сайттар қол жетімді болды.
By the 1980s, both the diversity and the number of installed systems for machine translation had increased. A number of systems relying on mainframe technology were in use, such as SYSTRAN, Logos, Ariane G5, and Metal. As a result of the improved availability of microcomputers, there was a market for lower end machine translation systems. Many companies took advantage of this in Europe, Japan, and the USA. Systems were also brought onto the market in China, Eastern Europe, Korea, and the Soviet Union. During the 1980s there was a lot of activity in MT in Japan especially. With the fifth generation computer, Japan intended to leap over its competition in computer hardware and software, and one project that many large Japanese electronics firms found themselves involved in was creating software for translating into and from English (Fujitsu, Toshiba, NTT, Brother, Catena, Matsushita, Mitsubishi, Sharp, Sanyo, Hitachi, NEC, Panasonic, Kodensha, Nova, Oki). Research during the 1980s typically relied on translation through some variety of intermediary linguistic representation involving morphological, syntactic, and semantic analysis. At the end of the 1980s, there was a large surge in a number of novel methods for machine translation. One system was developed at IBM that was based on statistical methods. Makoto Nagao and his group used methods based on large numbers of translation examples, a technique that is now termed example based machine translation. A defining feature of both of these approaches was the neglect of syntactic and semantic rules and reliance instead on the manipulation of large text corpora. During the 1990s, encouraged by successes in speech recognition and speech synthesis, research began into speech translation with the development of the German Verbmobil project. The Forward Area Language Converter (FALCon) system, a machine translation technology designed by the Army Research Laboratory, was fielded 1997 to translate documents for soldiers in Bosnia. There was significant growth in the use of machine translation as a result of the advent of low cost and more powerful computers. It was in the early 1990s that machine translation began to make the transition away from large mainframe computers toward personal computers and workstations. Two companies that led the PC market for a time were Globalink and MicroTac, following which a merger of the two companies (in December 1994) was found to be in the corporate interest of both. Intergraph and Systran also began to offer PC versions around this time. Sites also became available on the internet, such as AltaVista's Babel Fish (using Systran technology) and Google Language Tools (also initially using Systran technology exclusively).
2000 жылдар
Машиналық аударма саласында 2000-шы жылдары үлкен өзгерістер болды. Статистикалық машиналық аударма және мысалға негізделген машиналық аударма бойынша көптеген зерттеулер жүргізілді. Сөйлеу аудармалары саласында зерттеулер аударма жүйелерін домендік шектеуге емес, домендік шектеуге бағытталды. Еуропадағы (мысалы, TC STAR) және АҚШ-тағы (STR DUST және DARPA Global автономды тілдерді пайдалану бағдарламасы) түрлі зерттеу жобаларында Парламент баяндамаларын және хабарларды автоматты түрде аудару үшін шешімдер жасалды. Бұл сценарийлерде мазмұнның ауқымы белгілі бір саламен шектелмейді, керісінше аударма жасалатын баяндамалар әр түрлі тақырыптарды қамтиды. Француз-неміс тіліндегі Quaero жобасы көп тілді интернет үшін машиналық аудармаларды пайдалану мүмкіндігін зерттеді. Бұл жоба веб-беттерді ғана емес, сонымен қатар интернеттегі бейне және аудио файлдарды аударуға бағытталған.
The field of machine translation has seen major changes in the 2000s. A large amount of research was done into statistical machine translation and example based machine translation. In the area of speech translation, research was focused on moving from domain limited systems to domain unlimited translation systems. In different research projects in Europe (like TC STAR) and in the United States (STR DUST and DARPA Global autonomous language exploitation program), solutions for automatically translating Parliamentary speeches and broadcast news was developed. In these scenarios the domain of the content was no longer limited to any special area, but rather the speeches to be translated cover a variety of topics. The French–German project Quaero investigated the possibility of making use of machine translations for a multi lingual internet. The project sought to translate not only webpages, but also videos and audio files on the internet.
2010 жылдар
Соңғы онжылдықта статистикалық машиналық аударманың орнын жүйкелік машиналық аударма (НМТ) әдістері алды. Нейрологиялық машиналық аударма термині 2014 жылы осы тақырыпқа қатысты алғашқы зерттеуді жариялаған Бахданау және басқалар және Суцкевер және басқалар жасаған. Нейро желілеріне статистикалық модельдерге қажетті жадтың тек бір бөлігі ғана қажет болды және бүтін сөйлемдерді интеграцияланған түрде модельдеуге болады. Бірінші ірі масштабтағы НМТ-ны Baidu 2015 жылы іске қосты, содан кейін 2016 жылы Google Neural Machine Translation (GNMT) іске қосылды. Одан кейін DeepL Translator сияқты басқа аударма қызметтері пайда болды және NMT технологиясын Microsoft Translator сияқты ескі аударма қызметтерінде қабылдады. Нейрондық желілер бірден бір нервтік желі архитектурасын (seq2seq) пайдаланып, екі қайталанатын нейрондық желілерді (RNN) қолданады. Кодтаушы RNN және декодер RNN. Кодтаушы RNN бастапқы сөйлемде кодтау векторларын қолданады, ал декодер RNN алдыңғы кодтау векторының негізінде мақсатты сөйлемді жасайды. Көңіл бөлу қабатының, трансформация мен кері таралу техникасының одан әрі жетілуі НМТ-ны икемді етіп, көптеген машиналық аударма, қорытындылау және чат-бот технологияларында қабылдады.
The past decade witnessed neural machine translation (NMT) methods replace statistical machine translation. The term neural machine translation was coined by Bahdanau et al and Sutskever et al who also published the first research regarding this topic in 2014. Neural networks only needed a fraction of the memory needed by statistical models and whole sentences could be modeled in an integrated manner. The first large scale NMT was launched by Baidu in 2015, followed by Google Neural Machine Translation (GNMT) in 2016. This was followed by other translation services like DeepL Translator and the adoption of NMT technology in older translation services like Microsoft translator. Neural networks use a single end to end neural network architecture known as sequence to sequence (seq2seq) which uses two recurrent neural networks (RNN). An encoder RNN and a decoder RNN. Encoder RNN uses encoding vectors on the source sentence and the decoder RNN generates the target sentence based on the previous encoding vector. Further advancements in the attention layer, transformation and back propagation techniques have made NMTs flexible and adopted in most machine translation, summarization and chatbot technologies.