Ағылшыншамен салыстырыңыз: абзацты басыңыз — түпнұсқа терезеде ашылады. Абзац астындағы EN түймесі оны мәтін ішінде көрсетеді.
Мазмұны
Кіріспе
Сөздердің бір-біріне жиі қолданылуы
корпустық лингвистика ұғымы
Frequent occurrence of words next to each other
the corpus linguistics notion
Корпустық лингвистикада, тіркес – бұл сөздердің немесе терминдердің қатар келу жиілігі, кездейсоқ күтілгеннен артық болатын тізбегі. Фразеологияда тіркес – құрамына кіретін сөздерден түсінілетін композициялық фраземаның бір түрі. Бұл идиомадан өзгеше, онда тұтастың мағынасы оның бөліктерінен шығарылмайды және мүлдем байланыссыз болуы мүмкін. Тіркестердің жеті негізгі түрі бар: сын есім + есім, есім + есім (мысалы, жиын есімдер), есім + етістік, етістік + есім, үстеу + сын есім, етістік + сөз тіркесі (фразалық етістіктер) және етістік + үстеу. Тіркес табу – деректерді өндіруге ұқсас түрлі есептеу лингвистикасы элементтерін қолдана отырып, құжат немесе корпуста тіркестерді анықтайтын есептеу әдісі.
In corpus linguistics, a collocation is a series of words or terms that co occur more often than would be expected by chance. In phraseology, a collocation is a type of compositional phraseme, meaning that it can be understood from the words that make it up. This contrasts with an idiom, where the meaning of the whole cannot be inferred from its parts, and may be completely unrelated. There are about seven main types of collocations: adjective + noun, noun + noun (such as collective nouns), noun + verb, verb + noun, adverb + adjective, verbs + prepositional phrase (phrasal verbs), and verb + adverb. Collocation extraction is a computational technique that finds collocations in a document or corpus, using various computational linguistics elements resembling data mining.
Сөздіктерде
1933 жылы Гарольд Палмердің ағылшын тіліндегі сөз тіркестері туралы екінші аралық баяндамасы шет тілін үйренуші кез келген адам үшін сөз тіркестерінің табиғи дыбысты тілде сөйлеудің кілті екенін көрсетті. Осылайша, 1940 жылдардан бастап, жиі кездесетін сөз тіркестері туралы ақпарат бір тілдік оқушылар сөздіктерінің маңызды бөлігіне айналды. Сөздіктер «сөздерге баса назар бермей, сөз тіркестеріне көбірек көңіл бөлгенде», сөз тіркестеріне назар аудару арта түсті. Бұл тенденцияны 21-ші ғасырдың басынан бастап ірі мәтіндік корпус пен интеллектуалды корпус іздеу бағдарламалық құралдарының қолжетімділігімен күшейтті, бұл сөздіктерде сөз тіркестерін жүйелі түрде қарастыруға мүмкіндік берді. Осы құралдарды пайдаланып, «Макмиллан ағылшын сөздігі» және «Лонгман қазіргі заманғы ағылшын тілінің сөздігі» сияқты сөздіктерге жиі кездесетін сөз тіркестерінің тізімдерімен қораптар немесе бөлімдер қосылды. Сонымен қатар, тілдегі жиі кездесетін сөз тіркестерін сипаттауға арналған бірнеше мамандандырылған сөздіктер де бар. Оларға (испан тілінде) Redes: Diccionario combinatorio del español contemporaneo (2004), (француз тілінде) Le Robert: Dictionnaire des combinaisons de mots (2007) және (ағылшын тілінде) LTP Dictionary of Selected Collocations (1997) мен Macmillan Collocations Dictionary (2010) кіреді.
In 1933, Harold Palmer's Second Interim Report on English Collocations highlighted the importance of collocation as a key to producing natural sounding language, for anyone learning a foreign language. Thus from the 1940s onwards, information about recurrent word combinations became a standard feature of monolingual learner's dictionaries. As these dictionaries became "less word centred and more phrase centred", more attention was paid to collocation. This trend was supported, from the beginning of the 21st century, by the availability of large text corpora and intelligent corpus querying software, making it possible to provide a more systematic account of collocation in dictionaries. Using these tools, dictionaries such as the Macmillan English Dictionary and the Longman Dictionary of Contemporary English included boxes or panels with lists of frequent collocations. There are also a number of specialized dictionaries devoted to describing the frequent collocations in a language. These include (for Spanish) Redes: Diccionario combinatorio del español contemporaneo (2004), (for French) Le Robert: Dictionnaire des combinaisons de mots (2007), and (for English) the LTP Dictionary of Selected Collocations (1997) and the Macmillan Collocations Dictionary (2010).
Статистикалық маңызды орналасу
Студенттің t тестісі корпустағы тіркестің кездесуінің статистикалық маңыздылығын анықтау үшін қолданылуы мүмкін. Биграмма үшін, корпустың өлшемімен белгіленетін корпуста кездесуінің шартсыз ықтималдығы болсын, ал корпуста кездесуінің шартсыз ықтималдығы болсын. Биграмма үшін t көрсеткіші былай есептеледі:
Student's t test can be used to determine whether the occurrence of a collocation in a corpus is statistically significant. For a bigram , let be the unconditional probability of occurrence of in a corpus with size , and let be the unconditional probability of occurrence of in the corpus. The t score for the bigram is calculated as:
мұнда – кездесуінің еңбектік орташасы, – кездесулер саны, – текстте және тәуелсіз түрде пайда болады деген нөлдік гипотеза бойынша ықтималдығы, ал – еңбектік дисперсия. Үлкен болған жағдайда, t тестісі Z тестісімен тең.
where is the sample mean of the occurrence of , is the number of occurrences of , is the probability of under the null hypothesis that and appear independently in the text, and is the sample variance. With a large , the t test is equivalent to a Z test.