Сұрақ-жауап жүйелері: Компьютер ғылымы, ақпаратты іздеу, тілді өңдеу саласы. BASEBALL, LUNAR сияқты алғашқы жүйелерге шолу. Автоматты жауап беру технологиясы.
Ағылшыншамен салыстырыңыз: абзацты басыңыз — түпнұсқа терезеде ашылады. Абзац астындағы EN түймесі оны мәтін ішінде көрсетеді.
Мазмұны
Кіріспе
Компьютерлік ғылым саласы
Сұрақ-жауап (QA) – ақпаратты іздеу және табиғи тілді өңдеу (NLP) салаларындағы компьютерлік ғылым саласы. Ол адамдардың табиғи тілде қойған сұрақтарға автоматты түрде жауап беретін жүйелерді құрумен айналысады.
Computer science discipline
Question answering (QA) is a computer science discipline within the fields of information retrieval and natural language processing (NLP) that is concerned with building systems that automatically answer questions that are posed by humans in a natural language.
Тарих
Алғашқы екі сұрақ-жауап жүйесі – BASEBALL және LUNAR болды. BASEBALL бір жыл бойы Үлкен лига бейсболы туралы сұрақтарға жауап берді. LUNAR «Аполлон» ғарыш миссияларымен Айдан әкелінген тау жыныстарының геологиялық талдауы туралы сұрақтарға жауап берді. Екі сұрақ-жауап жүйесі де өздері таңдаған салаларында өте тиімді болды. LUNAR 1971 жылы Ай ғылымы конвенциясында көрсетілді және жүйе туралы білімі жоқ адамдар қойған өз саласы бойынша сұрақтардың 90%-ына жауап бере алды. Келесі жылдары тағы да шектеулі салалық сұрақ-жауап жүйелері әзірленді. Бұл жүйелердің ортақ ерекшелігі – олардың таңдалған саланың мамандары қолмен жазған негізгі деректер қоры немесе білім жүйесінің болуы. BASEBALL және LUNAR тілдік мүмкіндіктері ELIZA және DOCTOR сияқты алғашқы чатбот бағдарламаларына ұқсас техникаларды қолданды. SHRDLU – Терри Виноградтың 1960 жылдардың соңы мен 1970 жылдардың басында әзірлеген сәтті сұрақ-жауап беру бағдарламасы. Ол ойыншық әлеміндегі роботтың жұмысын («блоктар әлемі») симуляциялады және роботқа әлем туралы сұрақтар қою мүмкіндігін ұсынды. Бұл жүйенің артықшылығы – өте нақты доменді және физика ережелері компьютерлік бағдарламаға кодтау оңай болған өте қарапайым әлемді таңдауы болды. 1970 жылдары білімнің тар салаларын бағытталған білім базалары жасалды. Осы сарапшы жүйелермен байланысу үшін әзірленген сұрақ-жауап жүйелері білім саласындағы сұрақтарға дұрыс және нақты жауаптар берді. Бұл сарапшы жүйелер ішкі архитектурасынан басқа қазіргі заманғы сұрақ-жауап жүйелеріне өте ұқсас болды. Сарапшы жүйелер сарапшылар құрастырған және ұйымдастырған білім базаларына көп сүйенеді, ал көптеген қазіргі заманғы сұрақ-жауап жүйелері үлкен, құрылымдалмаған, табиғи тілдік мәтін корпусының статистикалық өңдеуіне сүйенеді. 1970 және 1980 жылдары есептеу лингвистикасының жан-жақты теориялары дамыды, бұл мәтінді түсіну және сұрақтарға жауап беру бойынша амбициялық жобалардың дамуына әкелді. Бір мысал – 1980 жылдардың соңында Роберт Виленскинің Беркли университетінде әзірлеген Unix Consultant (UC). Жүйе Unix операциялық жүйесіне қатысты сұрақтарға жауап берді. Ол өз саласының жан-жақты, қолмен жасалған білім базасына ие болды және әртүрлі пайдаланушыларға жауап беру үшін жауапты бейімдеуге бағытталды. Тағы бір жоба – LILOG, Германиядағы бір қаладағы туристік ақпарат саласында жұмыс істейтін мәтінді түсіндіру жүйесі. UC және LILOG жобаларында әзірленген жүйелер қарапайым демонстрация кезеңінен ешқашан шықпады, бірақ олар есептеу лингвистикасы және ойлау теориясының дамуына көмектесті. Денсаулық сақтау және өмірлік ғылымдар үшін EAGLi сияқты табиғи тілмен сұрақтарға жауап берудің арнайы жүйелері жасалды.
Two early question answering systems were BASEBALL and LUNAR. BASEBALL answered questions about Major League Baseball over a period of one year. LUNAR answered questions about the geological analysis of rocks returned by the Apollo Moon missions. Both question answering systems were very effective in their chosen domains. LUNAR was demonstrated at a lunar science convention in 1971 and it was able to answer 90% of the questions in its domain that were posed by people untrained on the system. Further restricted domain question answering systems were developed in the following years. The common feature of all these systems is that they had a core database or knowledge system that was hand written by experts of the chosen domain. The language abilities of BASEBALL and LUNAR used techniques similar to ELIZA and DOCTOR, the first chatterbot programs. SHRDLU was a successful question answering program developed by Terry Winograd in the late 1960s and early 1970s. It simulated the operation of a robot in a toy world (the "blocks world"), and it offered the possibility of asking the robot questions about the state of the world. The strength of this system was the choice of a very specific domain and a very simple world with rules of physics that were easy to encode in a computer program. In the 1970s, knowledge bases were developed that targeted narrower domains of knowledge. The question answering systems developed to interface with these expert systems produced and valid responses to questions within an area of knowledge. These expert systems closely resembled modern question answering systems except in their internal architecture. Expert systems rely heavily on expert constructed and organized knowledge bases, whereas many modern question answering systems rely on statistical processing of a large, unstructured, natural language text corpus. The 1970s and 1980s saw the development of comprehensive theories in computational linguistics, which led to the development of ambitious projects in text comprehension and question answering. One example was the Unix Consultant (UC), developed by Robert Wilensky at U. C. Berkeley in the late 1980s. The system answered questions pertaining to the Unix operating system. It had a comprehensive, hand crafted knowledge base of its domain, and it aimed at phrasing the answer to accommodate various types of users. Another project was LILOG, a text understanding system that operated on the domain of tourism information in a German city. The systems developed in the UC and LILOG projects never went past the stage of simple demonstrations, but they helped the development of theories on computational linguistics and reasoning. Specialized natural language question answering systems have been developed, such as EAGLi for health and life scientists.
Сәулет
2001 жылдан бастап сұрақ-жауап жүйелері әдетте сұрақтың түрін және жауаптың түрін анықтайтын сұрақ жіктегіш модулін қамтиды. Сұрақ-жауап жүйелерінің әртүрлі түрлері әртүрлі архитектураларды қолданады. Мысалы, қазіргі заманғы ашық домендегі сұрақтарға жауап беру жүйелері ретривер-оқырман архитектурасын пайдалануы мүмкін. Ретривердің мақсаты – берілген сұраққа қатысты тиісті құжаттарды іздеп табу, ал оқырман ізделіп алынған құжаттардан жауапты анықтау үшін қолданылады. GPT 3, T5 және BART сияқты жүйелер трансформаторға негізделген архитектурадағы үлкен көлемдегі мәтіндік деректерді негізгі параметрлерде сақтайтын, аяғынан аяғына дейінгі архитектураны қолданады. Мұндай модельдер сыртқы білім көздерін пайдаланбастан сұрақтарға жауап бере алады.
as of 2001, question answering systems typically included a question classifier module that determined the type of question and the type of answer. Different types of question answering systems employ different architectures. For example, modern open domain question answering systems may use a retriever reader architecture. The retriever is aimed at retrieving relevant documents related to a given question, while the reader is used to infer the answer from the retrieved documents. Systems such as GPT 3, T5, and BART use an end to end architecture in which a transformer based architecture stores large scale textual data in the underlying parameters. Such models can answer questions without accessing any external knowledge sources.
Сұрақ-жауап әдістері
Сұрақтарға жауап беру үшін сапалы іздеу корпусы қажет; жауаптары бар құжаттар болмаса, ешқандай сұрақ-жауап жүйесі көп нәрсе істей алмайды. Көлемді жинақтар, әдетте, сұрақтарға жауап берудің тиімділігін арттырады, егер сұрақтың саласы жинақтан ерекшеленбесе. Веб сияқты үлкен жинақтардағы деректердің қайталануы ақпараттың әртүрлі контекстерде және құжаттарда түрліше айтылуы мүмкін екенін көрсетеді, бұл екі артықшылыққа алып келеді:
Question answering is dependent on a good search corpus; without documents containing the answer, there is little any question answering system can do. Larger collections generally mean better question answering performance, unless the question domain is orthogonal to the collection. Data redundancy in massive collections, such as the web, means that nuggets of information are likely to be phrased in many different ways in differing contexts and documents, leading to two benefits:
Егер дұрыс ақпарат түрлі форматта ұсынылса, сұрақ-жауап жүйесі мәтінді түсіну үшін аз күрделі тілдік өңдеу (NLP) техникаларын қолдануға тура келеді. Дұрыс жауаптарды қате жауаптардан бөліп алуға болады, себебі жүйе дұрыс жауаптың корпустағы дұрыс емес жауаптардан көбірек кездесетін нұсқаларына сүйене алады. Кейбір сұрақ-жауап жүйелері автоматты қорытуға көп мөлшерде сенеді.
If the right information appears in many forms, the question answering system needs to perform fewer complex NLP techniques to understand the text. Correct answers can be filtered from false positives because the system can rely on versions of the correct answer appearing more times in the corpus than incorrect ones. Some question answering systems rely heavily on automated reasoning.
Ашық домендегі сұрақ-жауап
Ақпаратты іздеуде ашық доменді сұрақ-жауап жүйесі пайдаланушының сұрағына жауап беруге тырысады. Берілген жауап тиісті құжаттар тізімі емес, қысқа мәтін түрінде болады. Жүйе есептік лингвистика, ақпаратты іздеу және білімді ұсыну әдістерін үйлестіре отырып жауаптарды табады. Жүйе кілт сөздер жиынтығының орнына табиғи тілдегі сұрақты қабылдайды, мысалы: «Қытайдың ұлттық күні қашан?» Содан кейін жүйе осы кіріс сөйлемді логикалық формадағы сұранысқа түрлендіреді. Табиғи тілдегі сұрақтарды қабылдау жүйені пайдаланушыға ыңғайлырақ етеді, бірақ оны іске асыру қиын, себебі сұрақтардың түрлері әртүрлі және жүйе дұрыс жауап беру үшін дұрыс сұрақ түрін анықтауы керек. Сұраққа сұрақ түрін тағайындау – маңызды міндет; жауапты табу процесінің өзі дұрыс сұрақ түрін және одан кейін дұрыс жауап түрін анықтауға негізделген. Кілт сөздерді іздеу – кіріс сұрақтың түрін анықтаудың алғашқы қадамы. Кейбір жағдайларда сөздер сұрақ түрін анық көрсетеді, мысалы, «Кім», «Қайда», «Қашан» немесе «Қанша» – бұл сөздер жүйеге жауаптар тиісінше «Адам», «Орын», «Күн» немесе «Сан» түрінде болуы керек екенін ұғындыруы мүмкін. POS (сөз табының) белгілеу және синтаксистік талдау әдістері де жауап түрін анықтай алады. Жоғарыдағы мысалда тақырып – «Қытайдың ұлттық күні», предикат – «болу», ал анықтауыш – «қашан», демек, жауап түрі – «Күн». Өкінішке орай, «Қайсы», «Не» немесе «Қалай» сияқты кейбір сұрақтар нақты жауап түрлеріне сәйкес келмейді: әрқайсысы бірнеше түрді білдіре алады. Мұндай жағдайларда сұраққа басқа сөздерді де қарастыру қажет. Контексті түсіну үшін WordNet сияқты лексикалық сөздік қолданылуы мүмкін. Жүйе сұрақтың түрін анықтағаннан кейін, дұрыс кілт сөздерді қамтитын құжаттар жиынтығын табу үшін ақпаратты іздеу жүйесін пайдаланады. Тэггер және NP/Verb Group chunker табылған құжаттарда дұрыс объектілер мен қатынастардың көрсетілгенін тексереді. «Кім» немесе «Қайда» сияқты сұрақтар үшін атаулы тұлғаны танушы алынған құжаттардан тиісті «Адам» және «Орын» атауларын табады. Векторлық кеңістік моделі кандидат жауаптарды жіктеуге мүмкіндік береді. Сұрақ түрін талдау кезеңінде анықталғандай, жауаптың дұрыс түрін тексеріңіз. Дәйектеу әдісі кандидат жауаптарды растай алады. Содан кейін әр кандидатқа сұрақтағы сөздердің санына және олардың кандидатқа қаншалықты жақын екеніне қарай баға беріледі – қанша көп және жақын болса, соғұрлым жақсы. Жауап содан кейін талдау арқылы ықшам және мағыналы түрге аударылады. Алдыңғы мысалда күтілетін жауап – «1 қазан».
In information retrieval, an open domain question answering system tries to return an answer in response to the user's question. The returned answer is in the form of short texts rather than a list of relevant documents. The system finds answers by using a combination of techniques from computational linguistics, information retrieval, and knowledge representation. The system takes a natural language question as an input rather than a set of keywords, for example: "When is the national day of China?" It then transforms this input sentence into a query in its logical form. Accepting natural language questions makes the system more user friendly, but harder to implement, as there are a variety of question types and the system will have to identify the correct one in order to give a sensible answer. Assigning a question type to the question is a crucial task; the entire answer extraction process relies on finding the correct question type and hence the correct answer type. Keyword extraction is the first step in identifying the input question type. In some cases, words clearly indicate the question type, e. g., "Who", "Where", "When", or "How many"—these words might suggest to the system that the answers should be of type "Person", "Location", "Date", or "Number", respectively. POS (part of speech) tagging and syntactic parsing techniques can also determine the answer type. In the example above, the subject is "Chinese National Day", the predicate is "is" and the adverbial modifier is "when", therefore the answer type is "Date". Unfortunately, some interrogative words like "Which", "What", or "How" do not correspond to unambiguous answer types: Each can represent more than one type. In situations like this, other words in the question need to be considered. A lexical dictionary such as WordNet can be used for understanding the context. Once the system identifies the question type, it uses an information retrieval system to find a set of documents that contain the correct keywords. A tagger and NP/Verb Group chunker can verify whether the correct entities and relations are mentioned in the found documents. For questions such as "Who" or "Where", a named entity recogniser finds relevant "Person" and "Location" names from the retrieved documents. A vector space model can classify the candidate answers. Check if the answer is of the correct type as determined in the question type analysis stage. An inference technique can validate the candidate answers. A score is then given to each of these candidates according to the number of question words it contains and how close these words are to the candidate—the more and the closer the better. The answer is then translated by parsing into a compact and meaningful representation. In the previous example, the expected output answer is "1st Oct."
Математикалық сұрақтарға жауап беру
2018 жылы Ask Platypus және Wikidata негізінде құрылған, ашық дереккөзді, математиканы түсінетін MathQA сұрақ-жауап жүйесі жарияланды. MathQA ағылшын немесе хинди тіліндегі сұрақты қабылдап, Wikidata-дан алынған математикалық формуланы қысқа жауап ретінде қайтарады. Бұл формула пайдаланушыға айнымалылардың мәндерін енгізуге мүмкіндік беретін есептеу нысанына аударылады. Жүйе, егер қолжетімді болса, айнымалылардың және жиі қолданылатын тұрақтылардың атауларын және мәндерін Wikidata-дан алады. Мақұлданған мәліметтерге сәйкес, жүйе коммерциялық есептеу математикалық білімдер базасындағы тест жиынтығында жақсы нәтижелер көрсетеді. MathQA Wikimedia ұйымында https://mathqa.wmflabs.org/ мекенжайында орналасқан. 2022 жылы жүйе 15 түрлі математикалық сұраққа жауап беруге мүмкіндік алды. MathQA әдістері табиғи тіл мен формула тілін үйлестіруі тиіс. Мүмкін болатын тәсілдердің бірі – Entity Linking арқылы бақыланатын түсіндірулер жасау. CLEF 2020 жарысындағы "ARQMath Task" Math Stack Exchange платформасынан жаңадан қойылған сұрақтарды, қауымдастық жауап берген бұрынғы сұрақтармен байланыстыру мақсатымен ұйымдастырылды. Бұрын жауапталған, семантикалық жақын сұрақтарға гиперсілтемелер беру пайдаланушыларға тезірек жауап алуға көмектеседі, бірақ семантикалық байланыстылықтың анықталуы қиындық тудырады. Зертхананың бастамасы математикалық сұраныстардың 20%-ы жалпы мақсаттағы іздеу жүйелерінде дұрыс құрылған сұрақтар түрінде берілетінімен байланысты. Бұл сынақ екі бөлек тапсырманы қамтиды: 1-тапсырма – "Жауапты іздеу" (жаңа сұрақтарға ескі жазбалардың жауаптарын сәйкестендіру), 2-тапсырма – "Формуланы іздеу" (жаңа сұрақтарға ескі жазбалардың формулаларын сәйкестендіру). Математика саласынан бастап, формула тілін қолдана отырып, мақсат – кейіннен бұл тапсырманы басқа салаларға (мысалы, химия, биология сияқты STEM пәндеріне) кеңейту, оларда басқа арнайы жазу жүйелері қолданылады (мысалы, химиялық формулалар). Содан кейін формулалар формуланың әртүрлі нұсқаларын жасау үшін қайта құрылады. Айнымалылар кездейсоқ мәндермен алмастырылып, жеке студенттерге арналған тесттерге қолайлы көптеген әртүрлі сұрақтар жасалады. PhysWikiquiz Wikimedia ұйымында https://physwikiquiz.wmflabs.org/ мекенжайында орналасқан.
An open source, math aware, question answering system called MathQA, based on Ask Platypus and Wikidata, was published in 2018. MathQA takes an English or Hindi natural language question as input and returns a mathematical formula retrieved from Wikidata as a succinct answer, translated into a computable form that allows the user to insert values for the variables. The system retrieves names and values of variables and common constants from Wikidata if those are available. It is claimed that the system outperforms a commercial computational mathematical knowledge engine on a test set. MathQA is hosted by Wikimedia at https://mathqa. wmflabs. org/. In 2022, it was extended to answer 15 math question types. MathQA methods need to combine natural and formula language. One possible approach is to perform supervised annotation via Entity Linking. The "ARQMath Task" at CLEF 2020 was launched to address the problem of linking newly posted questions from the platform Math Stack Exchange to existing ones that were already answered by the community. Providing hyperlinks to already answered, semantically related questions helps users to get answers earlier but is a challenging problem because semantic relatedness is not trivial. The lab was motivated by the fact that 20% of mathematical queries in general purpose search engines are expressed as well formed questions. The challenge contained two separate sub tasks. Task 1: "Answer retrieval" matching old post answers to newly posed questions, and Task 2: "Formula retrieval" matching old post formulae to new questions. Starting with the domain of mathematics, which involves formula language, the goal is to later extend the task to other domains (e. g., STEM disciplines, such as chemistry, biology, etc. ), which employ other types of special notation (e. g., chemical formulae). The formulae are then rearranged to generate a set of formula variants. Subsequently, the variables are substituted with random values to generate a large number of different questions suitable for individual student tests. PhysWikiquiz is hosted by Wikimedia at https://physwikiquiz. wmflabs. org/.