Ағылшыншамен салыстырыңыз: абзацты басыңыз — түпнұсқа терезеде ашылады. Абзац астындағы EN түймесі оны мәтін ішінде көрсетеді.
Кіріспе
Статистика мен психометриядағы өлшемнің жалпы тұрақтылығы
Overall consistency of a measure in statistics and psychometrics
Статистика мен психометрияда сенімділік – өлшемнің жалпы тұрақтылығы. Егер өлшем тұрақты жағдайларда ұқсас нәтижелер көрсетсе, онда оның сенімділігі жоғары деп есептеледі: "Бұл – тест нәтижелері жиынтығының сипаттамасы, ол нәтижелерге енген өлшеу процесінен туындаған кездейсоқ қателік мөлшерімен байланысты. Жоғары сенімді нәтижелер дәл, қайталанатын және бір сынақтан екіншісіне сәйкес келеді. Яғни, егер сынақ процесі сол топпен қайталанса, негізінен бірдей нәтижелер алынар еді. Нәтижелердегі қателік мөлшерін көрсету үшін әдетте 0,00 (көп қателік) мен 1,00 (қателік жоқ) аралығындағы мәндермен әртүрлі сенімділік коэффициенттері қолданылады. Мысалы, адамдардың бойы мен салмағын өлшеу әдетте өте сенімді болады.
In statistics and psychometrics, reliability is the overall consistency of a measure. A measure is said to have a high reliability if it produces similar results under consistent conditions:"It is the characteristic of a set of test scores that relates to the amount of random error from the measurement process that might be embedded in the scores. Scores that are highly reliable are precise, reproducible, and consistent from one testing occasion to another. That is, if the testing process were repeated with a group of test takers, essentially the same results would be obtained. Various kinds of reliability coefficients, with values ranging between 0.00 (much error) and 1.00 (no error), are usually used to indicate the amount of error in the scores." For example, measurements of people's height and weight are often extremely reliable.
Жалпы үлгі
Іс жүзінде, сынақ шаралары ешқашан толық сәйкес келмейді. Өлшемнің дәлдігіне сәйкессіздіктің әсерін бағалау үшін сынақ сенімділігі теориялары жасалған. Тест сенімділігінің дерлік барлық теорияларының бастапқы негізі – тест нәтижелері екі түрлі фактордың әсерін көрсетеді деген идея:
In practice, testing measures are never perfectly consistent. Theories of test reliability have been developed to estimate the effects of inconsistency on the accuracy of measurement. The basic starting point for almost all theories of test reliability is the idea that test scores reflect the influence of two sorts of factors:
1. Өлшемнің орташа қатесі = 0
1. Mean error of measurement = 0
2. Шынайы нәтижелер мен қателер өзара байланысты емес
2. True scores and errors are uncorrelated
3. Әртүрлі өлшемдердегі қателер өзара байланысты емес
3. Errors on different measures are uncorrelated
Сенімділік теориясы көрсеткендей, алынған нәтижелердің дисперсиясы – шынайы нәтижелердің дисперсиясы мен өлшем қателерінің дисперсиясының қосындысы. Кронбахтың альфасы – ішкі тұрақтылықты бағалаудың бұрынғы түрі, Кудер-Ричардсонның 20-формуласының жалпылануы. Бұл сенімділік өлшемдері әртүрлі қате көздеріне сезімталдықтарымен ерекшеленеді, сондықтан олар міндетті түрде бірдей болуы керек емес. Сондай-ақ, сенімділік – өлшемнің өзінің емес, оның нәтижелерінің қасиеті, демек, ол үлгіге тәуелді болып саналады. Бір үлгіде алынған сенімділік бағалаулары екінші үлгідегіден (егер екінші үлгі басқа популяциядан алынған болса) өзгеше болуы мүмкін, себебі екінші популяциядағы шынайы өзгергіштік басқаша. (Бұл барлық типтегі өлшемдерге қатысты: аула таяқтары үйлерді жақсы өлшесе де, жәндіктердің ұзындығын өлшегенде сенімділігі төмен болуы мүмкін.) Сенімділікті жазбаша бағалауларда сөздің анықтығын арттыру, өлшемді ұзарту және басқа да бейресми жолдармен жақсартуға болады. Дегенмен, формалды психометриялық талдау, яғни элементтік талдау, сенімділікті арттырудың ең тиімді тәсілі саналады. Бұл талдау элементтердің қиындық деңгейін және элементтерді ажырата алу индексін есептеуден тұрады, соңғы индекс элементтер арасындағы корреляцияны және бүкіл тест бойынша элементтердің жиынтық нәтижелерін есептеуді қамтиды. Егер тым қиын, тым оңай немесе нөлге жақын немесе теріс ажырата алуға ие элементтер жақсырақ элементтермен ауыстырылса, өлшемнің сенімділігі артады. мұнда қателік деңгейі көрсетілген.
Reliability theory shows that the variance of obtained scores is simply the sum of the variance of true scores plus the variance of errors of measurement. Cronbach's alpha is a generalization of an earlier form of estimating internal consistency, Kuder–Richardson Formula 20. These measures of reliability differ in their sensitivity to different sources of error and so need not be equal. Also, reliability is a property of the scores of a measure rather than the measure itself and are thus said to be sample dependent. Reliability estimates from one sample might differ from those of a second sample (beyond what might be expected due to sampling variations) if the second sample is drawn from a different population because the true variability is different in this second population. (This is true of measures of all types—yardsticks might measure houses well yet have poor reliability when used to measure the lengths of insects.) Reliability may be improved by clarity of expression (for written assessments), lengthening the measure, and other informal means. However, formal psychometric analysis, called item analysis, is considered the most effective way to increase reliability. This analysis consists of computation of item difficulties and item discrimination indices, the latter index involving computation of correlations between the items and sum of the item scores of the entire test. If items that are too difficult, too easy, and/or have near zero or negative discrimination are replaced with better items, the reliability of the measure will increase. where is the failure rate.