Кіріспе
Ықтималдық үлестірімі және гамма-үлестірімінің ерекше жағдайы – хи-квадратты үлестірімінің математикасы. Ықтималдық теориясы мен статистикада, еркіндік дәрежесі бар хи-квадратты үлестірімі (сондай-ақ хи-квадрат немесе үлестірім) тәуелсіз стандартты қалыпты кездейсоқ шамалардың квадраттарының қосындысының үлестірімі болып табылады. Хи-квадратты үлестірімі – гамма-үлестірімінің ерекше жағдайы және ол қорытынды статистикада, әсіресе гипотезаларды тексеруде және сенімділік интервалдарын құруда кеңінен қолданылатын ықтималдық үлестірімдерінің бірі болып табылады. Бұл үлестірім кейде орталық хи-квадратты үлестірім деп аталады, ол жалпы орталық емес хи-квадратты үлестірімінің ерекше жағдайы. Хи-квадратты үлестірімі, байқалған үлестірудің теориялық үлестіруге сәйкестігін, сапалық деректерді жіктеудің екі өлшемінің тәуелсіздігін және үлгідегі стандартты ауытқудан қалыпты үлестірудің популяциялық стандартты ауытқуын бағалау үшін сенімділік интервалын табу үшін қолданылатын хи-квадраттық сынақтарда қолданылады. Көптеген басқа статистикалық сынақтар да осы үлестіруді пайдаланады, мысалы, Фридманның разряд бойынша дисперсиялық талдауы.
the mathematics of the chi squared distribution
In probability theory and statistics, the chi squared distribution (also chi square or distribution) with degrees of freedom is the distribution of a sum of the squares of independent standard normal random variables. The chi squared distribution is a special case of the gamma distribution and is one of the most widely used probability distributions in inferential statistics, notably in hypothesis testing and in construction of confidence intervals. This distribution is sometimes called the central chi squared distribution, a special case of the more general noncentral chi squared distribution. The chi squared distribution is used in the common chi squared tests for goodness of fit of an observed distribution to a theoretical one, the independence of two criteria of classification of qualitative data, and in finding the confidence interval for estimating the population standard deviation of a normal distribution from a sample standard deviation. Many other statistical tests also use this distribution, such as Friedman's analysis of variance by ranks.
Анықтамалар
Егер Z1, ..., Zk тәуелсіз, стандартты қалыпты кездейсоқ шамалар болса, онда олардың квадраттарының қосындысы k еркіндік дәрежесімен χ² (хи-квадрат) үлестіріміне сәйкес келеді. Бұл әдетте былай белгіленеді:
χ² (k)
Хи-квадрат үлестірімінің бір параметрі бар: k – еркіндік дәрежесінің санын анықтайтын оң бүтін сан (қосылатын кездейсоқ шамалардың саны, Zi).
Кіріспе
Хи-квадратты үлестірімі негізінен гипотезаларды тексеруде, ал аздаған дәрежеде негізгі үлестірімі қалыпты болған кезде популяциялық дисперсияның сенімділік интервалдарын табу үшін қолданылады. Қалыпты және экспоненциалды үлестірімдер сияқты кеңінен танымал үлестірімдерден өзгеше, хи-квадратты үлестірім табиғи құбылыстарды тікелей модельдеуде жиі қолданылмайды. Ол келесі гипотезалық сынақтарда қолданылады:
Chi squared test of independence in contingency tables
Chi squared test of goodness of fit of observed data to hypothetical distributions
Likelihood ratio test for nested models
Log rank test in survival analysis
Cochran–Mantel–Haenszel test for stratified contingency tables
Wald test
Score test
It is also a component of the definition of the t distribution and the F distribution used in t tests, analysis of variance, and regression analysis. The primary reason for which the chi squared distribution is extensively used in hypothesis testing is its relationship to the normal distribution. Many hypothesis tests use a test statistic, such as the t statistic in a t test. For these hypothesis tests, as the sample size, n, increases, the sampling distribution of the test statistic approaches the normal distribution (central limit theorem). Because the test statistic (such as t) is asymptotically normally distributed, provided the sample size is sufficiently large, the distribution used for hypothesis testing may be approximated by a normal distribution. Testing hypotheses using a normal distribution is well understood and relatively easy. The simplest chi squared distribution is the square of a standard normal distribution. So wherever a normal distribution could be used for a hypothesis test, a chi squared distribution could be used. Suppose that is a random variable sampled from the standard normal distribution, where the mean is and the variance is : Now, consider the random variable The distribution of the random variable is an example of a chi squared distribution: The subscript 1 indicates that this particular chi squared distribution is constructed from only 1 standard normal distribution. A chi squared distribution constructed by squaring a single standard normal distribution is said to have 1 degree of freedom. Thus, as the sample size for a hypothesis test increases, the distribution of the test statistic approaches a normal distribution. Just as extreme values of the normal distribution have low probability (and give small p values), extreme values of the chi squared distribution have low probability. An additional reason that the chi squared distribution is widely used is that it turns up as the large sample distribution of generalized likelihood ratio tests (LRT). LRTs have several desirable properties; in particular, simple LRTs commonly provide the highest power to reject the null hypothesis (Neyman–Pearson lemma) and this leads also to optimality properties of generalised LRTs. However, the normal and chi squared approximations are only valid asymptotically. For this reason, it is preferable to use the t distribution rather than the normal approximation or the chi squared approximation for a small sample size. Similarly, in analyses of contingency tables, the chi squared approximation will be poor for a small sample size, and it is preferable to use Fisher's exact test. Ramsey shows that the exact binomial test is always more powerful than the normal approximation. Lancaster shows the connections among the binomial, normal, and chi squared distributions, as follows. De Moivre and Laplace established that a binomial distribution could be approximated by a normal distribution. Specifically they showed the asymptotic normality of the random variable
where is the observed number of successes in trials, where the probability of success is , and
Squaring both sides of the equation gives
Using , , and , this equation can be rewritten as
The expression on the right is of the form that Karl Pearson would generalize to the form
where
= Pearson's cumulative test statistic, which asymptotically approaches a distribution;
= the number of observations of type ;
= the expected (theoretical) frequency of type , asserted by the null hypothesis that the fraction of type in the population is ; and
= the number of cells in the table. In the case of a binomial outcome (flipping a coin), the binomial distribution may be approximated by a normal distribution (for sufficiently large ). Because the square of a standard normal distribution is the chi squared distribution with one degree of freedom, the probability of a result such as 1 heads in 10 trials can be approximated either by using the normal distribution directly, or the chi squared distribution for the normalised, squared difference between observed and expected value. However, many problems involve more than the two possible outcomes of a binomial, and instead require 3 or more categories, which leads to the multinomial distribution. Just as de Moivre and Laplace sought for and found the normal approximation to the binomial, Pearson sought for and found a degenerate multivariate normal approximation to the multinomial distribution (the numbers in each category add up to the total sample size, which is considered fixed). Pearson showed that the chi squared distribution arose from such a multivariate normal approximation to the multinomial distribution, taking careful account of the statistical dependence (negative correlations) between numbers of observations in different categories. For the cases when (which include all of the cases when this CDF is less than half):
The tail bound for the cases when , similarly, is
For another approximation for the CDF modeled after the cube of a Gaussian, see under Noncentral chi squared distribution.
* Кестелік кестелерде тәуелсіздіктің хи-квадраттық сынағы
* Гипотетикалық үлестірімдерге сәйкес келетін байқалатын деректердің жақсылығын бағалау сынағы
* Ұялы модельдер үшін ықтималдық қатынасы сынағы
* Тіршілік ету талдауындағы логарифмдік ранг сынағы
* Қабатталған кестелік кестелер үшін Кохран-Мантель-Хензель сынағы
* Вальд сынағы
* Скор сынағы
Chi squared test of independence in contingency tables
Chi squared test of goodness of fit of observed data to hypothetical distributions
Likelihood ratio test for nested models
Log rank test in survival analysis
Cochran–Mantel–Haenszel test for stratified contingency tables
Wald test
Score test
It is also a component of the definition of the t distribution and the F distribution used in t tests, analysis of variance, and regression analysis. The primary reason for which the chi squared distribution is extensively used in hypothesis testing is its relationship to the normal distribution. Many hypothesis tests use a test statistic, such as the t statistic in a t test. For these hypothesis tests, as the sample size, n, increases, the sampling distribution of the test statistic approaches the normal distribution (central limit theorem). Because the test statistic (such as t) is asymptotically normally distributed, provided the sample size is sufficiently large, the distribution used for hypothesis testing may be approximated by a normal distribution. Testing hypotheses using a normal distribution is well understood and relatively easy. The simplest chi squared distribution is the square of a standard normal distribution. So wherever a normal distribution could be used for a hypothesis test, a chi squared distribution could be used. Suppose that is a random variable sampled from the standard normal distribution, where the mean is and the variance is : Now, consider the random variable The distribution of the random variable is an example of a chi squared distribution: The subscript 1 indicates that this particular chi squared distribution is constructed from only 1 standard normal distribution. A chi squared distribution constructed by squaring a single standard normal distribution is said to have 1 degree of freedom. Thus, as the sample size for a hypothesis test increases, the distribution of the test statistic approaches a normal distribution. Just as extreme values of the normal distribution have low probability (and give small p values), extreme values of the chi squared distribution have low probability. An additional reason that the chi squared distribution is widely used is that it turns up as the large sample distribution of generalized likelihood ratio tests (LRT). LRTs have several desirable properties; in particular, simple LRTs commonly provide the highest power to reject the null hypothesis (Neyman–Pearson lemma) and this leads also to optimality properties of generalised LRTs. However, the normal and chi squared approximations are only valid asymptotically. For this reason, it is preferable to use the t distribution rather than the normal approximation or the chi squared approximation for a small sample size. Similarly, in analyses of contingency tables, the chi squared approximation will be poor for a small sample size, and it is preferable to use Fisher's exact test. Ramsey shows that the exact binomial test is always more powerful than the normal approximation. Lancaster shows the connections among the binomial, normal, and chi squared distributions, as follows. De Moivre and Laplace established that a binomial distribution could be approximated by a normal distribution. Specifically they showed the asymptotic normality of the random variable
where is the observed number of successes in trials, where the probability of success is , and
Squaring both sides of the equation gives
Using , , and , this equation can be rewritten as
The expression on the right is of the form that Karl Pearson would generalize to the form
where
= Pearson's cumulative test statistic, which asymptotically approaches a distribution;
= the number of observations of type ;
= the expected (theoretical) frequency of type , asserted by the null hypothesis that the fraction of type in the population is ; and
= the number of cells in the table. In the case of a binomial outcome (flipping a coin), the binomial distribution may be approximated by a normal distribution (for sufficiently large ). Because the square of a standard normal distribution is the chi squared distribution with one degree of freedom, the probability of a result such as 1 heads in 10 trials can be approximated either by using the normal distribution directly, or the chi squared distribution for the normalised, squared difference between observed and expected value. However, many problems involve more than the two possible outcomes of a binomial, and instead require 3 or more categories, which leads to the multinomial distribution. Just as de Moivre and Laplace sought for and found the normal approximation to the binomial, Pearson sought for and found a degenerate multivariate normal approximation to the multinomial distribution (the numbers in each category add up to the total sample size, which is considered fixed). Pearson showed that the chi squared distribution arose from such a multivariate normal approximation to the multinomial distribution, taking careful account of the statistical dependence (negative correlations) between numbers of observations in different categories. For the cases when (which include all of the cases when this CDF is less than half):
The tail bound for the cases when , similarly, is
For another approximation for the CDF modeled after the cube of a Gaussian, see under Noncentral chi squared distribution.
Сонымен қатар, ол t-тестілерде, дисперсиялық талдауда және регрессиялық талдауда қолданылатын t- және F-үлестірімдерінің анықтамасының құрауышысы болып табылады. Хи-квадратты үлестірімінің гипотезаларды тексеруде кеңінен қолданылуының басты себебі – оның қалыпты үлестірумен байланысы. Көптеген гипотезалық сынақтарда t-тестіндегі t-статистикасы сияқты сынақ статистикасы қолданылады. Бұл гипотезалық сынақтар үшін, үлгінің көлемі n артқан сайын, сынақ статистикасының үлестірілуі орталық шек теоремасына сәйкес қалыпты үлестіруге жақындайды. Сынақ статистикасы (мысалы, t) асимптотикалық түрде қалыпты үлестірілгендіктен, егер үлгінің көлемі жеткілікті үлкен болса, гипотезаларды тексеру үшін қолданылатын үлестіруді қалыпты үлестірумен жуықтауға болады. Қалыпты үлестіруді пайдалана отырып гипотезаларды тексеру жақсы түсініледі және салыстырмалы түрде оңай. Ең қарапайым хи-квадратты үлестірімі стандартты қалыпты үлестірудің квадраты болып табылады. Сондықтан, гипотезалық сынақ үшін қалыпты үлестіруді қолдануға болатын жағдайда, хи-квадратты үлестіруді де қолдануға болады. Стандартты қалыпты үлестіруден алынған кездейсоқ айнымалыны қарастырайық, онда орташа мәні және дисперсиясы : Енді кездейсоқ айнымалыны қарастырайық. Кездейсоқ айнымалының үлестірілуі хи-квадратты үлестірудің мысалы болып табылады: 1 индексі осы хи-квадратты үлестірудің тек 1 стандартты қалыпты үлестіруден құралғандығын көрсетеді. Бір стандартты қалыпты үлестіруді квадраттау арқылы құрастырылған хи-квадратты үлестіру 1 еркіндік дәрежесіне ие болады. Осылайша, гипотезалық сынақтың үлгісінің көлемі артқан сайын, сынақ статистикасының үлестірілуі қалыпты үлестіруге жақындайды. Қалыпты үлестірудің шекті мәндерінің төмен ықтималдылығы бар (және кішкентай p-мәндерін береді), хи-квадратты үлестірудің де шекті мәндерінің төмен ықтималдылығы бар. Хи-квадратты үлестірімінің кеңінен қолданылуының қосымша себебі – оның жалпыланған ықтималдық қатынасы сынақтарының (LRT) үлкен үлгілік үлестірімі ретінде пайда болуы. LRT-лерде бірнеше қасиеттер бар; атап айтқанда, қарапайым LRT-лер нөлдік гипотезаны қабылдамау үшін ең жоғары қуатты қамтамасыз етеді (Нейман-Пирсон леммасы), бұл жалпыланған LRT-лердің оптималдық қасиеттеріне әкеледі. Алайда, қалыпты және хи-квадратты жуықтаулар тек асимптотикалық түрде ғана жарамды. Осы себепті, үлгінің көлемі кішкентай болса, қалыпты жуықтау немесе хи-квадратты жуықтау емес, t-үлестірімін қолдану дұрыс. Сол сияқты, кестелік кестелерді талдау кезінде кішкентай үлгі үшін хи-квадратты жуықтау нашар болады, сондықтан Фишердің нақты сынағын қолдану дұрыс. Рамзи нақты биномдық тесттің әрқашан қалыпты жуықтаудан гөрі күшті екенін көрсетеді. Ланкастер биномдық, қалыпты және хи-квадратты үлестірімдер арасындағы байланыстарды төмендегідей көрсетеді. Де Мойвр мен Лаплас биномдық үлестіруді қалыпты үлестірумен жуықтауға болатынын анықтады. Нақты айтқанда, олар кездейсоқ айнымалының асимптотикалық қалыптылығын көрсетті, мұнда сынақтардағы сәттіліктердің байқалатын саны, сәттілік ықтималдығы және теңдеудің екі жағын квадраттау арқылы теңдеуді қайта жазуға болады. Оң жақтағы өрнекті Карл Пирсон жалпылайтын формада = Пирсонның жиынтық тест статистикасы, ол асимптотикалық түрде үлестіруге жақындайды; = типті байқаулардың саны; = типтің күтілетін (теориялық) жиілігі, популяцияда типтің үлесі; және = кестедегі жасушалардың саны деген нөлдік гипотезамен мәлімделеді. Биномдық нәтиже жағдайында (сатымен ақшаны лақтырып тастау) биномдық үлестіруді қалыпты үлестірумен (жеткілікті үлкен үшін) жуықтауға болады. Стандартты қалыпты үлестірудің квадраты бір еркіндік дәрежесі бар хи-квадратты үлестіру болып табылады, сондықтан 10 сынақта 1 бас нәтижесінің ықтималдығын қалыпты үлестіруді тікелей пайдаланып немесе байқалатын және күтілетін мәндер арасындағы нормаланған, квадратталған айырмашылық үшін хи-квадратты үлестіруді пайдаланып жуықтауға болады. Алайда, көптеген мәселелерде биномдық екі мүмкін нәтижеден гөрі көбірек нәтижелер бар, және олар 3 немесе одан көп санаттарды қажет етеді, бұл көпмүшелік үлестіруге әкеледі. Де Мойвр мен Лаплас биномдыққа қалыпты жуықтауды іздеп тапқандай, Пирсон көпмүшелік үлестіруге дегенеративті көпөлшемді қалыпты жуықтауды іздеп тапты (әр санаттағы сандардың жалпы үлгі көлеміне қосылуы, ол тұрақты деп есептеледі). Пирсон хи-квадратты үлестіруінің көпмүшелік үлестіруге көпөлшемді қалыпты жуықтаудан туындайтынын көрсетті, әртүрлі санаттардағы байқаулар саны арасындағы статистикалық тәуелділікті (теріс корреляцияларды) ескере отырып. жағдайлар үшін (бұл CDF-нің жартысынан кем болған барлық жағдайларды қамтиды):
Chi squared test of independence in contingency tables
Chi squared test of goodness of fit of observed data to hypothetical distributions
Likelihood ratio test for nested models
Log rank test in survival analysis
Cochran–Mantel–Haenszel test for stratified contingency tables
Wald test
Score test
It is also a component of the definition of the t distribution and the F distribution used in t tests, analysis of variance, and regression analysis. The primary reason for which the chi squared distribution is extensively used in hypothesis testing is its relationship to the normal distribution. Many hypothesis tests use a test statistic, such as the t statistic in a t test. For these hypothesis tests, as the sample size, n, increases, the sampling distribution of the test statistic approaches the normal distribution (central limit theorem). Because the test statistic (such as t) is asymptotically normally distributed, provided the sample size is sufficiently large, the distribution used for hypothesis testing may be approximated by a normal distribution. Testing hypotheses using a normal distribution is well understood and relatively easy. The simplest chi squared distribution is the square of a standard normal distribution. So wherever a normal distribution could be used for a hypothesis test, a chi squared distribution could be used. Suppose that is a random variable sampled from the standard normal distribution, where the mean is and the variance is : Now, consider the random variable The distribution of the random variable is an example of a chi squared distribution: The subscript 1 indicates that this particular chi squared distribution is constructed from only 1 standard normal distribution. A chi squared distribution constructed by squaring a single standard normal distribution is said to have 1 degree of freedom. Thus, as the sample size for a hypothesis test increases, the distribution of the test statistic approaches a normal distribution. Just as extreme values of the normal distribution have low probability (and give small p values), extreme values of the chi squared distribution have low probability. An additional reason that the chi squared distribution is widely used is that it turns up as the large sample distribution of generalized likelihood ratio tests (LRT). LRTs have several desirable properties; in particular, simple LRTs commonly provide the highest power to reject the null hypothesis (Neyman–Pearson lemma) and this leads also to optimality properties of generalised LRTs. However, the normal and chi squared approximations are only valid asymptotically. For this reason, it is preferable to use the t distribution rather than the normal approximation or the chi squared approximation for a small sample size. Similarly, in analyses of contingency tables, the chi squared approximation will be poor for a small sample size, and it is preferable to use Fisher's exact test. Ramsey shows that the exact binomial test is always more powerful than the normal approximation. Lancaster shows the connections among the binomial, normal, and chi squared distributions, as follows. De Moivre and Laplace established that a binomial distribution could be approximated by a normal distribution. Specifically they showed the asymptotic normality of the random variable
where is the observed number of successes in trials, where the probability of success is , and
Squaring both sides of the equation gives
Using , , and , this equation can be rewritten as
The expression on the right is of the form that Karl Pearson would generalize to the form
where
= Pearson's cumulative test statistic, which asymptotically approaches a distribution;
= the number of observations of type ;
= the expected (theoretical) frequency of type , asserted by the null hypothesis that the fraction of type in the population is ; and
= the number of cells in the table. In the case of a binomial outcome (flipping a coin), the binomial distribution may be approximated by a normal distribution (for sufficiently large ). Because the square of a standard normal distribution is the chi squared distribution with one degree of freedom, the probability of a result such as 1 heads in 10 trials can be approximated either by using the normal distribution directly, or the chi squared distribution for the normalised, squared difference between observed and expected value. However, many problems involve more than the two possible outcomes of a binomial, and instead require 3 or more categories, which leads to the multinomial distribution. Just as de Moivre and Laplace sought for and found the normal approximation to the binomial, Pearson sought for and found a degenerate multivariate normal approximation to the multinomial distribution (the numbers in each category add up to the total sample size, which is considered fixed). Pearson showed that the chi squared distribution arose from such a multivariate normal approximation to the multinomial distribution, taking careful account of the statistical dependence (negative correlations) between numbers of observations in different categories. For the cases when (which include all of the cases when this CDF is less than half):
The tail bound for the cases when , similarly, is
For another approximation for the CDF modeled after the cube of a Gaussian, see under Noncentral chi squared distribution.
Құйрық шегі жағдайлар үшін, сонымен қатар,
Chi squared test of independence in contingency tables
Chi squared test of goodness of fit of observed data to hypothetical distributions
Likelihood ratio test for nested models
Log rank test in survival analysis
Cochran–Mantel–Haenszel test for stratified contingency tables
Wald test
Score test
It is also a component of the definition of the t distribution and the F distribution used in t tests, analysis of variance, and regression analysis. The primary reason for which the chi squared distribution is extensively used in hypothesis testing is its relationship to the normal distribution. Many hypothesis tests use a test statistic, such as the t statistic in a t test. For these hypothesis tests, as the sample size, n, increases, the sampling distribution of the test statistic approaches the normal distribution (central limit theorem). Because the test statistic (such as t) is asymptotically normally distributed, provided the sample size is sufficiently large, the distribution used for hypothesis testing may be approximated by a normal distribution. Testing hypotheses using a normal distribution is well understood and relatively easy. The simplest chi squared distribution is the square of a standard normal distribution. So wherever a normal distribution could be used for a hypothesis test, a chi squared distribution could be used. Suppose that is a random variable sampled from the standard normal distribution, where the mean is and the variance is : Now, consider the random variable The distribution of the random variable is an example of a chi squared distribution: The subscript 1 indicates that this particular chi squared distribution is constructed from only 1 standard normal distribution. A chi squared distribution constructed by squaring a single standard normal distribution is said to have 1 degree of freedom. Thus, as the sample size for a hypothesis test increases, the distribution of the test statistic approaches a normal distribution. Just as extreme values of the normal distribution have low probability (and give small p values), extreme values of the chi squared distribution have low probability. An additional reason that the chi squared distribution is widely used is that it turns up as the large sample distribution of generalized likelihood ratio tests (LRT). LRTs have several desirable properties; in particular, simple LRTs commonly provide the highest power to reject the null hypothesis (Neyman–Pearson lemma) and this leads also to optimality properties of generalised LRTs. However, the normal and chi squared approximations are only valid asymptotically. For this reason, it is preferable to use the t distribution rather than the normal approximation or the chi squared approximation for a small sample size. Similarly, in analyses of contingency tables, the chi squared approximation will be poor for a small sample size, and it is preferable to use Fisher's exact test. Ramsey shows that the exact binomial test is always more powerful than the normal approximation. Lancaster shows the connections among the binomial, normal, and chi squared distributions, as follows. De Moivre and Laplace established that a binomial distribution could be approximated by a normal distribution. Specifically they showed the asymptotic normality of the random variable
where is the observed number of successes in trials, where the probability of success is , and
Squaring both sides of the equation gives
Using , , and , this equation can be rewritten as
The expression on the right is of the form that Karl Pearson would generalize to the form
where
= Pearson's cumulative test statistic, which asymptotically approaches a distribution;
= the number of observations of type ;
= the expected (theoretical) frequency of type , asserted by the null hypothesis that the fraction of type in the population is ; and
= the number of cells in the table. In the case of a binomial outcome (flipping a coin), the binomial distribution may be approximated by a normal distribution (for sufficiently large ). Because the square of a standard normal distribution is the chi squared distribution with one degree of freedom, the probability of a result such as 1 heads in 10 trials can be approximated either by using the normal distribution directly, or the chi squared distribution for the normalised, squared difference between observed and expected value. However, many problems involve more than the two possible outcomes of a binomial, and instead require 3 or more categories, which leads to the multinomial distribution. Just as de Moivre and Laplace sought for and found the normal approximation to the binomial, Pearson sought for and found a degenerate multivariate normal approximation to the multinomial distribution (the numbers in each category add up to the total sample size, which is considered fixed). Pearson showed that the chi squared distribution arose from such a multivariate normal approximation to the multinomial distribution, taking careful account of the statistical dependence (negative correlations) between numbers of observations in different categories. For the cases when (which include all of the cases when this CDF is less than half):
The tail bound for the cases when , similarly, is
For another approximation for the CDF modeled after the cube of a Gaussian, see under Noncentral chi squared distribution.
CDF-ді модельдеу үшін Гаусс кубынан жасалған тағы бір жуықтау үшін, қалыпты емес хи-квадратты үлестіру бөліміне қараңыз.
Chi squared test of independence in contingency tables
Chi squared test of goodness of fit of observed data to hypothetical distributions
Likelihood ratio test for nested models
Log rank test in survival analysis
Cochran–Mantel–Haenszel test for stratified contingency tables
Wald test
Score test
It is also a component of the definition of the t distribution and the F distribution used in t tests, analysis of variance, and regression analysis. The primary reason for which the chi squared distribution is extensively used in hypothesis testing is its relationship to the normal distribution. Many hypothesis tests use a test statistic, such as the t statistic in a t test. For these hypothesis tests, as the sample size, n, increases, the sampling distribution of the test statistic approaches the normal distribution (central limit theorem). Because the test statistic (such as t) is asymptotically normally distributed, provided the sample size is sufficiently large, the distribution used for hypothesis testing may be approximated by a normal distribution. Testing hypotheses using a normal distribution is well understood and relatively easy. The simplest chi squared distribution is the square of a standard normal distribution. So wherever a normal distribution could be used for a hypothesis test, a chi squared distribution could be used. Suppose that is a random variable sampled from the standard normal distribution, where the mean is and the variance is : Now, consider the random variable The distribution of the random variable is an example of a chi squared distribution: The subscript 1 indicates that this particular chi squared distribution is constructed from only 1 standard normal distribution. A chi squared distribution constructed by squaring a single standard normal distribution is said to have 1 degree of freedom. Thus, as the sample size for a hypothesis test increases, the distribution of the test statistic approaches a normal distribution. Just as extreme values of the normal distribution have low probability (and give small p values), extreme values of the chi squared distribution have low probability. An additional reason that the chi squared distribution is widely used is that it turns up as the large sample distribution of generalized likelihood ratio tests (LRT). LRTs have several desirable properties; in particular, simple LRTs commonly provide the highest power to reject the null hypothesis (Neyman–Pearson lemma) and this leads also to optimality properties of generalised LRTs. However, the normal and chi squared approximations are only valid asymptotically. For this reason, it is preferable to use the t distribution rather than the normal approximation or the chi squared approximation for a small sample size. Similarly, in analyses of contingency tables, the chi squared approximation will be poor for a small sample size, and it is preferable to use Fisher's exact test. Ramsey shows that the exact binomial test is always more powerful than the normal approximation. Lancaster shows the connections among the binomial, normal, and chi squared distributions, as follows. De Moivre and Laplace established that a binomial distribution could be approximated by a normal distribution. Specifically they showed the asymptotic normality of the random variable
where is the observed number of successes in trials, where the probability of success is , and
Squaring both sides of the equation gives
Using , , and , this equation can be rewritten as
The expression on the right is of the form that Karl Pearson would generalize to the form
where
= Pearson's cumulative test statistic, which asymptotically approaches a distribution;
= the number of observations of type ;
= the expected (theoretical) frequency of type , asserted by the null hypothesis that the fraction of type in the population is ; and
= the number of cells in the table. In the case of a binomial outcome (flipping a coin), the binomial distribution may be approximated by a normal distribution (for sufficiently large ). Because the square of a standard normal distribution is the chi squared distribution with one degree of freedom, the probability of a result such as 1 heads in 10 trials can be approximated either by using the normal distribution directly, or the chi squared distribution for the normalised, squared difference between observed and expected value. However, many problems involve more than the two possible outcomes of a binomial, and instead require 3 or more categories, which leads to the multinomial distribution. Just as de Moivre and Laplace sought for and found the normal approximation to the binomial, Pearson sought for and found a degenerate multivariate normal approximation to the multinomial distribution (the numbers in each category add up to the total sample size, which is considered fixed). Pearson showed that the chi squared distribution arose from such a multivariate normal approximation to the multinomial distribution, taking careful account of the statistical dependence (negative correlations) between numbers of observations in different categories. For the cases when (which include all of the cases when this CDF is less than half):
The tail bound for the cases when , similarly, is
For another approximation for the CDF modeled after the cube of a Gaussian, see under Noncentral chi squared distribution.
Қосымшалық
Хи-квадратты үлестірудің анықтамасынан тәуелсіз хи-квадратты айнымалылардың қосындысы да хи-квадратты үлестіріледі. Нақтырақ айтқанда, егер тәуелсіз хи-квадратты айнымалылар болса, олардың әрқайсысы сәйкесінше , еркіндік дәрежесіне ие болса, онда хи-квадратты үлестіріледі, оның еркіндік дәрежесі -қа тең болады.
Концентрациясы
Хи-квадратты үлестірімі өзінің орташа мәнінің төңірегінде күшті шоғырлануға ие. Стандартты Лорент Массарт шектері:
Бір салдары – егер - гаусс кездейсоқ векторы болса, онда өлшем артаған сайын вектордың ұзындығының квадраты шамасында тығыз шоғырланады, ал ені : мұнда көрсеткішті аралығындағы кез келген мән ретінде таңдауға болады.
Орталық емес хи-квадраттық үлестірімі
Орталық емес хи-квадратты үлестірімі бірлік дисперсиясы және нөлден өзгеше орташа мәні бар тәуелсіз Гаусс кездейсоқ шамаларының квадраттарының қосындысынан шығарылады.
Жалпыланған хи-квадраттық үлестірімі
Жалпыланған хи-квадратты үлестірімі z'Az квадраттық түрінен шығарылады, мұнда z – кез келген ковариациялық матрицасы бар орташасы нөлдік Гаусс векторы, ал A – кез келген матрица.
Құндылықтар мен құнсыздықтар кестесі
Құны – сынақ статистикасының хи-квадраттық үлестірілімде кем дегенде сондай-ақ экстремалды мәнді байқау ықтималдығы. Сәйкесінше, тиісті еркіндік дәрежесі (df) үшін жинақталған үлестіру функциясы (CDF) осы нүктеден кем экстремалды мән алу ықтималдығын көрсетеді, ал CDF мәнін 1-ден шығару p мәнін береді. Таңдалған маңыздылық деңгейінен төмен p мәні статистикалық маңыздылықты көрсетеді, яғни нөлдік гипотезаны жоққа шығаруға жеткілікті дәлелдер бар екенін білдіреді. 0,05 маңыздылық деңгейі маңызды және маңызсыз нәтижелерді ажырату үшін жиі қолданылады. Төмендегі кесте алғашқы 10 еркіндік дәрежесі үшін сәйкес p мәндерін көрсетеді.
These values can be calculated evaluating the quantile function (also known as "inverse CDF" or "ICDF") of the chi squared distribution; e. g., the χ^(2) ICDF for 1=p = 0.05 and 1=df = 7 yields 2.1673 ≈ 2.17 as in the table above, noticing that 1 – p is the p value from the table.
| Еркіндік дәрежесі (df) | p мәні (ықтималдық) | 0.95 | 0.90 | 0.80 | 0.70 | 0.50 | 0.30 | 0.20 | 0.10 | 0.05 | 0.01 | 0.001 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | 0.004 | 0.02 | 0.06 | 0.15 | 0.46 | 1.07 | 1.64 | 2.71 | 3.84 | 6.63 | 10.83 |
| 2 | 0.10 | 0.21 | 0.45 | 0.71 | 1.39 | 2.41 | 3.22 | 4.61 | 5.99 | 9.21 | 13.82 |
| 3 | 0.35 | 0.58 | 1.01 | 1.42 | 2.37 | 3.66 | 4.64 | 6.25 | 7.81 | 11.34 | 16.27 |
| 4 | 0.71 | 1.06 | 1.65 | 2.20 | 3.36 | 4.88 | 5.99 | 7.78 | 9.49 | 13.28 | 18.47 |
| 5 | 1.14 | 1.61 | 2.34 | 3.00 | 4.35 | 6.06 | 7.29 | 9.24 | 11.07 | 15.09 | 20.52 |
| 6 | 1.63 | 2.20 | 3.07 | 3.83 | 5.35 | 7.23 | 8.56 | 10.64 | 12.59 | 16.81 | 22.46 |
| 7 | 2.17 | 2.83 | 3.82 | 4.67 | 6.35 | 8.38 | 9.80 | 12.02 | 14.07 | 18.48 | 24.32 |
| 8 | 2.73 | 3.49 | 4.59 | 5.53 | 7.34 | 9.52 | 11.03 | 13.36 | 15.51 | 20.09 | 26.12 |
| 9 | 3.32 | 4.17 | 5.38 | 6.39 | 8.34 | 10.66 | 12.24 | 14.68 | 16.92 | 21.67 | 27.88 |
| 10 | 3.94 | 4.87 | 6.18 | 7.27 | 9.34 | 11.78 | 13.44 | 15.99 | 18.31 | 23.21 | 29.59 |
These values can be calculated evaluating the quantile function (also known as "inverse CDF" or "ICDF") of the chi squared distribution; e. g., the χ^(2) ICDF for 1=p = 0.05 and 1=df = 7 yields 2.1673 ≈ 2.17 as in the table above, noticing that 1 – p is the p value from the table.
Бұл мәндерді хи-квадраттық үлестірілімінің квантильдік функциясын (сонымен қатар "кері CDF" немесе "ICDF" деп белгілі) есептеу арқылы табуға болады; мысалы, 1=p = 0.05 және 1=df = 7 үшін χ^(2) ICDF 2.1673 ≈ 2.17 нәтижесін береді, кестедегідей, ескеріңіз, 1 – p кестеден алынған p мәні.
These values can be calculated evaluating the quantile function (also known as "inverse CDF" or "ICDF") of the chi squared distribution; e. g., the χ^(2) ICDF for 1=p = 0.05 and 1=df = 7 yields 2.1673 ≈ 2.17 as in the table above, noticing that 1 – p is the p value from the table.
Тарих
Бұл үлестіруді алғаш рет неміс геодезисі және статистигі Фридрих Роберт Гельмерт 1875–6 жылғы еңбектерінде сипаттады, онда ол қалыпты популяцияның үлгілік дисперсиясының үлгілік үлестіруін есептеп шығарды. Осылайша, неміс тілінде бұл дәстүрлі түрде Helmert'sche ("Helmertian") немесе "Helmert үлестірімі" деп аталды. Үлестіруді ағылшын математигі Карл Пирсон тәуелсіз түрде қайта ашты, ол үшін ол 1900 жылы жарияланған Пирсонның хи-квадраттық сынағын әзірледі, ал есептелген мәндердің кестесі 1900 жылы жинақталды. "Хи-квадрат" атауы ақыр соңында Пирсонның көпөлшемді қалыпты үлестірудегі экспонентаның грек әрпі хи арқылы қысқартуынан шыққан, ол −½χ2 деп жазды, бұл қазіргі нотацияда −½'x'^(T)Σ^(−1)'x' (Σ – ковариациялық матрица) болар еді. Дегенмен, "хи-квадраттық үлестірулер" отбасы идеясы Пирсонға емес, 1920 жылдары Фишердің одан әрі дамытуының нәтижесінде пайда болды.