Введение
Распределение вероятностей и частный случай гамма-распределения – математика распределения хи-квадрат.
the mathematics of the chi squared distribution
В теории вероятностей и статистике распределение хи-квадрат (также хи-квадрат или χ²-распределение) с *k* степенями свободы представляет собой распределение суммы квадратов *k* независимых стандартных нормальных случайных величин. Распределение хи-квадрат является частным случаем гамма-распределения и является одним из наиболее широко используемых распределений вероятностей в математической статистике, особенно при проверке гипотез и построении доверительных интервалов. Это распределение иногда называют центральным распределением хи-квадрат, являющимся частным случаем более общего нецентрального распределения хи-квадрат. Распределение хи-квадрат используется в стандартных критериях хи-квадрат для проверки соответствия наблюдаемого распределения теоретическому, независимости двух критериев классификации качественных данных, а также для определения доверительного интервала для оценки стандартного отклонения генеральной совокупности по выборочному стандартному отклонению. Многие другие статистические тесты также используют это распределение, например, дисперсионный анализ Фридмана по рангам.
Определения
Если Z1, …, Zk – независимые стандартные нормально распределенные случайные величины, то сумма их квадратов распределена по закону хи-квадрат с k степенями свободы. Это обычно обозначается как χ².
is distributed according to the chi squared distribution with k degrees of freedom. This is usually denoted as
The chi squared distribution has one parameter: a positive integer k that specifies the number of degrees of freedom (the number of random variables being summed, Zi s).
Распределение хи-квадрат имеет один параметр: положительное целое число k, определяющее число степеней свободы (количество суммируемых случайных величин Zi).
is distributed according to the chi squared distribution with k degrees of freedom. This is usually denoted as
The chi squared distribution has one parameter: a positive integer k that specifies the number of degrees of freedom (the number of random variables being summed, Zi s).
Введение
Распределение хи-квадрат используется преимущественно в проверке гипотез и в меньшей степени для построения доверительных интервалов для дисперсии генеральной совокупности, когда базовое распределение является нормальным. В отличие от более известных распределений, таких как нормальное и экспоненциальное, распределение хи-квадрат реже применяется для непосредственного моделирования природных явлений. Оно возникает в следующих тестах гипотез, в частности:
Chi squared test of independence in contingency tables
Chi squared test of goodness of fit of observed data to hypothetical distributions
Likelihood ratio test for nested models
Log rank test in survival analysis
Cochran–Mantel–Haenszel test for stratified contingency tables
Wald test
Score test
It is also a component of the definition of the t distribution and the F distribution used in t tests, analysis of variance, and regression analysis. The primary reason for which the chi squared distribution is extensively used in hypothesis testing is its relationship to the normal distribution. Many hypothesis tests use a test statistic, such as the t statistic in a t test. For these hypothesis tests, as the sample size, n, increases, the sampling distribution of the test statistic approaches the normal distribution (central limit theorem). Because the test statistic (such as t) is asymptotically normally distributed, provided the sample size is sufficiently large, the distribution used for hypothesis testing may be approximated by a normal distribution. Testing hypotheses using a normal distribution is well understood and relatively easy. The simplest chi squared distribution is the square of a standard normal distribution. So wherever a normal distribution could be used for a hypothesis test, a chi squared distribution could be used. Suppose that is a random variable sampled from the standard normal distribution, where the mean is and the variance is : Now, consider the random variable The distribution of the random variable is an example of a chi squared distribution: The subscript 1 indicates that this particular chi squared distribution is constructed from only 1 standard normal distribution. A chi squared distribution constructed by squaring a single standard normal distribution is said to have 1 degree of freedom. Thus, as the sample size for a hypothesis test increases, the distribution of the test statistic approaches a normal distribution. Just as extreme values of the normal distribution have low probability (and give small p values), extreme values of the chi squared distribution have low probability. An additional reason that the chi squared distribution is widely used is that it turns up as the large sample distribution of generalized likelihood ratio tests (LRT). LRTs have several desirable properties; in particular, simple LRTs commonly provide the highest power to reject the null hypothesis (Neyman–Pearson lemma) and this leads also to optimality properties of generalised LRTs. However, the normal and chi squared approximations are only valid asymptotically. For this reason, it is preferable to use the t distribution rather than the normal approximation or the chi squared approximation for a small sample size. Similarly, in analyses of contingency tables, the chi squared approximation will be poor for a small sample size, and it is preferable to use Fisher's exact test. Ramsey shows that the exact binomial test is always more powerful than the normal approximation. Lancaster shows the connections among the binomial, normal, and chi squared distributions, as follows. De Moivre and Laplace established that a binomial distribution could be approximated by a normal distribution. Specifically they showed the asymptotic normality of the random variable
where is the observed number of successes in trials, where the probability of success is , and
Squaring both sides of the equation gives
Using , , and , this equation can be rewritten as
The expression on the right is of the form that Karl Pearson would generalize to the form
where
= Pearson's cumulative test statistic, which asymptotically approaches a distribution;
= the number of observations of type ;
= the expected (theoretical) frequency of type , asserted by the null hypothesis that the fraction of type in the population is ; and
= the number of cells in the table. In the case of a binomial outcome (flipping a coin), the binomial distribution may be approximated by a normal distribution (for sufficiently large ). Because the square of a standard normal distribution is the chi squared distribution with one degree of freedom, the probability of a result such as 1 heads in 10 trials can be approximated either by using the normal distribution directly, or the chi squared distribution for the normalised, squared difference between observed and expected value. However, many problems involve more than the two possible outcomes of a binomial, and instead require 3 or more categories, which leads to the multinomial distribution. Just as de Moivre and Laplace sought for and found the normal approximation to the binomial, Pearson sought for and found a degenerate multivariate normal approximation to the multinomial distribution (the numbers in each category add up to the total sample size, which is considered fixed). Pearson showed that the chi squared distribution arose from such a multivariate normal approximation to the multinomial distribution, taking careful account of the statistical dependence (negative correlations) between numbers of observations in different categories. For the cases when (which include all of the cases when this CDF is less than half):
The tail bound for the cases when , similarly, is
For another approximation for the CDF modeled after the cube of a Gaussian, see under Noncentral chi squared distribution.
Тест хи-квадрат на независимость в таблицах сопряженности
Тест хи-квадрат на соответствие наблюдаемых данных гипотетическим распределениям
Тест отношения правдоподобия для вложенных моделей
Log-rank тест в анализе выживаемости
Тест Кохрана — Мантеля — Хензеля для стратифицированных таблиц сопряженности
Wald-тест
Score-тест
Chi squared test of independence in contingency tables
Chi squared test of goodness of fit of observed data to hypothetical distributions
Likelihood ratio test for nested models
Log rank test in survival analysis
Cochran–Mantel–Haenszel test for stratified contingency tables
Wald test
Score test
It is also a component of the definition of the t distribution and the F distribution used in t tests, analysis of variance, and regression analysis. The primary reason for which the chi squared distribution is extensively used in hypothesis testing is its relationship to the normal distribution. Many hypothesis tests use a test statistic, such as the t statistic in a t test. For these hypothesis tests, as the sample size, n, increases, the sampling distribution of the test statistic approaches the normal distribution (central limit theorem). Because the test statistic (such as t) is asymptotically normally distributed, provided the sample size is sufficiently large, the distribution used for hypothesis testing may be approximated by a normal distribution. Testing hypotheses using a normal distribution is well understood and relatively easy. The simplest chi squared distribution is the square of a standard normal distribution. So wherever a normal distribution could be used for a hypothesis test, a chi squared distribution could be used. Suppose that is a random variable sampled from the standard normal distribution, where the mean is and the variance is : Now, consider the random variable The distribution of the random variable is an example of a chi squared distribution: The subscript 1 indicates that this particular chi squared distribution is constructed from only 1 standard normal distribution. A chi squared distribution constructed by squaring a single standard normal distribution is said to have 1 degree of freedom. Thus, as the sample size for a hypothesis test increases, the distribution of the test statistic approaches a normal distribution. Just as extreme values of the normal distribution have low probability (and give small p values), extreme values of the chi squared distribution have low probability. An additional reason that the chi squared distribution is widely used is that it turns up as the large sample distribution of generalized likelihood ratio tests (LRT). LRTs have several desirable properties; in particular, simple LRTs commonly provide the highest power to reject the null hypothesis (Neyman–Pearson lemma) and this leads also to optimality properties of generalised LRTs. However, the normal and chi squared approximations are only valid asymptotically. For this reason, it is preferable to use the t distribution rather than the normal approximation or the chi squared approximation for a small sample size. Similarly, in analyses of contingency tables, the chi squared approximation will be poor for a small sample size, and it is preferable to use Fisher's exact test. Ramsey shows that the exact binomial test is always more powerful than the normal approximation. Lancaster shows the connections among the binomial, normal, and chi squared distributions, as follows. De Moivre and Laplace established that a binomial distribution could be approximated by a normal distribution. Specifically they showed the asymptotic normality of the random variable
where is the observed number of successes in trials, where the probability of success is , and
Squaring both sides of the equation gives
Using , , and , this equation can be rewritten as
The expression on the right is of the form that Karl Pearson would generalize to the form
where
= Pearson's cumulative test statistic, which asymptotically approaches a distribution;
= the number of observations of type ;
= the expected (theoretical) frequency of type , asserted by the null hypothesis that the fraction of type in the population is ; and
= the number of cells in the table. In the case of a binomial outcome (flipping a coin), the binomial distribution may be approximated by a normal distribution (for sufficiently large ). Because the square of a standard normal distribution is the chi squared distribution with one degree of freedom, the probability of a result such as 1 heads in 10 trials can be approximated either by using the normal distribution directly, or the chi squared distribution for the normalised, squared difference between observed and expected value. However, many problems involve more than the two possible outcomes of a binomial, and instead require 3 or more categories, which leads to the multinomial distribution. Just as de Moivre and Laplace sought for and found the normal approximation to the binomial, Pearson sought for and found a degenerate multivariate normal approximation to the multinomial distribution (the numbers in each category add up to the total sample size, which is considered fixed). Pearson showed that the chi squared distribution arose from such a multivariate normal approximation to the multinomial distribution, taking careful account of the statistical dependence (negative correlations) between numbers of observations in different categories. For the cases when (which include all of the cases when this CDF is less than half):
The tail bound for the cases when , similarly, is
For another approximation for the CDF modeled after the cube of a Gaussian, see under Noncentral chi squared distribution.
Оно также является компонентом определения t-распределения и F-распределения, используемых в t-тестах, дисперсионном анализе и регрессионном анализе. Основная причина широкого использования распределения хи-квадрат в проверке гипотез заключается в его связи с нормальным распределением. Многие тесты гипотез используют тестовую статистику, например t-статистику в t-тесте. Для этих тестов, при увеличении размера выборки n, распределение выборочной статистики приближается к нормальному распределению (центральная предельная теорема). Поскольку тестовая статистика (например, t) асимптотически нормально распределена, при достаточно большом размере выборки распределение, используемое для проверки гипотез, может быть аппроксимировано нормальным распределением. Проверка гипотез с использованием нормального распределения хорошо изучена и относительно проста. Самое простое распределение хи-квадрат – это квадрат стандартного нормального распределения. Таким образом, везде, где можно использовать нормальное распределение для проверки гипотезы, можно использовать и распределение хи-квадрат. Предположим, что – это случайная величина, взятая из стандартного нормального распределения, где среднее равно , а дисперсия равна . Теперь рассмотрим случайную величину . Распределение случайной величины является примером распределения хи-квадрат: нижний индекс 1 указывает, что это конкретное распределение хи-квадрат построено только из 1 стандартного нормального распределения. Распределение хи-квадрат, построенное путем возведения в квадрат одного стандартного нормального распределения, имеет 1 степень свободы. Следовательно, по мере увеличения размера выборки для проверки гипотезы распределение тестовой статистики приближается к нормальному распределению. Как и экстремальные значения нормального распределения имеют низкую вероятность (и дают малые p-значения), так и экстремальные значения распределения хи-квадрат имеют низкую вероятность. Дополнительная причина широкого использования распределения хи-квадрат заключается в том, что оно возникает как асимптотическое распределение обобщенных тестов отношения правдоподобия (LRT). LRT обладают несколькими желательными свойствами; в частности, простые LRT обычно обеспечивают наибольшую мощность для отклонения нулевой гипотезы (лемма Неймана — Пирсона), что также приводит к свойствам оптимальности обобщенных LRT. Однако нормальные и хи-квадратные аппроксимации верны только асимптотически. По этой причине предпочтительнее использовать t-распределение, а не нормальную или хи-квадратную аппроксимацию для небольшого размера выборки. Аналогично, при анализе таблиц сопряженности хи-квадратная аппроксимация будет плохой для небольшого размера выборки, и предпочтительнее использовать точный тест Фишера. Рамси показывает, что точный биномиальный тест всегда мощнее, чем нормальная аппроксимация. Ланкастер показывает связи между биномиальным, нормальным и хи-квадратным распределениями следующим образом. Де Муавр и Лаплас установили, что биномиальное распределение можно аппроксимировать нормальным распределением. В частности, они показали асимптотическую нормальность случайной величины , где – наблюдаемое число успехов в испытаниях, где вероятность успеха равна , и . Возведение в квадрат обеих частей уравнения дает . Используя , , и , это уравнение можно переписать как . Выражение справа имеет вид, который Карл Пирсон обобщит до формы , где
Chi squared test of independence in contingency tables
Chi squared test of goodness of fit of observed data to hypothetical distributions
Likelihood ratio test for nested models
Log rank test in survival analysis
Cochran–Mantel–Haenszel test for stratified contingency tables
Wald test
Score test
It is also a component of the definition of the t distribution and the F distribution used in t tests, analysis of variance, and regression analysis. The primary reason for which the chi squared distribution is extensively used in hypothesis testing is its relationship to the normal distribution. Many hypothesis tests use a test statistic, such as the t statistic in a t test. For these hypothesis tests, as the sample size, n, increases, the sampling distribution of the test statistic approaches the normal distribution (central limit theorem). Because the test statistic (such as t) is asymptotically normally distributed, provided the sample size is sufficiently large, the distribution used for hypothesis testing may be approximated by a normal distribution. Testing hypotheses using a normal distribution is well understood and relatively easy. The simplest chi squared distribution is the square of a standard normal distribution. So wherever a normal distribution could be used for a hypothesis test, a chi squared distribution could be used. Suppose that is a random variable sampled from the standard normal distribution, where the mean is and the variance is : Now, consider the random variable The distribution of the random variable is an example of a chi squared distribution: The subscript 1 indicates that this particular chi squared distribution is constructed from only 1 standard normal distribution. A chi squared distribution constructed by squaring a single standard normal distribution is said to have 1 degree of freedom. Thus, as the sample size for a hypothesis test increases, the distribution of the test statistic approaches a normal distribution. Just as extreme values of the normal distribution have low probability (and give small p values), extreme values of the chi squared distribution have low probability. An additional reason that the chi squared distribution is widely used is that it turns up as the large sample distribution of generalized likelihood ratio tests (LRT). LRTs have several desirable properties; in particular, simple LRTs commonly provide the highest power to reject the null hypothesis (Neyman–Pearson lemma) and this leads also to optimality properties of generalised LRTs. However, the normal and chi squared approximations are only valid asymptotically. For this reason, it is preferable to use the t distribution rather than the normal approximation or the chi squared approximation for a small sample size. Similarly, in analyses of contingency tables, the chi squared approximation will be poor for a small sample size, and it is preferable to use Fisher's exact test. Ramsey shows that the exact binomial test is always more powerful than the normal approximation. Lancaster shows the connections among the binomial, normal, and chi squared distributions, as follows. De Moivre and Laplace established that a binomial distribution could be approximated by a normal distribution. Specifically they showed the asymptotic normality of the random variable
where is the observed number of successes in trials, where the probability of success is , and
Squaring both sides of the equation gives
Using , , and , this equation can be rewritten as
The expression on the right is of the form that Karl Pearson would generalize to the form
where
= Pearson's cumulative test statistic, which asymptotically approaches a distribution;
= the number of observations of type ;
= the expected (theoretical) frequency of type , asserted by the null hypothesis that the fraction of type in the population is ; and
= the number of cells in the table. In the case of a binomial outcome (flipping a coin), the binomial distribution may be approximated by a normal distribution (for sufficiently large ). Because the square of a standard normal distribution is the chi squared distribution with one degree of freedom, the probability of a result such as 1 heads in 10 trials can be approximated either by using the normal distribution directly, or the chi squared distribution for the normalised, squared difference between observed and expected value. However, many problems involve more than the two possible outcomes of a binomial, and instead require 3 or more categories, which leads to the multinomial distribution. Just as de Moivre and Laplace sought for and found the normal approximation to the binomial, Pearson sought for and found a degenerate multivariate normal approximation to the multinomial distribution (the numbers in each category add up to the total sample size, which is considered fixed). Pearson showed that the chi squared distribution arose from such a multivariate normal approximation to the multinomial distribution, taking careful account of the statistical dependence (negative correlations) between numbers of observations in different categories. For the cases when (which include all of the cases when this CDF is less than half):
The tail bound for the cases when , similarly, is
For another approximation for the CDF modeled after the cube of a Gaussian, see under Noncentral chi squared distribution.
= кумулятивная тестовая статистика Пирсона, которая асимптотически приближается к распределению;
= число наблюдений типа ;
= ожидаемая (теоретическая) частота типа , утверждаемая нулевой гипотезой о том, что доля типа в популяции равна ; и
= число ячеек в таблице. В случае биномиального исхода (подбрасывание монеты) биномиальное распределение может быть аппроксимировано нормальным распределением (при достаточно большом ). Поскольку квадрат стандартного нормального распределения является распределением хи-квадрат с одной степенью свободы, вероятность такого результата, как 1 орел в 10 испытаниях, можно аппроксимировать либо с использованием нормального распределения напрямую, либо с использованием распределения хи-квадрат для нормализованной, возведенной в квадрат разницы между наблюдаемым и ожидаемым значением. Однако многие задачи включают в себя более двух возможных исходов биномиального распределения и вместо этого требуют 3 или более категорий, что приводит к многочленному распределению. Как Де Муавр и Лаплас искали и нашли нормальную аппроксимацию для биномиального распределения, так и Пирсон искал и нашел вырожденную многомерную нормальную аппроксимацию для многочленного распределения (числа в каждой категории в сумме дают размер выборки, который считается фиксированным). Пирсон показал, что распределение хи-квадрат возникает из такой многомерной нормальной аппроксимации к многочленному распределению, тщательно учитывая статистическую зависимость (отрицательную корреляцию) между числами наблюдений в разных категориях. Для случаев, когда (что включает в себя все случаи, когда эта CDF меньше половины):
Chi squared test of independence in contingency tables
Chi squared test of goodness of fit of observed data to hypothetical distributions
Likelihood ratio test for nested models
Log rank test in survival analysis
Cochran–Mantel–Haenszel test for stratified contingency tables
Wald test
Score test
It is also a component of the definition of the t distribution and the F distribution used in t tests, analysis of variance, and regression analysis. The primary reason for which the chi squared distribution is extensively used in hypothesis testing is its relationship to the normal distribution. Many hypothesis tests use a test statistic, such as the t statistic in a t test. For these hypothesis tests, as the sample size, n, increases, the sampling distribution of the test statistic approaches the normal distribution (central limit theorem). Because the test statistic (such as t) is asymptotically normally distributed, provided the sample size is sufficiently large, the distribution used for hypothesis testing may be approximated by a normal distribution. Testing hypotheses using a normal distribution is well understood and relatively easy. The simplest chi squared distribution is the square of a standard normal distribution. So wherever a normal distribution could be used for a hypothesis test, a chi squared distribution could be used. Suppose that is a random variable sampled from the standard normal distribution, where the mean is and the variance is : Now, consider the random variable The distribution of the random variable is an example of a chi squared distribution: The subscript 1 indicates that this particular chi squared distribution is constructed from only 1 standard normal distribution. A chi squared distribution constructed by squaring a single standard normal distribution is said to have 1 degree of freedom. Thus, as the sample size for a hypothesis test increases, the distribution of the test statistic approaches a normal distribution. Just as extreme values of the normal distribution have low probability (and give small p values), extreme values of the chi squared distribution have low probability. An additional reason that the chi squared distribution is widely used is that it turns up as the large sample distribution of generalized likelihood ratio tests (LRT). LRTs have several desirable properties; in particular, simple LRTs commonly provide the highest power to reject the null hypothesis (Neyman–Pearson lemma) and this leads also to optimality properties of generalised LRTs. However, the normal and chi squared approximations are only valid asymptotically. For this reason, it is preferable to use the t distribution rather than the normal approximation or the chi squared approximation for a small sample size. Similarly, in analyses of contingency tables, the chi squared approximation will be poor for a small sample size, and it is preferable to use Fisher's exact test. Ramsey shows that the exact binomial test is always more powerful than the normal approximation. Lancaster shows the connections among the binomial, normal, and chi squared distributions, as follows. De Moivre and Laplace established that a binomial distribution could be approximated by a normal distribution. Specifically they showed the asymptotic normality of the random variable
where is the observed number of successes in trials, where the probability of success is , and
Squaring both sides of the equation gives
Using , , and , this equation can be rewritten as
The expression on the right is of the form that Karl Pearson would generalize to the form
where
= Pearson's cumulative test statistic, which asymptotically approaches a distribution;
= the number of observations of type ;
= the expected (theoretical) frequency of type , asserted by the null hypothesis that the fraction of type in the population is ; and
= the number of cells in the table. In the case of a binomial outcome (flipping a coin), the binomial distribution may be approximated by a normal distribution (for sufficiently large ). Because the square of a standard normal distribution is the chi squared distribution with one degree of freedom, the probability of a result such as 1 heads in 10 trials can be approximated either by using the normal distribution directly, or the chi squared distribution for the normalised, squared difference between observed and expected value. However, many problems involve more than the two possible outcomes of a binomial, and instead require 3 or more categories, which leads to the multinomial distribution. Just as de Moivre and Laplace sought for and found the normal approximation to the binomial, Pearson sought for and found a degenerate multivariate normal approximation to the multinomial distribution (the numbers in each category add up to the total sample size, which is considered fixed). Pearson showed that the chi squared distribution arose from such a multivariate normal approximation to the multinomial distribution, taking careful account of the statistical dependence (negative correlations) between numbers of observations in different categories. For the cases when (which include all of the cases when this CDF is less than half):
The tail bound for the cases when , similarly, is
For another approximation for the CDF modeled after the cube of a Gaussian, see under Noncentral chi squared distribution.
Граничное значение для случаев, когда , аналогично, равно
Chi squared test of independence in contingency tables
Chi squared test of goodness of fit of observed data to hypothetical distributions
Likelihood ratio test for nested models
Log rank test in survival analysis
Cochran–Mantel–Haenszel test for stratified contingency tables
Wald test
Score test
It is also a component of the definition of the t distribution and the F distribution used in t tests, analysis of variance, and regression analysis. The primary reason for which the chi squared distribution is extensively used in hypothesis testing is its relationship to the normal distribution. Many hypothesis tests use a test statistic, such as the t statistic in a t test. For these hypothesis tests, as the sample size, n, increases, the sampling distribution of the test statistic approaches the normal distribution (central limit theorem). Because the test statistic (such as t) is asymptotically normally distributed, provided the sample size is sufficiently large, the distribution used for hypothesis testing may be approximated by a normal distribution. Testing hypotheses using a normal distribution is well understood and relatively easy. The simplest chi squared distribution is the square of a standard normal distribution. So wherever a normal distribution could be used for a hypothesis test, a chi squared distribution could be used. Suppose that is a random variable sampled from the standard normal distribution, where the mean is and the variance is : Now, consider the random variable The distribution of the random variable is an example of a chi squared distribution: The subscript 1 indicates that this particular chi squared distribution is constructed from only 1 standard normal distribution. A chi squared distribution constructed by squaring a single standard normal distribution is said to have 1 degree of freedom. Thus, as the sample size for a hypothesis test increases, the distribution of the test statistic approaches a normal distribution. Just as extreme values of the normal distribution have low probability (and give small p values), extreme values of the chi squared distribution have low probability. An additional reason that the chi squared distribution is widely used is that it turns up as the large sample distribution of generalized likelihood ratio tests (LRT). LRTs have several desirable properties; in particular, simple LRTs commonly provide the highest power to reject the null hypothesis (Neyman–Pearson lemma) and this leads also to optimality properties of generalised LRTs. However, the normal and chi squared approximations are only valid asymptotically. For this reason, it is preferable to use the t distribution rather than the normal approximation or the chi squared approximation for a small sample size. Similarly, in analyses of contingency tables, the chi squared approximation will be poor for a small sample size, and it is preferable to use Fisher's exact test. Ramsey shows that the exact binomial test is always more powerful than the normal approximation. Lancaster shows the connections among the binomial, normal, and chi squared distributions, as follows. De Moivre and Laplace established that a binomial distribution could be approximated by a normal distribution. Specifically they showed the asymptotic normality of the random variable
where is the observed number of successes in trials, where the probability of success is , and
Squaring both sides of the equation gives
Using , , and , this equation can be rewritten as
The expression on the right is of the form that Karl Pearson would generalize to the form
where
= Pearson's cumulative test statistic, which asymptotically approaches a distribution;
= the number of observations of type ;
= the expected (theoretical) frequency of type , asserted by the null hypothesis that the fraction of type in the population is ; and
= the number of cells in the table. In the case of a binomial outcome (flipping a coin), the binomial distribution may be approximated by a normal distribution (for sufficiently large ). Because the square of a standard normal distribution is the chi squared distribution with one degree of freedom, the probability of a result such as 1 heads in 10 trials can be approximated either by using the normal distribution directly, or the chi squared distribution for the normalised, squared difference between observed and expected value. However, many problems involve more than the two possible outcomes of a binomial, and instead require 3 or more categories, which leads to the multinomial distribution. Just as de Moivre and Laplace sought for and found the normal approximation to the binomial, Pearson sought for and found a degenerate multivariate normal approximation to the multinomial distribution (the numbers in each category add up to the total sample size, which is considered fixed). Pearson showed that the chi squared distribution arose from such a multivariate normal approximation to the multinomial distribution, taking careful account of the statistical dependence (negative correlations) between numbers of observations in different categories. For the cases when (which include all of the cases when this CDF is less than half):
The tail bound for the cases when , similarly, is
For another approximation for the CDF modeled after the cube of a Gaussian, see under Noncentral chi squared distribution.
Для другой аппроксимации для CDF, смоделированной на основе куба гауссиана, см. раздел «Нецентральное распределение хи-квадрат».
Chi squared test of independence in contingency tables
Chi squared test of goodness of fit of observed data to hypothetical distributions
Likelihood ratio test for nested models
Log rank test in survival analysis
Cochran–Mantel–Haenszel test for stratified contingency tables
Wald test
Score test
It is also a component of the definition of the t distribution and the F distribution used in t tests, analysis of variance, and regression analysis. The primary reason for which the chi squared distribution is extensively used in hypothesis testing is its relationship to the normal distribution. Many hypothesis tests use a test statistic, such as the t statistic in a t test. For these hypothesis tests, as the sample size, n, increases, the sampling distribution of the test statistic approaches the normal distribution (central limit theorem). Because the test statistic (such as t) is asymptotically normally distributed, provided the sample size is sufficiently large, the distribution used for hypothesis testing may be approximated by a normal distribution. Testing hypotheses using a normal distribution is well understood and relatively easy. The simplest chi squared distribution is the square of a standard normal distribution. So wherever a normal distribution could be used for a hypothesis test, a chi squared distribution could be used. Suppose that is a random variable sampled from the standard normal distribution, where the mean is and the variance is : Now, consider the random variable The distribution of the random variable is an example of a chi squared distribution: The subscript 1 indicates that this particular chi squared distribution is constructed from only 1 standard normal distribution. A chi squared distribution constructed by squaring a single standard normal distribution is said to have 1 degree of freedom. Thus, as the sample size for a hypothesis test increases, the distribution of the test statistic approaches a normal distribution. Just as extreme values of the normal distribution have low probability (and give small p values), extreme values of the chi squared distribution have low probability. An additional reason that the chi squared distribution is widely used is that it turns up as the large sample distribution of generalized likelihood ratio tests (LRT). LRTs have several desirable properties; in particular, simple LRTs commonly provide the highest power to reject the null hypothesis (Neyman–Pearson lemma) and this leads also to optimality properties of generalised LRTs. However, the normal and chi squared approximations are only valid asymptotically. For this reason, it is preferable to use the t distribution rather than the normal approximation or the chi squared approximation for a small sample size. Similarly, in analyses of contingency tables, the chi squared approximation will be poor for a small sample size, and it is preferable to use Fisher's exact test. Ramsey shows that the exact binomial test is always more powerful than the normal approximation. Lancaster shows the connections among the binomial, normal, and chi squared distributions, as follows. De Moivre and Laplace established that a binomial distribution could be approximated by a normal distribution. Specifically they showed the asymptotic normality of the random variable
where is the observed number of successes in trials, where the probability of success is , and
Squaring both sides of the equation gives
Using , , and , this equation can be rewritten as
The expression on the right is of the form that Karl Pearson would generalize to the form
where
= Pearson's cumulative test statistic, which asymptotically approaches a distribution;
= the number of observations of type ;
= the expected (theoretical) frequency of type , asserted by the null hypothesis that the fraction of type in the population is ; and
= the number of cells in the table. In the case of a binomial outcome (flipping a coin), the binomial distribution may be approximated by a normal distribution (for sufficiently large ). Because the square of a standard normal distribution is the chi squared distribution with one degree of freedom, the probability of a result such as 1 heads in 10 trials can be approximated either by using the normal distribution directly, or the chi squared distribution for the normalised, squared difference between observed and expected value. However, many problems involve more than the two possible outcomes of a binomial, and instead require 3 or more categories, which leads to the multinomial distribution. Just as de Moivre and Laplace sought for and found the normal approximation to the binomial, Pearson sought for and found a degenerate multivariate normal approximation to the multinomial distribution (the numbers in each category add up to the total sample size, which is considered fixed). Pearson showed that the chi squared distribution arose from such a multivariate normal approximation to the multinomial distribution, taking careful account of the statistical dependence (negative correlations) between numbers of observations in different categories. For the cases when (which include all of the cases when this CDF is less than half):
The tail bound for the cases when , similarly, is
For another approximation for the CDF modeled after the cube of a Gaussian, see under Noncentral chi squared distribution.
Добавленность
Из определения распределения хи-квадрат следует, что сумма независимых хи-квадратных переменных также имеет распределение хи-квадрат. В частности, если являются независимыми хи-квадратными переменными с , степенями свободы соответственно, то имеет распределение хи-квадрат с степенями свободы.
Концентрация
Распределение хи-квадрат демонстрирует сильную концентрацию вокруг своего среднего значения. Стандартные границы Лорана Массара:
Одним из следствий является то, что если – гауссовский случайный вектор в , то при увеличении размерности квадрат длины вектора сильно концентрируется вокруг с шириной : где показатель может быть выбран как любое значение в .
Нецентральное распределение чи-квадрат
Нецентральное хи-квадрат распределение получается из суммы квадратов независимых нормально распределённых случайных величин с единичной дисперсией и ненулевыми математическими ожиданиями.
Обобщенное распределение чи-квадрат
Обобщенное хи-квадрат распределение получается из квадратичной формы z'Az, где z — гауссовский вектор с нулевым средним и произвольной матрицей ковариации, а A — произвольная матрица.
Таблица значений против значений
Значение представляет собой вероятность наблюдения тестовой статистики, по крайней мере, столь же экстремальной в распределении хи-квадрат. Соответственно, поскольку функция кумулятивного распределения (CDF) для соответствующих степеней свободы (df) дает вероятность получения значения, менее экстремального, чем данная точка, вычитание значения CDF из 1 дает p-значение. Низкое p-значение, ниже выбранного уровня значимости, указывает на статистическую значимость, то есть достаточно доказательств для отклонения нулевой гипотезы. Уровень значимости 0,05 часто используется в качестве границы между значимыми и незначимыми результатами. В таблице ниже приведены p-значения, соответствующие для первых 10 степеней свободы.
These values can be calculated evaluating the quantile function (also known as "inverse CDF" or "ICDF") of the chi squared distribution; e. g., the χ^(2) ICDF for 1=p = 0.05 and 1=df = 7 yields 2.1673 ≈ 2.17 as in the table above, noticing that 1 – p is the p value from the table.
Степени свободы (df) | Значение | 0.95 | 0.90 | 0.80 | 0.70 | 0.50 | 0.30 | 0.20 | 0.10 | 0.05 | 0.01 | 0.001
---|---|---|---|---|---|---|---|---|---|---|---|---
1 | 0.004 | 0.02 | 0.06 | 0.15 | 0.46 | 1.07 | 1.64 | 2.71 | 3.84 | 6.63 | 10.83
2 | 0.10 | 0.21 | 0.45 | 0.71 | 1.39 | 2.41 | 3.22 | 4.61 | 5.99 | 9.21 | 13.82
3 | 0.35 | 0.58 | 1.01 | 1.42 | 2.37 | 3.66 | 4.64 | 6.25 | 7.81 | 11.34 | 16.27
4 | 0.71 | 1.06 | 1.65 | 2.20 | 3.36 | 4.88 | 5.99 | 7.78 | 9.49 | 13.28 | 18.47
5 | 1.14 | 1.61 | 2.34 | 3.00 | 4.35 | 6.06 | 7.29 | 9.24 | 11.07 | 15.09 | 20.52
6 | 1.63 | 2.20 | 3.07 | 3.83 | 5.35 | 7.23 | 8.56 | 10.64 | 12.59 | 16.81 | 22.46
7 | 2.17 | 2.83 | 3.82 | 4.67 | 6.35 | 8.38 | 9.80 | 12.02 | 14.07 | 18.48 | 24.32
8 | 2.73 | 3.49 | 4.59 | 5.53 | 7.34 | 9.52 | 11.03 | 13.36 | 15.51 | 20.09 | 26.12
9 | 3.32 | 4.17 | 5.38 | 6.39 | 8.34 | 10.66 | 12.24 | 14.68 | 16.92 | 21.67 | 27.88
10 | 3.94 | 4.87 | 6.18 | 7.27 | 9.34 | 11.78 | 13.44 | 15.99 | 18.31 | 23.21 | 29.59
These values can be calculated evaluating the quantile function (also known as "inverse CDF" or "ICDF") of the chi squared distribution; e. g., the χ^(2) ICDF for 1=p = 0.05 and 1=df = 7 yields 2.1673 ≈ 2.17 as in the table above, noticing that 1 – p is the p value from the table.
Эти значения можно вычислить, оценивая квантильную функцию (также известную как "обратная CDF" или "ICDF") распределения хи-квадрат; например, χ^(2) ICDF для p = 0,05 и df = 7 дает 2,1673 ≈ 2,17, как в таблице выше, учитывая, что 1 – p является p-значением из таблицы.
These values can be calculated evaluating the quantile function (also known as "inverse CDF" or "ICDF") of the chi squared distribution; e. g., the χ^(2) ICDF for 1=p = 0.05 and 1=df = 7 yields 2.1673 ≈ 2.17 as in the table above, noticing that 1 – p is the p value from the table.
История
Это распределение впервые было описано немецким геодезистом и статистиком Фридрихом Робертом Гельмертом в работах 1875–1876 годов, где он вычислил асимптотическое распределение выборочной дисперсии нормальной совокупности. Таким образом, в немецком языке оно традиционно было известно как Helmert'sche ("Гельмертовское") или "распределение Гельмерта". Распределение было независимо повторно открыто английским математиком Карлом Пирсоном в контексте проверки гипотез о соответствии, для которой он разработал свой критерий хи-квадрат, опубликованный в 1900 году, с таблицей вычисленных значений, опубликованной в , собранной в 1999 году. Название "хи-квадрат" в конечном итоге происходит от сокращения, использованного Пирсоном для обозначения показателя степени в многомерном нормальном распределении с использованием греческой буквы Chi, записывая −½χ² для того, что в современной нотации выглядит как −½'x'^(T)Σ^(−1)'x' (где Σ – матрица ковариации). Однако идея семейства "распределений хи-квадрат" не принадлежит Пирсону, а возникла как дальнейшая разработка, выполненная Фишером в 1920-х годах.