Сравнивайте с английским: нажмите на абзац — оригинал откроется в окне. Кнопка EN под абзацем показывает его прямо в тексте.
Содержание
Введение
Статистический тест
Statistical test
В статистике G-тесты — это статистические тесты отношения правдоподобия или максимального правдоподобия, которые все чаще применяются в ситуациях, где ранее рекомендовалось использовать критерий хи-квадрат.
In statistics, G tests are likelihood ratio or maximum likelihood statistical significance tests that are increasingly being used in situations where chi squared tests were previously recommended.
Вывод
Мы можем вывести значение G-теста из теста отношения правдоподобия, где базовой моделью является мультиномиальная модель. Предположим, у нас есть выборка, в которой каждый – это количество раз, когда объект типа был наблюдаем. Кроме того, пусть – общее количество наблюдаемых объектов. Если мы предполагаем, что базовая модель является мультиномиальной, то статистика теста определяется как, где – нулевая гипотеза, а – оценка максимального правдоподобия (MLE) параметров, полученная по данным. Вспомним, что для мультиномиальной модели MLE определяется как: Кроме того, мы можем представить каждый параметр нулевой гипотезы как: Таким образом, подставляя представления и в отношение правдоподобия в логарифмической форме, уравнение упрощается до: Переобозначим переменные через и через . Наконец, умножим на коэффициент (используется для того, чтобы сделать формулу G-теста асимптотически эквивалентной формуле хи-квадрат Пирсона) для получения следующего вида:
We can derive the value of the G test from the log likelihood ratio test where the underlying model is a multinomial model. Suppose we had a sample where each is the number of times that an object of type was observed. Furthermore, let be the total number of objects observed. If we assume that the underlying model is multinomial, then the test statistic is defined bywhere is the null hypothesis and is the maximum likelihood estimate (MLE) of the parameters given the data. Recall that for the multinomial model, the MLE of given some data is defined byFurthermore, we may represent each null hypothesis parameter asThus, by substituting the representations of and in the log likelihood ratio, the equation simplifies toRelabel the variables with and with Finally, multiply by a factor of (used to make the G test formula asymptotically equivalent to the Pearson's chi squared test formula) to achieve the form
Эвристически, можно представить как непрерывную величину, стремящуюся к нулю, в этом случае и члены с нулевыми наблюдениями можно просто отбросить. Однако ожидаемое количество в каждой ячейке должно быть строго больше нуля для каждой ячейки, чтобы применять данный метод.
Heuristically, one can imagine as continuous and approaching zero, in which case and terms with zero observations can simply be dropped. However the expected count in each cell must be strictly greater than zero for each cell to apply the method.
Распространение и использование
При нулевой гипотезе о том, что наблюдаемые частоты являются результатом случайной выборки из распределения с заданными ожидаемыми частотами, распределение статистики G приблизительно соответствует хи-квадрат распределению с таким же числом степеней свободы, как и в соответствующем хи-квадрат тесте. Для очень малых выборок предпочтительнее многочленный тест на соответствие, точный тест Фишера для таблиц сопряженности или даже байесовский выбор гипотез, чем G-тест. Макдональд рекомендует всегда использовать точный тест (точный тест на соответствие, точный тест Фишера), если общий размер выборки меньше 1000. Нет ничего особенного в размере выборки 1000, это просто удобное округленное число, которое находится в диапазоне, где точный тест, хи-квадрат тест и G-тест дадут почти идентичные p-значения. Электронные таблицы, веб-калькуляторы и SAS не должны испытывать проблем с проведением точного теста при размере выборки 1000. — John H. McDonald
Given the null hypothesis that the observed frequencies result from random sampling from a distribution with the given expected frequencies, the distribution of G is approximately a chi squared distribution, with the same number of degrees of freedom as in the corresponding chi squared test. For very small samples the multinomial test for goodness of fit, and Fisher's exact test for contingency tables, or even Bayesian hypothesis selection are preferable to the G test. McDonald recommends to always use an exact test (exact test of goodness of fit, Fisher's exact test) if the total sample size is less than 1 000 There is nothing magical about a sample size of 1 000, it's just a nice round number that is well within the range where an exact test, chi square test, and G–test will give almost identical p values. Spreadsheets, web page calculators, and SAS shouldn't have any problem doing an exact test on a sample size of 1 000 — John H. McDonald
Применение
Тест Макдональд — Крайтмана в статистической генетике является применением G-теста. Даннинг представил этот тест сообществу компьютерной лингвистики, где он сейчас широко используется. Программа R scape (используемая Rfam) применяет G-тест для выявления ковариации между позициями выравнивания последовательностей РНК.
The McDonald–Kreitman test in statistical genetics is an application of the G test. Dunning introduced the test to the computational linguistics community where it is now widely used. The R scape program (used by Rfam) uses G test to detect co variation between RNA sequence alignment positions.
Статистическое программное обеспечение
В R быстрые реализации можно найти в пакетах AMR и Rfast. Для пакета AMR команда g.test работает точно так же, как chisq.test из базового R. В R также есть функция likelihood.test в пакете Deducer. Важно отметить: G-тест Фишера в пакете GeneCycle языка программирования R (fisher.g.test) не реализует G-тест, как описано в этой статье, а скорее точный тест Фишера для гауссовского белого шума во временном ряду. Другая реализация в R для вычисления статистики G и соответствующих p-значений предоставляется пакетом entropy. Команды: Gstat для стандартной статистики G и связанного с ней p-значения, и Gstatindep для статистики G, применяемой для сравнения совместных и произведенных распределений для проверки независимости. В SAS можно провести G-тест, используя опцию /chisq после команды proc freq. В Stata можно провести G-тест, используя опцию lr после команды tabulate. В Java используйте org.apache.commons.math3.stat.inference.GTest. В Python используйте scipy.stats.power_divergence с lambda = 0.
In R fast implementations can be found in the AMR and Rfast packages. For the AMR package, the command is g. test which works exactly like chisq. test from base R. R also has the likelihood. test function in the Deducer package. Note: Fisher's G test in the GeneCycle Package of the R programming language (fisher. g. test) does not implement the G test as described in this article, but rather Fisher's exact test of Gaussian white noise in a time series. Another R implementation to compute the G statistic and corresponding p values is provided by the R package entropy. The commands are Gstat for the standard G statistic and the associated p value and Gstatindep for the G statistic applied to comparing joint and product distributions to test independence. In SAS, one can conduct G test by applying the /chisq option after the proc freq. In Stata, one can conduct a G test by applying the lr option after the tabulate command. In Java, use org. apache. commons. math3. stat. inference. GTest. In Python, use scipy. stats. power divergence with lambda =0.