Введение
Неравенство, применимое к случайным величинам
В теории информации неравенство Фано (также известное как теорема Фано и лемма Фано) связывает среднюю потерю информации в зашумлённом канале с вероятностью ошибки классификации. Оно было выведено Робертом Фано в начале 1950-х годов, когда он вёл Ph.D. семинар по теории информации в Массачусетском технологическом институте (MIT), а затем опубликовано в его учебнике 1961 года. Оно используется для нахождения нижней границы вероятности ошибки любого декодера, а также нижней границы минимаксных рисков при оценке плотности. Пусть случайные величины и обозначают входное и выходное сообщения с совместным распределением вероятностей. Пусть обозначает событие ошибки, то есть , где – приближённая форма неравенства Фано, при этом обозначает носитель распределения ,
In information theory, Fano's inequality (also known as the Fano converse and the Fano lemma) relates the average information lost in a noisy channel to the probability of the categorization error. It was derived by Robert Fano in the early 1950s while teaching a Ph. D. seminar in information theory at MIT, and later recorded in his 1961 textbook. It is used to find a lower bound on the error probability of any decoder as well as the lower bounds for minimax risks in density estimation. Let the random variables and represent input and output messages with a joint probability Let represent an occurrence of error; i. e., that , with being an approximate version of Fano's inequality is
where denotes the support of ,
is the conditional entropy,
is the probability of the communication error, and
is the corresponding binary entropy.
является условной энтропией,
In information theory, Fano's inequality (also known as the Fano converse and the Fano lemma) relates the average information lost in a noisy channel to the probability of the categorization error. It was derived by Robert Fano in the early 1950s while teaching a Ph. D. seminar in information theory at MIT, and later recorded in his 1961 textbook. It is used to find a lower bound on the error probability of any decoder as well as the lower bounds for minimax risks in density estimation. Let the random variables and represent input and output messages with a joint probability Let represent an occurrence of error; i. e., that , with being an approximate version of Fano's inequality is
where denotes the support of ,
is the conditional entropy,
is the probability of the communication error, and
is the corresponding binary entropy.
– вероятность ошибки передачи, а
In information theory, Fano's inequality (also known as the Fano converse and the Fano lemma) relates the average information lost in a noisy channel to the probability of the categorization error. It was derived by Robert Fano in the early 1950s while teaching a Ph. D. seminar in information theory at MIT, and later recorded in his 1961 textbook. It is used to find a lower bound on the error probability of any decoder as well as the lower bounds for minimax risks in density estimation. Let the random variables and represent input and output messages with a joint probability Let represent an occurrence of error; i. e., that , with being an approximate version of Fano's inequality is
where denotes the support of ,
is the conditional entropy,
is the probability of the communication error, and
is the corresponding binary entropy.
– соответствующая бинарная энтропия.
In information theory, Fano's inequality (also known as the Fano converse and the Fano lemma) relates the average information lost in a noisy channel to the probability of the categorization error. It was derived by Robert Fano in the early 1950s while teaching a Ph. D. seminar in information theory at MIT, and later recorded in his 1961 textbook. It is used to find a lower bound on the error probability of any decoder as well as the lower bounds for minimax risks in density estimation. Let the random variables and represent input and output messages with a joint probability Let represent an occurrence of error; i. e., that , with being an approximate version of Fano's inequality is
where denotes the support of ,
is the conditional entropy,
is the probability of the communication error, and
is the corresponding binary entropy.
Интуиция
Неравенство Фано можно интерпретировать как способ разделения неопределенности условного распределения на два вопроса, задаваемых произвольным предиктором. Первый вопрос, соответствующий члену , относится к неопределенности, связанной с самим предиктором. Если предсказание верно, то дальнейшая неопределенность отсутствует. Если предсказание неверно, то неопределенность любого дискретного распределения ограничена сверху энтропией равномерного распределения по всем возможным вариантам, исключая неверное предсказание. Эта энтропия равна… Рассматривая крайние случаи, если предиктор всегда верен, то первый и второй члены неравенства равны 0, и существование идеального предиктора означает, что полностью определяется , а следовательно, и . Если же предиктор всегда ошибается, то первый член равен 0, и можно ограничить только сверху, используя равномерное распределение по оставшимся вариантам.