Введение
Ожидаемое количество информации, необходимое для определения результата стохастического источника данных.
In information theory, the entropy of a random variable is the average level of "information", "surprise", or "uncertainty" inherent to the variable's possible outcomes. Given a discrete random variable , which takes values in the alphabet and is distributed according to , the entropy is
where denotes the sum over the variable's possible values. The choice of base for , the logarithm, varies for different applications. Base 2 gives the unit of bits (or "shannons"), while base e gives "natural units" nat, and base 10 gives units of "dits", "bans", or "hartleys". An equivalent definition of entropy is the expected value of the self information of a variable. The concept of information entropy was introduced by Claude Shannon in his 1948 paper "A Mathematical Theory of Communication", and is also referred to as Shannon entropy. Shannon's theory defines a data communication system composed of three elements: a source of data, a communication channel, and a receiver. The "fundamental problem of communication" – as expressed by Shannon – is for the receiver to be able to identify what data was generated by the source, based on the signal it receives through the channel. In fact, log is the only function that satisfies а specific set of conditions defined in section
Hence, we can define the information, or surprisal, of an event by
or equivalently,
Entropy measures the expected (i. e., average) amount of information conveyed by identifying the outcome of a random trial. This implies that rolling a die has higher entropy than tossing a coin because each outcome of a die toss has smaller probability than each outcome of a coin toss
Consider a coin with probability p of landing on heads and probability 1 − p of landing on tails. The maximum surprise is when , for which one outcome is not expected over the other. In this case a coin flip has an entropy of one bit. (Similarly, one trit with equiprobable values contains (about 1.58496) bits of information because it can have one of three values.) The minimum surprise is when or , when the event outcome is known ahead of time, and the entropy is zero bits. When the entropy is zero bits, this is sometimes referred to as unity, where there is no uncertainty at all – no freedom of choice – no information. Other values of p give entropies between zero and one bits.
В теории информации энтропия случайной величины представляет собой средний уровень "информации", "удивления" или "неопределенности", присущий возможным исходам этой величины. Для дискретной случайной величины, принимающей значения из алфавита и распределенной согласно , энтропия определяется как
In information theory, the entropy of a random variable is the average level of "information", "surprise", or "uncertainty" inherent to the variable's possible outcomes. Given a discrete random variable , which takes values in the alphabet and is distributed according to , the entropy is
where denotes the sum over the variable's possible values. The choice of base for , the logarithm, varies for different applications. Base 2 gives the unit of bits (or "shannons"), while base e gives "natural units" nat, and base 10 gives units of "dits", "bans", or "hartleys". An equivalent definition of entropy is the expected value of the self information of a variable. The concept of information entropy was introduced by Claude Shannon in his 1948 paper "A Mathematical Theory of Communication", and is also referred to as Shannon entropy. Shannon's theory defines a data communication system composed of three elements: a source of data, a communication channel, and a receiver. The "fundamental problem of communication" – as expressed by Shannon – is for the receiver to be able to identify what data was generated by the source, based on the signal it receives through the channel. In fact, log is the only function that satisfies а specific set of conditions defined in section
Hence, we can define the information, or surprisal, of an event by
or equivalently,
Entropy measures the expected (i. e., average) amount of information conveyed by identifying the outcome of a random trial. This implies that rolling a die has higher entropy than tossing a coin because each outcome of a die toss has smaller probability than each outcome of a coin toss
Consider a coin with probability p of landing on heads and probability 1 − p of landing on tails. The maximum surprise is when , for which one outcome is not expected over the other. In this case a coin flip has an entropy of one bit. (Similarly, one trit with equiprobable values contains (about 1.58496) bits of information because it can have one of three values.) The minimum surprise is when or , when the event outcome is known ahead of time, and the entropy is zero bits. When the entropy is zero bits, this is sometimes referred to as unity, where there is no uncertainty at all – no freedom of choice – no information. Other values of p give entropies between zero and one bits.
где обозначает суммирование по всем возможным значениям величины. Выбор основания для логарифма варьируется в зависимости от области применения. Основание 2 дает единицу измерения в битах (или "шаннонах"), основание *e* – "натуральные единицы" (наты), а основание 10 – единицы "дитов", "бандов" или "хартли". Эквивалентное определение энтропии – это математическое ожидание самоинформации величины. Концепция информационной энтропии была введена Клодом Шенноном в его статье 1948 года «Математическая теория связи» и также называется энтропией Шеннона. Теория Шеннона определяет систему передачи данных, состоящую из трех элементов: источника данных, канала связи и приемника. «Фундаментальная проблема связи», как выразился Шеннон, заключается в том, чтобы приемник мог идентифицировать данные, сгенерированные источником, на основе сигнала, полученного по каналу. Фактически, логарифм – единственная функция, удовлетворяющая определенному набору условий, определенных в разделе.
In information theory, the entropy of a random variable is the average level of "information", "surprise", or "uncertainty" inherent to the variable's possible outcomes. Given a discrete random variable , which takes values in the alphabet and is distributed according to , the entropy is
where denotes the sum over the variable's possible values. The choice of base for , the logarithm, varies for different applications. Base 2 gives the unit of bits (or "shannons"), while base e gives "natural units" nat, and base 10 gives units of "dits", "bans", or "hartleys". An equivalent definition of entropy is the expected value of the self information of a variable. The concept of information entropy was introduced by Claude Shannon in his 1948 paper "A Mathematical Theory of Communication", and is also referred to as Shannon entropy. Shannon's theory defines a data communication system composed of three elements: a source of data, a communication channel, and a receiver. The "fundamental problem of communication" – as expressed by Shannon – is for the receiver to be able to identify what data was generated by the source, based on the signal it receives through the channel. In fact, log is the only function that satisfies а specific set of conditions defined in section
Hence, we can define the information, or surprisal, of an event by
or equivalently,
Entropy measures the expected (i. e., average) amount of information conveyed by identifying the outcome of a random trial. This implies that rolling a die has higher entropy than tossing a coin because each outcome of a die toss has smaller probability than each outcome of a coin toss
Consider a coin with probability p of landing on heads and probability 1 − p of landing on tails. The maximum surprise is when , for which one outcome is not expected over the other. In this case a coin flip has an entropy of one bit. (Similarly, one trit with equiprobable values contains (about 1.58496) bits of information because it can have one of three values.) The minimum surprise is when or , when the event outcome is known ahead of time, and the entropy is zero bits. When the entropy is zero bits, this is sometimes referred to as unity, where there is no uncertainty at all – no freedom of choice – no information. Other values of p give entropies between zero and one bits.
Следовательно, мы можем определить информацию, или удивление, от события как
In information theory, the entropy of a random variable is the average level of "information", "surprise", or "uncertainty" inherent to the variable's possible outcomes. Given a discrete random variable , which takes values in the alphabet and is distributed according to , the entropy is
where denotes the sum over the variable's possible values. The choice of base for , the logarithm, varies for different applications. Base 2 gives the unit of bits (or "shannons"), while base e gives "natural units" nat, and base 10 gives units of "dits", "bans", or "hartleys". An equivalent definition of entropy is the expected value of the self information of a variable. The concept of information entropy was introduced by Claude Shannon in his 1948 paper "A Mathematical Theory of Communication", and is also referred to as Shannon entropy. Shannon's theory defines a data communication system composed of three elements: a source of data, a communication channel, and a receiver. The "fundamental problem of communication" – as expressed by Shannon – is for the receiver to be able to identify what data was generated by the source, based on the signal it receives through the channel. In fact, log is the only function that satisfies а specific set of conditions defined in section
Hence, we can define the information, or surprisal, of an event by
or equivalently,
Entropy measures the expected (i. e., average) amount of information conveyed by identifying the outcome of a random trial. This implies that rolling a die has higher entropy than tossing a coin because each outcome of a die toss has smaller probability than each outcome of a coin toss
Consider a coin with probability p of landing on heads and probability 1 − p of landing on tails. The maximum surprise is when , for which one outcome is not expected over the other. In this case a coin flip has an entropy of one bit. (Similarly, one trit with equiprobable values contains (about 1.58496) bits of information because it can have one of three values.) The minimum surprise is when or , when the event outcome is known ahead of time, and the entropy is zero bits. When the entropy is zero bits, this is sometimes referred to as unity, where there is no uncertainty at all – no freedom of choice – no information. Other values of p give entropies between zero and one bits.
или, эквивалентно,
In information theory, the entropy of a random variable is the average level of "information", "surprise", or "uncertainty" inherent to the variable's possible outcomes. Given a discrete random variable , which takes values in the alphabet and is distributed according to , the entropy is
where denotes the sum over the variable's possible values. The choice of base for , the logarithm, varies for different applications. Base 2 gives the unit of bits (or "shannons"), while base e gives "natural units" nat, and base 10 gives units of "dits", "bans", or "hartleys". An equivalent definition of entropy is the expected value of the self information of a variable. The concept of information entropy was introduced by Claude Shannon in his 1948 paper "A Mathematical Theory of Communication", and is also referred to as Shannon entropy. Shannon's theory defines a data communication system composed of three elements: a source of data, a communication channel, and a receiver. The "fundamental problem of communication" – as expressed by Shannon – is for the receiver to be able to identify what data was generated by the source, based on the signal it receives through the channel. In fact, log is the only function that satisfies а specific set of conditions defined in section
Hence, we can define the information, or surprisal, of an event by
or equivalently,
Entropy measures the expected (i. e., average) amount of information conveyed by identifying the outcome of a random trial. This implies that rolling a die has higher entropy than tossing a coin because each outcome of a die toss has smaller probability than each outcome of a coin toss
Consider a coin with probability p of landing on heads and probability 1 − p of landing on tails. The maximum surprise is when , for which one outcome is not expected over the other. In this case a coin flip has an entropy of one bit. (Similarly, one trit with equiprobable values contains (about 1.58496) bits of information because it can have one of three values.) The minimum surprise is when or , when the event outcome is known ahead of time, and the entropy is zero bits. When the entropy is zero bits, this is sometimes referred to as unity, where there is no uncertainty at all – no freedom of choice – no information. Other values of p give entropies between zero and one bits.
Энтропия измеряет ожидаемое (т.е. среднее) количество информации, получаемое при определении исхода случайного испытания. Это означает, что бросок игральной кости имеет более высокую энтропию, чем подбрасывание монеты, поскольку каждый исход броска кости имеет меньшую вероятность, чем каждый исход подбрасывания монеты. Рассмотрим монету с вероятностью *p* выпадения орлом и вероятностью 1 − *p* выпадения решкой. Максимальное удивление достигается при , когда исходы равновероятны. В этом случае подбрасывание монеты имеет энтропию в один бит. (Аналогично, один трит с равновероятными значениями содержит (около 1,58496) битов информации, поскольку он может принимать одно из трех значений.) Минимальное удивление достигается при или , когда исход события известен заранее, и энтропия равна нулю битов. Когда энтропия равна нулю битов, это иногда называют состоянием определенности, где нет никакой неопределенности – нет свободы выбора – нет информации. Другие значения *p* дают энтропии между нулем и одним битом.
In information theory, the entropy of a random variable is the average level of "information", "surprise", or "uncertainty" inherent to the variable's possible outcomes. Given a discrete random variable , which takes values in the alphabet and is distributed according to , the entropy is
where denotes the sum over the variable's possible values. The choice of base for , the logarithm, varies for different applications. Base 2 gives the unit of bits (or "shannons"), while base e gives "natural units" nat, and base 10 gives units of "dits", "bans", or "hartleys". An equivalent definition of entropy is the expected value of the self information of a variable. The concept of information entropy was introduced by Claude Shannon in his 1948 paper "A Mathematical Theory of Communication", and is also referred to as Shannon entropy. Shannon's theory defines a data communication system composed of three elements: a source of data, a communication channel, and a receiver. The "fundamental problem of communication" – as expressed by Shannon – is for the receiver to be able to identify what data was generated by the source, based on the signal it receives through the channel. In fact, log is the only function that satisfies а specific set of conditions defined in section
Hence, we can define the information, or surprisal, of an event by
or equivalently,
Entropy measures the expected (i. e., average) amount of information conveyed by identifying the outcome of a random trial. This implies that rolling a die has higher entropy than tossing a coin because each outcome of a die toss has smaller probability than each outcome of a coin toss
Consider a coin with probability p of landing on heads and probability 1 − p of landing on tails. The maximum surprise is when , for which one outcome is not expected over the other. In this case a coin flip has an entropy of one bit. (Similarly, one trit with equiprobable values contains (about 1.58496) bits of information because it can have one of three values.) The minimum surprise is when or , when the event outcome is known ahead of time, and the entropy is zero bits. When the entropy is zero bits, this is sometimes referred to as unity, where there is no uncertainty at all – no freedom of choice – no information. Other values of p give entropies between zero and one bits.
Пример
Теория информации полезна для расчета минимального объема информации, необходимого для передачи сообщения, как, например, при сжатии данных. Рассмотрим, к примеру, передачу последовательностей, состоящих из 4 символов – 'A', 'B', 'C' и 'D' – по двоичному каналу. Если все 4 буквы равновероятны (по 25%), то оптимальным решением будет использование двух битов для кодирования каждой буквы. 'A' можно закодировать как '00', 'B' как '01', 'C' как '10', а 'D' как '11'. Однако, если вероятности каждой буквы различны, скажем, 'A' встречается с вероятностью 70%, 'B' – с 26%, а 'C' и 'D' – с 2% каждая, можно использовать коды переменной длины. В этом случае 'A' будет кодироваться как '0', 'B' как '10', 'C' как '110', а 'D' как '111'. При таком представлении 70% времени потребуется передавать только один бит, 26% времени – два бита и лишь 4% времени – три бита. В среднем требуется меньше 2 бит, поскольку энтропия ниже (благодаря высокой частоте встречаемости 'A' и 'B' – вместе они составляют 96% всех символов). Вычисление суммы взвешенных вероятностей логарифмов вероятностей позволяет измерить и учесть этот эффект. Английский текст, рассматриваемый как последовательность символов, обладает относительно низкой энтропией, то есть он достаточно предсказуем. Можно с уверенностью предположить, что, например, буква 'e' встречается гораздо чаще, чем 'z', сочетание 'qu' – гораздо чаще, чем любое другое сочетание с 'q', а сочетание 'th' – чаще, чем 'z', 'q' или 'qu'. После первых нескольких букв часто можно угадать остальную часть слова. Энтропия английского текста составляет от 0,6 до 1,3 бита на символ сообщения.
Альтернативная характеристика
Другая характеристика энтропии использует следующие свойства. Обозначим и
Непрерывность: H должен быть непрерывным, то есть небольшое изменение значений вероятностей должно приводить лишь к небольшому изменению энтропии. Симметрия: H должен оставаться неизменным при перестановке исходов xi. То есть, для любой перестановки
Максимум: должен быть максимальным, если все исходы равновероятны, то есть
Увеличение числа исходов: для равновероятных событий энтропия должна возрастать с увеличением числа исходов, то есть
Аддитивность: для ансамбля из n равномерно распределенных элементов, разделенного на k ящиков (подсистем) с b1, …, bk элементами в каждом, энтропия всего ансамбля должна быть равна сумме энтропии системы ящиков и индивидуальных энтропий ящиков, каждая из которых взвешена вероятностью нахождения в соответствующем ящике.
Continuity: H should be continuous, so that changing the values of the probabilities by a very small amount should only change the entropy by a small amount. Symmetry: H should be unchanged if the outcomes xi are re ordered. That is, for any permutation of Maximum: should be maximal if all the outcomes are equally likely i. e. Increasing number of outcomes: for equiprobable events, the entropy should increase with the number of outcomes i. e.
Additivity: given an ensemble of n uniformly distributed elements that are partitioned into k boxes (sub systems) with b1, , bk elements each, the entropy of the whole ensemble should be equal to the sum of the entropy of the system of boxes and the individual entropies of the boxes, each weighted with the probability of being in that particular box.
Ограничения энтропии в криптографии
В криптоанализе энтропия часто используется как приблизительная мера непредсказуемости криптографического ключа, хотя его истинная неопределенность не поддается измерению. Например, 128-битный ключ, сгенерированный равномерно и случайным образом, имеет 128 бит энтропии. Для его взлома методом грубой силы потребуется (в среднем) 2<sup>128</sup> попыток. Энтропия не позволяет оценить необходимое количество попыток, если возможные ключи не выбираются равномерно. Вместо этого для измерения усилий, необходимых для атаки методом грубой силы, можно использовать показатель, называемый "угадыванием". Другие проблемы могут возникнуть из-за неравномерных распределений, используемых в криптографии. Например, рассмотрим одноразовый блок из 1 000 000 двоичных цифр, использующий операцию исключающего ИЛИ. Если блок имеет 1 000 000 бит энтропии, он идеален. Если блок имеет 999 999 бит энтропии, равномерно распределенных (каждый отдельный бит блока имеет 0,999999 бит энтропии), он может обеспечить хорошую безопасность. Но если блок имеет 999 999 бит энтропии, при этом первый бит фиксирован, а остальные 999 999 бит совершенно случайны, первый бит шифротекста вообще не будет зашифрован.