Ағылшыншамен салыстырыңыз: абзацты басыңыз — түпнұсқа терезеде ашылады. Абзац астындағы EN түймесі оны мәтін ішінде көрсетеді.
Мазмұны
Кіріспе
Видеоны компрессиялау форматы, MPEG 1-ге жаңадан келген.
Video compression format, succeeds MPEG 1
H.262 немесе MPEG 2 Part 2 (ресми түрде ITU-T Recommendation H.262 және ISO/IEC 13818-2, сондай-ақ MPEG 2 Video деп белгілі) – ITU-T Study Group 16 Video Coding Experts Group (VCEG) және ISO/IEC Moving Picture Experts Group (MPEG) бірлесіп стандарттаған және қолдауды жалғастырып жатқан бейне кодтау форматы. Ол көптеген компаниялардың қатысуымен әзірленді. Бұл ISO/IEC MPEG 2 стандартының екінші бөлігі. ITU-T Recommendation H.262 және ISO/IEC 13818-2 құжаттары толық сәйкес келеді. Стандартты ITU-дан ақы төлеп алуға болады.
H.262 or MPEG 2 Part 2 (formally known as ITU T Recommendation H.262 and ISO/IEC 13818 2, also known as MPEG 2 Video) is a video coding format standardised and jointly maintained by ITU T Study Group 16 Video Coding Experts Group (VCEG) and ISO/IEC Moving Picture Experts Group (MPEG), and developed with the involvement of many companies. It is the second part of the ISO/IEC MPEG 2 standard. The ITU T Recommendation H.262 and ISO/IEC 13818 2 documents are identical. The standard is available for a fee from the ITU T
Суреттерден үлгі алу
8 биттік дискретизацияны қолданатын HDTV камерасы 25 кадр/секундық бейне үшін 25 × 1920 × 1080 × 3 = 155,520,000 байт/секундық шикі бейне ағынын жасайды (4:4:4 дискретизация форматын пайдалана отырып). Бұл дерек ағыны цифрлық теледидарды қолданыстағы телеарналардың еніне және фильмдерді DVD-ге сыйдыру үшін қысылуы тиіс. Бейнені қысу ыңғайлы, өйткені суреттердегі деректер көбінесе кеңістікте де, уақытта да артық болады. Мысалы, суреттің жоғарғы бөлігінде аспан көк болуы мүмкін, ал сол көк аспан кадрдан кадрға жалғаса береді. Сонымен қатар, көздің жұмыс істеу принципіне сәйкес, бейне суреттерінен кейбір деректерді жою немесе жуықтау мүмкін, бұл сурет сапасының айтарлықтай нашарлауына әкелмейді. Дерек көлемін азайтудың кең таралған (және ескі) тәсілі – әрбір толық бейне "кадрын" хабар тарату/кодирование кезінде екі "өріске" бөлу: "жоғарғы өріс" – тақ нөмірленген көлденең жолдар және "төменгі өріс" – жұп нөмірленген жолдар. Қабылдау/декодирование кезінде екі өріс кезекпен көрсетіледі, бір өрістің жолдары алдыңғы өрістің жолдары арасында өріледі; бұл формат интерлейс видео деп аталады. Типтік өріс жылдамдығы 50 (Еуропа/PAL) немесе 59,94 (АҚШ/NTSC) өріс/секунд, бұл 25 (Еуропа/PAL) немесе 29,97 (Солтүстік Америка/NTSC) толық кадр/секундқа сәйкес келеді. Егер бейне интерлейс болмаса, онда ол прогрессивті сканерлеу бейнесі деп аталады және әрбір сурет толық кадр болып табылады. MPEG 2 екі опцияны да қолдайды. Цифрлық теледидар осы суреттерді компьютерлік жабдықпен өңдеу үшін цифрлық форматқа келтіруді талап етеді. Әрбір сурет элементі (пиксель) бір жарықтық санымен және екі түс санымен бейнеленеді. Бұл пикселдің жарықтығы мен түсін сипаттайды (YCbCr қараңыз). Осылайша, әрбір цифрлық сурет бастапқыда үш тіктөртбұрышты сандар массивімен бейнеленеді. Өңделетін дерек көлемін азайту үшін тағы бір кең таралған тәсіл – екі түстік жазықтықты қосалқы дискретизациялау (алиасингтің алдын алу үшін төмен өткізгіш сүзгіден өткеннен кейін). Бұл жұмыс істейді, өйткені адамның көру жүйесі түстердің тоны мен қанықтығына қарағанда жарықтықтың егжей-тегжейін жақсырақ анықтайды. 4:2:2 термині түстік жазықтықтың көлденең бағытта 2:1 қатынасымен қосалқы дискретизацияланған бейне үшін қолданылады, ал 4:2:0 термині түстік жазықтықтың көлденең және тік бағытта 2:1 қатынасымен қосалқы дискретизацияланған бейне үшін қолданылады. Жарықтық пен түстің бірдей ажыратымдылығы бар бейне 4:4:4 деп аталады. MPEG 2 бейне құжаты барлық үш дискретизация түрін қарастырады, бірақ 4:2:0 тұтынушылық бейне үшін ең көп таралған, ал MPEG 2 4:4:4 бейнесі үшін анықталған "профильдер" жоқ (профильдер туралы толық ақпарат алу үшін төменде қараңыз). Бұл бөлімдегі талқылау негізінен MPEG 2 бейне қысылуын сипаттайды, бірақ өрістер, түстік форматтар, көрініс өзгерістеріне жауаптар, бит ағынының бөліктерін белгілейтін арнайы кодтар және басқа да ақпараттар сияқты талқыланбаған көптеген егжей-тегжейлі мәліметтер бар. Интерлейс кодтау үшін өрістерді өңдеу мүмкіндіктерін қоспағанда, MPEG 2 бейнесі MPEG 1 бейнесіне (тіпті бұрынғы H.261 стандартына да) өте ұқсас, сондықтан төмендегі сипаттама MPEG 1-ге де бірдей қолданылады.
An HDTV camera with 8 bit sampling generates a raw video stream of 25 × 1920 × 1080 × 3 = 155,520,000 bytes per second for 25 frame per second video (using the 4:4:4 sampling format). This stream of data must be compressed if digital TV is to fit in the bandwidth of available TV channels and if movies are to fit on DVDs. Video compression is practical because the data in pictures is often redundant in space and time. For example, the sky can be blue across the top of a picture and that blue sky can persist for frame after frame. Also, because of the way the eye works, it is possible to delete or approximate some data from video pictures with little or no noticeable degradation in image quality. A common (and old) trick to reduce the amount of data is to separate each complete "frame" of video into two "fields" upon broadcast/encoding: the "top field", which is the odd numbered horizontal lines, and the "bottom field", which is the even numbered lines. Upon reception/decoding, the two fields are displayed alternately with the lines of one field interleaving between the lines of the previous field; this format is called interlaced video. The typical field rate is 50 (Europe/PAL) or 59.94 (US/NTSC) fields per second, corresponding to 25 (Europe/PAL) or 29.97 (North America/NTSC) whole frames per second. If the video is not interlaced, then it is called progressive scan video and each picture is a complete frame. MPEG 2 supports both options. Digital television requires that these pictures be digitized so that they can be processed by computer hardware. Each picture element (a pixel) is then represented by one luma number and two chroma numbers. These describe the brightness and the color of the pixel (see YCbCr). Thus, each digitized picture is initially represented by three rectangular arrays of numbers. Another common practice to reduce the amount of data to be processed is to subsample the two chroma planes (after low pass filtering to avoid aliasing). This works because the human visual system better resolves details of brightness than details in the hue and saturation of colors. The term 4:2:2 is used for video with the chroma subsampled by a ratio of 2:1 horizontally, and 4:2:0 is used for video with the chroma subsampled by 2:1 both vertically and horizontally. Video that has luma and chroma at the same resolution is called 4:4:4. The MPEG 2 Video document considers all three sampling types, although 4:2:0 is by far the most common for consumer video, and there are no defined "profiles" of MPEG 2 for 4:4:4 video (see below for further discussion of profiles). While the discussion below in this section generally describes MPEG 2 video compression, there are many details that are not discussed, including details involving fields, chrominance formats, responses to scene changes, special codes that label the parts of the bitstream, and other pieces of information. Aside from features for handling fields for interlaced coding, MPEG 2 Video is very similar to MPEG 1 Video (and even quite similar to the earlier H.261 standard), so the entire description below applies equally well to MPEG 1.
I-каркастар, P-каркастар және B-каркастар
MPEG 2 кодталған кадрлардың үш негізгі түрін қамтиды: ішкі кодталған кадрлар (I кадрлар), болжамды кодталған кадрлар (P кадрлар) және екі жақты болжамды кодталған кадрлар (B кадрлар). I кадр – бұл бір ғана сығымдалмаған (шыкі) кадрдың жеке сығымдалған нұсқасы. I кадрды кодтау кеңістіктік артықшылықты және көзге кескіндегі кейбір өзгерістерді байқау мүмкін еместігін пайдаланады. P және B кадрларынан айырмашылығы, I кадрлар алдыңғы немесе келесі кадрлардағы деректерге тәуелді емес, сондықтан олардың кодтамасы тұрақты суретті кодтауға өте ұқсас (JPEG суреттік кодтамасына ұқсас). Қысқаша айтқанда, шикі кадр 8 пикселге 8 пикселдік блоктарға бөлінеді. Әр блоктан алынған деректер дискретті косинус түрлендіруімен (DCT) түрлендіріледі. Нәтижесінде 8×8 матрица пайда болады, онда нақты сандық мәндер болады. Түрлендіру кеңістіктік өзгерістерді жиіліктік өзгерістерге айналдырады, бірақ блоктан алынған ақпаратты өзгертпейді; егер түрлендіру өте дәл есептелсе, бастапқы блок кері косинус түрлендіруін қолдану арқылы дәл қайта жаңартылуы мүмкін (сондай-ақ өте дәлдікпен). 8 биттік бүтін сандарды нақты бағаланған түрлендіру коэффициенттеріне түрлендіру осы өңдеу кезеңінде қолданылатын деректердің көлемін ұлғайтады, бірақ түрлендірудің артықшылығы – сурет деректерін коэффициенттерді кванттау арқылы жуықтауға болады. Кванттаудан кейін көптеген түрлендіру коэффициенттері, әдетте жоғары жиілікті компоненттер, нөлге тең болады, бұл негізінен дөңгелектеу операциясы. Бұл қадамның кемшілігі – жарықтық пен түстің кейбір ең ұсақ айырмашылықтарының жоғалуы. Кванттау енгізуші таңдағанша, ол дөңгелек немесе ұсақ болуы мүмкін. Егер кванттау тым дөңгелек болмаса және кванттаудан кейін матрицаға кері түрлендіру қолданса, бастапқы суретке өте ұқсас, бірақ толыққанды бірдей емес сурет алынады. Содан кейін квантталған коэффициенттер матрицасы өзі сығылады. Әдетте, 8×8 коэффициенттер массивінің бір бұрышында кванттау қолданғаннан кейін тек нөлдер болады. Матрицаның қарама-қарсы бұрышынан бастап, содан кейін матрицаны зигзаг тәрізді кесіп, коэффициенттерді тізбекке біріктіріп, содан кейін тізбектегі қатарынан келетін нөлдерді орындалу ұзындығы кодтарымен алмастырып, содан кейін осы нәтижеге Хаффман кодтамасын қолдану арқылы матрица кішірек көлемдегі деректерге дейін азайтылады. Осылайша алынған энтропиялық кодталған деректер таратылады немесу DVD дискісіне жазылады. Қабылдағышта немесе ойнатқышта бүкіл процесс кері қайтарылады, бұл қабылдағышқа бастапқы кадрды шамамен қалпына келтіруге мүмкіндік береді. B кадрларын өңдеу P кадрларын өңдеуге ұқсас, бірақ B кадрлары келесі анықтамалық кадрдағы суретті, сондай-ақ алдыңғы анықтамалық кадрдағы суретті пайдаланады. Нәтижесінде, B кадрлары әдетте P кадрларына қарағанда көбірек сығылуды қамтамасыз етеді. B кадрлары MPEG 2 бейнесінде сілтеме кадрлары болып табылмайды. Әдетте, әрбір 15-ші кадр немесе одан да көп I кадрға айналады. P және B кадрлары I кадрдан кейін осындай тізбекпен келе алады: IBBPBBPBBPBB(I), суреттер тобын (GOP) құру үшін; алайда стандарт мұнда икемді. Кодтаушы I, P және B кадрлары ретінде қай суреттерді кодтау керектігін таңдайды.
MPEG 2 includes three basic types of coded frames: intra coded frames (I frames), predictive coded frames (P frames), and bidirectionally predictive coded frames (B frames). An I frame is a separately compressed version of a single uncompressed (raw) frame. The coding of an I frame takes advantage of spatial redundancy and of the inability of the eye to detect certain changes in the image. Unlike P frames and B frames, I frames do not depend on data in the preceding or the following frames, and so their coding is very similar to how a still photograph would be coded (roughly similar to JPEG picture coding). Briefly, the raw frame is divided into 8 pixel by 8 pixel blocks. The data in each block is transformed by the discrete cosine transform (DCT). The result is an 8×8 matrix of coefficients that have real number values. The transform converts spatial variations into frequency variations, but it does not change the information in the block; if the transform is computed with perfect precision, the original block can be recreated exactly by applying the inverse cosine transform (also with perfect precision). The conversion from 8 bit integers to real valued transform coefficients actually expands the amount of data used at this stage of the processing, but the advantage of the transformation is that the image data can then be approximated by quantizing the coefficients. Many of the transform coefficients, usually the higher frequency components, will be zero after the quantization, which is basically a rounding operation. The penalty of this step is the loss of some subtle distinctions in brightness and color. The quantization may either be coarse or fine, as selected by the encoder. If the quantization is not too coarse and one applies the inverse transform to the matrix after it is quantized, one gets an image that looks very similar to the original image but is not quite the same. Next, the quantized coefficient matrix is itself compressed. Typically, one corner of the 8×8 array of coefficients contains only zeros after quantization is applied. By starting in the opposite corner of the matrix, then zigzagging through the matrix to combine the coefficients into a string, then substituting run length codes for consecutive zeros in that string, and then applying Huffman coding to that result, one reduces the matrix to a smaller quantity of data. It is this entropy coded data that is broadcast or that is put on DVDs. In the receiver or the player, the whole process is reversed, enabling the receiver to reconstruct, to a close approximation, the original frame. The processing of B frames is similar to that of P frames except that B frames use the picture in a subsequent reference frame as well as the picture in a preceding reference frame. As a result, B frames usually provide more compression than P frames. B frames are never reference frames in MPEG 2 Video. Typically, every 15th frame or so is made into an I frame. P frames and B frames might follow an I frame like this, IBBPBBPBBPBB(I), to form a Group of Pictures (GOP); however, the standard is flexible about this. The encoder selects which pictures are coded as I , P , and B frames.
Макроблоктар
P-кадрлар I-кадрлардан көбірек қысуды қамтамасыз етеді, себебі олар алдыңғы I-кадрдың немесе P-кадрдың деректерін пайдаланады – эталондық кадр. P-кадрды жасау үшін алдыңғы эталондық кадр қайта құрылады, дәл теледидар қабылдағышында немесе DVD ойнатқышында болатындай. Қысылатын кадр 16 пикселге 16 пикселдік макроблоктарға бөлінеді. Содан кейін, әрбір макроблок үшін қайта құрылған эталондық кадрда, қысылатын макроблоктың мазмұнына ең жақын келетін 16х16 аймақ ізделеді. Бұл қадамның өзгеруі "қозғалыс векторы" ретінде кодталады. Көбінесе, өзгеру нөлге тең болады, бірақ егер суретте қозғалыс болса, өзгеру оңға 23 пиксел және жоғарыға 4,5 пиксел болуы мүмкін. MPEG 1 және MPEG 2 стандартында қозғалыс векторлары толық сан немесе жарты сан түрінде болуы мүмкін. Екі аймақ арасындағы сәйкестік көбінесе толық болмайды. Бұл қателікті түзету үшін, кодтаушы екі аймақтың барлық сәйкес пикселдерінің айырмасын есептейді, содан кейін осы макроблок айырмасы бойынша жоғарыда сипатталғандай, 16х16 макроблоктағы төрт 8х8 аймақ үшін DCT және коэффициенттер тізбегін есептейді. Бұл "қалдық" қозғалыс векторына қосылады және нәтиже қабылдағышқа жіберіледі немесе әрбір қысылған макроблок үшін DVD-ге сақталады. Кейде қолайлы сәйкестік табылмайды. Онда макроблок I-кадр макроблогы сияқты қарастырылады.
P frames provide more compression than I frames because they take advantage of the data in a previous I frame or P frame – a reference frame. To generate a P frame, the previous reference frame is reconstructed, just as it would be in a TV receiver or DVD player. The frame being compressed is divided into 16 pixel by 16 pixel macroblocks. Then, for each of those macroblocks, the reconstructed reference frame is searched to find a 16 by 16 area that closely matches the content of the macroblock being compressed. The offset is encoded as a "motion vector". Frequently, the offset is zero, but if something in the picture is moving, the offset might be something like 23 pixels to the right and 4 and a half pixels up. In MPEG 1 and MPEG 2, motion vector values can either represent integer offsets or half integer offsets. The match between the two regions will often not be perfect. To correct for this, the encoder takes the difference of all corresponding pixels of the two regions, and on that macroblock difference then computes the DCT and strings of coefficient values for the four 8×8 areas in the 16×16 macroblock as described above. This "residual" is appended to the motion vector and the result sent to the receiver or stored on the DVD for each macroblock being compressed. Sometimes no suitable match is found. Then, the macroblock is treated like an I frame macroblock.
Патент иелері
MPEG LA тізімінде көрсетілгендей, MPEG 2 бейне технологиясына патенттерге ие болған ұйымдар төменде келтірілген. Бұл патенттердің бәрі АҚШ-та және көптеген басқа елдерде енді қолданылмайды.
The following organizations have held patents for MPEG 2 video technology, as listed at MPEG LA. All of these patents are now expired in the US and most other territories.