Ағылшыншамен салыстырыңыз: абзацты басыңыз — түпнұсқа терезеде ашылады. Абзац астындағы EN түймесі оны мәтін ішінде көрсетеді.
Мазмұны
Кіріспе
Инструкциялық құбыржол
Instruction pipeline
Компьютерлік аппараттар тарихында, кейбір ертедегі қысқартылған нұсқаулар жиынтығы компьютерлік орталық процессорлары (RISC CPU) қазір классикалық RISC құбыржолы деп аталатын өте ұқсас архитектуралық шешімді қолданды. Осы процессорлар: MIPS, SPARC, Motorola 88000, және кейіннен білім беру мақсатында жасалған DLX процессоры. Бұл классикалық скалярлық RISC жүйелерінің әрқайсысы цикл сайын бір нұсқауды іздеп, орындауға тырысады. Әрбір жобаның негізгі ортақ ұғымы – бес кезеңді нұсқауларды орындау құбыржолы. Жұмыс істеу кезінде құбыржолдың әрбір кезеңі бір уақытта бір нұсқаумен жұмыс жасайды. Осы кезеңдердің әрқайсысы күйді сақтауға арналған триггерлер жиынтығынан және осы триггерлердің шығыстарымен жұмыс істейтін комбинациялық логикадан тұрады.
In the history of computer hardware, some early reduced instruction set computer central processing units (RISC CPUs) used a very similar architectural solution, now called a classic RISC pipeline. Those CPUs were: MIPS, SPARC, Motorola 88000, and later the notional CPU DLX invented for education. Each of these classic scalar RISC designs fetches and tries to execute one instruction per cycle. The main common concept of each design is a five stage execution instruction pipeline. During operation, each pipeline stage works on one instruction at a time. Each of these stages consists of a set of flip flops to hold state, and combinational logic that operates on the outputs of those flip flops.
Нұсқауды алу
Нұсқаулар бір циклді оқуға қажетті жадта сақталады. Бұл жад SRAM-ға немесе нұсқаулық кэшке арналуы мүмкін. Компьютерлік ғылымда "латенция" термині жиі қолданылады және ол операция басталғаннан аяқталғанға дейінгі уақытты білдіреді. Осылайша, нұсқауды алудың латенциясы бір сағат циклына тең (бір циклдік SRAM қолдансаңыз немесе нұсқау кэште болса). Осылайша, Instruction Fetch кезеңінде 32 биттік нұсқаулық жадтан алынады. Бағдарлама санағышы (Program Counter), немесе PC – нұсқаулық жадқа ұсынылатын мекенжайды сақтайтын регистр. Адрес циклдің басында нұсқаулық жадқа беріледі. Содан кейін цикл ішінде нұсқаулық жадтан нұсқау оқылады, және бір уақытта келесі PC-ні анықтау үшін есептеу жүргізіледі. Келесі PC, PC-ні 4-ке арттыру арқылы және осыны келесі PC ретінде қабылдау немесе тармақтану/секіру есептеуінің нәтижесін келесі PC ретінде қабылдау арқылы есептеледі. Классикалық RISC-де барлық нұсқаулардың ұзындығы бірдей болады. (Бұл RISC-ті CISC-тен ажырататын бір ерекшелік). Бастапқы RISC жобаларында нұсқаудың көлемі 4 байт, сондықтан нұсқаудың мекенжайына әрқашан 4 қосыңыз, бірақ тармақтану, секіру немесе қателік болған жағдайда PC + 4 қолданбаңыз (төмендегі кідірілген тармақтарды қараңыз). (Кейбір заманауи құрылғылар келесі нұсқаулық мекенжайын болжау үшін күрделі алгоритмдерді (тармақ болжау және тармақ мақсаты болжау) қолданатынын ескеріңіз.)
The instructions reside in memory that takes one cycle to read. This memory can be dedicated to SRAM, or an Instruction Cache. The term "latency" is used in computer science often and means the time from when an operation starts until it completes. Thus, instruction fetch has a latency of one clock cycle (if using single cycle SRAM or if the instruction was in the cache). Thus, during the Instruction Fetch stage, a 32 bit instruction is fetched from the instruction memory. The Program Counter, or PC is a register that holds the address that is presented to the instruction memory. The address is presented to instruction memory at the start of a cycle. Then during the cycle, the instruction is read out of instruction memory, and at the same time, a calculation is done to determine the next PC. The next PC is calculated by incrementing the PC by 4, and by choosing whether to take that as the next PC or to take the result of a branch/jump calculation as the next PC. Note that in classic RISC, all instructions have the same length. (This is one thing that separates RISC from CISC ). In the original RISC designs, the size of an instruction is 4 bytes, so always add 4 to the instruction address, but don't use PC + 4 for the case of a taken branch, jump, or exception (see delayed branches, below). (Note that some modern machines use more complicated algorithms (branch prediction and branch target prediction) to guess the next instruction address.)
Нұсқаулық кодтау
Тағы бір мәселе, алғашқы RISC машиналарын ескі CISC машиналарыннан ерекшелендіретін нәрсе – RISC-те микрокодтың болмауы. CISC микрокодталған нұсқауларының жағдайында, нұсқаулар кэшінен алынғаннан кейін, нұсқау биттері құбыр арқылы жылжытылады, онда құбырдың әрбір сатысындағы қарапайым комбинациялық логика тікелей нұсқау биттерінен дерек жолы үшін басқару сигналдарын жасайды. CISC құрылымдарында, дәстүрлі түрде «декодирование» сатысы деп аталатын кезеңде өте аз декодирование жасалады. Декодированиедің жеткіліксіздігінен нұсқаудың не істейтінін көрсету үшін көбірек нұсқау биттерін пайдалану қажет. Бұл регистр индекстері сияқты мәліметтерге азырақ биттерді қалдырады. Барлық MIPS, SPARC және DLX нұсқауларында ең көп дегенде екі регистрлік кіріс болады. Декодирование сатысында осы екі регистрдің индекстері нұсқаудың ішінде анықталады және индекстер регистр жадына адрес ретінде жіберіледі. Осылайша, аталған екі регистр тізілім файлынан оқылады. MIPS құрылымында тізілім файлында 32 жазба бар. Тізілім файлы оқылғанда, осы сатызда нұсқауды орындауға құбыр дайын екенін анықтайтын нұсқауды беру логикасы жұмыс істейді. Егер дайын болмаса, нұсқауды беру логикасы нұсқауды алу және декодирование сатысын тоқтатуға себеп болады. Тоқтату циклында кіріс флип-флоптары жаңа биттерді қабылдамайды, сондықтан сол циклде жаңа есептеулер орындалмайды. Егер декодированиеленген нұсқау тармақтану немесе секіру болса, тармақтанудың немесе секірудің мақсатты адресі тізілім файлын оқумен бірге есептеледі. Тармақтану шарты келесі циклда есептеледі (тізілім файлы оқылғаннан кейін), және егер тармақтану орын алса немесе нұсқау секіру болса, бірінші сатыздағы PC есептелген PC-ге емес, тармақтану мақсатына тағайындалады. Кейбір архитектуралар Арифметикалық-логикалық құрылғыны (ALU) Орындау сатысында пайдаланды, бірақ бұл нұсқаулардың өнімділігін сәл төмендетті. Декодирование сатысы көптеген аппараттық құралдармен аяқталды: MIPS екі регистр тең болса тармақтану мүмкіндігіне ие, сондықтан 32 биттік AND ағашы тізілім файлын оқығаннан кейін тізбектей жұмыс істейді, бұл сатыздан өте ұзын сындық жолға әкеледі (яғни секундқа азырақ циклдер). Сонымен қатар, тармақтану мақсатын есептеу әдетте 16 биттік қосу және 14 биттік инкрементаторды қажет етеді. Декодирование сатысында тармақтануды шешу бір циклдық тармақ қателігін болжауға мүмкіндік берді. Тармақтанулар жиі орын алғандықтан (осылайша қате болжау жиі кездесетін), бұл жазаны төмен ұстау өте маңызды болды.
Another thing that separates the first RISC machines from earlier CISC machines, is that RISC has no microcode. In the case of CISC micro coded instructions, once fetched from the instruction cache, the instruction bits are shifted down the pipeline, where simple combinational logic in each pipeline stage produces control signals for the datapath directly from the instruction bits. In those CISC designs, very little decoding is done in the stage traditionally called the decode stage. A consequence of this lack of decoding is that more instruction bits have to be used to specifying what the instruction does. That leaves fewer bits for things like register indices. All MIPS, SPARC, and DLX instructions have at most two register inputs. During the decode stage, the indexes of these two registers are identified within the instruction, and the indexes are presented to the register memory, as the address. Thus the two registers named are read from the register file. In the MIPS design, the register file had 32 entries. At the same time the register file is read, instruction issue logic in this stage determines if the pipeline is ready to execute the instruction in this stage. If not, the issue logic causes both the Instruction Fetch stage and the Decode stage to stall. On a stall cycle, the input flip flops do not accept new bits, thus no new calculations take place during that cycle. If the instruction decoded is a branch or jump, the target address of the branch or jump is computed in parallel with reading the register file. The branch condition is computed in the following cycle (after the register file is read), and if the branch is taken or if the instruction is a jump, the PC in the first stage is assigned the branch target, rather than the incremented PC that has been computed. Some architectures made use of the Arithmetic logic unit (ALU) in the Execute stage, at the cost of slightly decreased instruction throughput. The decode stage ended up with quite a lot of hardware: MIPS has the possibility of branching if two registers are equal, so a 32 bit wide AND tree runs in series after the register file read, making a very long critical path through this stage (which means fewer cycles per second). Also, the branch target computation generally required a 16 bit add and a 14 bit incrementer. Resolving the branch in the decode stage made it possible to have just a single cycle branch mis predict penalty. Since branches were very often taken (and thus mis predicted), it was very important to keep this penalty low.
Жадына қатынау
Егер дерек жадына қол жеткізу қажет болса, ол осы кезеңде жүзеге асырылады. Осы кезеңде бір циклдік кешігу нұсқаулары өз нәтижелерін тікелей келесі кезеңге жібереді. Бұл жіберу бір және екі циклдік нұсқаулардың нәтижелерін құбырдың бір кезеңіне жазуын қамтамасыз етеді, соның арқасында тіркелімдер жиымына тек бір жазу порты қолданылады және ол әрқашан бос болады. Тікелей бейнелеу және виртуалды таңбалау схемасын қолданатын дерек кэштеуінде, көптеген дерек кэші ұйымдарының ішіндегі ең қарапайымында, екі SRAM пайдаланылады – біреуі деректерді, екіншісі таңбаларды сақтайды.
If data memory needs to be accessed, it is done in this stage. During this stage, single cycle latency instructions simply have their results forwarded to the next stage. This forwarding ensures that both one and two cycle instructions always write their results in the same stage of the pipeline so that just one write port to the register file can be used, and it is always available. For direct mapped and virtually tagged data caching, the simplest by far of the numerous data cache organizations, two SRAMs are used, one storing data and the other storing tags.
Қайта жазу
Бұл кезеңде бір циклді және екі циклді нұсқаулар өз нәтижелерін регистрлік файлға жазады. Екі түрлі кезең бір уақытта регистрлік файлға қол жеткізеді: декодтау кезеңі екі бастапқы регистрді оқиды, ал кері жазу кезеңі бұрынғы нұсқаудың мақсаттық регистрін жазады. Нағыз кремнийде бұл қауіп тудыруы мүмкін (қауіптер туралы толығырақ ақпарат алу үшін төмен қараңыз). Себебі декодтау кезеңінде оқылатын бастапқы регистрлердің бірі, кері жазу кезеңінде жазылатын мақсаттық регистрмен сәйкес келуі мүмкін. Ондай жағдайда регистрлік файлдағы жад жасушалары бір уақытта оқылып та, жазылып та тұрады. Кремнийдегі жад жасушаларының көптеген түрлері бір уақытта оқылып, жазылғанда дұрыс жұмыс істемейді.
During this stage, both single cycle and two cycle instructions write their results into the register file. Note that two different stages are accessing the register file at the same time—the decode stage is reading two source registers, at the same time that the writeback stage is writing a previous instruction's destination register. On real silicon, this can be a hazard (see below for more on hazards). That is because one of the source registers being read in decode might be the same as the destination register being written in writeback. When that happens, then the same memory cells in the register file are being both read and written the same time. On silicon, many implementations of memory cells will not operate correctly when read and written at the same time.
Қауіптер
Хеннесси мен Паттерсон құбыржолдағы командалар қате нәтижелерге әкелетін жағдайлар үшін «қауіп» терминін енгізді.
Hennessy and Patterson coined the term hazard for situations where instructions in a pipeline would produce wrong answers.
Құрылымдық қауіптер
Құрылымдық қауіптер екі нұсқау бір уақытта бірдей ресурстарды пайдалануға тырысқанда туындайды. Классикалық RISC құбырлары осы қауіптерден аппараттық құралдарды көбейту арқылы сақтануға тырысты. Атап айтқанда, тармақталу нұсқаулары тармақтың мақсатты мекенжайын есептеу үшін ALU-ды пайдаланатын. Егер ALU осы мақсатта декодтау кезеңінде қолданылса, онда ALU нұсқауын ұстанғаннан кейін тармақталу нұсқаулары екеуі де бір уақытта ALU-ды пайдалануға тырысқан болар еді. Бұл қақтығысты декодтау сатысына арнайы тармақ мақсатты қосу құрылғысын енгізу арқылы оңай шешуге болады.
Structural hazards occur when two instructions might attempt to use the same resources at the same time. Classic RISC pipelines avoided these hazards by replicating hardware. In particular, branch instructions could have used the ALU to compute the target address of the branch. If the ALU were used in the decode stage for that purpose, an ALU instruction followed by a branch would have seen both instructions attempt to use the ALU simultaneously. It is simple to resolve this conflict by designing a specialized branch target adder into the decode stage.
Ерекшеліктер
32 биттік RISC екі үлкен санды қосу үшін ADD нұсқауын өңдейді, ал нәтиже 32 битке сыймайды. Көптеген архитектуралар ұсынатын ең қарапайым шешім – оралымдық арифметика. Мүмкіндігінше ең жоғары кодталған мәннен үлкен сандардың жоғары разряды кесіліп, сыюға дейін азайтылады. Кәдімгі бүтін сандар жүйесінде 3000000000+3000000000=6000000000 болады. Ал 32 биттік оралымдық арифметикада 3000000000+3000000000=1705032704 (6000000000 mod 2^32) болады. Бұл аса пайдалы көрінбесе де, оралымдық арифметиканың ең үлкен артықшылығы – әрбір операцияның нақты анықталған нәтижесі бар. Бірақ бағдарламашы, әсіресе үлкен бүтін сандарды қолдайтын тілде (мысалы, Lisp немесе Scheme) бағдарламаласа, оралымдық арифметиканы қаламайды. Кейбір архитектуралар (мысалы, MIPS) нәтижені ораудың орнына, ағын кезінде ерекше жағдайларда арнайы орындарға өтетін қосымша операцияларды анықтайды. Осы жағдайда, мақсатты жердегі бағдарламалық қамтамасыз ету мәселені шешуге жауапты болады. Бұл ерекше өту ерекше жағдай деп аталады. Ерекше жағдайлар кәдімгі өтулерден өзгеше, себебі мақсатты мекенжай нұсқаудың өзінде көрсетілмейді және өту шешімі нұсқаудың нәтижесіне байланысты болады. Классикалық RISC машиналарының біріндегі бағдарламалық қамтамасыз етуде ең көп кездесетін ерекше жағдай – TLB қатесі. Ерекше жағдайлар өтулер мен секірулерден де өзгеше, себебі басқа басқару ағыны өзгерістері кодтау кезеңінде шешіледі. Ал ерекше жағдайлар қайта жазу кезеңінде шешіледі. Ерекше жағдай анықталғанда, одан кейінгі нұсқаулар (құбырдың басында) жарамсыз деп белгіленеді және құбырдың соңына жеткенде олардың нәтижелері жойылады. Бағдарлама санағы ерекше жағдайды өңдеушінің мекенжайына орнатылады, ал ерекше жағдайдың орналасуы мен себебі арнайы тіркегіштерге жазылады. Бағдарламалық қамтамасыз етудің мәселені оңай (және жылдам) шешіп, бағдарламаны қайта іске қосуы үшін процессор нақты ерекше жағдайды анықтауы керек. Нақты ерекше жағдай дегеніміз, ерекше жағдайға дейінгі барлық нұсқаулар орындалған, ал ерекше жағдай және одан кейінгі нұсқаулар орындалмаған. Нақты ерекше жағдайларды анықтау үшін процессор бағдарламалық тәртіппен бағдарламалық қамтамасыз етудің көрінетін күйіне өзгерістер енгізуі керек. Бұл тәртіп классикалық RISC құбыржолында өте табиғи түрде жүзеге асырылады. Көптеген нұсқаулар өз нәтижелерін тіркегіштерге қайта жазу кезеңінде жазады, сондықтан бұл жазулар автоматты түрде бағдарламалық тәртіппен орындалады. Дегенмен, сақтау нұсқаулары өз нәтижелерін қол жеткізу кезеңінде сақтау деректері кезегіне жазады. Егер сақтау нұсқаулығы ерекше жағдайға тап болса, сақтау деректері кезегінің жазбасы жарамсыз деп танылады, сондықтан ол кейінірек SRAM деректеріне жазылмайды.
Suppose a 32 bit RISC processes an ADD instruction that adds two large numbers, and the result does not fit in 32 bits. The simplest solution, provided by most architectures, is wrapping arithmetic. Numbers greater than the maximum possible encoded value have their most significant bits chopped off until they fit. In the usual integer number system, 3000000000+3000000000=6000000000. With unsigned 32 bit wrapping arithmetic, 3000000000+3000000000=1705032704 (6000000000 mod 2^32). This may not seem terribly useful. The largest benefit of wrapping arithmetic is that every operation has a well defined result. But the programmer, especially if programming in a language supporting large integers (e. g. Lisp or Scheme), may not want wrapping arithmetic. Some architectures (e. g. MIPS), define special addition operations that branch to special locations on overflow, rather than wrapping the result. Software at the target location is responsible for fixing the problem. This special branch is called an exception. Exceptions differ from regular branches in that the target address is not specified by the instruction itself, and the branch decision is dependent on the outcome of the instruction. The most common kind of software visible exception on one of the classic RISC machines is a TLB miss. Exceptions are different from branches and jumps, because those other control flow changes are resolved in the decode stage. Exceptions are resolved in the writeback stage. When an exception is detected, the following instructions (earlier in the pipeline) are marked as invalid, and as they flow to the end of the pipe their results are discarded. The program counter is set to the address of a special exception handler, and special registers are written with the exception location and cause. To make it easy (and fast) for the software to fix the problem and restart the program, the CPU must take a precise exception. A precise exception means that all instructions up to the excepting instruction have been executed, and the excepting instruction and everything afterwards have not been executed. To take precise exceptions, the CPU must commit changes to the software visible state in the program order. This in order commit happens very naturally in the classic RISC pipeline. Most instructions write their results to the register file in the writeback stage, and so those writes automatically happen in program order. Store instructions, however, write their results to the Store Data Queue in the access stage. If the store instruction takes an exception, the Store Data Queue entry is invalidated so that it is not written to the cache data SRAM later.
Кэш қатесімен жұмыс істеу
Кейде дерек кэшінде немесе нұсқаулық кэшінде қажетті дерек немесе нұсқаулық болмайды. Мұндай жағдайларда, процессор кэш қажетті деректермен толтырылғанға дейін жұмысын тоқтатуы керек, содан кейін орындауды қайта бастауы керек. Кэшке қажетті деректерді толтыру (және мүмкін, кэш жолын жадқа қайта жазу) мәселесі құбыр желісінің ұйымдасқандығына тән емес және бұл жерде талқыланбайды. Тоқтату/қайта бастау мәселесін шешудің екі стратегиясы бар. Біріншісі – жаһандық тоқтату сигналы. Бұл сигнал қосылғанда, нұсқаулардың құбыр бойымен жылжуына кедерес келтіреді, әдетте әр кезеңнің басындағы триггерлерге (flip-flops) сағатты тоқтату арқылы. Бұл стратегияның кемшілігі – триггерлердің көп саны болғандықтан, жаһандық тоқтату сигналының таралуына көп уақыт қажет. Машина, әдетте, тоқтатуды қажет ететін жағдайды анықтаған циклде тоқтатуға тиіс болғандықтан, тоқтату сигналы жылдамдық шектеуші маңызды тізбекке айналады. Тоқтатуды/қайта бастауды басқарудың тағы бір стратегиясы – ерекше жағдай логикасын қайта пайдалану. Машина кінәлі нұсқаулық бойынша ерекше жағдайды тудырады, ал қалған барлық нұсқаулар күшін жояды. Кэш қажетті деректермен толтырылғаннан кейін, кэш қатесіне себеп болған нұсқаулық қайта іске қосылады. Дерек кэшіндегі қателерді өңдеуді жеделдету үшін, нұсқаулықты дерек кэші толтырылғаннан кейін бір циклде қайта іске қосуға болады.
Occasionally, either the data or instruction cache does not contain a required datum or instruction. In these cases, the CPU must suspend operation until the cache can be filled with the necessary data, and then must resume execution. The problem of filling the cache with the required data (and potentially writing back to memory the evicted cache line) is not specific to the pipeline organization, and is not discussed here. There are two strategies to handle the suspend/resume problem. The first is a global stall signal. This signal, when activated, prevents instructions from advancing down the pipeline, generally by gating off the clock to the flip flops at the start of each stage. The disadvantage of this strategy is that there are a large number of flip flops, so the global stall signal takes a long time to propagate. Since the machine generally has to stall in the same cycle that it identifies the condition requiring the stall, the stall signal becomes a speed limiting critical path. Another strategy to handle suspend/resume is to reuse the exception logic. The machine takes an exception on the offending instruction, and all further instructions are invalidated. When the cache has been filled with the necessary data, the instruction that caused the cache miss restarts. To expedite data cache miss handling, the instruction can be restarted so that its access cycle happens one cycle after the data cache is filled.