Ағылшыншамен салыстырыңыз: абзацты басыңыз — түпнұсқа терезеде ашылады. Абзац астындағы EN түймесі оны мәтін ішінде көрсетеді.
Мазмұны
Кіріспе
Жоғары жұмыс уақытымен жүйелер, ә. і. "әрқашан қосулы"
Systems with high up time, a. k. a. "always on"
Жоғары қолжетімділік (HA) – жүйенің сипаттамасы, ол келісілген деңгейдегі операциялық өнімділікті, әдетте жұмыс уақытын, қалыптыдан ұзақ мерзімге қамтамасыз етуге бағытталған. Модернизация осы жүйелерге тәуелділікті арттырды. Мысалы, ауруханалар мен деректер орталықтары күнделікті қызметтерін орындау үшін жүйелерінің жоғары қолжетімділігін қажет етеді. Қолжетімділік – пайдаланушылардың қызмет немесе тауар алуға, жүйеге кіруге, жаңа тапсырма беруге, қолданыстағы тапсырманы жаңартуға немесе өзгертуге, алдыңғы тапсырманың нәтижелерін алуға мүмкіндігін білдіреді. Егер пайдаланушы жүйеге кіре алмаса, ол пайдаланушының көзқарасы бойынша қолжетімсіз болып табылады. Әдетте, "тоқтау уақыты" термині жүйенің қолжетімсіз кезеңдерін білдіреді.
High availability (HA) is a characteristic of a system that aims to ensure an agreed level of operational performance, usually uptime, for a higher than normal period. Modernization has resulted in an increased reliance on these systems. For example, hospitals and data centers require high availability of their systems to perform routine daily activities. Availability refers to the ability of the user community to obtain a service or good, access the system, whether to submit new work, update or alter existing work, or collect the results of previous work. If a user cannot access the system, it is – from the user's point of view – unavailable. Generally, the term downtime is used to refer to periods when a system is unavailable.
Негізгі қағидалар
Сенімділік инженериясында жоғары қолжетімділікке қол жеткізуге көмектесетін жүйелерді жобалаудың үш қағидасы бар. Жеке сәтсіздік нүктелерін жою. Бұл жүйеге қосымша мүмкіндіктерді қосуды немесе құруды білдіреді, сонда компонент бұзылған жағдайда бүкіл жүйе бұзылмайды. Сенімді ауысу. Қосымша мүмкіндігі бар жүйелерде ауысу нүктесінің өзі жеке сәтсіздік нүктесіне айналуы мүмкін. Сенімді жүйелер сенімді ауысуды қамтамасыз етуі керек. Сәтсіздіктерді олар пайда болатын кезде анықтау. Егер жоғарыда аталған екі қағида сақталса, пайдаланушы ешқашан сәтсіздікті байқамауы мүмкін, бірақ техникалық қызмет көрсету оны міндетті түрде анықтауы керек.
There are three principles of systems design in reliability engineering which can help achieve high availability. Elimination of single points of failure. This means adding or building redundancy into the system so that failure of a component does not mean failure of the entire system. Reliable crossover. In redundant systems, the crossover point itself tends to become a single point of failure. Reliable systems must provide for reliable crossover. Detection of failures as they occur. If the two principles above are observed, then a user may never see a failure – but the maintenance activity must.
Жоспарланған және жоспарланбаған жұмыс істемеу уақыты
Жоспарланған және жоспарланбаған жұмыс тоқтауын ажыратуға болады. Әдетте, жоспарланған жұмыс тоқтауы жүйе жұмысын бұзатын және қазіргі орнатылған жүйе дизайнымен болдырмау мүмкін емес техникалық қызмет көрсетудің нәтижесі болып табылады. Жоспарланған жұмыс тоқтауына жүйелік бағдарламалық жасақтаманы жаңартулар, қайта жүктеуді қажет ететін немесе жүйе конфигурациясының өзгерістері, қайта жүктеуден кейін ғана күшіне енетін өзгерістер кіруі мүмкін. Жалпы, жоспарланған жұмыс тоқтауы әдетте логикалық, басқарушылық бастамамен туындаған оқиғаның нәтижесі болып табылады. Жоспарланбаған жұмыс тоқтауы әдетте аппараттық немесе бағдарламалық жасақтаманың істен шығуы сияқты физикалық оқиғалардан немесе қоршаған ортаның аномалиясынан туындайды. Жоспарланбаған жұмыс тоқтауының мысалдарына қуат көзінің үзілуі, процессордың (CPU) немесе жедел жадтың (RAM) компоненттерінің (немесе басқа аппараттық компоненттердің) істен шығуы, температураның артуына байланысты өшірілу, логикалық немесе физикалық желілік қосылымның үзілуі, қауіпсіздік бұзушылықтары, сондай-ақ әртүрлі қосымшалардың, аралық бағдарламалық қамтамасыз етудің және операциялық жүйелердің істен шығуы жатады. Егер пайдаланушыларды жоспарланған жұмыс тоқтауы туралы ескертуге болады, онда бұл айырмашылық пайдалы. Бірақ егер нақты жоғары қолжетімділік қажет болса, жұмыс тоқтауы жоспарланған немесе жоспарланбаған болсын, жұмыс тоқтауы болып саналады. Көптеген есептеу орталықтары қолжетімділік есептеулерінен жоспарланған жұмыс тоқтауын алып тастайды, оның есептеу пайдаланушыларына тигізер әсері аз немесе мүлдем жоқ деп есептейді. Осылайша, олар өте жоғары қолжетімділікке қол жеткізгенін мәлімдей алады, бұл үздіксіз қолжетімділік туралы иллюзия тудыруы мүмкін. Шынымен үздіксіз қолжетімділікті қамтамасыз ететін жүйелер салыстырмалы түрде сирек кездеседі және құны жоғары, көбінесе кез келген бір сәтсіздік нүктесін жоятын және аппараттық, желілік, операциялық жүйелерді, аралық бағдарламалық қамтамасыз етуді және қосымшаларды онлайн режимінде жаңартуға, түзетуге және алмастыруға мүмкіндік беретін арнайы, мұқият жобаланған жүйелерді пайдаланады. Кейбір жүйелер үшін жоспарланған жұмыс тоқтауы маңызды емес, мысалы, кеңсе ғимаратындағы жүйе түнде барлығы үйлеріне кеткеннен кейін тоқтатылса.
A distinction can be made between scheduled and unscheduled downtime. Typically, scheduled downtime is a result of maintenance that is disruptive to system operation and usually cannot be avoided with a currently installed system design. Scheduled downtime events might include patches to system software that require a reboot or system configuration changes that only take effect upon a reboot. In general, scheduled downtime is usually the result of some logical, management initiated event. Unscheduled downtime events typically arise from some physical event, such as a hardware or software failure or environmental anomaly. Examples of unscheduled downtime events include power outages, failed CPU or RAM components (or possibly other failed hardware components), an over temperature related shutdown, logically or physically severed network connections, security breaches, or various application, middleware, and operating system failures. If users can be warned away from scheduled downtimes, then the distinction is useful. But if the requirement is for true high availability, then downtime is downtime whether or not it is scheduled. Many computing sites exclude scheduled downtime from availability calculations, assuming that it has little or no impact upon the computing user community. By doing this, they can claim to have phenomenally high availability, which might give the illusion of continuous availability. Systems that exhibit truly continuous availability are comparatively rare and higher priced, and most have carefully implemented specialty designs that eliminate any single point of failure and allow online hardware, network, operating system, middleware, and application upgrades, patches, and replacements. For certain systems, scheduled downtime does not matter, for example system downtime at an office building after everybody has gone home for the night.
Проценттік есептеу
Қолжетімділік әдетте белгілі бір жылдағы жұмыс істеу уақытының пайызы ретінде көрсетіледі. Төмендегі кестеде жүйе үздіксіз жұмыс істеуі керек екенін ескере отырып, нақты бір қолжетімділік пайызы үшін рұқсат етілетін тоқтау уақыты көрсетілген. Қызмет деңгейі туралы келісімдер көбінесе айлық есеп айырысу циклдарына сәйкес қызметтік кредиттерді есептеу үшін айлық тоқтау уақытына немесе қолжетімділікке сілтеме жасайды. Төмендегі кестеде берілген қолжетімділік пайызына сәйкес жүйе қанша уақытқа қолжетімсіз болатыны көрсетілген.
Availability is usually expressed as a percentage of uptime in a given year. The following table shows the downtime that will be allowed for a particular percentage of availability, presuming that the system is required to operate continuously. Service level agreements often refer to monthly downtime or availability in order to calculate service credits to match monthly billing cycles. The following table shows the translation from a given availability percentage to the corresponding amount of time a system would be unavailable. Availability %Downtime per yearDowntime per quarterDowntime per monthDowntime per weekDowntime per day (24 hours) 90% ("one nine")36.53 days9.13 days73.05 hours16.80 hours2.40 hours95% ("one nine five")18.26 days4.56 days36.53 hours8.40 hours1.20 hours97% ("one nine seven")10.96 days2.74 days21.92 hours5.04 hours43.20 minutes98% ("one nine eight")7.31 days43.86 hours14.61 hours3.36 hours28.80 minutes99% ("two nines")3.65 days21.9 hours7.31 hours1.68 hours14.40 minutes99.5% ("two nines five")1.83 days10.98 hours3.65 hours50.40 minutes7.20 minutes99.8% ("two nines eight")17.53 hours4.38 hours87.66 minutes20.16 minutes2.88 minutes99.9% ("three nines")8.77 hours2.19 hours43.83 minutes10.08 minutes1.44 minutes99.95% ("three nines five")4.38 hours65.7 minutes21.92 minutes5.04 minutes43.20 seconds99.99% ("four nines")52.60 minutes13.15 minutes4.38 minutes1.01 minutes8.64 seconds99.995% ("four nines five")26.30 minutes6.57 minutes2.19 minutes30.24 seconds4.32 seconds99.999% ("five nines")5.26 minutes1.31 minutes26.30 seconds6.05 seconds864.00 milliseconds99.9999% ("six nines")31.56 seconds7.89 seconds2.63 seconds604.80 milliseconds86.40 milliseconds99.99999% ("seven nines")3.16 seconds0.79 seconds262.98 milliseconds60.48 milliseconds8.64 milliseconds99.999999% ("eight nines")315.58 milliseconds78.89 milliseconds26.30 milliseconds6.05 milliseconds864.00 microseconds99.9999999% ("nine nines")31.56 milliseconds7.89 milliseconds2.63 milliseconds604.80 microseconds86.40 microseconds99.99999999% ("ten nines")3.16 milliseconds788.40 microseconds262.80 microseconds60.48 microseconds8.64 microseconds99.999999999% ("eleven nines")315.58 microseconds78.84 microseconds26.28 microseconds6.05 microseconds864.00 nanoseconds99.9999999999% ("twelve nines")31.56 microseconds7.88 microseconds2.63 microseconds604.81 nanoseconds86.40 nanoseconds
The terms uptime and availability are often used interchangeably but do not always refer to the same thing. For example, a system can be "up" with its services not "available" in the case of a network outage. Or a system undergoing software maintenance can be "available" to be worked on by a system administrator, but its services do not appear "up" to the end user or customer. The subject of the terms is thus important here: whether the focus of a discussion is the server hardware, server OS, functional service, software service/process, or similar, it is only if there is a single, consistent subject of the discussion that the words uptime and availability can be used synonymously.
Availability is usually expressed as a percentage of uptime in a given year. The following table shows the downtime that will be allowed for a particular percentage of availability, presuming that the system is required to operate continuously. Service level agreements often refer to monthly downtime or availability in order to calculate service credits to match monthly billing cycles. The following table shows the translation from a given availability percentage to the corresponding amount of time a system would be unavailable. Availability %Downtime per yearDowntime per quarterDowntime per monthDowntime per weekDowntime per day (24 hours) 90% ("one nine")36.53 days9.13 days73.05 hours16.80 hours2.40 hours95% ("one nine five")18.26 days4.56 days36.53 hours8.40 hours1.20 hours97% ("one nine seven")10.96 days2.74 days21.92 hours5.04 hours43.20 minutes98% ("one nine eight")7.31 days43.86 hours14.61 hours3.36 hours28.80 minutes99% ("two nines")3.65 days21.9 hours7.31 hours1.68 hours14.40 minutes99.5% ("two nines five")1.83 days10.98 hours3.65 hours50.40 minutes7.20 minutes99.8% ("two nines eight")17.53 hours4.38 hours87.66 minutes20.16 minutes2.88 minutes99.9% ("three nines")8.77 hours2.19 hours43.83 minutes10.08 minutes1.44 minutes99.95% ("three nines five")4.38 hours65.7 minutes21.92 minutes5.04 minutes43.20 seconds99.99% ("four nines")52.60 minutes13.15 minutes4.38 minutes1.01 minutes8.64 seconds99.995% ("four nines five")26.30 minutes6.57 minutes2.19 minutes30.24 seconds4.32 seconds99.999% ("five nines")5.26 minutes1.31 minutes26.30 seconds6.05 seconds864.00 milliseconds99.9999% ("six nines")31.56 seconds7.89 seconds2.63 seconds604.80 milliseconds86.40 milliseconds99.99999% ("seven nines")3.16 seconds0.79 seconds262.98 milliseconds60.48 milliseconds8.64 milliseconds99.999999% ("eight nines")315.58 milliseconds78.89 milliseconds26.30 milliseconds6.05 milliseconds864.00 microseconds99.9999999% ("nine nines")31.56 milliseconds7.89 milliseconds2.63 milliseconds604.80 microseconds86.40 microseconds99.99999999% ("ten nines")3.16 milliseconds788.40 microseconds262.80 microseconds60.48 microseconds8.64 microseconds99.999999999% ("eleven nines")315.58 microseconds78.84 microseconds26.28 microseconds6.05 microseconds864.00 nanoseconds99.9999999999% ("twelve nines")31.56 microseconds7.88 microseconds2.63 microseconds604.81 nanoseconds86.40 nanoseconds
The terms uptime and availability are often used interchangeably but do not always refer to the same thing. For example, a system can be "up" with its services not "available" in the case of a network outage. Or a system undergoing software maintenance can be "available" to be worked on by a system administrator, but its services do not appear "up" to the end user or customer. The subject of the terms is thus important here: whether the focus of a discussion is the server hardware, server OS, functional service, software service/process, or similar, it is only if there is a single, consistent subject of the discussion that the words uptime and availability can be used synonymously.
Жұмыс істеу уақыты мен қолжетімділік жиі бірінің орнына бірін қолданылады, бірақ олар әрқашан бірдей нәрсені білдірмейді. Мысалы, жүйе желілік ақаулық жағдайында қызметтері "қолжетімді" болмаса да, "жұмыс істеп" тұруы мүмкін. Немесе бағдарламалық қамтамасыз етуді күтіп-ұстау жүйесі жүйе әкімшісімен жұмыс істеуге "қолжетімді" болуы мүмкін, бірақ оның қызметтері соңғы пайдаланушы немесе клиент үшін "жұмыс істемей" көрінуі мүмкін. Сондықтан терминдердің нысаны маңызды: талқылау серверлік жабдық, серверлік операциялық жүйе, функционалдық қызмет, бағдарламалық қызмет/процесс немесе осыған ұқсас нәрсеге қатысты болса, жұмыс істеу уақыты мен қолжетімділік сөздерін синоним ретінде қолдануға болады, егер талқылаудың нысаны біртұтас және тұрақты болса ғана.
Availability is usually expressed as a percentage of uptime in a given year. The following table shows the downtime that will be allowed for a particular percentage of availability, presuming that the system is required to operate continuously. Service level agreements often refer to monthly downtime or availability in order to calculate service credits to match monthly billing cycles. The following table shows the translation from a given availability percentage to the corresponding amount of time a system would be unavailable. Availability %Downtime per yearDowntime per quarterDowntime per monthDowntime per weekDowntime per day (24 hours) 90% ("one nine")36.53 days9.13 days73.05 hours16.80 hours2.40 hours95% ("one nine five")18.26 days4.56 days36.53 hours8.40 hours1.20 hours97% ("one nine seven")10.96 days2.74 days21.92 hours5.04 hours43.20 minutes98% ("one nine eight")7.31 days43.86 hours14.61 hours3.36 hours28.80 minutes99% ("two nines")3.65 days21.9 hours7.31 hours1.68 hours14.40 minutes99.5% ("two nines five")1.83 days10.98 hours3.65 hours50.40 minutes7.20 minutes99.8% ("two nines eight")17.53 hours4.38 hours87.66 minutes20.16 minutes2.88 minutes99.9% ("three nines")8.77 hours2.19 hours43.83 minutes10.08 minutes1.44 minutes99.95% ("three nines five")4.38 hours65.7 minutes21.92 minutes5.04 minutes43.20 seconds99.99% ("four nines")52.60 minutes13.15 minutes4.38 minutes1.01 minutes8.64 seconds99.995% ("four nines five")26.30 minutes6.57 minutes2.19 minutes30.24 seconds4.32 seconds99.999% ("five nines")5.26 minutes1.31 minutes26.30 seconds6.05 seconds864.00 milliseconds99.9999% ("six nines")31.56 seconds7.89 seconds2.63 seconds604.80 milliseconds86.40 milliseconds99.99999% ("seven nines")3.16 seconds0.79 seconds262.98 milliseconds60.48 milliseconds8.64 milliseconds99.999999% ("eight nines")315.58 milliseconds78.89 milliseconds26.30 milliseconds6.05 milliseconds864.00 microseconds99.9999999% ("nine nines")31.56 milliseconds7.89 milliseconds2.63 milliseconds604.80 microseconds86.40 microseconds99.99999999% ("ten nines")3.16 milliseconds788.40 microseconds262.80 microseconds60.48 microseconds8.64 microseconds99.999999999% ("eleven nines")315.58 microseconds78.84 microseconds26.28 microseconds6.05 microseconds864.00 nanoseconds99.9999999999% ("twelve nines")31.56 microseconds7.88 microseconds2.63 microseconds604.81 nanoseconds86.40 nanoseconds
The terms uptime and availability are often used interchangeably but do not always refer to the same thing. For example, a system can be "up" with its services not "available" in the case of a network outage. Or a system undergoing software maintenance can be "available" to be worked on by a system administrator, but its services do not appear "up" to the end user or customer. The subject of the terms is thus important here: whether the focus of a discussion is the server hardware, server OS, functional service, software service/process, or similar, it is only if there is a single, consistent subject of the discussion that the words uptime and availability can be used synonymously.
Бес-бес жадыға салу
Қарапайым есте сақтау ережесі бойынша, 5 тоғыздық жүйе жылына шамамен 5 минут тоқтау уақытын қамтамасыз етеді. Одан басқа нұсқаларды 10-ға көбейту немесе бөлу арқылы алуға болады: 4 тоғыздық – 50 минут, ал 3 тоғыздық – 500 минут. Керісінше, 6 тоғыздық – 0,5 минут (30 секунд), ал 7 тоғыздық – 3 секунд.
A simple mnemonic rule states that 5 nines allows approximately 5 minutes of downtime per year. Variants can be derived by multiplying or dividing by 10: 4 nines is 50 minutes and 3 nines is 500 minutes. In the opposite direction, 6 nines is 0.5 minutes (30 sec) and 7 nines is 3 seconds.
"Керектіліктері 10" трюгі
"Тоғыздық" қолжетімділік пайызы үшін рұқсат етілген істен шығу уақытын есептеу үшін тағы бір есімдік тәсіл – күніне секундтар формуласын қолдану. Мысалы, 90% ("бір тоғыз") көрсеткішін береді, демек рұқсат етілген істен шығу уақыты күніне секунд. Сондай-ақ, 99,999% ("бес тоғыз") көрсеткішін береді, демек рұқсат етілген істен шығу уақыты күніне секунд.
Another memory trick to calculate the allowed downtime duration for an " nines" availability percentage is to use the formula seconds per day. For example, 90% ("one nine") yields the exponent , and therefore the allowed downtime is seconds per day. Also, 99.999% ("five nines") gives the exponent , and therefore the allowed downtime is seconds per day.
Өлшем және түсіндіру
Қолжетімділікті өлшеу белгілі бір дәрежеде түсіндіруді қажет етеді. Жыл емес жылда 365 күн жұмыс істеген жүйе, ең жоғалаған сәтте 9 сағатқа созылған желілік ақаулықтан зардап шегуі мүмкін; пайдаланушылар қауымдастығы жүйені қолжетімсіз деп қарастырса, жүйе әкімшісі 100% жұмыс уақытын мәлімдейді. Дегенмен, қолжетімділіктің нақты анықтамасын ескере отырып, жүйенің қолжетімділігі шамамен 99,9% немесе үш тоғызды (8760 сағаттың ішінде 8751 сағат қолжетімді уақыт) құрайды. Сондай-ақ, жүйелердің өнімділігінде мәселелер туындағанда, жүйелер жұмыс істеуін жалғастырғанның өзінде де пайдаланушылар оларды толық немесе ішінара қолжетімсіз деп санауы мүмкін. Сол сияқты, белгілі бір қолданба функцияларының қолжетімсіздігі әкімшілерге байқалмауы мүмкін, бірақ пайдаланушылар үшін өте ауыр салдарға әкелуі мүмкін – нағыз қолжетімділікті өлшеу кешенді болуы керек. Қолжетімділікті анықтау үшін өлшеу қажет, мүмкіндігінше, өте жоғары қолжетімділікке ие жан-жақты бақылау құралдарымен ("инструментациямен") өлшеу керек. Егер инструментация жеткіліксіз болса, күндіз-түні үлкен көлемде транзакцияларды өңдейтін жүйелер, мысалы, кредиттік карталарды өңдеу жүйелері немесе телефон станциялары, сұраныстың уақытша азаюына ұшырайтын жүйелерге қарағанда, кем дегенде пайдаланушылар тарапынан жақсырақ бақыланады. Баламалы өлшем – еңіс-төмендік арасындағы орташа уақыт (MTBF).
Availability measurement is subject to some degree of interpretation. A system that has been up for 365 days in a non leap year might have been eclipsed by a network failure that lasted for 9 hours during a peak usage period; the user community will see the system as unavailable, whereas the system administrator will claim 100% uptime. However, given the true definition of availability, the system will be approximately 99.9% available, or three nines (8751 hours of available time out of 8760 hours per non leap year). Also, systems experiencing performance problems are often deemed partially or entirely unavailable by users, even when the systems are continuing to function. Similarly, unavailability of select application functions might go unnoticed by administrators yet be devastating to users – a true availability measure is holistic. Availability must be measured to be determined, ideally with comprehensive monitoring tools ("instrumentation") that are themselves highly available. If there is a lack of instrumentation, systems supporting high volume transaction processing throughout the day and night, such as credit card processing systems or telephone switches, are often inherently better monitored, at least by the users themselves, than systems which experience periodic lulls in demand. An alternative metric is mean time between failures (MTBF).
Жақын байланысты ұғымдар
Қалпына келтіру уақыты (немесе жөндеудің болжамды уақыты (ETR), сондай-ақ қалпына келтіру уақыты мақсаты (RTO) қолжетімділікпен тығыз байланысты, яғни жоспарланған тоқтау үшін қажетті жалпы уақыт немесе жоспарланбаған тоқтаудан толық қалпына келтіруге қажетті уақыт. Тағы бір өлшем – қалпына келтірудің орташа уақыты (MTTR). Кейбір жүйелік құрылымдар мен қателерде қалпына келтіру уақыты шексіз болуы мүмкін, яғни толық қалпына келтіру мүмкін емес. Мұндай мысалдың бірі – екінші апаттық қалпына келтіру дерек орталығы болмаған жағдайда дерек орталығы мен оның жүйелерін жоятын өрт немесе су тасқыны. Тағы бір байланысты ұғым – деректердің қолжетімділігі, яғни дерекқорыштар мен басқа да ақпарат сақтау жүйелері жүйелік транзакцияларды дұрыс тіркеп, хабарлай алады. Ақпаратты басқару көбінесе әртүрлі қателік жағдайлардағы қабылданатын (немесе нақты) деректер жоғалуын анықтау үшін деректердің қолжетімділігіне немесе қалпына келтіру нүктесі мақсатына жеке назар аударады. Кейбір пайдаланушылар қолданба қызметтерінің тоқтауына төтеп бере алады, бірақ деректерді жоғалтуға төтеп бере алмайды. Қызмет деңгейі туралы келісім ("SLA") ұйымның қолжетімділік мақсаттары мен талаптарын ресми түрде бекітеді.
Recovery time (or estimated time of repair (ETR), also known as recovery time objective (RTO) is closely related to availability, that is the total time required for a planned outage or the time required to fully recover from an unplanned outage. Another metric is mean time to recovery (MTTR). Recovery time could be infinite with certain system designs and failures, i. e. full recovery is impossible. One such example is a fire or flood that destroys a data center and its systems when there is no secondary disaster recovery data center. Another related concept is data availability, that is the degree to which databases and other information storage systems faithfully record and report system transactions. Information management often focuses separately on data availability, or Recovery Point Objective, in order to determine acceptable (or actual) data loss with various failure events. Some users can tolerate application service interruptions but cannot tolerate data loss. A service level agreement ("SLA") formalizes an organization's availability objectives and requirements.
Әскери басқару жүйелері
Жоғары қолжетімділік – пилоттықсыз көлік құралдары мен автономды теңіз кемелерінің басқару жүйелері үшін басты талаптардың бірі. Егер басқару жүйесі істен шығып қалса, Ground Combat Vehicle (GCV) немесе ASW Continuous Trail Unmanned Vessel (ACTUV) жоғалып қалуы мүмкін.
High availability is one of the primary requirements of the control systems in unmanned vehicles and autonomous maritime vessels. If the controlling system becomes unavailable, the Ground Combat Vehicle (GCV) or ASW Continuous Trail Unmanned Vessel (ACTUV) would be lost.
Жүйенің құрылысы
Жалпы жүйелік жобаға қосымша компоненттерді қосу жоғары қолжетімділікке қол жеткізуге бағытталған күш-жігерді әлсіретуі мүмкін, өйткені күрделі жүйелерде өзінен-өзі көбірек ықтимал сәтсіздіктер бар және оларды дұрыс іске асыру қиын. Кейбір сарапшылар ең жоғары қолжетімді жүйелер қарапайым архитектураға (бір ғана, жоғары сапалы, көп мақсатты физикалық жүйе, кешенді ішкі аппараттық артықшылықпен) сәйкес келеді деген теорияны алға тартса да, мұндай архитектураның бүкіл жүйені жаңарту және операциялық жүйені жаңарту үшін тоқтату қажеттігінен кемшілігі бар. Жүйелердің жетілдірілген жобалары қызмет көрсетуді тоқтатпай жүйелерді жөндеуге және жаңартуға мүмкіндік береді (мысалы, жүктеме теңгерімдеу және автоматты ауыстыру). Жоғары қолжетімділік күрделі жүйелерде жұмысын қалпына келтіру үшін адамның араласуын азайтады, себебі істен шығудың ең көп таралған себебі – адам қателігі. Артықшылық жоғары деңгейдегі қолжетімділікпен жүйелерді құру үшін қолданылады (мысалы, ұшақтардың ұшу компьютерлері). Мұндай жағдайда ақауларды жоғары дәрежеде анықтау және ортақ себептерге байланысты ақауларды болдырмау қажет. Артықшылықтың екі түрі бар: белсенді артықшылық және пассивті артықшылық. Пассивті артықшылық өнімділіктің төмендеуін өтеу үшін жобаға жеткілікті қосымша қуатты қосу арқылы жоғары қолжетімділікке қол жеткізуге мүмкіндік береді. Ең қарапайым мысал – екі бөлек қозғалтқышы және екі бөлек винті бар қайық. Қозғалтқыш немесе винтінің біреуі істен шыққан жағдайда да қайық бағытына қарай қозғалады. Күрделі мысалға электр қуатын беруді қамтитын ірі жүйедегі бірнеше артық қуат өндіру қондырғыларын жатқызуға болады. Жеке компоненттердің істен шығуы, егер оның нәтижесіндегі өнімділіктің төмендеуі жүйенің жалпы спецификация шегінен аспаса, ақау болып есептелмейді. Активті артықшылық күрделі жүйелерде өнімділіктің төмендеуіне жол бермей жоғары қолжетімділікке қол жеткізу үшін қолданылады. Бірнеше бірдей элементтер бір жобаға енгізіледі, сондай-ақ ақауларды анықтау және дауыс беру схемасын қолдана отырып, істен шыққан элементтерді автоматты түрде айналып өтуге мүмкіндік беретін жүйені қайта конфигурациялау әдісі қарастырылады. Бұл әдіс байланысқан күрделі есептеу жүйелерінде қолданылады. Интернет маршрутизациясы осы саладағы Бирман және Джозефтің алғашқы жұмыстарының нәтижесі болып табылады. Активті артықшылық жүйеге күрделі ақаулық түрлерін енгізе алады, мысалы, қате дауыс беру логикасының салдарынан жүйенің үнемі қайта конфигурациялануы. Нөлдік тоқтау уақытымен жүйелерді жобалау, модельдеу және симуляция нәтижесіндегі орташа істен шығу аралығы жоспарланған техникалық қызмет көрсету, жаңарту шаралары немесе жүйенің қызмет ету мерзімі аралығынан әлдеқайда асып түсетінін білдіреді. Нөлдік тоқтау уақыты кейбір ұшақ түрлері мен көптеген байланыс спутниктері үшін қажетті үлкен артықшылықты қамтиды. Глобалды позициялау жүйесі – нөлдік тоқтау уақытымен жүйеге мысал. Ақаулықты диагностикалау жүйесі шектеулі артықшылығы бар жүйелерде жоғары қолжетімділікке қол жеткізу үшін қолданылады. Жөндеу жұмыстары ақаулық индикаторы іске қосылғаннан кейін ғана қысқа уақыт ішінде жүргізіледі. Ақаулық тек қана миссияның маңызды кезеңінде орын алса ғана маңызды болып саналады. Ірі жүйелердің теориялық сенімділігін бағалау үшін модельдеу және симуляция қолданылады. Мұндай модельдердің нәтижелері әртүрлі жобалау нұсқаларын бағалау үшін пайдаланылады. Жүйенің толыққанды моделі құрылады және компоненттерді алып тастау арқылы модель сынақтан өтеді. Артықшылықты симуляциялау N x критерийлерін қамтиды. N жүйедегі компоненттердің жалпы санын білдіреді. x жүйені сынау үшін қолданылатын компоненттердің санын білдіреді. N-1 дегеніміз, модель бір компоненті істен шыққан барлық мүмкін комбинациялар бойынша жұмысын бағау арқылы сынақтан өтеді. N-2 дегеніміз, модель бір мезгілде екі компоненті істен шыққан барлық мүмкін комбинациялар бойынша жұмысын бағау арқылы сынақтан өтеді.
Adding more components to an overall system design can undermine efforts to achieve high availability because complex systems inherently have more potential failure points and are more difficult to implement correctly. While some analysts would put forth the theory that the most highly available systems adhere to a simple architecture (a single, high quality, multi purpose physical system with comprehensive internal hardware redundancy), this architecture suffers from the requirement that the entire system must be brought down for patching and operating system upgrades. More advanced system designs allow for systems to be patched and upgraded without compromising service availability (see load balancing and failover). High availability requires less human intervention to restore operation in complex systems; the reason for this being that the most common cause for outages is human error. Redundancy is used to create systems with high levels of availability (e. g. aircraft flight computers). In this case it is required to have high levels of failure detectability and avoidance of common cause failures. Two kinds of redundancy are passive redundancy and active redundancy. Passive redundancy is used to achieve high availability by including enough excess capacity in the design to accommodate a performance decline. The simplest example is a boat with two separate engines driving two separate propellers. The boat continues toward its destination despite failure of a single engine or propeller. A more complex example is multiple redundant power generation facilities within a large system involving electric power transmission. Malfunction of single components is not considered to be a failure unless the resulting performance decline exceeds the specification limits for the entire system. Active redundancy is used in complex systems to achieve high availability with no performance decline. Multiple items of the same kind are incorporated into a design that includes a method to detect failure and automatically reconfigure the system to bypass failed items using a voting scheme. This is used with complex computing systems that are linked. Internet routing is derived from early work by Birman and Joseph in this area. Active redundancy may introduce more complex failure modes into a system, such as continuous system reconfiguration due to faulty voting logic. Zero downtime system design means that modeling and simulation indicates mean time between failures significantly exceeds the period of time between planned maintenance, upgrade events, or system lifetime. Zero downtime involves massive redundancy, which is needed for some types of aircraft and for most kinds of communications satellites. Global Positioning System is an example of a zero downtime system. Fault instrumentation can be used in systems with limited redundancy to achieve high availability. Maintenance actions occur during brief periods of down time only after a fault indicator activates. Failure is only significant if this occurs during a mission critical period. Modeling and simulation is used to evaluate the theoretical reliability for large systems. The outcome of this kind of model is used to evaluate different design options. A model of the entire system is created, and the model is stressed by removing components. Redundancy simulation involves the N x criteria. N represents the total number of components in the system. x is the number of components used to stress the system. N 1 means the model is stressed by evaluating performance with all possible combinations where one component is faulted. N 2 means the model is stressed by evaluating performance with all possible combinations where two component are faulted simultaneously.
Қол жеткізбеу шығындары
1998 жылғы IBM Global Services компаниясының баяндамасында, 1996 жылы жұмыс істемей қалған жүйелер американдық кәсіпорындарға өнімділік пен кірістің жоғалтуынан 4,54 миллиард долларлық шығын келтіргені бағаланды.
In a 1998 report from IBM Global Services, unavailable systems were estimated to have cost American businesses $4.54 billion in 1996, due to lost productivity and revenues.