Введение
Научный инструмент, используемый для автоматизации процесса секвенирования ДНК. Секвенсор ДНК – это научный инструмент, используемый для автоматизации процесса секвенирования ДНК. При наличии образца ДНК секвенсор ДНК используется для определения последовательности четырех оснований: G (гуанин), C (цитозин), A (аденин) и T (тимин). Результат представляется в виде текстовой строки, называемой «ридом». Некоторые секвенсоры ДНК также можно рассматривать как оптические приборы, поскольку они анализируют световые сигналы, исходящие от флуорохромов, присоединенных к нуклеотидам. Первый автоматизированный секвенсор ДНК, изобретенный Ллойдом М. Смитом, был представлен компанией Applied Biosystems в 1987 году. Он использовал метод секвенирования Сангера – технологию, которая легла в основу секвенсоров ДНК «первого поколения» и позволила завершить проект «Геном человека» в 2001 году. Секвенсоры ДНК первого поколения, по сути, представляют собой автоматизированные системы электрофореза, которые обнаруживают миграцию меченых фрагментов ДНК. Поэтому эти секвенсоры также могут использоваться для генотипирования генетических маркеров, где необходимо определить только длину фрагмента ДНК (например, микросателлитов, AFLP). Проект «Геном человека» стимулировал разработку более дешевых, высокопроизводительных и точных платформ, известных как секвенсоры следующего поколения (NGS), для секвенирования генома человека. К ним относятся платформы секвенирования ДНК 454, SOLiD и Illumina. Секвенсоры следующего поколения значительно увеличили скорость секвенирования ДНК по сравнению с предыдущими методами Сангера. Подготовка образцов ДНК может быть автоматизирована всего за 90 минут.
A DNA sequencer is a scientific instrument used to automate the DNA sequencing process. Given a sample of DNA, a DNA sequencer is used to determine the order of the four bases: G (guanine), C (cytosine), A (adenine) and T (thymine). This is then reported as a text string, called a read. Some DNA sequencers can be also considered optical instruments as they analyze light signals originating from fluorochromes attached to nucleotides. The first automated DNA sequencer, invented by Lloyd M. Smith, was introduced by Applied Biosystems in 1987. It used the Sanger sequencing method, a technology which formed the basis of the "first generation" of DNA sequencers and enabled the completion of the human genome project in 2001. This first generation of DNA sequencers are essentially automated electrophoresis systems that detect the migration of labelled DNA fragments. Therefore, these sequencers can also be used in the genotyping of genetic markers where only the length of a DNA fragment(s) needs to be determined (e. g. microsatellites, AFLPs). The Human Genome Project spurred the development of cheaper, high throughput and more accurate platforms known as Next Generation Sequencers (NGS) to sequence the human genome. These include the 454, SOLiD and Illumina DNA sequencing platforms. Next generation sequencing machines have increased the rate of DNA sequencing substantially, as compared with the previous Sanger methods. DNA samples can be prepared automatically in as little as 90 mins,
More recent, third generation DNA sequencers such as PacBio SMRT and Oxford Nanopore offer the possibility of sequencing long molecules, compared to short read technologies such as Illumina SBS or MGI Tech's DNBSEQ. Because of limitations in DNA sequencer technology, the reads of many of these technologies are short, compared to the length of a genome therefore the reads must be assembled into longer contigs. The data may also contain errors, caused by limitations in the DNA sequencing technique or by errors during PCR amplification. DNA sequencer manufacturers use a number of different methods to detect which DNA bases are present. The specific protocols applied in different sequencing platforms have an impact in the final data that is generated. Therefore, comparing data quality and cost across different technologies can be a daunting task. Each manufacturer provides their own ways to inform sequencing errors and scores. However, errors and scores between different platforms cannot always be compared directly. Since these systems rely on different DNA sequencing approaches, choosing the best DNA sequencer and method will typically depend on the experiment objectives and available budget. and Sanger (1975). Gilbert introduced a sequencing method based on chemical modification of DNA followed by cleavage at specific bases whereas Sanger's technique is based on dideoxynucleotide chain termination. The Sanger method became popular due to its increased efficiency and low radioactivity. The first automated DNA sequencer was the AB370A, introduced in 1986 by Applied Biosystems. The AB370A was able to sequence 96 samples simultaneously, 500 kilobases per day, and reaching read lengths up to 600 bases. This was the beginning of the "first generation" of DNA sequencers, It was developed by 454 Life Sciences and purchased by Roche in 2007. 454 utilizes the detection of pyrophosphate released by the DNA polymerase reaction when adding a nucleotide to the template strain. Roche currently manufactures two systems based on their pyrosequencing technology: the GS FLX+ and the GS Junior System. The GS FLX+ System promises read lengths of approximately 1000 base pairs while the GS Junior System promises 400 base pair reads. A predecessor to GS FLX+, the 454 GS FLX Titanium system was released in 2008, achieving an output of 0.7G of data per run, with 99.9% accuracy after quality filter, and a read length of up to 700bp. In 2009, Roche launched the GS Junior, a bench top version of the 454 sequencer with read length up to 400bp, and simplified library preparation and data processing. One of the advantages of 454 systems is their running speed. Manpower can be reduced with automation of library preparation and semi automation of emulsion PCR. A disadvantage of the 454 system is that it is prone to errors when estimating the number of bases in a long string of identical nucleotides. This is referred to as a homopolymer error and occurs when there are 6 or more identical bases in row. Another disadvantage is that the price of reagents is relatively more expensive compared with other next generation sequencers. In 2013 Roche announced that they would be shutting down development of 454 technology and phasing out 454 machines completely in 2016 when its technology became noncompetitive. Roche produces a number of software tools which are optimised for the analysis of 454 sequencing data. Such as,
GS Run Processor converts raw images generated by a sequencing run into intensity values. The process consists of two main steps: image processing and signal processing. The software also applies normalization, signal correction, base calling and quality scores for individual reads. The software outputs data in Standard Flowgram Format (or SFF) files to be used in data analysis applications (GS De Novo Assembler, GS Reference Mapper or GS Amplicon Variant Analyzer). GS De Novo Assembler is a tool for de novo assembly of whole genomes up to 3GB in size from shotgun reads alone or combined with paired end data generated by 454 sequencers. It also supports de novo assembly of transcripts (including analysis), and also isoform variant detection. Illumina makes a number of next generation sequencing machines using this technology including the HiSeq, Genome Analyzer IIx, MiSeq and the HiScanSQ, which can also process microarrays. The technology leading to these DNA sequencers was first released by Solexa in 2006 as the Genome Analyzer. and Sanger based DNA sequencers such as the 3500 Genetic Analyzer. Under the Ion Torrent brand, Applied Biosystems produces four next generation sequencers: the Ion PGM System, Ion Proton System, Ion S5 and Ion S5xl systems. The company is also believed to be developing their new capillary DNA sequencer called SeqStudio that will be released early 2018. SOLiD systems was acquired by Applied Biosystems in 2006. SOLiD applies sequencing by ligation and dual base encoding. The first SOLiD system was launched in 2007, generating reading lengths of 35bp and 3G data per run. After five upgrades, the 5500xl sequencing system was released in 2010, considerably increasing read length to 85bp, improving accuracy up to 99.99% and producing 30G per 7 day run. and has to some extent limited its use to experiments where read length is less vital such as resequencing and transcriptome analysis and more recently ChIP Seq and methylation experiments. a data analysis package for resequencing, ChiP Seq and transcriptome analysis. It uses the MaxMapper algorithm to map the colour space reads.
Более новые секвенсоры ДНК третьего поколения, такие как PacBio SMRT и Oxford Nanopore, предлагают возможность секвенирования длинных молекул по сравнению с технологиями короткого чтения, такими как Illumina SBS или DNBSEQ от MGI Tech. Из-за ограничений технологии секвенирования ДНК, риды многих из этих технологий коротки по сравнению с длиной генома, поэтому риды необходимо собирать в более длинные контиги. Данные также могут содержать ошибки, вызванные ограничениями в технике секвенирования ДНК или ошибками при амплификации ПЦР. Производители секвенсоров ДНК используют различные методы для определения присутствующих оснований ДНК. Конкретные протоколы, применяемые на различных платформах секвенирования, влияют на генерируемые конечные данные. Поэтому сравнение качества данных и стоимости различных технологий может быть сложной задачей. Каждый производитель предоставляет свои собственные способы информирования об ошибках и оценках секвенирования. Однако ошибки и оценки между различными платформами не всегда можно напрямую сравнивать. Поскольку эти системы основаны на различных подходах к секвенированию ДНК, выбор наилучшего секвенсора ДНК и метода обычно зависит от целей эксперимента и доступного бюджета. Сангер (1975) и Гилберт представили методы секвенирования, основанные на химической модификации ДНК с последующим расщеплением в определенных основаниях, в то время как техника Сангера основана на обрыве цепи дидеоксинуклеотидами. Метод Сангера стал популярным благодаря своей повышенной эффективности и низкой радиоактивности. Первым автоматизированным секвенсором ДНК был AB370A, представленный в 1986 году компанией Applied Biosystems. AB370A мог секвенировать 96 образцов одновременно, 500 килобаз в день, достигая длины ридов до 600 оснований. Это стало началом секвенсоров ДНК «первого поколения». Он был разработан компанией 454 Life Sciences и приобретен компанией Roche в 2007 году. 454 использует обнаружение пирофосфата, высвобождаемого реакцией ДНК-полимеразы при добавлении нуклеотида к матричной цепи. В настоящее время Roche производит две системы на основе своей технологии пиросеквенирования: GS FLX+ и GS Junior System. Система GS FLX+ обещает длину ридов примерно 1000 пар оснований, а система GS Junior – 400 пар оснований. Система 454 GS FLX Titanium, предшественник GS FLX+, была выпущена в 2008 году, обеспечивая выход 0,7 Г данных за запуск, точность 99,9% после качественного фильтра и длину ридов до 700 п.н. В 2009 году Roche выпустила GS Junior – настольную версию секвенсора 454 с длиной ридов до 400 п.н. и упрощенной подготовкой библиотек и обработкой данных. Одним из преимуществ систем 454 является их скорость работы. Затраты рабочей силы могут быть снижены за счет автоматизации подготовки библиотек и полуавтоматизации эмульсионной ПЦР. Недостатком системы 454 является ее склонность к ошибкам при оценке количества оснований в длинной последовательности идентичных нуклеотидов. Это называется гомополимерной ошибкой и возникает при наличии 6 или более идентичных оснований подряд. Другой недостаток заключается в том, что стоимость реагентов относительно выше по сравнению с другими секвенсорами следующего поколения. В 2013 году Roche объявила о прекращении разработки технологии 454 и полном выводе машин 454 из эксплуатации в 2016 году, когда ее технология стала неконкурентоспособной. Roche производит ряд программных инструментов, оптимизированных для анализа данных секвенирования 454. Например, GS Run Processor преобразует необработанные изображения, генерируемые в ходе секвенирования, в значения интенсивности. Процесс состоит из двух основных этапов: обработка изображений и обработка сигналов. Программное обеспечение также применяет нормализацию, коррекцию сигналов, определение оснований и оценки качества отдельных ридов. Программное обеспечение выводит данные в файлы стандартного формата Flowgram (SFF) для использования в приложениях анализа данных (GS De Novo Assembler, GS Reference Mapper или GS Amplicon Variant Analyzer). GS De Novo Assembler – это инструмент для de novo сборки геномов размером до 3 ГБ только из shotgun-ридов или в сочетании с парными данными, генерируемыми секвенсорами 454. Он также поддерживает de novo сборку транскриптов (включая анализ) и обнаружение вариантов изоформ. Illumina производит ряд секвенсоров следующего поколения, использующих эту технологию, включая HiSeq, Genome Analyzer IIx, MiSeq и HiScanSQ, которые также могут обрабатывать микроматрицы. Технология, лежащая в основе этих секвенсоров ДНК, была впервые выпущена Solexa в 2006 году как Genome Analyzer. И секвенсоры ДНК на основе Сангера, такие как 3500 Genetic Analyzer. Под брендом Ion Torrent компания Applied Biosystems производит четыре секвенсора следующего поколения: Ion PGM System, Ion Proton System, Ion S5 и Ion S5xl systems. Сообщается, что компания также разрабатывает новый капиллярный секвенсор ДНК под названием SeqStudio, который будет выпущен в начале 2018 года. Системы SOLiD были приобретены компанией Applied Biosystems в 2006 году. SOLiD применяет секвенирование посредством лигирования и двойного кодирования оснований. Первая система SOLiD была запущена в 2007 году, генерируя длины чтения 35 п.н. и 3 Г данных за запуск. После пяти обновлений в 2010 году была выпущена система секвенирования 5500xl, значительно увеличившая длину чтения до 85 п.н., повысившая точность до 99,99% и производящая 30 Г за 7-дневный запуск. И в некоторой степени ограничила ее использование в экспериментах, где длина чтения менее важна, таких как ресеквенирование и транскриптомика, и в последнее время ChIP Seq и эксперименты по метилированию. Пакет для анализа данных для ресеквенирования, ChiP Seq и транскриптомики. Он использует алгоритм MaxMapper для сопоставления ридов в цветовом пространстве.
A DNA sequencer is a scientific instrument used to automate the DNA sequencing process. Given a sample of DNA, a DNA sequencer is used to determine the order of the four bases: G (guanine), C (cytosine), A (adenine) and T (thymine). This is then reported as a text string, called a read. Some DNA sequencers can be also considered optical instruments as they analyze light signals originating from fluorochromes attached to nucleotides. The first automated DNA sequencer, invented by Lloyd M. Smith, was introduced by Applied Biosystems in 1987. It used the Sanger sequencing method, a technology which formed the basis of the "first generation" of DNA sequencers and enabled the completion of the human genome project in 2001. This first generation of DNA sequencers are essentially automated electrophoresis systems that detect the migration of labelled DNA fragments. Therefore, these sequencers can also be used in the genotyping of genetic markers where only the length of a DNA fragment(s) needs to be determined (e. g. microsatellites, AFLPs). The Human Genome Project spurred the development of cheaper, high throughput and more accurate platforms known as Next Generation Sequencers (NGS) to sequence the human genome. These include the 454, SOLiD and Illumina DNA sequencing platforms. Next generation sequencing machines have increased the rate of DNA sequencing substantially, as compared with the previous Sanger methods. DNA samples can be prepared automatically in as little as 90 mins,
More recent, third generation DNA sequencers such as PacBio SMRT and Oxford Nanopore offer the possibility of sequencing long molecules, compared to short read technologies such as Illumina SBS or MGI Tech's DNBSEQ. Because of limitations in DNA sequencer technology, the reads of many of these technologies are short, compared to the length of a genome therefore the reads must be assembled into longer contigs. The data may also contain errors, caused by limitations in the DNA sequencing technique or by errors during PCR amplification. DNA sequencer manufacturers use a number of different methods to detect which DNA bases are present. The specific protocols applied in different sequencing platforms have an impact in the final data that is generated. Therefore, comparing data quality and cost across different technologies can be a daunting task. Each manufacturer provides their own ways to inform sequencing errors and scores. However, errors and scores between different platforms cannot always be compared directly. Since these systems rely on different DNA sequencing approaches, choosing the best DNA sequencer and method will typically depend on the experiment objectives and available budget. and Sanger (1975). Gilbert introduced a sequencing method based on chemical modification of DNA followed by cleavage at specific bases whereas Sanger's technique is based on dideoxynucleotide chain termination. The Sanger method became popular due to its increased efficiency and low radioactivity. The first automated DNA sequencer was the AB370A, introduced in 1986 by Applied Biosystems. The AB370A was able to sequence 96 samples simultaneously, 500 kilobases per day, and reaching read lengths up to 600 bases. This was the beginning of the "first generation" of DNA sequencers, It was developed by 454 Life Sciences and purchased by Roche in 2007. 454 utilizes the detection of pyrophosphate released by the DNA polymerase reaction when adding a nucleotide to the template strain. Roche currently manufactures two systems based on their pyrosequencing technology: the GS FLX+ and the GS Junior System. The GS FLX+ System promises read lengths of approximately 1000 base pairs while the GS Junior System promises 400 base pair reads. A predecessor to GS FLX+, the 454 GS FLX Titanium system was released in 2008, achieving an output of 0.7G of data per run, with 99.9% accuracy after quality filter, and a read length of up to 700bp. In 2009, Roche launched the GS Junior, a bench top version of the 454 sequencer with read length up to 400bp, and simplified library preparation and data processing. One of the advantages of 454 systems is their running speed. Manpower can be reduced with automation of library preparation and semi automation of emulsion PCR. A disadvantage of the 454 system is that it is prone to errors when estimating the number of bases in a long string of identical nucleotides. This is referred to as a homopolymer error and occurs when there are 6 or more identical bases in row. Another disadvantage is that the price of reagents is relatively more expensive compared with other next generation sequencers. In 2013 Roche announced that they would be shutting down development of 454 technology and phasing out 454 machines completely in 2016 when its technology became noncompetitive. Roche produces a number of software tools which are optimised for the analysis of 454 sequencing data. Such as,
GS Run Processor converts raw images generated by a sequencing run into intensity values. The process consists of two main steps: image processing and signal processing. The software also applies normalization, signal correction, base calling and quality scores for individual reads. The software outputs data in Standard Flowgram Format (or SFF) files to be used in data analysis applications (GS De Novo Assembler, GS Reference Mapper or GS Amplicon Variant Analyzer). GS De Novo Assembler is a tool for de novo assembly of whole genomes up to 3GB in size from shotgun reads alone or combined with paired end data generated by 454 sequencers. It also supports de novo assembly of transcripts (including analysis), and also isoform variant detection. Illumina makes a number of next generation sequencing machines using this technology including the HiSeq, Genome Analyzer IIx, MiSeq and the HiScanSQ, which can also process microarrays. The technology leading to these DNA sequencers was first released by Solexa in 2006 as the Genome Analyzer. and Sanger based DNA sequencers such as the 3500 Genetic Analyzer. Under the Ion Torrent brand, Applied Biosystems produces four next generation sequencers: the Ion PGM System, Ion Proton System, Ion S5 and Ion S5xl systems. The company is also believed to be developing their new capillary DNA sequencer called SeqStudio that will be released early 2018. SOLiD systems was acquired by Applied Biosystems in 2006. SOLiD applies sequencing by ligation and dual base encoding. The first SOLiD system was launched in 2007, generating reading lengths of 35bp and 3G data per run. After five upgrades, the 5500xl sequencing system was released in 2010, considerably increasing read length to 85bp, improving accuracy up to 99.99% and producing 30G per 7 day run. and has to some extent limited its use to experiments where read length is less vital such as resequencing and transcriptome analysis and more recently ChIP Seq and methylation experiments. a data analysis package for resequencing, ChiP Seq and transcriptome analysis. It uses the MaxMapper algorithm to map the colour space reads.
Бекман Колтер
Beckman Coulter (ныне Danaher) ранее выпускала секвенаторы ДНК, основанные на методе прерывания цепи и капиллярном электрофорезе, под торговой маркой CEQ, включая модель CEQ 8000. В настоящее время компания производит систему генетического анализа GeXP, использующую метод секвенирования с использованием окрашенных терминаторов. Этот метод использует термоциклер аналогично ПЦР для денатурации, отжига и элонгации фрагментов ДНК, что позволяет амплифицировать секвенируемые фрагменты.
Тихоокеанские биологические науки
Pacific Biosciences производит системы секвенирования PacBio RS и Sequel, используя метод секвенирования отдельных молекул в реальном времени, или SMRT. Эта система позволяет получать длины прочтений в несколько тысяч пар оснований. Повышенная частота ошибок в исходных данных корректируется либо с помощью циркулярного консенсуса, при котором одна и та же цепь считывается многократно, либо с использованием оптимизированных стратегий сборки генома. Ученые сообщают о достижении точности 99,9999% при использовании этих стратегий. Система Sequel была представлена в 2015 году с увеличенной производительностью и более низкой ценой.
Оксфордская нанопора
Секвенсор MinION от Oxford Nanopore Technologies основан на развивающейся технологии секвенирования через нанопоры для анализа нуклеиновых кислот. Устройство имеет длину четыре дюйма и питается от USB-порта. MinION декодирует ДНК непосредственно по мере прохождения молекулы через нанопору, подвешенную в мембране, со скоростью 450 оснований в секунду. Изменения в электрическом токе указывают на присутствующее основание. Первоначально точность устройства составляла от 60 до 85 процентов, в то время как у традиционных приборов – 99,9 процента. Даже неточные результаты могут быть полезны благодаря возможности получения длинных прочтений. В начале 2021 года исследователи из Университета Британской Колумбии, используя специальные молекулярные метки, смогли снизить погрешность устройства с 5–15 процентов до менее 0,005 процента, даже при одновременном секвенировании множества длинных фрагментов ДНК. Существуют еще две модификации продукта на базе MinION: первая – GridION, немного больший секвенсор, способный обрабатывать до пяти проточных ячеек MinION одновременно, а вторая – PromethION, использующий до 100 000 пор параллельно, что делает его более подходящим для высокопроизводительного секвенирования.
МГИ
MGI производит высокопроизводительные секвенаторы для научных исследований и клинических применений, такие как DNBSEQ G50, DNBSEQ G400 и DNBSEQ T7, на основе запатентованной технологии DNBSEQ. Она основана на технологиях секвенирования ДНК-наношариков и комбинаторного синтеза зондов-якорей, в которых ДНК-наношарики (DNB) загружаются на чип с матричным расположением через жидкостную систему, а затем к адаптерной области DNB добавляется праймер для гибридизации. DNBSEQ T7 способен генерировать короткие риды в очень больших масштабах – до 60 геномов человека в день. DNBSEQ T7 использовался для получения парных ридов длиной 150 п.н., с глубиной покрытия 30X, для секвенирования генома SARS-CoV-2 или COVID-19 с целью выявления генетических вариантов, предрасполагающих к тяжелому течению COVID-19. С применением новой методики исследователи из Китайского национального генного банка секвенировали библиотеки, свободные от ПЦР, на массивах DNBSEQ без ПЦР от MGI, впервые получив истинное секвенирование всего генома без ПЦР. MGISEQ 2000 использовался в секвенировании одноклеточной РНК для изучения патогенеза и механизмов восстановления у пациентов с COVID-19, о чем опубликовано в журнале Nature Medicine.
Сравнение
Текущие предложения в технологии секвенирования ДНК демонстрируют доминирующего игрока: Illumina (декабрь 2019 года), за которым следуют PacBio, MGI и Oxford Nanopore. + Сравнение характеристик и производительности секвенаторов ДНК следующего поколения. Секвенатор Ion Torrent PGM 454 GS FLX Sanger 3730xl Производитель Ion Torrent (Life Technologies) 454 Life Sciences (Roche) Illumina Applied Biosystems (Life Technologies) Pacific Biosciences Applied Biosystems (Life Technologies) MGI Химия секвенирования Ионная полупроводниковая секвенция Пиросеквенирование Полимеразное секвенирование на основе синтеза Лигационное секвенирование Фосфолинк-флуоресцентные нуклеотиды Прерывание дидезокси-цепи Полимеразное секвенирование на основе синтеза Подход к амплификации Эмульсионная ПЦР Эмульсионная ПЦР Амплификация мостиком Эмульсионная ПЦР Одиночная молекула; без амплификации ПЦР Генерация ДНК-наношариков (DNB) Выход данных за прогон 100–200 Мб 0,7 Гб 600 Гб 120 Гб 0,5–1,0 Гб 1,9–84 Кб 1440 Гб / 1500–1800 млн ридов Точность 99% 99,9% 99,9% 99,94% 88,0% (>99,9999% CCS или HGAP) 99,999% 99,90% Время прогона 2 часа 24 часа 3–10 дней 7–14 дней 2–4 часа 20 минут – 3 часа 3–5 дней Длина чтения 200–400 п.н. 700 п.н. 100x100 п.н. парные концы 50x50 п.н. парные концы 14 000 п.н. (N50) 400–900 п.н. 100/150/200 п.н. парные концы Стоимость прогона 350 долл. США 7000 долл. США 6000 долл. США (геном человека с покрытием 30x) 4000 долл. США 125–300 долл. США 4 долл. США (одно чтение/реакция) Н/Д Стоимость за Мб 1,00 долл. США 10 долл. США 0,07 долл. США 0,13 долл. США 0,13 долл. США – 0,60 долл. США 2400 долл. США 0,007 долл. США Стоимость прибора 80 000 долл. США 500 000 долл. США 690 000 долл. США 495 000 долл. США 695 000 долл. США 95 000 долл. США Н/Д