Введение
База данных геномики крыс
База данных генома крыс (RGD) – это база данных геномики, генетики, физиологии и функциональных данных крыс, а также данных для сравнительной геномики крысы, человека и мыши. RGD отвечает за прикрепление биологической информации к геному крысы посредством структурированной терминологии, или онтологии, аннотаций, присвоенных генам и количественным локусам признаков (QTL), а также за консолидацию данных о штаммах крыс и предоставление их научному сообществу. Они также разрабатывают набор инструментов для извлечения и анализа геномных, физиологических и функциональных данных крысы, а также сравнительных данных для крысы, мыши, человека и пяти других видов. RGD возникла как совместная работа между исследовательскими институтами, участвующими в генетических и геномных исследованиях крыс. Её цель, как указано в запросе Национального института здоровья на получение гранта: HL 99 013, заключается в создании базы данных генома крысы для сбора, консолидации и интеграции данных, полученных в ходе продолжающихся генетических и геномных исследований крысы, и предоставления этих данных широкому научному сообществу. Вторичная, но важная цель – обеспечить курацию отображенных позиций для количественных локусов признаков, известных мутаций и других фенотипических данных. Крыса продолжает широко использоваться исследователями в качестве модельного организма для изучения фармакологии, токсикологии, общей физиологии, биологии и патофизиологии заболеваний. В последние годы наблюдается быстрый рост генетических и геномных данных о крысах. Кроме того, база данных генома крысы стала центральным источником информации о крысах для исследований и теперь содержит информацию не только о генетике и геномике, но и о физиологии и молекулярной биологии. Для всех этих областей доступны инструменты и страницы данных, которые курируются сотрудниками RGD.
Данные
Данные RGD состоят из ручных аннотаций исследователей RGD, а также импортированных аннотаций из различных источников. RGD также экспортирует собственные аннотации для обмена с другими. На странице данных RGD перечислены восемь типов данных, хранящихся в базе данных: гены, QTL, маркеры, карты, штаммы, онтологии, последовательности и ссылки. Из них шесть активно используются и регулярно обновляются. Тип данных RGD Maps относится к устаревшим генетическим и радиационным гибридным картам. Эти данные в значительной степени заменены последовательностью всего генома крысы. Тип данных Sequences не представляет собой полный список геномных, транскриптных или белковых последовательностей, а содержит в основном последовательности ПЦР-праймеров, определяющие простой полиморфизм длины последовательности (SSLP) и маркеры экспрессированных последовательностей (EST). Такие последовательности полезны прежде всего для исследователей, которые все еще используют эти маркеры для генотипирования своих животных и для различения маркеров с одинаковым названием. Шесть основных типов данных в RGD следующие:
Genes: Initial gene records are imported and updated from the National Center for Biotechnology Information's (NCBI's) Gene database on a weekly basis. Data imported during this process includes the Gene ID, Genbank/RefSeq nucleotide and protein sequence identifiers, HomoloGene group IDs and Ensembl Gene, Transcript and Protein IDs. Additional protein related data is imported from the UniProtKB database. RGD curators review the literature and manually curate Gene Ontology (GO), diseases, phenotypes and pathways for rat genes, diseases and pathways for mouse genes, and diseases, phenotypes and pathways for human genes. In addition, the site imports GO annotations for mouse and human genes from the GO Consortium, rat electronic annotations from UniProt and mouse phenotype annotations from the Mouse Genome Database/Mouse Genome Informatics (MGD/MGI). QTLs: RGD's staff manually curates data for rat and human QTLs from the literature where such publications exist or from records directly submitted by researchers. Mouse QTL records, including Mammalian Phenotype (MP) ontology assignments, are imported directly from MGI. For rat and human QTLs, curation includes assigning MP, HP, and disease ontology annotations. QTL positions are automatically assigned based on the genomic positions of peak and/or flanking markers or single nucleotide polymorphisms (SNPs). QTL records link to information about related strains, candidate genes, associated markers and related QTLs. Strains: Like QTL records, RGD strain records are either manually curated from the literature or submitted by researchers. Strain records include information about the official symbol of the strain, origin and availability of the strain, associated phenotypes, whether the strain is a model for a human disease, and any information that is available about breeding, behavior, husbandry, etc. Strain records link to information about related genes, alleles, and QTLs, associated strains (e. g. parental strains or substrains) and, where available, strain specific damaging nucleotide variants. For congenic and mutant strains, genomic positions are assigned for the introgressed region (congenic strains) or the location of the mutated sequence (mutant strains). Markers: Because genetic markers such as SSLPs and ESTs have been, and continue to be, used for QTLs and strains, RGD stores marker data for rat, human and mouse. Marker data includes the sequences of the associated forward and reverse PCR primers, genomic positions and links to NCBI's Probe database. Marker records link to associated QTL, strain and gene records. Cell lines: RGD stores cell line records based on imports from Cellosaurus. Although the largest numbers of these are human and mouse cell lines, records are also available for rat, bonobo, dog, squirrel, pig, green monkey and naked mole rat. Ontologies: In order to make RGD's data both human readable and available for computational analysis and retrieval, RGD relies on the use of multiple ontologies. As of July 2021, RGD used 19 different ontologies to express the various types of data applicable to RGD's diverse datatypes. Ontology annotations are assigned manually by curators Ontologies which are imported from outside sources are updated weekly. References: RGD references are scientific publications and resources that have been used for curation of information into the database, and are sources for data objects such as QTLs and strains. For references accessed via NCBI's PubMed, imported data includes the title, authors, citation and PubMed ID, and an RGD ID is generated. In some cases, references are generated as internal records, such as bulk uploads from automated pipelines or personal communications with data sources. These additional references give RGD users an identification of the source of particular pieces and types of data for which PubMed records are not available. Both types of reference records provide links to all of the data curated from that article or source, including genes, QTLs, strains, disease and other ontology annotations. The resources curated for information can be retrieved from the database using the reference search page or links on an object page. Uncurated references are also available, which are known to contain relevant data but have not yet been manually reviewed. These are found as PubMed links listed in the ‘References – uncurated’ section of an object report (e. g. a gene report).
Гены: начальные записи генов импортируются и обновляются из базы данных генов Национального центра биотехнологической информации (NCBI) еженедельно. Данные, импортируемые в ходе этого процесса, включают идентификатор гена, идентификаторы нуклеотидных и белковых последовательностей Genbank/RefSeq, идентификаторы групп HomoloGene и идентификаторы генов, транскриптов и белков Ensembl. Дополнительные данные, связанные с белками, импортируются из базы данных UniProtKB. Кураторы RGD просматривают литературу и вручную курируют генную онтологию (GO), заболевания, фенотипы и пути для генов крыс, заболевания и пути для генов мыши, а также заболевания, фенотипы и пути для генов человека. Кроме того, сайт импортирует аннотации GO для генов мыши и человека из консорциума GO, электронные аннотации для крыс из UniProt и аннотации фенотипов мыши из базы данных генома мыши/информатики генома мыши (MGD/MGI).
Genes: Initial gene records are imported and updated from the National Center for Biotechnology Information's (NCBI's) Gene database on a weekly basis. Data imported during this process includes the Gene ID, Genbank/RefSeq nucleotide and protein sequence identifiers, HomoloGene group IDs and Ensembl Gene, Transcript and Protein IDs. Additional protein related data is imported from the UniProtKB database. RGD curators review the literature and manually curate Gene Ontology (GO), diseases, phenotypes and pathways for rat genes, diseases and pathways for mouse genes, and diseases, phenotypes and pathways for human genes. In addition, the site imports GO annotations for mouse and human genes from the GO Consortium, rat electronic annotations from UniProt and mouse phenotype annotations from the Mouse Genome Database/Mouse Genome Informatics (MGD/MGI). QTLs: RGD's staff manually curates data for rat and human QTLs from the literature where such publications exist or from records directly submitted by researchers. Mouse QTL records, including Mammalian Phenotype (MP) ontology assignments, are imported directly from MGI. For rat and human QTLs, curation includes assigning MP, HP, and disease ontology annotations. QTL positions are automatically assigned based on the genomic positions of peak and/or flanking markers or single nucleotide polymorphisms (SNPs). QTL records link to information about related strains, candidate genes, associated markers and related QTLs. Strains: Like QTL records, RGD strain records are either manually curated from the literature or submitted by researchers. Strain records include information about the official symbol of the strain, origin and availability of the strain, associated phenotypes, whether the strain is a model for a human disease, and any information that is available about breeding, behavior, husbandry, etc. Strain records link to information about related genes, alleles, and QTLs, associated strains (e. g. parental strains or substrains) and, where available, strain specific damaging nucleotide variants. For congenic and mutant strains, genomic positions are assigned for the introgressed region (congenic strains) or the location of the mutated sequence (mutant strains). Markers: Because genetic markers such as SSLPs and ESTs have been, and continue to be, used for QTLs and strains, RGD stores marker data for rat, human and mouse. Marker data includes the sequences of the associated forward and reverse PCR primers, genomic positions and links to NCBI's Probe database. Marker records link to associated QTL, strain and gene records. Cell lines: RGD stores cell line records based on imports from Cellosaurus. Although the largest numbers of these are human and mouse cell lines, records are also available for rat, bonobo, dog, squirrel, pig, green monkey and naked mole rat. Ontologies: In order to make RGD's data both human readable and available for computational analysis and retrieval, RGD relies on the use of multiple ontologies. As of July 2021, RGD used 19 different ontologies to express the various types of data applicable to RGD's diverse datatypes. Ontology annotations are assigned manually by curators Ontologies which are imported from outside sources are updated weekly. References: RGD references are scientific publications and resources that have been used for curation of information into the database, and are sources for data objects such as QTLs and strains. For references accessed via NCBI's PubMed, imported data includes the title, authors, citation and PubMed ID, and an RGD ID is generated. In some cases, references are generated as internal records, such as bulk uploads from automated pipelines or personal communications with data sources. These additional references give RGD users an identification of the source of particular pieces and types of data for which PubMed records are not available. Both types of reference records provide links to all of the data curated from that article or source, including genes, QTLs, strains, disease and other ontology annotations. The resources curated for information can be retrieved from the database using the reference search page or links on an object page. Uncurated references are also available, which are known to contain relevant data but have not yet been manually reviewed. These are found as PubMed links listed in the ‘References – uncurated’ section of an object report (e. g. a gene report).
QTL: сотрудники RGD вручную курируют данные для QTL крыс и человека из литературы, где такие публикации существуют, или из записей, непосредственно представленных исследователями. Записи QTL мыши, включая назначения онтологии фенотипа млекопитающих (MP), импортируются непосредственно из MGI. Для QTL крыс и человека курирование включает назначение MP, HP и аннотаций онтологии заболеваний. Позиции QTL автоматически присваиваются на основе геномных позиций пиковых и/или фланкирующих маркеров или однонуклеотидных полиморфизмов (SNP). Записи QTL связаны с информацией о родственных штаммах, кандидатных генах, ассоциированных маркерах и родственных QTL.
Genes: Initial gene records are imported and updated from the National Center for Biotechnology Information's (NCBI's) Gene database on a weekly basis. Data imported during this process includes the Gene ID, Genbank/RefSeq nucleotide and protein sequence identifiers, HomoloGene group IDs and Ensembl Gene, Transcript and Protein IDs. Additional protein related data is imported from the UniProtKB database. RGD curators review the literature and manually curate Gene Ontology (GO), diseases, phenotypes and pathways for rat genes, diseases and pathways for mouse genes, and diseases, phenotypes and pathways for human genes. In addition, the site imports GO annotations for mouse and human genes from the GO Consortium, rat electronic annotations from UniProt and mouse phenotype annotations from the Mouse Genome Database/Mouse Genome Informatics (MGD/MGI). QTLs: RGD's staff manually curates data for rat and human QTLs from the literature where such publications exist or from records directly submitted by researchers. Mouse QTL records, including Mammalian Phenotype (MP) ontology assignments, are imported directly from MGI. For rat and human QTLs, curation includes assigning MP, HP, and disease ontology annotations. QTL positions are automatically assigned based on the genomic positions of peak and/or flanking markers or single nucleotide polymorphisms (SNPs). QTL records link to information about related strains, candidate genes, associated markers and related QTLs. Strains: Like QTL records, RGD strain records are either manually curated from the literature or submitted by researchers. Strain records include information about the official symbol of the strain, origin and availability of the strain, associated phenotypes, whether the strain is a model for a human disease, and any information that is available about breeding, behavior, husbandry, etc. Strain records link to information about related genes, alleles, and QTLs, associated strains (e. g. parental strains or substrains) and, where available, strain specific damaging nucleotide variants. For congenic and mutant strains, genomic positions are assigned for the introgressed region (congenic strains) or the location of the mutated sequence (mutant strains). Markers: Because genetic markers such as SSLPs and ESTs have been, and continue to be, used for QTLs and strains, RGD stores marker data for rat, human and mouse. Marker data includes the sequences of the associated forward and reverse PCR primers, genomic positions and links to NCBI's Probe database. Marker records link to associated QTL, strain and gene records. Cell lines: RGD stores cell line records based on imports from Cellosaurus. Although the largest numbers of these are human and mouse cell lines, records are also available for rat, bonobo, dog, squirrel, pig, green monkey and naked mole rat. Ontologies: In order to make RGD's data both human readable and available for computational analysis and retrieval, RGD relies on the use of multiple ontologies. As of July 2021, RGD used 19 different ontologies to express the various types of data applicable to RGD's diverse datatypes. Ontology annotations are assigned manually by curators Ontologies which are imported from outside sources are updated weekly. References: RGD references are scientific publications and resources that have been used for curation of information into the database, and are sources for data objects such as QTLs and strains. For references accessed via NCBI's PubMed, imported data includes the title, authors, citation and PubMed ID, and an RGD ID is generated. In some cases, references are generated as internal records, such as bulk uploads from automated pipelines or personal communications with data sources. These additional references give RGD users an identification of the source of particular pieces and types of data for which PubMed records are not available. Both types of reference records provide links to all of the data curated from that article or source, including genes, QTLs, strains, disease and other ontology annotations. The resources curated for information can be retrieved from the database using the reference search page or links on an object page. Uncurated references are also available, which are known to contain relevant data but have not yet been manually reviewed. These are found as PubMed links listed in the ‘References – uncurated’ section of an object report (e. g. a gene report).
Штаммы: как и записи QTL, записи штаммов RGD либо вручную курируются из литературы, либо представляются исследователями. Записи штаммов включают информацию об официальном символе штамма, происхождении и доступности штамма, связанных с ним фенотипах, является ли штамм моделью для заболевания человека, а также любую доступную информацию о разведении, поведении, содержании и т. д. Записи штаммов связаны с информацией о родственных генах, аллелях и QTL, ассоциированных штаммах (например, родительских штаммах или подштаммах) и, при наличии, специфичных для штамма вариантах нуклеотидов, вызывающих повреждения. Для конгенных и мутантных штаммов геномные позиции присваиваются интрогрессирующей области (конгенные штаммы) или местоположению мутировавшей последовательности (мутантные штаммы).
Genes: Initial gene records are imported and updated from the National Center for Biotechnology Information's (NCBI's) Gene database on a weekly basis. Data imported during this process includes the Gene ID, Genbank/RefSeq nucleotide and protein sequence identifiers, HomoloGene group IDs and Ensembl Gene, Transcript and Protein IDs. Additional protein related data is imported from the UniProtKB database. RGD curators review the literature and manually curate Gene Ontology (GO), diseases, phenotypes and pathways for rat genes, diseases and pathways for mouse genes, and diseases, phenotypes and pathways for human genes. In addition, the site imports GO annotations for mouse and human genes from the GO Consortium, rat electronic annotations from UniProt and mouse phenotype annotations from the Mouse Genome Database/Mouse Genome Informatics (MGD/MGI). QTLs: RGD's staff manually curates data for rat and human QTLs from the literature where such publications exist or from records directly submitted by researchers. Mouse QTL records, including Mammalian Phenotype (MP) ontology assignments, are imported directly from MGI. For rat and human QTLs, curation includes assigning MP, HP, and disease ontology annotations. QTL positions are automatically assigned based on the genomic positions of peak and/or flanking markers or single nucleotide polymorphisms (SNPs). QTL records link to information about related strains, candidate genes, associated markers and related QTLs. Strains: Like QTL records, RGD strain records are either manually curated from the literature or submitted by researchers. Strain records include information about the official symbol of the strain, origin and availability of the strain, associated phenotypes, whether the strain is a model for a human disease, and any information that is available about breeding, behavior, husbandry, etc. Strain records link to information about related genes, alleles, and QTLs, associated strains (e. g. parental strains or substrains) and, where available, strain specific damaging nucleotide variants. For congenic and mutant strains, genomic positions are assigned for the introgressed region (congenic strains) or the location of the mutated sequence (mutant strains). Markers: Because genetic markers such as SSLPs and ESTs have been, and continue to be, used for QTLs and strains, RGD stores marker data for rat, human and mouse. Marker data includes the sequences of the associated forward and reverse PCR primers, genomic positions and links to NCBI's Probe database. Marker records link to associated QTL, strain and gene records. Cell lines: RGD stores cell line records based on imports from Cellosaurus. Although the largest numbers of these are human and mouse cell lines, records are also available for rat, bonobo, dog, squirrel, pig, green monkey and naked mole rat. Ontologies: In order to make RGD's data both human readable and available for computational analysis and retrieval, RGD relies on the use of multiple ontologies. As of July 2021, RGD used 19 different ontologies to express the various types of data applicable to RGD's diverse datatypes. Ontology annotations are assigned manually by curators Ontologies which are imported from outside sources are updated weekly. References: RGD references are scientific publications and resources that have been used for curation of information into the database, and are sources for data objects such as QTLs and strains. For references accessed via NCBI's PubMed, imported data includes the title, authors, citation and PubMed ID, and an RGD ID is generated. In some cases, references are generated as internal records, such as bulk uploads from automated pipelines or personal communications with data sources. These additional references give RGD users an identification of the source of particular pieces and types of data for which PubMed records are not available. Both types of reference records provide links to all of the data curated from that article or source, including genes, QTLs, strains, disease and other ontology annotations. The resources curated for information can be retrieved from the database using the reference search page or links on an object page. Uncurated references are also available, which are known to contain relevant data but have not yet been manually reviewed. These are found as PubMed links listed in the ‘References – uncurated’ section of an object report (e. g. a gene report).
Маркеры: поскольку генетические маркеры, такие как SSLP и EST, использовались и продолжают использоваться для QTL и штаммов, RGD хранит данные маркеров для крыс, людей и мышей. Данные маркеров включают последовательности соответствующих прямых и обратных ПЦР-праймеров, геномные позиции и ссылки на базу данных Probe NCBI. Записи маркеров связаны с ассоциированными записями QTL, штаммов и генов.
Genes: Initial gene records are imported and updated from the National Center for Biotechnology Information's (NCBI's) Gene database on a weekly basis. Data imported during this process includes the Gene ID, Genbank/RefSeq nucleotide and protein sequence identifiers, HomoloGene group IDs and Ensembl Gene, Transcript and Protein IDs. Additional protein related data is imported from the UniProtKB database. RGD curators review the literature and manually curate Gene Ontology (GO), diseases, phenotypes and pathways for rat genes, diseases and pathways for mouse genes, and diseases, phenotypes and pathways for human genes. In addition, the site imports GO annotations for mouse and human genes from the GO Consortium, rat electronic annotations from UniProt and mouse phenotype annotations from the Mouse Genome Database/Mouse Genome Informatics (MGD/MGI). QTLs: RGD's staff manually curates data for rat and human QTLs from the literature where such publications exist or from records directly submitted by researchers. Mouse QTL records, including Mammalian Phenotype (MP) ontology assignments, are imported directly from MGI. For rat and human QTLs, curation includes assigning MP, HP, and disease ontology annotations. QTL positions are automatically assigned based on the genomic positions of peak and/or flanking markers or single nucleotide polymorphisms (SNPs). QTL records link to information about related strains, candidate genes, associated markers and related QTLs. Strains: Like QTL records, RGD strain records are either manually curated from the literature or submitted by researchers. Strain records include information about the official symbol of the strain, origin and availability of the strain, associated phenotypes, whether the strain is a model for a human disease, and any information that is available about breeding, behavior, husbandry, etc. Strain records link to information about related genes, alleles, and QTLs, associated strains (e. g. parental strains or substrains) and, where available, strain specific damaging nucleotide variants. For congenic and mutant strains, genomic positions are assigned for the introgressed region (congenic strains) or the location of the mutated sequence (mutant strains). Markers: Because genetic markers such as SSLPs and ESTs have been, and continue to be, used for QTLs and strains, RGD stores marker data for rat, human and mouse. Marker data includes the sequences of the associated forward and reverse PCR primers, genomic positions and links to NCBI's Probe database. Marker records link to associated QTL, strain and gene records. Cell lines: RGD stores cell line records based on imports from Cellosaurus. Although the largest numbers of these are human and mouse cell lines, records are also available for rat, bonobo, dog, squirrel, pig, green monkey and naked mole rat. Ontologies: In order to make RGD's data both human readable and available for computational analysis and retrieval, RGD relies on the use of multiple ontologies. As of July 2021, RGD used 19 different ontologies to express the various types of data applicable to RGD's diverse datatypes. Ontology annotations are assigned manually by curators Ontologies which are imported from outside sources are updated weekly. References: RGD references are scientific publications and resources that have been used for curation of information into the database, and are sources for data objects such as QTLs and strains. For references accessed via NCBI's PubMed, imported data includes the title, authors, citation and PubMed ID, and an RGD ID is generated. In some cases, references are generated as internal records, such as bulk uploads from automated pipelines or personal communications with data sources. These additional references give RGD users an identification of the source of particular pieces and types of data for which PubMed records are not available. Both types of reference records provide links to all of the data curated from that article or source, including genes, QTLs, strains, disease and other ontology annotations. The resources curated for information can be retrieved from the database using the reference search page or links on an object page. Uncurated references are also available, which are known to contain relevant data but have not yet been manually reviewed. These are found as PubMed links listed in the ‘References – uncurated’ section of an object report (e. g. a gene report).
Клеточные линии: RGD хранит записи клеточных линий на основе импорта из Cellosaurus. Хотя наибольшее количество из них составляют клеточные линии человека и мыши, доступны также записи для крыс, бонобо, собак, белок, свиней, зеленых обезьян и голых землекопов.
Genes: Initial gene records are imported and updated from the National Center for Biotechnology Information's (NCBI's) Gene database on a weekly basis. Data imported during this process includes the Gene ID, Genbank/RefSeq nucleotide and protein sequence identifiers, HomoloGene group IDs and Ensembl Gene, Transcript and Protein IDs. Additional protein related data is imported from the UniProtKB database. RGD curators review the literature and manually curate Gene Ontology (GO), diseases, phenotypes and pathways for rat genes, diseases and pathways for mouse genes, and diseases, phenotypes and pathways for human genes. In addition, the site imports GO annotations for mouse and human genes from the GO Consortium, rat electronic annotations from UniProt and mouse phenotype annotations from the Mouse Genome Database/Mouse Genome Informatics (MGD/MGI). QTLs: RGD's staff manually curates data for rat and human QTLs from the literature where such publications exist or from records directly submitted by researchers. Mouse QTL records, including Mammalian Phenotype (MP) ontology assignments, are imported directly from MGI. For rat and human QTLs, curation includes assigning MP, HP, and disease ontology annotations. QTL positions are automatically assigned based on the genomic positions of peak and/or flanking markers or single nucleotide polymorphisms (SNPs). QTL records link to information about related strains, candidate genes, associated markers and related QTLs. Strains: Like QTL records, RGD strain records are either manually curated from the literature or submitted by researchers. Strain records include information about the official symbol of the strain, origin and availability of the strain, associated phenotypes, whether the strain is a model for a human disease, and any information that is available about breeding, behavior, husbandry, etc. Strain records link to information about related genes, alleles, and QTLs, associated strains (e. g. parental strains or substrains) and, where available, strain specific damaging nucleotide variants. For congenic and mutant strains, genomic positions are assigned for the introgressed region (congenic strains) or the location of the mutated sequence (mutant strains). Markers: Because genetic markers such as SSLPs and ESTs have been, and continue to be, used for QTLs and strains, RGD stores marker data for rat, human and mouse. Marker data includes the sequences of the associated forward and reverse PCR primers, genomic positions and links to NCBI's Probe database. Marker records link to associated QTL, strain and gene records. Cell lines: RGD stores cell line records based on imports from Cellosaurus. Although the largest numbers of these are human and mouse cell lines, records are also available for rat, bonobo, dog, squirrel, pig, green monkey and naked mole rat. Ontologies: In order to make RGD's data both human readable and available for computational analysis and retrieval, RGD relies on the use of multiple ontologies. As of July 2021, RGD used 19 different ontologies to express the various types of data applicable to RGD's diverse datatypes. Ontology annotations are assigned manually by curators Ontologies which are imported from outside sources are updated weekly. References: RGD references are scientific publications and resources that have been used for curation of information into the database, and are sources for data objects such as QTLs and strains. For references accessed via NCBI's PubMed, imported data includes the title, authors, citation and PubMed ID, and an RGD ID is generated. In some cases, references are generated as internal records, such as bulk uploads from automated pipelines or personal communications with data sources. These additional references give RGD users an identification of the source of particular pieces and types of data for which PubMed records are not available. Both types of reference records provide links to all of the data curated from that article or source, including genes, QTLs, strains, disease and other ontology annotations. The resources curated for information can be retrieved from the database using the reference search page or links on an object page. Uncurated references are also available, which are known to contain relevant data but have not yet been manually reviewed. These are found as PubMed links listed in the ‘References – uncurated’ section of an object report (e. g. a gene report).
Онтологии: чтобы сделать данные RGD читаемыми для человека и доступными для вычислительного анализа и извлечения, RGD полагается на использование нескольких онтологий. По состоянию на июль 2021 года RGD использовал 19 различных онтологий для выражения различных типов данных, применимых к разнообразным типам данных RGD. Аннотации онтологий присваиваются вручную кураторами. Онтологии, импортируемые из внешних источников, обновляются еженедельно.
Genes: Initial gene records are imported and updated from the National Center for Biotechnology Information's (NCBI's) Gene database on a weekly basis. Data imported during this process includes the Gene ID, Genbank/RefSeq nucleotide and protein sequence identifiers, HomoloGene group IDs and Ensembl Gene, Transcript and Protein IDs. Additional protein related data is imported from the UniProtKB database. RGD curators review the literature and manually curate Gene Ontology (GO), diseases, phenotypes and pathways for rat genes, diseases and pathways for mouse genes, and diseases, phenotypes and pathways for human genes. In addition, the site imports GO annotations for mouse and human genes from the GO Consortium, rat electronic annotations from UniProt and mouse phenotype annotations from the Mouse Genome Database/Mouse Genome Informatics (MGD/MGI). QTLs: RGD's staff manually curates data for rat and human QTLs from the literature where such publications exist or from records directly submitted by researchers. Mouse QTL records, including Mammalian Phenotype (MP) ontology assignments, are imported directly from MGI. For rat and human QTLs, curation includes assigning MP, HP, and disease ontology annotations. QTL positions are automatically assigned based on the genomic positions of peak and/or flanking markers or single nucleotide polymorphisms (SNPs). QTL records link to information about related strains, candidate genes, associated markers and related QTLs. Strains: Like QTL records, RGD strain records are either manually curated from the literature or submitted by researchers. Strain records include information about the official symbol of the strain, origin and availability of the strain, associated phenotypes, whether the strain is a model for a human disease, and any information that is available about breeding, behavior, husbandry, etc. Strain records link to information about related genes, alleles, and QTLs, associated strains (e. g. parental strains or substrains) and, where available, strain specific damaging nucleotide variants. For congenic and mutant strains, genomic positions are assigned for the introgressed region (congenic strains) or the location of the mutated sequence (mutant strains). Markers: Because genetic markers such as SSLPs and ESTs have been, and continue to be, used for QTLs and strains, RGD stores marker data for rat, human and mouse. Marker data includes the sequences of the associated forward and reverse PCR primers, genomic positions and links to NCBI's Probe database. Marker records link to associated QTL, strain and gene records. Cell lines: RGD stores cell line records based on imports from Cellosaurus. Although the largest numbers of these are human and mouse cell lines, records are also available for rat, bonobo, dog, squirrel, pig, green monkey and naked mole rat. Ontologies: In order to make RGD's data both human readable and available for computational analysis and retrieval, RGD relies on the use of multiple ontologies. As of July 2021, RGD used 19 different ontologies to express the various types of data applicable to RGD's diverse datatypes. Ontology annotations are assigned manually by curators Ontologies which are imported from outside sources are updated weekly. References: RGD references are scientific publications and resources that have been used for curation of information into the database, and are sources for data objects such as QTLs and strains. For references accessed via NCBI's PubMed, imported data includes the title, authors, citation and PubMed ID, and an RGD ID is generated. In some cases, references are generated as internal records, such as bulk uploads from automated pipelines or personal communications with data sources. These additional references give RGD users an identification of the source of particular pieces and types of data for which PubMed records are not available. Both types of reference records provide links to all of the data curated from that article or source, including genes, QTLs, strains, disease and other ontology annotations. The resources curated for information can be retrieved from the database using the reference search page or links on an object page. Uncurated references are also available, which are known to contain relevant data but have not yet been manually reviewed. These are found as PubMed links listed in the ‘References – uncurated’ section of an object report (e. g. a gene report).
Ссылки: ссылки RGD — это научные публикации и ресурсы, которые использовались для курирования информации в базу данных и являются источниками объектов данных, таких как QTL и штаммы. Для ссылок, доступных через NCBI PubMed, импортируемые данные включают название, авторов, цитирование и идентификатор PubMed, а также генерируется идентификатор RGD. В некоторых случаях ссылки генерируются как внутренние записи, такие как массовая загрузка из автоматизированных конвейеров или личное общение с источниками данных. Эти дополнительные ссылки предоставляют пользователям RGD идентификацию источника определенных фрагментов и типов данных, для которых записи PubMed недоступны. Оба типа записей ссылок предоставляют ссылки на все данные, курированные из этой статьи или источника, включая гены, QTL, штаммы, заболевания и другие аннотации онтологий. Ресурсы, курируемые для информации, можно получить из базы данных с помощью страницы поиска ссылок или ссылок на странице объекта. Также доступны некурируемые ссылки, которые, как известно, содержат соответствующие данные, но еще не были проверены вручную. Они находятся в виде ссылок PubMed, перечисленных в разделе «Ссылки — некурируемые» отчета об объекте (например, отчета о гене).
Genes: Initial gene records are imported and updated from the National Center for Biotechnology Information's (NCBI's) Gene database on a weekly basis. Data imported during this process includes the Gene ID, Genbank/RefSeq nucleotide and protein sequence identifiers, HomoloGene group IDs and Ensembl Gene, Transcript and Protein IDs. Additional protein related data is imported from the UniProtKB database. RGD curators review the literature and manually curate Gene Ontology (GO), diseases, phenotypes and pathways for rat genes, diseases and pathways for mouse genes, and diseases, phenotypes and pathways for human genes. In addition, the site imports GO annotations for mouse and human genes from the GO Consortium, rat electronic annotations from UniProt and mouse phenotype annotations from the Mouse Genome Database/Mouse Genome Informatics (MGD/MGI). QTLs: RGD's staff manually curates data for rat and human QTLs from the literature where such publications exist or from records directly submitted by researchers. Mouse QTL records, including Mammalian Phenotype (MP) ontology assignments, are imported directly from MGI. For rat and human QTLs, curation includes assigning MP, HP, and disease ontology annotations. QTL positions are automatically assigned based on the genomic positions of peak and/or flanking markers or single nucleotide polymorphisms (SNPs). QTL records link to information about related strains, candidate genes, associated markers and related QTLs. Strains: Like QTL records, RGD strain records are either manually curated from the literature or submitted by researchers. Strain records include information about the official symbol of the strain, origin and availability of the strain, associated phenotypes, whether the strain is a model for a human disease, and any information that is available about breeding, behavior, husbandry, etc. Strain records link to information about related genes, alleles, and QTLs, associated strains (e. g. parental strains or substrains) and, where available, strain specific damaging nucleotide variants. For congenic and mutant strains, genomic positions are assigned for the introgressed region (congenic strains) or the location of the mutated sequence (mutant strains). Markers: Because genetic markers such as SSLPs and ESTs have been, and continue to be, used for QTLs and strains, RGD stores marker data for rat, human and mouse. Marker data includes the sequences of the associated forward and reverse PCR primers, genomic positions and links to NCBI's Probe database. Marker records link to associated QTL, strain and gene records. Cell lines: RGD stores cell line records based on imports from Cellosaurus. Although the largest numbers of these are human and mouse cell lines, records are also available for rat, bonobo, dog, squirrel, pig, green monkey and naked mole rat. Ontologies: In order to make RGD's data both human readable and available for computational analysis and retrieval, RGD relies on the use of multiple ontologies. As of July 2021, RGD used 19 different ontologies to express the various types of data applicable to RGD's diverse datatypes. Ontology annotations are assigned manually by curators Ontologies which are imported from outside sources are updated weekly. References: RGD references are scientific publications and resources that have been used for curation of information into the database, and are sources for data objects such as QTLs and strains. For references accessed via NCBI's PubMed, imported data includes the title, authors, citation and PubMed ID, and an RGD ID is generated. In some cases, references are generated as internal records, such as bulk uploads from automated pipelines or personal communications with data sources. These additional references give RGD users an identification of the source of particular pieces and types of data for which PubMed records are not available. Both types of reference records provide links to all of the data curated from that article or source, including genes, QTLs, strains, disease and other ontology annotations. The resources curated for information can be retrieved from the database using the reference search page or links on an object page. Uncurated references are also available, which are known to contain relevant data but have not yet been manually reviewed. These are found as PubMed links listed in the ‘References – uncurated’ section of an object report (e. g. a gene report).
Инструменты генома
Инструменты генома RGD включают в себя как программное обеспечение, разработанное в RGD, так и инструменты из внешних источников.
Инструменты генома, разработанные в RGD
RGD разрабатывает веб-инструменты, предназначенные для использования данных, хранящихся в базе данных RGD, для анализа у крыс и между видами. Они включают в себя:
OntoMate: OntoMate is an ontology driven, concept based literature search engine that has been developed by RGD as an alternative for the basic PubMed search engine in the gene curation workflow. Converting data from free text in the scientific literature to a structured searchable format is one of the main tasks of all model organism databases. OntoMate tags abstracts with gene names, gene mutations, organism names, disease, and other terms from the ontologies/vocabularies used at RGD. All terms/ entities tagged to an abstract are listed with the abstract in the search results. OntoMate also provides user activated filters for species, date and other parameters relevant to the literature search, which has streamlined the process compared to using PubMed. Besides its usefulness for RGD internal curation processes, the tool is available to all RGD users. Gene Annotator: The Gene Annotator or GA tool takes as input a list of gene symbols, RGD IDs, GenBank accession numbers, Ensembl identifiers, or a chromosomal region and retrieves gene orthologs, external database identifiers and ontology annotations for the corresponding genes in RGD. The data can be downloaded into an Excel spreadsheet or analyzed in the tool. The Annotation Distribution function displays a list of terms in each of seven categories with the percentage of genes from the input list with annotations to each term. The Comparison Heat Map function allows comparisons of annotations for genes in the input list across two ontologies or across two branches of the same ontology. Variant Visualizer: Variant Visualizer (VV) is a viewing and analysis tool for rat strain specific sequence polymorphisms. VV takes as input a list of gene symbols or a genomic region as defined by chromosome, start and stop positions or by two gene or marker symbols. The user must also select their strains of interest from a list of strains for which whole genome sequences exist and can set parameters for the variants in the result set. Output is a heatmap type display of variants. Additional information for individual variants can be viewed in a detail pane display. Multi Ontology Enrichment Tool (MOET): MOET is a web based ontology analysis tool used to identify terms from any or all of the ontologies used by RGD for gene curation (Disease, Pathway, Phenotype, GO, ChEBI) that are over represented in the annotations for those genes, or for orthologs in other species. It outputs a downloadable graph and a list of statistically overrepresented terms in the user’s list of genes using hypergeometric distribution. MOET also displays the corresponding Bonferroni correction and odds ratio on the results page. Gene Ortholog Location Finder (GOLF): GOLF is used to compare genes or positions within regions of interest across RGD species or assemblies. Results are displayed with the corresponding genes/positions in both species or on both assemblies in a side by side tabular view. Inputs and outputs to GOLF can be exported to other RGD tools for analysis or downloaded using the links on the GOLF results page. InterViewer: InterViewer is a protein protein interactive viewer that displays the appropriate information about types of interactions and links to associated genes pertaining to the user’s input. PhenoMiner: PhenoMiner combines phenotypic data from different rat strains, so researchers can use filters to find the quantitative phenotypic data they are looking for. OLGA Object List Generator & Analyzer: OLGA is a search engine designed to allow users to run multiple queries, generate a list of objects from each query and flexibly combine the results using Boolean specifications. OLGA takes as input either a list of object symbols or search parameters based on ontology annotations or position. The final list of genes, QTLs or strains can be downloaded or submitted to the GA Tool, the Variant Visualizer, the Genome Viewer or other RGD tools. Genome Viewer: The Genome Viewer (GViewer) tool provides users with complete genome views of genes, QTLs and mapped strains annotated to a function, biological process, cellular component, phenotype, disease, pathway, or chemical interaction. GViewer allows Boolean searches across multiple ontologies. Output is displayed against a karyotype of the rat genome. Overgo Probe Designer: Overgo probes are pairs of partially overlapping 22mer oligonucleotides derived from repeat masked genomic sequence and used as high specific activity probes for genome mapping. The Overgo Probe Designer tool takes as input a nucleotide sequence and outputs a list of optimized probe sequences containing the requisite 8 nucleotide overlap on their 3' ends. ACP Haplotyper: The ACP Haplotyper creates a visual haplotype that can be used to identify conserved and non conserved chromosomal regions between any of the 48 rat strains characterized as part of the ACP project. For the selected chromosome and between the selected strains, the tool compares the allele size data for microsatellite markers on the selected genetic or RH map.
OntoMate: OntoMate – это поисковая система по научной литературе, основанная на онтологиях и концепциях, разработанная RGD как альтернатива стандартному поиску в PubMed в процессе курирования генов. Преобразование данных из свободного текста в научной литературе в структурированный формат, пригодный для поиска, является одной из основных задач всех баз данных модельных организмов. OntoMate помечает аннотации названиями генов, мутациями генов, названиями организмов, заболеваниями и другими терминами из онтологий и словарей, используемых в RGD. Все термины/сущности, связанные с аннотацией, отображаются вместе с аннотацией в результатах поиска. OntoMate также предоставляет пользователям фильтры по виду, дате и другим параметрам, релевантным для поиска литературы, что упрощает процесс по сравнению с использованием PubMed. Помимо использования в процессах внутренней курирования RGD, инструмент доступен всем пользователям RGD.
OntoMate: OntoMate is an ontology driven, concept based literature search engine that has been developed by RGD as an alternative for the basic PubMed search engine in the gene curation workflow. Converting data from free text in the scientific literature to a structured searchable format is one of the main tasks of all model organism databases. OntoMate tags abstracts with gene names, gene mutations, organism names, disease, and other terms from the ontologies/vocabularies used at RGD. All terms/ entities tagged to an abstract are listed with the abstract in the search results. OntoMate also provides user activated filters for species, date and other parameters relevant to the literature search, which has streamlined the process compared to using PubMed. Besides its usefulness for RGD internal curation processes, the tool is available to all RGD users. Gene Annotator: The Gene Annotator or GA tool takes as input a list of gene symbols, RGD IDs, GenBank accession numbers, Ensembl identifiers, or a chromosomal region and retrieves gene orthologs, external database identifiers and ontology annotations for the corresponding genes in RGD. The data can be downloaded into an Excel spreadsheet or analyzed in the tool. The Annotation Distribution function displays a list of terms in each of seven categories with the percentage of genes from the input list with annotations to each term. The Comparison Heat Map function allows comparisons of annotations for genes in the input list across two ontologies or across two branches of the same ontology. Variant Visualizer: Variant Visualizer (VV) is a viewing and analysis tool for rat strain specific sequence polymorphisms. VV takes as input a list of gene symbols or a genomic region as defined by chromosome, start and stop positions or by two gene or marker symbols. The user must also select their strains of interest from a list of strains for which whole genome sequences exist and can set parameters for the variants in the result set. Output is a heatmap type display of variants. Additional information for individual variants can be viewed in a detail pane display. Multi Ontology Enrichment Tool (MOET): MOET is a web based ontology analysis tool used to identify terms from any or all of the ontologies used by RGD for gene curation (Disease, Pathway, Phenotype, GO, ChEBI) that are over represented in the annotations for those genes, or for orthologs in other species. It outputs a downloadable graph and a list of statistically overrepresented terms in the user’s list of genes using hypergeometric distribution. MOET also displays the corresponding Bonferroni correction and odds ratio on the results page. Gene Ortholog Location Finder (GOLF): GOLF is used to compare genes or positions within regions of interest across RGD species or assemblies. Results are displayed with the corresponding genes/positions in both species or on both assemblies in a side by side tabular view. Inputs and outputs to GOLF can be exported to other RGD tools for analysis or downloaded using the links on the GOLF results page. InterViewer: InterViewer is a protein protein interactive viewer that displays the appropriate information about types of interactions and links to associated genes pertaining to the user’s input. PhenoMiner: PhenoMiner combines phenotypic data from different rat strains, so researchers can use filters to find the quantitative phenotypic data they are looking for. OLGA Object List Generator & Analyzer: OLGA is a search engine designed to allow users to run multiple queries, generate a list of objects from each query and flexibly combine the results using Boolean specifications. OLGA takes as input either a list of object symbols or search parameters based on ontology annotations or position. The final list of genes, QTLs or strains can be downloaded or submitted to the GA Tool, the Variant Visualizer, the Genome Viewer or other RGD tools. Genome Viewer: The Genome Viewer (GViewer) tool provides users with complete genome views of genes, QTLs and mapped strains annotated to a function, biological process, cellular component, phenotype, disease, pathway, or chemical interaction. GViewer allows Boolean searches across multiple ontologies. Output is displayed against a karyotype of the rat genome. Overgo Probe Designer: Overgo probes are pairs of partially overlapping 22mer oligonucleotides derived from repeat masked genomic sequence and used as high specific activity probes for genome mapping. The Overgo Probe Designer tool takes as input a nucleotide sequence and outputs a list of optimized probe sequences containing the requisite 8 nucleotide overlap on their 3' ends. ACP Haplotyper: The ACP Haplotyper creates a visual haplotype that can be used to identify conserved and non conserved chromosomal regions between any of the 48 rat strains characterized as part of the ACP project. For the selected chromosome and between the selected strains, the tool compares the allele size data for microsatellite markers on the selected genetic or RH map.
Gene Annotator: Инструмент Gene Annotator (GA) принимает на вход список символов генов, идентификаторов RGD, номеров доступа GenBank, идентификаторов Ensembl или хромосомный регион и извлекает ортологи генов, идентификаторы внешних баз данных и аннотации онтологий для соответствующих генов в RGD. Данные можно загрузить в электронную таблицу Excel или проанализировать непосредственно в инструменте. Функция «Распределение аннотаций» отображает список терминов в каждой из семи категорий с указанием процента генов из входного списка, имеющих аннотации к каждому термину. Функция «Тепловая карта сравнения» позволяет сравнивать аннотации генов из входного списка в двух онтологиях или в двух ветвях одной и той же онтологии.
OntoMate: OntoMate is an ontology driven, concept based literature search engine that has been developed by RGD as an alternative for the basic PubMed search engine in the gene curation workflow. Converting data from free text in the scientific literature to a structured searchable format is one of the main tasks of all model organism databases. OntoMate tags abstracts with gene names, gene mutations, organism names, disease, and other terms from the ontologies/vocabularies used at RGD. All terms/ entities tagged to an abstract are listed with the abstract in the search results. OntoMate also provides user activated filters for species, date and other parameters relevant to the literature search, which has streamlined the process compared to using PubMed. Besides its usefulness for RGD internal curation processes, the tool is available to all RGD users. Gene Annotator: The Gene Annotator or GA tool takes as input a list of gene symbols, RGD IDs, GenBank accession numbers, Ensembl identifiers, or a chromosomal region and retrieves gene orthologs, external database identifiers and ontology annotations for the corresponding genes in RGD. The data can be downloaded into an Excel spreadsheet or analyzed in the tool. The Annotation Distribution function displays a list of terms in each of seven categories with the percentage of genes from the input list with annotations to each term. The Comparison Heat Map function allows comparisons of annotations for genes in the input list across two ontologies or across two branches of the same ontology. Variant Visualizer: Variant Visualizer (VV) is a viewing and analysis tool for rat strain specific sequence polymorphisms. VV takes as input a list of gene symbols or a genomic region as defined by chromosome, start and stop positions or by two gene or marker symbols. The user must also select their strains of interest from a list of strains for which whole genome sequences exist and can set parameters for the variants in the result set. Output is a heatmap type display of variants. Additional information for individual variants can be viewed in a detail pane display. Multi Ontology Enrichment Tool (MOET): MOET is a web based ontology analysis tool used to identify terms from any or all of the ontologies used by RGD for gene curation (Disease, Pathway, Phenotype, GO, ChEBI) that are over represented in the annotations for those genes, or for orthologs in other species. It outputs a downloadable graph and a list of statistically overrepresented terms in the user’s list of genes using hypergeometric distribution. MOET also displays the corresponding Bonferroni correction and odds ratio on the results page. Gene Ortholog Location Finder (GOLF): GOLF is used to compare genes or positions within regions of interest across RGD species or assemblies. Results are displayed with the corresponding genes/positions in both species or on both assemblies in a side by side tabular view. Inputs and outputs to GOLF can be exported to other RGD tools for analysis or downloaded using the links on the GOLF results page. InterViewer: InterViewer is a protein protein interactive viewer that displays the appropriate information about types of interactions and links to associated genes pertaining to the user’s input. PhenoMiner: PhenoMiner combines phenotypic data from different rat strains, so researchers can use filters to find the quantitative phenotypic data they are looking for. OLGA Object List Generator & Analyzer: OLGA is a search engine designed to allow users to run multiple queries, generate a list of objects from each query and flexibly combine the results using Boolean specifications. OLGA takes as input either a list of object symbols or search parameters based on ontology annotations or position. The final list of genes, QTLs or strains can be downloaded or submitted to the GA Tool, the Variant Visualizer, the Genome Viewer or other RGD tools. Genome Viewer: The Genome Viewer (GViewer) tool provides users with complete genome views of genes, QTLs and mapped strains annotated to a function, biological process, cellular component, phenotype, disease, pathway, or chemical interaction. GViewer allows Boolean searches across multiple ontologies. Output is displayed against a karyotype of the rat genome. Overgo Probe Designer: Overgo probes are pairs of partially overlapping 22mer oligonucleotides derived from repeat masked genomic sequence and used as high specific activity probes for genome mapping. The Overgo Probe Designer tool takes as input a nucleotide sequence and outputs a list of optimized probe sequences containing the requisite 8 nucleotide overlap on their 3' ends. ACP Haplotyper: The ACP Haplotyper creates a visual haplotype that can be used to identify conserved and non conserved chromosomal regions between any of the 48 rat strains characterized as part of the ACP project. For the selected chromosome and between the selected strains, the tool compares the allele size data for microsatellite markers on the selected genetic or RH map.
Variant Visualizer: Variant Visualizer (VV) – это инструмент для просмотра и анализа полиморфизмов последовательности, специфичных для линий крыс. VV принимает на вход список символов генов или геномный регион, определяемый хромосомой, начальной и конечной позициями, или двумя символами генов или маркеров. Пользователь также должен выбрать интересующие его линии крыс из списка линий, для которых существуют полные геномные последовательности, и задать параметры для вариантов в результирующем наборе. Результатом является отображение вариантов в виде тепловой карты. Дополнительная информация об отдельных вариантах может быть просмотрена в панели подробностей.
OntoMate: OntoMate is an ontology driven, concept based literature search engine that has been developed by RGD as an alternative for the basic PubMed search engine in the gene curation workflow. Converting data from free text in the scientific literature to a structured searchable format is one of the main tasks of all model organism databases. OntoMate tags abstracts with gene names, gene mutations, organism names, disease, and other terms from the ontologies/vocabularies used at RGD. All terms/ entities tagged to an abstract are listed with the abstract in the search results. OntoMate also provides user activated filters for species, date and other parameters relevant to the literature search, which has streamlined the process compared to using PubMed. Besides its usefulness for RGD internal curation processes, the tool is available to all RGD users. Gene Annotator: The Gene Annotator or GA tool takes as input a list of gene symbols, RGD IDs, GenBank accession numbers, Ensembl identifiers, or a chromosomal region and retrieves gene orthologs, external database identifiers and ontology annotations for the corresponding genes in RGD. The data can be downloaded into an Excel spreadsheet or analyzed in the tool. The Annotation Distribution function displays a list of terms in each of seven categories with the percentage of genes from the input list with annotations to each term. The Comparison Heat Map function allows comparisons of annotations for genes in the input list across two ontologies or across two branches of the same ontology. Variant Visualizer: Variant Visualizer (VV) is a viewing and analysis tool for rat strain specific sequence polymorphisms. VV takes as input a list of gene symbols or a genomic region as defined by chromosome, start and stop positions or by two gene or marker symbols. The user must also select their strains of interest from a list of strains for which whole genome sequences exist and can set parameters for the variants in the result set. Output is a heatmap type display of variants. Additional information for individual variants can be viewed in a detail pane display. Multi Ontology Enrichment Tool (MOET): MOET is a web based ontology analysis tool used to identify terms from any or all of the ontologies used by RGD for gene curation (Disease, Pathway, Phenotype, GO, ChEBI) that are over represented in the annotations for those genes, or for orthologs in other species. It outputs a downloadable graph and a list of statistically overrepresented terms in the user’s list of genes using hypergeometric distribution. MOET also displays the corresponding Bonferroni correction and odds ratio on the results page. Gene Ortholog Location Finder (GOLF): GOLF is used to compare genes or positions within regions of interest across RGD species or assemblies. Results are displayed with the corresponding genes/positions in both species or on both assemblies in a side by side tabular view. Inputs and outputs to GOLF can be exported to other RGD tools for analysis or downloaded using the links on the GOLF results page. InterViewer: InterViewer is a protein protein interactive viewer that displays the appropriate information about types of interactions and links to associated genes pertaining to the user’s input. PhenoMiner: PhenoMiner combines phenotypic data from different rat strains, so researchers can use filters to find the quantitative phenotypic data they are looking for. OLGA Object List Generator & Analyzer: OLGA is a search engine designed to allow users to run multiple queries, generate a list of objects from each query and flexibly combine the results using Boolean specifications. OLGA takes as input either a list of object symbols or search parameters based on ontology annotations or position. The final list of genes, QTLs or strains can be downloaded or submitted to the GA Tool, the Variant Visualizer, the Genome Viewer or other RGD tools. Genome Viewer: The Genome Viewer (GViewer) tool provides users with complete genome views of genes, QTLs and mapped strains annotated to a function, biological process, cellular component, phenotype, disease, pathway, or chemical interaction. GViewer allows Boolean searches across multiple ontologies. Output is displayed against a karyotype of the rat genome. Overgo Probe Designer: Overgo probes are pairs of partially overlapping 22mer oligonucleotides derived from repeat masked genomic sequence and used as high specific activity probes for genome mapping. The Overgo Probe Designer tool takes as input a nucleotide sequence and outputs a list of optimized probe sequences containing the requisite 8 nucleotide overlap on their 3' ends. ACP Haplotyper: The ACP Haplotyper creates a visual haplotype that can be used to identify conserved and non conserved chromosomal regions between any of the 48 rat strains characterized as part of the ACP project. For the selected chromosome and between the selected strains, the tool compares the allele size data for microsatellite markers on the selected genetic or RH map.
Multi Ontology Enrichment Tool (MOET): MOET – это инструмент для анализа онтологий, используемый для выявления терминов из любой или всех онтологий, используемых RGD для курирования генов (Заболевание, Путь, Фенотип, GO, ChEBI), которые представлены в аннотациях для этих генов или для ортологов у других видов. Он выдает загружаемый график и список статистически значимых терминов в списке генов пользователя, рассчитанных с использованием гипергеометрического распределения. MOET также отображает соответствующую поправку Бонферрони и отношение шансов на странице результатов.
OntoMate: OntoMate is an ontology driven, concept based literature search engine that has been developed by RGD as an alternative for the basic PubMed search engine in the gene curation workflow. Converting data from free text in the scientific literature to a structured searchable format is one of the main tasks of all model organism databases. OntoMate tags abstracts with gene names, gene mutations, organism names, disease, and other terms from the ontologies/vocabularies used at RGD. All terms/ entities tagged to an abstract are listed with the abstract in the search results. OntoMate also provides user activated filters for species, date and other parameters relevant to the literature search, which has streamlined the process compared to using PubMed. Besides its usefulness for RGD internal curation processes, the tool is available to all RGD users. Gene Annotator: The Gene Annotator or GA tool takes as input a list of gene symbols, RGD IDs, GenBank accession numbers, Ensembl identifiers, or a chromosomal region and retrieves gene orthologs, external database identifiers and ontology annotations for the corresponding genes in RGD. The data can be downloaded into an Excel spreadsheet or analyzed in the tool. The Annotation Distribution function displays a list of terms in each of seven categories with the percentage of genes from the input list with annotations to each term. The Comparison Heat Map function allows comparisons of annotations for genes in the input list across two ontologies or across two branches of the same ontology. Variant Visualizer: Variant Visualizer (VV) is a viewing and analysis tool for rat strain specific sequence polymorphisms. VV takes as input a list of gene symbols or a genomic region as defined by chromosome, start and stop positions or by two gene or marker symbols. The user must also select their strains of interest from a list of strains for which whole genome sequences exist and can set parameters for the variants in the result set. Output is a heatmap type display of variants. Additional information for individual variants can be viewed in a detail pane display. Multi Ontology Enrichment Tool (MOET): MOET is a web based ontology analysis tool used to identify terms from any or all of the ontologies used by RGD for gene curation (Disease, Pathway, Phenotype, GO, ChEBI) that are over represented in the annotations for those genes, or for orthologs in other species. It outputs a downloadable graph and a list of statistically overrepresented terms in the user’s list of genes using hypergeometric distribution. MOET also displays the corresponding Bonferroni correction and odds ratio on the results page. Gene Ortholog Location Finder (GOLF): GOLF is used to compare genes or positions within regions of interest across RGD species or assemblies. Results are displayed with the corresponding genes/positions in both species or on both assemblies in a side by side tabular view. Inputs and outputs to GOLF can be exported to other RGD tools for analysis or downloaded using the links on the GOLF results page. InterViewer: InterViewer is a protein protein interactive viewer that displays the appropriate information about types of interactions and links to associated genes pertaining to the user’s input. PhenoMiner: PhenoMiner combines phenotypic data from different rat strains, so researchers can use filters to find the quantitative phenotypic data they are looking for. OLGA Object List Generator & Analyzer: OLGA is a search engine designed to allow users to run multiple queries, generate a list of objects from each query and flexibly combine the results using Boolean specifications. OLGA takes as input either a list of object symbols or search parameters based on ontology annotations or position. The final list of genes, QTLs or strains can be downloaded or submitted to the GA Tool, the Variant Visualizer, the Genome Viewer or other RGD tools. Genome Viewer: The Genome Viewer (GViewer) tool provides users with complete genome views of genes, QTLs and mapped strains annotated to a function, biological process, cellular component, phenotype, disease, pathway, or chemical interaction. GViewer allows Boolean searches across multiple ontologies. Output is displayed against a karyotype of the rat genome. Overgo Probe Designer: Overgo probes are pairs of partially overlapping 22mer oligonucleotides derived from repeat masked genomic sequence and used as high specific activity probes for genome mapping. The Overgo Probe Designer tool takes as input a nucleotide sequence and outputs a list of optimized probe sequences containing the requisite 8 nucleotide overlap on their 3' ends. ACP Haplotyper: The ACP Haplotyper creates a visual haplotype that can be used to identify conserved and non conserved chromosomal regions between any of the 48 rat strains characterized as part of the ACP project. For the selected chromosome and between the selected strains, the tool compares the allele size data for microsatellite markers on the selected genetic or RH map.
Gene Ortholog Location Finder (GOLF): GOLF используется для сравнения генов или позиций в интересующих регионах у различных видов или сборок RGD. Результаты отображаются в виде таблицы, содержащей соответствующие гены/позиции для обоих видов или обеих сборок. Входные и выходные данные GOLF могут быть экспортированы в другие инструменты RGD для анализа или загружены по ссылкам на странице результатов GOLF.
OntoMate: OntoMate is an ontology driven, concept based literature search engine that has been developed by RGD as an alternative for the basic PubMed search engine in the gene curation workflow. Converting data from free text in the scientific literature to a structured searchable format is one of the main tasks of all model organism databases. OntoMate tags abstracts with gene names, gene mutations, organism names, disease, and other terms from the ontologies/vocabularies used at RGD. All terms/ entities tagged to an abstract are listed with the abstract in the search results. OntoMate also provides user activated filters for species, date and other parameters relevant to the literature search, which has streamlined the process compared to using PubMed. Besides its usefulness for RGD internal curation processes, the tool is available to all RGD users. Gene Annotator: The Gene Annotator or GA tool takes as input a list of gene symbols, RGD IDs, GenBank accession numbers, Ensembl identifiers, or a chromosomal region and retrieves gene orthologs, external database identifiers and ontology annotations for the corresponding genes in RGD. The data can be downloaded into an Excel spreadsheet or analyzed in the tool. The Annotation Distribution function displays a list of terms in each of seven categories with the percentage of genes from the input list with annotations to each term. The Comparison Heat Map function allows comparisons of annotations for genes in the input list across two ontologies or across two branches of the same ontology. Variant Visualizer: Variant Visualizer (VV) is a viewing and analysis tool for rat strain specific sequence polymorphisms. VV takes as input a list of gene symbols or a genomic region as defined by chromosome, start and stop positions or by two gene or marker symbols. The user must also select their strains of interest from a list of strains for which whole genome sequences exist and can set parameters for the variants in the result set. Output is a heatmap type display of variants. Additional information for individual variants can be viewed in a detail pane display. Multi Ontology Enrichment Tool (MOET): MOET is a web based ontology analysis tool used to identify terms from any or all of the ontologies used by RGD for gene curation (Disease, Pathway, Phenotype, GO, ChEBI) that are over represented in the annotations for those genes, or for orthologs in other species. It outputs a downloadable graph and a list of statistically overrepresented terms in the user’s list of genes using hypergeometric distribution. MOET also displays the corresponding Bonferroni correction and odds ratio on the results page. Gene Ortholog Location Finder (GOLF): GOLF is used to compare genes or positions within regions of interest across RGD species or assemblies. Results are displayed with the corresponding genes/positions in both species or on both assemblies in a side by side tabular view. Inputs and outputs to GOLF can be exported to other RGD tools for analysis or downloaded using the links on the GOLF results page. InterViewer: InterViewer is a protein protein interactive viewer that displays the appropriate information about types of interactions and links to associated genes pertaining to the user’s input. PhenoMiner: PhenoMiner combines phenotypic data from different rat strains, so researchers can use filters to find the quantitative phenotypic data they are looking for. OLGA Object List Generator & Analyzer: OLGA is a search engine designed to allow users to run multiple queries, generate a list of objects from each query and flexibly combine the results using Boolean specifications. OLGA takes as input either a list of object symbols or search parameters based on ontology annotations or position. The final list of genes, QTLs or strains can be downloaded or submitted to the GA Tool, the Variant Visualizer, the Genome Viewer or other RGD tools. Genome Viewer: The Genome Viewer (GViewer) tool provides users with complete genome views of genes, QTLs and mapped strains annotated to a function, biological process, cellular component, phenotype, disease, pathway, or chemical interaction. GViewer allows Boolean searches across multiple ontologies. Output is displayed against a karyotype of the rat genome. Overgo Probe Designer: Overgo probes are pairs of partially overlapping 22mer oligonucleotides derived from repeat masked genomic sequence and used as high specific activity probes for genome mapping. The Overgo Probe Designer tool takes as input a nucleotide sequence and outputs a list of optimized probe sequences containing the requisite 8 nucleotide overlap on their 3' ends. ACP Haplotyper: The ACP Haplotyper creates a visual haplotype that can be used to identify conserved and non conserved chromosomal regions between any of the 48 rat strains characterized as part of the ACP project. For the selected chromosome and between the selected strains, the tool compares the allele size data for microsatellite markers on the selected genetic or RH map.
InterViewer: InterViewer – это интерактивный просмотрщик белок-белковых взаимодействий, отображающий соответствующую информацию о типах взаимодействий и ссылках на связанные гены, относящиеся к введенным пользователем данным.
OntoMate: OntoMate is an ontology driven, concept based literature search engine that has been developed by RGD as an alternative for the basic PubMed search engine in the gene curation workflow. Converting data from free text in the scientific literature to a structured searchable format is one of the main tasks of all model organism databases. OntoMate tags abstracts with gene names, gene mutations, organism names, disease, and other terms from the ontologies/vocabularies used at RGD. All terms/ entities tagged to an abstract are listed with the abstract in the search results. OntoMate also provides user activated filters for species, date and other parameters relevant to the literature search, which has streamlined the process compared to using PubMed. Besides its usefulness for RGD internal curation processes, the tool is available to all RGD users. Gene Annotator: The Gene Annotator or GA tool takes as input a list of gene symbols, RGD IDs, GenBank accession numbers, Ensembl identifiers, or a chromosomal region and retrieves gene orthologs, external database identifiers and ontology annotations for the corresponding genes in RGD. The data can be downloaded into an Excel spreadsheet or analyzed in the tool. The Annotation Distribution function displays a list of terms in each of seven categories with the percentage of genes from the input list with annotations to each term. The Comparison Heat Map function allows comparisons of annotations for genes in the input list across two ontologies or across two branches of the same ontology. Variant Visualizer: Variant Visualizer (VV) is a viewing and analysis tool for rat strain specific sequence polymorphisms. VV takes as input a list of gene symbols or a genomic region as defined by chromosome, start and stop positions or by two gene or marker symbols. The user must also select their strains of interest from a list of strains for which whole genome sequences exist and can set parameters for the variants in the result set. Output is a heatmap type display of variants. Additional information for individual variants can be viewed in a detail pane display. Multi Ontology Enrichment Tool (MOET): MOET is a web based ontology analysis tool used to identify terms from any or all of the ontologies used by RGD for gene curation (Disease, Pathway, Phenotype, GO, ChEBI) that are over represented in the annotations for those genes, or for orthologs in other species. It outputs a downloadable graph and a list of statistically overrepresented terms in the user’s list of genes using hypergeometric distribution. MOET also displays the corresponding Bonferroni correction and odds ratio on the results page. Gene Ortholog Location Finder (GOLF): GOLF is used to compare genes or positions within regions of interest across RGD species or assemblies. Results are displayed with the corresponding genes/positions in both species or on both assemblies in a side by side tabular view. Inputs and outputs to GOLF can be exported to other RGD tools for analysis or downloaded using the links on the GOLF results page. InterViewer: InterViewer is a protein protein interactive viewer that displays the appropriate information about types of interactions and links to associated genes pertaining to the user’s input. PhenoMiner: PhenoMiner combines phenotypic data from different rat strains, so researchers can use filters to find the quantitative phenotypic data they are looking for. OLGA Object List Generator & Analyzer: OLGA is a search engine designed to allow users to run multiple queries, generate a list of objects from each query and flexibly combine the results using Boolean specifications. OLGA takes as input either a list of object symbols or search parameters based on ontology annotations or position. The final list of genes, QTLs or strains can be downloaded or submitted to the GA Tool, the Variant Visualizer, the Genome Viewer or other RGD tools. Genome Viewer: The Genome Viewer (GViewer) tool provides users with complete genome views of genes, QTLs and mapped strains annotated to a function, biological process, cellular component, phenotype, disease, pathway, or chemical interaction. GViewer allows Boolean searches across multiple ontologies. Output is displayed against a karyotype of the rat genome. Overgo Probe Designer: Overgo probes are pairs of partially overlapping 22mer oligonucleotides derived from repeat masked genomic sequence and used as high specific activity probes for genome mapping. The Overgo Probe Designer tool takes as input a nucleotide sequence and outputs a list of optimized probe sequences containing the requisite 8 nucleotide overlap on their 3' ends. ACP Haplotyper: The ACP Haplotyper creates a visual haplotype that can be used to identify conserved and non conserved chromosomal regions between any of the 48 rat strains characterized as part of the ACP project. For the selected chromosome and between the selected strains, the tool compares the allele size data for microsatellite markers on the selected genetic or RH map.
PhenoMiner: PhenoMiner объединяет фенотипические данные из разных линий крыс, позволяя исследователям использовать фильтры для поиска количественных фенотипических данных.
OntoMate: OntoMate is an ontology driven, concept based literature search engine that has been developed by RGD as an alternative for the basic PubMed search engine in the gene curation workflow. Converting data from free text in the scientific literature to a structured searchable format is one of the main tasks of all model organism databases. OntoMate tags abstracts with gene names, gene mutations, organism names, disease, and other terms from the ontologies/vocabularies used at RGD. All terms/ entities tagged to an abstract are listed with the abstract in the search results. OntoMate also provides user activated filters for species, date and other parameters relevant to the literature search, which has streamlined the process compared to using PubMed. Besides its usefulness for RGD internal curation processes, the tool is available to all RGD users. Gene Annotator: The Gene Annotator or GA tool takes as input a list of gene symbols, RGD IDs, GenBank accession numbers, Ensembl identifiers, or a chromosomal region and retrieves gene orthologs, external database identifiers and ontology annotations for the corresponding genes in RGD. The data can be downloaded into an Excel spreadsheet or analyzed in the tool. The Annotation Distribution function displays a list of terms in each of seven categories with the percentage of genes from the input list with annotations to each term. The Comparison Heat Map function allows comparisons of annotations for genes in the input list across two ontologies or across two branches of the same ontology. Variant Visualizer: Variant Visualizer (VV) is a viewing and analysis tool for rat strain specific sequence polymorphisms. VV takes as input a list of gene symbols or a genomic region as defined by chromosome, start and stop positions or by two gene or marker symbols. The user must also select their strains of interest from a list of strains for which whole genome sequences exist and can set parameters for the variants in the result set. Output is a heatmap type display of variants. Additional information for individual variants can be viewed in a detail pane display. Multi Ontology Enrichment Tool (MOET): MOET is a web based ontology analysis tool used to identify terms from any or all of the ontologies used by RGD for gene curation (Disease, Pathway, Phenotype, GO, ChEBI) that are over represented in the annotations for those genes, or for orthologs in other species. It outputs a downloadable graph and a list of statistically overrepresented terms in the user’s list of genes using hypergeometric distribution. MOET also displays the corresponding Bonferroni correction and odds ratio on the results page. Gene Ortholog Location Finder (GOLF): GOLF is used to compare genes or positions within regions of interest across RGD species or assemblies. Results are displayed with the corresponding genes/positions in both species or on both assemblies in a side by side tabular view. Inputs and outputs to GOLF can be exported to other RGD tools for analysis or downloaded using the links on the GOLF results page. InterViewer: InterViewer is a protein protein interactive viewer that displays the appropriate information about types of interactions and links to associated genes pertaining to the user’s input. PhenoMiner: PhenoMiner combines phenotypic data from different rat strains, so researchers can use filters to find the quantitative phenotypic data they are looking for. OLGA Object List Generator & Analyzer: OLGA is a search engine designed to allow users to run multiple queries, generate a list of objects from each query and flexibly combine the results using Boolean specifications. OLGA takes as input either a list of object symbols or search parameters based on ontology annotations or position. The final list of genes, QTLs or strains can be downloaded or submitted to the GA Tool, the Variant Visualizer, the Genome Viewer or other RGD tools. Genome Viewer: The Genome Viewer (GViewer) tool provides users with complete genome views of genes, QTLs and mapped strains annotated to a function, biological process, cellular component, phenotype, disease, pathway, or chemical interaction. GViewer allows Boolean searches across multiple ontologies. Output is displayed against a karyotype of the rat genome. Overgo Probe Designer: Overgo probes are pairs of partially overlapping 22mer oligonucleotides derived from repeat masked genomic sequence and used as high specific activity probes for genome mapping. The Overgo Probe Designer tool takes as input a nucleotide sequence and outputs a list of optimized probe sequences containing the requisite 8 nucleotide overlap on their 3' ends. ACP Haplotyper: The ACP Haplotyper creates a visual haplotype that can be used to identify conserved and non conserved chromosomal regions between any of the 48 rat strains characterized as part of the ACP project. For the selected chromosome and between the selected strains, the tool compares the allele size data for microsatellite markers on the selected genetic or RH map.
OLGA – Object List Generator & Analyzer: OLGA – это поисковая система, позволяющая пользователям выполнять несколько запросов, генерировать список объектов из каждого запроса и гибко объединять результаты с использованием булевых спецификаций. OLGA принимает на вход либо список символов объектов, либо параметры поиска, основанные на аннотациях онтологии или позиции. Окончательный список генов, QTL или линий можно загрузить или отправить в инструмент GA, Variant Visualizer, Genome Viewer или другие инструменты RGD.
OntoMate: OntoMate is an ontology driven, concept based literature search engine that has been developed by RGD as an alternative for the basic PubMed search engine in the gene curation workflow. Converting data from free text in the scientific literature to a structured searchable format is one of the main tasks of all model organism databases. OntoMate tags abstracts with gene names, gene mutations, organism names, disease, and other terms from the ontologies/vocabularies used at RGD. All terms/ entities tagged to an abstract are listed with the abstract in the search results. OntoMate also provides user activated filters for species, date and other parameters relevant to the literature search, which has streamlined the process compared to using PubMed. Besides its usefulness for RGD internal curation processes, the tool is available to all RGD users. Gene Annotator: The Gene Annotator or GA tool takes as input a list of gene symbols, RGD IDs, GenBank accession numbers, Ensembl identifiers, or a chromosomal region and retrieves gene orthologs, external database identifiers and ontology annotations for the corresponding genes in RGD. The data can be downloaded into an Excel spreadsheet or analyzed in the tool. The Annotation Distribution function displays a list of terms in each of seven categories with the percentage of genes from the input list with annotations to each term. The Comparison Heat Map function allows comparisons of annotations for genes in the input list across two ontologies or across two branches of the same ontology. Variant Visualizer: Variant Visualizer (VV) is a viewing and analysis tool for rat strain specific sequence polymorphisms. VV takes as input a list of gene symbols or a genomic region as defined by chromosome, start and stop positions or by two gene or marker symbols. The user must also select their strains of interest from a list of strains for which whole genome sequences exist and can set parameters for the variants in the result set. Output is a heatmap type display of variants. Additional information for individual variants can be viewed in a detail pane display. Multi Ontology Enrichment Tool (MOET): MOET is a web based ontology analysis tool used to identify terms from any or all of the ontologies used by RGD for gene curation (Disease, Pathway, Phenotype, GO, ChEBI) that are over represented in the annotations for those genes, or for orthologs in other species. It outputs a downloadable graph and a list of statistically overrepresented terms in the user’s list of genes using hypergeometric distribution. MOET also displays the corresponding Bonferroni correction and odds ratio on the results page. Gene Ortholog Location Finder (GOLF): GOLF is used to compare genes or positions within regions of interest across RGD species or assemblies. Results are displayed with the corresponding genes/positions in both species or on both assemblies in a side by side tabular view. Inputs and outputs to GOLF can be exported to other RGD tools for analysis or downloaded using the links on the GOLF results page. InterViewer: InterViewer is a protein protein interactive viewer that displays the appropriate information about types of interactions and links to associated genes pertaining to the user’s input. PhenoMiner: PhenoMiner combines phenotypic data from different rat strains, so researchers can use filters to find the quantitative phenotypic data they are looking for. OLGA Object List Generator & Analyzer: OLGA is a search engine designed to allow users to run multiple queries, generate a list of objects from each query and flexibly combine the results using Boolean specifications. OLGA takes as input either a list of object symbols or search parameters based on ontology annotations or position. The final list of genes, QTLs or strains can be downloaded or submitted to the GA Tool, the Variant Visualizer, the Genome Viewer or other RGD tools. Genome Viewer: The Genome Viewer (GViewer) tool provides users with complete genome views of genes, QTLs and mapped strains annotated to a function, biological process, cellular component, phenotype, disease, pathway, or chemical interaction. GViewer allows Boolean searches across multiple ontologies. Output is displayed against a karyotype of the rat genome. Overgo Probe Designer: Overgo probes are pairs of partially overlapping 22mer oligonucleotides derived from repeat masked genomic sequence and used as high specific activity probes for genome mapping. The Overgo Probe Designer tool takes as input a nucleotide sequence and outputs a list of optimized probe sequences containing the requisite 8 nucleotide overlap on their 3' ends. ACP Haplotyper: The ACP Haplotyper creates a visual haplotype that can be used to identify conserved and non conserved chromosomal regions between any of the 48 rat strains characterized as part of the ACP project. For the selected chromosome and between the selected strains, the tool compares the allele size data for microsatellite markers on the selected genetic or RH map.
Genome Viewer: Инструмент Genome Viewer (GViewer) предоставляет пользователям полные геномные представления генов, QTL и картированных линий, аннотированных функцией, биологическим процессом, клеточным компонентом, фенотипом, заболеванием, путем или химическим взаимодействием. GViewer позволяет выполнять булевы поиски по нескольким онтологиям. Результаты отображаются на фоне кариотипа генома крысы.
OntoMate: OntoMate is an ontology driven, concept based literature search engine that has been developed by RGD as an alternative for the basic PubMed search engine in the gene curation workflow. Converting data from free text in the scientific literature to a structured searchable format is one of the main tasks of all model organism databases. OntoMate tags abstracts with gene names, gene mutations, organism names, disease, and other terms from the ontologies/vocabularies used at RGD. All terms/ entities tagged to an abstract are listed with the abstract in the search results. OntoMate also provides user activated filters for species, date and other parameters relevant to the literature search, which has streamlined the process compared to using PubMed. Besides its usefulness for RGD internal curation processes, the tool is available to all RGD users. Gene Annotator: The Gene Annotator or GA tool takes as input a list of gene symbols, RGD IDs, GenBank accession numbers, Ensembl identifiers, or a chromosomal region and retrieves gene orthologs, external database identifiers and ontology annotations for the corresponding genes in RGD. The data can be downloaded into an Excel spreadsheet or analyzed in the tool. The Annotation Distribution function displays a list of terms in each of seven categories with the percentage of genes from the input list with annotations to each term. The Comparison Heat Map function allows comparisons of annotations for genes in the input list across two ontologies or across two branches of the same ontology. Variant Visualizer: Variant Visualizer (VV) is a viewing and analysis tool for rat strain specific sequence polymorphisms. VV takes as input a list of gene symbols or a genomic region as defined by chromosome, start and stop positions or by two gene or marker symbols. The user must also select their strains of interest from a list of strains for which whole genome sequences exist and can set parameters for the variants in the result set. Output is a heatmap type display of variants. Additional information for individual variants can be viewed in a detail pane display. Multi Ontology Enrichment Tool (MOET): MOET is a web based ontology analysis tool used to identify terms from any or all of the ontologies used by RGD for gene curation (Disease, Pathway, Phenotype, GO, ChEBI) that are over represented in the annotations for those genes, or for orthologs in other species. It outputs a downloadable graph and a list of statistically overrepresented terms in the user’s list of genes using hypergeometric distribution. MOET also displays the corresponding Bonferroni correction and odds ratio on the results page. Gene Ortholog Location Finder (GOLF): GOLF is used to compare genes or positions within regions of interest across RGD species or assemblies. Results are displayed with the corresponding genes/positions in both species or on both assemblies in a side by side tabular view. Inputs and outputs to GOLF can be exported to other RGD tools for analysis or downloaded using the links on the GOLF results page. InterViewer: InterViewer is a protein protein interactive viewer that displays the appropriate information about types of interactions and links to associated genes pertaining to the user’s input. PhenoMiner: PhenoMiner combines phenotypic data from different rat strains, so researchers can use filters to find the quantitative phenotypic data they are looking for. OLGA Object List Generator & Analyzer: OLGA is a search engine designed to allow users to run multiple queries, generate a list of objects from each query and flexibly combine the results using Boolean specifications. OLGA takes as input either a list of object symbols or search parameters based on ontology annotations or position. The final list of genes, QTLs or strains can be downloaded or submitted to the GA Tool, the Variant Visualizer, the Genome Viewer or other RGD tools. Genome Viewer: The Genome Viewer (GViewer) tool provides users with complete genome views of genes, QTLs and mapped strains annotated to a function, biological process, cellular component, phenotype, disease, pathway, or chemical interaction. GViewer allows Boolean searches across multiple ontologies. Output is displayed against a karyotype of the rat genome. Overgo Probe Designer: Overgo probes are pairs of partially overlapping 22mer oligonucleotides derived from repeat masked genomic sequence and used as high specific activity probes for genome mapping. The Overgo Probe Designer tool takes as input a nucleotide sequence and outputs a list of optimized probe sequences containing the requisite 8 nucleotide overlap on their 3' ends. ACP Haplotyper: The ACP Haplotyper creates a visual haplotype that can be used to identify conserved and non conserved chromosomal regions between any of the 48 rat strains characterized as part of the ACP project. For the selected chromosome and between the selected strains, the tool compares the allele size data for microsatellite markers on the selected genetic or RH map.
Overgo Probe Designer: Overgo-зонды – это пары частично перекрывающихся олигонуклеотидов длиной 22 основания, полученных из геномной последовательности с маскировкой повторов, и используемых в качестве высокоспецифичных зондов для картирования генома. Инструмент Overgo Probe Designer принимает на вход последовательность нуклеотидов и выдает список оптимизированных последовательностей зондов, содержащих необходимый перехлест в 8 нуклеотидов на 3'-концах.
OntoMate: OntoMate is an ontology driven, concept based literature search engine that has been developed by RGD as an alternative for the basic PubMed search engine in the gene curation workflow. Converting data from free text in the scientific literature to a structured searchable format is one of the main tasks of all model organism databases. OntoMate tags abstracts with gene names, gene mutations, organism names, disease, and other terms from the ontologies/vocabularies used at RGD. All terms/ entities tagged to an abstract are listed with the abstract in the search results. OntoMate also provides user activated filters for species, date and other parameters relevant to the literature search, which has streamlined the process compared to using PubMed. Besides its usefulness for RGD internal curation processes, the tool is available to all RGD users. Gene Annotator: The Gene Annotator or GA tool takes as input a list of gene symbols, RGD IDs, GenBank accession numbers, Ensembl identifiers, or a chromosomal region and retrieves gene orthologs, external database identifiers and ontology annotations for the corresponding genes in RGD. The data can be downloaded into an Excel spreadsheet or analyzed in the tool. The Annotation Distribution function displays a list of terms in each of seven categories with the percentage of genes from the input list with annotations to each term. The Comparison Heat Map function allows comparisons of annotations for genes in the input list across two ontologies or across two branches of the same ontology. Variant Visualizer: Variant Visualizer (VV) is a viewing and analysis tool for rat strain specific sequence polymorphisms. VV takes as input a list of gene symbols or a genomic region as defined by chromosome, start and stop positions or by two gene or marker symbols. The user must also select their strains of interest from a list of strains for which whole genome sequences exist and can set parameters for the variants in the result set. Output is a heatmap type display of variants. Additional information for individual variants can be viewed in a detail pane display. Multi Ontology Enrichment Tool (MOET): MOET is a web based ontology analysis tool used to identify terms from any or all of the ontologies used by RGD for gene curation (Disease, Pathway, Phenotype, GO, ChEBI) that are over represented in the annotations for those genes, or for orthologs in other species. It outputs a downloadable graph and a list of statistically overrepresented terms in the user’s list of genes using hypergeometric distribution. MOET also displays the corresponding Bonferroni correction and odds ratio on the results page. Gene Ortholog Location Finder (GOLF): GOLF is used to compare genes or positions within regions of interest across RGD species or assemblies. Results are displayed with the corresponding genes/positions in both species or on both assemblies in a side by side tabular view. Inputs and outputs to GOLF can be exported to other RGD tools for analysis or downloaded using the links on the GOLF results page. InterViewer: InterViewer is a protein protein interactive viewer that displays the appropriate information about types of interactions and links to associated genes pertaining to the user’s input. PhenoMiner: PhenoMiner combines phenotypic data from different rat strains, so researchers can use filters to find the quantitative phenotypic data they are looking for. OLGA Object List Generator & Analyzer: OLGA is a search engine designed to allow users to run multiple queries, generate a list of objects from each query and flexibly combine the results using Boolean specifications. OLGA takes as input either a list of object symbols or search parameters based on ontology annotations or position. The final list of genes, QTLs or strains can be downloaded or submitted to the GA Tool, the Variant Visualizer, the Genome Viewer or other RGD tools. Genome Viewer: The Genome Viewer (GViewer) tool provides users with complete genome views of genes, QTLs and mapped strains annotated to a function, biological process, cellular component, phenotype, disease, pathway, or chemical interaction. GViewer allows Boolean searches across multiple ontologies. Output is displayed against a karyotype of the rat genome. Overgo Probe Designer: Overgo probes are pairs of partially overlapping 22mer oligonucleotides derived from repeat masked genomic sequence and used as high specific activity probes for genome mapping. The Overgo Probe Designer tool takes as input a nucleotide sequence and outputs a list of optimized probe sequences containing the requisite 8 nucleotide overlap on their 3' ends. ACP Haplotyper: The ACP Haplotyper creates a visual haplotype that can be used to identify conserved and non conserved chromosomal regions between any of the 48 rat strains characterized as part of the ACP project. For the selected chromosome and between the selected strains, the tool compares the allele size data for microsatellite markers on the selected genetic or RH map.
ACP Haplotyper: ACP Haplotyper создает визуальный гаплотип, который можно использовать для идентификации консервативных и некосервативных хромосомных регионов между любой из 48 линий крыс, охарактеризованных в рамках проекта ACP. Для выбранной хромосомы и между выбранными линиями инструмент сравнивает данные о размере аллелей для микросателлитных маркеров на выбранной генетической или RH-карте.
OntoMate: OntoMate is an ontology driven, concept based literature search engine that has been developed by RGD as an alternative for the basic PubMed search engine in the gene curation workflow. Converting data from free text in the scientific literature to a structured searchable format is one of the main tasks of all model organism databases. OntoMate tags abstracts with gene names, gene mutations, organism names, disease, and other terms from the ontologies/vocabularies used at RGD. All terms/ entities tagged to an abstract are listed with the abstract in the search results. OntoMate also provides user activated filters for species, date and other parameters relevant to the literature search, which has streamlined the process compared to using PubMed. Besides its usefulness for RGD internal curation processes, the tool is available to all RGD users. Gene Annotator: The Gene Annotator or GA tool takes as input a list of gene symbols, RGD IDs, GenBank accession numbers, Ensembl identifiers, or a chromosomal region and retrieves gene orthologs, external database identifiers and ontology annotations for the corresponding genes in RGD. The data can be downloaded into an Excel spreadsheet or analyzed in the tool. The Annotation Distribution function displays a list of terms in each of seven categories with the percentage of genes from the input list with annotations to each term. The Comparison Heat Map function allows comparisons of annotations for genes in the input list across two ontologies or across two branches of the same ontology. Variant Visualizer: Variant Visualizer (VV) is a viewing and analysis tool for rat strain specific sequence polymorphisms. VV takes as input a list of gene symbols or a genomic region as defined by chromosome, start and stop positions or by two gene or marker symbols. The user must also select their strains of interest from a list of strains for which whole genome sequences exist and can set parameters for the variants in the result set. Output is a heatmap type display of variants. Additional information for individual variants can be viewed in a detail pane display. Multi Ontology Enrichment Tool (MOET): MOET is a web based ontology analysis tool used to identify terms from any or all of the ontologies used by RGD for gene curation (Disease, Pathway, Phenotype, GO, ChEBI) that are over represented in the annotations for those genes, or for orthologs in other species. It outputs a downloadable graph and a list of statistically overrepresented terms in the user’s list of genes using hypergeometric distribution. MOET also displays the corresponding Bonferroni correction and odds ratio on the results page. Gene Ortholog Location Finder (GOLF): GOLF is used to compare genes or positions within regions of interest across RGD species or assemblies. Results are displayed with the corresponding genes/positions in both species or on both assemblies in a side by side tabular view. Inputs and outputs to GOLF can be exported to other RGD tools for analysis or downloaded using the links on the GOLF results page. InterViewer: InterViewer is a protein protein interactive viewer that displays the appropriate information about types of interactions and links to associated genes pertaining to the user’s input. PhenoMiner: PhenoMiner combines phenotypic data from different rat strains, so researchers can use filters to find the quantitative phenotypic data they are looking for. OLGA Object List Generator & Analyzer: OLGA is a search engine designed to allow users to run multiple queries, generate a list of objects from each query and flexibly combine the results using Boolean specifications. OLGA takes as input either a list of object symbols or search parameters based on ontology annotations or position. The final list of genes, QTLs or strains can be downloaded or submitted to the GA Tool, the Variant Visualizer, the Genome Viewer or other RGD tools. Genome Viewer: The Genome Viewer (GViewer) tool provides users with complete genome views of genes, QTLs and mapped strains annotated to a function, biological process, cellular component, phenotype, disease, pathway, or chemical interaction. GViewer allows Boolean searches across multiple ontologies. Output is displayed against a karyotype of the rat genome. Overgo Probe Designer: Overgo probes are pairs of partially overlapping 22mer oligonucleotides derived from repeat masked genomic sequence and used as high specific activity probes for genome mapping. The Overgo Probe Designer tool takes as input a nucleotide sequence and outputs a list of optimized probe sequences containing the requisite 8 nucleotide overlap on their 3' ends. ACP Haplotyper: The ACP Haplotyper creates a visual haplotype that can be used to identify conserved and non conserved chromosomal regions between any of the 48 rat strains characterized as part of the ACP project. For the selected chromosome and between the selected strains, the tool compares the allele size data for microsatellite markers on the selected genetic or RH map.
Третьи лица адаптировали инструменты геномной обработки для использования с данными RGD
RGD предлагает несколько сторонних программных инструментов, адаптированных для использования на веб-сайте с использованием данных, хранящихся в базе данных RGD. К ним относятся:
JBrowse: JBrowse – это бесплатный, интерактивный инструмент для анализа данных, специфичный для баз данных. Программное обеспечение было разработано и в настоящее время поддерживается проектом базы данных модельных организмов. Через JBrowse можно получить доступ к генетическим и фенотипическим типам данных, включая основные наборы данных и данные о химическом взаимодействии генов, а также к их связи с геномной последовательностью. RatMine: RatMine – это версия программного обеспечения InterMine, ориентированная на крыс. Она позволяет пользователям извлекать и анализировать данные о крысах из различных баз данных, включая RGD, NCBI, UniProtKB и Ensembl, в одном месте, используя единый формат. Платформа InterMine была адаптирована для нескольких видов в других базах данных и разработана для обеспечения совместимости между экземплярами, чтобы пользователи могли выполнять запросы по разным видам непосредственно из интерфейса RatMine.
JBrowse: JBrowse is a free, interactive, and database specific data analysis tool. The software was created and is currently maintained by the Generic Model Organism Database project, Genetic and phenotypic data types, including fundamental datasets and gene chemical interaction data, and their relationship to the genomic sequence can be accessed through JBrowse. RatMine: RatMine is a rat centric version of the InterMine software. It enables users to mine and analyze rat data from diverse databases including RGD, NCBI, UniProtKB and Ensembl in a single location using a consistent format. The InterMine platform has been adapted for multiple species in other databases and is designed to be interoperable between instances so that users can query across species from the RatMine interface.
Пути
Ресурсы Pathway RGD включают в себя онтологию терминов Pathway (разработанную и поддерживаемую в RGD, охватывающую не только метаболические пути, но и пути, связанные с заболеваниями, лекарственными средствами, регуляцией и передачей сигналов), а также интерактивные диаграммы, отображающие компоненты и взаимодействия выбранных путей. На страницах диаграмм представлены описание, списки генов-участников пути и дополнительные элементы, таблицы аннотаций заболеваний, путей и фенотипов, сделанных для генов-участников пути, соответствующие ссылки и диаграмма онтологического пути. Также представлены наборы и сети путей, то есть объединения связанных путей, которые вносят вклад в более крупный процесс, такой как гомеостаз глюкозы или регуляция экспрессии генов, а также диаграммы физиологических путей, отображающие сети органов, тканей, клеток и молекулярных путей на уровне целого организма или системы.
Нокауты
До недавнего времени прямые, специфические геномные манипуляции на крысах были невозможны. Однако с развитием таких технологий, как нуклеазы цинковых пальцев и методов мутагенеза на основе CRISPR, ситуация изменилась. К группам, создающим нокаутные модели генов крыс и другие типы генетически модифицированных крыс, относится Центр человеческой и молекулярной генетики при MCW. RGD предоставляет доступ к информации о штаммах крыс, полученных в ходе этих исследований, через страницы, посвященные проекту PhysGen Knockout и Центру геномного редактирования крыс MCW (GERRC), ссылки на которые находятся в заголовках страниц RGD. Финансирование проектов PhysGenKO и GERRC осуществлялось Национальным институтом сердца, легких и крови (NHLBI). Основной целью обоих проектов являлось создание крыс с изменениями в одном или нескольких конкретных генах, связанных с задачами NHLBI. Гены предлагались исследователями, работающими с крысами. Предложения рассматривались внешним консультативным советом. В рамках проекта PhysGenKO многие из созданных крыс были фенотипированы с использованием стандартизированного протокола высокопроизводительного фенотипирования, а данные доступны в инструменте PhenoMiner, разработанном RGD.
Общественная пропаганда и образование
RGD взаимодействует с исследовательским сообществом, изучающим крыс, различными способами, включая email-форум, страницу новостей, страницу в Facebook, аккаунт в Twitter, а также регулярное участие и выступления на научных встречах и конференциях. Дополнительные образовательные мероприятия включают создание обучающих видеороликов, демонстрирующих как использование инструментов и данных RGD, так и более общие темы, такие как биомедицинские онтологии и биологическая (т.е. генная, QTL и штаммовая) номенклатура. Эти видеоролики доступны для просмотра на нескольких онлайн-видеохостингах, включая YouTube.
Финансирование
RGD финансируется грантом R01HL64541 от Национального института сердца, легких и крови (NHLBI) от имени Национальных институтов здоровья (NIH). Главным исследователем по гранту является Энн Э. Квитек, доктор философии, которая сменила Мэри Э. Шимояму, доктор философии, на этой руководящей должности в марте 2020 года. Мелинда Р. Дуинелл, доктор философии, является соисследователем.
Сборка нового генома
Новая сборка генома крысы, mRatBN7.2, была получена в рамках проекта «Древо жизни Дарвина» в Институте Велкома Сэнгера и принята Консорциумом геномных ссылок. mRatBN7.2 получена из самца крысы BN/NHsdMcwi, являющегося прямым потомком ранее секвенированной самки крысы BN. Новый эталонный геном крысы BN был создан с использованием различных технологий, включая длинные прочтения PacBio, прочтения 10X Genomics с линковкой, карты Bionano и Arima Hi-C. Его контигуозность сопоставима с эталонными сборками геномов человека и мыши. Он доступен в GenBank NCBI и RefSeq и в ближайшем будущем станет основной сборкой в RGD.