Биоинформатика в иммунологии: анализ геномных данных, вычислительные методы для понимания иммунной системы и разработки лекарств. Исследования иммунитета.
Сравнивайте с английским: нажмите на абзац — оригинал откроется в окне. Кнопка EN под абзацем показывает его прямо в тексте.
Содержание
Введение
Биоинформатические подходы к иммунологии
В академической среде вычислительная иммунология – это область науки, объединяющая высокопроизводительные геномные и биоинформатические подходы в иммунологии. Главная цель этой области заключается в преобразовании иммунологических данных в вычислительные задачи, решении этих задач с помощью математических и вычислительных методов и последующей интерпретации полученных результатов с иммунологической точки зрения.
Bioinformatics approaches to immunology
In academia, computational immunology is a field of science that encompasses high throughput genomic and bioinformatics approaches to immunology. The field's main aim is to convert immunological data into computational problems, solve these problems using mathematical and computational approaches and then convert these results into immunologically meaningful interpretations.
Введение
Иммунная система – сложная система человеческого организма, и её понимание является одной из самых трудных задач в биологии. Исследования в области иммунологии важны для понимания механизмов, лежащих в основе защиты организма, а также для разработки лекарств от иммунологических заболеваний и поддержания здоровья. Недавние достижения в геномных и протеомных технологиях кардинально изменили иммунологические исследования. Секвенирование геномов человека и других модельных организмов привело к экспоненциальному росту объемов данных, релевантных для иммунологии, и одновременно в научной литературе и клинических записях накапливаются огромные массивы функциональных и клинических данных. Последние успехи в биоинформатике или вычислительной биологии помогли осмыслить и организовать эти крупномасштабные данные, что дало начало новой области, известной как вычислительная иммунология или иммуноинформатика. Вычислительная иммунология является ветвью биоинформатики и основывается на схожих концепциях и инструментах, таких как выравнивание последовательностей и инструменты для предсказания структуры белка. Иммуномика – это дисциплина, аналогичная геномике и протеомике. Это наука, которая объединяет иммунологию с компьютерными науками, математикой, химией и биохимией для масштабного анализа функций иммунной системы. Она направлена на изучение сложных белок-белковых взаимодействий и сетей, что позволяет лучше понять иммунные ответы и их роль в нормальном, патологическом и восстановительном состояниях. Вычислительная иммунология является частью иммуномики, специализирующейся на анализе крупномасштабных экспериментальных данных.
The immune system is a complex system of the human body and understanding it is one of the most challenging topics in biology. Immunology research is important for understanding the mechanisms underlying the defense of human body and to develop drugs for immunological diseases and maintain health. Recent findings in genomic and proteomic technologies have transformed the immunology research drastically. Sequencing of the human and other model organism genomes has produced increasingly large volumes of data relevant to immunology research and at the same time huge amounts of functional and clinical data are being reported in the scientific literature and stored in clinical records. Recent advances in bioinformatics or computational biology were helpful to understand and organize these large scale data and gave rise to new area that is called Computational immunology or immunoinformatics. Computational immunology is a branch of bioinformatics and it is based on similar concepts and tools, such as sequence alignment and protein structure prediction tools. Immunomics is a discipline like genomics and proteomics. It is a science, which specifically combines immunology with computer science, mathematics, chemistry, and biochemistry for large scale analysis of immune system functions. It aims to study the complex protein–protein interactions and networks and allows a better understanding of immune responses and their role during normal, diseased and reconstitution states. Computational immunology is a part of immunomics, which is focused on analyzing large scale experimental data.
История
Компьютерная иммунология зародилась более 90 лет назад с теоретического моделирования эпидемиологии малярии. В то время акцент делался на использовании математических методов для изучения распространения заболеваний. С тех пор эта область расширилась, включив в себя все остальные аспекты процессов, происходящих в иммунной системе, и связанных с ними заболеваний.
Computational immunology began over 90 years ago with the theoretic modeling of malaria epidemiology. At that time, the emphasis was on the use of mathematics to guide the study of disease transmission. Since then, the field has expanded to cover all other aspects of immune system processes and diseases.
Иммунологическая база данных
После недавних достижений в области секвенирования и протеомики наблюдается многократное увеличение объема генерируемых молекулярных и иммунологических данных. Эти данные настолько разнообразны, что могут быть классифицированы в различных базах данных в зависимости от их использования в исследованиях. На сегодняшний день в коллекции баз данных Nucleic Acids Research (NAR) зарегистрировано 31 различная иммунологическая база данных, представленных в следующей таблице, а также некоторые другие базы данных, связанные с иммунитетом. Информация в таблице взята из описаний баз данных в коллекции NAR Database Collection.
After the recent advances in sequencing and proteomics technology, there have been many fold increase in generation of molecular and immunological data. The data are so diverse that they can be categorized in different databases according to their use in the research. Until now there are total 31 different immunological databases noted in the Nucleic Acids Research (NAR) Database Collection, which are given in the following table, together with some more immune related databases. The information given in the table is taken from the database descriptions in NAR Database Collection. Database Description ALPSbase Autoimmune lymphoproliferative syndrome database AntigenDB Sequence, structure, and other data on pathogen antigens. AntiJen Quantitative binding data for peptides and proteins of immunological interest. BCIpep This database stores information of all experimentally determined B cell epitopes of antigenic proteins. This is a curated database where detailed information about the epitopes are collected and compiled from published literature and existing databases. It covers a wide range of pathogenic organisms like virus, bacteria, protozoa and fungi. Each entry in database provides full information about a B cell epitope that includes amino acid sequences, source of the antigenic protein, immunogenicity, model organism and antibody generation/neutralization test. dbMHC dbMHC provides access to HLA sequences, tools to support genetic testing of HLA loci, HLA allele and haplotype frequencies of over 90 populations worldwide, as well as clinical datasets on hematopoietic stem cell transplantation, and insulin dependent diabetes mellitus (IDDM), Rheumatoid Arthritis (RA), Narcolepsy and Spondyloarthropathy. For more information go to this link http://www. oxfordjournals. org/nar/database/summary/604|| DIGIT Database of ImmunoGlobulin sequences and Integrated Tools. FIMM FIMM is an integrated database of functional molecular immunology that focuses on the T cell response to disease specific antigens. FIMM provides fully referenced information integrated with data retrieval and sequence analysis tools on HLA, peptides, T cell epitopes, antigens, diseases and constitutes one backbone of future computational immunology research. Antigen protein data have been enriched with more than 27,000 sequences derived from the non redundant SwissProt TREMBL TREMBL NEW (SPTR) database of antigens similar or related FIMM antigens across various species to facilitate a comprehensive analysis of conserved or variable T cell epitopes. GPX Macrophage Expression Atlas The GPX Macrophage Expression Atlas (GPX MEA) is an online resource for expression based studies of a range of macrophage cell types following treatment with pathogens and immune modulators. GPX Macrophage Expression Atlas (GPX MEA) follows the MIAME standard and includes an objective quality score with each experiment. It places special emphasis on rigorously capturing the experimental design and enables the statistical analysis of expression data from different micro array experiments. This is the first example of a focussed macrophage gene expression database that allows efficient identification of transcriptional patterns, which provide novel insights into biology of this cell system. HaptenDB It is a comprehensive database of hapten molecules. This is a curated database where information is collected and compiled from published literature and web resources. Presently database has more than 1700 entries where each entry provides comprehensive detail about a hapten molecule that includes: i) nature of the hapten; ii) methods of anti hapten antibody production; iii) information about carrier protein; iv) coupling method; v) assay method (used for characterization) and vi) specificities of antibodies. The Haptendb covers wide array of haptens ranging from antibiotics of biomedical importance to pesticides. This database will be very useful for studying the serological reactions and production of antibodies. HPTAA HPTAA is a database of potential tumor associated antigens that uses expression data from various expression platforms, including carefully chosen publicly available microarray expression data, GEO SAGE data and Unigene expression data. IEDB 3D Structural data within the Immune Epitope Database. IL2Rgbase X linked severe combined immunodeficiency mutations. IMGT IMGT is an integrated knowledge resource specialized in IG, TR, MHC, IG superfamily, MHC superfamily and related proteins of the immune system of human and other vertebrate species. IMGTW comprises 6 databases, 15 on line tools for sequence, gene and 3D structure analysis, and more than 10,000 pages of resources Web. Data standardization, based on IMGT ONTOLOGY, has been approved by WHO/IUIS. IMGT GENE DB IMGT/GENE DB is the IMGT® comprehensive genome database for immunoglobulins (IG) and T cell receptors (TR) genes from human and mouse, and, in development, from other vertebrate species (e. g. rat). IMGT/GENE DB is part of IMGT®, the international ImMunoGeneTics information system®, the high quality integrated knowledge resource specialized in IG, TR, major histocompatibility complex (MHC) of human and other vertebrate species, and related proteins of the immune system (RPI) that belong to the immunoglobulin superfamily (IgSF) and to the MHC superfamily (MhcSF). IMGT/HLA There are currently over 1600 officially recognised HLA alleles and these sequences are made available to the scientific community through the IMGT/HLA database. In 1998, the IMGT/HLA database was publicly released. Since this time, the database has grown and is the primary source of information for the study of sequences of the human major histocompatibility complex. The initial release of the database contained allele reports, alignment tools, submission tools as well as detailed descriptions of the source cells. The database is updated quarterly with all the new and confirmatory sequences submitted to the WHO Nomenclature Committee and on average an additional 75 new and confirmatory sequences are included in each quarterly release. The IMGT/HLA database provides a centralized resource for everybody interested, either centrally or peripherally, in the HLA system. IMGT/LIGM DB IMGT/LIGM DB is the IMGT® comprehensive database of immunoglobulin (IG) and T cell receptor (TR) nucleotide sequences, from human and other vertebrate species, with translation for fully annotated sequences, created in 1989 by LIGM http://www. imgt. org/textes/IMGTinformation/LIGM. html), Montpellier, France, on the Web since July 1995. IMGT/LIGM DB is the first and the largest database of IMGT®, the international ImMunoGeneTics information system®, the high quality integrated knowledge resource specialized in IG, TR, major histocompatibility complex (MHC) of human and other vertebrate species, and related proteins of the immune system (RPI) that belong to the immunoglobulin superfamily (IgSF) and to the MHC superfamily (MhcSF). IMGT/LIGM DB sequence data are identified by the EMBL/GenBank/DDBJ accession number. The unique source of data for IMGT/LIGM DB is EMBL which shares data with GenBank and DDBJ. Interferon Stimulated Gene Database Interferons (IFN) are a family of multifunctional cytokines that activate transcription of a subset of genes. The gene products induced by IFN are responsible for the antiviral, antiproliferative and immunomodulatory properties of this cytokine. In order to obtain a more comprehensive understanding of the genes regulated by IFNs we have used different microarray formats to identify over 400 interferon stimulated genes (ISG). To facilitate the dissemination of this data we have compiled a database comprising the ISGs assigned into functional categories. The database is fully searchable and contains links to sequence and Unigene information. The database and the array data are accessible via the World Wide Web at (http://www. lerner. ccf. org/labs/williams/ ). We intend to add published ISG sequences and those discovered by further transcript profiling to the database to eventually compile a complete list of ISGs. IPD ESTDAB The Immuno Polymorphism Database (IPD) is a set of specialist databases related to the study of polymorphic genes in the immune system. IPD ESTDAB is a database of immunologically characterised melanoma cell lines. The database works in conjunction with the European Searchable Tumour Cell Line Database (ESTDAB) cell bank, which is housed in TÜbingen, Germany and provides immunologically characterised tumour cells. IPD HPA Human Platelet Antigens Human platelet antigens are alloantigens expressed only on platelets, specifically on platelet membrane glycoproteins. These platelet specific antigens are immunogenic and can result in pathological reactions to transfusion therapy. The IPD HPA section contains nomenclature information and additional background material about Human platelet antigen. The different genes in the HPA system have not been sequenced to the same level as some of the other projects and so currently only single nucleotide polymorphisms (SNP) are used to determine alleles. This information is presented in a grid of SNP for each gene The IPD and HPA nomenclature committee hope to expand this to provide full sequence alignments when possible. MHCPEP This database contains list of MHC binding peptides. MPID T2 MPID T2 (https://web. archive. org/web/20120902154345/http://biolinfo. org/mpid t2/) is a highly curated database for sequence structure function information on MHC peptide interactions. It contains all structures of major histocompatibility complex proteins (MHC) containing bound peptides, with emphasis on the structural characterization of these complexes. Database entries have been grouped into fully referenced redundant and non redundant categories. The MHC peptide interactions have been presented in terms of a set of sequence and structural parameters representative of molecular recognition. MPID will facilitate the development of algorithms to predict whether a query peptide sequence will bind to a specific MHC allele. MPID data has been sorted primarily on the basis of MHC Class, followed by organism (MHC source), next by allele type and finally by the length of peptide in the binding groove (peptide residues within 5 Å of the MHC). Data on inter molecular hydrogen bonds, gap volume and gap index available in MPID are pre computed and the interface area due to complex formation is calculated based on accessible surface area calculations. The available MHC peptide databases have addressed sequence information as well as binding (or the lack thereof) of peptide sequences. MUGEN Mouse Database Murine models of immune processes and immunological diseases. Protegen Protective antigen database and analysis system. SuperHapten SuperHapten is a manually curated hapten database integrating information from literature and web resources. The current version of the database compiles 2D/3D structures, physicochemical properties and references for about 7,500 haptens and 25,000 synonyms. The commercial availability is documented for about 6,300 haptens and 450 related antibodies, enabling experimental approaches on cross reactivity. The haptens are classified regarding their origin: pesticides, herbicides, insecticides, drugs, natural compounds, etc. Queries allow identification of haptens and associated antibodies according to functional class, carrier protein, chemical scaffold, composition or structural similarity. The Immune Epitope Database (IEDB) The Immune Epitope Database (IEDB, www. iedb. org), provides a catalog of experimentally characterized B and T cell epitopes, as well as data on MHC binding and MHC ligand elution experiments. The database represents the molecular structures recognized by adaptive immune receptors and the experimental contexts in which these molecules were determined to be immune epitopes. Epitopes recognized in humans, non human primates, rodents, pigs, cats and all other tested species are included. Both positive and negative experimental results are captured. Over the course of four years, the data from 180,978 experiments were curated manually from the literature, covering about 99% of all publicly available information on peptide epitopes mapped in infectious agents (excluding HIV) and 93% of those mapped in allergens. TmaDB To analyse TMA output a relational database (known as TmaDB) has been developed to collate all aspects of information relating to TMAs. These data include the TMA construction protocol, experimental protocol and results from the various immunocytological and histochemical staining experiments including the scanned images for each of the TMA cores. Furthermore, the database contains pathological information associated with each of the specimens on the TMA slide, the location of the various TMAs and the individual specimen blocks (from which cores were taken) in the laboratory and their current status. TmaDB has been designed to incorporate and extend many of the published common data elements and the XML format for TMA experiments and is therefore compatible with the TMA data exchange specifications developed by the Association for Pathology Informatics community. VBASE2 VBASE2 is an integrative database of germ line V genes from the immunoglobulin loci of human and mouse. It presents V gene sequences from the EMBL database and Ensembl together with the corresponding links to the source data. The VBASE2 dataset is generated in an automatic process based on a BLAST search of V genes against EMBL and the Ensembl dataset. The BLAST hits are evaluated with the DNAPLOT program, which allows immunoglobulin sequence alignment and comparison, RSS recognition and analysis of the V(D)J rearrangements. As a result of the BLAST hit evaluation, the VBASE2 entries are classified into 3 different classes: class 1 holds sequences for which a genomic reference and a rearranged sequence is known. Class 2 contains sequences, which have not been found in a rearrangement, thus lacking evidence of functionality. Class 3 contains sequences which have been found in different V(D)J rearrangements but lack a genomic reference. All VBASE2 sequences are compared with the datasets from the VBASE , IMGT and KABAT databases (latest published versions), and the respective references are provided in each VBASE2 sequence entry. The VBASE2 database can be accessed by either a text based query form or by a sequence alignment with the DNAPLOT program. A DAS server shows the VBASE2 dataset within the Ensembl Genome Browser and links to the database. Epitome Epitome is a database of all known antigenic residues and the antibodies that interact with them, including a detailed description of the residues involved in the interaction and their sequence/structure environments. Each entry in the database describes one interaction between a residue on an antigenic protein and a residue on an antibody chain. Every interaction is described using the following parameters: PDB identifier, antigen chain ID PDB position of the antigenic residue, type of antigenic residue and its sequence environment, antigen residue secondary structure state, antigen residue solvent accessibility, antibody chain ID, type of antibody chain (heavy or light), CDR number, PDB position of the antibody residue, and type of antibody residue and its sequence environment. Additionally, interactions can be visualized using an interface to Jmol. ImmGen The Immunological Genome consortium database includes expression profiles for more than 250 mouse immune cell types, and several data browsers to study the dataset. ImmPort ImmPort, the Immunology Database and Analysis Portal, is a comprehensive, highly curated and standardized database of more than 400 publicly shared clinical and research studies funded by NIAID/DAIT (National Institutes of Allergy and Infectious Disease/Division of Allergy, Immunology and Transplantation). Shared data includes study metadata, over thirty types of mechanistic assays (e. g. flow cytometry, mass cytometry, ELISA, HAI, MBAA, etc ) as well as clinical assessments, lab tests and adverse events. ImmPort is a recommended data repository for Nature Scientific Data – Cytometry & Immunology and PLOS ONE. ImmPort has also been awarded the CoreTrust Seal as a trustworthy data repository. All shared data is available for download. Online resources for allergy information are also available on http://www. allergen. org. Such data is valuable for investigation of cross reactivity between known allergens and analysis of potential allergenicity in proteins. The Structural Database of Allergen Proteins (SDAP) stores information of allergenic proteins. The Food Allergy Research and Resource Program (FARRP) Protein Allergen Online Database contains sequences of known and putative allergens derived from scientific literature and public databases. Allergome emphasizes the annotation of allergens that result in an IgE mediated disease.
Описание базы данных | База данных
------- | --------
ALPSbase | База данных аутоиммунного лимфопролиферативного синдрома
AntigenDB | Последовательности, структура и другие данные о патогенных антигенах
AntiJen | Количественные данные о связывании пептидов и белков, представляющих иммунологический интерес
BCIpep | Эта база данных хранит информацию обо всех экспериментально определенных B-клеточных эпитопах антигенных белков. Это курируемая база данных, в которой подробная информация об эпитопах собирается и компилируется из опубликованной литературы и существующих баз данных. Она охватывает широкий спектр патогенных организмов, таких как вирусы, бактерии, простейшие и грибки. Каждая запись в базе данных содержит полную информацию о B-клеточном эпитопе, включая последовательности аминокислот, источник антигенного белка, иммуногенность, модельный организм и тест на генерацию/нейтрализацию антител.
dbMHC | dbMHC предоставляет доступ к последовательностям HLA, инструментам для поддержки генетического тестирования локусов HLA, частотам аллелей HLA и гаплотипов более чем 90 популяций по всему миру, а также к клиническим наборам данных о трансплантации гемопоэтических стволовых клеток, инсулинозависимом сахарном диабете (IDDM), ревматоидном артрите (RA), нарколепсии и спондилоартропатии. Дополнительную информацию можно найти по ссылке http://www.oxfordjournals.org/nar/database/summary/604
DIGIT | База данных последовательностей иммуноглобулинов и интегрированных инструментов
FIMM | FIMM – это интегрированная база данных функциональной молекулярной иммунологии, ориентированная на T-клеточный ответ на специфические для заболевания антигены. FIMM предоставляет полностью процитированную информацию, интегрированную с инструментами поиска данных и анализа последовательностей для HLA, пептидов, T-клеточных эпитопов, антигенов, заболеваний и является одной из основ будущих исследований в области вычислительной иммунологии. Данные о белках антигенов обогащены более чем 27 000 последовательностями из не избыточной базы данных SwissProt TREMBL TREMBL NEW (SPTR) антигенов, аналогичных или связанных с антигенами FIMM у различных видов, чтобы облегчить всесторонний анализ консервативных или переменных T-клеточных эпитопов.
GPX | Атлас экспрессии макрофагов GPX (GPX MEA) – это онлайн-ресурс для исследований экспрессии различных типов клеток макрофагов после обработки патогенами и иммуномодуляторами. GPX MEA соответствует стандарту MIAME и включает в себя объективную оценку качества для каждого эксперимента. Особое внимание уделяется строгому учету дизайна эксперимента и позволяет проводить статистический анализ данных экспрессии из различных экспериментов с использованием микроматриц. Это первый пример базы данных экспрессии генов макрофагов, который позволяет эффективно идентифицировать паттерны транскрипции, предоставляя новые сведения о биологии этой клеточной системы.
HaptenDB | Это всеобъемлющая база данных молекул гаптенов. Это курируемая база данных, в которой информация собирается и компилируется из опубликованной литературы и веб-ресурсов. В настоящее время база данных содержит более 1700 записей, каждая из которых предоставляет подробную информацию о молекуле гаптена, включая: i) природу гаптена; ii) методы производства антигаптенных антител; iii) информацию о белке-носителе; iv) метод конъюгации; v) метод анализа (используемый для характеристики) и vi) специфичность антител. Haptendb охватывает широкий спектр гаптенов, от антибиотиков, имеющих биомедицинское значение, до пестицидов. Эта база данных будет очень полезна для изучения серологических реакций и производства антител.
HPTAA | HPTAA – это база данных потенциальных опухолевых антигенов, использующая данные экспрессии из различных платформ, включая тщательно отобранные общедоступные данные экспрессии микроматриц, данные GEO SAGE и данные экспрессии Unigene.
IEDB 3D | Структурные данные в базе данных иммунных эпитопов
IL2Rgbase | Мутации, вызывающие тяжелый комбинированный иммунодефицит, сцепленный с X-хромосомой
IMGT | IMGT – это интегрированный ресурс знаний, специализирующийся на генах IG, TR, MHC, суперсемействе IG, суперсемействе MHC и связанных белках иммунной системы человека и других позвоночных. IMGTW включает 6 баз данных, 15 онлайн-инструментов для анализа последовательностей, генов и 3D-структур, а также более 10 000 страниц ресурсов в Интернете. Стандартизация данных, основанная на IMGT ONTOLOGY, была одобрена ВОЗ/IUIS.
IMGT/GENE DB | IMGT/GENE DB – это исчерпывающая геномная база данных IMGT® для генов иммуноглобулинов (IG) и Т-клеточных рецепторов (TR) человека и мыши, а также, в стадии разработки, других позвоночных (например, крысы). IMGT/GENE DB является частью IMGT®, международной информационной системы ImMunoGeneTics®, высококачественного интегрированного ресурса знаний, специализирующегося на IG, TR, комплексе главного гистосовместимости (MHC) человека и других позвоночных, а также связанных белках иммунной системы (RPI), принадлежащих к суперсемейству иммуноглобулинов (IgSF) и суперсемейству MHC (MhcSF).
IMGT/HLA | В настоящее время существует более 1600 официально признанных аллелей HLA, и эти последовательности доступны научному сообществу через базу данных IMGT/HLA. В 1998 году база данных IMGT/HLA была опубликована. С тех пор база данных расширилась и является основным источником информации для изучения последовательностей комплекса главного гистосовместимости человека. Первоначальный выпуск базы данных содержал отчеты об аллелях, инструменты выравнивания, инструменты отправки, а также подробные описания исходных клеток. База данных обновляется ежеквартально всеми новыми и подтверждающими последовательностями, представленными в Комитет номенклатуры ВОЗ, и в среднем в каждом ежеквартальном выпуске добавляется дополнительно 75 новых и подтверждающих последовательностей. База данных IMGT/HLA предоставляет централизованный ресурс для всех заинтересованных, центрально или периферически, в системе HLA.
IMGT/LIGM DB | IMGT/LIGM DB – это исчерпывающая база данных IMGT® нуклеотидных последовательностей иммуноглобулинов (IG) и Т-клеточных рецепторов (TR) человека и других позвоночных, с трансляцией для полностью аннотированных последовательностей, созданная в 1989 году LIGM (http://www.imgt.org/textes/IMGTinformation/LIGM.html), Монпелье, Франция, в Интернете с июля 1995 года. IMGT/LIGM DB является первой и крупнейшей базой данных IMGT®, международной информационной системы ImMunoGeneTics®, высококачественного интегрированного ресурса знаний, специализирующегося на IG, TR, комплексе главного гистосовместимости (MHC) человека и других позвоночных, а также связанных белках иммунной системы (RPI), принадлежащих к суперсемейству иммуноглобулинов (IgSF) и суперсемейству MHC (MhcSF). Данные последовательностей IMGT/LIGM DB идентифицируются по номеру доступа EMBL/GenBank/DDBJ. Единственным источником данных для IMGT/LIGM DB является EMBL, который обменивается данными с GenBank и DDBJ.
Interferon Stimulated Gene Database | Интерфероны (IFN) – это семейство мультифункциональных цитокинов, которые активируют транскрипцию подмножества генов. Продукты генов, индуцированные IFN, отвечают за противовирусные, антипролиферативные и иммуномодулирующие свойства этого цитокина. Чтобы получить более полное представление о генах, регулируемых IFN, мы использовали различные форматы микроматриц для идентификации более 400 генов, стимулированных интерфероном (ISG). Чтобы облегчить распространение этих данных, мы составили базу данных, состоящую из ISG, отнесенных к функциональным категориям. База данных полностью доступна для поиска и содержит ссылки на информацию о последовательностях и Unigene. База данных и данные микроматриц доступны в World Wide Web по адресу (http://www.lerner.ccf...
After the recent advances in sequencing and proteomics technology, there have been many fold increase in generation of molecular and immunological data. The data are so diverse that they can be categorized in different databases according to their use in the research. Until now there are total 31 different immunological databases noted in the Nucleic Acids Research (NAR) Database Collection, which are given in the following table, together with some more immune related databases. The information given in the table is taken from the database descriptions in NAR Database Collection. Database Description ALPSbase Autoimmune lymphoproliferative syndrome database AntigenDB Sequence, structure, and other data on pathogen antigens. AntiJen Quantitative binding data for peptides and proteins of immunological interest. BCIpep This database stores information of all experimentally determined B cell epitopes of antigenic proteins. This is a curated database where detailed information about the epitopes are collected and compiled from published literature and existing databases. It covers a wide range of pathogenic organisms like virus, bacteria, protozoa and fungi. Each entry in database provides full information about a B cell epitope that includes amino acid sequences, source of the antigenic protein, immunogenicity, model organism and antibody generation/neutralization test. dbMHC dbMHC provides access to HLA sequences, tools to support genetic testing of HLA loci, HLA allele and haplotype frequencies of over 90 populations worldwide, as well as clinical datasets on hematopoietic stem cell transplantation, and insulin dependent diabetes mellitus (IDDM), Rheumatoid Arthritis (RA), Narcolepsy and Spondyloarthropathy. For more information go to this link http://www. oxfordjournals. org/nar/database/summary/604|| DIGIT Database of ImmunoGlobulin sequences and Integrated Tools. FIMM FIMM is an integrated database of functional molecular immunology that focuses on the T cell response to disease specific antigens. FIMM provides fully referenced information integrated with data retrieval and sequence analysis tools on HLA, peptides, T cell epitopes, antigens, diseases and constitutes one backbone of future computational immunology research. Antigen protein data have been enriched with more than 27,000 sequences derived from the non redundant SwissProt TREMBL TREMBL NEW (SPTR) database of antigens similar or related FIMM antigens across various species to facilitate a comprehensive analysis of conserved or variable T cell epitopes. GPX Macrophage Expression Atlas The GPX Macrophage Expression Atlas (GPX MEA) is an online resource for expression based studies of a range of macrophage cell types following treatment with pathogens and immune modulators. GPX Macrophage Expression Atlas (GPX MEA) follows the MIAME standard and includes an objective quality score with each experiment. It places special emphasis on rigorously capturing the experimental design and enables the statistical analysis of expression data from different micro array experiments. This is the first example of a focussed macrophage gene expression database that allows efficient identification of transcriptional patterns, which provide novel insights into biology of this cell system. HaptenDB It is a comprehensive database of hapten molecules. This is a curated database where information is collected and compiled from published literature and web resources. Presently database has more than 1700 entries where each entry provides comprehensive detail about a hapten molecule that includes: i) nature of the hapten; ii) methods of anti hapten antibody production; iii) information about carrier protein; iv) coupling method; v) assay method (used for characterization) and vi) specificities of antibodies. The Haptendb covers wide array of haptens ranging from antibiotics of biomedical importance to pesticides. This database will be very useful for studying the serological reactions and production of antibodies. HPTAA HPTAA is a database of potential tumor associated antigens that uses expression data from various expression platforms, including carefully chosen publicly available microarray expression data, GEO SAGE data and Unigene expression data. IEDB 3D Structural data within the Immune Epitope Database. IL2Rgbase X linked severe combined immunodeficiency mutations. IMGT IMGT is an integrated knowledge resource specialized in IG, TR, MHC, IG superfamily, MHC superfamily and related proteins of the immune system of human and other vertebrate species. IMGTW comprises 6 databases, 15 on line tools for sequence, gene and 3D structure analysis, and more than 10,000 pages of resources Web. Data standardization, based on IMGT ONTOLOGY, has been approved by WHO/IUIS. IMGT GENE DB IMGT/GENE DB is the IMGT® comprehensive genome database for immunoglobulins (IG) and T cell receptors (TR) genes from human and mouse, and, in development, from other vertebrate species (e. g. rat). IMGT/GENE DB is part of IMGT®, the international ImMunoGeneTics information system®, the high quality integrated knowledge resource specialized in IG, TR, major histocompatibility complex (MHC) of human and other vertebrate species, and related proteins of the immune system (RPI) that belong to the immunoglobulin superfamily (IgSF) and to the MHC superfamily (MhcSF). IMGT/HLA There are currently over 1600 officially recognised HLA alleles and these sequences are made available to the scientific community through the IMGT/HLA database. In 1998, the IMGT/HLA database was publicly released. Since this time, the database has grown and is the primary source of information for the study of sequences of the human major histocompatibility complex. The initial release of the database contained allele reports, alignment tools, submission tools as well as detailed descriptions of the source cells. The database is updated quarterly with all the new and confirmatory sequences submitted to the WHO Nomenclature Committee and on average an additional 75 new and confirmatory sequences are included in each quarterly release. The IMGT/HLA database provides a centralized resource for everybody interested, either centrally or peripherally, in the HLA system. IMGT/LIGM DB IMGT/LIGM DB is the IMGT® comprehensive database of immunoglobulin (IG) and T cell receptor (TR) nucleotide sequences, from human and other vertebrate species, with translation for fully annotated sequences, created in 1989 by LIGM http://www. imgt. org/textes/IMGTinformation/LIGM. html), Montpellier, France, on the Web since July 1995. IMGT/LIGM DB is the first and the largest database of IMGT®, the international ImMunoGeneTics information system®, the high quality integrated knowledge resource specialized in IG, TR, major histocompatibility complex (MHC) of human and other vertebrate species, and related proteins of the immune system (RPI) that belong to the immunoglobulin superfamily (IgSF) and to the MHC superfamily (MhcSF). IMGT/LIGM DB sequence data are identified by the EMBL/GenBank/DDBJ accession number. The unique source of data for IMGT/LIGM DB is EMBL which shares data with GenBank and DDBJ. Interferon Stimulated Gene Database Interferons (IFN) are a family of multifunctional cytokines that activate transcription of a subset of genes. The gene products induced by IFN are responsible for the antiviral, antiproliferative and immunomodulatory properties of this cytokine. In order to obtain a more comprehensive understanding of the genes regulated by IFNs we have used different microarray formats to identify over 400 interferon stimulated genes (ISG). To facilitate the dissemination of this data we have compiled a database comprising the ISGs assigned into functional categories. The database is fully searchable and contains links to sequence and Unigene information. The database and the array data are accessible via the World Wide Web at (http://www. lerner. ccf. org/labs/williams/ ). We intend to add published ISG sequences and those discovered by further transcript profiling to the database to eventually compile a complete list of ISGs. IPD ESTDAB The Immuno Polymorphism Database (IPD) is a set of specialist databases related to the study of polymorphic genes in the immune system. IPD ESTDAB is a database of immunologically characterised melanoma cell lines. The database works in conjunction with the European Searchable Tumour Cell Line Database (ESTDAB) cell bank, which is housed in TÜbingen, Germany and provides immunologically characterised tumour cells. IPD HPA Human Platelet Antigens Human platelet antigens are alloantigens expressed only on platelets, specifically on platelet membrane glycoproteins. These platelet specific antigens are immunogenic and can result in pathological reactions to transfusion therapy. The IPD HPA section contains nomenclature information and additional background material about Human platelet antigen. The different genes in the HPA system have not been sequenced to the same level as some of the other projects and so currently only single nucleotide polymorphisms (SNP) are used to determine alleles. This information is presented in a grid of SNP for each gene The IPD and HPA nomenclature committee hope to expand this to provide full sequence alignments when possible. MHCPEP This database contains list of MHC binding peptides. MPID T2 MPID T2 (https://web. archive. org/web/20120902154345/http://biolinfo. org/mpid t2/) is a highly curated database for sequence structure function information on MHC peptide interactions. It contains all structures of major histocompatibility complex proteins (MHC) containing bound peptides, with emphasis on the structural characterization of these complexes. Database entries have been grouped into fully referenced redundant and non redundant categories. The MHC peptide interactions have been presented in terms of a set of sequence and structural parameters representative of molecular recognition. MPID will facilitate the development of algorithms to predict whether a query peptide sequence will bind to a specific MHC allele. MPID data has been sorted primarily on the basis of MHC Class, followed by organism (MHC source), next by allele type and finally by the length of peptide in the binding groove (peptide residues within 5 Å of the MHC). Data on inter molecular hydrogen bonds, gap volume and gap index available in MPID are pre computed and the interface area due to complex formation is calculated based on accessible surface area calculations. The available MHC peptide databases have addressed sequence information as well as binding (or the lack thereof) of peptide sequences. MUGEN Mouse Database Murine models of immune processes and immunological diseases. Protegen Protective antigen database and analysis system. SuperHapten SuperHapten is a manually curated hapten database integrating information from literature and web resources. The current version of the database compiles 2D/3D structures, physicochemical properties and references for about 7,500 haptens and 25,000 synonyms. The commercial availability is documented for about 6,300 haptens and 450 related antibodies, enabling experimental approaches on cross reactivity. The haptens are classified regarding their origin: pesticides, herbicides, insecticides, drugs, natural compounds, etc. Queries allow identification of haptens and associated antibodies according to functional class, carrier protein, chemical scaffold, composition or structural similarity. The Immune Epitope Database (IEDB) The Immune Epitope Database (IEDB, www. iedb. org), provides a catalog of experimentally characterized B and T cell epitopes, as well as data on MHC binding and MHC ligand elution experiments. The database represents the molecular structures recognized by adaptive immune receptors and the experimental contexts in which these molecules were determined to be immune epitopes. Epitopes recognized in humans, non human primates, rodents, pigs, cats and all other tested species are included. Both positive and negative experimental results are captured. Over the course of four years, the data from 180,978 experiments were curated manually from the literature, covering about 99% of all publicly available information on peptide epitopes mapped in infectious agents (excluding HIV) and 93% of those mapped in allergens. TmaDB To analyse TMA output a relational database (known as TmaDB) has been developed to collate all aspects of information relating to TMAs. These data include the TMA construction protocol, experimental protocol and results from the various immunocytological and histochemical staining experiments including the scanned images for each of the TMA cores. Furthermore, the database contains pathological information associated with each of the specimens on the TMA slide, the location of the various TMAs and the individual specimen blocks (from which cores were taken) in the laboratory and their current status. TmaDB has been designed to incorporate and extend many of the published common data elements and the XML format for TMA experiments and is therefore compatible with the TMA data exchange specifications developed by the Association for Pathology Informatics community. VBASE2 VBASE2 is an integrative database of germ line V genes from the immunoglobulin loci of human and mouse. It presents V gene sequences from the EMBL database and Ensembl together with the corresponding links to the source data. The VBASE2 dataset is generated in an automatic process based on a BLAST search of V genes against EMBL and the Ensembl dataset. The BLAST hits are evaluated with the DNAPLOT program, which allows immunoglobulin sequence alignment and comparison, RSS recognition and analysis of the V(D)J rearrangements. As a result of the BLAST hit evaluation, the VBASE2 entries are classified into 3 different classes: class 1 holds sequences for which a genomic reference and a rearranged sequence is known. Class 2 contains sequences, which have not been found in a rearrangement, thus lacking evidence of functionality. Class 3 contains sequences which have been found in different V(D)J rearrangements but lack a genomic reference. All VBASE2 sequences are compared with the datasets from the VBASE , IMGT and KABAT databases (latest published versions), and the respective references are provided in each VBASE2 sequence entry. The VBASE2 database can be accessed by either a text based query form or by a sequence alignment with the DNAPLOT program. A DAS server shows the VBASE2 dataset within the Ensembl Genome Browser and links to the database. Epitome Epitome is a database of all known antigenic residues and the antibodies that interact with them, including a detailed description of the residues involved in the interaction and their sequence/structure environments. Each entry in the database describes one interaction between a residue on an antigenic protein and a residue on an antibody chain. Every interaction is described using the following parameters: PDB identifier, antigen chain ID PDB position of the antigenic residue, type of antigenic residue and its sequence environment, antigen residue secondary structure state, antigen residue solvent accessibility, antibody chain ID, type of antibody chain (heavy or light), CDR number, PDB position of the antibody residue, and type of antibody residue and its sequence environment. Additionally, interactions can be visualized using an interface to Jmol. ImmGen The Immunological Genome consortium database includes expression profiles for more than 250 mouse immune cell types, and several data browsers to study the dataset. ImmPort ImmPort, the Immunology Database and Analysis Portal, is a comprehensive, highly curated and standardized database of more than 400 publicly shared clinical and research studies funded by NIAID/DAIT (National Institutes of Allergy and Infectious Disease/Division of Allergy, Immunology and Transplantation). Shared data includes study metadata, over thirty types of mechanistic assays (e. g. flow cytometry, mass cytometry, ELISA, HAI, MBAA, etc ) as well as clinical assessments, lab tests and adverse events. ImmPort is a recommended data repository for Nature Scientific Data – Cytometry & Immunology and PLOS ONE. ImmPort has also been awarded the CoreTrust Seal as a trustworthy data repository. All shared data is available for download. Online resources for allergy information are also available on http://www. allergen. org. Such data is valuable for investigation of cross reactivity between known allergens and analysis of potential allergenicity in proteins. The Structural Database of Allergen Proteins (SDAP) stores information of allergenic proteins. The Food Allergy Research and Resource Program (FARRP) Protein Allergen Online Database contains sequences of known and putative allergens derived from scientific literature and public databases. Allergome emphasizes the annotation of allergens that result in an IgE mediated disease.
Инструменты
Различные вычислительные, математические и статистические методы доступны и описаны в литературе. Эти инструменты полезны для сбора, анализа и интерпретации иммунологических данных. К ним относятся текстовая добыча (text mining), управление информацией, анализ последовательностей, анализ молекулярных взаимодействий и математические модели, позволяющие проводить расширенное моделирование иммунной системы и иммунологических процессов. Предпринимаются попытки извлечения интересных и сложных закономерностей из неструктурированных текстовых документов в иммунологической области, например, категоризация информации о перекрестной реактивности аллергенов, использование программ BLAST и TreeView, а также специализированные инструменты иммуноинформатики, такие как EpiMatrix, IMGT/V QUEST для анализа последовательностей IG и TR, IMGT/Collier de Perles и IMGT/StructuralQuery для анализа структуры вариабельных доменов IG. Методы, основанные на сравнении последовательностей, разнообразны и применяются для анализа консервации последовательностей HLA, подтверждения происхождения последовательностей вируса иммунодефицита человека (ВИЧ) и построения гомологичных моделей для анализа резистентности полимеразы вируса гепатита B к ламивудину и эмтрицитабину. Также существуют вычислительные модели, ориентированные на белок-белковые взаимодействия и сети. Доступны инструменты для картирования Т- и В-клеточных эпитопов, прогнозирования участков расщепления протеасомой и прогнозирования пептидов, транспортируемых TAP. Экспериментальные данные имеют решающее значение для разработки и обоснования моделей, предназначенных для прогнозирования различных молекулярных мишеней. Инструменты вычислительной иммунологии представляют собой сочетание экспериментальных данных и математически разработанных вычислительных методов.
A variety of computational, mathematical and statistical methods are available and reported. These tools are helpful for collection, analysis, and interpretation of immunological data. They include text mining, information management, sequence analysis, analysis of molecular interactions, and mathematical models that enable advanced simulations of immune system and immunological processes. Attempts are being made for the extraction of interesting and complex patterns from non structured text documents in the immunological domain. Such as categorization of allergen cross reactivity information, BLAST, and TreeView, as well as specialized immunoinformatics tools, such as EpiMatrix, IMGT/V QUEST for IG and TR sequence analysis, IMGT/ Collier de Perles and IMGT/StructuralQuery for IG variable domain structure analysis. Methods that rely on sequence comparison are diverse and have been applied to analyze HLA sequence conservation, help verify the origins of human immunodeficiency virus (HIV) sequences, and construct homology models for the analysis of hepatitis B virus polymerase resistance to lamivudine and emtricitabine. There are also some computational models which focus on protein–protein interactions and networks. There are also tools which are used for T and B cell epitope mapping, proteasomal cleavage site prediction, and TAP– peptide prediction. The experimental data is very much important to design and justify the models to predict various molecular targets. Computational immunology tools is the game between experimental data and mathematically designed computational tools.
Аллергия
Аллергия, хотя и является важным разделом иммунологии, также существенно различается у разных людей, а иногда даже у генетически близких особей. Оценка аллергенного потенциала белка сосредоточена на трех основных аспектах: (i) иммуногенности; (ii) перекрестной реактивности; и (iii) клинических проявлениях. Иммуногенность обусловлена реакцией B-клеток, продуцирующих IgE-антитела, и/или T-клеток на конкретный аллерген. Поэтому исследования иммуногенности в основном направлены на выявление участков распознавания аллергенов B- и T-клетками. Трехмерные структурные свойства аллергенов определяют их аллергенность. Использование инструментов иммуноинформатики может быть полезным для прогнозирования аллергенности белков и будет играть все более важную роль в скрининге новых пищевых продуктов перед их широким внедрением для потребления человеком. Таким образом, прилагаются значительные усилия для создания надежных и всеобъемлющих баз данных по аллергии и их объединения с хорошо валидированными инструментами прогнозирования, чтобы обеспечить выявление потенциальных аллергенов в генетически модифицированных лекарствах и продуктах питания. Хотя разработки находятся на начальной стадии, Всемирная организация здравоохранения и Продовольственная и сельскохозяйственная организация ООН предложили руководства по оценке аллергенности генетически модифицированных продуктов питания. Согласно Codex Alimentarius, белок считается потенциально аллергенным, если он имеет идентичность в ≥6 последовательных аминокислот или ≥35% сходства последовательности в окне из 80 аминокислот с известным аллергеном. Несмотря на наличие этих правил, их внутренние ограничения становятся все более очевидными, и о случаях, не соответствующих правилам, сообщалось неоднократно.
Allergies, while a critical subject of immunology, also vary considerably among individuals and sometimes even among genetically similar individuals. The assessment of protein allergenic potential focuses on three main aspects: (i) immunogenicity; (ii) cross reactivity; and (iii) clinical symptoms. Immunogenicity is due to responses of an IgE antibody producing B cell and/or of a T cell to a particular allergen. Therefore, immunogenicity studies focus mainly on identifying recognition sites of B cells and T cells for allergens. The three dimensional structural properties of allergens control their allergenicity. The use of immunoinformatics tools can be useful to predict protein allergenicity and will become increasingly important in the screening of novel foods before their wide scale release for human use. Thus, there are major efforts under way to make reliable broad based allergy databases and combine these with well validated prediction tools in order to enable the identification of potential allergens in genetically modified drugs and foods. Though the developments are on primary stage, the World Health organization and Food and Agriculture Organization have proposed guidelines for evaluating allergenicity of genetically modified foods. According to the Codex alimentarius, a protein is potentially allergenic if it possesses an identity of ≥6 contiguous amino acids or ≥35% sequence similarity over an 80 amino acid window with a known allergen. Though there are rules, their inherent limitations have started to become apparent and exceptions to the rules have been well reported
Функции иммунной системы
Используя эту технологию, можно понять модель, лежащую в основе иммунной системы. Она была использована для моделирования Т-клеточной супрессии, миграции периферических лимфоцитов, памяти Т-клеток, толерантности, функции тимуса и сетей антител. Модели помогают предсказывать динамику токсичности патогенов и Т-клеточной памяти в ответ на различные стимулы. Существуют также несколько моделей, полезных для понимания природы специфичности в иммунной сети и иммуногенности. Например, это позволило изучить функциональную связь между транспортом пептидов TAP и представлением антигена HLA класса I. TAP – это трансмембранный белок, ответственный за транспорт антигенных пептидов в эндоплазматический ретикулум, где молекулы MHC класса I могут связываться с ними и представлять их Т-клеткам. Поскольку TAP не связывает все пептиды с одинаковой силой, аффинность связывания TAP может влиять на способность конкретного пептида получить доступ к пути MHC класса I. Искусственная нейронная сеть (ANN), компьютерная модель, была использована для изучения связывания пептидов с человеческим TAP и его взаимосвязи со связыванием MHC класса I. С помощью этого метода было обнаружено, что аффинность пептидов, связывающихся с HLA через TAP, различается в зависимости от соответствующего супертипа HLA. Это исследование может иметь важные последствия для разработки иммунотерапевтических препаратов и вакцин на основе пептидов. Это демонстрирует возможности моделирования для понимания сложных иммунных взаимодействий.
Using this technology it is possible to know the model behind immune system. It has been used to model T cell mediated suppression, peripheral lymphocyte migration, T cell memory, tolerance, thymic function, and antibody networks. Models are helpful to predicts dynamics of pathogen toxicity and T cell memory in response to different stimuli. There are also several models which are helpful in understanding the nature of specificity in immune network and immunogenicity. For example, it was useful to examine the functional relationship between TAP peptide transport and HLA class I antigen presentation. TAP is a transmembrane protein responsible for the transport of antigenic peptides into the endoplasmic reticulum, where MHC class I molecules can bind them and presented to T cells. As TAP does not bind all peptides equally, TAP binding affinity could influence the ability of a particular peptide to gain access to the MHC class I pathway. Artificial neural network (ANN), a computer model was used to study peptide binding to human TAP and its relationship with MHC class I binding. The affinity of HLA binding peptides for TAP was found to differ according to the HLA supertype concerned using this method. This research could have important implications for the design of peptide based immuno therapeutic drugs and vaccines. It shows the power of the modeling approach to understand complex immune interactions.
Информатика рака
Рак является результатом соматических мутаций, которые дают раковым клеткам селективное преимущество в росте. В последнее время крайне важно выявлять новые мутации. Методы геномики и протеомики используются во всем мире для идентификации мутаций, связанных с каждым конкретным видом рака, и их лечением. Вычислительные инструменты применяются для прогнозирования роста и поверхностных антигенов на раковых клетках. Имеются публикации, описывающие целенаправленный подход к оценке мутаций и риска развития рака. Алгоритм CanPredict использовался для определения степени сходства конкретного гена с известными генами, вызывающими рак. Иммунология рака приобрела такую значимость, что объём данных в этой области стремительно растёт. Сети взаимодействия белков предоставляют ценную информацию о tumorigenesis у человека. Раковые белки демонстрируют сетевую топологию, отличную от топологии нормальных белков в человеческой интерактоме. Иммуноинформатика оказалась полезной для повышения эффективности опухолевой вакцинации. Недавно были проведены новаторские исследования, посвященные анализу динамики иммунной системы организма в ответ на искусственно индуцированный иммунитет, вызванный стратегиями вакцинации. Другие инструменты моделирования используют предсказанные раковые пептиды для прогнозирования иммуноспецифических противораковых реакций, зависящих от конкретного HLA. Вероятно, эти ресурсы значительно расширятся в ближайшем будущем, и иммуноинформатика станет одной из ключевых областей развития в данной сфере.
Cancer is the result of somatic mutations which provide cancer cells with a selective growth advantage. Recently it has been very important to determine the novel mutations. Genomics and proteomics techniques are used worldwide to identify mutations related to each specific cancer and their treatments. Computational tools are used to predict growth and surface antigens on cancerous cells. There are publications explaining a targeted approach for assessing mutations and cancer risk. Algorithm CanPredict was used to indicate how closely a specific gene resembles known cancer causing genes. Cancer immunology has been given so much importance that the data related to it is growing rapidly. Protein–protein interaction networks provide valuable information on tumorigenesis in humans. Cancer proteins exhibit a network topology that is different from normal proteins in the human interactome. Immunoinformatics have been useful in increasing success of tumour vaccination. Recently, pioneering works have been conducted to analyse the host immune system dynamics in response to artificial immunity induced by vaccination strategies. Other simulation tools use predicted cancer peptides to forecast immune specific anticancer responses that is dependent on the specified HLA. These resources are likely to grow significantly in the near future and immunoinformatics will be a major growth area in this domain.