Skip to main content

Curation

Curation process​

Association data indexed in the website is extracted from the scientific literature (PubMed indexed publications). Detailed eligibility criteria can be found on the accompanying page. Full genomewide summary statistics & accompanying metadata may also be submitted directly by authors (see Summary statistics methods), either before or after journal publication. This page focuses on the literature-extracted data. The GWAS Catalog does not generate any of its own data.

Extracted information includes publication information, study cohort information such as cohort size, country of recruitment and subject ancestry, and SNP-disease association information including SNP identifier (i.e. RSID), p-value, gene and risk allele. Each study is also assigned a trait that best represents the phenotype under investigation. When multiple traits are analysed in the same publication multiple entries are created, or (for older publications) individual SNPs are annotated with their specific traits. Each trait is annotated with one or more ontology terms which are used to summarise, query and visualise the data in the Catalog’s interfaces.

Data extraction and curation for the GWAS Catalog is an expert activity; each step is performed by scientists supported by a web-based tracking and data entry system which allows multiple curators to search, annotate, verify and publish the Catalog data. Papers that qualify for inclusion in the Catalog are identified through a weekly literature search.

First all data, including association information for SNPs, traits, and study and sample metadata, are extracted by a curator. In the case of author submissions, the curator performs QC of study and sample metadata with respect to the published manuscript, trait annotation, and extraction of any lead associations reported in the paper. Once data extraction is complete, an automated pipeline performs validation and mapping of the extracted data (see the Quality control and SNP mapping section below for more details), before release of the data for inclusion in the Catalog.

The current extraction methods are described below. Pilots to investigate extending Catalog scope and alternative methods of data acquisition are described on our pilot projects page. For a full list of publications by the GWAS Catalog team, and other useful background reading, please visit our Resources page

Data extraction​

GWAS Catalog data are manually extracted from the literature (either the main text or supporting information) by expert scientists, or submitted directly by authors and QC’d against the published manuscript.

The GWAS Catalog contains at least one entry (study) for each publication. However when a publication includes multiple GWAS analyses, for example on different traits or distinct sample cohorts, these are split into multiple entries (studies). In cases where publications containing multiple GWAS analyses are not split (typically for older publications), individual SNP associations are annotated with the specific trait, sample population or analysis method as "p-value text".

Information on the following study-level fields is extracted. Capitalised titles correspond to column headings in tab delimited file:

  • DATE ADDED TO CATALOG: Date a study is published in the catalog

  • PUBMEDID: PubMed identification number

  • FIRST AUTHOR: Last name and initials of first author

  • DATE: Publication date, (online (epub) date if available)

  • JOURNAL: Abbreviated journal name

  • LINK: PubMed URL

  • STUDY: Title of paper

  • DISEASE/TRAIT: Description of disease/trait examined in the study

  • INITIAL SAMPLE DESCRIPTION: Sample size and ancestry description for stage 1 of GWAS (summing across multiple Stage 1 populations, if applicable)

  • REPLICATION SAMPLE DESCRIPTION: Sample size and ancestry description for subsequent replication(s) (summing across multiple populations, if applicable)

  • GENOTYPING TECHNOLOGY: Genome-wide genotyping array, exome-wide genotyping array, targeted genotyping array (with an optional field for additional array information, for example ImmunoChip or MetaboChip).

  • PLATFORM (SNPS PASSING QC): Manufacturer of genotyping array used in GWAS stage; Number of SNPs passing quality control metrics and included in the GWAS analysis; also includes notation of “imputed” if imputation is used and "pooled" to denote studies of pooled DNA.

For each identified SNP, we extract:

  • REPORTED GENE: Gene(s) reported by the author; "intergenic" is used to denote a reported intergenic location (or lack of gene if it appeared that gene information was sought); “NR” is used to denote that no gene location information was reported.

  • STRONGEST SNP-RISK ALLELE: SNP(s) most strongly associated with trait + risk/effect allele (? for unknown risk allele). May also refer to a haplotype.

  • SNPS: Strongest SNP; If a haplotype is reported above, this field may include more than one rs number (multiple SNPs comprising the haplotype). Multiple SNPs may also be included if proxy SNPs are reported.

  • RISK ALLELE FREQUENCY: Reported risk/effect allele frequency associated with strongest SNP in controls (if not available among all controls, among the control group with the largest sample size). If the associated locus is a haplotype the haplotype frequency will be extracted.

  • P-VALUE: Reported p-value for strongest SNP risk allele. Note that p-values are rounded to 1 significant digit (for example, a published p-value of 4.8 x 10-7 is rounded to 5 x 10-7).

  • P-VALUE (TEXT): Information describing context of p-value (e.g. females, smokers).

  • OR or BETA: Reported odds ratio or beta-coefficient associated with strongest SNP risk allele. Note that, for papers curated prior to January 2021, if an OR <1 is reported this was inverted, along with the reported allele, so that all ORs included in the Catalog were >1. From January 2021 onwards, the OR and risk/effect allele are extracted exactly as reported in the paper, without inversion. Appropriate unit and increase/decrease are included for beta coefficients.

  • 95% CI (TEXT): Reported 95% confidence interval associated with strongest SNP risk allele, along with unit in the case of beta-coefficients. If 95% CIs are not published, we estimate these using the standard error, where available.

Ancestry data extraction​

Sample ancestry information is available in two distinct forms; a free text sample description and structured ancestry and recruitment information. The free text descriptions of the initial and replication stages of the GWAS provide summary descriptions of the samples analysed in each stage, based on the language used in the paper and including population descriptors used by the authors. The structured information, including an "ancestry label", is designed to represent data using controlled terms, enabling searching, visualisation and integration.

Full details of the methods and rationale used to represent population data in the GWAS Catalog can be found in our publication by Morales et al. and on our Population Descriptors page.

Quality control and SNP mapping​

An automated pipeline adds additional SNP specific information associated with the rsID extracted. This information includes the SNP’s base pair and cytogenetic location(s) in the current human genome reference assembly, mapped genes, mapped gene’s distance and positioning, and SNP function. This information is then used for queries of the search interface and in the production of the diagram. The pipeline also performs checks for consistency and missing information, such as SNP identifiers, existence of SNPs in dbSNP, validation of gene names and confirmation that the reported SNP and gene are in the same chromosomal region. This information is retrieved using the Ensembl API and the source of the data is both Ensembl and NCBI.

Additional information added by this pipeline. Capitalised titles correspond to column headings in tab delimited file:

  • REGION: Cytogenetic region associated with rs number.

  • CHR_ID: Chromosome number associated with rs number.

  • CHR_POS: Chromosomal position, in base pairs, associated with rs number (dbSNP Build 144, Genome Assembly GRCh38.p5, NCBI).

  • MAPPED GENE(S): Gene(s) mapped to the strongest SNP. If the SNP is located within a gene, that gene is listed, with multiple overlapping genes separated by “, ”. If the SNP is intergenic, the upstream and downstream genes are listed, separated by “ - ”.

  • UPSTREAM_GENE_ID: Entrez Gene ID for nearest upstream gene to rs number, if not within gene.

  • DOWNSTREAM_GENE_ID: Entrez Gene ID for nearest downstream gene to rs number, if not within gene.

  • SNP_GENE_IDS: Entrez Gene ID, if rs number within gene; multiple IDs denote overlapping genes.

  • UPSTREAM_GENE_DISTANCE: Distance in kb for nearest upstream gene to rs number, if not within gene.

  • DOWNSTREAM_GENE_DISTANCE: Distance in kb for nearest downstream gene to rs number, if not within gene.

  • MERGED: Denotes whether the SNP has been merged into a subsequent rs record (0 = no; 1 = yes).

  • SNP_ID_CURRENT: Current rs number (will differ from strongest SNP when merged = 1).

  • CONTEXT: SNP functional class.

  • INTERGENIC: Denotes whether SNP is in intergenic region (0 = no; 1 = yes).

Additional guidelines for data extraction and interpretation​

  • Missing or not applicable fields are denoted as follows: ?, allele not reported; NS, not significant (no associations at p<1.0 x 10-5 identified); NR, not reported.

  • Where multiple genetic models are available, effect sizes (ORs or beta-coefficients) are prioritized as follows: 1) genotypic model, per-allele estimate; 2) genotypic model, heterozygote estimate, 3) allelic model, allelic estimate.

  • If more than one SNP within a gene, or within a genomic region of 100kb upstream and downstream, meets the above extraction criteria, we report one SNP, unless there was evidence for an independent association.

  • Associations attributed to a combination of one or more genetic variants are denoted as such in the “Strongest SNP-Risk Allele” (e.g."3-SNP haplotype 1"). If available, rs numbers for SNPs comprising the haplotype are included in the “SNPs” field so that they are indexed and searchable using the SNP search features.

  • If the p-value, OR, and 95% CI fields are not available for the combined population, we extract estimates from the population group with the largest sample size.