Skip to main content

Submitting summary statistics plus metadata

For publications that are not yet included in the GWAS Catalog, or for pre-publication submissions, we ask you to submit metadata in addition to the summary statistics files.

Instructions are provided below. When you have completed the submission form, return to the main submission instructions.

WARNING: When filling in the metadata submission template, do not change the structure of the file. Do not rename or delete sheets, and do not alter hidden properties. Only enter your study details in the provided fields. Any structural edits may cause the submission to fail validation.

Data for you to enter​

There are 2 tabs in the submission form for you to complete:

  1. Studies

  2. Samples

In each tab, mandatory columns are highlighted in orange. Grey columns are optional. However, we encourage you to submit as much information as you can.

1. Study tab​

Overview​

In the Study tab, please add one row for each separate GWAS analysis (study) in the submission. In particular, please make sure that there is one row for each set of summary statistics (i.e. each full set of p-values) that you have uploaded. For example, if you performed GWAS analyses for 100 different metabolite measurements, there should be 100 rows - one for each metabolite.

If you performed multiple GWA studies but are only sharing summary statistics for a subset of the studies, you can enter metadata rows without summary statistics. However you must submit summary statistics for at least one study (including filename, md5sum, genome assembly).

Most studies submitted to the GWAS Catalog are for SNP-based analyses. If your submission also includes gene-based or CNV-based studies, please represent these as additional rows in the Study tab. Do not include multiple variant types (i.e. SNPs, genes, CNVs) in the same summary statistics file. For further assistance with complex submissions, please contact gwas-subs@ebi.ac.uk. And for details on validation criteria for CNV and gene-based files, please see https://www.ebi.ac.uk/gwas/apps/beyond-snps/.

Column instructions​

There are 24 columns for you to fill in:

warning

Do not upload a separate README file to Globus. Copy the contents of your README into the Readme text column of the submission form. The contents are stored as author notes in the metadata YAML. If you enter only README.txt, the author notes will contain only README.txt.

HeaderDescriptionMandatoryValidationExample
Study tagEach genome-wide association study in the submission must have a unique free-text label. You can use any string of characters that will help you identify each individual GWAS.yesFree textWHR_unadj
Genotyping technologyMethod(s) used to genotype variants in the discovery stage. Separate multiple methods by pipes "|".yesMust match one of the following options:; Genome-wide genotyping array; Targeted genotyping array; Exome genotyping array; Whole genome sequencing; Exome-wide sequencingGenome-wide genotyping array
Array manufacturerManufacturer of the genotyping array used for the discovery stage. Separate multiple manufacturers by pipes "|".noMust match one of the following options:; Illumina; Affymetrix; PerlegenIllumina|Affymetrix
Array informationAdditional information about the genotyping array. For example, for targeted arrays, please provide the specific type of array.noFree textImmunochip
Analysis softwareSoftware and version used for the association analysis.yes if any p-values in sumstats file = 0Free textPLINK 1.9
ImputationWere SNPs imputed for the discovery GWAS?yesMust match one of the following options:; Yes; NoYes
Imputation panelPanel used for imputationnoFree text1000 Genomes Phase 3
Imputation softwareImputation softwarenoFree textIMPUTE
Variant countThe number of variants analysed in the discovery stage (after QC)yesAn integer525000
Statistical modelA brief description of the statistical model used to determine association significance. Important to distinguish studies that would otherwise appear identical (e.g. the same trait analysed using additive, dominant and recessive models).noFree textadditive model
Study descriptionAdditional information about the studynoFree text…​
Adjusted covariatesAny covariates the GWAS is adjusted for. Multiple values can be listed separated by '|'.noFree textage|sex
Reported traitThe trait under investigation. Please describe the trait concisely but with enough detail to be clear to a non-specialist. Avoid use of abbreviations; if these are necessary, please define them or their source in the readme file.yesFree textReticulocyte count
Background traitAny background trait(s) shared by all individuals in the GWAS (e.g. in both cases and controls)noFree textNicotine dependence
Summary statistics fileThe name of the summary statistics file uploaded via Globus. Summary statistics must be submitted for at least one study. Enter "NR" for any additional studies without summary statistics.yesA valid filenameexample.tsv
md5 sumThe md5 checksum of the summary statistics file. Must be entered for all studies with summary statistics. Enter "NR" for any studies without summary statistics. See how to calculate checksums here.yesA valid md5 checksum (32-digit hexadecimal number)49ea8cf53801c7f1e2f11336fb8a29c8
ReadmeThe readme text that accompanies your analysis. Please copy the text into this cell, rather than uploading a separate readme file. If the same readme applies to all studies in the submission, please copy the text into each row. Leave blank for any studies without summary statistics.noA standard readme fileSee readme instructions here.
Summary statistics assemblyGenome assembly for the summary statistics. Must be entered for every row that includes summary statistics. Enter "NR" for any additional studies without summary statistics.yesMust match one of the following options:; GRCh38; GRCh37; NCBI36; NCBI35; NCBI34GRCh38
Neg Log10 p-valuesEnter yes if the summary statistics p-values are given in the negative log10 form.noMust match one of the following options:; Yes; Noyes
MAF lower limitLowest possible allele frequency given in summary statistics *nonumeric0.0001
Cohort(s)List of cohort(s) represented in the discovery sample, separated by pipes "|". Enter only if the specific named cohorts are used in the analysis.noFree textUKBB|FINRISK
Cohort specific referenceList of cohort specific identifier(s) issued to this research study, separated by pipes "|". For example, an ANID issued by UK Biobank.noFree textANID45956
SexTo indicate a sex-stratified analysis, enter the sex of participants as M or F. For non-sex-stratified analyses enter "combined", or NR if unknown.noFree textcombined
Coordinate systemCoordinate system used for the summary statistics: 1-based or 0-based (More info).yesFree text1-based

* Effect allele frequency is a mandatory field. However, where privacy concerns might otherwise be a barrier to sharing the data (e.g. in small cohorts with sensitive phenotypes), a cutoff may be specified so that allele frequencies below that cutoff are rounded-up to mask their true values. For example, 0.01 here indicates the lowest possible value for the minor allele frequency in this file is 0.01, and anything below this threshold has been rounded up to 0.01 (noting that the minor allele is not necessarily the effect allele). Since masking allele frequency limits the downstream utility of the data, please submit full EAF data where possible.

2. Sample tab​

Overview​

In the Sample tab, enter information about the samples included in each GWAS. Please create a new row for each GWAS and for each for each group of individuals assigned an ancestry category label within each GWAS. For multi-ancestry GWAS, please create a new row for each included ancestry category label (study tags can be reused multiple times on this sheet). See the Column Instructions below for the list of standardised ancestry categories used by the GWAS Catalog. For more information on why this is necessary, refer to our documentation.

For example, for 2 studies on different traits, analysed in 2 ancestry categories, create 2 x 2 = 4 rows:

Study tagStageNumber of individuals…​Ancestry category…​
s1_LDLdiscovery500…​East Asian…​
s1_LDLdiscovery500…​Sub-Saharan African…​
s2_HDLdiscovery500…​East Asian…​
s2_HDLdiscovery500…​Sub-Saharan African…​

Samples that contributed to the genome-wide analysis reported in your summary statistics should be annotated as “discovery” samples in the Stage column.

Your study design may have also included replication samples that were not analysed genome-wide, and therefore do not directly relate to your submitted summary statistics. However, information about these samples will eventually be included in the curated GWAS Catalog entry for your studies, so providing details at this stage will be of great help to our curators. You can add replication samples as additional rows, again separated by ancestry category, with “replication” in the Stage column.

For example, for a single study in 2 ancestry groups, with 2 stages (discovery and replication), create 2 x 2 = 4 rows:

Study tagStageNumber of individuals…​Ancestry category…​
s1_LDLdiscovery500…​East Asian…​
s1_LDLdiscovery500…​Sub-Saharan African…​
s1_LDLreplication200…​East Asian…​
s1_LDLreplication200…​Sub-Saharan African…​

Column Instructions​

There are 12 columns for you to fill in:

HeaderDescriptionMandatoryValidationExample
Study tagA unique free-text label for each genome-wide association study in the submission. This should match the study tag that you have provided in the “study” tab. This will allow the sample information to be linked to the correct study. You must provide at least one sample row for each study.yesFree textWHR_unadj
StageStage of the experimental designyesMust match one of the following options:; discovery; replicationdiscovery
Number of individualsNumber of individuals in this groupyesAn integer2000
Case control studyIs this a case control study?no (default is false)Must match one of the following options:; Yes; Noyes
Number of casesNumber of cases in this groupnoAn integer1000
Number of controlsNumber of controls in this groupnoAn integer1000
Sample descriptionAdditional information required for the interpretation of results, e.g. sex (males/females), age (adults/children), ordinal variables, or multiple traits analysed together ("or" traits).; Please do not enter ancestry information in this column (see other columns below).noFree text1000 males, 1000 females; 700 severe cases, 700 moderate cases, 600 mild cases; 1200 major depression cases, 800 bipolar disorder cases
Ancestry categoryAn ancestry category label that is appropriate for the sample. For more information about each category label, see Table 1, Morales et al., 2018.; You should create a new row for each ancestry category label. Providing group-specific sample numbers is important for reusability of the data. However, if separate sample numbers are unavailable for each group, you may enter multiple labels in the same row, separated by pipes "|".; If it is not possible to assign an ancestry category label to a group of samples, enter 'NR' (Not Reported).yesMust match one of the following options:; Aboriginal Australian; African American or Afro-Caribbean; African unspecified; Asian unspecified; Central Asian; Circumpolar peoples; East Asian; European; Greater Middle Eastern (Middle Eastern, North African or Persian); Hispanic or Latin American; Native American; NR; Oceanian; Other; Other admixed ancestry; South Asian; South East Asian; Sub-Saharan AfricanEast Asian
AncestryThe most detailed available population descriptor(s) for the sample. Separate multiple descriptors by pipes "|".noFree textHan Chinese
Founder/Genetically isolated population descriptionFor founder or genetically isolated population, provide description. If multiple founder/genetically isolated populations are included for the same ancestry category, separate using pipes "|".noFree textKorculan(founder/genetic isolate)|Vis(founder/genetic isolate)
Ancestry methodName the method used to determine sample ancestry. For consistency, we recommend you choose between the terms “self-reported” or “genetically determined” where appropriate, but other text is permissible if these do not apply. Multiple values can be listed separated by pipes "|".noFree textself-reported|genetically determined
Country of recruitmentList of country/countries where samples were recruited, separated by pipes "|".yesPlease copy country names exactly as written in this list.Japan|China

Additional information​

Some cells in Excel may display a "Number Stored as Text" error. Please ignore this, as it will not affect the template validation.