The calculation, the evidence, and the limits
Scientific transparency is part of the product.
A genomic score without provenance is just a number.
The two data paths
Genodex separates personal exploration from public contribution.
In the private path, a compatible raw DNA file is parsed on the user’s device. Supported phenotype calculations use the variants available in that local file. Importing and calculating stay in the private path until the person explicitly chooses publication.
In the public path, a person deliberately completes the publish-genome flow. A genomic archive and related record are uploaded and made available through Genodex’s public infrastructure. Public genome and phenotype pages present that contributed data alongside available context and provenance. Public genomic data is identifying; removing a display name or direct account identifier does not de-identify a genome.
What Genodex calculates
Genodex maps variants available in a genome to the variant definitions in a supported phenotype model. The calculation engine applies the model’s stored allele directions, coefficients, normalizations, and reference values to produce the available result fields. Depending on the model, these can include a score, z-score, percentile, normalized effect, curve position, or SNP coverage.
The deterministic Genodex engine calculates the scientific score. AI begins after calculation. Given the same supported input, model data, and software version, the engine follows the same calculation path. Reproducibility describes the calculation; scientific validity and population fit depend on the underlying evidence.
From variants to context
A calculated value is only the beginning of the presentation. Where records are available, Genodex connects the estimate to:
the phenotype and related traits;
the variants expected by the model and those found in the imported genome;
effect or alternate alleles and model direction;
connected genes and variant identifiers;
studies, DOI records, PubMed records, and dbSNP pages;
population or ancestry context supplied by the model or source project;
SNP coverage, uncertainty fields, and other model limitations; and
the project, organization, license, and recommended citation for public source data when these can be verified.
A linked paper establishes provenance for an association or source record. Clinical interpretation of an individual result requires separate clinical evidence and professional review.
Percentiles, curves, and uncertainty
A percentile locates a result within the reference distribution used by a model. Outcome probability and clinical diagnosis require different evidence. Environment, development, behavior, measurement, and variants outside the model remain part of the broader context around every result.
Coverage also matters. Raw DNA files differ by provider and genotyping platform, so a file may omit variants expected by a model. Missing variants, ambiguous identifiers, strand orientation, imputation assumptions, linkage disequilibrium, effect-size construction, sample size, and source-association quality shape the meaning and reliability of a result. Genodex displays the available coverage and evidence context. Missing evidence remains visible as missing evidence.
Population transfer and representation
GWAS findings reflect the people, recruitment criteria, measurements, and analytical choices in the source study. Many genomic datasets overrepresent particular ancestry groups. An effect estimated in one population may be weaker, stronger, absent, or differently calibrated in another.
Population labels are contextual summaries rather than fixed biological categories. Genodex presents the available population context so users can judge transfer limitations. Reference-group fit remains visible as one dimension of the evidence, never a measure of identity or worth.
Public sources and provenance
Genodex distinguishes the creator or source of an external dataset from Genodex’s role as provider of this educational presentation. A genome page identifies its source project when known. Verified official sources govern which licenses and recommended citations appear. Unknown fields remain omitted, and external open licenses appear only with supporting evidence.
Public contributions made directly through Genodex follow the publication terms presented in the app. Their public status carries re-identification risk and remains subject to third-party rights.
How AI is used
AI reports and Genome Chat form an optional explanation layer after calculation. For an AI request, Genodex assembles a constrained fact set from selected calculated values and connected scientific records, together with relevant genome metadata and the user’s request. Genome Chat can also include recent conversation context.
OpenAI translates that supplied context into explanatory language. Generated language can vary, omit a qualification, misread context, or contain an error. The displayed calculation, coverage, provenance, linked evidence, and source limitations remain the source of record for every explanation.
Updates and corrections
Catalog records can change when source projects update, a publication or identifier is corrected, provenance is verified, or Genodex changes a transformation or model. A recalculation made with a different source-data, model, or software version may produce a different result.
Where trustworthy timestamps exist, Genodex uses them to communicate data freshness. A timestamp records an update to a record or source; individual clinical review is a separate process. Questions or correction requests can be sent to joe@genodex.org.
Designed for genomic literacy
Genodex serves genomic literacy and evidence navigation through educational estimates, transparent calculations, and connected sources. Medical and other high-impact decisions belong to qualified professionals working with appropriate clinical information and validated methods.
A genotype is not a diagnosis. Your genome is a strong signal, never a verdict.