Live · free to use
Stop borrowing European allele frequencies for Malaysian patients.
A free allele-frequency reference built on 9,580 Malaysian people and 756,022 markers, with frequencies broken down by ancestry group. Type a marker, a gene or a condition. Aggregate counts only, and no personal data of any kind. Looking up a marker by its rsID needs no account.
Why this exists
For 70,000 markers, one “Malaysian frequency” describes nobody.
differ by 20 percentage points or more in allele frequency between Chinese, Malay and Indian Malaysians, and 21,537 of those differ by more than 30 points. For any of them a single pooled national figure is meaningless, and the US or European reference value a researcher would otherwise reach for is simply the wrong number for most patients in front of them.
That is roughly one marker in eleven of everything we can compare, counted across all 740,809 markers that carry an ancestry breakdown in this release. It is the reason to use a local denominator rather than a global one. Here is the difference on one well-understood marker:
1 in 2
1 in 19
1 in 95
If you work on a Malaysian cohort you have almost certainly reached for a global or largely European allele frequency because there was no local number to cite. For most markers that is a rough approximation. For pharmacogenomics it can be badly wrong, because several of the variants that matter most in clinical dosing sit at very different frequencies across Malay, Chinese and Indian ancestries than they do in European reference sets.
Look it up
Search by marker, gene or condition.
An rsID goes straight to the marker, and always will, with no account. A gene name returns every published marker in or near it, and anything else searches the 1,071 markers we have curated a clinical meaning for. Those broader searches show the first two matches; a free account opens the full list.
Try rs671, rs4149056, CYP2C19, MTHFR, drug response
The cohort
Who these numbers come from.
Ancestry groups in this release
Ancestry groups are inferred from genetic data by principal component analysis over roughly 17,000 markers, followed by clustering, without using any labels. Self-reported ethnicity is deliberately not collected. The grouping is measured; the group names are provisional, pending confirmation against reference populations.
Why the numbers differ
Three denominators appear on this page and they mean different things. The cohort overall is 12,000 or more participants and growing. This release reports frequencies from the 9,580 genotyped on a genome-wide array. And every individual marker carries its own number of people, because the arrays changed over the collection period, so a marker on a newer array was never measured in people run on an older one.
A further 1,567 participants were run on a targeted PCR panel instead. They are deliberately excluded from every frequency here: that panel measures a few hundred pre-selected markers rather than the genome, and pooling it would distort those markers' denominators.
Sex composition of this release
Self-reported at kit activation and matched to the array cohort by barcode, so no name or date of birth is involved. Matched for 8,961 of the 9,580 participants in this release, or 93.5%. Mismatches against genetic sex are almost always a data-entry error at activation rather than a biological finding.
Age at sample collection
This is how old each participant was when their sample was collected, not how old they are now. Computed for the 8,898 participants, 92.9% of this release, who have both a usable date of birth and a collection date. That denominator is not interchangeable with the 9,580 behind the frequencies. Age is derived from a restricted record that never enters the research archive and was never written to disk in producing these counts.
Headline participant figures are rounded down. Roughly 3,000 further community samples are in the processing queue and join the published counts as they clear quality control.
Provenance. This cohort comes entirely from Prima Nexus's own testing activity and from community participation, collected continuously since April 2019. It is not derived from, and carries no connection to, any client research project or any government or public-sector programme. Nothing generated under a client or public-sector engagement is used here, and nothing here is offered on any third party's behalf.
About the data
Everything here is measured, and you can check all of it.
A dataset you cannot interrogate is a dataset you should not cite. These are the figures a reviewer would ask for, stated before being asked.
Platform
All genotyping is performed on Illumina bead arrays by our sequencing partners in Singapore, Hong Kong, Japan and South Korea: GSA v3 with Malaysian custom content, GSA v2/v3, and the Asian Screening Array.
The array for each batch is established from the number of markers in the delivered genotype file, not from the text of the accompanying report, because a marker count cannot be mislabelled and the reports are inconsistent.
Measured quality
98.48% median call rate. 9,249 of 9,580 individuals sit at or above a 95% call rate; the 331 below it are flagged for review before use rather than quietly included.
Where the laboratory's own figure was not retained, the call rate was recomputed directly from that person's genotype file, and the record says which.
Chain of custody
One folder per laboratory delivery, 302 of them, each carrying its order number, date, platform, the evidence for that platform, per-individual quality, and a SHA-256 for every file in it.
Raw scanner output is retained wherever it was supplied, so genotypes can be re-called from the original instrument measurement if a manifest is ever revised.
Integrity
All 28,171 archive files were fingerprinted with SHA-256. Re-running the verification re-reads every file and reports anything that differs, is missing, or is new.
That is the difference between saying the data is unaltered and demonstrating it.
Ancestry labels
The grouping is measured; the names are provisional. Groups come from principal component analysis over roughly 17,000 markers followed by clustering, with no labels used. The names attached to them are inferred from two anchor markers.
Confirming them means projecting reference populations into the same space, and no Malay or Austronesian population exists in 1000 Genomes, so a separate reference is needed. Sound for deciding which markers merit attention. Not sufficient on its own for publication.
Two Chinese subgroups
The split is genuine, not an artefact. Both subgroups carry the East Asian anchor markers equally, and the division is not explained by which array was used or which year the sample was collected. It is consistent with dialect-group structure within Malaysian Chinese.
The page combines them into one Chinese figure by default and shows the two underneath, because either presentation on its own would mislead.
Privacy
The research archive holds no personal identifiers. Individuals appear only as barcodes, and that is verified after every change rather than assumed. Records that do identify people are held in a separate restricted area excluded from every research output.
Nothing derived from fewer than 20 individuals is published anywhere, following the practice of the NIH All of Us programme. No individual-level data exists in the files behind this page.
Genome build
Positions are given in GRCh38 by default and labelled as such, with the GRCh37 coordinate shown alongside.
The same marker has different coordinates in the two builds, and mixing them is a common and silent source of wrong results.
What you get
Built to be cited, not just browsed.
- A permanent address per marker. Append
?marker=rs671to this page's URL and it opens on that record, so a methods section or supplementary table can point straight at it. - Genotype counts and allele frequency with the exact number of people that marker was measured in, so you can compute your own confidence intervals rather than trusting ours.
- Carrier rates by ancestry group on 98% of published markers, each with its own denominator, so a comparison is checkable rather than asserted.
- The whole table as a download, below, for work that needs more than a lookup.
- Honest quality flags. 645,680 markers are withheld rather than published with a number we would not defend, and a marker nobody in the cohort carries is shown as exactly that, because a zero is a citable result too.
Privacy and governance
Counts, never people.
The browser contains no personal information of any kind. No names, no dates of birth, no individual results, nothing that can be traced to a participant. It reports how many people carry a variant, the way a census reports how many people live in a state without naming them.
Nothing based on fewer than 20 people is ever displayed. Markers below that threshold are withheld from the published data entirely rather than shown with a suppressed count, and the page never computes a smaller subset of its own. That is the same small cell rule used by major national genomic programmes, so the standard is an accepted one rather than something we invented.
Collaboration
The datasets that matter most in Malaysia sit inside institutions. We can help you get to the right table.
A good local cohort is usually held by a university, a hospital or a national research institution, under their own governance and their own ethics approval. We do not hold those datasets and we do not speak for the people who do. What we can do is convene the right group around a question and support the science once a collaboration is agreed.
Tell us the question
Send us the research question and the population you need, not a data request. We keep a running picture of who in the local ecosystem works on what.
We group the interest
Where several groups are circling the same question, we bring them together, because a consortium is far more likely to clear a data access committee than a lone request.
We make the introduction
Where we know the institution holding a relevant cohort, we introduce you directly and step back. The conversation, and the decision, belong to them.
We run the platform work
Once a collaboration is agreed, we do what we are actually for: genotyping and assays, bioinformatics, sample logistics, ethics and grant documentation, and project coordination.
What we are not offering. We do not hold, broker, resell or grant access to any dataset we do not own, and nothing on this page should be read as an offer of access to a third party's data. Where we have processed samples for another organisation, we acted as a laboratory only and claim no rights over the resulting data. Access to any institutional cohort is decided by the institution that holds it and by its ethics committee, never by us.
Shape what comes next
The browser is free and open. The roadmap is decided by the people using it.
A free Prima Nexus account gets you the things an account is actually for: nominate the markers and genes your work depends on so they are curated next, hear when a refresh lands and what changed, and get the ancestry-label confirmation work as it completes. It takes a minute and costs nothing.
Already have an account? Nominate markers from the portal.