Prima Nexus / Malaysian Cohort Browser

Live · free to use

Stop borrowing European allele frequencies for Malaysian patients.

A free allele-frequency reference built on 9,580 Malaysian people and 756,022 markers, with frequencies broken down by ancestry group. Type a marker, a gene or a condition. Aggregate counts only, and no personal data of any kind. Looking up a marker by its rsID needs no account.

Free to use rsID lookup, no account Aggregate counts only Per marker denominators CSV download

Why this exists

For 70,000 markers, one “Malaysian frequency” describes nobody.

70,623markers

differ by 20 percentage points or more in allele frequency between Chinese, Malay and Indian Malaysians, and 21,537 of those differ by more than 30 points. For any of them a single pooled national figure is meaningless, and the US or European reference value a researcher would otherwise reach for is simply the wrong number for most patients in front of them.

That is roughly one marker in eleven of everything we can compare, counted across all 740,809 markers that carry an ancestry breakdown in this release. It is the reason to use a local denominator rather than a global one. Here is the difference on one well-understood marker:

ALDH2 rs671 · carriers of the flush variant, which also affects nitroglycerin response see the full record →
49.2%Chinese carry a copy
1 in 2
5.2%Malay
1 in 19
1.1%Indian
1 in 95

If you work on a Malaysian cohort you have almost certainly reached for a global or largely European allele frequency because there was no local number to cite. For most markers that is a rough approximation. For pharmacogenomics it can be badly wrong, because several of the variants that matter most in clinical dosing sit at very different frequencies across Malay, Chinese and Indian ancestries than they do in European reference sets.

Look it up

Search by marker, gene or condition.

An rsID goes straight to the marker, and always will, with no account. A gene name returns every published marker in or near it, and anything else searches the 1,071 markers we have curated a clinical meaning for. Those broader searches show the first two matches; a free account opens the full list.

Try rs671, rs4149056, CYP2C19, MTHFR, drug response

Reading these figures. The large percentage on each card is the share of people carrying at least one copy of the minor allele, derived from the genotype counts shown beneath it. The ancestry bars are the same quantity for each group, with the allele frequency and the raw counts given beneath each one. Every figure carries the number of people that specific marker was measured in, because the arrays changed over the years and a bare percentage is not a checkable claim.

The cohort

Who these numbers come from.

12,000+participants in the cohort overall
9,580genotyped on a genome-wide array and included in this release
756,022markers published
740,809markers with an ancestry breakdown, 98% of those published
+400new participants a month, and rising

Ancestry groups in this release

Ancestry groups are inferred from genetic data by principal component analysis over roughly 17,000 markers, followed by clustering, without using any labels. Self-reported ethnicity is deliberately not collected. The grouping is measured; the group names are provisional, pending confirmation against reference populations.

Why the numbers differ

Three denominators appear on this page and they mean different things. The cohort overall is 12,000 or more participants and growing. This release reports frequencies from the 9,580 genotyped on a genome-wide array. And every individual marker carries its own number of people, because the arrays changed over the collection period, so a marker on a newer array was never measured in people run on an older one.

A further 1,567 participants were run on a targeted PCR panel instead. They are deliberately excluded from every frequency here: that panel measures a few hundred pre-selected markers rather than the genome, and pooling it would distort those markers' denominators.

Sex composition of this release

Self-reported at kit activation and matched to the array cohort by barcode, so no name or date of birth is involved. Matched for 8,961 of the 9,580 participants in this release, or 93.5%. Mismatches against genetic sex are almost always a data-entry error at activation rather than a biological finding.

Age at sample collection

This is how old each participant was when their sample was collected, not how old they are now. Computed for the 8,898 participants, 92.9% of this release, who have both a usable date of birth and a collection date. That denominator is not interchangeable with the 9,580 behind the frequencies. Age is derived from a restricted record that never enters the research archive and was never written to disk in producing these counts.

Headline participant figures are rounded down. Roughly 3,000 further community samples are in the processing queue and join the published counts as they clear quality control.

Provenance. This cohort comes entirely from Prima Nexus's own testing activity and from community participation, collected continuously since April 2019. It is not derived from, and carries no connection to, any client research project or any government or public-sector programme. Nothing generated under a client or public-sector engagement is used here, and nothing here is offered on any third party's behalf.

About the data

Everything here is measured, and you can check all of it.

A dataset you cannot interrogate is a dataset you should not cite. These are the figures a reviewer would ask for, stated before being asked.

Platform

All genotyping is performed on Illumina bead arrays by our sequencing partners in Singapore, Hong Kong, Japan and South Korea: GSA v3 with Malaysian custom content, GSA v2/v3, and the Asian Screening Array.

The array for each batch is established from the number of markers in the delivered genotype file, not from the text of the accompanying report, because a marker count cannot be mislabelled and the reports are inconsistent.

Measured quality

98.48% median call rate. 9,249 of 9,580 individuals sit at or above a 95% call rate; the 331 below it are flagged for review before use rather than quietly included.

Where the laboratory's own figure was not retained, the call rate was recomputed directly from that person's genotype file, and the record says which.

Chain of custody

One folder per laboratory delivery, 302 of them, each carrying its order number, date, platform, the evidence for that platform, per-individual quality, and a SHA-256 for every file in it.

Raw scanner output is retained wherever it was supplied, so genotypes can be re-called from the original instrument measurement if a manifest is ever revised.

Integrity

All 28,171 archive files were fingerprinted with SHA-256. Re-running the verification re-reads every file and reports anything that differs, is missing, or is new.

That is the difference between saying the data is unaltered and demonstrating it.

Ancestry labels

The grouping is measured; the names are provisional. Groups come from principal component analysis over roughly 17,000 markers followed by clustering, with no labels used. The names attached to them are inferred from two anchor markers.

Confirming them means projecting reference populations into the same space, and no Malay or Austronesian population exists in 1000 Genomes, so a separate reference is needed. Sound for deciding which markers merit attention. Not sufficient on its own for publication.

Two Chinese subgroups

The split is genuine, not an artefact. Both subgroups carry the East Asian anchor markers equally, and the division is not explained by which array was used or which year the sample was collected. It is consistent with dialect-group structure within Malaysian Chinese.

The page combines them into one Chinese figure by default and shows the two underneath, because either presentation on its own would mislead.

Privacy

The research archive holds no personal identifiers. Individuals appear only as barcodes, and that is verified after every change rather than assumed. Records that do identify people are held in a separate restricted area excluded from every research output.

Nothing derived from fewer than 20 individuals is published anywhere, following the practice of the NIH All of Us programme. No individual-level data exists in the files behind this page.

Genome build

Positions are given in GRCh38 by default and labelled as such, with the GRCh37 coordinate shown alongside.

The same marker has different coordinates in the two builds, and mixing them is a common and silent source of wrong results.

What you get

Built to be cited, not just browsed.

  • A permanent address per marker. Append ?marker=rs671 to this page's URL and it opens on that record, so a methods section or supplementary table can point straight at it.
  • Genotype counts and allele frequency with the exact number of people that marker was measured in, so you can compute your own confidence intervals rather than trusting ours.
  • Carrier rates by ancestry group on 98% of published markers, each with its own denominator, so a comparison is checkable rather than asserted.
  • The whole table as a download, below, for work that needs more than a lookup.
  • Honest quality flags. 645,680 markers are withheld rather than published with a number we would not defend, and a marker nobody in the cohort carries is shown as exactly that, because a zero is a citable result too.
Download the full table. Every published marker with its counts, frequencies, gene, both genome-build positions and its per-marker denominator. The download needs a free account, so we know who is using the dataset and can tell you when it is refreshed or a figure is corrected. Looking up a single marker by its rsID stays open to everyone, with no account. Broad searches show the first two matches, and an account opens the rest. An institutional email is approved instantly; a personal one takes a quick manual check.
Create a free account to download 39 MB compressed, 108 MB unpacked. Please cite it as the release of 30 July 2026; the data is refreshed monthly, so a later download will carry larger denominators.

Privacy and governance

Counts, never people.

The browser contains no personal information of any kind. No names, no dates of birth, no individual results, nothing that can be traced to a participant. It reports how many people carry a variant, the way a census reports how many people live in a state without naming them.

Nothing based on fewer than 20 people is ever displayed. Markers below that threshold are withheld from the published data entirely rather than shown with a suppressed count, and the page never computes a smaller subset of its own. That is the same small cell rule used by major national genomic programmes, so the standard is an accepted one rather than something we invented.

Collaboration

The datasets that matter most in Malaysia sit inside institutions. We can help you get to the right table.

A good local cohort is usually held by a university, a hospital or a national research institution, under their own governance and their own ethics approval. We do not hold those datasets and we do not speak for the people who do. What we can do is convene the right group around a question and support the science once a collaboration is agreed.

01

Tell us the question

Send us the research question and the population you need, not a data request. We keep a running picture of who in the local ecosystem works on what.

02

We group the interest

Where several groups are circling the same question, we bring them together, because a consortium is far more likely to clear a data access committee than a lone request.

03

We make the introduction

Where we know the institution holding a relevant cohort, we introduce you directly and step back. The conversation, and the decision, belong to them.

04

We run the platform work

Once a collaboration is agreed, we do what we are actually for: genotyping and assays, bioinformatics, sample logistics, ethics and grant documentation, and project coordination.

What we are not offering. We do not hold, broker, resell or grant access to any dataset we do not own, and nothing on this page should be read as an offer of access to a third party's data. Where we have processed samples for another organisation, we acted as a laboratory only and claim no rights over the resulting data. Access to any institutional cohort is decided by the institution that holds it and by its ethics committee, never by us.

Shape what comes next

The browser is free and open. The roadmap is decided by the people using it.

A free Prima Nexus account gets you the things an account is actually for: nominate the markers and genes your work depends on so they are curated next, hear when a refresh lands and what changed, and get the ancestry-label confirmation work as it completes. It takes a minute and costs nothing.

Already have an account? Nominate markers from the portal.