Public genomic metadata you can trust — and trace.
OpenBioData recovers missing or inconsistent sample metadata in NCBI records by tracing every field back to its source paper — with a citation and a confidence score you can check yourself.
The record exists. It just can't be trusted.
NCBI records are often missing exactly the fields that matter for real research — even when the answer is sitting in plain text in the paper that deposited the data.
A researcher needs 1,000+ public samples to study a disease population. He assumes the public record is accurate — that's the default assumption everyone starts with.
He opens the database. Country: blank. Sample type: wrong. Collection date: missing. So he reads the original papers by hand, one at a time, for weeks — before analysis even starts.
Run against his own hand-curated, trusted dataset, the same checks surface systematic errors — samples mislabeled by origin, era, or species. If a careful lab's own curation has this, the public commons has it too.
One pipeline. Two directions.
The same four-step process works on public records by default — or point it at your own internal data and it becomes an audit layer.
Trace public + internal sources
Give it an accession or a paper link. It finds the source publication and, if you connect one, your own internal records too — public sources are the default, internal is opt-in.
Cross-check
It doesn't stop at the original paper — it finds every publication that cites or reuses the same sample and checks them all against each other for agreement.
Recover & verify
Missing fields get filled in, existing fields get checked against the evidence, and every value ships with a citation and a confidence score — never a black box.
Validate internal data
Point the same pipeline at your own sample metadata instead of NCBI's, and recovery becomes auditing — surfacing exactly where your internal records disagree with the evidence.
Real records. Real gaps found.
A metadata-curation lab at a major public health graduate school is using OpenBioData to speed up manual curation and flag likely errors before they enter their reference dataset.
A genomic surveillance program is running its antimicrobial-resistance sample set through the tool to recover context missing from the original NCBI deposits.
An AI-native bioinformatics platform is piloting OpenBioData as the curated dataset layer feeding its own analysis agents.
Open for researchers. Private for teams.
Same engine, two ways to use it — depending on whether your data is public or yours alone.
Free & open source
MIT-licensed, self-hostable, and free to try — built for anyone tracing public NCBI records.
- 10 samples without an account, 30 signed in
- Confidence score and source citation on every field
- Self-host with your own API key — nothing sent to us
- Open pipeline: source-fetching and scoring logic is auditable on GitHub
Private validation, at your scale
For teams whose product or research depends on data quality they can't fully verify by hand.
- Cross-check your own internal metadata, not just public records
- Private deployment — your data stays yours
- Schema-aligned output for your existing pipelines
- Volume pricing and a scoped pilot before any commitment
Built by people who ran into this problem firsthand.
Vy Khanh Phung
B.S. Computer Science & Biochemistry, Dickinson College. Ran into public metadata quality problems firsthand during a research internship at Oxford, which became the starting point for OpenBioData.
Gowtham Gopalakrishnan
M.S. Data Science, University of Arizona. AI/ML engineer with hands-on research in computational drug repurposing and epidemic modeling — architects and leads engineering for BioMetadataAudit.
See what's missing in your data.
Try it free on public records, or talk to us about validating your own.