Share sensitive data without sharing a single real record
Zero-Knowledge Synthetic Data generates a differentially-private clone of a sensitive dataset — patient records, PII — that matches the real statistical distributions while containing zero genuine records. Safe to share, train on, and export, with a mathematical privacy guarantee.
The data you need to use is the data you can't touch
Your most valuable datasets — patient outcomes, claims, customer PII — are locked behind HIPAA and GDPR, so partners, vendors, and even your own ML teams can't use them without months of legal review and re-identification risk. Zero-Knowledge Synthetic Data reads the real records only to compute noise-protected aggregate statistics, then samples an entirely new dataset that behaves like the original but contains none of its rows — so you can collaborate, train, and demo freely with a differential-privacy guarantee you can put in front of a regulator.
Why teams use it
Built to move risk off your team and speed to market up.
Differential privacy
Records are read only to compute Laplace-noised aggregates; you choose the ε privacy budget with your privacy officer.
Statistically faithful
Synthetic columns match the real distributions, with a per-column fidelity report so you can trust the output.
Safe to share & export
Every output row is freshly sampled — there are zero genuine records to re-identify, so the clone is safe to hand to partners.
How it works
Three steps, seconds each.
Provide a dataset
Paste a CSV of the sensitive data (or load the sample patient dataset).
Set the privacy budget
Choose ε — the differential-privacy strength — for the generation run.
Download the clone
Get a synthetic CSV plus a per-column fidelity + privacy-guarantee report.
Zero
Real records in output
ε-DP
Formal guarantee
Self-hosted
Runs on your data
A private clone with a fidelity report
See the synthetic output, the differential-privacy budget, and per-column fidelity — zero real records.
| age | systolic_bp | cholesterol | has_disease |
|---|---|---|---|
| 47 | 119 | 188 | 0 |
| 66 | 148 | 242 | 1 |
| 52 | 124 | 201 | 0 |
ε-differential privacy guarantee: every row freshly sampled from noise-protected aggregates. No genuine record is reproduced.
Product preview — representative of the shipping experience.
Frequently asked questions
The details teams ask before they deploy.
What is the privacy guarantee?+
Generation is differentially private: the real records are read only to compute Laplace-noised aggregate statistics, and every output row is freshly sampled from those aggregates. You set the ε (epsilon) privacy budget with your privacy officer — lower ε means stronger privacy.
Can any real record be re-identified from the output?+
No genuine rows are copied into the output — the synthetic dataset contains zero real records. Combined with the differential-privacy noise, that removes the row-level re-identification risk that blocks sharing raw data.
Will the synthetic data still be useful?+
Yes. Synthetic columns are calibrated to match the real statistical distributions, and every run returns a per-column fidelity report so you can confirm the clone behaves like the original before you rely on it.
What formats can I use?+
Provide a CSV (header row plus rows) by paste or upload, choose ε and the number of rows, and download a synthetic CSV plus a PDF fidelity-and-guarantee report. Each run is logged to your organization audit trail.
Where does generation run?+
On your own Inferix infrastructure. The sensitive source data never leaves your boundary for a third-party service.
Unlock without sharing a single real record for your team
Deploy on your own infrastructure with SSO, audit logs, and a named solutions engineer. Talk to sales to activate it on your plan.
- Self-hosted
- SSO & audit logs
- Volume pricing