dbTwin provides a platform that generates synthetic clinico‑genomic datasets preserving the underlying biology, allowing researchers to share high‑dimensional RNA‑seq and clinical data without patient privacy risks. The tool runs inside a user’s private cloud, requiring only an expression counts matrix and metadata, and can produce synthetic data in days instead of months, eliminating IRB delays and data use agreements.
Funding
Funding not disclosed
Founders
Product
Problem
Researchers need to share high‑dimensional RNA‑seq and clinical metadata for translational studies, but patient privacy regulations, IRB approvals, and data‑use agreements create lengthy delays and legal risk.
Solution
dbTwin offers a private‑cloud platform that generates synthetic clinico‑genomic datasets preserving the original biological signals. Users upload their raw expression matrix and paired clinical metadata, and the system creates synthetic RNA‑seq counts and matching metadata that retain disease‑relevant patterns while removing identifiable patient information. Because the tool runs entirely within the user’s own environment, no data leave the secure infrastructure, eliminating the need for IRB review or data‑use contracts. Synthetic cohorts can be produced in days, enabling rapid collaboration across biobanks, disease foundations, cancer centers, and biopharma partners without compromising privacy.
Target Audience
Primary customers are biobanks, disease foundations, cancer and rare‑disease centers, biopharma data teams, and academic researchers who need to exchange molecular datasets without exposing patient privacy.
Features
- Bulk ingestion of RNA‑seq count matrices and associated clinical metadata with minimal preprocessing
- Synthetic data generation that maintains high‑dimensional biological relationships and disease signals
- Fully on‑premise deployment on private cloud infrastructure; no external data transfer, GPUs, or deep‑learning training required
- Three‑step workflow (Ingest → Generate → Share) with automated schema matching and validation
- API and web console for programmatic access and batch export of synthetic CSV files
- Role‑based sharing controls allowing institutions to distribute synthetic cohorts to external collaborators safely