DataJoint provides a programmable data‑science platform that links laboratory instruments, code, and multimodal datasets into version‑controlled, relational workflows. It automatically tracks provenance, runs containerized analysis pipelines on local, HPC, or cloud resources, and offers built‑in quality‑control, fine‑grained access controls, and modular “Elements” for common biomedical assays.
Funding
$4.9M raised to dateRaised to date based on public sources. This may differ from the amount the company actually raised and is based only on what is publicly available on the internet.

CFICFounders
Product
Problem
Scientific labs generate massive, multimodal datasets that are stored in disparate systems, making it difficult to maintain provenance, enforce quality control, and reproduce analyses. Manual stitching of code, metadata, and compute environments leads to errors, duplicated effort, and delays in publishing. Without a unified framework, scaling experiments across instruments, collaborators, and time becomes unsustainable.
Solution
DataJoint delivers a programmable data‑science platform that binds instruments, code, and data into a single, version‑controlled workflow. By defining relational schemas and automatic dependency graphs, the system records every transformation, ensuring full provenance and reproducible results. Pipelines are triggered on data arrival, executing containerized code in parallel and handling heterogeneous modalities—from electrophysiology to imaging and omics. Integrated Jupyter notebooks, Plotly Dash dashboards, and native Python/Matlab APIs let researchers query, visualize, and iterate on data without leaving their analysis environment. Built‑in quality‑control modules flag anomalies in real time, while role‑based permissions and encrypted storage keep data secure and compliant with institutional policies. The platform’s modular “Elements” library provides ready‑to‑use reference implementations for common neuroscience and biomedical assays, which can be customized or extended to fit any experimental design.
Target Audience
Primary users are research laboratories and core facilities in neuroscience, oncology, and other biomedical domains that need to manage large, multimodal datasets and run reproducible analysis pipelines. The platform also serves data scientists and bioinformaticians who require programmatic access to curated, provenance‑rich data for AI and machine‑learning workflows.
Features
- Relational schema engine with automatic dependency tracking and full metadata lineage for every data object.
- Container‑orchestrated execution engine that launches Python, MATLAB, or custom binaries on local, HPC, or cloud resources.
- Real‑time quality‑control and anomaly detection pipelines that generate alerts and audit logs.
- Integrated JupyterLab workbench with Plotly Dash visualizations and seamless Weights & Biases logging for machine‑learning experiments.
- Modular “DataJoint Elements” library offering pre‑built pipelines for electrophysiology, calcium imaging, optogenetics, multi‑omics, and more.
- Fine‑grained access control, end‑to‑end encryption, and compliance‑ready audit trails for NIH and GDPR requirements.
- RESTful API and Python client that support ad‑hoc queries without writing SQL, enabling rapid data exploration and reporting.
- Scalable metadata catalog that links raw files, processed results, and model artifacts across projects and institutions.