Code Ocean is a computational research platform that enables researchers to create reproducible environments and manage data through automated Dockerfile generation and integrated ML workflows. It addresses the challenges of slow collaboration and inconsistent result reproduction across diverse compute environments in computational science and bioinformatics.
Funding
$16.5M raised to dateRaised to date based on public sources. This may differ from the amount the company actually raised and is based only on what is publicly available on the internet.


MMFounders
Product
Problem
Computational scientists face challenges in collaborating effectively and reproducing research results consistently due to diverse compute environments and manual processes. Building and maintaining bespoke systems for computational biology and bioinformatics is time-consuming and resource-intensive.
Solution
Code Ocean provides a computational research platform that enables researchers to create reproducible environments and manage data through automated Dockerfile generation and integrated ML workflows. The platform streamlines collaboration by providing a shared environment for developing and sharing compute capsules, pipelines, and data assets. It ensures result provenance and reproducibility by tracking all development with Git and offering a lineage graph that visualizes every step of a computation. Code Ocean installs into existing cloud architectures, allowing users to set up environments, provision compute resources, and deploy to the cloud with a single click.
Target Audience
The primary users are computational scientists, bioinformaticians, and IT/engineering teams in biotech and pharmaceutical companies who need a secure, cloud-based environment for reproducible research and collaborative development.
Features
- Automated Dockerfile generation for creating reproducible compute environments
- Visual builder for creating and monitoring bioinformatics pipelines, with one-click import from nf-core
- Integration with MLflow for tracking machine learning models, parameters, and lineage
- Scalable compute and storage resources for analyzing large, multimodal datasets
- Automated result provenance and lineage graph generation for tracking data and results
- Support for various data analysis workflows using open-source software and containerization
- Tools for managing data access and ensuring compliance with FAIR principles
- Cloud management features for provisioning compute resources (CPUs, GPUs, RAM) and reducing cloud costs