Lamin provides an open data framework for biology that enables querying, tracing, and validation of datasets and models at scale. It unifies lakehouse capabilities, data lineage tracking, feature store management, and bio-registry integration through a single API. This platform automates context generation for agents and researchers by linking data, models, and reports into queryable feature spaces.
Funding
$500K raised to dateRaised to date based on public sources. This may differ from the amount the company actually raised and is based only on what is publicly available on the internet.

Founders
Product
Problem
Biological research generates large-scale datasets and metadata that are often difficult to manage and standardize across different labs and experiments. Tracking data lineage across various notebooks, scripts, and pipelines can be challenging, hindering reproducibility and transparency. Existing tools often require extensive coding to validate, standardize, and annotate biological data, creating barriers to collaboration between computational and experimental researchers.
Solution
Lamin provides an open-source data infrastructure designed to streamline the management of biological data and metadata through a unified API. The platform tracks data lineage across notebooks, scripts, and pipelines, integrating with workflow managers like Redun and Nextflow to ensure reproducibility. It standardizes, validates, and annotates biological data with minimal code, facilitating collaboration between dry and wet labs. Lamin leverages distributed databases for zero-copy data transfer and enables scalable learning by transforming artifacts into queryable datasets, predictive models, and analytical insights. The system unifies metadata and ontologies, allowing users to manage genes, proteins, cell types, tissues, samples, and experiments while importing data from public ontologies.
Target Audience
Lamin is designed for biological researchers in both computational and experimental labs who need a scalable and reproducible data management solution.
Features
- Unified API for accessing data at scale, creating a "lakehouse" environment
- Automated tracking of data lineage across notebooks, scripts, and pipelines
- Integration with workflow managers like Redun and Nextflow
- Tools for validating, standardizing, and annotating data with minimal coding
- Support for distributed databases with zero-copy data transfer
- Ability to manage genes, proteins, cell types, tissues, samples, and experiments
- Import functionality for public ontologies
- Tools for transforming artifacts into queryable datasets, predictive models, and analytical insights