Poindexter Labs offers a GDPR‑first platform that creates and curates olympiad‑level mathematical datasets with fully worked solutions, providing stepwise annotations, error taxonomies, and dual LaTeX and JSON/Parquet outputs. Their secure workflow includes variant‑rich, cheat‑resistant assessments for contributor vetting and an automated evaluation pipeline that combines exact‑match checks, rubric scoring, and programmatic verifiers to deliver clean, verifiable signals for model fine‑tuning or reinforcement learning. The service is aimed at AI research labs and enterprises needing high‑quality, auditable math training data for instruction‑tuning, benchmark creation, or safety testing.
Funding
£2M raised to dateRaised to date based on public sources. This may differ from the amount the company actually raised and is based only on what is publicly available on the internet.
Founders
Product
Problem
AI models for mathematics often suffer from low-quality training data, ambiguous reasoning steps, and insecure annotation pipelines, leading to poor performance on complex problem solving and limited trust in generated solutions.
Solution
Poindexter Labs provides a secure, GDPR‑first platform for creating and curating olympiad‑level mathematical datasets with fully worked solutions. Their workflow breaks problems into precise, stepwise reasoning components, adds error taxonomies, and generates both LaTeX and structured JSON/Parquet outputs for easy integration. Evaluators apply exact‑match checks, rubric scoring, and programmatic verifiers to produce clean, verifiable signals for model fine‑tuning or reinforcement learning. The service also includes variant‑rich, cheat‑resistant assessments for vetting contributors and ensures data governance through encrypted storage, least‑privilege access, and audit trails.
Target Audience
Primary customers are AI research labs, machine‑learning teams, and enterprises that require high‑quality, verifiable mathematical training data for instruction‑tuning, benchmark creation, or safety testing.
Features
- Generation of olympiad‑grade problems and complete solutions with stepwise annotations and dependency tracking
- Dual output format: publication‑ready LaTeX and machine‑readable JSON/Parquet with rich metadata
- Automated evaluation pipeline combining exact‑match, rubric scoring, and programmatic verifiers for reliable reward signals
- Secure contributor workflow with hidden‑gold checks, time‑to‑solve metrics, and style consistency analysis
- GDPR‑first data governance: UK ICO registration, encrypted storage, least‑privilege access, and full audit trails
- Contributor vetting using variant‑rich math assessments designed by former IMO/IOI/IPhO medalists