Proximal provides high‑quality, complex datasets designed to train and evaluate frontier AI models, treating data as a first‑class research problem. By publishing open benchmarks and research on model performance across various harnesses, they enable developers to assess and improve AI agents that solve technical challenges without relying on domain‑expert bottlenecks.
Funding
Funding not disclosed

Founders
Product
Problem
Training frontier AI models for autonomous technical problem solving requires large, high-quality, complex datasets, yet such data is often scarce, fragmented, or dependent on costly domain expert curation.
Solution
Proximal treats data generation as a research discipline, producing and openly publishing extensive benchmark datasets focused on code generation and technical tasks. By releasing detailed implementation notes, leaderboards, and performance metrics, they enable developers to evaluate and improve models without building data pipelines from scratch. Their open‑source approach fosters transparent comparison of leading models such as Claude, Grok, and GPT, accelerating progress toward autonomous agents. The company continuously updates its datasets and benchmarks, ensuring relevance to evolving model capabilities and research needs.
Target Audience
Primary users are AI research labs, model developers, and engineering teams building autonomous agents that require high‑quality training data for technical problem solving.
Features
- Curated, large‑scale code‑generation datasets designed for training frontier language models
- Publicly available benchmarks and leaderboards that rank models across multiple performance dimensions
- Detailed implementation documentation and research notes accompanying each dataset release
- Open‑source distribution allowing unrestricted access for academic and commercial developers
- Continuous updates and new dataset releases aligned with emerging model architectures