Skip to main content
FA

Felafax AI (YC S24)

Provides an open-source AI platform that optimizes machine learning training on non-NVIDIA GPUs, including Google TPUs and AWS Trainium, using a custom-built framework with XLA compiler and JAX. Achieves H100-level performance at 30% lower cost while enabling on-premises deployment, seamless scaling from 8 to 1024 chips, and automated ML operations for large models like Llama 3.1 405B.

Founded 20245500+ followers
Updated 20 months ago

Funding

$500K raised to dateRaised to date based on public sources. This may differ from the amount the company actually raised and is based only on what is publicly available on the internet.

Funding rounds are not available yet.

Founders

Product

Problem

Training large machine learning models requires significant computational resources, often relying on expensive NVIDIA GPUs. Organizations seeking to deploy AI on-premises or leverage alternative hardware accelerators face challenges in achieving optimal performance and cost-efficiency. Scaling training across multiple accelerators and managing the complexities of ML operations further compounds these issues.

Solution

Felafax provides an open-source AI platform designed to optimize machine learning training across diverse hardware accelerators, including Google TPUs, AWS Trainium, and AMD GPUs, in addition to NVIDIA. The platform utilizes a custom-built framework with an XLA compiler and JAX to achieve performance comparable to NVIDIA H100 GPUs at a reduced cost. It enables on-premises deployment, facilitating data security and privacy, while also offering seamless scaling from 8 to 1024 chips. Felafax automates ML operations, including model partitioning and multi-controller training, allowing users to focus on model innovation rather than infrastructure management.

Target Audience

The primary target audience includes enterprises and research institutions looking to train and deploy large AI models cost-effectively on diverse hardware infrastructure, while maintaining data privacy through on-premises deployment.

Features

  • Custom training platform built with XLA compiler and JAX for optimized performance on various accelerators
  • Support for on-premises deployment within a user's Virtual Private Cloud (VPC)
  • One-click cluster spin-up for scaling training from 8 to 1024 TPU chips
  • Automated ML operations, including model partitioning and multi-controller training and inference
  • Out-of-the-box templates with pre-configured environments for PyTorch XLA and JAX
  • No-code UI for fine-tuning models, with Jupyter notebook access for advanced customization
This profile is AI-generated and may contain inaccuracies.