Recursal.ai develops a post-transformer architecture that enables instant, serverless inference of Hugging Face models, achieving 100x cost efficiency for over 100 languages. Their platform allows users to effortlessly fine-tune and deploy the RWKV foundation model, making advanced AI accessible to a global audience.
Funding
Funding not disclosed
Founders
Product
Problem
Deploying and scaling large language models (LLMs) for inference can be computationally expensive, requiring significant infrastructure investment and specialized expertise. Existing solutions often lack cost-efficiency and accessibility, particularly for users working with a wide range of open-source models.
Solution
Recursal.ai offers Featherless, a serverless inference platform that enables instant deployment and scaling of Hugging Face models. The platform leverages a post-transformer architecture to achieve significant cost reductions, providing up to 100x cheaper inference for over 100 languages. Users can access a vast library of open-weight models and deploy them at scale for fine-tuning, testing, and production without the burden of server management or operational overhead.
Target Audience
The primary target audience includes AI teams, developers, and researchers who need cost-effective and scalable inference solutions for open-source LLMs.
Features
- Serverless architecture for instant and scalable LLM inference
- Support for a wide range of open-source models on Hugging Face, including Llama 2 and 3, Mistral, Qwen, and DeepSeek
- Compatibility with OpenAI SDK, LangChain, and LiteLLM
- Flat pricing with unlimited tokens and predictable billing
- No logging of prompts or completions for private, secure, and anonymous usage
- API access for fine-tuning, testing, and production
- Support for models up to 72B parameters