Fractile is developing specialized chips that perform all operations for running large language models directly in memory, eliminating the significant delays caused by moving model weights to the processor. This technology enables the fastest possible inference of the largest transformer networks, achieving speeds up to 100 times faster at one-tenth the cost of current systems.
Funding
$260M raised to dateRaised to date based on public sources. This may differ from the amount the company actually raised and is based only on what is publicly available on the internet.
PGFounders
Product
Problem
Existing hardware architectures create a significant bottleneck for large language model (LLM) inference because they require constant movement of model weights between memory and the processor, consuming the majority of compute time. This memory bottleneck limits the speed and efficiency of LLM inference, increasing costs and hindering real-time applications.
Solution
Fractile is developing specialized hardware that performs all operations required for LLM inference directly within memory, eliminating the need to move model weights to the processor. This in-memory computing approach drastically reduces the memory bottleneck, enabling faster and more efficient inference for large transformer networks. Fractile's technology aims to achieve significantly faster inference speeds at a fraction of the cost compared to current GPU-based systems. By removing the memory bottleneck, Fractile unlocks new possibilities for real-time LLM applications and enables the deployment of larger, more complex models.
Target Audience
The primary target audience includes organizations and researchers working with large language models who require faster, more efficient, and cost-effective inference solutions.
Features
- In-memory computing architecture that eliminates the memory bottleneck in LLM inference
- Specialized hardware designed for efficient execution of all operations required for LLM inference
- Optimized for large transformer networks
- Achieves significantly faster inference speeds compared to traditional GPU-based systems
- Reduces the cost of LLM inference