Skip to main content

General Compute

General Compute operates a neocloud providing dedicated racks of alternative inference silicon, such as Cerebras, SambaNova, Positron, and d-Matrix, for organizations running latency-sensitive AI workloads. The company splits model inference by running prefill on NVIDIA B300 GPUs while offloading decode phases to purpose-built ASICs, delivering up to 16.1x faster decode speeds. Customers can access the hardware via bare metal or an OpenAI-compatible API endpoint under a single contract.

San Francisco, United States · HQ
101K+ followers
  • Artificial Intelligence
  • AI Agents
  • Hardware
Updated 2 days ago

Funding

Funding rounds are not available yet.

Founders

Founder details are not available yet.

Product

Problem

Large language model inference for agentic workloads is bottlenecked by decode speed, as autoregressive token generation is memory-bound rather than compute-bound. Standard GPUs perform well at prefill but stall during sequential decode, causing latency to compound across the hundreds or thousands of tool calls an agent makes per run. This makes agent interactions feel sluggish and limits the viability of purpose-built AI applications that depend on rapid serial inference.

Solution

General Compute is a neocloud and deployment arm for heterogeneous compute, offering dedicated racks of alternative silicon designed specifically for high-speed decode. The company disaggregates inference workloads by keeping prefill on NVIDIA B300 GPUs while routing decode to fast ASICs such as SambaNova SN50, Cerebras, Positron, and d-Matrix. General Compute owns the hardware, handles siting, bring-up, and operations, and provides customers with dedicated capacity under contract—either as bare metal or through a single OpenAI-compatible API endpoint. By tailoring silicon to the memory-bound nature of autoregressive generation, the platform achieves up to 16.1x faster decode throughput compared to GPU-only infrastructure.

Target Audience

Primary customers are AI product teams and enterprises running latency-critical agentic applications—such as coding agents and tool-using LLM systems—that require sub-second decode latency and sustained serial inference throughput.

Features

  • Dedicated racks of purpose-built inference silicon (Cerebras, SambaNova, Positron, d-Matrix) sited with the customer's fleet under one contract
  • Disaggregated inference architecture that routes prefill to NVIDIA B300 GPUs and decode to ASICs optimized for memory-bound autoregressive generation
  • Up to 16.1x faster decode speeds and sustained throughput 100–1,000x above GPU baselines for agent workloads with serial tool calls
  • OpenAI-compatible API spanning prefill and decode, plus bare-metal access with root privileges on rented machines
  • Full model bring-up and software integration handled by General Compute, eliminating the need for teams to adapt to new chip software stacks
  • Transparent, apples-to-apples benchmarking against GPU cloud baselines across context sizes, with vendor specs cited honestly where unpublished
This profile is AI-generated and may contain inaccuracies.