Skip to main content
H

Hedgehog

Hedgehog provides an open‑source AI network platform that automates the management of GPU‑accelerated clusters using RoCEv2, ECN, and PFC to prevent congestion and maintain low latency. It offers zero‑touch lifecycle management, multi‑tenant virtual private clouds, and hardware‑agnostic deployment, enabling AI infrastructure teams to provision and monitor high‑performance fabrics without specialized InfiniBand expertise.

Seattle, United StatesFounded 2022262K+ followers
Updated 3 months ago

Funding

$3.8M raised to dateRaised to date based on public sources. This may differ from the amount the company actually raised and is based only on what is publicly available on the internet.

Funding rounds are not available yet.

Founders

Product

Problem

AI training, fine‑tuning, and inference workloads generate massive east‑west traffic across GPU clusters, leading to network congestion, packet loss, and high latency. Managing such high‑performance fabrics typically requires specialized InfiniBand expertise, which is scarce and costly. Consequently, organizations struggle to deliver fast job completion while keeping operational overhead low.

Solution

Hedgehog delivers an open‑source AI Network platform that automates the operation of backend GPU fabrics. The data plane uses RoCEv2, ECN, and PFC to detect and mitigate congestion, preserving effective bandwidth and minimizing latency for both training and inference. A cloud‑style user experience provides multi‑tenant virtual private clouds, gateway services for north‑south traffic, and zero‑touch lifecycle management, allowing existing cloud operations teams to provision and manage AI networks without dedicated network engineers. The solution is hardware‑agnostic, supporting equipment from multiple vendors, and integrates open observability tools for real‑time performance monitoring. An API‑driven provisioning model enables rapid tenant isolation and service scaling, while a virtual lab lets users test configurations without physical hardware.

Target Audience

The primary customers are AI infrastructure teams, cloud operators, and hyperscale or enterprise data centers that run GPU‑accelerated workloads and need high‑performance, low‑latency networking without dedicated InfiniBand staff.

Features

  • Congestion‑aware routing with RoCEv2, ECN, and PFC to maintain lossless, high‑throughput GPU interconnects
  • Virtual Private Cloud (VPC) service for north‑south traffic, providing tenant‑level network segmentation
  • Gateway services offering transit, load balancing, and security across public clouds and on‑premise resources
  • Zero‑Touch Lifecycle Management (ZTLm) that automates provisioning, upgrades, and health checks
  • Multi‑tenant isolation enforced via API‑driven VPC provisioning and role‑based access controls
  • Open observability stack delivering telemetry, latency histograms, and bandwidth utilization dashboards
  • Hardware‑agnostic deployment supporting switches and NICs from any vendor through open‑source drivers
  • Virtual lab environment for sandboxed testing of network topologies on standard VMs without physical equipment
This profile is AI-generated and may contain inaccuracies.