Skip to main content
P

Paris

Paris offers a distributed superintelligence platform that ensembles multiple large AI experts to improve inference performance. Their Paris‑2 system combines three 11 billion‑parameter video models with a lightweight router, delivering over 50% better Fréchet Video Distance and gains in CLIP, aesthetic, and motion quality compared to a monolithic baseline.

San FranciscoFounded 2023177K+ followers
Updated 1 month ago

Funding

Funding not disclosed

Funding rounds are not available yet.

Founders

Founder details are not available yet.

Product

Problem

Training large diffusion models requires homogeneous GPU superclusters and extensive compute resources, limiting accessibility for organizations with heterogeneous, commodity hardware. This centralization restricts the scalability and efficiency of generative AI for domains such as robotics, video synthesis, and world modeling.

Solution

Bagel Labs’ Paris platform implements Distributed Diffusion Models (DDM), which replace a monolithic diffusion model with an ensemble of smaller expert networks trained independently on data partitions without gradient synchronization. At inference, a lightweight router dynamically combines the outputs of these experts, delivering higher-quality results while using significantly less data and compute than traditional baselines. Paris‑1 demonstrates a 24% improvement in FID on image benchmarks, and Paris‑2 achieves over 50% reduction in Fréchet Video Distance for video generation, along with gains in CLIP alignment, aesthetics, and motion fidelity. The approach enables state‑of‑the‑art generative training on commodity hardware, unlocking compute capacity unavailable to conventional centralized training pipelines.

Target Audience

Primary customers are AI research labs, robotics developers, and media companies that need high‑quality generative models but lack access to large, homogeneous GPU farms.

Features

  • Decentralized training architecture that eliminates gradient synchronization across nodes
  • Ensemble of independent expert models (e.g., three 11B video experts) managed by a lightweight router
  • Significant metric improvements: 24% FID reduction for images, 50%+ FVD reduction for video
  • Compatibility with heterogeneous, commodity GPU clusters for both training and inference
  • Scalable to robotics, video synthesis, and world‑modeling applications without requiring homogeneous supercomputers
  • Open‑source release of Paris‑1 as the first publicly available DDM implementation
This profile is AI-generated and may contain inaccuracies.