Skip to main content
A

AGI

This applied AI lab focuses on developing Artificial General Intelligence for daily use cases. Their initial product, AGI-0, functions as a personalized, proactive co-worker accessible via smartphone. The company is dedicated to advancing human-AI interaction through agentic technology.

San Francisco, United StatesFounded 2025203K+ followers
Updated 4 months ago

Funding

Funding not disclosed

Funding rounds are not available yet.

Founders

Product

Problem

Evaluating the performance and reliability of AI agents designed to interact with web interfaces is challenging due to the lack of standardized, realistic testing environments. Existing benchmarks often fail to capture the complexities and dynamic nature of real-world websites, hindering the development and deployment of effective web agents.

Solution

AGI, Inc. provides REAL Bench, a comprehensive evaluation framework for AI web agents. REAL Bench offers a collection of sandbox replicas of popular websites across various domains, designed to mimic the structural and functional elements of their real-world counterparts. This realistic environment allows developers and researchers to rigorously assess agent capabilities in navigation, information retrieval, and task completion. The platform includes a public leaderboard for benchmarking agent performance and a detailed scoring system that provides granular insights into successes and failures. By utilizing the AGI SDK, users can easily integrate and test their agents within this standardized testbed, facilitating the development of more robust and reliable AI web agents.

Target Audience

The primary audience includes AI researchers, developers of AI web agents, and companies seeking to deploy or improve their AI agent solutions for web interaction.

Features

  • **REAL Bench**: A standardized evaluation framework featuring sandbox replicas of 11 popular websites across e-commerce, social media, news, and travel sectors.
  • **Realistic Web Environment**: Replicas maintain the structural and functional elements of real-world sites, providing a challenging testbed for AI agents.
  • **AGI SDK**: A Python package that simplifies the integration and execution of agent evaluations within the REAL Bench framework.
  • **Leaderboard and Community Engagement**: A public platform for submitting agent performance data, fostering competition and collaborative improvement.
  • **Comprehensive Evaluation Metrics**: Detailed scoring based on task success rates, information retrieval accuracy, and action-taking reliability, with granular breakdowns by task category and website.
  • **Model Benchmarking**: Comparative analysis of various AI models (e.g., Claude-3.7-Sonnet-Thinking, Gemini-2.5-Pro-Experimental, DeepSeek-V3) and agent frameworks (e.g., Browser-Use, StageHand) against the REAL Bench.
  • **Reinforcement Learning Foundation**: The framework supports training autonomous web agents through reinforcement learning by providing structured environments and well-defined tasks.
This profile is AI-generated and may contain inaccuracies.