Aurelius provides a decentralized platform that runs AI agents in persistent simulated worlds to evaluate their decision‑making under moral dilemmas and uncertainty. The system records full deliberation traces as immutable data units, generating high‑quality alignment datasets that labs and enterprises can use for benchmarking and training robust, accountable models.
Funding
Funding not disclosed
Founders
Product
Problem
Current AI alignment processes rely on limited human feedback and proprietary datasets, which are slow, expensive, and vulnerable to manipulation, leading to opaque standards and insufficient accountability for AI behavior in high‑impact domains.
Solution
Aurelius offers a decentralized infrastructure that uses persistent simulated worlds—world models—to place AI agents in moral dilemmas with competing interests, incomplete information, and real consequences. The platform scores agents on their behavior under pressure, capturing the full deliberation process as atomic “aene” data units. Top‑performing agents generate high‑fidelity alignment datasets, which are then used to train future models to reason robustly under uncertainty. By being open, permissionless, and multi‑agentic, Aurelius enables transparent benchmarking, scalable dataset creation, and community‑driven standards for AI alignment across sectors such as healthcare, finance, and hiring.
Target Audience
Primary users are AI research labs, safety teams, and enterprises developing models for high‑stakes applications who need transparent, scalable alignment data and benchmarking tools.
Features
- World‑model infrastructure that creates persistent simulated environments with social structures, resource constraints, and evolving agent relationships
- Behavioral scoring system that evaluates agents on decision‑making under pressure, not just answer correctness
- Automated generation of high‑quality alignment datasets from top‑performing agents’ full deliberation traces
- Open, permissionless protocol allowing any participant to contribute scenarios, agents, or evaluation metrics
- Multi‑agentic framework supporting diverse perspectives and competing goals to surface genuine moral tension
- Secure, immutable capture of observation, deliberation, and decision data as “aene” units for reproducible research