Skip to main content
C

CaML

CaML develops methods to align large language models with beneficial goals by shaping their values and personas through data curation and fine-tuning. Their Synthetic Document Finetuning (SDF) technique cultivates more compassionate AI assistants, ensuring these models internalize desirable traits even after further training.

Updated 2 months ago

Funding

Funding not disclosed

Funding rounds are not available yet.

Founders

Founder details are not available yet.

Product

Problem

Large language models (LLMs) can inadvertently adopt misaligned behaviors and personas from their training data, even after fine-tuning. This poses a risk as AI capabilities advance, potentially leading to undesirable outcomes if these models do not internalize beneficial goals.

Solution

CaML researches methods to shape the values and personas of future AI systems through data curation and fine-tuning techniques. The project utilizes Synthetic Document Finetuning (SDF) to cultivate more compassionate and morally thoughtful AI assistants. By focusing on the impact of pre-training data, CaML aims to ensure that advanced AI systems align with desirable goals, promoting a positive societal impact. Their work investigates how to generate and filter data to robustly instill these values, even after subsequent fine-tuning stages.

Target Audience

The primary audience includes AI researchers, developers, and organizations focused on AI safety and alignment, particularly those working with large language models.

Features

  • Synthetic Document Finetuning (SDF) for robustly shaping AI personas and values.
  • Research into data generation and filtering methods to improve AI compassion and open-mindedness.
  • Custom benchmarks, including the Animal Harms Assessment (AHA 2.0), to evaluate AI compassion and moral open-mindedness.
  • Analysis of persona vectors to understand and quantify shifts in model behavior.
  • Evidence demonstrating that SDF-induced compassion persists through subsequent Supervised Fine-Tuning (SFT) and Reinforcement Learning from AI Feedback (RLAIF).
  • Development of pretraining pipelines for generating diverse, compassionate synthetic data.
  • Open-source model and data releases on Huggingface for community access and further research.
This profile is AI-generated and may contain inaccuracies.