CaML develops methods to align large language models with beneficial goals by shaping their values and personas through data curation and fine-tuning. Their Synthetic Document Finetuning (SDF) technique cultivates more compassionate AI assistants, ensuring these models internalize desirable traits even after further training.
Funding
Funding not disclosed
Founders
Product
Problem
Large language models (LLMs) can inadvertently adopt misaligned behaviors and personas from their training data, even after fine-tuning. This poses a risk as AI capabilities advance, potentially leading to undesirable outcomes if these models do not internalize beneficial goals.
Solution
CaML researches methods to shape the values and personas of future AI systems through data curation and fine-tuning techniques. The project utilizes Synthetic Document Finetuning (SDF) to cultivate more compassionate and morally thoughtful AI assistants. By focusing on the impact of pre-training data, CaML aims to ensure that advanced AI systems align with desirable goals, promoting a positive societal impact. Their work investigates how to generate and filter data to robustly instill these values, even after subsequent fine-tuning stages.
Target Audience
The primary audience includes AI researchers, developers, and organizations focused on AI safety and alignment, particularly those working with large language models.
Features
- Synthetic Document Finetuning (SDF) for robustly shaping AI personas and values.
- Research into data generation and filtering methods to improve AI compassion and open-mindedness.
- Custom benchmarks, including the Animal Harms Assessment (AHA 2.0), to evaluate AI compassion and moral open-mindedness.
- Analysis of persona vectors to understand and quantify shifts in model behavior.
- Evidence demonstrating that SDF-induced compassion persists through subsequent Supervised Fine-Tuning (SFT) and Reinforcement Learning from AI Feedback (RLAIF).
- Development of pretraining pipelines for generating diverse, compassionate synthetic data.
- Open-source model and data releases on Huggingface for community access and further research.