Synth Labs is developing a hybrid approach that combines Reinforcement Learning from Human Feedback (RLHF) and Reinforcement Learning from AI Feedback (RLAIF) to enhance the alignment of large language models with human preferences. This method addresses the limitations of current alignment techniques by utilizing synthetic preferences to improve model performance and scalability.
Funding
Funding not disclosed

Founders
Product
Problem
Current methods for aligning large language models (LLMs) with human preferences, such as Reinforcement Learning from Human Feedback (RLHF), are resource-intensive and face scalability challenges. Relying solely on human preference labels can be limiting, while alternative approaches using AI-generated synthetic preferences may not accurately reflect human judgment.
Solution
Synth Labs is developing a hybrid approach that combines Reinforcement Learning from Human Feedback (RLHF) and Reinforcement Learning from AI Feedback (RLAIF) to improve the alignment of large language models with human preferences. This method leverages synthetic preferences to enhance model performance and scalability, addressing the limitations of current alignment techniques. By unifying RLHF and RLAIF methodologies, Synth Labs aims to create models that adapt and scale automatically, ensuring that AI's potential hinges on trust and interpretability. The company focuses on frontier challenges in AI post-training, using generative reward models to improve LLM behavior.
Target Audience
The primary audience includes AI researchers and developers working on large language models, particularly those focused on improving alignment, scalability, and trustworthiness.
Features
- Hybrid RLHF and RLAIF approach for enhanced LLM alignment
- Generative reward models for improved model behavior
- Utilization of synthetic preferences to improve model performance and scalability
- Focus on post-training techniques to unlock new capabilities in foundation models
- Research on controlling language models at inference time