Blackcatlabs offers an open‑source red‑team evaluation framework that tests chat models for resistance to adversarially induced emotional dependency and false claims of being human or a therapist. The tool is designed to be reusable, allowing researchers and developers to run, extend, and build upon the evaluation to ensure safer conversational AI.
Funding
Funding not disclosed
Founders
Product
Problem
Chat models can be manipulated through adversarial prompts to develop emotional dependency on users or to falsely present themselves as human beings or qualified therapists, posing risks of misuse and user harm.
Solution
Blackcat Labs offers an open red‑team evaluation framework that systematically probes chat models for susceptibility to emotional dependency and false human or therapist claims. The framework provides a suite of adversarial prompts and assessment criteria that can be executed by researchers or developers to measure these failure modes. Results are published openly, enabling the community to reproduce, extend, and benchmark model behavior across versions and architectures. By surfacing these vulnerabilities early, the evaluation helps model creators improve safety controls and informs users about the reliability of conversational agents.
Target Audience
Primary users are AI safety researchers, LLM developers, and organizations deploying conversational agents who need to assess and mitigate risks of emotional manipulation and false identity claims.
Features
- Library of curated adversarial prompts targeting emotional attachment and impersonation behaviors
- Automated scoring metrics that quantify dependency risk and false identity assertions
- Open-source implementation allowing anyone to run the evaluation on their own models
- Extensible design supporting addition of new test cases and integration with continuous testing pipelines
- Publicly shared benchmark results to facilitate comparative analysis across models