Deasy Labs provides an automated workflow for data and AI teams to manage unstructured content metadata at scale. The platform automatically derives schemas, tags content using LLMs, and allows for human-in-the-loop validation. This enables teams to curate relevant data slices and export enriched metadata directly to vector databases or storage systems for improved AI retrieval performance.
Funding
Funding not disclosed
Founders
Product
Problem
Enterprises struggle to efficiently manage and retrieve information from unstructured data sources due to the lack of standardized metadata. Ensuring data quality and relevance for generative AI applications is further complicated by the difficulty in matching use cases with the best possible data sets.
Solution
Deasy Labs provides a metadata tagging solution that automates the generation and standardization of metadata from unstructured data, enabling efficient data retrieval and enhanced data governance. The platform connects to vector databases and can generate metadata from embeddings or underlying documents. It backward engineers metadata schemas from document corpora and extracts metadata at both chunk and document levels, creating hierarchical, multi-modal, and standardized metadata. Deasy Labs' retrieval agent uses the generated metadata to select the most relevant information and agents for specific tasks.
Target Audience
Deasy Labs targets enterprises seeking to improve knowledge management, data governance, and the performance of their AI applications by leveraging high-quality metadata for unstructured data.
Features
- Auto-suggested metadata generation through analysis of large document and image sets
- Customizable metadata definition via LLM-powered labeling
- Hierarchical metadata inference to build relationships between documents
- Automatic standardization and grouping of similar metadata values for easy filtering and updates
- Human-in-the-loop validation workflow for testing, analyzing, and refining metadata through reinforcement learning
- Quality scores and evidence generation for all metadata to provide easy validation
- Direct connection of metadata back into underlying vector databases
- Intelligent selection and filtering of relevant information
- Continuous and automated metadata maintenance, including dynamic taxonomies
- Integration with data sources such as Sharepoint, S3, AzureBlob, and Dropbox
- On-prem deployment within private clouds
- API access for platform integration
- Easy import and export of metadata to connect with existing data systems and MDM tools
- User permission management and controls