
Marvin Systems provides an AI-powered document anonymization and pseudonymization platform that helps organizations protect sensitive information across PDFs, Word documents, PowerPoint files, and emails. The platform detects 30 types of personal and organizational data, then replaces it with consistent pseudonyms or generic tags while preserving the original document format. It includes an editor for review, de-pseudonymization capabilities, and API access for workflow automation.
Funding
Funding not disclosed
Founders
Product
Problem
Organizations handling sensitive documents face significant challenges in protecting personal data across various file formats, including PDFs, Office documents, and email threads. Manual redaction is time-consuming, error-prone, and often fails to catch all identifying information, especially when documents contain embedded images, comments, or version-tracking metadata that may expose personal data.
Target Audience
Primary customers are organizations handling sensitive documents, including legal firms, HR departments, financial institutions, and healthcare providers that need to comply with data protection regulations while processing client records, contracts, or internal documentation.
Features
- Detection of 30 data types including personal, organizational, and location-based information
- Dual-mode pseudonymization with strict matching (exact entity grouping) and dynamic matching (cross-type and typo-tolerant grouping)
- Support for PDF, Word, PowerPoint, and .msg email files with format preservation
- Integrated document editor with preview toggle, automatic zoom, and Word document rendering
- Configurable data retention policies including zero-data-retention options and custom deletion timelines
- API access for automation and third-party integrations, available from the Pro plan
- Session management supporting up to 50 files per case with consistent pseudonym preservation
- Guest mode access allowing document processing without account creation
- Image anonymization within Office documents with three processing options (skip, remove, or process)
- Detection of identifying data in document comments, author information, and version tracking blocks