Skip to main content
C

Chonkie

Chonkie provides proactive research automation by monitoring sources and summarizing key signals relevant to user-defined topics. The platform integrates private internal documents with public data to generate contextual reports and visualizations from disparate data sources. Users can ask follow-up questions on insights, receive cited answers, and deploy the system with enterprise-grade security on their own infrastructure.

San Francisco, United States2300+ followers
Updated 2 months ago

Funding

Funding not disclosed

Funding rounds are not available yet.

Founders

Product

Problem

Preparing unstructured documents for AI applications often involves complex data ingestion, cleaning, and transformation processes. This can be time-consuming and requires specialized expertise, hindering the efficient development and deployment of AI models that rely on accurate and well-formatted data.

Solution

Chonkie provides an open-source data ingestion pipeline designed to streamline the preparation of unstructured documents for AI applications. The platform automates the cleaning, chunking, and enrichment of data, transforming raw text into AI-ready formats. It facilitates secure connections to vector databases and supports flexible data export, enabling developers to build more accurate and efficient AI models with reduced manual effort. Chonkie offers both a downloadable library and a cloud-based service for flexible deployment.

Target Audience

The primary users are AI/ML engineers, data scientists, and developers building RAG (Retrieval-Augmented Generation) pipelines and other AI applications that require efficient data preparation and transformation.

Features

  • **Data Ingestion:** Supports ingestion from various document formats including TXT, PDF, CSV, Markdown, and code files (JS/TSX, Python, Java, C/C++, Rust).
  • **Data Cleaning & Preprocessing:** Includes a "Chef" stage for text cleaning, normalization, PII removal, and punctuation standardization.
  • **Advanced Chunking Strategies:** Offers multiple chunking methods such as `TokenChunker`, `SentenceChunker`, `RecursiveChunker`, `SemanticChunker`, `SDPMChunker`, `LateChunker`, `CodeChunker`, `NeuralChunker`, and `SlumberChunker`.
  • **Text Enrichment:** Features a "Refinery" stage to add metadata like embeddings, summaries, and topics to processed chunks.
  • **Vector Database Integration:** Provides "Handshakes" for direct ingestion into popular vector databases including Chroma, Qdrant, Turbopuffer, and Pgvector.
  • **Data Export:** Includes "Porters" for exporting processed data in formats like JSON and Datasets.
  • **Tokenizer Flexibility:** Supports multiple tokenizers including Hugging Face `tokenizers`, `tiktoken`, and custom token counting functions.
  • **Embedding Provider Support:** Integrates with embedding models via `AutoEmbeddings` from providers like Sentence Transformers, OpenAI, Cohere, Gemini, Jina, and VoyageAI.
  • **LLM Integration:** Offers "Genies" for interacting with LLMs (e.g., `OpenAIGenie`, `GeminiGenie`) to support advanced chunking strategies.
  • **Lightweight & Fast:** Designed for minimal installation size (15MB default) and high processing speeds, outperforming competitors in benchmarks for size and speed.
  • **Open-Source:** Available under the MIT license, fostering community contributions and transparency.
  • **Chonkie Cloud:** A hosted service for cloud-based data processing and API access.
This profile is AI-generated and may contain inaccuracies.