The startup offers a generative artificial intelligence platform that enables enterprises to process and analyze data in all Indian languages. By integrating multi-modal route planning and traffic signal interconnectivity, the platform enhances real-time data access and collaboration, fostering a robust data economy for businesses.
Funding
$10.1M raised to dateRaised to date based on public sources. This may differ from the amount the company actually raised and is based only on what is publicly available on the internet.
Founders
Product
Problem
Many AI models and natural language processing tools lack adequate support for the diverse set of Indian languages, hindering effective communication and data analysis for a significant portion of the global population. This gap limits the potential for AI to address local needs and participate in the digital economy.
Solution
Soket AI is developing foundational large language models (LLMs) specifically for Indic languages, aiming to create ethical and inclusive AI. Their work focuses on building multilingual models, efficient tokenization methods, and techniques for cross-lingual knowledge transfer. Soket AI's models are designed to understand and generate text in multiple Indian languages, enabling more effective AI applications in these regions. The company is also researching methods for faster and more efficient inference, as well as fact alignment to ensure the reliability of the generated content.
Target Audience
The primary target audience includes AI researchers, developers, and businesses seeking to build applications and services that support Indian languages.
Features
- Multilingual tokenizer for efficient encoding of Indic languages.
- BHASHA-7B series of transformer models pre-trained on 22 scheduled languages of India, with an 8K context length.
- Instruction-tuned versions of BHASHA-7B models aligned with human intents.
- Realtime Speech API for building natural voice agents that understand and respond across any language or platform.
- Pragna-1B: A 1.25 billion parameter open-source multilingual foundational model supporting Hindi, Gujarati, Bangla, and English.
- Open-source COOM framework for scalable optimizations of language models.