uListen is a Chrome extension that applies speech‑to‑text and speaker diarization to YouTube podcast videos, producing time‑coded captions and searchable transcripts. It also generates AI‑summarized key takeaways and two‑minute overviews, enabling users to jump to specific moments or skip ads and intros. The service offers a free catalog of pre‑processed episodes and a credit‑based premium tier for on‑demand processing and advanced fact‑checking features.
Funding
Funding not disclosed
Founders
Product
Problem
Listeners of YouTube podcasts often lose valuable insights because episodes lack searchable transcripts, speaker‑identified captions, and concise summaries, making it difficult to revisit specific points or verify information later.
Solution
uListen delivers a Chrome extension that applies large‑language‑model‑based speech‑to‑text and summarization pipelines to YouTube podcast videos, generating real‑time, speaker‑tagged captions and searchable transcripts. The extension overlays AI‑driven key takeaways and concise summaries directly in the video player, allowing users to jump to any moment with a single click. Core AI features—including karaoke‑style highlighting, ad/intro skipping, and chapter navigation—are offered free for a curated catalog of over 6,000 pre‑processed episodes. Premium plans unlock on‑demand processing for any YouTube podcast, deeper fact‑checking, and higher‑resolution summaries, all billed via a credit system. All data remains client‑side and is never sold, with optional Google sign‑in required only for paid upgrades.
Target Audience
The primary users are knowledge‑workers, professionals, and lifelong learners who consume YouTube podcast content and need fast retrieval of specific ideas.
Features
- Automatic speech recognition (ASR) with speaker diarization to produce accurate, time‑coded captions.
- Large‑language‑model summarization that extracts key insights and generates two‑minute “at‑a‑glance” overviews.
- Full‑text indexing of transcripts enabling Ctrl + F‑style search and click‑to‑jump navigation within the video.
- Smart playback controls that detect and skip ads, intros, and chapter markers using audio fingerprinting.
- Karaoke‑style word highlighting synchronized to audio for enhanced accessibility.
- Cloud‑native processing pipeline with end‑to‑end encryption; free tier runs on a pre‑cached episode library, premium tier processes arbitrary videos on demand.
- Credit‑based billing for on‑demand episodes and premium AI modules such as fact‑checking and deep‑insight extraction.