Skip to main content
Memvid processes audio and video files using Whisper for transcription, making spoken content searchable. Video files also support key frame extraction and playback from within the CLI.

How It Works

Key features:
  • Automatic transcription - Whisper converts speech to text
  • Timestamp alignment - Text segments linked to audio/video timecodes
  • Key frame extraction - Important frames from video
  • In-CLI playback - Play segments directly from terminal
  • Visual search - CLIP embeddings for video frames

Supported Formats

Audio

Video


Ingesting Audio

Basic Usage

Transcription Options

Whisper Models

Language Support

Whisper supports 99 languages. Specify for better accuracy:

Ingesting Video

Basic Usage

Frame Extraction

Visual Embeddings

Enable CLIP embeddings for visual search:

Searching Transcribed Content

Ask Questions

Visual Search (Video)


Playback

Playing Audio

Playing Video

Playback Controls


Use Cases

Meeting Recordings

Podcast Library

Video Tutorials

Lecture Archive

Voicemail/Call Logs


GPU Acceleration

Transcription is CPU-intensive. Enable GPU for faster processing:

macOS (Apple Silicon)

Linux/Windows (NVIDIA CUDA)

Performance Comparison


Batch Processing

Parallel Ingestion

Large Libraries


Frame Metadata

Each transcribed segment includes metadata:
Access metadata:

Troubleshooting

No Transcription Output

Poor Transcription Quality

Playback Issues

Out of Memory


Limitations


SDK Support

Audio/video processing is currently CLI-only. SDK support planned. Workaround:

Next Steps

Visual Embeddings

CLIP search for images and video frames

Memory Cards

Extract entities from transcriptions