Turn any long-form video (podcasts, streams, interviews, debates) into viral 9:16 vertical shorts with AI-powered clipping, auto-framing, kinetic captions, and one-click export β 100% free, forever, no watermarks, no limits.
OpenClip Studio is a fully local, self-hosted video repurposing platform that does everything paid tools like Opus Clip, Vizard, and Descript do β but completely free, running on your own hardware. No cloud costs, no API quotas, no watermarks.
Drop in a video (or paste a YouTube link) β Get 5 ranked viral clips β Customize captions & framing β Export in 1080Γ1920.
| Feature | Details |
|---|---|
| π§ AI Viral Clip Discovery | Context-aware algorithm finds the most engaging, complete-narrative moments (30β45s) with hook detection, sentiment scoring, engagement-pause analysis, and speaker-turn tracking |
| ποΈ Local Whisper Transcription | Faster-Whisper (INT8 quantized) runs entirely on CPU β no cloud APIs needed. Word-level timestamps for pixel-perfect subtitle sync |
| π― AI Subject Tracking | YOLOv8n ONNX model detects speakers and tracks their position across frames for intelligent 9:16 reframing |
| π₯ Dual-Speaker Split-Screen | Opus Clip-style stacked top/bottom layout β automatically crops each speaker into their own panel with a styled divider |
| βοΈ Smart Jump-Cut Engine | Detects dead air (>0.5s silences), generates jump-cut manifests, and maps SFX cue points (vine boom, whoosh, etc.) |
| π¬ Kinetic Captions | Multiple subtitle styles (Hormozi, MrBeast, Minimal, Neon, Typewriter) with word-by-word highlight animation and emoji triggers |
| π₯ YouTube/URL Import | Paste any YouTube, Vimeo, or direct video URL β yt-dlp handles the download automatically |
| π¨ Full Studio Editor | Real-time preview, caption styling, reframe mode selection, SFX mixing, and export controls in a premium dark-mode UI |
| π€ One-Click Export | Renders production-ready 1080Γ1920 MP4 with burned subtitles, H.264/AAC encoding, and optional SFX overlay |
| π No Limits | No watermarks, no video length caps, no monthly quotas. Process as many videos as you want |
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β CLIENT (React + Vite) β
β β
β IngestionZone β ProcessingStatus β ClipList β Studio β
β β
β Components: β
β β’ IngestionZone.jsx β Upload / YouTube URL / Samples β
β β’ ProcessingStatus.jsx β Real-time pipeline progress β
β β’ ClipList.jsx β Ranked viral clips with scores β
β β’ StudioEditor.jsx β Full editor with live preview β
β β’ ExportModal.jsx β Render progress & download β
β β’ ApiSettingsModal.jsx β Optional AI key configuration β
β β’ Header.jsx β App header & navigation β
ββββββββββββββββββββββββ¬ββββββββββββββββββββββββββββββββββββ
β HTTP REST API (port 5000)
ββββββββββββββββββββββββΌββββββββββββββββββββββββββββββββββββ
β SERVER (Express.js) β
β β
β server.js β API routes & middleware β
β β
β Services: β
β βββ ffmpegService.js β Video probe, normalize, β
β β crop, render, preview β
β βββ transcribeService.js β Whisper orchestration, β
β β Groq API fallback β
β βββ transcribe_local.py β Faster-Whisper INT8 on CPU β
β βββ viralityService.js β AI clip discovery engine β
β βββ subtitleService.js β ASS subtitle generation β
β βββ jumpCutService.js β Dead-air removal & SFX β
β βββ downloaderService.js β yt-dlp URL downloads β
β βββ tracker_local.py β YOLOv8n ONNX tracking β
β βββ generate_sfx.py β Procedural SFX synthesis β
β β
β Config: β
β βββ hookPhrases.json β Custom hook phrase config β
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
The core differentiator of OpenClip Studio is its context-aware viral clip discovery engine. Here's exactly how it selects engaging moments:
The input video's audio is extracted to 16 kHz WAV and fed to Faster-Whisper (CTranslate2 INT8 quantization) running locally on CPU. This produces word-level timestamps with ~50ms accuracy β no cloud APIs required.
"What" β 88.12s β 88.34s
"does" β 88.34s β 88.52s
"your" β 88.52s β 88.71s
"sign" β 88.71s β 88.95s
"mean?" β 88.95s β 89.20s
Words are grouped into complete grammatical thoughts using:
- Terminal punctuation (
.?!) - Natural pause boundaries (>0.85s gap between words)
- Dangling-word protection β never breaks on connectors like "and", "but", "the"
Each candidate clip window (30β45 seconds) is scored on multiple axes:
| Signal | Weight | Description |
|---|---|---|
| Hook Type | +12 to +18 | Question hooks, bold statements, personal stories, thought-provoking topics |
| Speaker Turns | +1 to +5 | Dialogue exchanges (pronoun detection: "you", "I", "we") indicate conversation |
| Pause Emphasis | +2 per pause | Pauses >0.8s suggest dramatic emphasis or turn-taking |
| Sentiment | Β±2 per word | Positive words (amazing, love, wow) boost; negative words (hate, terrible) add intensity |
| Payoff Quality | +6 to +12 | Does the clip end with a resolution? ("believe", "peace", "solution", period/question mark) |
| Dialogue Completeness | +10 | Contains both a question AND an answer/explanation |
| Pacing (WPM) | +6 | Optimal speech rate: 115β180 words per minute |
| Banter Penalty | β20 to skip | Stream setup, mic checks, small talk ("check check", "can you hear me") are discarded |
| Generic Intro Penalty | β20 | "What's up everybody", "welcome back" in first 60 seconds |
After identifying a strong clip, the engine attempts to expand it by adding one sentence before (for better hook setup) or after (for payoff/conclusion) β as long as the result stays within the 30β45s target window.
Overlapping candidates (>30% time overlap) are merged, keeping only the highest-scoring version.
If a Groq or Gemini API key is provided, the engine can alternatively use Llama-3.3 70B or Gemini 1.5 Flash for even smarter clip selection β but the local algorithm produces great results without any API keys.
OpenClip Studio supports four reframing modes for converting 16:9 source video into 9:16 vertical shorts:
| Mode | Description |
|---|---|
| π― Smart Track | AI tracks the primary speaker's horizontal position and dynamically crops a 1080px-wide window around them |
| π₯ Split Stacked | Opus Clip-style dual-speaker layout: two 1080Γ960 panels stacked vertically with a styled divider |
| π Center Crop | Simple center crop β fast and reliable for talking-head content |
| π«οΈ Blur Fill | Frosted glass background (10x downscale + boxblur) with the original video overlaid in the center |
| Style | Look |
|---|---|
| Hormozi | Bold white text with yellow keyword highlights and black outline β the Alex Hormozi signature look |
| MrBeast | Heavy impact font with colored word-by-word pop animations |
| Minimal | Clean, thin white text β modern and unobtrusive |
| Neon | Glowing colored text with soft shadow β eye-catching on dark backgrounds |
| Typewriter | Monospaced font with character-by-character reveal animation |
All styles support emoji triggers β when specific keywords appear (money πΈ, AI π€, fire π₯, brain π§ , etc.), the corresponding emoji is automatically inserted.
| Tool | Required | Install |
|---|---|---|
| Node.js | v18+ | nodejs.org |
| Python 3 | 3.9+ | Usually pre-installed on Linux/macOS |
| FFmpeg | Latest | sudo apt install ffmpeg or brew install ffmpeg |
| yt-dlp | Latest | pip install yt-dlp (for YouTube URL imports) |
# 1. Clone the repository
git clone https://github.com/paragbml/open-clip-studio.git
cd open-clip-studio
# 2. Install server dependencies
npm install
# 3. Install client dependencies
npm --prefix client install
# 4. Install Python AI dependencies
pip install faster-whisper onnxruntime numpy
# 5. (Optional) Download YOLOv8n model for subject tracking
python3 -c "
from huggingface_hub import hf_hub_download
hf_hub_download('s1777/yolo-v8n-onnx', 'yolov8n.onnx')
"
# 6. Start development servers (client + server)
npm run devThe app will be available at http://localhost:5173 (client) with the API server on http://localhost:5000.
Create a .env file in the project root for enhanced clip discovery:
# Optional β local AI works great without these
GROQ_API_KEY=gsk_your_groq_key_here
GEMINI_API_KEY=your_gemini_key_hereopen-clip-studio/
βββ package.json # Root package with dev scripts
βββ .gitignore
βββ .env # (Optional) API keys
β
βββ client/ # React + Vite frontend
β βββ package.json
β βββ vite.config.js
β βββ index.html
β βββ src/
β βββ main.jsx # App entry point
β βββ App.jsx # Root component & view routing
β βββ App.css # App-level styles
β βββ index.css # Global design system
β βββ components/
β βββ Header.jsx # Top navigation bar
β βββ IngestionZone.jsx # Video upload / URL paste
β βββ ProcessingStatus.jsx # Pipeline progress UI
β βββ ClipList.jsx # Ranked clip cards
β βββ StudioEditor.jsx # Full clip editor
β βββ ExportModal.jsx # Render & download dialog
β βββ ApiSettingsModal.jsx # API key configuration
β
βββ server/ # Express.js backend
βββ server.js # API routes & middleware
βββ config/
β βββ hookPhrases.json # Customizable hook phrases
βββ services/
β βββ ffmpegService.js # Video processing pipeline
β βββ transcribeService.js # Transcription orchestrator
β βββ transcribe_local.py # Faster-Whisper engine
β βββ viralityService.js # AI clip discovery engine
β βββ subtitleService.js # ASS subtitle generator
β βββ jumpCutService.js # Dead-air & SFX engine
β βββ downloaderService.js # yt-dlp URL downloader
β βββ tracker_local.py # YOLOv8n subject tracker
β βββ generate_sfx.py # Procedural SFX synthesis
βββ uploads/ # (gitignored) Uploaded videos
βββ exports/ # (gitignored) Rendered clips
βββ samples/ # (gitignored) Sample videos
βββ assets/
βββ sfx/ # Sound effect WAV files
All endpoints are served from http://localhost:5000.
| Method | Endpoint | Description |
|---|---|---|
POST |
/api/upload |
Upload a video file (up to 4GB) |
POST |
/api/download-url |
Download from YouTube/Vimeo URL |
POST |
/api/process |
Run full pipeline: probe β transcribe β discover clips |
POST |
/api/render-clip |
Render a final 1080Γ1920 short with captions & SFX |
| Method | Endpoint | Description |
|---|---|---|
GET |
/api/status |
Health check & capability report |
GET |
/api/samples |
List available sample videos |
GET |
/api/stream?path=... |
HTTP 206 range-request video streaming |
GET |
/api/clip-preview?filePath=...&startTime=...&duration=... |
Generate lightweight clip preview MP4 |
GET |
/api/clip-tracking?filePath=...&startTime=...&duration=... |
Run YOLOv8n subject tracking |
POST |
/api/jump-cuts |
Calculate jump-cut segments & SFX events |
# 1. Upload a video
curl -X POST http://localhost:5000/api/upload \
-F "video=@my_podcast.mp4"
# Response: { "success": true, "filePath": "/path/to/upload.mp4", ... }
# 2. Process (transcribe + find viral clips)
curl -X POST http://localhost:5000/api/process \
-H "Content-Type: application/json" \
-d '{ "filePath": "/path/to/upload.mp4" }'
# Response: { "clips": [{ "start": 88.0, "end": 124.0, "viralityScore": 98, ... }] }
# 3. Render a clip
curl -X POST http://localhost:5000/api/render-clip \
-H "Content-Type: application/json" \
-d '{
"filePath": "/path/to/upload.mp4",
"startTime": 88.0,
"duration": 36.0,
"reframeMode": "smart_track",
"burnSubtitles": true,
"words": [...]
}'
# Response: { "downloadUrl": "/exports/openclip_xxx.mp4" }Customize which phrases the AI considers as strong hooks. Clips starting with these phrases get a +12 virality score boost:
[
"can you explain",
"what does",
"the truth is",
"here's the thing",
"nobody talks about",
"let me tell you"
]| Variable | Default | Description |
|---|---|---|
PORT |
5000 |
API server port |
GROQ_API_KEY |
β | Optional Groq API key for Llama-3 clip analysis |
GEMINI_API_KEY |
β | Optional Gemini API key for enhanced analysis |
FFMPEG_PATH |
ffmpeg |
Custom FFmpeg binary path |
FFPROBE_PATH |
ffprobe |
Custom FFprobe binary path |
YTDLP_PATH |
~/.local/bin/yt-dlp |
Custom yt-dlp binary path |
| Feature | Opus Clip (Pro $19/mo) | OpenClip Studio (Free) |
|---|---|---|
| Clip Discovery | AI (cloud) | AI (local + optional LLM) |
| Virality Scoring | β | β Context + sentiment + engagement |
| Dual-Speaker Split | β | β Stacked top/bottom |
| Subject Tracking | β (cloud GPU) | β YOLOv8n on CPU |
| Kinetic Captions | β | β 5 styles + emoji triggers |
| YouTube Import | β | β via yt-dlp |
| Monthly Limit | 100β300 mins | Unlimited |
| Watermark | On free plan | Never |
| Data Privacy | Cloud upload | 100% local |
| Price | $19β$39/month | $0 forever |
- Frontend: React 19, Vite 8, Lucide React icons, Canvas Confetti
- Backend: Node.js, Express 4, Multer
- Video Processing: FFmpeg (probe, crop, render, subtitle burn)
- Transcription: Faster-Whisper (CTranslate2 INT8 on CPU)
- Object Detection: YOLOv8n via ONNX Runtime
- URL Downloads: yt-dlp
- Optional LLMs: Groq (Llama-3.3 70B), Google Gemini 1.5 Flash
MIT License β do whatever you want with it. Free forever.
Built with β€οΈ as a free alternative to expensive video clipping tools.
If this helped you, consider giving it a β on GitHub!