Private, local-first software that turns long-form video into ranked, captioned vertical clips. It combines transcription, local LLM analysis, vision-assisted framing, review, rendering, and campaign planning behind one engine used by the CLI, desktop app, and web dashboard.
Local Shorts Studio is a self-hosted, offline-capable AI clipping tool for converting podcasts, interviews, webinars, wedding footage, and long YouTube videos into vertical Shorts, Reels, and TikTok clips on Apple Silicon.
- Creates non-destructive projects from videos already on your machine
- Inspects media and extracts audio with FFmpeg
- Transcribes locally with faster-whisper
- Chunks transcripts and asks Ollama or an explicitly enabled provider for clip candidates
- Validates, scores, and ranks candidates
- Detects and follows faces for vertical reframing
- Supports center crop, primary-face tracking, and two-person split layouts
- Produces ASS captions and optional call-to-action overlays
- Renders 9:16 clips without modifying the source
- Builds reusable campaign plans for batch execution
- Provides a Typer CLI, PySide6 desktop app, and Gradio web dashboard
- Includes an optional multi-provider chat workspace
Cloud providers are disabled by default. The normal transcription, analysis, framing, caption, and render workflow can run on-device.
The supported product target is Apple Silicon macOS with Python 3.12+, uv, FFmpeg/ffprobe, and Ollama.
git clone <your-repository-url>
cd local-shorts-studio
./scripts/bootstrap_macos.sh
source .venv/bin/activate
local-shorts doctorStart an interface:
local-shorts-dashboard # local web app on http://127.0.0.1:7861
local-shorts-gui # native PySide6 desktop app
local-shorts --help # automation-friendly CLIFor a manual developer install:
uv sync \
--extra dev \
--extra transcription \
--extra intelligence \
--extra vision \
--extra gui \
--extra dashboardImport video
↓
Inspect → extract audio → transcribe
↓
Chunk transcript → analyze → validate/rank candidates
↓
Review candidate → choose crop mode → create captions
↓
Render vertical clip → inspect output
All interfaces call the same shorts_engine services, so a project created in the CLI is visible in the desktop app and dashboard.
Rendering needs an FFmpeg build with the ass and drawtext filters. Check before processing a large project:
ffmpeg -filters | grep -E '(^| )ass|drawtext'Face-aware framing requires the vision extra. Local transcription requires the transcription extra and downloads the selected Whisper model on first use.
shorts_engine/
├── domain/ entities, schemas, validation
├── services/ pipeline, analysis, ranking, campaigns
├── media/ FFmpeg/ffprobe, captions, rendering
├── transcription/ faster-whisper integration
└── vision/ face detection and tracking
cli/ Typer command interface
app/ PySide6 desktop application
dashboard/ unified local web dashboard
chat_dashboard/ optional model chat interface
tests/ unit, integration, smoke, and GUI tests
docs/milestones/ implemented feature plans and decisions
See docs/current_status.md for implementation notes and AGENTS.md for the engineering contract.
uv run ruff check .
uv run ruff format --check .
uv run mypy .
uv run pytestVerified in the current workspace:
- Ruff lint: passing
- Ruff formatting: passing
- mypy: passing across 199 source files
- pytest: 474 passing, 1 skipped
- Optional-environment failures remain when OpenCV is absent or FFmpeg lacks
ass/drawtext
- Source media is never overwritten.
- User workspaces under
projects/are ignored by Git. - Subprocesses use argument arrays rather than shell strings.
- Cloud providers require an explicit privacy setting and credentials.
- Review generated clip boundaries and captions before publishing.
Yes. It transcribes the source, divides it into analyzable chunks, generates candidate moments, validates their boundaries, removes overlaps, and ranks the remaining clips.
The core workflow can run locally with FFmpeg, faster-whisper, Ollama, and OpenCV. Model files must be downloaded before a completely offline session.
Yes. It supports center framing, primary-face tracking, two-person split crops, ASS captions, and text overlays for 9:16 output.
No. The same engine powers a Typer CLI, a native PySide6 desktop application, and a unified local Gradio dashboard.
Any supported local video can become a project, including podcasts, interviews, lectures, webinars, event recordings, wedding footage, and creator videos.
Recommended GitHub description, topics, target search intents, and social-preview settings are maintained in GITHUB_METADATA.md.
The package is currently marked proprietary. Choose a public license before accepting contributions or redistributing it.
