Turn a 60-minute workshop into a week of viral content β automatically.
Demo Link: https://drive.google.com/file/d/1ZBfFm8Yr2ZhZA5N2Tjnm4bRpFJSWz3bY/view?usp=sharing
Mentors, educators, and creators produce hours of high-value long-form video. But modern audiences consume content in 60-second bursts. The most profound insights β a framework that could change someone's career, a story that reframes everything β are buried inside 60-minute recordings that most people never finish watching.
The Wisdom Gap: Valuable knowledge exists. Audiences exist. The bridge doesn't.
AttentionX closes that gap. Upload one session β get a week's worth of viral, vertical, caption-ready clips.
βΆ Watch the full demo on Google Drive
Direct Link:
https://drive.google.com/file/d/1ZBfFm8Yr2ZhZA5N2Tjnm4bRpFJSWz3bY/view?usp=sharing
The demo walks through:
- Uploading a 60-minute mentorship session
- Live emotional peak detection on the waveform timeline
- Auto-generated virality scores and hook headlines
- Smart 9:16 crop with face tracking
- Karaoke-style caption export
| Feature | What it does |
|---|---|
| Emotional Peak Detection | Fuses Librosa audio energy + Gemini sentiment scoring to find the most impactful 60-second windows |
| Virality Score | Each detected clip gets a 0β100% score based on audio intensity, sentiment profundity, and content type |
| Smart Vertical Crop | MediaPipe face tracking keeps the speaker centered in a 9:16 frame β no manual cropping |
| Karaoke Captions | Word-level Whisper timestamps drive animated, high-contrast caption overlays |
| Hook Headline Generator | Gemini generates a scroll-stopping 5β8 word headline for each clip automatically |
| One-Click Export | Burned-in captions + title card + 9:16 vertical MP4 ready for TikTok, Reels, and Shorts |
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β React Frontend (Vite) β
β Upload Zone β Peak Timeline β Clip Preview β
ββββββββββββββββββββββββββ¬ββββββββββββββββββββββββββββββββββ
β REST API
ββββββββββββββββββββββββββΌββββββββββββββββββββββββββββββββββ
β FastAPI Backend (Python) β
β Job Queue β Status Updates β Storage β
ββββββββ¬βββββββββββββββββββββββββββββββββββββββ¬βββββββββββββ
β β
ββββββββΌββββββββββββ βββββββββββββΌβββββββββββββ
β AI Analysis β β Video Processing β
β Pipeline β β Engine β
β β β β
β β’ OpenAI Whisperβ β β’ MediaPipe face track β
β β’ Librosa RMS β β β’ MoviePy crop/clip β
β β’ Gemini Flash β β β’ Caption burn-in β
β β’ Virality Fuse β β β’ 9:16 export β
ββββββββββββββββββββ βββββββββββββββββββββββββββ
β
ββββββββββββββββββββββββββΌββββββββββββββββββββββββββββββββββ
β Supabase (Storage + Postgres) β
β Job tracking Β· Clip storage Β· User data β
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
AttentionX doesn't guess β it fuses three independent AI signals to find the golden nuggets:
Signal A β Audio Energy (Librosa) Extracts RMS energy per frame, normalizes, and smooths with a rolling average. Detects where the speaker is most passionate and energized.
Signal B β Sentiment & Profundity (Gemini 1.5 Flash) Sends the full transcript (Gemini's 1M token context window handles 60+ minute sessions in one call) and scores each segment for: personal vulnerability, counterintuitive claims, actionable frameworks, and quotable one-liners.
Signal C β Timestamps (Whisper) Word-level timestamps from OpenAI Whisper power karaoke-style captions β each word highlights exactly as it's spoken.
Fusion:
virality_score = (0.4 Γ audio_energy) + (0.6 Γ gemini_score)
Top 5 non-overlapping windows (min 90s gap) become your clips.
16:9 source (1920Γ1080) β 9:16 output (608Γ1080)
- MediaPipe detects the speaker's face center X coordinate per frame (every 3rd frame for speed)
- A 30-frame rolling average smooths the crop window to eliminate jitter
- The crop window is clamped to frame bounds so it never goes out of range
- MoviePy applies the per-frame crop function at export time
The result: the speaker stays centered in frame throughout the entire clip, even if they move.
- Python 3.10+
- Node.js 18+
- ffmpeg installed (
brew install ffmpegorapt install ffmpeg) - API keys: Google Gemini, OpenAI, Supabase
# Clone the repository
git clone https://github.com/Piyusha942007/AttentionX-AI.git
cd attentionx
# Backend setup
cd backend
pip install -r requirements.txt
cp .env.example .env
# Fill in your API keys in .env
# Frontend setup
cd ../frontend
npm install
cp .env.example .env.local
# Fill in your Supabase URL and anon keyBackend .env:
GEMINI_API_KEY=your_gemini_api_key
OPENAI_API_KEY=your_openai_api_key
SUPABASE_URL=your_supabase_project_url
SUPABASE_SERVICE_KEY=your_supabase_service_keyFrontend .env.local:
VITE_SUPABASE_URL=your_supabase_project_url
VITE_SUPABASE_ANON_KEY=your_supabase_anon_key
VITE_API_URL=http://localhost:8000# Terminal 1 β Backend
cd backend
uvicorn main:app --reload --port 8000
# Terminal 2 β Frontend
cd frontend
npm run devattentionx/
βββ frontend/
β βββ src/
β β βββ pages/
β β β βββ Dashboard.jsx # Main 3-panel layout
β β βββ components/
β β β βββ UploadZone.jsx # Drag-and-drop file input
β β β βββ PeakTimeline.jsx # Waveform + peak markers
β β β βββ NuggetCard.jsx # Individual clip card
β β β βββ PreviewPane.jsx # Side-by-side 16:9 / 9:16
β β βββ App.jsx
β βββ package.json
β
βββ backend/
β βββ main.py # FastAPI app + routes
β βββ models.py # Pydantic schemas
β βββ db.py # Supabase client
β βββ worker.py # Background job processor
β βββ services/
β βββ transcriber.py # OpenAI Whisper wrapper
β βββ analyzer.py # Gemini sentiment analysis
β βββ scorer.py # Librosa + fusion scoring
β βββ face_tracker.py # MediaPipe face detection
β βββ cropper.py # MoviePy vertical crop
β βββ captioner.py # Karaoke caption burn-in
β
βββ README.md
Frontend
- React 18 + Vite
- Tailwind CSS
- Recharts (waveform visualization)
Backend
- FastAPI (Python)
- asyncio background workers
AI / ML
- Google Gemini 1.5 Flash β sentiment analysis & hook generation
- OpenAI Whisper β transcription with word-level timestamps
- Librosa β audio energy / RMS extraction
- MediaPipe β real-time face detection & tracking
Video Processing
- MoviePy β clip cutting, cropping, export
- ffmpeg β encoding and caption burn-in
Infrastructure
- Supabase β Postgres job tracking + object storage
- Vercel β frontend hosting
- Railway / Render β backend hosting
| Criterion | How AttentionX delivers |
|---|---|
| Impact (20%) | Turns 1 hour of content into 5 ready-to-publish clips in under 5 minutes |
| Innovation (20%) | 3-signal virality fusion (audio + semantic AI + timing) is novel; no existing tool does this |
| Technical Execution (20%) | Clean modular Python services, typed FastAPI endpoints, React component architecture |
| User Experience (25%) | Premium dark dashboard, real-time waveform, one-click export, demo mode for instant wow |
| Presentation (15%) | Full demo video linked above showing end-to-end flow on a real 60-min session |
The demo video is hosted on Google Drive:
βΆ Watch Demo β Google Drive
Link: https://drive.google.com/file/d/1ZBfFm8Yr2ZhZA5N2Tjnm4bRpFJSWz3bY/view?usp=sharing
Replace the link above with your actual Google Drive share link before submission.
Built for the AttentionX AI Hackathon by UnsaidTalks Education
| Role | Responsibility |
|---|---|
| Full Stack | React dashboard, FastAPI backend, Supabase integration |
| AI/ML | Gemini pipeline, Whisper transcription, virality scoring |
| Media Engineering | MediaPipe tracking, MoviePy crop, caption burn-in |
MIT License β see LICENSE for details.
Built with β‘ for the AttentionX AI Hackathon Β· UnsaidTalks Education Β· 2026