Skill Detail

Turn long video or audio into reviewed clips and subtitles with FunClip

Use FunClip when an agent needs a repeatable local workflow for transcribing media, proposing clip timestamps, and producing reviewable video segments with subtitle files.

Media & TranscriptionMulti-Framework
Media & Transcription Multi-Framework Security Reviewed
⭐ 5.9k GitHub stars
COPY SKILL INSTRUCTIONS (OPTIONAL)
npx skills add agentskillexchange/skills --skill turn-long-video-or-audio-into-reviewed-clips-and-subtitles-with-funclip Copy
Uses the third-party skills CLI, not an ASE-owned installer. Check your agent’s compatibility. This copies instructions; complete the upstream tool setup below separately.
At a glance
Tools required
Python environment, FunClip source checkout, FunASR-supported models, local or server Gradio runtime, and media files to transcribe and clip.
Install & setup
Clone the upstream FunClip repository, install its Python dependencies as documented, launch the local Gradio app with the documented python funclip/launch.py command, then load media files for transcription, subtitle generation, LLM-assisted timestamp selection, and reviewed clipping.
Author
ModelScope / FunASR
Publisher
Open Source
Last updated
Jul 2, 2026
Quick brief

FunClip is a local, open-source workflow for turning long video or audio into transcripts, subtitles, and selected clips. The operator gives the agent a media file and a clipping goal, then the agent runs FunClip’s FunASR-backed transcription flow, reviews or edits the generated transcript and SRT output, uses the LLM-assisted clipping path to identify candidate timestamps, and exports only the approved segments.

How it works

What this skill actually does

Invoke this skill when the task is media production or review: podcast highlights, meeting excerpts, interview clips, lecture segments, translated subtitle prep, or any workflow where the agent should propose clips but a human still needs to inspect the transcript and final cut. Use the product normally when a person just wants to interact with the Gradio UI manually.

The boundary is media-intake to reviewable clip output. This is not a generic video editor listing or a transcription SDK card; the agent workflow is to ingest a local media asset, generate transcript and subtitle artifacts, select candidate ranges from the text, and hand back reproducible clip decisions plus files for review.