Images, video & audio
Generate, convert, and analyze image, video, and audio content. 990 voci corrispondenti nel registro Forge — sono mostrate le prime 150.
- modelscope/FunASRIndustrial-grade speech recognition toolkit: 170x realtime, 50+ languages, speaker diarization, emotion detect
- video-shotcraftCreate cinematic product videos from shot recipe cards, a validated template, and code/audio assets (Remotion
- SamurAIGPT/Generative-Media-SkillsMulti-modal Generative Media Skills for AI Agents (Claude Code, Cursor, Gemini CLI). High-quality image, video
- MiniMax-AI/MiniMax-MCPOfficial MiniMax Model Context Protocol (MCP) server that enables interaction with powerful Text to Speech, im
- jau123/MeiGen-AI-Design-MCPSupports GPT Image 2, Nanobanana & ComfyUI, with a 1,400+ prompt library, carefully crafted hooks and a multi-
- waybarrios/vllm-mlxOpenAI and Anthropic compatible server for Apple Silicon. Run LLMs and vision-language models (Llama, Qwen-VL,
- gyoridavid/short-video-makerCreates short videos for TikTok, Instagram Reels, and YouTube Shorts using the Model Context Protocol (MCP) an
- yzfly/douyin-mcp-server提取抖音无水印视频链接,视频文案,douyin-mcp-server,mcp,claude skill,支持龙虾
- SamurAIGPT/muapi-cliOfficial CLI for muapi.ai — generate images, videos & audio from the terminal. MCP server, 14 AI models, npm +
- jordanrendric/claude-video-visionGive Claude the ability to watch and understand videos — Claude Code plugin with frame extraction and multimod
- bitwize-music-studio/claude-ai-music-skillsHuman + AI music production workflow for Suno - skills, templates, and tools
- fpv-immersive-video-promptingUse when the user wants to turn a static scene, character images, aerial map, drawn path, or route-control ima
- watch-videoSLASH-COMMAND-ONLY. Invoke ONLY when the user explicitly types the literal `/watch-video` slash command. Never
- video-highlight-skillLong-video agent workflow for analyzing videos, building timestamped content indexes, selecting highlights, cr
- vox-style-collage-videoTurn a script or fact list into a finished Vox-style paper-collage explainer video: flat torn-edge photographi
- image-audit对图片进行鉴黄、政治、暴恐内容审核。先将图片压缩到 500px/JPEG 后直传 NX API 审核,以表格汇总结果。适用于用户提到图片审核、内容检查、鉴黄、政治识别、暴恐识别、违规扫描、图片安全、JPG/PNG/Web
- remotion-clone-videoPixel-perfect clone of any video (promo, product demo, motion-graphics piece, UI walkthrough) as a Remotion pr
- image-compress对图片进行智能压缩优化。支持本地路径、文件夹和远程 URL,直传 NX API 压缩后返回 CDN 地址和压缩率。适用于用户提到图片压缩、图片优化、减小图片体积、TinyPNG、JPG/PNG/WebP 压缩的场景。
- full-page-screenshot在浏览器中对网页做完整长截图(full-page screenshot),生成一张包含页面所有内容的 PNG 长图。 触发场景:用户要求对网页截长图、全页截图、完整截图、long screenshot、full page
- seedance-video-script把商业意图 + 素材转成 Seedance 2.5 可直接使用的视频分镜提示词。覆盖产品广告 / TVC 复刻 / 短剧 / 电商 / 数字人 / 科普等商业场景。内嵌 2.5 知识底座(时间戳分段、@引用语法、白模参考
- fb-ad-video-studioProduce high-converting Facebook/Instagram/TikTok ad VIDEOS — motion-graphics spots and talking-head founder a
- video-notesUse when the user gives a YouTube link or a local video file and wants an automatic transcript plus key-points
- xiaohongshu-imagesTransform markdown/HTML into styled 3:4 ratio images for Xiaohongshu
- smart-video-editor把多段原始素材剪成一条成片——先用视觉模型看懂每段画面拍了什么、哪几秒可用,再决定取舍、顺序、时长、调色和 BGM,最后用 ffmpeg 渲染。适用于"我给你几段视频,帮我剪一条小红书/朋友圈/vlog"这类需求。当用户
- ai-video-sannong端到端制作竖屏 AI 妙招短视频(默认 3 镜 × 5 秒 = 15 秒,9:16)。题材通用:三农园艺、阳台种菜、美食、家居清洁收纳、宠物、育儿等都行,开场先问用户选题方向,不预设。工作流:方向确认 → 热门检索/对标
- bilibili-audio下载B站视频音频。支持单个视频下载和批量下载UP主所有视频。当用户提供B站视频链接(bilibili.com、b23.tv)、要求下载B站音频、下载B站视频、提取B站音频、或批量下载某个UP主的视频时使用。触发词:B站下
- talking-head-videoTurn a raw 口播 / talking-head selfie video plus its script into a finished explainer, vertical 9:16 or landscap
- screenshot-to-htmlReplicate a UI screenshot/image/mockup into a high-fidelity, self-contained HTML page. Faithful image-to-HTML
- NoLegitHere/capture-windows-application-mcp-deletion_scheduled-81732965MCP server that launches Windows executables, captures window screenshots, and returns inline images to any LL
- Ashot72/multi-agent-a2aExploring Google’s A2A AI System: Agent-to-Agent Workflows, Routing, and Conversation History
- ai.pictomancer/image-processingImage processing for AI agents. Resize, convert, compress, and pipeline images.
- ai.router/ai-gatewayOne key, 100+ models — chat with any LLM and generate video, images, speech. Free trial at 370.ai.
- ai.smithery/Artin0123-gemini-image-mcp-serverAnalyze images and videos with Gemini to get fast, reliable visual insights. Handle content from U…
- ai.smithery/aryankeluskar-poke-video-mcpSearch your Flashback video library with natural language to instantly find relevant moments. Get…
- ai.smithery/ctaylor86-mcp-video-download-serverConnect your video workflows to cloud storage. Organize and access video assets across projects wi…
- com.audioscrape/audio-intelligenceThe audio intelligence layer. Search podcast transcripts, speakers, and entities across 250K+ shows.
- com.snapog/snapogGive agents instant OG image generation, social metadata audits, and rendering guidance.
- io.github.Br0ski777/image-generatorAI image generation from text prompts via Gemini. x402 micropayment.
- io.github.Br0ski777/image-resizeResize images from URL — PNG, JPEG, WebP output. x402 micropayment.
- io.github.Br0ski777/screenshot-pdfCapture screenshots (PNG/JPEG/WebP) and generate PDFs from any URL. x402 payments.
- io.github.BrightWayAI/video-analyzerAnalyze videos: extract frames, transcribe audio, generate storyboard breakdowns.
- io.github.CSOAI-ORG/image-metadata-ai-mcpimage-metadata-ai-mcp MCP server by MEOK AI Labs
- io.github.CSOAI-ORG/video-editing-ai-mcpVideo Editing Ai MCP server. Tools: split scenes, generate subtitles, thumbnail data. Built by MEOK
- io.github.CSOAI-ORG/voice-audio-mcpVoice Audio MCP server. Tools: text to speech, list voices, transcribe. Built by MEOK AI Labs.
- io.github.Createya-ai/createya-mcp100+ AI models: FLUX, Sora, Veo, Kling, Runway, Suno. OAuth or Bearer. No VPN. RU billing.
- io.github.Digital-Defiance/mcp-screenshotScreenshot capture with PII masking and cross-platform support for AI agents
- io.github.JimothySnicket/gemini-imageGoogle Gemini image generation, editing, and local processing via MCP
- io.github.JuzzyDee/audio-analyzerGives LLMs ears. Spectral, harmonic, rhythm, stereo, and structural audio analysis.
- io.github.KamaruSama/mcp-screenshotTake Wayland screenshots via grim/slurp. Full screen, region, or interactive crop.
- io.github.KyaniteLabs/mcp-videoGuardrailed FFmpeg video MCP server with Hyperframes and repurposing tools.
- io.github.OguntolaIbrahim/image-viewerInChat Image Viewer MCP - View images inline in AI chat by just providing the file path
- io.github.PedroMarianoAlmeida/ffmpeg-mcpFFmpeg video/audio tools: cut, convert, concat, remove silence, and raw commands.
- io.github.Wolflangis/video-loomTurn audio or text into a finished AI music video: narrative, multi-provider clips, final cut.
- io.github.ahmaddioxide/image-resolverMCP server for context-aware royalty-free image search from Pexels and Unsplash
- io.github.bookyo/nsfw-image-detectorNSFWJS-based image safety detector (GIF/APNG/WebP/JPEG/PNG).
- io.github.botmonster/image2svg-mcpMCP server that converts raster images (PNG/JPG/WEBP) to SVG vector format
- io.github.burningion/video-editing-mcpMCP Server for Video Jungle - Analyze, Search, Generate, and Edit Videos
- io.github.chaehomie/aether-studioGenerate, edit, and explore AI images. Flux, Imagen, LoRA identity swap, upscale, and more.
- io.github.copperline-labs/rendex-mcpRender HTML, Markdown, or any URL to images or PDF, plus reader-mode extraction. MCP-native.
- io.github.counterpoint-studio/audio-file-mcp-appInspect local audio files — playback, metadata, loudness, spectrogram.
- io.github.davidmosiah/short-video-agent-kitProvider-neutral short-form AI video toolkit for agents: Sora, Gemini Veo, xAI/Grok, Seedance.
- io.github.fasuizu-br/image-toolsBackground removal, 4x upscaling, and face restoration via GPU
- io.github.hifarrer/ffmpegapiHosted MCP tools for FFmpeg-style video and audio processing through FFMPEG API.
- io.github.htmlcsstoimage/html-css-to-imageAn MCP server for generating images from HTML & CSS or screenshots of URLs using htmlcsstoimage.com.
- io.github.jxoesneon/gemini-audio-mcpHigh-performance audio, music, and voice generation MCP server for Gemini 2.5 and Lyria 3.
- io.github.kaitoy/nasa-imagesMCP server with UI for NASA images
- io.github.ludmila-omlopes/youtube-video-analyzerMCP stdio server for analyzing YouTube videos with Google Gemini
- io.github.pachote/nira-video-mcpVideo production MCP: 55 tools — edit, grade, TTS captions, beat-sync, master, AI generate.
- io.github.pijusz/mcp-sanity-imagesMCP server for uploading local images to Sanity CMS
- io.github.pvliesdonk/image-generation-mcpMCP server for AI image generation via OpenAI, Stable Diffusion (SD WebUI), or placeholders.
- io.github.remoas/appstore-screenshotsGenerate professional App Store screenshots by matching any top app's style.
- io.github.rocnubie/aiimageeditor-mcpBest Image AI Editor free online. Free AI photo editor for picture & image editing.
- io.github.rocnubie/bestaiimageprompt-mcpThe best AI image prompt library online. Latest trending prompts copy & paste for GPT image 2,
- io.github.rocnubie/videotoaudioconverter-mcpVideo to Audio Converter is a free, browser-based tool that extracts audio tracks from common video
- io.github.rocnubie/z-image-mcpZ-Image is an online interface for the Z-Image generation model with fast inference and prompt
- io.github.rog0x/imageImage metadata, favicon, OG image for AI agents
- io.github.sethbang/screenshot-serverMCP server for web page and cross-platform system screenshots (macOS, Linux, Windows)
- io.github.shalevshalit/image-recognition-mcpMCP server for AI-powered image recognition and description using OpenAI vision models.
- io.github.shalevshalit/image-recongnition-mcpMCP server for AI-powered image recognition and description using OpenAI vision models.
- io.github.shinpr/mcp-imageAI image generation and editing with prompt optimization and quality presets
- io.github.sidart10/runway-mcp-serverAI video generation with Gen-4, Veo 3, and Aleph editing. Text-to-video, image-to-video, 4K upscale
- io.github.therealMrFunGuy/screenshot-apiWebsite screenshots as PNG, JPEG, or PDF with custom viewports and full-page capture.
- io.github.u2n4/video-url-analyzer-mcpAnalyze YouTube/TikTok/Instagram videos: transcripts, AI insights, tutorial extraction
- io.github.udhaykumarbala/gemini-image-studioAI image generation and editing with Google Gemini. Structured JSON editing.
- io.github.usnmax88/reel25-mcpVideo analytics for TikTok, Instagram, and YouTube. Track, analyze, and discover content.
- io.github.ykshah1309/live-audio-intelligence-mcpLive webcast transcription + vocal stress analysis (F0 jitter, hesitation) for earnings calls.
- io.studiomeyer/videoVideo production: recording, editing, effects, captions, TTS, screenshots. 8 tools.
- io.veed/fabric-mcpGenerate AI talking-head videos with custom characters and voices.
- space.studiosphere/pulsePrivacy-first audio intelligence: BPM, key, waveform. Audio never stored. Pay per second.
- video.future/future-video-studioCreate and manage cinematic AI video renders through the Future Video Studio Agent API.
- video.open/open-videoAI-powered video publishing, channel management, and monetization via open.video
- seo-imagesImage optimization analysis for SEO and performance. Checks alt text, file sizes, formats, responsive images,
- system-audioCapture system audio (what the computer is playing) using crystal-audio. Covers ProcessTap (macOS 14.2+), Scre
- muapi-animal-video-generatorCreate a hilarious and ultra-realistic video of an anthropomorphic animal acting like a human vlogger in a rea
- muapi-award-ceremony-videoGenerate a 15-second cinematic awards-ceremony video — a host announces a winner from the stage, a spotlight f
- video-studioBuild finished MP4 videos whose visuals are synced to audio — from a ready-made song, from a pre-recorded voic
- blender-vfx-audioAudio with Blender VFX. audio editing.
- audio-designImplement game audio practice — bus/mixer architecture and gain in decibels, ducking (sidechain), adaptive/dyn
- godot-audioPlay and mix audio in Godot 4.7: AudioStreamPlayer (2D/3D variants), audio buses with volume/mute and effects,
- venice-audio-musicAsync music / audio-track generation via Venice. Covers the /audio/quote + /audio/queue + /audio/retrieve + /a
- venice-audio-speechGenerate speech from text via POST /audio/speech, and clone a voice via POST /audio/voices. Covers TTS models
- venice-audio-transcriptionTranscribe audio files to text via POST /audio/transcriptions. Covers supported models (Parakeet, Whisper, Wiz
- venice-image-editTransform existing images with Venice. Covers POST /image/edit (prompt-driven single-image edit), /image/multi
- venice-image-generateGenerate images with Venice. Covers POST /image/generate (Venice-native), POST /images/generations (OpenAI-com
- venice-videoGenerate and transcribe videos via Venice. Covers the async /video/quote + /video/queue + /video/retrieve + /v
- ai-video-generatorGenerate short-form videos with AI — script writing, text-to-speech narration, stock footage selection, subtit
- codex-gpt-imageGenerate or edit images with gpt-image-2 through Codex/ChatGPT subscription authentication instead of OPENAI_A
- avenox-videoAvenox Studio — local-first YouTube video production pipeline (ROUTER, read first). Use for ANY request to edi
- aso-appstore-screenshotsA curated guide to convention files AI agents read, write, and act on: AGENTS.md, CLAUDE.md, SKILL.md, llms.tx
- nano-image-generatorA curated guide to convention files AI agents read, write, and act on: AGENTS.md, CLAUDE.md, SKILL.md, llms.tx
- audio-producer-agentUse this skill to create single-voice audio content like audiobooks, voiceovers, narrations, jingles, and audi
- image-generationGenerate and optimize images for social media posts across all platforms
- video-generationGenerate videos through an available or configured video backend using structured prompts, explicit output pat
- video-producer-agentUse this skill to create complete videos with voiceover and music. Triggers: "create video", "product video",
- design-image-to-codeTurn a reference image or mockup into a working local frontend that matches it closely. Inspects the image, ex
- video-to-skillWatch a tutorial or demo video and generate a Claude Code skill from it. Activated when user says "create a sk
- audio-transcribeThis skill should be used when the user explicitly asks to "transcribe a meeting", "transcribe audio", "transc
- remotion-video-creationBest practices for Remotion - Video creation in React. 29 domain-specific rules covering 3D, animations, audio
- video-content-extractorExtract key frames and text content from video files using ffmpeg and Tesseract OCR.
- automotive-audioExpert skill in active noise cancellation focusing on audio domain applications. Covers 40 topics across audio
- browser-screenshot-diffVisual + DOM diff between two recorded sessions at matching trajectory step ids; used for visual regression an
- video-downloadDownload videos from Instagram, YouTube, TikTok, Twitter/X, Facebook, and 1000+ other platforms as MP4 files u
- video-subtitle-extractorExtract and save subtitles from video URLs as plain .txt files. Supports Bilibili (B站) and YouTube. Use when t
- surgepix-image-translateTranslate text on images to a target language using SurgePix API, keeping the background unchanged. Use when t
- gpt-image-2面向 GPT Image 2 的图像生成 / 编辑技能。可在 3 种环境下使用:(A) Garden 本地模式,通过 OpenAI 兼容接口直接出图并落盘;(B) Host-Native 模式,把本 Skill 当作提示
- grok-imageGenerate images with xAI's Grok image models (grok-2-image / grok-2-image-1212). Use when the user asks for Gr
- grok-imagine-videoGenerate, edit, and extend videos with xAI's grok-imagine-video model. Use when the user wants to create a vid
- agentverse-image-genGenerate images using AI agents on Fetch.ai's Agentverse. Sends a text prompt to an image generation agent and
- generate-imageGenerate, edit, or batch-create SEO, marketing, social, and document images with Gemini or OpenAI through the
- imageCreate or optimize marketing images, social graphics, product mockups, banners, cover art, listing visuals, br
- image-to-codeBuild or reconstruct a visually important website through an image-first workflow: generate or inspect section
- manim-videoManim CE animations: 3Blue1Brown math/algo videos.
- videoPlan and produce video with available AI tools or programmatic frameworks. Use for video prompts, avatars, exp
- video-editingAI-assisted video editing workflows for cutting, structuring, and augmenting real footage. Covers the full pip
- browseract-google-image-api-skillThis skill helps users automatically extract structured image data from Google Images via BrowserAct API. Agen
- browseract-tiktok-hashtag-videosTikTok hashtag video scraper: input a hashtag name → output paginated video list with full metadata (author pr
- browseract-tiktok-profile-videosTikTok user profile video scraper: input a TikTok username → output the user's profile info plus paginated vid
- browseract-tiktok-search-videosTikTok keyword search video scraper: input search keyword → output paginated video list with full metadata (au
- browseract-tiktok-video-detailTikTok single video detail scraper: input a TikTok video URL → output full video metadata (author profile, eng
- browseract-youtube-video-api-skillThis skill helps users automatically extract channel-level and video detail data from a specific YouTube chann
- screenshot-docsCapture static screenshots or short GIF demonstrations of a GUI, CLI, or web app and embed them in README or d
- prizmad-video-adsGenerate AI UGC video ads from any public product URL. Use when the user wants to produce, iterate on, or A/B
- mindstudio-generate-image-block-skillWrite, configure, and optimize prompts for MindStudio's Generate Image block. Use this skill whenever someone
- poster-image-to-htmlUse when a user provides a reference image (poster, banner, promo graphic) and wants faithful HTML recreation,
- video-ads-structureVideo Ads Structure — Skill especializada para video ads structure
- architecture-imageCreate crisp architecture, integration, workflow, and system diagrams as editable SVG plus rendered PNG assets
- image-to-editable-ppt当用户提供一张或多张幻灯片图片、图片版 PPT/PPTX 或 PDF,并要求转成可编辑 PowerPoint/PPTX、重建幻灯片对象、保留页面备注或做可编辑化复刻时使用。
- docker-imagesBuild small, reproducible images: layer ordering for cache hits, multi-stage builds, pinning by digest, and wh
- ai-blending-reference-imagesHow to combine two images or ideas into one, using multi-image reference input, the modern replacement for pho
- ai-image-controls-comparedThe same handful of dials (aspect ratio, style strength, seed, reference images) mapped across every major ima
Altri argomenti