Grounded SAM: Marrying Grounding DINO with Segment Anything & Stable Diffusion & Recognize Anything - Automatically Detect , Segment and Generate Anything
-
Updated
Sep 5, 2024 - Jupyter Notebook
Grounded SAM: Marrying Grounding DINO with Segment Anything & Stable Diffusion & Recognize Anything - Automatically Detect , Segment and Generate Anything
PyTorch tutorials and fun projects including neural talk, neural style, poem writing, anime generation (《深度学习框架PyTorch:入门与实战》)
Automagically synchronize subtitles with video.
SD.Next: All-in-one WebUI for AI generative image and video creation, captioning and processing
Automatic subtitle synchronization tool.
A simple yet powerful JQuery star rating plugin with fractional rating support.
新一代 AI 专业字幕软件,剪映字幕、elevenlabs 语音转文本的最佳本地版平替之一,也是加强版。中英转录识别准确率超过 97%,词语音频对齐率 98%,说话人分割与识别准确率 96%,带有最先进的 ASR 开源模型,100+字幕动画(支持导出透明背景的字幕动画)。说话人识别、专业字幕编辑器、命令行工具、Skill,达芬奇字幕插件(含字幕动画插件),PR 字幕插件,本地转录、远程转录、文稿匹配、智能拆行、AI校正、AI 智能热词、翻译、双语字幕、专业字幕编辑器、字幕合成、自定义大模型 API
OSINT Tool gets a range of information from an Instagram account 🛠
An official implementation for " UniVL: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation"
🍠 Access YouTube from R
Vega-Lite Chart Dataset and NL Generation Framework using LLMs
Audio Captioning datasets for PyTorch.
A Snapchat-inspired caption integrated within a regular UIImageView.
Browser extension for creating & uploading community captions for YouTube, niconico and other video sharing sites.
[ICLR'26] Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMs
Scalable annotation pipeline for action-aglined fine-grained instruciton for Visual-language-Action model
To associate your repository with the caption topic, visit your repo's landing page and select "manage topics."